feat(lcm): honor compression.target_ratio via a token-budgeted fresh tail - #168
Conversation
|
| Filename | Overview |
|---|---|
| plugins/context_engine/lcm/config.py | Adds _hermes_lcm_value/_hermes_lcm_int/_lcm_config_bool helpers and four new LCMConfig fields; env override range guard for target_ratio correctly clamps out-of-range values; consistent with existing fail-safe-ON doctrine. |
| plugins/context_engine/lcm/engine.py | Core logic: adds _refresh_fresh_tail_token_budget, _dynamic_fresh_tail_count, _fresh_tail_start; frozen-K computed once at compress() entry; D-7 rotate floor and D-10 fail-open target_ratio fix are correct; all degenerate paths fall back to legacy count. |
| tests/context_engine/test_lcm_fresh_tail_token_budget.py | 45 tests covering all spec invariants; AST guard now includes AugAssign negative control; env range guard test class added; test isolation via _clean_env is thorough. |
| plugins/context_engine/lcm/tools.py | lcm_status surfaces five new budget-related fields via getattr with safe defaults; straightforward observability addition with no logic change. |
| scripts/lcm_tail_budget_shakedown.py | Offline A/B shakedown script; gates match spec AC-2/AC-9; uses a throwaway HERMES_HOME with proper cleanup. |
| tests/context_engine/test_lcm_failopen_observability.py | Pins fresh_tail_token_budget_enabled=False to preserve existing tiny-corpus compaction fixture; no behavior change to existing tests. |
| scripts/probe_hermes_lcm_isolated.py | Pins fresh_tail_token_budget_enabled=False for the toy corpus probe with a clear explanatory comment. |
Flowchart
%%{init: {'theme': 'neutral'}}%%
flowchart TD
A[update_model] --> B[_refresh_fresh_tail_token_budget]
B --> C{enabled AND threshold > 0?}
C -- No --> D[budget = 0, legacy tail]
C -- Yes --> E{explicit override?}
E -- Yes --> F[budget = explicit]
E -- No --> G[budget = target_ratio x threshold_tokens]
F --> H[caps = budget, 0.9x threshold, max_tokens]
G --> H
H --> K[_fresh_tail_token_budget = min caps]
K --> L[compress called]
L --> M[_dynamic_fresh_tail_count on working_messages]
M --> N[Walk reversed, accumulate tokens]
N --> O{over budget AND count >= base?}
O -- No --> N
O -- Yes --> P[count = max base count]
P --> Q[Frozen-K locked for pass]
Q --> R[compress completes]
R --> S[_last_fresh_tail_count = frozen_K]
S --> U[rotate: max static _last_fresh_tail_count]
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
A[update_model] --> B[_refresh_fresh_tail_token_budget]
B --> C{enabled AND threshold > 0?}
C -- No --> D[budget = 0, legacy tail]
C -- Yes --> E{explicit override?}
E -- Yes --> F[budget = explicit]
E -- No --> G[budget = target_ratio x threshold_tokens]
F --> H[caps = budget, 0.9x threshold, max_tokens]
G --> H
H --> K[_fresh_tail_token_budget = min caps]
K --> L[compress called]
L --> M[_dynamic_fresh_tail_count on working_messages]
M --> N[Walk reversed, accumulate tokens]
N --> O{over budget AND count >= base?}
O -- No --> N
O -- Yes --> P[count = max base count]
P --> Q[Frozen-K locked for pass]
Q --> R[compress completes]
R --> S[_last_fresh_tail_count = frozen_K]
S --> U[rotate: max static _last_fresh_tail_count]
Reviews (2): Last reviewed commit: "fix(lcm): Greptile review — env-path (0,..." | Re-trigger Greptile
…tail
LCM previously ignored compression.target_ratio entirely and kept a fixed
32-message fresh tail (~6.5K tokens on live traffic) regardless of context
size. This maps the fleet-standard knob to a token budget:
budget = target_ratio × threshold_tokens
(= target_ratio × threshold × context_length)
capped at lcm.fresh_tail_max_tokens (default 60K, guards 1M-window models)
and 0.9 × threshold_tokens (convergence clamp — compaction can never no-op
loop even at ratio=1.0).
At compress() entry the dynamic count K is computed ONCE (frozen-K): walk
backward accumulating count_message_tokens until the budget is spent,
floored at the legacy fresh_tail_count (the tail can only widen, never
shrink below today's behavior). All in-pass cut sites use len(list) - K —
sound because the leaf-loop span is removal-only (the only list-growing
transform, stub insertion via _sanitize_active_context_messages, runs
strictly after the last cut; an AST source-contract test with planted-
violation negative controls guards this invariant against future edits).
Out-of-pass readers (preflight, _raw_backlog_messages) compute fresh.
Also:
- protect_last_n now tracks the dynamic K after each compaction so the
compaction-stats consumer (find_inturn_kept_cut slack window) stays
centered; pre-first-compaction value unchanged.
- rotate_active_session preserves max(static, last-dynamic-K) rows so the
lifecycle frontier can never advance past the wider in-memory tail.
- fail-open fallback compressor now reads the configured target_ratio
(was a hardcoded 0.20).
- lcm_status exposes the resolved budget + last dynamic count for live
verification.
- disable switch: lcm.fresh_tail_token_budget_enabled: false (or env
LCM_FRESH_TAIL_TOKEN_BUDGET_ENABLED=0), fail-safe ON for garbage values.
Degenerate inputs (no context window, flag off, invalid ratio, estimator
exception) reproduce the legacy fixed-count cut exactly.
Ingest/replay/dedup paths untouched (PR #107 territory): zero edits to
_reconcile_ingest_cursor_from_store, the scaffold-overlap guard, or the
min-evidence bar at the replay-skip site.
Verification: 45 new tests (config plumbing, budget resolution, chokepoint
helper, frozen-K boundary consistency, AST guard + planted violations,
rotate floor incl. window-grow, fail-open ratio); full context_engine +
compaction-stats suite 348 passed; offline A/B shakedown on a mixed-density
corpus at 200K: ON-arm tail 98 msgs / 36,946 tokens vs budget 37,500,
legacy arm 32 msgs, 3 fires per arm, cadence_ratio 1.01 (gate ≤ 2.0).
Spec: ~/.hermes/plans/2026-07-01_lcm-target-ratio-token-budget-tail-SPEC.md
(v0.4, three Opus review passes: BLOCK → AWC → AWC, all blockers folded)
e4d7a4a to
cef9f62
Compare
… + AugAssign list-growth coverage in the frozen-K AST guard - LCM_TARGET_RATIO env override now applies the same (0,1] range guard as the config-file path (out-of-range → default 0.20), with tests. - The frozen-K AST guard also flags augmented-assignment list growth (working_messages += [...]) inside the leaf loop — scoped to message-list vars / list-literal RHS so integer counters don't false-positive — with a planted += negative control.
…25 rows (t_caabb1fa) 49/84 rows. LCM rows graded against stephenschoettler/hermes-lcm @ 8d1b1e6 (v1.0.0-rc.1) per D1a and reconciled with the lcm-upstream-divergence-ledger (DD-1 -> #168 SUPERSEDED, DD-2 -> #107 SUPERSEDED, DC-1/DC-2/DC-3 KEEP). Three generic fixes upstream still lacks by read: #412, #902, dc4245a -> UPSTREAM. Route disclosure corrected in plugins.md header (judgment on claude-fable-5-1).
What
LCM ignored
compression.target_ratioand kept a fixed 32-message fresh tail (~6.5K tokens measured on live traffic). This maps the fleet-standard knob to a token-budgeted tail:capped at
lcm.fresh_tail_max_tokens(60K default) and 0.9×threshold_tokens (convergence clamp). At current fleet config (0.25 × 0.75 × 200K) that's 37,500 tokens ≈ 229 messages of verbatim recency vs ~6.5K today — density-adaptive (chatty turns keep many messages, tool-heavy turns keep fewer).Design highlights
len(list) − K. Sound because the leaf-loop span is removal-only — guarded by an AST source-contract test with planted-violation negative controls.fresh_tail_count— the tail can only widen; every degenerate input (no window, flag off, bad ratio, estimator exception) reproduces the legacy cut exactly.lcm.fresh_tail_token_budget_enabled: false+ restart, no revert.Verification