Skip to content

feat(lcm): honor compression.target_ratio via a token-budgeted fresh tail - #168

Merged
Kyzcreig merged 2 commits into
mainfrom
feat/lcm-target-ratio-tail-budget
Jul 2, 2026
Merged

Kyzcreig merged 2 commits into
mainfrom
feat/lcm-target-ratio-tail-budget

Conversation

@Kyzcreig

@Kyzcreig Kyzcreig commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator

What

LCM ignored compression.target_ratio and kept a fixed 32-message fresh tail (~6.5K tokens measured on live traffic). This maps the fleet-standard knob to a token-budgeted tail:

budget = target_ratio × threshold × context_length

capped at lcm.fresh_tail_max_tokens (60K default) and 0.9×threshold_tokens (convergence clamp). At current fleet config (0.25 × 0.75 × 200K) that's 37,500 tokens ≈ 229 messages of verbatim recency vs ~6.5K today — density-adaptive (chatty turns keep many messages, tool-heavy turns keep fewer).

Design highlights

  • Frozen-K per compress() pass: dynamic count computed once at entry; all in-pass cuts use len(list) − K. Sound because the leaf-loop span is removal-only — guarded by an AST source-contract test with planted-violation negative controls.
  • Floor at legacy fresh_tail_count — the tail can only widen; every degenerate input (no window, flag off, bad ratio, estimator exception) reproduces the legacy cut exactly.
  • Zero edits to ingest/replay/dedup (PR fix(lcm): preserve real per-message timestamps + stop dup-on-replay #107 territory).
  • Rollback: lcm.fresh_tail_token_budget_enabled: false + restart, no revert.
  • Fail-open fallback compressor now reads configured target_ratio (was hardcoded 0.20).

Verification

  • 45 new tests; full context_engine + compaction-stats suite 348 passed
  • Offline A/B shakedown (mixed-density corpus, 200K window): ON tail 98 msgs / 36,946 tok vs budget 37,500; legacy arm 32 msgs; 3 fires/arm; cadence_ratio 1.01 (gate ≤2.0) — 7/7 gates PASS
  • Spec v0.4, 3 Opus review passes (BLOCK → AWC → AWC), all blockers folded with live ground-truth

@greptile-apps

greptile-apps Bot commented Jul 2, 2026 •

Copy link
Copy Markdown

Greptile Summary

This PR maps compression.target_ratio to a token-budgeted fresh tail in LCM, replacing the legacy hard-coded 32-message fixed count with a dynamically computed target_ratio × threshold_tokens budget (capped at fresh_tail_max_tokens and a 0.9× convergence clamp). The frozen-K discipline ensures the per-pass dynamic count is computed once and all in-loop cuts reference that constant, keeping count-arithmetic sound.

  • Config plumbing (config.py): Three new lcm.* knobs and compression.target_ratio sourcing, with env-override range guard; both previously-flagged gaps (AugAssign AST guard, env range validation) are addressed.
  • Engine changes (engine.py): _refresh_fresh_tail_token_budget computed on every _set_context_length; frozen-K locked once per compress() pass; D-7 rotate floor and D-10 fail-open target_ratio fix are correct; all degenerate paths fall back to legacy count.
  • Tests: 45 new tests covering config plumbing, budget resolution, chokepoint helper, frozen-K boundary consistency, rotate-floor invariants, and env range guard.

Confidence Score: 5/5

Safe to merge; all degenerate inputs fall back to the legacy fixed-count tail, the convergence clamp prevents the fresh tail from consuming more than 90% of the threshold budget, and the rollback path is a single config flag plus restart.

Both issues raised in the previous review round — the AugAssign gap in the AST frozen-K guard and the missing env-override range validation for target_ratio — are fully addressed with implementation fixes and dedicated test cases. The frozen-K discipline is sound: count is frozen once per compress() pass on the post-ingest list, the leaf loop is removal-only (AST-verified), and every failure path reproduces the legacy behavior exactly.

No files require special attention; engine.py carries the most logic but is backed by 45 targeted tests and an offline A/B shakedown.

Important Files Changed

Filename Overview
plugins/context_engine/lcm/config.py Adds _hermes_lcm_value/_hermes_lcm_int/_lcm_config_bool helpers and four new LCMConfig fields; env override range guard for target_ratio correctly clamps out-of-range values; consistent with existing fail-safe-ON doctrine.
plugins/context_engine/lcm/engine.py Core logic: adds _refresh_fresh_tail_token_budget, _dynamic_fresh_tail_count, _fresh_tail_start; frozen-K computed once at compress() entry; D-7 rotate floor and D-10 fail-open target_ratio fix are correct; all degenerate paths fall back to legacy count.
tests/context_engine/test_lcm_fresh_tail_token_budget.py 45 tests covering all spec invariants; AST guard now includes AugAssign negative control; env range guard test class added; test isolation via _clean_env is thorough.
plugins/context_engine/lcm/tools.py lcm_status surfaces five new budget-related fields via getattr with safe defaults; straightforward observability addition with no logic change.
scripts/lcm_tail_budget_shakedown.py Offline A/B shakedown script; gates match spec AC-2/AC-9; uses a throwaway HERMES_HOME with proper cleanup.
tests/context_engine/test_lcm_failopen_observability.py Pins fresh_tail_token_budget_enabled=False to preserve existing tiny-corpus compaction fixture; no behavior change to existing tests.
scripts/probe_hermes_lcm_isolated.py Pins fresh_tail_token_budget_enabled=False for the toy corpus probe with a clear explanatory comment.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A[update_model] --> B[_refresh_fresh_tail_token_budget]
    B --> C{enabled AND threshold > 0?}
    C -- No --> D[budget = 0, legacy tail]
    C -- Yes --> E{explicit override?}
    E -- Yes --> F[budget = explicit]
    E -- No --> G[budget = target_ratio x threshold_tokens]
    F --> H[caps = budget, 0.9x threshold, max_tokens]
    G --> H
    H --> K[_fresh_tail_token_budget = min caps]
    K --> L[compress called]
    L --> M[_dynamic_fresh_tail_count on working_messages]
    M --> N[Walk reversed, accumulate tokens]
    N --> O{over budget AND count >= base?}
    O -- No --> N
    O -- Yes --> P[count = max base count]
    P --> Q[Frozen-K locked for pass]
    Q --> R[compress completes]
    R --> S[_last_fresh_tail_count = frozen_K]
    S --> U[rotate: max static _last_fresh_tail_count]
Loading
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
    A[update_model] --> B[_refresh_fresh_tail_token_budget]
    B --> C{enabled AND threshold > 0?}
    C -- No --> D[budget = 0, legacy tail]
    C -- Yes --> E{explicit override?}
    E -- Yes --> F[budget = explicit]
    E -- No --> G[budget = target_ratio x threshold_tokens]
    F --> H[caps = budget, 0.9x threshold, max_tokens]
    G --> H
    H --> K[_fresh_tail_token_budget = min caps]
    K --> L[compress called]
    L --> M[_dynamic_fresh_tail_count on working_messages]
    M --> N[Walk reversed, accumulate tokens]
    N --> O{over budget AND count >= base?}
    O -- No --> N
    O -- Yes --> P[count = max base count]
    P --> Q[Frozen-K locked for pass]
    Q --> R[compress completes]
    R --> S[_last_fresh_tail_count = frozen_K]
    S --> U[rotate: max static _last_fresh_tail_count]
Loading

Reviews (2): Last reviewed commit: "fix(lcm): Greptile review — env-path (0,..." | Re-trigger Greptile

…tail

LCM previously ignored compression.target_ratio entirely and kept a fixed
32-message fresh tail (~6.5K tokens on live traffic) regardless of context
size. This maps the fleet-standard knob to a token budget:

  budget = target_ratio × threshold_tokens
         (= target_ratio × threshold × context_length)

capped at lcm.fresh_tail_max_tokens (default 60K, guards 1M-window models)
and 0.9 × threshold_tokens (convergence clamp — compaction can never no-op
loop even at ratio=1.0).

At compress() entry the dynamic count K is computed ONCE (frozen-K): walk
backward accumulating count_message_tokens until the budget is spent,
floored at the legacy fresh_tail_count (the tail can only widen, never
shrink below today's behavior). All in-pass cut sites use len(list) - K —
sound because the leaf-loop span is removal-only (the only list-growing
transform, stub insertion via _sanitize_active_context_messages, runs
strictly after the last cut; an AST source-contract test with planted-
violation negative controls guards this invariant against future edits).
Out-of-pass readers (preflight, _raw_backlog_messages) compute fresh.

Also:
- protect_last_n now tracks the dynamic K after each compaction so the
  compaction-stats consumer (find_inturn_kept_cut slack window) stays
  centered; pre-first-compaction value unchanged.
- rotate_active_session preserves max(static, last-dynamic-K) rows so the
  lifecycle frontier can never advance past the wider in-memory tail.
- fail-open fallback compressor now reads the configured target_ratio
  (was a hardcoded 0.20).
- lcm_status exposes the resolved budget + last dynamic count for live
  verification.
- disable switch: lcm.fresh_tail_token_budget_enabled: false (or env
  LCM_FRESH_TAIL_TOKEN_BUDGET_ENABLED=0), fail-safe ON for garbage values.

Degenerate inputs (no context window, flag off, invalid ratio, estimator
exception) reproduce the legacy fixed-count cut exactly.

Ingest/replay/dedup paths untouched (PR #107 territory): zero edits to
_reconcile_ingest_cursor_from_store, the scaffold-overlap guard, or the
min-evidence bar at the replay-skip site.

Verification: 45 new tests (config plumbing, budget resolution, chokepoint
helper, frozen-K boundary consistency, AST guard + planted violations,
rotate floor incl. window-grow, fail-open ratio); full context_engine +
compaction-stats suite 348 passed; offline A/B shakedown on a mixed-density
corpus at 200K: ON-arm tail 98 msgs / 36,946 tokens vs budget 37,500,
legacy arm 32 msgs, 3 fires per arm, cadence_ratio 1.01 (gate ≤ 2.0).

Spec: ~/.hermes/plans/2026-07-01_lcm-target-ratio-token-budget-tail-SPEC.md
(v0.4, three Opus review passes: BLOCK → AWC → AWC, all blockers folded)
@Kyzcreig
Kyzcreig enabled auto-merge (squash) July 2, 2026 05:40
@Kyzcreig
Kyzcreig force-pushed the feat/lcm-target-ratio-tail-budget branch from e4d7a4a to cef9f62 Compare July 2, 2026 05:40
… + AugAssign list-growth coverage in the frozen-K AST guard

- LCM_TARGET_RATIO env override now applies the same (0,1] range guard as
  the config-file path (out-of-range → default 0.20), with tests.
- The frozen-K AST guard also flags augmented-assignment list growth
  (working_messages += [...]) inside the leaf loop — scoped to message-list
  vars / list-literal RHS so integer counters don't false-positive — with a
  planted += negative control.
@Kyzcreig
Kyzcreig merged commit 01dd1a7 into main Jul 2, 2026
32 checks passed
@Kyzcreig
Kyzcreig deleted the feat/lcm-target-ratio-tail-budget branch July 2, 2026 05:48
Kyzcreig added a commit that referenced this pull request Sep 25, 2026
…25 rows (t_caabb1fa)

49/84 rows. LCM rows graded against stephenschoettler/hermes-lcm @ 8d1b1e6
(v1.0.0-rc.1) per D1a and reconciled with the lcm-upstream-divergence-ledger
(DD-1 -> #168 SUPERSEDED, DD-2 -> #107 SUPERSEDED, DC-1/DC-2/DC-3 KEEP).
Three generic fixes upstream still lacks by read: #412, #902, dc4245a -> UPSTREAM.
Route disclosure corrected in plugins.md header (judgment on claude-fable-5-1).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant