Skip to content

fix(model-metadata): local ctx probe must not read max_tokens as the context window - #263

Merged
Kyzcreig merged 1 commit into
mainfrom
fix/local-ctx-probe-max-tokens
Jul 10, 2026
Merged

Kyzcreig merged 1 commit into
mainfrom
fix/local-ctx-probe-max-tokens

Conversation

@Kyzcreig

Copy link
Copy Markdown
Collaborator

The second half of the early-auto-compaction bug

Companion to #260. After the lm-studio-misdetection fix reclassified the fleet's Claude proxy as a generic local server, the local context-length probe (_query_local_context_length_uncached) hit the proxy's Anthropic passthrough GET /v1/models/{id}, which returns BOTH:

  • max_input_tokens: 1000000 — the context window
  • max_tokens: 128000 — the max output tokens

The alias chain read max_model_len or context_length or max_tokens, so it picked up max_tokens=128000 and cached it as the window → collapsed a 1M model to 128K → auto-compaction at ~96K (0.75 × 128K). This is why sessions kept compacting early even with the proxy correctly advertising context_length: 1000000.

Fix

Both the /v1/models/{id} and the /v1/models list branches now read genuine context-window fields only — max_model_len, context_length, context_window, max_input_tokens, max_position_embeddings — and drop max_tokens (output tokens, never the context window).

Verified live end-to-end against the running proxy: _query_local_context_length and the full get_model_context_length now resolve claude-fable-5 → 1000000, claude-haiku-4-5 → 200000.

Tests

Regression on both branches: an Anthropic-style passthrough with max_input_tokens+max_tokens must resolve to the input/context value, never the output cap. 36 pass in the local-ctx suite.

…context window

Second half of the early-auto-compaction bug. After the lm-studio-misdetection
fix reclassified the Claude proxy as a generic local server, the local
context-length probe (_query_local_context_length_uncached) hit the proxy's
Anthropic passthrough GET /v1/models/{id}, which returns BOTH max_input_tokens
(1000000, the context window) AND max_tokens (128000, the max OUTPUT tokens).
The alias chain read `max_model_len or context_length or max_tokens`, so it
picked up max_tokens=128000 and cached it as the window — collapsing a 1M model
to 128K and firing auto-compaction at ~96K (0.75*128K).

Fix: both the /v1/models/{id} and the /v1/models-list branches now read genuine
context-WINDOW fields only — max_model_len, context_length, context_window,
max_input_tokens, max_position_embeddings — and DROP max_tokens (which is output
tokens, never the context window). Verified live end-to-end: fable@:18801 now
resolves to 1000000, haiku to 200000.

Tests: Anthropic-passthrough max_input_tokens-wins regression on both the detail
and list branches.
@greptile-apps

greptile-apps Bot commented Jul 10, 2026

Copy link
Copy Markdown

Greptile Summary

This PR fixes premature auto-compaction caused by _query_local_context_length_uncached mistakenly reading max_tokens (the max output token cap) as the context window size on Anthropic-style proxy endpoints. The fix removes max_tokens from the alias chain in both the /v1/models/{id} detail branch and the /v1/models list branch, and replaces it with the correct context-window fields (context_window, max_input_tokens, max_position_embeddings).

  • agent/model_metadata.py: Both GET branches now probe only genuine context-window fields (max_model_len, context_length, context_window, max_input_tokens, max_position_embeddings) and explicitly exclude max_tokens, with a detailed comment explaining why.
  • tests/agent/test_model_metadata_local_ctx.py: Two regression tests are added — one for the detail-endpoint branch and one for the list-endpoint branch — each asserting that max_input_tokens (1 000 000) wins over max_tokens (128 000) in an Anthropic-style passthrough response.

Confidence Score: 5/5

Safe to merge — the change is a narrow two-line alias-chain edit that removes a misread field and adds semantically correct alternatives, backed by targeted regression tests for both affected code paths.

The root cause is clearly identified and the fix is minimal: max_tokens is removed from both the detail-endpoint and list-endpoint alias chains, replaced with fields that actually represent the context window. Both changed branches have dedicated regression tests with correctly ordered mock side-effects.

No files require special attention. Both changed files are straightforward and the test coverage directly mirrors the production failure scenario.

Important Files Changed

Filename Overview
agent/model_metadata.py Removes max_tokens from both alias chains and adds context_window, max_input_tokens, max_position_embeddings; change is minimal, well-commented, and correctly targeted
tests/agent/test_model_metadata_local_ctx.py Adds two regression tests covering both the detail-endpoint and list-endpoint branches; mocks are set up correctly and assertions are specific

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A[_query_local_context_length_uncached] --> B{server_type}
    B -->|ollama| C[POST /api/show\nread num_ctx or model_info]
    B -->|lm-studio| D[GET /api/v1/models\nread context_length]
    B -->|vllm or generic| E[GET /v1/models/model-id]

    E -->|200 OK| F["Context-window fields only\nmax_model_len OR context_length\nOR context_window OR max_input_tokens\nOR max_position_embeddings\nmax_tokens EXCLUDED"]
    F -->|value found| Z[return int ctx]
    F -->|no value| G[GET /v1/models list]
    E -->|non-200| G

    G -->|200 OK| H["Same field chain\nper matching model entry"]
    H -->|value found| Z
    H -->|not found| Y[return None]

    C --> Z
    D --> Z
Loading
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
    A[_query_local_context_length_uncached] --> B{server_type}
    B -->|ollama| C[POST /api/show\nread num_ctx or model_info]
    B -->|lm-studio| D[GET /api/v1/models\nread context_length]
    B -->|vllm or generic| E[GET /v1/models/model-id]

    E -->|200 OK| F["Context-window fields only\nmax_model_len OR context_length\nOR context_window OR max_input_tokens\nOR max_position_embeddings\nmax_tokens EXCLUDED"]
    F -->|value found| Z[return int ctx]
    F -->|no value| G[GET /v1/models list]
    E -->|non-200| G

    G -->|200 OK| H["Same field chain\nper matching model entry"]
    H -->|value found| Z
    H -->|not found| Y[return None]

    C --> Z
    D --> Z
Loading

Reviews (1): Last reviewed commit: "fix(model-metadata): local ctx probe mus..." | Re-trigger Greptile

@Kyzcreig
Kyzcreig merged commit ecec3e9 into main Jul 10, 2026
35 checks passed
@Kyzcreig
Kyzcreig deleted the fix/local-ctx-probe-max-tokens branch July 10, 2026 16:03
Kyzcreig added a commit that referenced this pull request Jul 26, 2026
…00 summary-role pin) (#431)

* vendor(lcm): cherry-pick 9 upstream fixes from hermes-lcm (03b74f8 -> selected from 49e99a2)

Upstream hermes-lcm went MIT (LICENSE added 2026-06-26, f7ae61f) — vendoring
and cherry-picking now permitted with attribution. Tier-1 picks per
/tmp/lcm-refresh-out/CHERRY-PICK-LIST.md, code-only (their tests/docs/changelog
excluded; our vendored copy carries no tests dir — coverage rides
tests/context_engine/):

  #263 preserve source lineage after long sessions
  #264 perf: aggregate DAG status stats
  #265 harden externalized payload durability
  #269 preserve raw session ownership across compression rollover
  #278 avoid payload integrity false positives from log examples
  #280 pin summary role to user after system anchor (Anthropic HTTP 400) <- highest value
  #285 make context engine deepcopy clone-safe (subagent spawn safety)
  #282 strip injected context before compaction
  (+) discard reasoning-only summaries (unclosed <think> = quality bug + prompt leak)

All 9 verified clean-apply by the refresh-analysis worker on a simulated copy
of our tree; re-applied here onto fork/main.

* test(compaction): regenerate in-turn reconcile fixture for the vendored sanitizer

The 9 LCM cherry-picks (3c1d61c) add one line to
_sanitize_active_context_messages (upstream pick #282, strip injected
context before compaction), which moves the fixture's
sanitizer_source_sha1 provenance hash and reds
test_fixture_sanitizer_provenance_current.

Regenerated via the committed generator:
  python tests/agent/fixtures/gen_inturn_reconcile_fixture.py

Diff is the provenance hash ONLY -- messages, compressed,
true_kept_count and fresh_tail_count are byte-identical. The new line
routes through _preserved_objective_context_content, which returns ""
unless a row starts with the preserved-objective prefix, so it is a
strict no-op on all 632 fixture rows (verified: 0 rows mutated) and the
real sanitizer output is unchanged. Only one of the four hashed
functions changed; the other three are byte-identical.

Follow-up candidate (not this PR): this provenance test is a
change-detector, which AGENTS.md discourages -- it should assert the
sanitizer's behaviour (live sanitize(raw_tail) == committed comp tail)
rather than pinning its source SHA.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
@Kyzcreig
Kyzcreig restored the fix/local-ctx-probe-max-tokens branch September 21, 2026 10:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant