Repository navigation
fix(model-metadata): local ctx probe must not read max_tokens as the context window - #263
Conversation
…context window
Second half of the early-auto-compaction bug. After the lm-studio-misdetection
fix reclassified the Claude proxy as a generic local server, the local
context-length probe (_query_local_context_length_uncached) hit the proxy's
Anthropic passthrough GET /v1/models/{id}, which returns BOTH max_input_tokens
(1000000, the context window) AND max_tokens (128000, the max OUTPUT tokens).
The alias chain read `max_model_len or context_length or max_tokens`, so it
picked up max_tokens=128000 and cached it as the window — collapsing a 1M model
to 128K and firing auto-compaction at ~96K (0.75*128K).
Fix: both the /v1/models/{id} and the /v1/models-list branches now read genuine
context-WINDOW fields only — max_model_len, context_length, context_window,
max_input_tokens, max_position_embeddings — and DROP max_tokens (which is output
tokens, never the context window). Verified live end-to-end: fable@:18801 now
resolves to 1000000, haiku to 200000.
Tests: Anthropic-passthrough max_input_tokens-wins regression on both the detail
and list branches.
|
| Filename | Overview |
|---|---|
| agent/model_metadata.py | Removes max_tokens from both alias chains and adds context_window, max_input_tokens, max_position_embeddings; change is minimal, well-commented, and correctly targeted |
| tests/agent/test_model_metadata_local_ctx.py | Adds two regression tests covering both the detail-endpoint and list-endpoint branches; mocks are set up correctly and assertions are specific |
Flowchart
%%{init: {'theme': 'neutral'}}%%
flowchart TD
A[_query_local_context_length_uncached] --> B{server_type}
B -->|ollama| C[POST /api/show\nread num_ctx or model_info]
B -->|lm-studio| D[GET /api/v1/models\nread context_length]
B -->|vllm or generic| E[GET /v1/models/model-id]
E -->|200 OK| F["Context-window fields only\nmax_model_len OR context_length\nOR context_window OR max_input_tokens\nOR max_position_embeddings\nmax_tokens EXCLUDED"]
F -->|value found| Z[return int ctx]
F -->|no value| G[GET /v1/models list]
E -->|non-200| G
G -->|200 OK| H["Same field chain\nper matching model entry"]
H -->|value found| Z
H -->|not found| Y[return None]
C --> Z
D --> Z
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
A[_query_local_context_length_uncached] --> B{server_type}
B -->|ollama| C[POST /api/show\nread num_ctx or model_info]
B -->|lm-studio| D[GET /api/v1/models\nread context_length]
B -->|vllm or generic| E[GET /v1/models/model-id]
E -->|200 OK| F["Context-window fields only\nmax_model_len OR context_length\nOR context_window OR max_input_tokens\nOR max_position_embeddings\nmax_tokens EXCLUDED"]
F -->|value found| Z[return int ctx]
F -->|no value| G[GET /v1/models list]
E -->|non-200| G
G -->|200 OK| H["Same field chain\nper matching model entry"]
H -->|value found| Z
H -->|not found| Y[return None]
C --> Z
D --> Z
Reviews (1): Last reviewed commit: "fix(model-metadata): local ctx probe mus..." | Re-trigger Greptile
…00 summary-role pin) (#431) * vendor(lcm): cherry-pick 9 upstream fixes from hermes-lcm (03b74f8 -> selected from 49e99a2) Upstream hermes-lcm went MIT (LICENSE added 2026-06-26, f7ae61f) — vendoring and cherry-picking now permitted with attribution. Tier-1 picks per /tmp/lcm-refresh-out/CHERRY-PICK-LIST.md, code-only (their tests/docs/changelog excluded; our vendored copy carries no tests dir — coverage rides tests/context_engine/): #263 preserve source lineage after long sessions #264 perf: aggregate DAG status stats #265 harden externalized payload durability #269 preserve raw session ownership across compression rollover #278 avoid payload integrity false positives from log examples #280 pin summary role to user after system anchor (Anthropic HTTP 400) <- highest value #285 make context engine deepcopy clone-safe (subagent spawn safety) #282 strip injected context before compaction (+) discard reasoning-only summaries (unclosed <think> = quality bug + prompt leak) All 9 verified clean-apply by the refresh-analysis worker on a simulated copy of our tree; re-applied here onto fork/main. * test(compaction): regenerate in-turn reconcile fixture for the vendored sanitizer The 9 LCM cherry-picks (3c1d61c) add one line to _sanitize_active_context_messages (upstream pick #282, strip injected context before compaction), which moves the fixture's sanitizer_source_sha1 provenance hash and reds test_fixture_sanitizer_provenance_current. Regenerated via the committed generator: python tests/agent/fixtures/gen_inturn_reconcile_fixture.py Diff is the provenance hash ONLY -- messages, compressed, true_kept_count and fresh_tail_count are byte-identical. The new line routes through _preserved_objective_context_content, which returns "" unless a row starts with the preserved-objective prefix, so it is a strict no-op on all 632 fixture rows (verified: 0 rows mutated) and the real sanitizer output is unchanged. Only one of the four hashed functions changed; the other three are byte-identical. Follow-up candidate (not this PR): this provenance test is a change-detector, which AGENTS.md discourages -- it should assert the sanitizer's behaviour (live sanitize(raw_tail) == committed comp tail) rather than pinning its source SHA. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
The second half of the early-auto-compaction bug
Companion to #260. After the lm-studio-misdetection fix reclassified the fleet's Claude proxy as a generic local server, the local context-length probe (
_query_local_context_length_uncached) hit the proxy's Anthropic passthroughGET /v1/models/{id}, which returns BOTH:max_input_tokens: 1000000— the context windowmax_tokens: 128000— the max output tokensThe alias chain read
max_model_len or context_length or max_tokens, so it picked upmax_tokens=128000and cached it as the window → collapsed a 1M model to 128K → auto-compaction at ~96K (0.75 × 128K). This is why sessions kept compacting early even with the proxy correctly advertisingcontext_length: 1000000.Fix
Both the
/v1/models/{id}and the/v1/modelslist branches now read genuine context-window fields only —max_model_len,context_length,context_window,max_input_tokens,max_position_embeddings— and dropmax_tokens(output tokens, never the context window).Verified live end-to-end against the running proxy:
_query_local_context_lengthand the fullget_model_context_lengthnow resolveclaude-fable-5 → 1000000,claude-haiku-4-5 → 200000.Tests
Regression on both branches: an Anthropic-style passthrough with
max_input_tokens+max_tokensmust resolve to the input/context value, never the output cap. 36 pass in the local-ctx suite.