fix(models): update Anthropic model contextLength to 1M - #7129
diegosouzapw merged 13 commits into
Conversation
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
…the default branch (diegosouzapw#7168)
… (queue_conditions alone are eligibility-only) (diegosouzapw#7179)
…uto_merge_conditions (rules-based path is EOL 2026-07-16) (diegosouzapw#7216)
…; free plan queue is serial) (diegosouzapw#7220)
…H-hosted build hang dequeued every attempt) (diegosouzapw#7225)
|
Thanks for catching that Anthropic's context windows moved — but this PR has a real routing bug as written. |
|
Thanks for the thorough review — you're absolutely right about the routing bug. Changes made1. Added
|
…copy hard-fails every PR (diegosouzapw#7341) main's copy of this test still does git I/O inside a unit test: const baseSrc = git(['show', 'origin/main:' + FILE]); Runners check out a shallow single ref, so origin/main does not resolve and the test dies with 'fatal: invalid object name origin/main'. Every PR into main fails Unit Tests (7/8) on it — today that is diegosouzapw#7313, diegosouzapw#7315, diegosouzapw#7316, diegosouzapw#7334, diegosouzapw#7336 and diegosouzapw#7337, six PRs red on a defect none of them introduced. diegosouzapw#7313 has no other red at all. release/v3.8.49 already carries a fix (2e42b8e, diegosouzapw#7174: try/catch, fetch origin/main on demand, t.skip() when unreachable), but it only reaches main at release time — so main stays broken for the whole cycle. Cherry-picking it would also import a new problem: PR Test Policy classifies t.skip() as a silenced assertion, which we watched it correctly catch on diegosouzapw#7300 today. This is the hermetic version instead (ported from diegosouzapw#7327, which does the same for the release branch): read the file straight off disk, compare against an empty base so baseTaut/baseExtTaut are 0 — the strictest possible comparison point — and call evaluateMasking() directly. No git ref, no fetch, no skip, nothing the runner's checkout depth can break. The diegosouzapw#6634 regression stays covered: the guard's logic lives in SELF_TEST_FIXTURE_RE (check-test-masking.mjs:337), not in the test. Proven both ways on main before committing — neutralise SELF_TEST_FIXTURE_RE to /$^/ and the test FAILS; restore it and it passes 2/2, with check-test-masking.mjs left byte-identical. Co-authored-by: growab <nekron@icloud.com>
…bers (diegosouzapw#7347) main's ratchet had been failing --require-tighten on every PR: 11 metrics improved but the baseline was never tightened. Same class as the diegosouzapw#6634 selfref guard — an infra fix that lands only on the release branch leaves main red for the whole cycle, and every PR into main pays for it. Values are the merged-coverage numbers from a run on main itself (a local run measures ~68% vs CI's ~80%; the baseline's own note warns about that gap). Only the 11 coverage values change — gitleaks and semgrepFindings keep main's own state. No changelog fragment: diegosouzapw#7326 carries it on release/v3.8.49, and a second one here would double the entry at release time.
Opus 4.6, Sonnet 4.6, and Sonnet 5 all have 1M context window since GA (2026-03-13). Update: - agyModels + antigravityModelAliases: 200000 -> 1048576 (binary 1M, matches existing convention in those files) - claude registry: 200000 -> 1000000 (decimal 1M, matches other entries in that file) Older models (opus-4-5, sonnet-4-5, haiku-4-5) and defaultContextLength: 200000 left untouched. Signed-off-by: Minxi Hou <houminxi@gmail.com>
Sonnet 4.6 has 1M context GA since 2026-02-17. Without this entry, the CC-compatible wire image omits the context-1m beta header for Sonnet 4.6 requests, causing large-context requests to be rejected by Anthropic.
Verify that every Claude model with contextLength > 200K has a matching entry in the beta header allowlist, and vice versa. Parses source files directly (no module imports needed). Signed-off-by: Minxi Hou <houminxi@gmail.com>
Add sanity check that parser finds at least one model. Handle type annotations in CONTEXT_1M_SUPPORTED_MODELS assignment. Remove weaker subagent test. Signed-off-by: Minxi Hou <houminxi@gmail.com>
be8627d to
9e0f0d2
Compare
# Conflicts: # .mergify.yml
…4.6 fix Three pre-existing tests hardcoded the old 200000 contextLength for claude-opus-4-6-thinking / claude-sonnet-4-6 / claude-sonnet-5, which this PR's own registry change (agyModels.ts, antigravityModelAliases.ts, claude registry) legitimately bumped to 1M GA: - antigravity-model-aliases.test.ts: deepEqual assertions for claude-opus-4-6-thinking and claude-sonnet-5 in the Antigravity catalog expected contextLength: 200000; now 1048576. Added an explicit contextLength assertion for claude-sonnet-4-6 too. - auto-combo-context-advertising.test.ts: resolveComboContextLimit regression test asserted the claude target's own limit as 200000; the test's intent (own limit wins over an 8k sibling, not compressed) is unchanged, only the registry's current correct value (1000000) needed updating. Also refreshed a stale comment in the MAX-of-candidates test (functionally unaffected, since gemini's 1048576 already wins either way). - models-catalog-route.test.ts: refreshed a stale inline comment (the actual assertion checks the MIN across combo targets, 128000, unaffected by the 1M bump). Verified via Anthropic's own docs (platform.claude.com/docs/en/build-with-claude/context-windows): "Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5, and Claude Sonnet 4.6 have a 1M-token context window ... on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry" -- matches this PR's claimed 2026-03-13 GA date exactly, and Google Cloud coverage corroborates the Antigravity-hosted ids. Independently corroborated for the Antigravity-specific path by unrelated third-party projects fixing the same gap (earendil-works/pi#2209, earendil-works/pi#2194). Swept the full test suite for other contextLength/contextWindow assertions against these three model ids (grep + targeted runs across models-catalog-route, model-capabilities-registry, t31-t33-t34-t38-model-specs, executor-antigravity, agy-provider, provider-models-config, model-metadata-registry, claude-web-sonnet5-registry, combo-routing-engine, command-code-executor, combo-lockout-quota-reset -- 157 tests green); the Bedrock-hosted registry and modelSpecs.ts module were already correctly at 1000000 for these models, reinforcing this PR's direction. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Thanks @HouMinXi! Merged into release/v3.8.49 after merge-train validation on the combined tree. The fallback-chain try/catch was split into the proFallbackChain module to clear the complexity gates — authorship preserved. |
#7129 added claude-sonnet-4-6 to CONTEXT_1M_SUPPORTED_MODELS (1M context GA'd 2026-02-17) but missed this test in its sweep. A non-CC anthropic-compatible target with extendedContext:true now legitimately receives the context-1m beta header for this model — updating the stale undefined expectation.
…#7129) * chore(ci): add .mergify.yml to main — Mergify only reads config from the default branch (diegosouzapw#7168) * fix(ci): add the auto-enqueue pull_request_rule to the Mergify config (queue_conditions alone are eligibility-only) (diegosouzapw#7179) * fix(ci): migrate Mergify auto-enqueue to merge_protections_settings.auto_merge_conditions (rules-based path is EOL 2026-07-16) (diegosouzapw#7216) * fix(ci): drop Mergify batch settings (batching is a paid-tier feature; free plan queue is serial) (diegosouzapw#7220) * fix(ci): merge queue tolerates the advisory dast-smoke failure (its GH-hosted build hang dequeued every attempt) (diegosouzapw#7225) * test(ci): make the diegosouzapw#6634 selfref guard hermetic — main's copy hard-fails every PR (diegosouzapw#7341) main's copy of this test still does git I/O inside a unit test: const baseSrc = git(['show', 'origin/main:' + FILE]); Runners check out a shallow single ref, so origin/main does not resolve and the test dies with 'fatal: invalid object name origin/main'. Every PR into main fails Unit Tests (7/8) on it — today that is diegosouzapw#7313, diegosouzapw#7315, diegosouzapw#7316, diegosouzapw#7334, diegosouzapw#7336 and diegosouzapw#7337, six PRs red on a defect none of them introduced. diegosouzapw#7313 has no other red at all. release/v3.8.49 already carries a fix (8bbd411, diegosouzapw#7174: try/catch, fetch origin/main on demand, t.skip() when unreachable), but it only reaches main at release time — so main stays broken for the whole cycle. Cherry-picking it would also import a new problem: PR Test Policy classifies t.skip() as a silenced assertion, which we watched it correctly catch on diegosouzapw#7300 today. This is the hermetic version instead (ported from diegosouzapw#7327, which does the same for the release branch): read the file straight off disk, compare against an empty base so baseTaut/baseExtTaut are 0 — the strictest possible comparison point — and call evaluateMasking() directly. No git ref, no fetch, no skip, nothing the runner's checkout depth can break. The diegosouzapw#6634 regression stays covered: the guard's logic lives in SELF_TEST_FIXTURE_RE (check-test-masking.mjs:337), not in the test. Proven both ways on main before committing — neutralise SELF_TEST_FIXTURE_RE to /$^/ and the test FAILS; restore it and it passes 2/2, with check-test-masking.mjs left byte-identical. Co-authored-by: growab <nekron@icloud.com> * chore(quality): tighten main's coverage baseline to the CI's real numbers (diegosouzapw#7347) main's ratchet had been failing --require-tighten on every PR: 11 metrics improved but the baseline was never tightened. Same class as the diegosouzapw#6634 selfref guard — an infra fix that lands only on the release branch leaves main red for the whole cycle, and every PR into main pays for it. Values are the merged-coverage numbers from a run on main itself (a local run measures ~68% vs CI's ~80%; the baseline's own note warns about that gap). Only the 11 coverage values change — gitleaks and semgrepFindings keep main's own state. No changelog fragment: diegosouzapw#7326 carries it on release/v3.8.49, and a second one here would double the entry at release time. * fix(models): update Anthropic model contextLength to 1M Opus 4.6, Sonnet 4.6, and Sonnet 5 all have 1M context window since GA (2026-03-13). Update: - agyModels + antigravityModelAliases: 200000 -> 1048576 (binary 1M, matches existing convention in those files) - claude registry: 200000 -> 1000000 (decimal 1M, matches other entries in that file) Older models (opus-4-5, sonnet-4-5, haiku-4-5) and defaultContextLength: 200000 left untouched. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(claude): add sonnet-4-6 to CONTEXT_1M_SUPPORTED_MODELS Sonnet 4.6 has 1M context GA since 2026-02-17. Without this entry, the CC-compatible wire image omits the context-1m beta header for Sonnet 4.6 requests, causing large-context requests to be rejected by Anthropic. * test: regression test for CONTEXT_1M_SUPPORTED_MODELS allowlist Verify that every Claude model with contextLength > 200K has a matching entry in the beta header allowlist, and vice versa. Parses source files directly (no module imports needed). Signed-off-by: Minxi Hou <houminxi@gmail.com> * test: improve regression test with empty-registry guard Add sanity check that parser finds at least one model. Handle type annotations in CONTEXT_1M_SUPPORTED_MODELS assignment. Remove weaker subagent test. Signed-off-by: Minxi Hou <houminxi@gmail.com> * test: update stale contextLength expectations for the 1M Sonnet/Opus 4.6 fix Three pre-existing tests hardcoded the old 200000 contextLength for claude-opus-4-6-thinking / claude-sonnet-4-6 / claude-sonnet-5, which this PR's own registry change (agyModels.ts, antigravityModelAliases.ts, claude registry) legitimately bumped to 1M GA: - antigravity-model-aliases.test.ts: deepEqual assertions for claude-opus-4-6-thinking and claude-sonnet-5 in the Antigravity catalog expected contextLength: 200000; now 1048576. Added an explicit contextLength assertion for claude-sonnet-4-6 too. - auto-combo-context-advertising.test.ts: resolveComboContextLimit regression test asserted the claude target's own limit as 200000; the test's intent (own limit wins over an 8k sibling, not compressed) is unchanged, only the registry's current correct value (1000000) needed updating. Also refreshed a stale comment in the MAX-of-candidates test (functionally unaffected, since gemini's 1048576 already wins either way). - models-catalog-route.test.ts: refreshed a stale inline comment (the actual assertion checks the MIN across combo targets, 128000, unaffected by the 1M bump). Verified via Anthropic's own docs (platform.claude.com/docs/en/build-with-claude/context-windows): "Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5, and Claude Sonnet 4.6 have a 1M-token context window ... on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry" -- matches this PR's claimed 2026-03-13 GA date exactly, and Google Cloud coverage corroborates the Antigravity-hosted ids. Independently corroborated for the Antigravity-specific path by unrelated third-party projects fixing the same gap (earendil-works/pi#2209, earendil-works/pi#2194). Swept the full test suite for other contextLength/contextWindow assertions against these three model ids (grep + targeted runs across models-catalog-route, model-capabilities-registry, t31-t33-t34-t38-model-specs, executor-antigravity, agy-provider, provider-models-config, model-metadata-registry, claude-web-sonnet5-registry, combo-routing-engine, command-code-executor, combo-lockout-quota-reset -- 157 tests green); the Bedrock-hosted registry and modelSpecs.ts module were already correctly at 1000000 for these models, reinforcing this PR's direction. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: growab <nekron@icloud.com> Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
…zapw#7129) diegosouzapw#7129 added claude-sonnet-4-6 to CONTEXT_1M_SUPPORTED_MODELS (1M context GA'd 2026-02-17) but missed this test in its sweep. A non-CC anthropic-compatible target with extendedContext:true now legitimately receives the context-1m beta header for this model — updating the stale undefined expectation.
…#7129) * chore(ci): add .mergify.yml to main — Mergify only reads config from the default branch (diegosouzapw#7168) * fix(ci): add the auto-enqueue pull_request_rule to the Mergify config (queue_conditions alone are eligibility-only) (diegosouzapw#7179) * fix(ci): migrate Mergify auto-enqueue to merge_protections_settings.auto_merge_conditions (rules-based path is EOL 2026-07-16) (diegosouzapw#7216) * fix(ci): drop Mergify batch settings (batching is a paid-tier feature; free plan queue is serial) (diegosouzapw#7220) * fix(ci): merge queue tolerates the advisory dast-smoke failure (its GH-hosted build hang dequeued every attempt) (diegosouzapw#7225) * test(ci): make the diegosouzapw#6634 selfref guard hermetic — main's copy hard-fails every PR (diegosouzapw#7341) main's copy of this test still does git I/O inside a unit test: const baseSrc = git(['show', 'origin/main:' + FILE]); Runners check out a shallow single ref, so origin/main does not resolve and the test dies with 'fatal: invalid object name origin/main'. Every PR into main fails Unit Tests (7/8) on it — today that is diegosouzapw#7313, diegosouzapw#7315, diegosouzapw#7316, diegosouzapw#7334, diegosouzapw#7336 and diegosouzapw#7337, six PRs red on a defect none of them introduced. diegosouzapw#7313 has no other red at all. release/v3.8.49 already carries a fix (83a7551, diegosouzapw#7174: try/catch, fetch origin/main on demand, t.skip() when unreachable), but it only reaches main at release time — so main stays broken for the whole cycle. Cherry-picking it would also import a new problem: PR Test Policy classifies t.skip() as a silenced assertion, which we watched it correctly catch on diegosouzapw#7300 today. This is the hermetic version instead (ported from diegosouzapw#7327, which does the same for the release branch): read the file straight off disk, compare against an empty base so baseTaut/baseExtTaut are 0 — the strictest possible comparison point — and call evaluateMasking() directly. No git ref, no fetch, no skip, nothing the runner's checkout depth can break. The diegosouzapw#6634 regression stays covered: the guard's logic lives in SELF_TEST_FIXTURE_RE (check-test-masking.mjs:337), not in the test. Proven both ways on main before committing — neutralise SELF_TEST_FIXTURE_RE to /$^/ and the test FAILS; restore it and it passes 2/2, with check-test-masking.mjs left byte-identical. Co-authored-by: growab <nekron@icloud.com> * chore(quality): tighten main's coverage baseline to the CI's real numbers (diegosouzapw#7347) main's ratchet had been failing --require-tighten on every PR: 11 metrics improved but the baseline was never tightened. Same class as the diegosouzapw#6634 selfref guard — an infra fix that lands only on the release branch leaves main red for the whole cycle, and every PR into main pays for it. Values are the merged-coverage numbers from a run on main itself (a local run measures ~68% vs CI's ~80%; the baseline's own note warns about that gap). Only the 11 coverage values change — gitleaks and semgrepFindings keep main's own state. No changelog fragment: diegosouzapw#7326 carries it on release/v3.8.49, and a second one here would double the entry at release time. * fix(models): update Anthropic model contextLength to 1M Opus 4.6, Sonnet 4.6, and Sonnet 5 all have 1M context window since GA (2026-03-13). Update: - agyModels + antigravityModelAliases: 200000 -> 1048576 (binary 1M, matches existing convention in those files) - claude registry: 200000 -> 1000000 (decimal 1M, matches other entries in that file) Older models (opus-4-5, sonnet-4-5, haiku-4-5) and defaultContextLength: 200000 left untouched. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(claude): add sonnet-4-6 to CONTEXT_1M_SUPPORTED_MODELS Sonnet 4.6 has 1M context GA since 2026-02-17. Without this entry, the CC-compatible wire image omits the context-1m beta header for Sonnet 4.6 requests, causing large-context requests to be rejected by Anthropic. * test: regression test for CONTEXT_1M_SUPPORTED_MODELS allowlist Verify that every Claude model with contextLength > 200K has a matching entry in the beta header allowlist, and vice versa. Parses source files directly (no module imports needed). Signed-off-by: Minxi Hou <houminxi@gmail.com> * test: improve regression test with empty-registry guard Add sanity check that parser finds at least one model. Handle type annotations in CONTEXT_1M_SUPPORTED_MODELS assignment. Remove weaker subagent test. Signed-off-by: Minxi Hou <houminxi@gmail.com> * test: update stale contextLength expectations for the 1M Sonnet/Opus 4.6 fix Three pre-existing tests hardcoded the old 200000 contextLength for claude-opus-4-6-thinking / claude-sonnet-4-6 / claude-sonnet-5, which this PR's own registry change (agyModels.ts, antigravityModelAliases.ts, claude registry) legitimately bumped to 1M GA: - antigravity-model-aliases.test.ts: deepEqual assertions for claude-opus-4-6-thinking and claude-sonnet-5 in the Antigravity catalog expected contextLength: 200000; now 1048576. Added an explicit contextLength assertion for claude-sonnet-4-6 too. - auto-combo-context-advertising.test.ts: resolveComboContextLimit regression test asserted the claude target's own limit as 200000; the test's intent (own limit wins over an 8k sibling, not compressed) is unchanged, only the registry's current correct value (1000000) needed updating. Also refreshed a stale comment in the MAX-of-candidates test (functionally unaffected, since gemini's 1048576 already wins either way). - models-catalog-route.test.ts: refreshed a stale inline comment (the actual assertion checks the MIN across combo targets, 128000, unaffected by the 1M bump). Verified via Anthropic's own docs (platform.claude.com/docs/en/build-with-claude/context-windows): "Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5, and Claude Sonnet 4.6 have a 1M-token context window ... on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry" -- matches this PR's claimed 2026-03-13 GA date exactly, and Google Cloud coverage corroborates the Antigravity-hosted ids. Independently corroborated for the Antigravity-specific path by unrelated third-party projects fixing the same gap (earendil-works/pi#2209, earendil-works/pi#2194). Swept the full test suite for other contextLength/contextWindow assertions against these three model ids (grep + targeted runs across models-catalog-route, model-capabilities-registry, t31-t33-t34-t38-model-specs, executor-antigravity, agy-provider, provider-models-config, model-metadata-registry, claude-web-sonnet5-registry, combo-routing-engine, command-code-executor, combo-lockout-quota-reset -- 157 tests green); the Bedrock-hosted registry and modelSpecs.ts module were already correctly at 1000000 for these models, reinforcing this PR's direction. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: growab <nekron@icloud.com> Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
…zapw#7129) diegosouzapw#7129 added claude-sonnet-4-6 to CONTEXT_1M_SUPPORTED_MODELS (1M context GA'd 2026-02-17) but missed this test in its sweep. A non-CC anthropic-compatible target with extendedContext:true now legitimately receives the context-1m beta header for this model — updating the stale undefined expectation.
Opus 4.6, Sonnet 4.6, and Sonnet 5 all have 1M context window since GA (2026-03-13).
Changes:
Older models (4.5 series) and defaultContextLength left untouched.