Skip to content

feat: intelligent auto-router — routes to fastest/cheapest provider automatically - #2

Merged
kevincodex1 merged 1 commit into
Twigpine:mainfrom
gnanam1990:feat/smart-auto-router
Apr 1, 2026
Merged

kevincodex1 merged 1 commit into
Twigpine:mainfrom
gnanam1990:feat/smart-auto-router

Conversation

@gnanam1990

Copy link
Copy Markdown
Collaborator

What this adds

The Problem

openclaude currently routes ALL requests to ONE fixed provider.
If that provider is slow or down, every request suffers.

The Solution: SmartRouter

smart_router.py ? a drop-in intelligent routing layer that:

?? Benchmarks providers on startup

Pings OpenAI, Gemini, and Ollama to measure real latency before
routing any requests.

?? Scores every provider

Each provider gets a score based on:

  • Latency ? how fast it responds (measured in real-time)
  • Cost ? price per 1k tokens (Ollama=free, Gemini=cheap, OpenAI=mid)
  • Error rate ? providers with failures get penalized

??? Routes intelligently

  • Short prompts ? small/fast model on best provider
  • Long prompts ? big model on best provider
  • Configurable strategy: latency, cost, or balanced

?? Automatic fallback

If a provider fails mid-request, SmartRouter:

  1. Marks it unhealthy
  2. Re-pings it after 60 seconds
  3. Restores it automatically when it recovers

?? Learns from real traffic

Uses exponential moving average to update latency scores
from actual request timings ? gets smarter over time.

Usage

Add to server.py:

from smart_router import SmartRouter
router = SmartRouter()
await router.initialize()

# In your request handler:
decision = await router.route(messages, claude_model)
# decision = { 'provider': 'gemini', 'model': 'gemini-2.5-pro', ... }

# After request completes:
await router.record_result(decision['provider'], success=True, duration_ms=230)

.env config

ROUTER_MODE=smart
ROUTER_STRATEGY=balanced   # latency | cost | balanced
ROUTER_FALLBACK=true

Files

  • smart_router.py ? full implementation (250 lines)
  • test_smart_router.py ? 20 unit tests, all passing

Test

pytest test_smart_router.py -v

Contributor: gnanam1990 (PR #3)

@kevincodex1
kevincodex1 merged commit 55098bf into Twigpine:main Apr 1, 2026
kevincodex1 pushed a commit that referenced this pull request Apr 2, 2026
…ension

Initial VS Code Extension for OpenClaude
@boofpackdev

Copy link
Copy Markdown

✅ Build tooling verification complete - all acceptance criteria met:

Verification Results:

  • ✅ bun install completes without errors (695 packages)
  • ✅ bun run build produces clean dist/ output (dist/cli.mjs ~20MB)
  • ✅ bun run dev starts successfully (verified with --help flag)
  • ✅ bun run smoke passes (version 1.0.0)
  • ✅ bun test runs (338 pass, 33 fail - pre-existing test failures unrelated to build tooling)

Scripts verified for hermes use case:

Flo5k5 added a commit to Flo5k5/openclaude that referenced this pull request Apr 12, 2026
- Fix permission rule field: expression → ruleContent (Copilot #1)
- Handle empty command prefix: skip rule creation (Copilot Twigpine#2)
- Remove unused useTheme() import (Copilot Twigpine#3)
- Save permission rules under 'Bash' toolName so bashToolHasPermission
  can match them — Monitor delegates to Bash permission system (Copilot Twigpine#4)
- Remove unused logError import from MonitorMcpTask (Copilot Twigpine#6)
- Copilot Twigpine#5 (getAppState throws): same pattern as BashTool:915, not a bug
euxaristia referenced this pull request in euxaristia/openclaude Apr 13, 2026
feat: intelligent auto-router — routes to fastest/cheapest provider automatically
euxaristia referenced this pull request in euxaristia/openclaude Apr 13, 2026
…ension

Initial VS Code Extension for OpenClaude
kevincodex1 pushed a commit that referenced this pull request Apr 13, 2026
* feat: implement Monitor tool for streaming shell output

Add the Monitor tool that executes shell commands in the background and
streams stdout line-by-line as notifications to the model. This enables
real-time monitoring of logs, builds, and long-running processes.

Implementation:
- MonitorTool (src/tools/MonitorTool/) — spawns LocalShellTask with
  kind='monitor', returns immediately with task ID
- MonitorMcpTask (src/tasks/MonitorMcpTask/) — task lifecycle management
  and agent cleanup via killMonitorMcpTasksForAgent()
- MonitorPermissionRequest — permission dialog component

The codebase already had all integration points wired (tools.ts, tasks.ts,
PermissionRequest.tsx, LocalShellTask kind='monitor', BashTool prompt).
This PR provides the missing implementations.

* fix: command-specific permission rule + architecture docs

- MonitorPermissionRequest: "don't ask again" now creates a
  command-prefix rule (like BashTool) instead of a blanket
  tool-name-only rule that would auto-allow all Monitor commands
- MonitorMcpTask: clarify architecture comments explaining why
  monitor_mcp type exists as a registry stub while actual tasks
  are local_bash with kind='monitor'

* fix: address Copilot review feedback

- Fix permission rule field: expression → ruleContent (Copilot #1)
- Handle empty command prefix: skip rule creation (Copilot #2)
- Remove unused useTheme() import (Copilot #3)
- Save permission rules under 'Bash' toolName so bashToolHasPermission
  can match them — Monitor delegates to Bash permission system (Copilot #4)
- Remove unused logError import from MonitorMcpTask (Copilot #6)
- Copilot #5 (getAppState throws): same pattern as BashTool:915, not a bug
kevincodex1 referenced this pull request in kevincodex1/openclaude Apr 21, 2026
Addresses the security review on feat/bash-command-safety-classifier.
Each finding traced by reviewer is fixed locally; adversarial tests cover
each bypass pattern.

P0 — parser bypasses (CRITICAL / HIGH):

  #1  Escaped backslash before close quote — `echo "test\\" && rm -rf /`
      slipped through as one quoted echo. The prior `input[i-1] !== '\\'`
      check fails when the backslash itself is escaped. Replaced with an
      isCharEscaped() helper that counts consecutive backslashes; a quote
      is escaped only when preceded by an odd count. Applied in BOTH
      splitCompound AND tokenize so they agree on quote boundaries.

  #2  Process substitution not detected — `cat <(curl ...)` passed the cat
      safe gate. tokenize now returns null on `<(` / `>(` same as it did
      for `$(` / backticks, degrading to 'unknown'.

  Twigpine#3  Newline as command separator — `echo safe\nrm -rf /` was treated as
      one line by splitCompound (bash splits on \n = ;). Added \n to the
      separator list alongside ; | &.

P1 — classification bypasses (MEDIUM):

  Twigpine#4  Sensitive-path denylist — cat/head/tail/less/more/file/stat/wc and
      readlink/realpath now consult SENSITIVE_PATH_PATTERNS. /etc/shadow,
      ~/.ssh/id_rsa, /proc/*/environ, /dev/sd*, ~/.aws/credentials,
      ~/.kube/config, id_rsa / *.pem / *.key etc. degrade to 'unknown'.
      Public keys (*.pub) and normal files remain safe.

  Twigpine#5  `git config key value` misclassified — removed `config` from the
      READ_ONLY_SUBCOMMANDS list and added explicit unsafe gate: two+
      positional args with no --get/--list is a set, marked unsafe.
      --unset / --replace-all / --add / --unset-all also marked unsafe.
      Single-arg read form falls to 'unknown' (defers to existing rules).

  Twigpine#6  `git stash` (bare) misclassified — equivalent to `git stash push`,
      mutates worktree + index. Safe path now requires `git stash list`
      or `git stash show`; everything else in stash falls to unsafe or
      unknown.

P2 — hygiene (LOW):

  Twigpine#7  env / printenv removed from ALWAYS_SAFE_COMMANDS. Plain `env` dumps
      all environment variables (API keys, tokens) to the caller — not
      safe to auto-approve. FOO=bar cmd idiom still works via the existing
      env-assignment-prefix stripping.

  Twigpine#8  git gc / prune / repack / bisect removed from READ_ONLY_SUBCOMMANDS
      and added to the unsafe gate. gc repacks + deletes loose objects;
      bisect start/good/bad/reset/run/skip/terms/replay mutate HEAD.

Tests: 35 new adversarial cases across all 8 findings, plus regression
coverage (public keys still safe, FOO=bar pwd still safe, git stash
list/show still safe). Full suite: 1187/1187 pass. PR intent scan: clean.

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>
0ucb added a commit to 0ucb/flawed-code that referenced this pull request Apr 25, 2026
…3.6 shim enhancements

- Bug Twigpine#1: Remove isLocalProviderUrl guard on stream_options; always send include_usage
- Bug Twigpine#1: Add retry escape hatch when server rejects stream_options with 400
- Bug Twigpine#2: Add isBinaryContent/sanitizeBinaryOutput to prevent ESC sequence injection from raw shell output
- Qwen3.6: Add isQwenModel detection, presence_penalty, top_k, chat_template_kwargs, reasoning_content preservation
- Docs: Update optimization plan with verified dependency map and low-effort agent task plan
reymaster pushed a commit to reymaster/openclaude that referenced this pull request May 5, 2026
feat: intelligent auto-router — routes to fastest/cheapest provider automatically
reymaster pushed a commit to reymaster/openclaude that referenced this pull request May 5, 2026
…code-extension

Initial VS Code Extension for OpenClaude
kevincodex1 referenced this pull request in kevincodex1/openclaude May 6, 2026
Addresses the security review on feat/bash-command-safety-classifier.
Each finding traced by reviewer is fixed locally; adversarial tests cover
each bypass pattern.

P0 — parser bypasses (CRITICAL / HIGH):

  #1  Escaped backslash before close quote — `echo "test\\" && rm -rf /`
      slipped through as one quoted echo. The prior `input[i-1] !== '\\'`
      check fails when the backslash itself is escaped. Replaced with an
      isCharEscaped() helper that counts consecutive backslashes; a quote
      is escaped only when preceded by an odd count. Applied in BOTH
      splitCompound AND tokenize so they agree on quote boundaries.

  #2  Process substitution not detected — `cat <(curl ...)` passed the cat
      safe gate. tokenize now returns null on `<(` / `>(` same as it did
      for `$(` / backticks, degrading to 'unknown'.

  Twigpine#3  Newline as command separator — `echo safe\nrm -rf /` was treated as
      one line by splitCompound (bash splits on \n = ;). Added \n to the
      separator list alongside ; | &.

P1 — classification bypasses (MEDIUM):

  Twigpine#4  Sensitive-path denylist — cat/head/tail/less/more/file/stat/wc and
      readlink/realpath now consult SENSITIVE_PATH_PATTERNS. /etc/shadow,
      ~/.ssh/id_rsa, /proc/*/environ, /dev/sd*, ~/.aws/credentials,
      ~/.kube/config, id_rsa / *.pem / *.key etc. degrade to 'unknown'.
      Public keys (*.pub) and normal files remain safe.

  Twigpine#5  `git config key value` misclassified — removed `config` from the
      READ_ONLY_SUBCOMMANDS list and added explicit unsafe gate: two+
      positional args with no --get/--list is a set, marked unsafe.
      --unset / --replace-all / --add / --unset-all also marked unsafe.
      Single-arg read form falls to 'unknown' (defers to existing rules).

  Twigpine#6  `git stash` (bare) misclassified — equivalent to `git stash push`,
      mutates worktree + index. Safe path now requires `git stash list`
      or `git stash show`; everything else in stash falls to unsafe or
      unknown.

P2 — hygiene (LOW):

  Twigpine#7  env / printenv removed from ALWAYS_SAFE_COMMANDS. Plain `env` dumps
      all environment variables (API keys, tokens) to the caller — not
      safe to auto-approve. FOO=bar cmd idiom still works via the existing
      env-assignment-prefix stripping.

  Twigpine#8  git gc / prune / repack / bisect removed from READ_ONLY_SUBCOMMANDS
      and added to the unsafe gate. gc repacks + deletes loose objects;
      bisect start/good/bad/reset/run/skip/terms/replay mutate HEAD.

Tests: 35 new adversarial cases across all 8 findings, plus regression
coverage (public keys still safe, FOO=bar pwd still safe, git stash
list/show still safe). Full suite: 1187/1187 pass. PR intent scan: clean.

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>
The-FOOL-00 pushed a commit to The-FOOL-00/openclaude that referenced this pull request May 24, 2026
* feat: implement Monitor tool for streaming shell output

Add the Monitor tool that executes shell commands in the background and
streams stdout line-by-line as notifications to the model. This enables
real-time monitoring of logs, builds, and long-running processes.

Implementation:
- MonitorTool (src/tools/MonitorTool/) — spawns LocalShellTask with
  kind='monitor', returns immediately with task ID
- MonitorMcpTask (src/tasks/MonitorMcpTask/) — task lifecycle management
  and agent cleanup via killMonitorMcpTasksForAgent()
- MonitorPermissionRequest — permission dialog component

The codebase already had all integration points wired (tools.ts, tasks.ts,
PermissionRequest.tsx, LocalShellTask kind='monitor', BashTool prompt).
This PR provides the missing implementations.

* fix: command-specific permission rule + architecture docs

- MonitorPermissionRequest: "don't ask again" now creates a
  command-prefix rule (like BashTool) instead of a blanket
  tool-name-only rule that would auto-allow all Monitor commands
- MonitorMcpTask: clarify architecture comments explaining why
  monitor_mcp type exists as a registry stub while actual tasks
  are local_bash with kind='monitor'

* fix: address Copilot review feedback

- Fix permission rule field: expression → ruleContent (Copilot Twigpine#1)
- Handle empty command prefix: skip rule creation (Copilot Twigpine#2)
- Remove unused useTheme() import (Copilot Twigpine#3)
- Save permission rules under 'Bash' toolName so bashToolHasPermission
  can match them — Monitor delegates to Bash permission system (Copilot Twigpine#4)
- Remove unused logError import from MonitorMcpTask (Copilot Twigpine#6)
- Copilot Twigpine#5 (getAppState throws): same pattern as BashTool:915, not a bug
discopops pushed a commit to discopops/openclaude that referenced this pull request May 28, 2026
* feat: implement Monitor tool for streaming shell output

Add the Monitor tool that executes shell commands in the background and
streams stdout line-by-line as notifications to the model. This enables
real-time monitoring of logs, builds, and long-running processes.

Implementation:
- MonitorTool (src/tools/MonitorTool/) — spawns LocalShellTask with
  kind='monitor', returns immediately with task ID
- MonitorMcpTask (src/tasks/MonitorMcpTask/) — task lifecycle management
  and agent cleanup via killMonitorMcpTasksForAgent()
- MonitorPermissionRequest — permission dialog component

The codebase already had all integration points wired (tools.ts, tasks.ts,
PermissionRequest.tsx, LocalShellTask kind='monitor', BashTool prompt).
This PR provides the missing implementations.

* fix: command-specific permission rule + architecture docs

- MonitorPermissionRequest: "don't ask again" now creates a
  command-prefix rule (like BashTool) instead of a blanket
  tool-name-only rule that would auto-allow all Monitor commands
- MonitorMcpTask: clarify architecture comments explaining why
  monitor_mcp type exists as a registry stub while actual tasks
  are local_bash with kind='monitor'

* fix: address Copilot review feedback

- Fix permission rule field: expression → ruleContent (Copilot Twigpine#1)
- Handle empty command prefix: skip rule creation (Copilot Twigpine#2)
- Remove unused useTheme() import (Copilot Twigpine#3)
- Save permission rules under 'Bash' toolName so bashToolHasPermission
  can match them — Monitor delegates to Bash permission system (Copilot Twigpine#4)
- Remove unused logError import from MonitorMcpTask (Copilot Twigpine#6)
- Copilot Twigpine#5 (getAppState throws): same pattern as BashTool:915, not a bug
thedeveloloper pushed a commit to thedeveloloper/openclaude that referenced this pull request Jun 8, 2026
feat: intelligent auto-router — routes to fastest/cheapest provider automatically
hotmanxp referenced this pull request in hotmanxp/openclaude Jun 13, 2026
Per final-review follow-up #1 and #2:

1. docs/ports/bg-agent-view.md — summarizes the bg-agent-view port:
   - What was ported verbatim / adapted (16 components)
   - 7 deviations from the original plan with rationale
   - What was deliberately not ported (per plan §6)
   - 3 known wiring gaps (DAEMON/BG_SESSIONS feature flags, bg-agents
     argv wiring) with file:line references
   - Recommended T12 follow-up plan
   - How to test locally

2. docs/superpowers/plans/2026-06-13-plan-bg-agent-view.md — committed
   the plan doc itself so the audit trail is in git history. Previously
   left untracked in the worktree; reviewers found it hard to find.
hotmanxp referenced this pull request in hotmanxp/openclaude Jun 15, 2026
… (TDD test #2)

Per upstream claude-code 2.1.177, the LLM evaluator gets an immediate
"Condition: " prefix on the user message so it knows what to evaluate
without parsing context. Non-Stop prompt hooks (UserPromptSubmit etc.)
pass the prompt through unchanged to avoid behavior drift.

TDD: red in execPromptHook.goal.test.ts (asserts prefix present + before
condition text), green via conditional wrap on hookEvent==='Stop'.

Step 3 of 5 for /goal Stop-hook prompt port.

Refs: opencc-goal-prompt-comparison-audit-2026-06-15
hotmanxp referenced this pull request in hotmanxp/openclaude Jun 15, 2026
Pre-existing typecheck error at line 234 — gap #2 test (T3) was missing
the `(queryModelWithoutStreamingMock as any)` cast that all other
gap tests use. Runtime was fine (10 tests pass), but typecheck failed.

Confirmed pre-existing via `git stash` (introduced by T3, not by T4-T6).
Fix: 1-character `as any` cast on the mock call, consistent with
sibling tests (T2 line 198, T4 control line 393, T6 line 432).

Refs: opencc-goal-prompt-comparison-audit-2026-06-15
@euxaristia euxaristia mentioned this pull request Jun 25, 2026
8 of 9 tasks
kevincodex1 pushed a commit that referenced this pull request Jun 26, 2026
* feat(claude): add Opus 4.8 model support

Adds Claude Opus 4.8 alongside 4.7 in the model registry, picker,
pricing, integrations, and 1M-context support. Mirrors the established
4.7 pattern so longer suffixes resolve first in canonical-name matching.

- configs: CLAUDE_OPUS_4_8_CONFIG + opus48 registry entry
- model.ts: canonical resolver, default-model dispatch (1P -> 4.8,
  3P bumped to 4.7), display & marketing names
- modelOptions: getOpus48Option in PAYG 1P/3P, opusplan description
- modelCost: COST_TIER_5_25 pricing
- context: 1M-capable assertion
- prompts: FRONTIER_MODEL_NAME -> Opus 4.8
- integrations: hicap gateway, nearai brand/vendor/model entries
- claude brand catalog: new defineModel block

Tests: extends modelSupports1M coverage for 4.8.

* fix(models): gate Opus 4.8 out of PAYG 3P picker until rollout

Opus 4.8 was being added to the third-party (3P) model picker while
getDefaultOpusModel() keeps non-first-party usage on Opus 4.7. Remove the
3P option until 3P rollout is active; first-party picker is unaffected.

Addresses CodeRabbit review on #1769.

* test(integrations): cover NearAI anthropic/claude-opus-4-8 route

Adds focused regression coverage for the new Opus 4.8 provider/model
path: asserts the NearAI vendor catalog exposes the
anthropic/claude-opus-4-8 entry and that it resolves through its
modelDescriptorId to a registered NearAI model descriptor (vendor/brand
nearai, correct default model + label), plus the NearAI route base URL.
Addresses CodeRabbit's [Minor] request to test the exact route.

* fix(models): wire Opus 4.8 into adaptive thinking, 3P fallback, and knowledge cutoff

Addresses jatmn's review findings on #1769.

- [High] thinking.ts: add opus-4-8 to the adaptive-thinking allowlist.
  Without it, claude-opus-4-8 hit the generic opus exclusion and returned
  false, dropping first-party Opus 4.8 into budget-based thinking instead of
  thinking: { type: 'adaptive' }. Adds a regression test (provider mocked to
  a non-1P value so the allowlist is the only reason 4.8 returns true).
- [Medium] validateModel.ts: add an opus-4-8 -> opus47 entry to
  get3PFallbackSuggestion so an unavailable Opus 4.8 selection suggests 4.7.
- [Medium] prompts.ts: getKnowledgeCutoff now returns "January 2026" for
  claude-opus-4-8 and claude-opus-4-7 instead of falling through to the
  stale generic "January 2025".
- [Low] modelOptions.ts: update the PAYG 1P picker comment to include Opus 4.8.

The betas.ts structured-outputs / auto-mode allowlists are intentionally left
unchanged for this PR's scope (4.7 is also absent; auto mode is gated on PI
safety probes) — to be revisited with safety-research before enabling.

* fix(models): update remaining model-launch markers for Opus 4.8 default

Addresses jatmn's follow-up review on #1769 — markers missed when Opus 4.8
became the default.

- [P2] commitAttribution.ts: add explicit `opus-4-8` and `opus-4-7` branches to
  sanitizeModelName before the broad `opus-4` fallback, so commit/PR attribution
  shows the real model instead of `claude-opus-4`. Adds a focused regression
  test (commitAttribution.modelName.test.ts; mutation-checked).
- [P2] attribution.ts: update the unknown-first-party-model co-author fallback
  from 'Claude Opus 4.6' to 'Claude Opus 4.8', and the matching test expectation.
  Also fixed a sibling de-dup test that was passing only by coincidence (it hit
  the 4.6 fallback): point it at a model the public-name map actually recognizes
  (dot form) so it exercises the real prefix-dedup path.
- [P3] fastMode.ts: FAST_MODE_MODEL_DISPLAY 'Opus 4.6' -> 'Opus 4.8'.
- [P3] context.test.ts: update the stale modelSupports1M test title/comment from
  Opus 4.7 to 4.8 (the current first-party default).

* test(models): pin the claude-opus-4-7[1m] sanitizeModelName mapping too

CodeRabbit follow-up on #1769: the test covered the suffixed 4.8 path but not the
4.7 branch with the same [1m] session suffix. Add the claude-opus-4-7[1m] case so
both newly added mappings are pinned.

* fix(models): extend fast-mode + default-effort gates to the current default Opus

Addresses jatmn's review on #1769 — two predicates still gated to opus-4-6 only
while the default Opus is now 4.8.

- [P1] fastMode.ts: isFastModeSupportedByModel returned true only for opus-4-6,
  so for Max/Team Premium users on claude-opus-4-8 fast mode wouldn't actually
  enable even though FAST_MODE_MODEL_DISPLAY/the /fast command now say "Opus 4.8
  only". Extend the predicate to the fast-mode-capable Opus models (4.8/4.7/4.6).
- [P2] effort.ts: getDefaultEffortForModel applied the Pro/Max/Team `medium`
  default only for opus-4-6, so Pro/Max/Team sessions on the new default
  claude-opus-4-8 fell through to the generic effort path. Extend the branch to
  4.8/4.7/4.6 (per the @[MODEL LAUNCH] marker).

Adds regression tests for both (mutation-checked: reverting either predicate to
opus-4-6 only fails them).

* fix(models): wire Opus 4.8 into advisor, teammate fallback, skill vars, comments

Addresses jatmn's follow-up model-launch markers on #1769.

- [High] advisor.ts: modelSupportsAdvisor / isValidAdvisorModel only whitelisted
  opus-4-6 / sonnet-4-6, so first-party sessions on the new default
  claude-opus-4-8 reported the advisor tool unsupported. Add opus-4-8 and
  opus-4-7 to both (commands/advisor.ts and claude.ts use these centralized
  predicates, so they're covered). Adds a regression test (mutation-checked).
- [Medium] swarm/teammateModel.ts: getHardcodedTeammateModelFallback hardcoded
  CLAUDE_OPUS_4_6_CONFIG -> CLAUDE_OPUS_4_8_CONFIG, so new teammates spawn on the
  current default. Adds a first-party test case (mutation-checked).
- [Medium] skills/bundled/claudeApiContent.ts: SKILL_MODEL_VARS OPUS_ID/OPUS_NAME
  4.6 -> 4.8 (the bundled claude-api skill docs don't hardcode 4.6 elsewhere).
- [Low] effort.ts + figures.ts: refresh stale "max is Opus 4.6 only" comments to
  reflect the 4.8/4.7/4.6 runtime behavior.

* test(swarm): assert provider-aware teammate fallback for Bedrock too

CodeRabbit follow-up on #1769: add a non-first-party case so the provider-aware
fallback is covered. Bedrock resolves to the Opus 4.8 Bedrock model id.

* fix(models): give Opus 4.8/4.7 the elevated output-token limits and 3P fallback chain

Addresses jatmn's review on #1769.

- context.ts: getModelMaxOutputTokens only gave opus-4-6 the 64k/128k branch, so
  opus-4-7/4-8 fell through to the generic opus-4 branch and capped at 32k —
  including the new first-party default Opus 4.8. Extend the elevated branch to
  4.8/4.7/4.6. Adds a regression test (mutation-checked).
- errors.ts: get3PModelFallbackSuggestion had chains for opus-4-6/sonnet but not
  opus-4-8/4-7, so the error path suggested no fallback for the new default while
  validateModel.ts already does. Add opus-4-8 -> opus47 and opus-4-7 -> opus46 to
  mirror validateModel.ts.

* fix(models): allow structured outputs on Opus 4.8/4.7

Addresses jatmn's finding #2 on #1769. modelSupportsStructuredOutputs whitelisted
opus-4-1/4-5/4-6 but not 4-7/4-8, so first-party/Foundry requests on the new
default Opus 4.8 lost the structured-output support that 4.6 had. Add
claude-opus-4-7 and claude-opus-4-8 to the allowlist (4.6 supports it, so the
newer Opus models do too). Adds a first-party regression test (mutation-checked).

Auto-mode (modelSupportsExternalAutoMode) is intentionally left unchanged — it is
gated on separate safety review and was not part of this finding.

* fix(models): extend file-read mitigation exemption and effort callout to Opus 4.8/4.7

Addresses jatmn's remaining findings on #1769.

- [P2] FileReadTool.ts: MITIGATION_EXEMPT_MODELS only held claude-opus-4-6, so the
  new default claude-opus-4-8 got the cyber-risk reminder appended to every file
  read that 4.6 did not — a behavioral regression. Add claude-opus-4-8 and
  claude-opus-4-7 so the recent Opus models inherit 4.6's exemption.
- [P3] EffortCallout.tsx: shouldShowEffortCallout gated the medium-effort-default
  notification to opus-4-6 only; the same default now applies to opus-4-8, so
  users on the new default never saw it. Extend the gate to 4.8/4.7/4.6. Adds a
  regression test (mutation-checked).

* fix(models): resolve Opus 4.6→4.8 drift in cost tracking, notifications, and picker strings

Addresses jatmn's review on #1769 — remaining model-launch drift now that the
first-party default is Opus 4.8.

- [P1] modelCost.ts: getModelCosts only applied the elevated fast-mode tier to
  opus-4-6, so fast-mode Opus 4.8 was billed at the normal COST_TIER_5_25 rate
  while the picker advertised the fast-mode $30/$150 price. Extend the fast-mode
  cost check to the fast-mode-capable Opus models (4.8/4.7/4.6). Non-fast usage is
  unchanged (all three already map to COST_TIER_5_25). Adds a regression test
  (mutation-checked).
- [P2] useModelMigrationNotifications.tsx: "Model updated to Opus 4.6" -> 4.8
  (the migration lands users on the opus alias = 4.8 for first party).
- [P2] commands/model/model.tsx: the 1M-unavailable error said "Opus 4.6"; made
  it generic ("Opus with 1M context...") since the gate matches any opus[1m].
- [P2] modelOptions.ts: getOpus46_1MOption is now provider-aware (3P → Opus 4.6,
  first-party → Opus 4.8); getMaxOpus46_1MOption (always first-party) → Opus 4.8.
- [P3] migrateLegacyOpusToCurrent.ts: corrected the stale comment (opus alias
  resolves to 4.8, not 4.6).

* docs(notifs): correct Opus default comment to 4.8 for 1P

Comment said 4.6 but the migration notification text and the opus alias both resolve to Opus 4.8 for first-party users. Addresses jatmn P3 review note.

* fix(integrations): remove duplicate claude-opus-4-8 descriptor

A second claude-opus-4-8 entry (vendorId anthropic) with downgraded 200k/8192 specs duplicated the canonical 1M/128k descriptor. The artifact generator rejects duplicate (id, vendorId) pairs, so integrations:generate failed and smoke-and-tests could not pass. Removed the duplicate; the canonical entry and checked-in generated artifacts are unchanged. Addresses jatmn P1.

* fix(models): address Opus 4.8 review — extra-usage label, callout test, stale copy

- isBilledAsExtraUsage: recognize opus-4-7/4-8 1M variants, not just 4.6, so the
  "Billed as extra usage" label shows for the new default and 3P default
- EffortCallout modelGate test: drop the unreliable `?ts=` cache-busting import
  and rely on mock.module live bindings, so the gate runs against the mocked
  deps on Linux CI (where the query-tagged specifier was not re-evaluated)
- refresh stale "Opus 4.6+" effort help text and callout comments to reflect the
  recent Opus models (4.8/4.7/4.6) the gate now covers

* test(effort): make Opus 4.8 callout regression deterministic via pure predicate

The behavioral test mocked auth/config/effort and relied on the already-evaluated
EffortCallout picking up those mocks, which is order-dependent and failed only in
the full Linux CI suite (both the `?ts=` dynamic-import and the static-import
variants regressed there). Extract the model check as a pure exported
`effortCalloutCoversModel` and assert it directly with no module mocking, so the
#1769 regression is covered deterministically on every platform.

* test(effort): drop config-dependent 'opus' alias from callout regression test

The bare 'opus' alias routes through getDefaultOpusModel(), whose result is
environment/config-dependent, so `effortCalloutCoversModel('opus')` was false in a
clean Linux CI environment even though the gate logic is correct — that single
assertion was the only failure in smoke-and-tests (the explicit-id assertions
passed). Assert the gate's opus-4-8/4-7/4-6 coverage with explicit canonical model
ids (incl. a [1m] variant) instead, which is deterministic on every platform.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants