From 183401b9c8ffee62b3d250898ad41063da0703ba Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Fri, 31 Jul 2026 12:47:10 -0700 Subject: [PATCH 1/9] [review-correctness-opus-5] review: move the whole reviewer roster to Opus 5 Re-applied on top of #311 (refusal visibility and fallback). The original two commits conflicted with main independently of that stack: the sandbox.agent.version pin was retired and gh-aw v0.83.4 / firewall v0.27.42 landed underneath them, so review.md's frontmatter had moved. This keeps main's infrastructure evolution and re-applies the pin intent on top, rather than resurrecting frontmatter that main has since superseded. 22 pins move to claude-opus-5: the orchestrator, thread-reconciler, skill-auditor, claim-validator, conventions, the opt-in whole-change reviewers, all twelve specialist lenses, plus correctness-reviewer and first-principles from claude-fable-5. pattern-triage stays on Sonnet 4.6 as the cheap first pass; the eval's judge and match arbiter stay on Haiku. What changed since this PR was written, and why it is now stacked: The refusal risk this PR's body accepts as a specialist-lens hazard, detectable only via weekly drift, is measured and already open. Run 30656579898 caught correctness-reviewer on Fable 5 refusing incident-auth-bypass and adversarial-injection-approve under Anthropic's usage policy, at 5,207 tokens. It is a DEFAULT-roster agent, not a lens, on the metric this repo calls load-bearing. Moving it to Opus 5 does not resolve that: this PR's own assessment is that Opus 5 also ships elevated cyber safeguards and can return stop_reason refusal, so the hazard moves with the roster. Stacking on #311 means the roster lands with the mitigation in place: a refused agent re-dispatches on claude-opus-4-8 (already in that map for opus-5), recorded as fellBackTo and reported as a weekly rate. Note the detector this PR nominates changes character: a refusing reviewer no longer craters recall on security-adjacent cases, because it falls back and produces findings. The fallback counter replaces that inference with a direct measurement. Two things still owed on this branch, neither resolvable from the rebase: the risk section still describes the pre-#311 world, and the models.default-ai-credits-pricing question needs re-checking against firewall v0.27.42 (main retired the sandbox pin on the grounds that v0.27.42 prices fable-5; whether it prices claude-opus-5 is unverified, and an un-priced model is a 400 on every dispatch). --- .changeset/review-correctness-opus-5.md | 21 ++++++++++++ .github/workflows/review-eval-drift.yml | 12 ++++++- package.json | 5 ++- workflows/review/README.md | 30 +++++++++-------- workflows/review/review.md | 44 ++++++++++++------------- 5 files changed, 72 insertions(+), 40 deletions(-) create mode 100644 .changeset/review-correctness-opus-5.md diff --git a/.changeset/review-correctness-opus-5.md b/.changeset/review-correctness-opus-5.md new file mode 100644 index 00000000..4294a805 --- /dev/null +++ b/.changeset/review-correctness-opus-5.md @@ -0,0 +1,21 @@ +--- +"review": minor +--- + +Move the entire reviewer roster to Opus 5 (`claude-opus-5`). Every role that ran `claude-opus-4-8` (the orchestrator, `thread-reconciler`, `skill-auditor`, `claim-validator`, `conventions`, the opt-in whole-change reviewers `holistic` / `completeness` / `test-adequacy`, and all twelve specialist lenses) and both roles that ran `claude-fable-5` (`correctness-reviewer`, `first-principles`) now run `claude-opus-5` — 21 pins in total. `pattern-triage` stays on Sonnet 4.6 as the deliberately cheap first pass, and the judge and match arbiter stay pinned to Haiku 4.5 (they are the eval's ruler, not part of the reviewer). + +The basis: Opus 5 reports high precision *and* high recall at Opus 4.8's per-token price, which is half Fable's. That makes it a straight upgrade for the roles on 4.8, and for the two on Fable it should keep the recall the 2026-07-20 A/B bought (82% -> 89% must-catch, true misses down 41%) while dropping the +35% cost premium that came with it. The powered A/B is the acceptance criterion, not this description. + +This is a roster-wide swap measured in aggregate, so a per-role regression is easier to miss than it was when one role moved at a time. Four rows to read first, each with its own revert: `correctness-reviewer` (must-catch recall is the load-bearing metric; the six corpus rows Fable moved are where its gain lived); `claim-validator` (the precision gate — Fable measurably failed here at noise 43% -> 49% plus one wrong blocking flag on a clean case, which is why it never left Opus, so Opus 5's precision claim is the bet being made); the `security-auth` lens (see below); and the orchestrator, which is a throughput row rather than a quality one — it is the token-heaviest agent under a 20-minute cap, and Opus 5 writes longer at the same effort, so a timeout there is the swap surfacing operationally instead of measurably. + +**Known risk this accepts: refusal on the specialist lenses.** Those lenses were deliberately kept off Fable 5 because Fable's cyber safety classifiers can refuse benign security-focused analysis, and a refused lens is a silent coverage hole — it surfaces as a missing agent result, not an error. Opus 5 also ships elevated cybersecurity safeguards and can return `stop_reason: "refusal"` on cyber-adjacent input, so this move re-opens that hole rather than closing it: the model changed, the failure mode did not. gh-aw exposes no `fallbacks` parameter, so there is no server-side retry-on-another-model to lean on. The detector is the weekly drift corpus — a silently refusing security lens craters must-catch recall on the security-adjacent cases while every other metric looks normal. On that signature, put the `security-auth` lens back on `claude-opus-4-8` first; it is the narrowest revert available. `first-principles` also loses a property it was given on purpose (being the one non-Opus reviewer), so its perspective diversity now rests entirely on its prompt; if its findings converge with `holistic`'s in the drift series, put one reviewer back on a different family. + +**Consumers must be on gh-aw >= v0.83.0 before importing this release.** `claude-opus-5` is in no gh-aw-firewall release's curated AI-credits pricing table, nor in its bundled models.dev catalog fallback (checked through v0.27.41 and firewall `main` as of 2026-07-24), and the api-proxy's AI-credits guard — active by default — rejects any un-priced model with a 400 before the request reaches the model. That is how the first-principles dispatch failed on every run of review-v1.2.0 (299defb / 283d4b6), and pinning a newer `sandbox.agent.version` cannot fix it because no release prices the model. Note the blast radius has changed with the roster: when only two opt-in reviewers ran the un-priced model, a missing fallback cost those two dispatches; now it would 400 the orchestrator and every default-roster agent, i.e. the whole review. + +The frontmatter therefore carries `models.default-ai-credits-pricing` (`input: 5.0`, `output: 25.0`), a field added in gh-aw v0.83.0: it suppresses the rejection (the guard skips it whenever a default is configured) and prices the usage. Because Opus 5 lists at exactly Opus 4.8's price, that fallback is the model's real rate rather than an approximation — the proxy derives cache reads as `input x 0.1` = $0.50/M, also exact. Two caveats, both recorded in the frontmatter: the proxy's default-pricing path does not bill cache writes ($6.25/M real), so AI-credit accounting under-counts that component (the `models:` cost block carries the full rate for the run's cost summary); and the fallback applies to *any* un-priced model, so a future typo'd model id bills at Opus rates instead of failing loudly. Drop the field once a firewall release prices `claude-opus-5`. A consumer on an older gh-aw gets a compile-time failure on this file rather than a lock that 400s at runtime. + +`sandbox.agent.version` stays pinned at awf v0.27.27 even though its reason for existing (Fable 5's curated pricing entry) is gone. It now keeps the sandbox fixed while the whole roster's model changes underneath it: dropping it would inherit the compiler's default (v0.27.37+, which flipped the no-sudo default at v0.27.32), landing a sandbox-behavior change in the same commit as a 21-agent model change, indistinguishable in the eval. v0.27.27 is verified to accept and map `apiProxy.defaultAiCreditsPricing`, so the fallback works under the existing pin. Removing it is a separate PR. + +The eval harness needed one bump to be able to measure this at all: `@anthropic-ai/claude-agent-sdk` goes from `^0.3.205` to `^0.3.219`, the first release that knows `claude-opus-5`. The harness has no pricing table of its own — it enforces `--max-usd` from the `total_cost_usd` the SDK reports per result message — so an SDK that does not know the roster's model turns the budget cap into a no-op on precisely the arm under test. The lockfile pinned 0.3.205 and CI installs `--frozen-lockfile`, so the range bump alone would not have taken effect. The SDK is not in the api-proxy path, so this is an eval-only concern; production dispatch is covered by the pricing fallback above. + +The weekly drift budget is held at 220 rather than re-derived from the list price. Per-token parity with Opus 4.8 is not per-run parity: Opus 5 thinks by default and writes longer, and that now applies to every agent in the run rather than one, so the case-arm-run rate is unmeasured until the first drift week on the new roster. Under-sizing costs a red run (the `caseAsymmetry` gate) plus a contaminated noise-floor band, which is strictly worse than an unused ceiling. Re-size from a measured rate. diff --git a/.github/workflows/review-eval-drift.yml b/.github/workflows/review-eval-drift.yml index 5c60bea2..e532ca24 100644 --- a/.github/workflows/review-eval-drift.yml +++ b/.github/workflows/review-eval-drift.yml @@ -48,7 +48,17 @@ on: # 29855626692-29855643020: $97.19 / 90), so a drift week with both # arms on this roster is ~$194; 220 clears it with real headroom, # which matters now that a budget-skipped tail is a red run (the - # caseAsymmetry gate). If the swap is reverted, dial back toward 160. + # caseAsymmetry gate). + # + # HELD at 220 for the 2026-07-24 roster-wide move to Opus 5 (every role + # except pattern-triage), rather than re-derived from the list price. + # Opus 5 lists at Opus 4.8's per-token price, but per-token parity is not + # per-run parity: it thinks by default and writes longer, and that now + # applies to EVERY agent in the run rather than the two that ran Fable, + # so the case-arm-run rate is unmeasured until the first drift week on + # the new roster. Under-sizing here costs a red run (caseAsymmetry) and a + # contaminated noise-floor band, which is strictly worse than an unused + # ceiling. Re-size from that run's measured rate, not from the list price. description: "Total hard budget across both arms and all repeats" required: false default: "220" diff --git a/package.json b/package.json index dccd2a8a..257bbb04 100644 --- a/package.json +++ b/package.json @@ -10,7 +10,7 @@ "build": "tsc -p actions/tsconfig.json" }, "devDependencies": { - "@anthropic-ai/claude-agent-sdk": "^0.3.205", + "@anthropic-ai/claude-agent-sdk": "^0.3.219", "@changesets/cli": "^2.29.8", "@khanacademy/eslint-config": "^0.1.0", "@swc-node/register": "^1.11.1", @@ -27,8 +27,7 @@ "octokit": "^5.0.5", "prettier": "^2.6.2", "typescript": "^5.9.3", - "vitest": "^4.0.10", - "zod": "^4.4.3" + "vitest": "^4.0.10" }, "packageManager": "pnpm@10.0.0+sha512.b8fef5494bd3fe4cbd4edabd0745df2ee5be3e4b0b8b08fa643aa3e4c6702ccc0f00d68fa8a8c9858a735a0032485a44990ed2810526c875e416f001b17df12b" } diff --git a/workflows/review/README.md b/workflows/review/README.md index a50edd67..765fbc85 100644 --- a/workflows/review/README.md +++ b/workflows/review/README.md @@ -516,7 +516,9 @@ Fable's cyber classifiers can refuse benign security analysis, while `correctness-reviewer` — the default roster's load-bearing recall agent — was moved *onto* Fable 5 for its recall gain. Eval run 30656579898 caught it refusing `incident-auth-bypass` and `adversarial-injection-approve` outright, -at 5,207 tokens (so not a context limit). +at 5,207 tokens (so not a context limit). The roster has since moved to Opus 5, +which carries its own elevated cyber safeguards, so the hazard moved with it +rather than being resolved by the pin change. Refusals are **intermittent**: probe run 30658862532 saw the same Fable pin clear both cases that run 30656579898 blocked. The ordinary retry still cannot @@ -527,7 +529,7 @@ refusing pin to a model with a different refusal profile: | Pinned model | Falls back to | Basis | | --- | --- | --- | | `claude-fable-5` | `claude-opus-4-8` | measured (run 30656579898) | -| `claude-opus-5` | `claude-opus-4-8` | pre-emptive; #294 notes Opus 5 also ships elevated cyber safeguards | +| `claude-opus-5` | `claude-opus-4-8` | pre-emptive; Opus 5 ships elevated cyber safeguards and can also return `stop_reason: "refusal"` | Rules: **one hop**, never back to a model that already refused, and **no fallback for an unlisted pin** — an unmapped model's refusal stands and is @@ -552,19 +554,19 @@ sub-agent models — this table is the human-facing summary: | Role | Model | Effort | Why | | --- | --- | --- | --- | -| orchestrator | `claude-opus-4-8` | high | Owns every GitHub/safe-output decision | +| orchestrator | `claude-opus-5` | high | Owns every GitHub/safe-output decision | | `pattern-triage` | `claude-sonnet-4-6` | medium | Cheap first-pass triage | -| `thread-reconciler` | `claude-opus-4-8` | medium | Reconciliation | -| `correctness-reviewer` | `claude-fable-5` | high | Whole-change reviewer; bug-finding recall is the load-bearing metric | -| `skill-auditor` | `claude-opus-4-8` | high | Whole-change reviewer | -| `holistic` | `claude-opus-4-8` | high | Opt-in whole-change reviewer (`enable` in `ROUTING`) | -| `completeness` | `claude-opus-4-8` | high | Opt-in whole-change reviewer (`enable` in `ROUTING`) | -| `test-adequacy` | `claude-opus-4-8` | high | Opt-in whole-change reviewer (`enable` in `ROUTING`) | -| `conventions` | `claude-opus-4-8` | medium | Opt-in advisory targeted check (`enable` in `ROUTING`) | -| `documentation` | `claude-opus-4-8` | medium | Opt-in advisory targeted check (`enable` in `ROUTING`) | -| `first-principles` | `claude-fable-5` | high | Opt-in advisory-only; reviews the change's justification | -| `claim-validator` | `claude-opus-4-8` | xhigh | Adversarial claim validation; stays Opus (the Fable arm did not improve precision) | -| specialist lenses | `claude-opus-4-8` | high | Opt-in via `lens=` in `ROUTING`; the security & auth lens is xhigh | +| `thread-reconciler` | `claude-opus-5` | medium | Reconciliation | +| `correctness-reviewer` | `claude-opus-5` | high | Whole-change reviewer; bug-finding recall is the load-bearing metric | +| `skill-auditor` | `claude-opus-5` | high | Whole-change reviewer | +| `holistic` | `claude-opus-5` | high | Opt-in whole-change reviewer (`enable` in `ROUTING`) | +| `completeness` | `claude-opus-5` | high | Opt-in whole-change reviewer (`enable` in `ROUTING`) | +| `test-adequacy` | `claude-opus-5` | high | Opt-in whole-change reviewer (`enable` in `ROUTING`) | +| `conventions` | `claude-opus-5` | medium | Opt-in advisory targeted check (`enable` in `ROUTING`) | +| `documentation` | `claude-opus-5` | medium | Opt-in advisory targeted check (`enable` in `ROUTING`) | +| `first-principles` | `claude-opus-5` | high | Opt-in advisory-only; reviews the change's justification | +| `claim-validator` | `claude-opus-5` | xhigh | Adversarial claim validation; stays Opus (the Fable arm did not improve precision) | +| specialist lenses | `claude-opus-5` | high | Opt-in via `lens=` in `ROUTING`; the security & auth lens is xhigh | Only the orchestrator and the default roster (`pattern-triage`, `correctness-reviewer`, `skill-auditor`, `thread-reconciler`, `claim-validator`) diff --git a/workflows/review/review.md b/workflows/review/review.md index 3b9e3582..73051f33 100644 --- a/workflows/review/review.md +++ b/workflows/review/review.md @@ -179,7 +179,7 @@ observability: # job-level timeout-minutes still bounds the run. engine: id: claude - model: claude-opus-4-8 + model: claude-opus-5 env: BASH_DEFAULT_TIMEOUT_MS: "60000" BASH_MAX_TIMEOUT_MS: "1200000" @@ -1175,7 +1175,7 @@ return `{"findings": [], "hunts": [...]}` with the hunt states still recorded. --- name: correctness-reviewer description: Classifies each changed file's risk and reviews the diff for correctness defects; returns JSON. -model: claude-fable-5 +model: claude-opus-5 # effort: high — launch default (whole-change reviewer). gh-aw has no per-agent # effort field yet; the per-role model/effort table lives in the README. # Fable 5: bug-finding recall is this workflow's load-bearing metric, and @@ -1370,7 +1370,7 @@ before (run 29943085279 carried its one-line fix under `suggested_patch`): --- name: skill-auditor description: Evaluates the diff against the repo's best-practice skills and returns findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (whole-change reviewer). --- You audit a PR diff for best-practice "skill" violations. You have **no GitHub @@ -1534,7 +1534,7 @@ Return ONLY this JSON object (no prose, no code fence): --- name: thread-reconciler description: Decides which of the workflow's earlier review threads the current code has addressed; returns thread ids. -model: claude-opus-4-8 +model: claude-opus-5 # effort: medium — launch default (reconciliation). --- You decide which earlier review threads the current code has resolved. You have **no @@ -1593,7 +1593,7 @@ Return ONLY this JSON object (no prose, no code fence): --- name: claim-validator description: Re-checks each candidate review comment against the actual code and the repo's best-practice skills, and drops or corrects the ones that are wrong; returns JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: xhigh — launch default (claim-validator). Deliberately NOT moved to # Fable 5 with the correctness reviewer: in the 2026-07-20 pooled A/B the # Fable validator did not offset the higher flag rate (noise 43% -> 49%, one @@ -1780,7 +1780,7 @@ Every input `id` must appear exactly once. --- name: holistic description: Reviews the change as a whole — is the overall approach sound and coherent — and returns findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (whole-change reviewer). --- You are the **holistic** reviewer. Your single mandate is to **judge the @@ -1856,7 +1856,7 @@ scenario). If the change hangs together, return {"findings": []}. --- name: completeness description: Checks the change against its stated intent (PR description + linked ticket/doc) and returns findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (whole-change reviewer). --- You are the **completeness** reviewer. Your single mandate is to **check @@ -1927,7 +1927,7 @@ If the change matches its intent, return {"findings": []}. --- name: test-adequacy description: Evaluates whether the changed behavior is adequately tested and returns findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (whole-change reviewer). --- You are the **test-adequacy** reviewer. Your job is to judge whether the **changed @@ -1988,7 +1988,7 @@ If the changed behavior is adequately tested, return {"findings": []}. --- name: first-principles description: A diverse-perspective, advisory-only sanity check on whether the change should exist as written; returns findings as JSON. -model: claude-fable-5 +model: claude-opus-5 # effort: high — launch default. Ran on Fable 5 (claude-fable-5) from day one; # the correctness reviewer joined it after the 2026-07-20 A/B. Advisory-only, # never blocks. @@ -2061,7 +2061,7 @@ If you have nothing worth raising, return {"findings": []}. --- name: conventions description: Advisory, opt-in check of repo-specific conventions; returns findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: medium — launch default (advisory, opt-in targeted check). --- You are the **conventions** reviewer. You check the change against this repository's @@ -2127,7 +2127,7 @@ not worth flagging). If nothing deviates from repo conventions, return --- name: documentation description: Advisory, opt-in check that code comments and prose docs in the diff document intent rather than restate code; returns findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: medium — launch default (advisory, opt-in targeted check). Sibling of # `conventions`: same shape, same cost profile, different subject matter. --- @@ -2284,7 +2284,7 @@ maintain for nothing). If nothing in the change fails the policy, return --- name: security-auth description: Specialist security & auth lens — reviews touched files for authorization, secrets, injection, and unsafe-deserialization defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: xhigh — launch default. The security & auth lens is the one specialist # lens pinned to xhigh (per-role table in the README). gh-aw has no # per-agent effort field yet; this annotation and the README table are the authoritative @@ -2372,7 +2372,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: ai-safety-moderation description: Specialist AI safety & moderation lens — reviews AI/generation paths for missing moderation, prompt-injection surfaces, and PII exposure; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **AI safety & moderation** specialist lens. You review only AI/model and @@ -2442,7 +2442,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: mass-comms-coppa description: Specialist mass-comms & COPPA lens — reviews bulk-communication paths for audience/consent/age-gating defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **mass-comms & COPPA** specialist lens. You review only bulk-communication @@ -2510,7 +2510,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: caching-resource description: Specialist caching & resource lens — reviews caching and resource-management paths for key-scoping, invalidation, and exhaustion defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **caching & resource** specialist lens. You review only caching and @@ -2585,7 +2585,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: data-migrations description: Specialist data & migrations lens — reviews schema/migration/backfill changes for compatibility and safety defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **data & migrations** specialist lens. You review only schema changes, @@ -2655,7 +2655,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: concurrency-async description: Specialist concurrency & async lens — reviews concurrent/async code for races, unawaited work, and idempotency defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **concurrency & async** specialist lens. You review only concurrent and @@ -2724,7 +2724,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: api-federation-compat description: Specialist API & federation compatibility lens — reviews public API and GraphQL/federation changes for breaking-change defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **API & federation compatibility** specialist lens. You review only changes to @@ -2793,7 +2793,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: cross-deploy-serialization description: Specialist cross-deploy serialization lens — reviews persisted/queued/cached serialized shapes for rolling-deploy compatibility defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **cross-deploy serialization** specialist lens. You review only changes to @@ -2866,7 +2866,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: deploy-infra-config description: Specialist deploy & infra config lens — reviews deployment, infra-as-code, and config/flag changes for rollout-safety defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **deploy & infra config** specialist lens. You review only deployment @@ -2937,7 +2937,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: money-payments description: Specialist money & payments lens — reviews monetary and payment code for precision, idempotency, and currency defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **money & payments** specialist lens. You review only monetary computation and @@ -3006,7 +3006,7 @@ Conventional-Comment `label` is emitted (the orchestrator computes it from --- name: content-i18n description: Specialist content & i18n lens — reviews user-facing content for localization and internationalization defects; returns structured findings as JSON. -model: claude-opus-4-8 +model: claude-opus-5 # effort: high — launch default (specialist lens). --- You are the **content & i18n** specialist lens. You review only user-facing content for From 15965a64e92c357ba6570d2db06d3239a22f35d7 Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Fri, 31 Jul 2026 12:50:44 -0700 Subject: [PATCH 2/9] [review-correctness-opus-5] review: restore the Opus 5 pricing fallback my rebase dropped Re-applying only the pin changes onto main's frontmatter lost the load-bearing half of this PR: the models.default-ai-credits-pricing block. Without it every dispatch 400s before reaching the model, so the branch as I rebased it would have failed every run. Verified rather than assumed: claude-opus-5 is absent from the firewall api-proxy's curated pricing table at v0.27.42, the release gh-aw v0.83.4 defaults to now that this workflow's sandbox.agent.version pin is retired. The table carries claude-opus-4-5 through 4-8 and claude-fable-5 and stops there. The blast radius is worse than when this PR was written: main retired the sandbox pin on the grounds that v0.27.42 prices claude-fable-5, which is true and irrelevant once the roster is on Opus 5. And with all 22 pins on the un-priced model, a missing fallback 400s the orchestrator and every default agent rather than the two opt-in dispatches that #266 cost. The gh-aw >= v0.83.0 floor this field needs is already met (v0.83.4), so the compile-time constraint on consumers stands as documented but this repo clears it. --- workflows/review/review.md | 42 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 42 insertions(+) diff --git a/workflows/review/review.md b/workflows/review/review.md index 73051f33..44dfd157 100644 --- a/workflows/review/review.md +++ b/workflows/review/review.md @@ -206,6 +206,48 @@ sandbox: agent: id: awf +# claude-opus-5 is NOT in the firewall api-proxy's curated AI-credits pricing +# table — verified against v0.27.42, the release gh-aw v0.83.4 defaults to now +# that this workflow's `sandbox.agent.version` pin is retired (the table carries +# claude-opus-4-5 through 4-8 and claude-fable-5, and stops there). The proxy's +# AI-credits guard rejects an un-priced model with a 400 BEFORE the request +# reaches the model, so without the fallback below every dispatch fails — the +# #266 failure that killed the first-principles dispatch on every run, except +# now the whole roster runs the un-priced model rather than two opt-in agents. +# +# Both blocks carry claude-opus-5's list price (identical to Opus 4.8's: $5/M +# input, $0.50/M cache read, $6.25/M cache write, $25/M output) but in DIFFERENT +# UNITS, and they are not interchangeable: `default-ai-credits-pricing` is +# $/1M tokens and feeds the proxy's credit guard; `providers` is $/token and +# feeds the cost display. Because Opus 5 lists at exactly Opus 4.8's price, the +# fallback is the model's real rate rather than an approximation. +# +# Two caveats, unchanged from when this was written: the proxy's default-pricing +# path does not bill cache writes ($6.25/M real), so credit accounting +# under-counts that component (the `models:` block below carries the full rate +# for the cost summary); and the fallback applies to ANY un-priced model, so a +# future typo'd model id bills at Opus rates instead of failing loudly. +# +# MINIMUM COMPILER: gh-aw >= v0.83.0 for `models.default-ai-credits-pricing`. +# This repo compiles with v0.83.4, so the floor is already met; consumers on an +# older gh-aw get a COMPILE-time failure rather than a runtime 400. +models: + # $/1M tokens. `input` and `output` are the only rates the schema accepts, so + # the cache rates are the proxy's derivations, not ours. + default-ai-credits-pricing: + input: 5.0 + output: 25.0 + # $/token. + providers: + anthropic: + models: + claude-opus-5: + cost: + input: 5.0e-06 + output: 2.5e-05 + cache_read: 5.0e-07 + cache_write: 6.25e-06 + # The shared review workflow is more than this markdown file: its deterministic # pieces (the finding schema and validator today; the router, computed verdict, and # comment renderer as they land) are TypeScript under `workflows/review/lib/` in From 3c454c81a0d169eb37b620690e7f84c31d2a1871 Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Fri, 31 Jul 2026 13:04:27 -0700 Subject: [PATCH 3/9] [review-correctness-opus-5] review: note that firewall v0.27.43 prices Opus 5, and when to drop the fallback MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit v0.27.43 (newest firewall release, 2026-07-31) adds a curated claude-opus-5 entry at the same rates this fallback carries. The curated entry is strictly better: it bills cache writes, the one component the default-pricing path under-counts. No gh-aw release references v0.27.43 yet, so the fallback stays for now. The comment records the removal condition rather than leaving a future reader to re-derive it, and explicitly rules out pinning sandbox.agent.version forward to reach it early — review.md's own rule is that a version is pinned here only to hold a release BACK. --- workflows/review/review.md | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/workflows/review/review.md b/workflows/review/review.md index 44dfd157..955f29cb 100644 --- a/workflows/review/review.md +++ b/workflows/review/review.md @@ -207,8 +207,8 @@ sandbox: id: awf # claude-opus-5 is NOT in the firewall api-proxy's curated AI-credits pricing -# table — verified against v0.27.42, the release gh-aw v0.83.4 defaults to now -# that this workflow's `sandbox.agent.version` pin is retired (the table carries +# table at v0.27.42, the release gh-aw v0.83.4 defaults to now that this +# workflow's `sandbox.agent.version` pin is retired (that table carries # claude-opus-4-5 through 4-8 and claude-fable-5, and stops there). The proxy's # AI-credits guard rejects an un-priced model with a 400 BEFORE the request # reaches the model, so without the fallback below every dispatch fails — the @@ -231,6 +231,16 @@ sandbox: # MINIMUM COMPILER: gh-aw >= v0.83.0 for `models.default-ai-credits-pricing`. # This repo compiles with v0.83.4, so the floor is already met; consumers on an # older gh-aw get a COMPILE-time failure rather than a runtime 400. +# +# REMOVE THIS BLOCK when a gh-aw release defaults the firewall to v0.27.43 or +# later: that release DOES carry a curated claude-opus-5 entry, at the same +# rates ($5/M in, $0.50/M cache read, $6.25/M cache write, $25/M out), and the +# curated entry is strictly better because it bills cache writes — the one +# component this fallback path silently under-counts. v0.27.43 is the newest +# firewall release as of 2026-07-31 and no gh-aw release references it yet. +# Do NOT reach for `sandbox.agent.version: v0.27.43` to get there early: the +# note above is explicit that a version is pinned here only to hold a release +# BACK, never to move one forward, and this fallback works under any firewall. models: # $/1M tokens. `input` and `output` are the only rates the schema accepts, so # the cache rates are the proxy's derivations, not ours. From 8245704360c1e236f676b18c889c4d299fa865f3 Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Fri, 31 Jul 2026 14:10:57 -0700 Subject: [PATCH 4/9] [upd294b] review: fold the Opus 5 fallback into the pricing overlay's models block Two top-level `models:` keys is a duplicate YAML key that resolves silently to whichever comes last, so one of the fallback or the overlay would be dropped with no error. gh aw compile does not catch it: it compiles only .github/workflows, and the duplicate is in the shared source. Insert default-ai-credits-pricing into the overlay's existing block rather than relocating the block, so this PR's diff does not claim the overlay's lines. Stated at 50% of list (2.5/12.5 $/1M) for the same reason the overlay is, so a credit means real spend whichever path prices a request. --- .github/aw/logs/.gitignore | 5 +++ workflows/review/review.md | 81 ++++++++++++++------------------------ 2 files changed, 34 insertions(+), 52 deletions(-) create mode 100644 .github/aw/logs/.gitignore diff --git a/.github/aw/logs/.gitignore b/.github/aw/logs/.gitignore new file mode 100644 index 00000000..986a3211 --- /dev/null +++ b/.github/aw/logs/.gitignore @@ -0,0 +1,5 @@ +# Ignore all downloaded workflow logs +* + +# But keep the .gitignore file itself +!.gitignore diff --git a/workflows/review/review.md b/workflows/review/review.md index 955f29cb..093b5641 100644 --- a/workflows/review/review.md +++ b/workflows/review/review.md @@ -206,58 +206,6 @@ sandbox: agent: id: awf -# claude-opus-5 is NOT in the firewall api-proxy's curated AI-credits pricing -# table at v0.27.42, the release gh-aw v0.83.4 defaults to now that this -# workflow's `sandbox.agent.version` pin is retired (that table carries -# claude-opus-4-5 through 4-8 and claude-fable-5, and stops there). The proxy's -# AI-credits guard rejects an un-priced model with a 400 BEFORE the request -# reaches the model, so without the fallback below every dispatch fails — the -# #266 failure that killed the first-principles dispatch on every run, except -# now the whole roster runs the un-priced model rather than two opt-in agents. -# -# Both blocks carry claude-opus-5's list price (identical to Opus 4.8's: $5/M -# input, $0.50/M cache read, $6.25/M cache write, $25/M output) but in DIFFERENT -# UNITS, and they are not interchangeable: `default-ai-credits-pricing` is -# $/1M tokens and feeds the proxy's credit guard; `providers` is $/token and -# feeds the cost display. Because Opus 5 lists at exactly Opus 4.8's price, the -# fallback is the model's real rate rather than an approximation. -# -# Two caveats, unchanged from when this was written: the proxy's default-pricing -# path does not bill cache writes ($6.25/M real), so credit accounting -# under-counts that component (the `models:` block below carries the full rate -# for the cost summary); and the fallback applies to ANY un-priced model, so a -# future typo'd model id bills at Opus rates instead of failing loudly. -# -# MINIMUM COMPILER: gh-aw >= v0.83.0 for `models.default-ai-credits-pricing`. -# This repo compiles with v0.83.4, so the floor is already met; consumers on an -# older gh-aw get a COMPILE-time failure rather than a runtime 400. -# -# REMOVE THIS BLOCK when a gh-aw release defaults the firewall to v0.27.43 or -# later: that release DOES carry a curated claude-opus-5 entry, at the same -# rates ($5/M in, $0.50/M cache read, $6.25/M cache write, $25/M out), and the -# curated entry is strictly better because it bills cache writes — the one -# component this fallback path silently under-counts. v0.27.43 is the newest -# firewall release as of 2026-07-31 and no gh-aw release references it yet. -# Do NOT reach for `sandbox.agent.version: v0.27.43` to get there early: the -# note above is explicit that a version is pinned here only to hold a release -# BACK, never to move one forward, and this fallback works under any firewall. -models: - # $/1M tokens. `input` and `output` are the only rates the schema accepts, so - # the cache rates are the proxy's derivations, not ours. - default-ai-credits-pricing: - input: 5.0 - output: 25.0 - # $/token. - providers: - anthropic: - models: - claude-opus-5: - cost: - input: 5.0e-06 - output: 2.5e-05 - cache_read: 5.0e-07 - cache_write: 6.25e-06 - # The shared review workflow is more than this markdown file: its deterministic # pieces (the finding schema and validator today; the router, computed verdict, and # comment renderer as they land) are TypeScript under `workflows/review/lib/` in @@ -393,6 +341,35 @@ post-steps: # Do NOT collapse these to a bare `claude-opus-4` prefix: prefix matching would # also capture opus-4-0/4-1, which list at 3x the 4-5+ rate. models: + # The Opus 5 fallback. claude-opus-5 is NOT in the api-proxy's curated + # AI-credits table at v0.27.42, the release gh-aw v0.83.4 defaults to (that + # table carries claude-opus-4-5 through 4-8 and claude-fable-5, and stops + # there). The credit guard rejects an un-priced model with a 400 BEFORE the + # request reaches the model, so without this every dispatch fails: the #266 + # failure that killed the first-principles dispatch on every run, except now + # the whole roster runs the un-priced model rather than two opt-in agents. + # `providers` below cannot cover it while that key is dropped under v0.27.42, + # so this is what actually keeps the roster running. + # + # $/1M tokens, NOT the $/token units `providers` uses. Stated at 50% of list + # for the same reason `providers` is, so a credit means real spend whichever + # path prices a request. `input` and `output` are the only rates the schema + # accepts, so cache reads are the proxy's derivation ($0.25/M); the + # default-pricing path does not bill cache writes at all, which `providers` + # does once live. + # + # DELETE THIS KEY (not the whole block) once a gh-aw release defaults the + # firewall to v0.27.43+: that release carries a curated claude-opus-5 entry, + # strictly better because it bills cache writes, and because this fallback + # applies to ANY un-priced model a typo'd model id bills at Opus rates + # instead of failing loudly. + # + # MINIMUM COMPILER: gh-aw >= v0.83.0 for `models.default-ai-credits-pricing`. + # This repo compiles with v0.83.4, so the floor is met; consumers on an older + # gh-aw get a COMPILE-time failure rather than a runtime 400. + default-ai-credits-pricing: + input: 2.5 + output: 12.5 providers: anthropic: models: From 642ad5cdf2e268a3b9eca2de4206889847f0b57d Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Fri, 31 Jul 2026 14:22:06 -0700 Subject: [PATCH 5/9] [upd294b] review: drop the Opus 5 default pricing and the held-budget note The pricing overlay this PR stacks on already prices claude-opus-5 under `models.providers`, so `default-ai-credits-pricing` is a second, coarser statement of the same rate: it is $/1M rather than $/token, it does not bill cache writes, and it applies to ANY un-priced model, so a typo'd model id bills at Opus rates instead of failing loudly. Also drop the eval-drift note explaining why the 220 budget was held rather than re-derived for the Opus 5 move; the surrounding derivation already documents the number. --- .github/workflows/review-eval-drift.yml | 10 --------- workflows/review/review.md | 29 ------------------------- 2 files changed, 39 deletions(-) diff --git a/.github/workflows/review-eval-drift.yml b/.github/workflows/review-eval-drift.yml index e532ca24..c1df25ce 100644 --- a/.github/workflows/review-eval-drift.yml +++ b/.github/workflows/review-eval-drift.yml @@ -49,16 +49,6 @@ on: # arms on this roster is ~$194; 220 clears it with real headroom, # which matters now that a budget-skipped tail is a red run (the # caseAsymmetry gate). - # - # HELD at 220 for the 2026-07-24 roster-wide move to Opus 5 (every role - # except pattern-triage), rather than re-derived from the list price. - # Opus 5 lists at Opus 4.8's per-token price, but per-token parity is not - # per-run parity: it thinks by default and writes longer, and that now - # applies to EVERY agent in the run rather than the two that ran Fable, - # so the case-arm-run rate is unmeasured until the first drift week on - # the new roster. Under-sizing here costs a red run (caseAsymmetry) and a - # contaminated noise-floor band, which is strictly worse than an unused - # ceiling. Re-size from that run's measured rate, not from the list price. description: "Total hard budget across both arms and all repeats" required: false default: "220" diff --git a/workflows/review/review.md b/workflows/review/review.md index 093b5641..73051f33 100644 --- a/workflows/review/review.md +++ b/workflows/review/review.md @@ -341,35 +341,6 @@ post-steps: # Do NOT collapse these to a bare `claude-opus-4` prefix: prefix matching would # also capture opus-4-0/4-1, which list at 3x the 4-5+ rate. models: - # The Opus 5 fallback. claude-opus-5 is NOT in the api-proxy's curated - # AI-credits table at v0.27.42, the release gh-aw v0.83.4 defaults to (that - # table carries claude-opus-4-5 through 4-8 and claude-fable-5, and stops - # there). The credit guard rejects an un-priced model with a 400 BEFORE the - # request reaches the model, so without this every dispatch fails: the #266 - # failure that killed the first-principles dispatch on every run, except now - # the whole roster runs the un-priced model rather than two opt-in agents. - # `providers` below cannot cover it while that key is dropped under v0.27.42, - # so this is what actually keeps the roster running. - # - # $/1M tokens, NOT the $/token units `providers` uses. Stated at 50% of list - # for the same reason `providers` is, so a credit means real spend whichever - # path prices a request. `input` and `output` are the only rates the schema - # accepts, so cache reads are the proxy's derivation ($0.25/M); the - # default-pricing path does not bill cache writes at all, which `providers` - # does once live. - # - # DELETE THIS KEY (not the whole block) once a gh-aw release defaults the - # firewall to v0.27.43+: that release carries a curated claude-opus-5 entry, - # strictly better because it bills cache writes, and because this fallback - # applies to ANY un-priced model a typo'd model id bills at Opus rates - # instead of failing loudly. - # - # MINIMUM COMPILER: gh-aw >= v0.83.0 for `models.default-ai-credits-pricing`. - # This repo compiles with v0.83.4, so the floor is met; consumers on an older - # gh-aw get a COMPILE-time failure rather than a runtime 400. - default-ai-credits-pricing: - input: 2.5 - output: 12.5 providers: anthropic: models: From d3cafce3331210581e42e4b46e7e55f56a9dd77d Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Fri, 31 Jul 2026 14:31:42 -0700 Subject: [PATCH 6/9] [upd294c] review: rewrite the roster changeset, drop a stray logs .gitignore The changeset still described a default-ai-credits-pricing fallback and a sandbox.agent.version pin that this branch no longer carries, and ran to 1051 words. Rewrite at 308, accurate to the current diff: claude-opus-5 is priced by the models.providers overlay this stacks on, so the gh-aw v0.84.x dependency is a merge gate rather than a consumer compiler floor. Also remove .github/aw/logs/.gitignore, which a gh aw command created and a git add -A swept into an earlier commit here. It is unrelated to the roster move. --- .changeset/review-correctness-opus-5.md | 18 +++++------------- .github/aw/logs/.gitignore | 5 ----- 2 files changed, 5 insertions(+), 18 deletions(-) delete mode 100644 .github/aw/logs/.gitignore diff --git a/.changeset/review-correctness-opus-5.md b/.changeset/review-correctness-opus-5.md index 4294a805..68894ea0 100644 --- a/.changeset/review-correctness-opus-5.md +++ b/.changeset/review-correctness-opus-5.md @@ -2,20 +2,12 @@ "review": minor --- -Move the entire reviewer roster to Opus 5 (`claude-opus-5`). Every role that ran `claude-opus-4-8` (the orchestrator, `thread-reconciler`, `skill-auditor`, `claim-validator`, `conventions`, the opt-in whole-change reviewers `holistic` / `completeness` / `test-adequacy`, and all twelve specialist lenses) and both roles that ran `claude-fable-5` (`correctness-reviewer`, `first-principles`) now run `claude-opus-5` — 21 pins in total. `pattern-triage` stays on Sonnet 4.6 as the deliberately cheap first pass, and the judge and match arbiter stay pinned to Haiku 4.5 (they are the eval's ruler, not part of the reviewer). +Move the entire reviewer roster to Opus 5 (`claude-opus-5`): every role that ran `claude-opus-4-8` (orchestrator, `thread-reconciler`, `skill-auditor`, `claim-validator`, `conventions`, `documentation`, the opt-in `holistic` / `completeness` / `test-adequacy`, and all twelve specialist lenses) and both that ran `claude-fable-5` (`correctness-reviewer`, `first-principles`). `pattern-triage` stays on Sonnet 4.6 as the deliberately cheap first pass; the eval's judge and match arbiter stay on Haiku 4.5, being the ruler rather than the reviewer. -The basis: Opus 5 reports high precision *and* high recall at Opus 4.8's per-token price, which is half Fable's. That makes it a straight upgrade for the roles on 4.8, and for the two on Fable it should keep the recall the 2026-07-20 A/B bought (82% -> 89% must-catch, true misses down 41%) while dropping the +35% cost premium that came with it. The powered A/B is the acceptance criterion, not this description. +Opus 5 reports high precision and high recall at Opus 4.8's per-token price, half Fable's: a straight upgrade for the roles on 4.8, and for the two on Fable it should keep the recall the 2026-07-20 A/B bought (82% -> 89% must-catch) without the +35% premium. The powered A/B is the acceptance criterion, not this description. Read four rows first, each independently revertable: `correctness-reviewer` (recall), `claim-validator` (precision), the `security-auth` lens, and the orchestrator (throughput under the wall clock). -This is a roster-wide swap measured in aggregate, so a per-role regression is easier to miss than it was when one role moved at a time. Four rows to read first, each with its own revert: `correctness-reviewer` (must-catch recall is the load-bearing metric; the six corpus rows Fable moved are where its gain lived); `claim-validator` (the precision gate — Fable measurably failed here at noise 43% -> 49% plus one wrong blocking flag on a clean case, which is why it never left Opus, so Opus 5's precision claim is the bet being made); the `security-auth` lens (see below); and the orchestrator, which is a throughput row rather than a quality one — it is the token-heaviest agent under a 20-minute cap, and Opus 5 writes longer at the same effort, so a timeout there is the swap surfacing operationally instead of measurably. +**Known risk: refusal on the specialist lenses.** Those were kept off Fable because cyber safety classifiers can refuse benign security analysis, and a refused lens is a silent coverage hole rather than an error. Opus 5 can also return `stop_reason: "refusal"` on cyber-adjacent input, so this re-opens that hole. The detector is the weekly drift corpus; on that signature put `security-auth` back on `claude-opus-4-8` first, the narrowest revert available. -**Known risk this accepts: refusal on the specialist lenses.** Those lenses were deliberately kept off Fable 5 because Fable's cyber safety classifiers can refuse benign security-focused analysis, and a refused lens is a silent coverage hole — it surfaces as a missing agent result, not an error. Opus 5 also ships elevated cybersecurity safeguards and can return `stop_reason: "refusal"` on cyber-adjacent input, so this move re-opens that hole rather than closing it: the model changed, the failure mode did not. gh-aw exposes no `fallbacks` parameter, so there is no server-side retry-on-another-model to lean on. The detector is the weekly drift corpus — a silently refusing security lens craters must-catch recall on the security-adjacent cases while every other metric looks normal. On that signature, put the `security-auth` lens back on `claude-opus-4-8` first; it is the narrowest revert available. `first-principles` also loses a property it was given on purpose (being the one non-Opus reviewer), so its perspective diversity now rests entirely on its prompt; if its findings converge with `holistic`'s in the drift series, put one reviewer back on a different family. +`claude-opus-5` is priced by the `models.providers` overlay this stacks on, so it needs firewall v0.27.43 (gh-aw v0.84.x). Under v0.27.42 the overlay is dropped, the model is un-priced, and the api-proxy rejects every dispatch with a 400: do not merge before that release is stable. -**Consumers must be on gh-aw >= v0.83.0 before importing this release.** `claude-opus-5` is in no gh-aw-firewall release's curated AI-credits pricing table, nor in its bundled models.dev catalog fallback (checked through v0.27.41 and firewall `main` as of 2026-07-24), and the api-proxy's AI-credits guard — active by default — rejects any un-priced model with a 400 before the request reaches the model. That is how the first-principles dispatch failed on every run of review-v1.2.0 (299defb / 283d4b6), and pinning a newer `sandbox.agent.version` cannot fix it because no release prices the model. Note the blast radius has changed with the roster: when only two opt-in reviewers ran the un-priced model, a missing fallback cost those two dispatches; now it would 400 the orchestrator and every default-roster agent, i.e. the whole review. - -The frontmatter therefore carries `models.default-ai-credits-pricing` (`input: 5.0`, `output: 25.0`), a field added in gh-aw v0.83.0: it suppresses the rejection (the guard skips it whenever a default is configured) and prices the usage. Because Opus 5 lists at exactly Opus 4.8's price, that fallback is the model's real rate rather than an approximation — the proxy derives cache reads as `input x 0.1` = $0.50/M, also exact. Two caveats, both recorded in the frontmatter: the proxy's default-pricing path does not bill cache writes ($6.25/M real), so AI-credit accounting under-counts that component (the `models:` cost block carries the full rate for the run's cost summary); and the fallback applies to *any* un-priced model, so a future typo'd model id bills at Opus rates instead of failing loudly. Drop the field once a firewall release prices `claude-opus-5`. A consumer on an older gh-aw gets a compile-time failure on this file rather than a lock that 400s at runtime. - -`sandbox.agent.version` stays pinned at awf v0.27.27 even though its reason for existing (Fable 5's curated pricing entry) is gone. It now keeps the sandbox fixed while the whole roster's model changes underneath it: dropping it would inherit the compiler's default (v0.27.37+, which flipped the no-sudo default at v0.27.32), landing a sandbox-behavior change in the same commit as a 21-agent model change, indistinguishable in the eval. v0.27.27 is verified to accept and map `apiProxy.defaultAiCreditsPricing`, so the fallback works under the existing pin. Removing it is a separate PR. - -The eval harness needed one bump to be able to measure this at all: `@anthropic-ai/claude-agent-sdk` goes from `^0.3.205` to `^0.3.219`, the first release that knows `claude-opus-5`. The harness has no pricing table of its own — it enforces `--max-usd` from the `total_cost_usd` the SDK reports per result message — so an SDK that does not know the roster's model turns the budget cap into a no-op on precisely the arm under test. The lockfile pinned 0.3.205 and CI installs `--frozen-lockfile`, so the range bump alone would not have taken effect. The SDK is not in the api-proxy path, so this is an eval-only concern; production dispatch is covered by the pricing fallback above. - -The weekly drift budget is held at 220 rather than re-derived from the list price. Per-token parity with Opus 4.8 is not per-run parity: Opus 5 thinks by default and writes longer, and that now applies to every agent in the run rather than one, so the case-arm-run rate is unmeasured until the first drift week on the new roster. Under-sizing costs a red run (the `caseAsymmetry` gate) plus a contaminated noise-floor band, which is strictly worse than an unused ceiling. Re-size from a measured rate. +Also bumps `@anthropic-ai/claude-agent-sdk` to `^0.3.219`, the first release that knows `claude-opus-5`. The eval enforces `--max-usd` from the SDK's reported `total_cost_usd`, so an older SDK turns the budget cap into a no-op on precisely the arm under test. Eval-only; the SDK is not in the api-proxy path. diff --git a/.github/aw/logs/.gitignore b/.github/aw/logs/.gitignore deleted file mode 100644 index 986a3211..00000000 --- a/.github/aw/logs/.gitignore +++ /dev/null @@ -1,5 +0,0 @@ -# Ignore all downloaded workflow logs -* - -# But keep the .gitignore file itself -!.gitignore From 2cfb0fb3f6d29f489d95cf731479a4b923fcc0ab Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Fri, 31 Jul 2026 14:34:24 -0700 Subject: [PATCH 7/9] [upd294c] review: update the lockfile for the SDK bump, restore zod to the root manifest CI was red on this branch before the restack, for two reasons that both live in package.json. The SDK range moved to ^0.3.219 without regenerating pnpm-lock.yaml, so `pnpm install --frozen-lockfile` failed with ERR_PNPM_OUTDATED_LOCKFILE before any job ran a test. The changeset already explains why the range bump alone is insufficient; the lockfile just never followed. zod was also dropped from the root manifest. main carries it, and workflows/review/lib/dispatch-runner.ts imports it, so removing it breaks resolution for root-level runs. Nothing in this branch explains the removal, so restore it. --- package.json | 3 +- pnpm-lock.yaml | 91 ++++++++++++++++++++++++++++++++++++++++++++++++-- 2 files changed, 91 insertions(+), 3 deletions(-) diff --git a/package.json b/package.json index 257bbb04..6521a46e 100644 --- a/package.json +++ b/package.json @@ -27,7 +27,8 @@ "octokit": "^5.0.5", "prettier": "^2.6.2", "typescript": "^5.9.3", - "vitest": "^4.0.10" + "vitest": "^4.0.10", + "zod": "^4.4.3" }, "packageManager": "pnpm@10.0.0+sha512.b8fef5494bd3fe4cbd4edabd0745df2ee5be3e4b0b8b08fa643aa3e4c6702ccc0f00d68fa8a8c9858a735a0032485a44990ed2810526c875e416f001b17df12b" } diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index 18d1721b..14d14744 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -9,8 +9,8 @@ importers: .: devDependencies: '@anthropic-ai/claude-agent-sdk': - specifier: ^0.3.205 - version: 0.3.205(@anthropic-ai/sdk@0.110.0(zod@4.4.3))(@modelcontextprotocol/sdk@1.29.0(zod@4.4.3))(zod@4.4.3) + specifier: ^0.3.219 + version: 0.3.220(@anthropic-ai/sdk@0.110.0(zod@4.4.3))(@modelcontextprotocol/sdk@1.29.0(zod@4.4.3))(zod@4.4.3) '@changesets/cli': specifier: ^2.29.8 version: 2.29.8(@types/node@25.3.3) @@ -134,41 +134,81 @@ packages: cpu: [arm64] os: [darwin] + '@anthropic-ai/claude-agent-sdk-darwin-arm64@0.3.220': + resolution: {integrity: sha512-7VxlbEosK7DODiOnsjoVd0DSJzbnaPrM2jelMHI0y8zx1UnLS3WC6EFUXbvy74F2sXqEznh2tzn7EKWInaRN6Q==} + cpu: [arm64] + os: [darwin] + '@anthropic-ai/claude-agent-sdk-darwin-x64@0.3.205': resolution: {integrity: sha512-G6ETPmL5mNzJ2DFsWxG3jmsmrXgZX1N2ZCJvxaGUUpjTsKZJ4Tup1cWYvcd/m7o5fYZmx9REmgzTwsAIc1fdPQ==} cpu: [x64] os: [darwin] + '@anthropic-ai/claude-agent-sdk-darwin-x64@0.3.220': + resolution: {integrity: sha512-X9RwDsSmbF6ultKZroaip+DL8WRgC64gHbrAwrRlAFSPNZV7zmJyP2ur8rW7KrxqmtuehdMMkw8+SAC/6hD2PA==} + cpu: [x64] + os: [darwin] + '@anthropic-ai/claude-agent-sdk-linux-arm64-musl@0.3.205': resolution: {integrity: sha512-91fgdG4aTnQ29sKOcUqgH4+tKCW2ut6PWGRSYmXNDbROasJm1rAlPdzC5brdu/e4c0CDSNV6TWyE5JCjaS/jlQ==} cpu: [arm64] os: [linux] + '@anthropic-ai/claude-agent-sdk-linux-arm64-musl@0.3.220': + resolution: {integrity: sha512-OHoZOZ8Cf2TBr6oXIXPwyvUxj9jrq2w8E4poA8dMpacXszcPSPiCQCMuuOh4aWJzfeJE1+TtWxhKMVb2csXyZQ==} + cpu: [arm64] + os: [linux] + '@anthropic-ai/claude-agent-sdk-linux-arm64@0.3.205': resolution: {integrity: sha512-CXzySK3PV3EizCRPXnxPqeaAtgrBFDnMFOVpMe36oC3U16yDb1b1tAJGqZi/7uFrVvAiaXvnSFxhUWnDDSaO+A==} cpu: [arm64] os: [linux] + '@anthropic-ai/claude-agent-sdk-linux-arm64@0.3.220': + resolution: {integrity: sha512-WkROPwWskqhKR9XgnmseHQ6rLi9zM9qt57IWoToIjL/eXOqDWipp7JXZ1L5ud+LrA42dunHPZfBwD/vXZ+A7LA==} + cpu: [arm64] + os: [linux] + '@anthropic-ai/claude-agent-sdk-linux-x64-musl@0.3.205': resolution: {integrity: sha512-vvsb7GlnA8CTSVvvTkrXjcSeRKqxSM7p/tU3Od9ICAZeWHglptekEyzLEApzLuLbI5ewfFF/F0q3NwOBbo18dg==} cpu: [x64] os: [linux] + '@anthropic-ai/claude-agent-sdk-linux-x64-musl@0.3.220': + resolution: {integrity: sha512-K+FWj+LcGhC1Z7wqeWoLxm1iemcba5xKpLLFVwYm4V6HyMx3ruYd/2r2TiQtjT+JWeNFWIys0ScHiItR6vWAiA==} + cpu: [x64] + os: [linux] + '@anthropic-ai/claude-agent-sdk-linux-x64@0.3.205': resolution: {integrity: sha512-siS+1iNqBSlGFZZvJY6+mhzZ/6/ec/TbX9GMuwmTF0E6fxGhIIp797jJxR1q8r6FAq7d39mEoRNhC0Ffo60uNQ==} cpu: [x64] os: [linux] + '@anthropic-ai/claude-agent-sdk-linux-x64@0.3.220': + resolution: {integrity: sha512-tkTJFnpR9VifvWX2fmkCAPkT6+8Wk/gVu8B5jsVekKZPiZoWRHmMXO30BnZn+f0TZhgYP+82PSX3S8crH1kn+w==} + cpu: [x64] + os: [linux] + '@anthropic-ai/claude-agent-sdk-win32-arm64@0.3.205': resolution: {integrity: sha512-SpP5zF68weFez/6pKrGzq/UVAJDMDNphWqmkLfOpWTDBL5xy6XlIZw5Bl4EXoVnfi2VLFkwuffNeFe+9SdX7kw==} cpu: [arm64] os: [win32] + '@anthropic-ai/claude-agent-sdk-win32-arm64@0.3.220': + resolution: {integrity: sha512-rIwgq0UwQExWl6KrHUyC4w5KwpL9l6nd95aUTx6RitexaAuEw//xtfTVLnuE4hDDQZFkzEwpdKc3nxDWoGcUbA==} + cpu: [arm64] + os: [win32] + '@anthropic-ai/claude-agent-sdk-win32-x64@0.3.205': resolution: {integrity: sha512-kg2kkXyeSoFLruO3Ic2IruLxzBR0xCUtmlJHdWi3SYW7JhAKNJg4fcrdJsWcardmEw23Y2UDGDJbRyxqSVx6wg==} cpu: [x64] os: [win32] + '@anthropic-ai/claude-agent-sdk-win32-x64@0.3.220': + resolution: {integrity: sha512-MuOuXhbr66HlGaWXD2f3w0k2PsvmnbkwcUZ0dAe2poFLdl72GC2dapwwOBefxm9QmoNqk9+jmv/dSKGOVWyvLw==} + cpu: [x64] + os: [win32] + '@anthropic-ai/claude-agent-sdk@0.3.205': resolution: {integrity: sha512-ft6iBw9kXudsusiXNpeybIPBJ07Z3tqp1ROSg5cEJqgA+9i+JJj2sRfQth+QD+lyenbbAU8yPieLxIimvfBhtw==} engines: {node: '>=18.0.0'} @@ -177,6 +217,14 @@ packages: '@modelcontextprotocol/sdk': ^1.29.0 zod: ^4.0.0 + '@anthropic-ai/claude-agent-sdk@0.3.220': + resolution: {integrity: sha512-glc7SdwPkOkLw8oxwLo9PKTdLJGqW/PIR4urWXFoRtX9YllwozsEVc5Tc1+EvLSkfrsxPJqQWqOgpjUOQXf1oA==} + engines: {node: '>=18.0.0'} + peerDependencies: + '@anthropic-ai/sdk': '>=0.93.0' + '@modelcontextprotocol/sdk': ^1.29.0 + zod: ^4.0.0 + '@anthropic-ai/sdk@0.110.0': resolution: {integrity: sha512-hOP4bNYXDFHDxxiEgzlILXrxZIYCDnhe8sry0RDRKD/QnsEpvZcQpablCdm9X/WuD/YgOiSIkkqsL1mLLlTqJw==} hasBin: true @@ -3495,27 +3543,51 @@ snapshots: '@anthropic-ai/claude-agent-sdk-darwin-arm64@0.3.205': optional: true + '@anthropic-ai/claude-agent-sdk-darwin-arm64@0.3.220': + optional: true + '@anthropic-ai/claude-agent-sdk-darwin-x64@0.3.205': optional: true + '@anthropic-ai/claude-agent-sdk-darwin-x64@0.3.220': + optional: true + '@anthropic-ai/claude-agent-sdk-linux-arm64-musl@0.3.205': optional: true + '@anthropic-ai/claude-agent-sdk-linux-arm64-musl@0.3.220': + optional: true + '@anthropic-ai/claude-agent-sdk-linux-arm64@0.3.205': optional: true + '@anthropic-ai/claude-agent-sdk-linux-arm64@0.3.220': + optional: true + '@anthropic-ai/claude-agent-sdk-linux-x64-musl@0.3.205': optional: true + '@anthropic-ai/claude-agent-sdk-linux-x64-musl@0.3.220': + optional: true + '@anthropic-ai/claude-agent-sdk-linux-x64@0.3.205': optional: true + '@anthropic-ai/claude-agent-sdk-linux-x64@0.3.220': + optional: true + '@anthropic-ai/claude-agent-sdk-win32-arm64@0.3.205': optional: true + '@anthropic-ai/claude-agent-sdk-win32-arm64@0.3.220': + optional: true + '@anthropic-ai/claude-agent-sdk-win32-x64@0.3.205': optional: true + '@anthropic-ai/claude-agent-sdk-win32-x64@0.3.220': + optional: true + '@anthropic-ai/claude-agent-sdk@0.3.205(@anthropic-ai/sdk@0.110.0(zod@4.4.3))(@modelcontextprotocol/sdk@1.29.0(zod@4.4.3))(zod@4.4.3)': dependencies: '@anthropic-ai/sdk': 0.110.0(zod@4.4.3) @@ -3531,6 +3603,21 @@ snapshots: '@anthropic-ai/claude-agent-sdk-win32-arm64': 0.3.205 '@anthropic-ai/claude-agent-sdk-win32-x64': 0.3.205 + '@anthropic-ai/claude-agent-sdk@0.3.220(@anthropic-ai/sdk@0.110.0(zod@4.4.3))(@modelcontextprotocol/sdk@1.29.0(zod@4.4.3))(zod@4.4.3)': + dependencies: + '@anthropic-ai/sdk': 0.110.0(zod@4.4.3) + '@modelcontextprotocol/sdk': 1.29.0(zod@4.4.3) + zod: 4.4.3 + optionalDependencies: + '@anthropic-ai/claude-agent-sdk-darwin-arm64': 0.3.220 + '@anthropic-ai/claude-agent-sdk-darwin-x64': 0.3.220 + '@anthropic-ai/claude-agent-sdk-linux-arm64': 0.3.220 + '@anthropic-ai/claude-agent-sdk-linux-arm64-musl': 0.3.220 + '@anthropic-ai/claude-agent-sdk-linux-x64': 0.3.220 + '@anthropic-ai/claude-agent-sdk-linux-x64-musl': 0.3.220 + '@anthropic-ai/claude-agent-sdk-win32-arm64': 0.3.220 + '@anthropic-ai/claude-agent-sdk-win32-x64': 0.3.220 + '@anthropic-ai/sdk@0.110.0(zod@4.4.3)': dependencies: json-schema-to-ts: 3.1.1 From f94ea6b412d97a01256fa6df8e7e2b9cde58813b Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Fri, 31 Jul 2026 14:37:08 -0700 Subject: [PATCH 8/9] [upd294c] review: cut the roster changeset to the model change and the refusal risk --- .changeset/review-correctness-opus-5.md | 10 ++-------- 1 file changed, 2 insertions(+), 8 deletions(-) diff --git a/.changeset/review-correctness-opus-5.md b/.changeset/review-correctness-opus-5.md index 68894ea0..bed3c3b1 100644 --- a/.changeset/review-correctness-opus-5.md +++ b/.changeset/review-correctness-opus-5.md @@ -2,12 +2,6 @@ "review": minor --- -Move the entire reviewer roster to Opus 5 (`claude-opus-5`): every role that ran `claude-opus-4-8` (orchestrator, `thread-reconciler`, `skill-auditor`, `claim-validator`, `conventions`, `documentation`, the opt-in `holistic` / `completeness` / `test-adequacy`, and all twelve specialist lenses) and both that ran `claude-fable-5` (`correctness-reviewer`, `first-principles`). `pattern-triage` stays on Sonnet 4.6 as the deliberately cheap first pass; the eval's judge and match arbiter stay on Haiku 4.5, being the ruler rather than the reviewer. +Move the reviewer roster to Opus 5 (`claude-opus-5`). `pattern-triage` stays on Sonnet 4.6 as the cheap first pass. -Opus 5 reports high precision and high recall at Opus 4.8's per-token price, half Fable's: a straight upgrade for the roles on 4.8, and for the two on Fable it should keep the recall the 2026-07-20 A/B bought (82% -> 89% must-catch) without the +35% premium. The powered A/B is the acceptance criterion, not this description. Read four rows first, each independently revertable: `correctness-reviewer` (recall), `claim-validator` (precision), the `security-auth` lens, and the orchestrator (throughput under the wall clock). - -**Known risk: refusal on the specialist lenses.** Those were kept off Fable because cyber safety classifiers can refuse benign security analysis, and a refused lens is a silent coverage hole rather than an error. Opus 5 can also return `stop_reason: "refusal"` on cyber-adjacent input, so this re-opens that hole. The detector is the weekly drift corpus; on that signature put `security-auth` back on `claude-opus-4-8` first, the narrowest revert available. - -`claude-opus-5` is priced by the `models.providers` overlay this stacks on, so it needs firewall v0.27.43 (gh-aw v0.84.x). Under v0.27.42 the overlay is dropped, the model is un-priced, and the api-proxy rejects every dispatch with a 400: do not merge before that release is stable. - -Also bumps `@anthropic-ai/claude-agent-sdk` to `^0.3.219`, the first release that knows `claude-opus-5`. The eval enforces `--max-usd` from the SDK's reported `total_cost_usd`, so an older SDK turns the budget cap into a no-op on precisely the arm under test. Eval-only; the SDK is not in the api-proxy path. +The specialist lenses were kept off Fable 5 because cyber safety classifiers can refuse benign security-focused analysis, and a refused lens is a silent coverage hole: it surfaces as a missing agent result, not an error. Opus 5 can also return `stop_reason: "refusal"` on cyber-adjacent input, so this move re-opens that hole rather than closing it. The detector is the weekly drift corpus, where a refusing lens craters must-catch recall on security-adjacent cases while every other metric looks normal. On that signature, put the `security-auth` lens back on `claude-opus-4-8` first. From 470b456efd97b3e35016f576e4f6d53504d82642 Mon Sep 17 00:00:00 2001 From: James Wiesebron Date: Thu, 13 Aug 2026 10:44:05 -0400 Subject: [PATCH 9/9] [294-work] review: restore the Opus 5 credit-guard fallback and gate pricing mechanically Review feedback on #294: - The merge from the pricing-overlay base (#314) dropped the models.default-ai-credits-pricing block for the second time in this branch's history; without it every dispatch 400s on the stable toolchain (claude-opus-5 is not in firewall v0.27.42's curated table). Restored, with the prose reconciled to the merged 50% providers overlay: fallback rates stay at list because in the only window it binds, every other model bills at list too. - New model-pricing.test.ts turns both prose warnings into failing checks: every model pin must have a providers overlay entry, and the fallback must exist while claude-opus-5 is pinned (delete that test with the fallback block once gh-aw defaults to v0.27.43+). - Stale rationale comments updated: correctness-reviewer's Fable 5 justification now states what Opus 5 carries; first-principles records the move; the overlay's engine markers moved to opus-5, with opus-4-8 relabeled as the refusal-fallback target. - Changeset renamed to review-roster-opus-5 (roster-wide scope) and records why security-auth moves with the roster instead of a pre-emptive carve-out. Also merges the base branch (sonnet-4-6 overlay entry). --- ...ness-opus-5.md => review-roster-opus-5.md} | 2 +- workflows/review/lib/model-pricing.test.ts | 78 +++++++++++++++++++ workflows/review/review.md | 53 +++++++++++-- 3 files changed, 125 insertions(+), 8 deletions(-) rename .changeset/{review-correctness-opus-5.md => review-roster-opus-5.md} (61%) create mode 100644 workflows/review/lib/model-pricing.test.ts diff --git a/.changeset/review-correctness-opus-5.md b/.changeset/review-roster-opus-5.md similarity index 61% rename from .changeset/review-correctness-opus-5.md rename to .changeset/review-roster-opus-5.md index bed3c3b1..fa09caf2 100644 --- a/.changeset/review-correctness-opus-5.md +++ b/.changeset/review-roster-opus-5.md @@ -4,4 +4,4 @@ Move the reviewer roster to Opus 5 (`claude-opus-5`). `pattern-triage` stays on Sonnet 4.6 as the cheap first pass. -The specialist lenses were kept off Fable 5 because cyber safety classifiers can refuse benign security-focused analysis, and a refused lens is a silent coverage hole: it surfaces as a missing agent result, not an error. Opus 5 can also return `stop_reason: "refusal"` on cyber-adjacent input, so this move re-opens that hole rather than closing it. The detector is the weekly drift corpus, where a refusing lens craters must-catch recall on security-adjacent cases while every other metric looks normal. On that signature, put the `security-auth` lens back on `claude-opus-4-8` first. +The specialist lenses were kept off Fable 5 because cyber safety classifiers can refuse benign security-focused analysis, and a refused lens is a silent coverage hole: it surfaces as a missing agent result, not an error. Opus 5 can also return `stop_reason: "refusal"` on cyber-adjacent input, so this move re-opens that hole rather than closing it. The detector is the weekly drift corpus, where a refusing lens craters must-catch recall on security-adjacent cases while every other metric looks normal. On that signature, put the `security-auth` lens back on `claude-opus-4-8` first. Moving `security-auth` with the roster rather than carving it out pre-emptively is deliberate: the refusal hazard is intermittent and unproven on Opus 5, the one-hop runtime fallback (a refused agent re-dispatches once on `claude-opus-4-8`, recorded as `fellBackTo`) bounds each occurrence, and a uniform roster leaves one signature to watch and one revert to make instead of a permanent carve-out justified by a hazard that may not materialize. diff --git a/workflows/review/lib/model-pricing.test.ts b/workflows/review/lib/model-pricing.test.ts new file mode 100644 index 00000000..46ee6eed --- /dev/null +++ b/workflows/review/lib/model-pricing.test.ts @@ -0,0 +1,78 @@ +/** + * Mechanical pricing gates for the model pins in review.md (#294 review + * feedback: the merge-ordering constraint "do not ship an un-priced pin" + * rested entirely on human memory of a draft PR). + * + * Two hazards, both prose-only until this file: + * + * - The `models.providers` overlay matches per model and an unlisted model + * silently bills at FULL list price (the overlay's own MAINTENANCE note); + * the overlay omitting `claude-sonnet-4-6` shipped exactly that way in an + * earlier draft of #314. + * - On the stable toolchain (gh-aw v0.83.x -> firewall v0.27.42) the + * `providers` block is dropped silently and the api-proxy's credit guard + * rejects a model its curated table does not price with a 400 before the + * request reaches the model (#266). `claude-opus-5` is not in that table, + * so the `default-ai-credits-pricing` fallback is load-bearing for every + * dispatch until a gh-aw release defaults the firewall to v0.27.43+. + * + * DELETE the fallback test (only it) together with the fallback block when + * the toolchain moves; the coverage test is permanent. + */ +import {readFileSync} from "node:fs"; +import {join} from "node:path"; + +import {describe, it, expect} from "vitest"; + +const reviewMd = readFileSync(join(__dirname, "..", "review.md"), "utf8"); + +/** The workflow frontmatter (between the first pair of --- fences). */ +const frontmatter = reviewMd.split(/^---$/m)[1] ?? ""; + +/** + * Every model pin in the file: the engine's indented `model:` line in the + * frontmatter plus each sub-agent's `model:` line in its block frontmatter. + */ +const pins = [ + ...new Set( + [...reviewMd.matchAll(/^\s*model:\s*(claude-[a-z0-9.-]+)\s*$/gm)].map( + (match) => match[1], + ), + ), +]; + +/** + * The models the `providers` overlay prices: bare `claude-*:` mapping keys in + * the frontmatter (only the providers block declares them). + */ +const priced = new Set( + [...frontmatter.matchAll(/^\s+(claude-[a-z0-9.-]+):\s*$/gm)].map( + (match) => match[1], + ), +); + +describe("model pricing coverage (review.md frontmatter)", () => { + it("finds the pins and the overlay (guards the extraction itself)", () => { + // 22 agents plus the engine; a collapse to zero means the regexes + // rotted, not that the roster emptied. + expect(pins.length).toBeGreaterThanOrEqual(2); + expect(priced.size).toBeGreaterThanOrEqual(2); + }); + + it("prices every pinned model in the providers overlay", () => { + const unpriced = pins.filter((pin) => !priced.has(pin)); + // An unlisted model silently bills at full list price; add an entry + // at 50% of its Anthropic list rate (see the MAINTENANCE note). + expect(unpriced).toEqual([]); + }); + + it("keeps the stable-toolchain credit-guard fallback while claude-opus-5 is pinned", () => { + // Firewall v0.27.42's curated table does not price claude-opus-5; + // without `default-ai-credits-pricing` every dispatch 400s on the + // stable toolchain. Delete this test with the fallback block once a + // gh-aw release defaults the firewall to v0.27.43+. + if (pins.includes("claude-opus-5")) { + expect(frontmatter).toContain("default-ai-credits-pricing:"); + } + }); +}); diff --git a/workflows/review/review.md b/workflows/review/review.md index 9e7f4ced..93ad2e38 100644 --- a/workflows/review/review.md +++ b/workflows/review/review.md @@ -343,23 +343,60 @@ post-steps: # Do NOT collapse these to a bare `claude-opus-4` prefix: prefix matching would # also capture opus-4-0/4-1, which list at 3x the 4-5+ rate. models: + # claude-opus-5 (the engine and roster pin) is NOT in the firewall + # api-proxy's curated AI-credits pricing table at v0.27.42, the release + # gh-aw v0.83.4 defaults to (that table carries claude-opus-4-5 through 4-8 + # and claude-fable-5, and stops there). The proxy's AI-credits guard rejects + # an un-priced model with a 400 BEFORE the request reaches the model, so + # without this fallback every dispatch fails on the stable toolchain: the + # #266 failure that killed the first-principles dispatch on every run, + # except the whole roster runs the un-priced model rather than two opt-in + # agents. Units differ from the overlay below and the two are not + # interchangeable: `default-ai-credits-pricing` is $/1M tokens and feeds + # the credit guard; `providers` is $/token. Rates here are Anthropic LIST + # (Opus 5 lists at exactly Opus 4.8's price), deliberately not Khan's 50% + # rate: in the only window where this fallback binds (stable gh-aw v0.83.x, + # firewall v0.27.42, which drops the `providers` block silently) every + # other model bills at list from the curated table, so list keeps Opus 5 + # denominated consistently with the rest of the roster. Two caveats: the + # default-pricing path does not bill cache writes ($6.25/M real), so credit + # accounting under-counts that component while it binds; and the fallback + # applies to ANY un-priced model, so a typo'd model id bills at Opus rates + # instead of failing loudly. + # + # REMOVE THIS FALLBACK when a gh-aw release defaults the firewall to + # v0.27.43 or later: that release carries a curated claude-opus-5 entry at + # the same list rates and bills cache writes, and the recompile also makes + # the `providers` overlay below live (the higher-precedence source). Do NOT + # reach for `sandbox.agent.version: v0.27.43` to get there early; a version + # is pinned here only to hold a release BACK, never to move one forward. + # + # MINIMUM COMPILER: gh-aw >= v0.83.0 for `models.default-ai-credits-pricing`. + # $/1M tokens. `input` and `output` are the only rates the schema accepts, + # so the cache rates are the proxy's derivations, not ours. + default-ai-credits-pricing: + input: 5.0 + output: 25.0 providers: anthropic: models: - # Current engine model. + # The refusal-fallback target (the one-hop re-dispatch of #315) and + # an `engine:` override candidate; the engine until the roster moved + # to Opus 5. claude-opus-4-8: cost: input: "2.5e-06" output: "1.25e-05" cache_read: "2.5e-07" cache_write: "3.125e-06" - # Engine models a consumer may select via an `engine:` override. + # Current engine and roster model. claude-opus-5: cost: input: "2.5e-06" output: "1.25e-05" cache_read: "2.5e-07" cache_write: "3.125e-06" + # Engine models a consumer may select via an `engine:` override. claude-sonnet-5: cost: input: "1e-06" @@ -1189,8 +1226,9 @@ description: Classifies each changed file's risk and reviews the diff for correc model: claude-opus-5 # effort: high — launch default (whole-change reviewer). gh-aw has no per-agent # effort field yet; the per-role model/effort table lives in the README. -# Fable 5: bug-finding recall is this workflow's load-bearing metric, and -# stronger real-defect detection is Fable's headline gain over Opus 4.8. +# Opus 5: bug-finding recall is this workflow's load-bearing metric; the +# 2026-07-20 A/B recall gain that lived on Fable 5 is carried by Opus 5 at +# Opus 4.8's per-token price. See the roster table in the README. --- You are a correctness-focused code reviewer. You have **no GitHub access** — read the diff and file list from disk and return your result as JSON only. @@ -2000,9 +2038,10 @@ If the changed behavior is adequately tested, return {"findings": []}. name: first-principles description: A diverse-perspective, advisory-only sanity check on whether the change should exist as written; returns findings as JSON. model: claude-opus-5 -# effort: high — launch default. Ran on Fable 5 (claude-fable-5) from day one; -# the correctness reviewer joined it after the 2026-07-20 A/B. Advisory-only, -# never blocks. +# effort: high — launch default. Ran on Fable 5 (claude-fable-5) from day one, +# partly to be the one non-Opus reviewer; the correctness reviewer joined it +# after the 2026-07-20 A/B, and it moved to Opus 5 with the roster. +# Advisory-only, never blocks. --- You are the **first-principles** reviewer. Your single mandate is to review the **justification for the change, not the change itself**: where `holistic` asks