Skip to content
Merged
Show file tree
Hide file tree
Changes from 82 commits
Commits
Show all changes
83 commits
Select commit Hold shift + click to select a range
441f63a
implement(documenter): R1 single-body submission + E2 pr-context stag…
Jul 3, 2026
c2279b6
implement(coder): R8(a) versioned structured finding schema + validator
Jul 3, 2026
6a85827
implement(tester): unit tests for R8(a) finding schema/validator (tas…
Jul 3, 2026
107e956
implement(tester): unit tests for R8(a) finding schema/validator (tas…
Jul 3, 2026
8826894
Persist BRC history for slice-1 (#2548)
Jul 3, 2026
c02da14
slice-1: prettier fixes + changeset
jwbron Jul 3, 2026
b119a01
slice-1: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
b9508a6
review: address slice-1 feedback (empty body, lib distribution, comme…
jwbron Jul 6, 2026
123a10f
review: drop lib install step; scripts run via npx tsx
jwbron Jul 7, 2026
e637ece
implement(coder): R8(b) computed verdict + R8(c) Conventional-Comment…
Jul 3, 2026
8819339
implement(tester): R8(b)/R2 verdict truth-table + R8(c) rendering sna…
Jul 3, 2026
545e711
implement(tester): prettier-format task-2-4 test files (lint gate)
Jul 3, 2026
40efb53
implement(coder): prettier 2.6.2 lint fixes for slice-2 + finding-sch…
Jul 3, 2026
f618d2a
Persist BRC history for slice-2 (#2548)
Jul 3, 2026
6058d86
slice-2: changeset
jwbron Jul 3, 2026
da5b450
slice-2: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
0950bf2
review: address slice-2 feedback (verdict precedence, hold UX, empty …
jwbron Jul 6, 2026
6cfe3e3
implement(documenter): wire deterministic router into Step 3; drop re…
Jul 3, 2026
7959787
implement(documenter): fix router invocation path to lib/router.ts (r…
Jul 3, 2026
d96789d
implement(coder): R10 deterministic router + tier-scaled budget (slic…
Jul 3, 2026
3f919f3
implement(tester): R10 router unit tests — classification, lens, tier…
Jul 3, 2026
f3130ec
implement(coder): router CLI entrypoint + routing.json serialization …
Jul 3, 2026
a6984f1
Persist BRC history for slice-3 (#2548)
Jul 3, 2026
c41a812
implement(tester): cover router v2 serialization + CLI (toRoutingJson…
Jul 3, 2026
cb8d714
slice-3: replace satisfies with prettier-2-compatible check, changeset
jwbron Jul 3, 2026
29f2b6f
slice-3: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
abdf46c
review: address slice-3 feedback (consumer-owned routing config)
jwbron Jul 6, 2026
06cbb2e
review: last-match-wins tier precedence; move ROUTING format doc to R…
jwbron Jul 7, 2026
06ac3ee
implement(documenter): E1/E3/E5/E6/E7/R3b reliability prompt edits (s…
Jul 3, 2026
5bb941d
Persist BRC history for slice-4 (#2548)
Jul 3, 2026
a21b4ce
slice-4: changeset
jwbron Jul 3, 2026
fdf05d2
slice-4: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
798016b
review: address slice-4 feedback (steering text, touched-lines scoping)
jwbron Jul 6, 2026
238a355
implement(documenter): R9 bounded-investigation instructions for revi…
Jul 3, 2026
e5d1fd2
implement(coder): R9 per-finding investigation tool-call cap (slice-5)
Jul 3, 2026
8074db2
implement(tester): R9 per-finding tool-call cap tests (task-5-3)
Jul 3, 2026
a01c756
implement(tester): cover check() refusal previews (per-finding + run-…
Jul 3, 2026
1340a42
slice-5: prettier fix + changeset
jwbron Jul 3, 2026
7599eea
review: address slice-5 feedback (wire the investigation cap to its c…
jwbron Jul 6, 2026
3f01827
review: investigation-cap call sites run via npx tsx
jwbron Jul 7, 2026
4e62371
implement(documenter): slice-6 roster framework — always-on reviewers…
Jul 3, 2026
8d21077
Persist BRC history for slice-6 (#2548)
Jul 3, 2026
481198d
slice-6: changeset
jwbron Jul 3, 2026
33de326
slice-6: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
46228e0
review: rework slice-6 roster to opt-in (no new default cost)
jwbron Jul 6, 2026
8f8cb7c
review: genericize orchestrator (uniform findings contract, roster de…
jwbron Jul 7, 2026
ad4bede
slice-7: eleven specialist lenses (rebased onto genericized orchestra…
jwbron Jul 7, 2026
3e484a6
implement(coder): slice-9 — smoke benchmark (corpus + no-post runner …
Jul 3, 2026
f4160f2
implement(tester): slice-9 smoke.test.ts — vitest CI gate over the sm…
Jul 3, 2026
6eb3bf8
Persist BRC history for slice-9 (#2548)
Jul 3, 2026
54c36e2
slice-9: prettier fixes + changeset
jwbron Jul 3, 2026
cc4b30b
slice-9: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
53d0d9a
review: drop planning identifiers from the smoke benchmark (slice-9)
jwbron Jul 6, 2026
35216aa
review: drop low-value smoke corpus size test
jwbron Jul 7, 2026
1ab96d8
implement(tester): slice-10 rebalance verification against smoke set …
Jul 3, 2026
c4f33b7
implement(documenter): slice-10 wave-2 rebalance — edits 8-13, refute…
Jul 3, 2026
f76c328
Persist BRC history for slice-10 (#2548)
Jul 3, 2026
23aba3b
slice-10: prettier fix + changeset
jwbron Jul 3, 2026
721f8a5
slice-10: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
fa0bfbb
review: strip remaining plan identifiers (slice-10 and lens prompts)
jwbron Jul 6, 2026
38f3e75
review: drop the webapp-40536 experiment record (moves to the PR desc…
jwbron Jul 7, 2026
5d9fee4
review: drop the blocking-claim refuter panel; gate blocking on valid…
jwbron Jul 7, 2026
766a8bc
review: strip a reintroduced plan identifier from the verdict-gate pa…
jwbron Jul 7, 2026
1f735fa
implement(coder): slice-11 — full eval suite (datasets, metrics, judg…
Jul 3, 2026
cf342d2
implement(coder): slice-11 judge.ts — decouple from slice-8 type; har…
Jul 3, 2026
992048c
implement(tester): slice-11 eval-suite self-tests + CI-gate guard (ta…
Jul 3, 2026
14c38fb
Persist BRC history for slice-11 (#2548)
Jul 3, 2026
b6db587
slice-11: staged scheduled full-eval workflow with live judge (task-1…
jwbron Jul 3, 2026
169a6b8
slice-11: prettier fixes + changeset
jwbron Jul 3, 2026
96b0895
slice-11: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
4eafaf6
review: strip planning identifiers from the eval suite (slice-11)
jwbron Jul 6, 2026
af7ad3a
review: eval CI rework; deterministic suite rides node-ci, live judge…
jwbron Jul 7, 2026
32e773c
implement(documenter): slice-12 — R13 per-finding re-review resolutio…
Jul 3, 2026
29d5d73
implement(documenter): slice-12 R14 — correct drift-guard doc (review…
Jul 3, 2026
42375de
implement(coder): slice-12 P2 — live counters, dismissal-learning, co…
Jul 3, 2026
a22462c
implement(coder): slice-12 — fix NUL-byte encoding in dismissal-learn…
Jul 3, 2026
dbca449
Persist BRC history for slice-12 (#2548)
Jul 3, 2026
cba87ea
slice-12: prettier fixes + changeset
jwbron Jul 3, 2026
a975f59
slice-12: drop BRC history artifacts from slice PR
jwbron Jul 3, 2026
24ad180
review: strip planning identifiers from slice-12 modules and remainin…
jwbron Jul 6, 2026
a8492aa
review: plain version marker for attribution (semver is the behavior …
jwbron Jul 7, 2026
0452c8a
review: read per-directory REVIEW.md contracts from the consuming repo
jwbron Jul 7, 2026
e683d5c
Merge origin/main into review/read-review-md-contracts (resolve stack…
jwbron Jul 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .changeset/review-computed-verdict.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Make the review verdict a deterministic computation instead of a model judgment: structured findings flow through a computed verdict (severity/confidence truth table, with a hold-for-human outcome for policy conflicts), templated comment rendering, and a missing-dimension gate that refuses to conclude when a required review dimension produced no findings and no explicit all-clear. The model authors findings; code decides the verdict and renders the comments.
5 changes: 5 additions & 0 deletions .changeset/review-deterministic-router.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Add the deterministic review router (`workflows/review/lib/router.ts`): pure TypeScript that classifies changed files (generated vs source via .gitattributes), selects which specialist lenses to spawn by touched paths, maps owning teams from REVIEWERS rules (subsuming the old reviewer-mapper sub-agent), assigns per-file risk tiers, and computes one run budget scaled by the highest touched tier with a misrouted-PR floor. Diff-direction-dependent tiers are deferred to a single small-model call instead of guessed.
5 changes: 5 additions & 0 deletions .changeset/review-finding-schema-foundations.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Add the versioned structured finding schema (`workflows/review/lib/finding-schema.ts`): every sub-agent finding now carries id, lens, anchor (line/range/file/pr-level), severity, confidence, evidence trace, optional suggested patch and pre-merge obligation, validated against an exported `FINDING_SCHEMA_VERSION`. Review submission is standardized on a single robust `submit-pull-request-review` call with a guaranteed non-empty body (the empty-body retry fallback is removed), and PR context (`pr-context.json`) is staged on disk once per run for all sub-agents, extending the existing diff staging.
5 changes: 5 additions & 0 deletions .changeset/review-full-eval-suite.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Add the full eval suite under `workflows/review/eval/`: a four-dataset corpus (incidents, synthetic mutations, golden PRs, known-clean, plus an adversarial holdout), five pure metrics (must-catch recall, golden precision, clean false-block, noise, confidence calibration with ECE), an LLM judge with an injected model seam pinned to Opus 4.8 (audit sampling and thumbs calibration included), and release gates (adversarial hard gate for automatic mode, overfitting advisory). The deterministic suite runs on every PR through the normal test runner (node-ci); the live judge runs weekly via `.github/workflows/review-eval-full.yml` and the committed `workflows/review/eval/live-judge.ts` entry point, reporting to the job summary.
5 changes: 5 additions & 0 deletions .changeset/review-investigation-cap.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Bound reviewer investigation (R9): sub-agents get explicit instructions for cheap targeted verification (grep callers, trace the call chain, one targeted check per finding), and `workflows/review/lib/investigation-cap.ts` enforces a deterministic per-finding and per-run tool-call cap sourced from the router's run budget. Over-cap calls are refused with fixed reason codes and refusals never mutate state.
5 changes: 5 additions & 0 deletions .changeset/review-md-contracts.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Read per-directory `REVIEW.md` review contracts from the consuming repo. When the checkout carries them (a root `REVIEW.md` plus per-directory ones, as in webapp's agent-doc surface), the `correctness-reviewer` reads the root contract plus the nearest `REVIEW.md` above each reviewed file to calibrate what is Important versus a nit in that sub-tree, and the `claim-validator` uses the same contracts to calibrate claim labels (never `verification`, which stays code-evidence-only). Contract text can adjust emphasis but never overrides the workflow's rules, and since these files are read from the PR head, an edited `REVIEW.md` is itself reviewed on its merits. Repos without `REVIEW.md` files are unaffected.
5 changes: 5 additions & 0 deletions .changeset/review-p2-items.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

P2 items: the thread-reconciler recognizes three terminal resolutions (fixed, deferred-to-filed-issue, disagreed-with-reason) before keeping a thread open; the guidance comment carries a plain version marker (release tag + finding-schema version) for attribution and rollback, with semver as the behavior contract (no drift-stamp machinery); `lib/counters.ts` mines run health metrics (validator drop rate, comments per PR, verdict mix, thumbs agreement, cost) from existing per-run artifacts; and APPROVE-with-obligations renders pre-merge obligations as a distinct comment from the finding schema. Dismissal-learning is deferred until the thumbs sweep is scheduled and has accumulated real dismissal signals.
5 changes: 5 additions & 0 deletions .changeset/review-roster-framework.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Restructure the reviewer roster: five always-on reviewers (holistic, completeness, test-adequacy, first-principles, conventions) with explicit trust and advisory constraints and named mandates; launch-default model and effort assignments per role (Opus 4.8 workhorse, xhigh for the security lens and claim validation, Fable 5 for first-principles); and the kept gates (pattern-triage, claim-validator, deterministic dedup, verdict bookends, thread-reconciler) wired to the new roster.
5 changes: 5 additions & 0 deletions .changeset/review-smoke-benchmark.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Add the smoke benchmark: a 13-case tagged subset of the eval corpus (incident repros, adversarial-injection PRs, known-clean PRs) in the shared dataset format, a no-post runner that replays cases through the real deterministic review path (router, labels, scope filter, verdict, render) with zero GitHub writes, and a vitest gate (`workflows/review/eval/smoke.test.ts`) that asserts each case's computed verdict against its expected block. A dedicated CI entry point is staged at `.github-staging/review-smoke.yml` pending a human `git mv` into `.github/workflows/`.
5 changes: 5 additions & 0 deletions .changeset/review-specialist-lenses.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Add the eleven path-routed specialist reviewer lenses: security-auth (single lens, xhigh effort), ai-safety-moderation, mass-comms-coppa, caching-resource, data-migrations, concurrency-async, api-federation-compat, cross-deploy-serialization, deploy-infra-config, money-payments, and content-i18n. Each lens runs a tri-state incident hunt (found / not-found / not-applicable) and emits schema-validated structured findings; the standalone skill-auditor sub-agent is folded into the correctness pass, keeping the repo-specific correctness-checks extension point.
5 changes: 5 additions & 0 deletions .changeset/review-wave1-prompt-edits.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": patch
---

Reliability and quality prompt edits (wave 1): high-risk triggers are named then judged in one line (E1); prompt-injection attempts in the diff or PR description are themselves blocking findings (E3); deletions are findings (E5); bot threads are judged against the whole reply chain and a conceded point is never re-raised (E6); open human threads are deferred to instead of duplicated (E7); and pre-existing bugs on touched lines are in scope (R3b).
5 changes: 5 additions & 0 deletions .changeset/review-wave2-rebalance.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": patch
---

Recall/precision rebalance (wave 2): coverage-first hunting before filtering (E8), every blocking finding must state a concrete failing scenario (E9), three-state claim validation with symmetric evidence duties — refuted claims are dropped, uncertain ones downgraded to non-blocking, and only a validator-confirmed claim may carry a blocking label into the verdict (E10), claims are confirmed against the code before assertion (E11), findings cite exact lines (E12), and a posting bar keeps inline comments at confidence 0.5 or higher with low-confidence findings collapsed into one section (E13). An author-disputed claim cannot re-block unless the re-check traces to actual usage; otherwise it posts as a question. The planned blocking-claim refuter panel is removed rather than shipped (a production audit of 319 blocking claims found its addressable class was 2 claims; the confirmed-only verdict gate covers the rest at zero extra agents), with the eval suite's false-block metric as the tripwire for revisiting. The no-post runner gains a validation-replay stage and the smoke corpus gains six audit-seeded cases pinning the gate: production false blocks must approve, confirmed blocks must keep blocking.
40 changes: 40 additions & 0 deletions .github-staging/review-smoke.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
# Review smoke benchmark (R5, task-9-4)
#
# STAGED FILE — NOT YET ACTIVE. It lives under `.github-staging/` because no
# producer role in this pipeline may push to `.github/` directly (gateway phase
# restriction, #2508). A human must move it into place for it to run:
#
# git mv .github-staging/review-smoke.yml .github/workflows/review-smoke.yml
#
# (See the PR body's "Pre-merge obligations" note.) Until then the smoke set is
# still enforced by the repo-wide `pnpm run test` gate in node-ci.yml, which runs
# the same vitest suite; this dedicated entry point just makes the smoke gate
# explicit and independently visible on every PR, and does not assume the
# existing CI wiring.
#
# It runs the slice-9 smoke benchmark: the tagged smoke subset of the eval corpus
# replayed through the no-post runner (workflows/review/eval/runner.ts), which
# exercises the real deterministic review path (router -> labels -> scope filter
# -> verdict -> render) and performs NO GitHub write. The vitest gate lives with
# the eval harness (workflows/review/eval/) and asserts each smoke case's computed
# verdict against its `expected` block.

name: Review Smoke

on:
pull_request:
paths:
- "workflows/review/**"
- ".github/workflows/review-smoke.yml"

jobs:
smoke:
name: PR-reviewer smoke benchmark
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@93cb6efe18208431cddfb8368fd83d5badbf9bfd # v5
- uses: ./actions/shared-node-cache
# Run only the review eval harness (loader + no-post runner + smoke cases).
# `--run` forces a single non-watch run; the path filter scopes vitest to
# the smoke gate so this job stays fast and focused.
- run: pnpm run test --run workflows/review/eval
39 changes: 39 additions & 0 deletions .github/workflows/review-eval-full.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# Review eval — scheduled live judge.
#
# The DETERMINISTIC eval suite (loadCorpus -> runCorpus -> computeMetrics ->
# evaluateGates, plus the stub-judge composition checks) needs no secrets and
# already runs on every PR via the normal test runner (`pnpm run test --run`
# in node-ci.yml), so it has no job here.
#
# What runs here is the one thing normal CI cannot do: the same corpus run
# scored by the LIVE pinned judge model (judge.ts PINNED_JUDGE_MODEL) through
# the pure JudgeModel seam. Judge calls cost real money and must never gate
# PRs, so this is deliberately off-PR on a weekly schedule. The entry point is
# the committed script `workflows/review/eval/live-judge.ts` (lint/typecheck/
# test-covered like any other source file); it appends the metrics/gates/judge
# report to the job summary so results are visible without opening the log.

name: Review Eval Live Judge

on:
schedule:
# Weekly, Monday 09:00 UTC.
- cron: "0 9 * * 1"
workflow_dispatch: {}

jobs:
live-judge:
name: Live judge over the full corpus
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@93cb6efe18208431cddfb8368fd83d5badbf9bfd # v5
- uses: ./actions/shared-node-cache
- name: Run the live judge over the full corpus
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
if [ -z "$ANTHROPIC_API_KEY" ]; then
echo "ANTHROPIC_API_KEY secret not configured; skipping the live-judge run." >&2
exit 0
fi
pnpm dlx tsx workflows/review/eval/live-judge.ts
7 changes: 2 additions & 5 deletions package.json
Original file line number Diff line number Diff line change
Expand Up @@ -23,11 +23,8 @@
"fast-glob": "^3.3.3",
"memfs": "^4.51.0",
"prettier": "^2.6.2",
"typescript": "^5.9.3",
"vitest": "^4.0.10"
},
"packageManager": "pnpm@10.0.0+sha512.b8fef5494bd3fe4cbd4edabd0745df2ee5be3e4b0b8b08fa643aa3e4c6702ccc0f00d68fa8a8c9858a735a0032485a44990ed2810526c875e416f001b17df12b",
"dependencies": {
"@swc-node/register": "^1.11.1",
"typescript": "^5.9.3"
}
"packageManager": "pnpm@10.0.0+sha512.b8fef5494bd3fe4cbd4edabd0745df2ee5be3e4b0b8b08fa643aa3e4c6702ccc0f00d68fa8a8c9858a735a0032485a44990ed2810526c875e416f001b17df12b"
}
13 changes: 6 additions & 7 deletions pnpm-lock.yaml

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading
Loading