Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
81 commits
Select commit Hold shift + click to select a range
2d746ac
[jwbron/live-eval-corpus] review: live-enabled corpus format and ten …
jwbron Jul 9, 2026
6727eb6
[jwbron/live-eval-producer-staging] review: live-producer prompt extr…
jwbron Jul 9, 2026
5c00cf3
[jwbron/live-eval-producer-staging] review: the live producer and SDK…
jwbron Jul 9, 2026
52073ef
[jwbron/live-eval-ab-runner] review: the live A/B runner (phase 3)
jwbron Jul 9, 2026
93770e6
[jwbron/live-eval-ab-ci] review: per-PR live A/B workflow (phase 4)
jwbron Jul 9, 2026
30e7527
[jwbron/live-eval-corpus] review: exclude eval-corpus trees from lint…
jwbron Jul 9, 2026
19f06cb
[jwbron/live-eval-producer-staging] Merge branch 'jwbron/live-eval-co…
jwbron Jul 9, 2026
d1b5288
[jwbron/live-eval-ab-runner] Merge branch 'jwbron/live-eval-producer-…
jwbron Jul 9, 2026
254dfc4
[jwbron/live-eval-ab-ci] Merge branch 'jwbron/live-eval-ab-runner' in…
jwbron Jul 9, 2026
917cc11
[jwbron/live-eval-producer-staging] review: namespace live finding id…
jwbron Jul 9, 2026
0ee295a
[jwbron/live-eval-ab-runner] Merge branch 'jwbron/live-eval-producer-…
jwbron Jul 9, 2026
58cd4ed
[jwbron/live-eval-ab-runner] review: judge failures degrade the A/B r…
jwbron Jul 9, 2026
b19b08d
[jwbron/live-eval-ab-ci] Merge branch 'jwbron/live-eval-ab-runner' in…
jwbron Jul 9, 2026
bced291
[jwbron/review-trial-skill] review: add the review-trial skill (live …
jwbron Jul 9, 2026
da1115c
[tmp-refresh] Merge remote-tracking branch 'origin/main' into tmp-ref…
jwbron Jul 9, 2026
a6883be
[tmp-refresh] Merge remote-tracking branch 'origin/jwbron/live-eval-c…
jwbron Jul 9, 2026
0d02672
[tmp-refresh] Merge remote-tracking branch 'origin/jwbron/live-eval-p…
jwbron Jul 9, 2026
93d8dec
[tmp-refresh] Merge remote-tracking branch 'origin/jwbron/live-eval-a…
jwbron Jul 9, 2026
e3b34eb
[tmp-refresh] Merge remote-tracking branch 'origin/jwbron/live-eval-a…
jwbron Jul 9, 2026
491a983
[jwbron/review-rereview-accountability] review: re-review accountabil…
jwbron Jul 9, 2026
7b5318c
[jwbron/review-out-of-lane] review: hand off out-of-lane observations…
jwbron Jul 9, 2026
5dd182b
[jwbron/review-out-artifact-upload] review: fix the out/ artifact upl…
jwbron Jul 9, 2026
92bffa2
[jwbron/review-rereview-accountability] review: prettier-format the r…
jwbron Jul 9, 2026
2be8ede
[jwbron/review-out-of-lane] Merge branch 'jwbron/review-rereview-acco…
jwbron Jul 9, 2026
3a8fc5c
[jwbron/review-out-of-lane] review: prettier-format finding-schema
jwbron Jul 9, 2026
813767c
[jwbron/review-dispatch-tax] review: dedupe lens discipline snippets …
jwbron Jul 9, 2026
2812679
[jwbron/live-eval-corpus] review: route the specialist lens on each l…
jwbron Jul 9, 2026
2ce35a0
[jwbron/live-eval-producer-staging] Merge branch 'jwbron/live-eval-co…
jwbron Jul 9, 2026
c0fece2
[jwbron/live-eval-ab-runner] Merge branch 'jwbron/live-eval-producer-…
jwbron Jul 9, 2026
e02ac40
[jwbron/live-eval-ab-runner] review: carry agent-failure reasons into…
jwbron Jul 9, 2026
d225bd4
[jwbron/live-eval-ab-ci] Merge branch 'jwbron/live-eval-ab-runner' in…
jwbron Jul 9, 2026
de239ee
[jwbron/rereview-mode-dial] review: the re-review mode dial: ROUTING …
jwbron Jul 9, 2026
1329297
[jwbron/rereview-mode-dial] review: expose keptBlockingCount from the…
jwbron Jul 9, 2026
45c6f6e
[jwbron/rereview-mode-dial] review: wire the re-review mode dial into…
jwbron Jul 9, 2026
976b925
[jwbron/rereview-mode-dial] review: prettier-format the mode-dial fil…
jwbron Jul 9, 2026
cbc838d
[jwbron/live-eval-ab-runner] review: close every A/B report with a pe…
jwbron Jul 9, 2026
63097f1
[jwbron/live-eval-corpus] Merge remote-tracking branch 'origin/jwbron…
jwbron Jul 9, 2026
9012508
[jwbron/live-eval-producer-staging] Merge remote-tracking branch 'ori…
jwbron Jul 9, 2026
25133b4
[jwbron/live-eval-producer-staging] Merge branch 'jwbron/live-eval-co…
jwbron Jul 9, 2026
0a3d212
[jwbron/live-eval-ab-runner] Merge remote-tracking branch 'origin/jwb…
jwbron Jul 9, 2026
a547972
[jwbron/live-eval-ab-runner] Merge branch 'jwbron/live-eval-producer-…
jwbron Jul 9, 2026
391151b
[jwbron/live-eval-ab-ci] Merge remote-tracking branch 'origin/jwbron/…
jwbron Jul 9, 2026
d2c4c70
[jwbron/live-eval-ab-ci] Merge branch 'jwbron/live-eval-ab-runner' in…
jwbron Jul 9, 2026
082f580
[jwbron/rereview-live-corpus] review: re-review coverage for the eval…
jwbron Jul 9, 2026
48cdc39
[jwbron/rereview-live-corpus] review: the re-review mode sweep and th…
jwbron Jul 9, 2026
284fa40
[jwbron/rereview-live-corpus] review: dispatchable CI home for the re…
jwbron Jul 9, 2026
7f9ae36
[jwbron/rereview-live-corpus] review: manual trigger surface for ever…
jwbron Jul 9, 2026
996766f
[jwbron/live-eval-ab-runner] review: identity short-circuit, gate-fli…
jwbron Jul 10, 2026
fb81be8
[jwbron/live-eval-ab-runner] review: judge economics (Haiku pin, retr…
jwbron Jul 10, 2026
a659be8
[jwbron/live-eval-ab-ci] review: document why the baseline is the bas…
jwbron Jul 10, 2026
b7c3786
[jwbron/live-eval-ab-ci] Merge branch 'jwbron/live-eval-ab-runner' in…
jwbron Jul 10, 2026
541e413
[jwbron/review-trial-skill] Merge branch 'jwbron/live-eval-ab-ci' int…
jwbron Jul 10, 2026
fd42efd
[jwbron/review-out-artifact-upload] Merge branch 'jwbron/review-trial…
jwbron Jul 10, 2026
6d3459a
[jwbron/review-rereview-accountability] Merge branch 'jwbron/review-o…
jwbron Jul 10, 2026
1438a67
[jwbron/review-out-of-lane] Merge branch 'jwbron/review-rereview-acco…
jwbron Jul 10, 2026
98157c2
[jwbron/review-dispatch-tax] Merge branch 'jwbron/review-out-of-lane'…
jwbron Jul 10, 2026
bbb624c
[jwbron/rereview-mode-dial] Merge branch 'jwbron/review-dispatch-tax'…
jwbron Jul 10, 2026
4d1ac62
[jwbron/rereview-live-corpus] Merge branch 'jwbron/rereview-mode-dial…
jwbron Jul 10, 2026
193ae69
[jwbron/live-eval-ab-runner] review: type the judge response via the …
jwbron Jul 10, 2026
7a2065c
[jwbron/live-eval-ab-ci] Merge branch 'jwbron/live-eval-ab-runner' in…
jwbron Jul 10, 2026
00ce9d4
[jwbron/review-trial-skill] Merge branch 'jwbron/live-eval-ab-ci' int…
jwbron Jul 10, 2026
4bd445c
[jwbron/review-out-artifact-upload] Merge branch 'jwbron/review-trial…
jwbron Jul 10, 2026
da1e1db
[jwbron/review-rereview-accountability] Merge branch 'jwbron/review-o…
jwbron Jul 10, 2026
1bc9e08
[jwbron/review-out-of-lane] Merge branch 'jwbron/review-rereview-acco…
jwbron Jul 10, 2026
0123a92
[jwbron/review-dispatch-tax] Merge branch 'jwbron/review-out-of-lane'…
jwbron Jul 10, 2026
64f8115
[jwbron/rereview-mode-dial] Merge branch 'jwbron/review-dispatch-tax'…
jwbron Jul 10, 2026
5631477
[jwbron/rereview-live-corpus] Merge branch 'jwbron/rereview-mode-dial…
jwbron Jul 10, 2026
6a1a8bc
[jwbron/eval-measurement-tool] review: repeat aggregation, --repeats …
jwbron Jul 10, 2026
1dcc7c3
[jwbron/eval-measurement-tool] review: implement the fallback match a…
jwbron Jul 10, 2026
18eda9f
[jwbron/eval-measurement-tool] review: split live-ab report shapes/re…
jwbron Jul 10, 2026
26c8526
[jwbron/eval-measurement-tool] review: checkpoint the repeats artifac…
jwbron Jul 10, 2026
a63648f
[jwbron/eval-measurement-tool] review: publish the measured noise flo…
jwbron Jul 10, 2026
149bf14
[jwbron/eval-measurement-tool] review: raise the drift run's default …
jwbron Jul 10, 2026
52c954d
[jwbron/eval-measurement-tool] review: cover the arbiter, checkpoints…
jwbron Jul 10, 2026
de525c3
[jwbron/eval-measurement-tool] review: eval operator README; drift re…
jwbron Jul 10, 2026
64f5b14
[jwbron/eval-measurement-tool] review: link the eval operator guide f…
jwbron Jul 10, 2026
2369b86
[jwbron/eval-measurement-tool] review: ruler provenance, honest noise…
jwbron Jul 10, 2026
59f78c7
[jwbron/anchor-snap-provenance] review: anchor-snap fallback in the c…
jwbron Jul 10, 2026
3f96569
[jwbron/anchor-snap-provenance] review: pin anchor-snap in a determin…
jwbron Jul 10, 2026
c7ae211
[jwbron/diff-line-annotation] review: line-number-annotated staged di…
jwbron Jul 10, 2026
6b8e3c1
[jwbron/diff-line-annotation] Merge remote-tracking branch 'origin/ma…
jwbron Jul 13, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .changeset/review-annotated-diffs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"review": minor
---

Line-number-annotated staged diffs: remove the anchor mis-counting at the source. The mis-anchor pathology anchor-snap repairs downstream (reviewers counting unified-diff text lines instead of file lines) exists because the staged diff makes the model count; the staging now prints the real number on every content line so anchors are read off the page, never counted. A new deterministic `annotateDiffLineNumbers` (`lib/diff.ts`) prefixes each hunk content line with its line number (`+`/context lines carry the NEW-file RIGHT-side number, `-` lines the OLD-file LEFT-side number) while keeping the diff marker in column one, so annotated text still splits into file sections. The provenance CLI writes `full-stripped-annotated.diff` beside the raw stripped diff, and a new `annotate <in> <out>` subcommand produces `pr-annotated.diff` after Phase 1 builds `pr.diff` (the scoped and flip-gated depths refresh the annotated copies the same way). Every finding-producing reviewer (correctness, skill-auditor, conventions, the four whole-change reviewers, and all eleven specialist lenses via the shared disciplines block) now reads the annotated copy, takes `anchor.line` from the printed number, and strips the prefix when quoting code or authoring a `suggested_patch`. Everything that PARSES a diff keeps reading the raw files: provenance, re-review hunk fingerprints (whose signatures must not shift), scoped staging, and pattern-triage/claim-validator are untouched. The eval stages the annotated siblings for both arms unconditionally, and only a review.md version that names them reads them, so the A/B against a pre-annotation baseline is a pure prompt delta with no staging flag. The measurement instrument rides along: per-case anchor-snap counts (`perCase.snapped`) in the arm report and a pooled "Findings anchor-snapped" row in the aggregate, version-tolerant of older artifacts — if annotation works, candidate-arm snaps fall to zero because anchors arrive correct, with the anchor-snap gate remaining as the deterministic backstop.
9 changes: 8 additions & 1 deletion workflows/review/eval/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -130,7 +130,14 @@ reviewer, was wrong).
prompts carry the rule, both arms snap and the A/B is back to measuring
prompt deltas alone. Snaps are recorded per run (`snappedByProvenance`
in the report's runs; `out/snapped.json` in production artifacts) for
audit. Each record carries the original and snapped anchors, so the
audit, counted per case in the report (`perCase.snapped`), and pooled in
the aggregate ("Findings anchor-snapped"). The snap count is the direct
anchor-fidelity observable: the line-number-annotated staged diffs exist
to drive it to zero at the source, with the snap as backstop. Staging
writes the annotated copies (`pr-annotated.diff`,
`full-stripped-annotated.diff`) for both arms unconditionally; only a
review.md version that names them reads them, so annotation A/Bs are
pure prompt deltas with no staging flag. Each record carries the original and snapped anchors, so the
window class is derivable: a from/to distance within 3 is a near-miss
snap, anything larger is the past-EOF overflow class (the observed
diff-text-counting pathology). Reviewing audited snaps over real PRs is
Expand Down
15 changes: 15 additions & 0 deletions workflows/review/eval/aggregate.ts
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,12 @@ export type SampleRun = {
missedSpecs: {specKey: string; droppedBy?: string}[];
unmatchedPosted: number;
posted: number;
/**
* Findings the provenance gate anchor-snapped (0 for reports predating
* the field). The anchor-fidelity observable: a prompt fix that anchors
* correctly at the source drives this to zero.
*/
snapped: number;
};

/** One arm-run: a single pass of one arm over its cases. */
Expand Down Expand Up @@ -155,6 +161,9 @@ const parseArm = (
missedSpecs,
unmatchedPosted: unmatched,
posted: asNumber(match["postedCount"]),
snapped: Array.isArray(result["snappedByProvenance"])
? result["snappedByProvenance"].length
: 0,
};
});
const judge = raw["judge"];
Expand Down Expand Up @@ -291,6 +300,8 @@ export type ArmAggregate = {
noise: RateStat;
trueMisses: number;
foundButDropped: Record<string, number>;
/** Total anchor-snapped findings across the arm's case-runs. */
snapped: number;
usd: number;
};
/** Mean of per-sample judge means, when any sample carried one. */
Expand Down Expand Up @@ -352,6 +363,7 @@ const aggregateArm = (
let caseRuns = 0;
let unmatched = 0;
let posted = 0;
let snapped = 0;
let usd = 0;
const judgeMeans: number[] = [];

Expand All @@ -374,6 +386,7 @@ const aggregateArm = (
}
unmatched += run.unmatchedPosted;
posted += run.posted;
snapped += run.snapped;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (non-blocking): This PR makes snapped a required field on SampleRun (aggregate.ts:59), but the sampleRun builder in eval/aggregate.test.ts (not touched by this PR) omits it, so this snapped += run.snapped evaluates to NaN and renderAggregateMarkdown renders | Findings anchor-snapped | NaN | ... | for reports built from that fixture. Production reports are safe (parseArm defaults it to 0), so this is test-fixture-only. Add snapped: 0 to the builder's defaults. Relatedly, every fixture exercises only the zero branch, so the populated-count path (parse at aggregate.ts:164 plus this sum) is untested; a fixture with a non-empty snappedByProvenance would cover it.

Lower-confidence observations
  • workflows/review/lib/diff.ts:344- lines are annotated with the OLD-file (LEFT) number, while the "take anchor.line from the printed number — never count" instruction (review.md:1597/1701) isn't qualified for - lines. The schema line ("line is a RIGHT-side number", review.md:1662) mitigates this, but consider stating explicitly that a deletion finding anchors on the adjacent RIGHT-side context/added number, not the - line's printed OLD number — otherwise a side-less deletion anchor can be dropped by the RIGHT-side provenance gate where OLD/NEW diverge.
  • workflows/review/review.md:1702 — the new "strip the NNN| prefix" obligation for quoted code / suggested_patch has no deterministic backstop (unlike anchor-snap for anchors); suggested_patch is validated only as a non-empty string, so a leaked prefix would flow verbatim into a posted suggestion block. Consider a code-level strip in Phase 3.

const spec = (key: string) => {
const s = entry.specs.get(key) ?? {
caught: 0,
Expand Down Expand Up @@ -455,6 +468,7 @@ const aggregateArm = (
noise: rateStat(unmatched, posted),
trueMisses,
foundButDropped,
snapped,
usd,
},
...(judgeMeans.length > 0
Expand Down Expand Up @@ -734,6 +748,7 @@ export const renderAggregateMarkdown = (report: AggregateReport): string => {
`| Misses (true / dropped) | ${dropSummary(
baseline,
)} | | ${dropSummary(candidate)} | |`,
`| Findings anchor-snapped | ${baseline.pooled.snapped} | | ${candidate.pooled.snapped} | |`,
...(baseline.judgeMeanQuality !== undefined &&
candidate.judgeMeanQuality !== undefined
? [
Expand Down
16 changes: 16 additions & 0 deletions workflows/review/eval/live-ab-report.ts
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,13 @@ export type ArmRunReport = {
expected: string;
caught: number;
missed: string[];
/**
* Findings the provenance gate anchor-snapped this run. The direct
* observable for anchor fidelity: a prompt change that fixes
* anchoring at the source (line-number-annotated diffs) shows up
* here as candidate-arm snaps falling to zero.
*/
snapped: number;
/** `<agent>: <reason>` per failed agent (diagnosable from the report). */
failedAgents: string[];
/** Present iff the case is an open-PR (rereview) case. */
Expand Down Expand Up @@ -170,6 +177,10 @@ export const renderMultiMarkdownReport = (report: MultiAbReport): string => {
return lines.join("\n");
};

/** Total anchor-snaps across an arm's case runs (see `perCase.snapped`). */
const snappedTotal = (arm: ArmRunReport): number =>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (non-blocking): snappedTotal and the "Findings anchor-snapped" report row are only ever exercised at zero — produceHit/produceMiss never trigger a gate snap and no test asserts a non-zero total, so a wrong-field or missing-accumulation bug would render 0 and pass. A runArm case whose candidate anchor-snaps (as live-match.test.ts already constructs) asserting perCase[0].snapped === 1 and the report row would close this.

arm.perCase.reduce((sum, c) => sum + c.snapped, 0);

/** `caseId:specKey` -> drop bucket, for every found-but-dropped miss. */
const dropClassByKey = (arm: ArmRunReport): Map<string, string> => {
const map = new Map<string, string>();
Expand Down Expand Up @@ -345,6 +356,11 @@ export const renderMarkdownReport = (report: AbReport): string => {
String(dropClassByKey(baseline).size),
String(dropClassByKey(candidate).size),
),
row(
"Findings anchor-snapped",
String(snappedTotal(baseline)),
String(snappedTotal(candidate)),
),
"",
];

Expand Down
1 change: 1 addition & 0 deletions workflows/review/eval/live-ab.ts
Original file line number Diff line number Diff line change
Expand Up @@ -187,6 +187,7 @@ export const runArm = async (
expected: corpusCase.expected.verdict,
caught: match.caught.length,
missed: match.missed,
snapped: result.snappedByProvenance.length,
failedAgents: produced.perAgent
.filter((a) => a.failed !== undefined)
.map((a) => `${a.name}: ${a.failed}`),
Expand Down
11 changes: 11 additions & 0 deletions workflows/review/eval/live-stage.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,17 @@ describe("stageCase", () => {
expect(read("/stage/context/pr.diff")).toBe(DIFF);
expect(read("/stage/context/full-stripped.diff")).toBe(DIFF);

// The annotated siblings are staged unconditionally; only a
// review.md version that names them reads them, so an A/B against a
// pre-annotation baseline is a pure prompt delta.
const annotated = read("/stage/context/pr-annotated.diff");
expect(annotated).toContain("+ 1| const a = 2;");
expect(annotated).toContain("- 1| const a = 1;");
expect(annotated).toContain(" 2| export {a};");
expect(read("/stage/context/full-stripped-annotated.diff")).toBe(
annotated,
);

const files = JSON.parse(read("/stage/context/files.json"));
expect(files).toEqual([
{path: "src/a.ts", status: "modified", hasPatch: true},
Expand Down
18 changes: 17 additions & 1 deletion workflows/review/eval/live-stage.ts
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,9 @@
* generated files to strip)
* <dest>/context/pr.diff = full.diff (no pattern-triage pass:
* every changed file is a review file)
* <dest>/context/full-stripped-annotated.diff, pr-annotated.diff
* line-number-annotated copies (read by
* review.md versions that name them)
* <dest>/context/files.json path/status/hasPatch per changed file
* <dest>/context/review-files.json = files.json entries (see pr.diff)
* <dest>/context/provenance.json the diff's changed-line map
Expand All @@ -32,6 +35,7 @@ import {
writeFileSync,
} from "node:fs";

import {annotateDiffLineNumbers} from "../lib/diff";
import {computeDiffProvenance} from "../lib/provenance";
import {
buildScopedDiff,
Expand Down Expand Up @@ -192,9 +196,15 @@ const stageRereview = (

if (plan.staging === "new-hunks") {
const scoped = buildScopedDiff(currentDiff, anchorHunks);
const scopedAnnotated = annotateDiffLineNumbers(scoped);
fs.writeFileSync(`${contextDir}/scoped.diff`, scoped);
fs.writeFileSync(`${contextDir}/full-stripped.diff`, scoped);
fs.writeFileSync(`${contextDir}/pr.diff`, scoped);
fs.writeFileSync(
`${contextDir}/full-stripped-annotated.diff`,
scopedAnnotated,
);
fs.writeFileSync(`${contextDir}/pr-annotated.diff`, scopedAnnotated);
}
return plan;
};
Expand Down Expand Up @@ -226,10 +236,16 @@ export const stageCase = (

// The diff surfaces. Corpus diffs carry no generated files, so the
// stripped diff equals the full one; with no pattern-triage pass, the
// review diff does too.
// review diff does too. The annotated siblings are staged for BOTH arms
// unconditionally: only a review.md version that names them reads them,
// so an A/B between a pre-annotation baseline and an annotated candidate
// is a pure prompt delta with no staging flag.
const annotated = annotateDiffLineNumbers(diff);
fs.writeFileSync(`${contextDir}/full.diff`, diff);
fs.writeFileSync(`${contextDir}/full-stripped.diff`, diff);
fs.writeFileSync(`${contextDir}/pr.diff`, diff);
fs.writeFileSync(`${contextDir}/full-stripped-annotated.diff`, annotated);
fs.writeFileSync(`${contextDir}/pr-annotated.diff`, annotated);

// files.json + review-files.json: path/status/hasPatch. `hasPatch` is
// whether the diff carries a section for the path (the completeness
Expand Down
77 changes: 77 additions & 0 deletions workflows/review/lib/diff.test.ts
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
import {describe, it, expect} from "vitest";

import {
annotateDiffLineNumbers,
computeChangedLines,
countOrphanHunkLines,
splitUnifiedDiff,
Expand Down Expand Up @@ -188,6 +189,82 @@ describe("countOrphanHunkLines", () => {
});
});

describe("annotateDiffLineNumbers", () => {
it("prefixes added and context lines with RIGHT-side numbers, removed with LEFT-side", () => {
const annotated = annotateDiffLineNumbers(GIT_DIFF).split("\n");
expect(annotated).toContain(" 10| const a = 1;");
expect(annotated).toContain("- 11| const b = legacy(a);");
expect(annotated).toContain("+ 11| const b = modern(a);");
expect(annotated).toContain("+ 12| const c = b + 1;");
// Hunk 2 resumes at the header's stated positions.
expect(annotated).toContain(" 41| cleanup();");
expect(annotated).toContain("- 41| releaseLock();");
expect(annotated).toContain(" 42| done();");
// The new file numbers from 1.
expect(annotated).toContain("+ 1| export const x = 1;");
expect(annotated).toContain("+ 2| export const y = 2;");
});

it("passes headers, hunk headers, and no-newline markers through verbatim", () => {
const withMarker = [
"diff --git a/a.ts b/a.ts",
"--- a/a.ts",
"+++ b/a.ts",
"@@ -1 +1 @@",
"-x",
"\\ No newline at end of file",
"+y",
"\\ No newline at end of file",
].join("\n");
const annotated = annotateDiffLineNumbers(withMarker).split("\n");
expect(annotated[0]).toBe("diff --git a/a.ts b/a.ts");
expect(annotated[1]).toBe("--- a/a.ts");
expect(annotated[3]).toBe("@@ -1 +1 @@");
expect(annotated[4]).toBe("- 1| x");
expect(annotated[5]).toBe("\\ No newline at end of file");
expect(annotated[6]).toBe("+ 1| y");
});

it("keeps the diff marker in column one, so sections still split", () => {
const annotated = annotateDiffLineNumbers(GIT_DIFF);
expect(splitUnifiedDiff(annotated).map((s) => s.path)).toEqual(
splitUnifiedDiff(GIT_DIFF).map((s) => s.path),
);
});

it("does not annotate text after a hunk's stated extent (trailing lines)", () => {
const trailing = [
"--- a/a.ts",
"+++ b/a.ts",
"@@ -1 +1 @@",
"-x",
"+y",
"",
].join("\n");
const annotated = annotateDiffLineNumbers(trailing).split("\n");
// The trailing empty line (a split artifact, outside the hunk's
// counted extent) stays empty instead of gaining a phantom number.
expect(annotated[annotated.length - 1]).toBe("");
});

it("widens the number column for large files and is empty-safe", () => {
const big = [
"--- a/a.ts",
"+++ b/a.ts",
"@@ -9998,3 +9998,3 @@",
" keep;",
"-old;",
"+new;",
" tail;",
].join("\n");
const annotated = annotateDiffLineNumbers(big).split("\n");
expect(annotated).toContain(" 9998| keep;");
expect(annotated).toContain("- 9999| old;");
expect(annotated).toContain("+ 9999| new;");
expect(annotateDiffLineNumbers("")).toBe("");
});
});

describe("stripDiffFiles", () => {
it("removes the named files' sections and keeps the rest verbatim", () => {
const stripped = stripDiffFiles(GIT_DIFF, new Set(["src/new.ts"]));
Expand Down
71 changes: 71 additions & 0 deletions workflows/review/lib/diff.ts
Original file line number Diff line number Diff line change
Expand Up @@ -286,3 +286,74 @@ export const stripDiffFiles = (
.map((section) => section.text);
return kept.join("\n");
};

/**
* Annotate a unified diff with explicit line numbers, for the model-read
* staged copies (`full-stripped-annotated.diff`, `pr-annotated.diff`).
*
* The observed anchor pathology is reviewers counting diff TEXT lines
* instead of file lines; printing the real number on every line removes the
* counting entirely. Format, per hunk content line:
*
* `+ 16| added line` RIGHT-side (new file) line number
* ` 17| context line` RIGHT-side (new file) line number
* `- 12| removed line` LEFT-side (old file) line number
*
* The diff marker stays in column one, so annotated text still splits into
* file sections ({@link splitUnifiedDiff} recognises the same headers and
* the same `+`/`-`/space first columns). Headers, hunk headers, and
* `\ No newline` markers pass through untouched. The annotated copy is for
* model eyes only: every code parser (provenance, fingerprints, scoped
* staging) keeps reading the raw diff.
*/
export const annotateDiffLineNumbers = (diff: string): string => {
const out: string[] = [];
let oldLine = 0;
let newLine = 0;
/** Remaining old/new line counts of the hunk being consumed. */
let hunkOld = 0;
let hunkNew = 0;
let width = 3;

for (const line of diff.split("\n")) {
const hunk = /^@@ -(\d+)(?:,(\d+))? \+(\d+)(?:,(\d+))? @@/.exec(line);
if (hunk !== null) {
oldLine = Number(hunk[1] ?? "1");
newLine = Number(hunk[3] ?? "1");
hunkOld = Number(hunk[2] ?? "1");
hunkNew = Number(hunk[4] ?? "1");
width = Math.max(
width,
String(Math.max(oldLine + hunkOld, newLine + hunkNew)).length,
);
out.push(line);
continue;
}
if (hunkOld <= 0 && hunkNew <= 0) {
// Outside hunk content (file headers, preamble, trailing text):
// pass through verbatim. The countdown, not a flag, decides —
// so a `--- ` header after an exhausted hunk is never mistaken
// for a removed line.
out.push(line);
continue;
}
if (line.startsWith("+")) {
out.push(`+${String(newLine).padStart(width)}| ${line.slice(1)}`);
newLine++;
hunkNew--;
} else if (line.startsWith("-")) {
out.push(`-${String(oldLine).padStart(width)}| ${line.slice(1)}`);
oldLine++;
hunkOld--;
} else if (line.startsWith("\\")) {
out.push(line);
} else {
out.push(` ${String(newLine).padStart(width)}| ${line.slice(1)}`);
oldLine++;
newLine++;
hunkOld--;
hunkNew--;
}
}
return out.join("\n");
};
Loading
Loading