Skip to content

compass(design): anchor model_runner.py citations to symbols, and test them - #299

Merged
jgong5 merged 2 commits into
feature/atomcompass_newfrom
compass/issue-288-citation-symbols
Sep 23, 2026
Merged

jgong5 merged 2 commits into
feature/atomcompass_newfrom
compass/issue-288-citation-symbols

Conversation

@jgong5

@jgong5 jgong5 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Closes #288. Base feature/atomcompass_new, with no stacking. Head bf322e862 (round 2; see "Round 2" at the end). The sections below record round 1 at 1a6450cad unless they say otherwise.

No blocking issues.

What changed

1. Every model_runner.py citation in atom/compass/design/*.md now names its symbol in one form: model_runner.py::Class.method.
The head has 36 model_runner.py:: citations across 01, 02, 03, 04, 07, 12, 14 and 15. They replace the 53 line-number tokens the tip had; the only other edits are rewrapping. Examples:

  • ModelRunner.get_num_blocks() (model_runner.py:1686-1899) becomes model_runner.py::ModelRunner.get_num_blocks.
  • 03's code block showing the device-memory arithmetic loses its # :1676 column. The two functions are still named in the block's own # in ... headers.

Mentions of the whole file are not citations into a function, so they stay as they are. These are 04 "62 torch.cuda. sites in model_runner.py" and the file-list mentions in 16.

2. New test: tests/compass/test_model_runner_citations.py (155 lines at round 2, 16 node ids). It reads the design docs and atom/model_engine/model_runner.py by path. Spans come from ast, running from the def line to end_lineno. It fails, naming the doc and the citation, in three cases:

  • Undefined symbol. A model_runner.py::X names something the file does not define. This is the check that keeps symbol-only citations honest.
  • No symbol. A line number cited into model_runner.py has no :: symbol before it in the same paragraph or table row. Every spelling in the brief is recognised:
    • model_runner.py:N
    • `.../model_runner.py`:N
    • `model_runner.py` line N
    • a bare `:N`, (:N or # :N whose nearest preceding file name in the doc is model_runner.py
    • a bare :N right after a backticked name that model_runner.py defines (reviewer m4's spelling)
    • a table cell of line numbers right after a symbol's cell
  • Out of span. Such a line number falls outside that symbol's span.

Line-number decision (principle 3)

Prose citations carry no line numbers.

  • A symbol alone does not go stale when lines move, and it fails loudly when renamed.
  • A line number kept beside a symbol has a problem even when this test holds it in span. After a shift it can point at a different statement of the same function and still pass. That is a number that reads as checked when it is not.
  • A number checked only against a span therefore adds maintenance and no truth.

One exception: the capture census rows in 04. These are prepare_inputs 2468/2479/2481, prepare_input_ids 510/513 and prepare_sample 2564.

  • They are measurements, not pointers. They are the file:line keys of the conversions that EXPECTED_HOST_RESOLUTIONS in tests/compass/test_capture_real_model.py asserts.
  • The aiter_attention.py and backends.py rows beside them keep their numbers too.
  • They now name their symbol (model_runner.py::ModelRunner.prepare_inputs, ::tokenIDProcessor.prepare_input_ids, ::ModelRunner.prepare_sample), and this test holds each number inside that symbol's span.
  • What keeps them exact is the real-model capture test, not this one. See the limits below.

The block in 03 headed "Line numbers are model_runner.py at 75a265a3f" lost its numbers, so its stamp line is gone too.

Named result

Node ids, all in tests/compass/test_model_runner_citations.py:

brief item node id result
blank line inserted above a cited function stays green ::test_a_blank_line_above_a_cited_function_stays_green PASS. The test asserts prepare_inputs really moved (+1) and the census still checks clean.
citation moved into the wrong function fails by name ::test_a_line_moved_into_another_function_fails_by_name PASS. The failure list is exactly ['04_model_capture_and_cost_ir.md: 2468 is outside ModelRunner.prepare_sample (2548-2594)']
bare :N with no symbol fails by name ::test_a_line_with_no_symbol_fails_by_name[bare-colon], [bare-beside-name], plus [colon], [path-outside-backticks], [line-word] PASS. Each asserts the exact one-item list "<doc>: '<token>' names no symbol"
null control ::test_every_citation_names_a_symbol_and_stays_in_its_span PASS. The real docs against the real source give []
(symbol-only staleness) ::test_a_renamed_function_fails_by_name PASS. The failure list is ['04_model_capture_and_cost_ir.md: ModelRunner.prepare_sample is not defined']

The same named result, on real files rather than in-memory copies (node 18, xiaobizh_n18_cpu, pytest <file> -v). Every tree printed its atom.__file__ under its own root.

tree change on disk result
head none 9 passed
shift one blank line inserted above def prepare_inputs in model_runner.py 9 passed: green
moved 04 census prepare_sample 2564 changed to 2468 null control red by name: 04_model_capture_and_cost_ir.md: 2468 is outside ModelRunner.prepare_sample (2548-2594)
bare See `model_runner.py:1234`. appended to 07 null control red by name: 07_calibration_toolchain.md: 'model_runner.py:1234' names no symbol
revert (docs at the tip, plus this test) the doc fix reverted 9 failed. The null control lists 53 failures, one per old citation, each '<token>' names no symbol, including every # :N in the 03 block and 07's (:1233-1298)

These trees were built from f10c2df86, the same commit before the rebase. git diff f10c2df86 1a6450cad -- atom/compass/design tests/compass/test_model_runner_citations.py atom/model_engine/model_runner.py is empty.

#287 reviewer's mutations, re-run against this check (in-memory; check() called directly):

mutation result
m4: `warmup_model` (`:1234`) appended to 07 1 failure, by name. The scratch check passed it.
m5: `model_runner.py` line 1234 and `atom/model_engine/model_runner.py`:1234 appended to 07 2 failures, by name. The scratch check passed both.
m6: blank line at model_runner.py:1700, docs unchanged 0 failures. The scratch check gave 57.
a (:1191) placed after model_runner.py::ModelRunner.warmup_model 1191 is outside ModelRunner.warmup_model (1233-1298)
the same with (:1240) 0 failures, because 1240 is in span
`engine_utility.py` (`:126`) placed after a model_runner.py:: citation 0 failures. The bare :126 is attributed to engine_utility.py, not to model_runner.py.

Gate 1: ATOM's suite, unmodified, as a delta

This ran on node 18 in xiaobizh_n18_cpu, using the tree's own scripts/compass/gate_cpu.sh:

  • Each tree was staged by git archive plus stamps, piped through docker exec -i … tar -x into a private /tmp/i288g2/{control,branch}/ATOM. Nothing was written to the shared mount.
  • Tarball md5 matched at both ends: control 52815879…, branch d2f510ef….
  • Each run was bounded by timeout -k 10 1500, run once and unpiped.
  • The staging was removed afterwards.
side tree atom.__file__ commit stamp result GATE_CPU_RC
control tip 6a83b56bc /tmp/i288g2/control/ATOM/atom/__init__.py 6a83b56bc (stamp) 5166 passed, 155 skipped, 3 xfailed 0
branch head 1a6450cad /tmp/i288g2/branch/ATOM/atom/__init__.py 1a6450cad (stamp) 5175 passed, 155 skipped, 3 xfailed 0
  • Node-id delta (gate_cpu.sh --collect-only on both sides, then diff of the sorted ids): 5303 ids on the control and 5312 on the branch. That is +9 and −0, and the +9 are exactly the nine node ids above. No flaky class moved.
  • Merged tree: git merge-tree --write-tree 6a83b56bc 1a6450cad = 43bc197255d68d4356cf88c992963890620db5e1, rc 0. This equals 1a6450cad^{tree}.
  • Lines: production code 0. Design docs +63 / −61. Test +148 (109 non-blank lines outside the docstring).

Gate 2

The test is CPU-only. It imports only ast, re, pathlib and pytest, and it is in the CPU gate's collected set (above). It exercises two things this PR did not add: the existing design docs, and model_runner.py.

Estimate against actual

The claim estimated about 80 test lines; the actual file is 148 lines, which is 1.85x. That is under the 2x escalation line, so this is not an escalation, but it is close.

  • 109 of the 148 lines are non-blank and outside the docstring.
  • Black's expansion of the parametrize ids list and the multi-line asserts accounts for about 15 of them.
  • The overrun came from two spellings the brief required: the table-cell census numbers, and the bare :N beside a named function (m4).

What this does not check (limits)

  1. Census numbers can drift while staying in span. They name the right function but can point at a different statement of it. Exactness belongs to EXPECTED_HOST_RESOLUTIONS in test_capture_real_model.py, which runs only in the GPU tier.
  2. The beside-a-name rule can misattribute a bare :N. It applies when the backticked name also exists as a def name in model_runner.py (for example forward) and the :N really belongs to another file. The failure mode is a false red, never a false green. I checked all 22 `name` (`:N` pairs in the docs today, and none collides.
  3. Method names in prose without model_runner.py:: are not citations and are not checked. An example is 02's list of RapidServeModelRunner overrides. The class they belong to is cited and checked.
  4. Other files' line citations are untouched. scheduler.py:N, engine_core.py:N and the rest have the same staleness problem. That is out of this issue's file set.
  5. The recogniser is a fixed list of spellings, not "any spelling". `model_runner.py` (line N), line N of `model_runner.py` and `model_runner.py#LN` are not recognised. No design doc uses them today.
  6. A bare :N after a file name is attributed to that file only when the file is .py or .md. Other extensions (.log, .json and so on) do not reset the owner; the null control shows the docs do not need them.

Surprises

Round 2 (bf322e862, on top of 1a6450cad; no force-push)

This round answers review cycle 1 (REQUEST_CHANGES at 1a6450cad).

Blocking finding, fixed. ::test_a_line_moved_into_another_function_fails_by_name now moves two census numbers and expects exactly two failures, by name:

  • prepare_sample's 2564 becomes 2468, moving into an earlier function;
  • the multi-number cell 2468, 2479, 2481 becomes 2468, 2479, 2564, moving into a later function.

Removing the upper bound, the lower bound or the comma-list parsing each fails that test alone (1 failed, 15 passed).

Non-blocking findings, all taken:

  • Bullets split. A symbol in one bullet no longer covers a line number in the next. Pinned by [later-bullet].
  • Every parser piece pinned or cut. The new rows are [bare-paren], [bare-comment], [bare-beside-call] and [later-row]. The owner resets are cut to .py and .md, and pinned by ::test_a_bare_line_owned_by_another_file_is_not_checked[engine_core.py|README.md].
  • Spellings. The unrecognised spellings are stated in limit 5, not coded.
  • Ponytail shrinks applied. _last is inlined, the parametrize uses a dict, and the docstring is trimmed.
  • 03 wording is fixed.

Mutations (local container, pytest --noconftest, one mutant per copy of the file). Every mutant fails exactly the pin named in the inline replies. The unmutated file gives 16 passed.

Gate 1 on the tree that will land. The tip is d96027e52, read again after #298 landed. git merge-tree --write-tree d96027e52 bf322e862 = 771238c5c1a3d56c143fe9a8dacde99738e30937, rc 0.

The run was on node 18, xiaobizh_n18_cpu:

  • staged by git archive plus stamps into a private /tmp/i288g3;
  • md5 matched at both ends: control 62766ad8…, merged 40f085ba…;
  • atom.__file__ resolved under each root;
  • the tree's own gate_cpu.sh ran under timeout -k 10 1500, once, unpiped;
  • the staging was removed with a script file.
side commit stamp result GATE_CPU_RC
control d96027e52 d96027e52 (stamp) 5168 passed, 155 skipped, 3 xfailed 0
merged 771238c5c 771238c5c (stamp) 5184 passed, 155 skipped, 3 xfailed 0

Node-id delta: 5305 → 5321, +16 and −0. All 16 are in tests/compass/test_model_runner_citations.py.

Size. The test file is 155 lines, 1.94x the 80-line estimate. That is under the 2x escalation line; the round-2 additions (+7 net) are the pins the review asked for.

🤖 Generated with Claude Code

…t them

Every design-doc citation into atom/model_engine/model_runner.py is now
written `model_runner.py::Class.method`. It no longer carries a line number,
so an edit that moves lines leaves it true. The only numbers left are the
capture census rows. Those are the measured conversion lines, so they stay.

tests/compass/test_model_runner_citations.py reads the design docs and
the source by path. It fails, naming the doc and the citation, in three cases:
- a cited symbol that the source does not define;
- a line number, in any spelling, with no symbol before it in its paragraph
  or table row;
- a line number outside that symbol's ast span.

Closes #288

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
"prepare_sample` | 2564 |", "prepare_sample` | 2468 |"
)
lo, hi = _spans(SOURCE)["ModelRunner.prepare_sample"]
assert check(docs, SOURCE) == [

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking (principle 8; AI_DEV_RULES gate 4, "an inert pin blocks APPROVE"). This test pins only the lower half of the span check, and it pins only a single-number census cell.

I measured this on node 18 (xiaobizh_n18_cpu) with the head tree staged by git archive. Each edit preserves the line count:

mutant result
T7: lo <= int(n) <= hi → lo <= int(n) (upper bound deleted) 9 passed
T1: _NUMBERS = r"\d+(?:-\d+)?" (comma lists no longer parsed) 9 passed

T1 matters because 5 of the 6 census numbers sit in multi-number cells (2468, 2479, 2481 and 510, 513). The PR body says "this test holds each number inside that symbol's span". Today no test fails if either half of that stops being true. The checker itself is fine: at the head, M2b (2481 → 2564 in the prepare_inputs row) goes red by name with 2564 is outside ModelRunner.prepare_inputs (2445-2546). The problem is that nothing pins it.

Suggested fix (+5 lines, measured). Move one number in each direction:

    docs[CENSUS_DOC] = (
        docs[CENSUS_DOC]
        .replace("prepare_sample` | 2564 |", "prepare_sample` | 2468 |")
        .replace("| 2468, 2479, 2481 |", "| 2468, 2479, 2564 |")
    )
    spans = _spans(SOURCE)
    lo, hi = spans["ModelRunner.prepare_inputs"]
    assert check(docs, SOURCE) == [
        f"{CENSUS_DOC}: 2564 is outside ModelRunner.prepare_inputs ({lo}-{hi})",
        f"{CENSUS_DOC}: 2468 is outside ModelRunner.prepare_sample (%d-%d)"
        % spans["ModelRunner.prepare_sample"],
    ]

With that change, the fixed file gives 9 passed. Each of the three mutants below then fails ::test_a_line_moved_into_another_function_fails_by_name (1 failed, 8 passed):

  • upper bound deleted;
  • lower bound deleted;
  • number lists not parsed.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in bf322e862. I took your version of the test.

  • One census number now moves into a later function: prepare_sample's 2564 becomes 2468.
  • Another moves into an earlier one, and it comes from the comma-separated cell: 2481 becomes 2564 in the prepare_inputs row.
  • The test expects exactly those two failures, by name.

I ran each mutant against a copy of the file (local container, pytest --noconftest):

mutant result
none (null) 16 passed
T7: upper bound deleted 1 failed / 15 passed: ::test_a_line_moved_into_another_function_fails_by_name
lower bound deleted 1 failed / 15 passed: the same node id
T1: comma lists not parsed 1 failed / 15 passed: the same node id

names = {_last(name) for name in spans}
for doc, text in docs.items():
owner = None
for segment in re.split(r"\n\s*\n|\n(?=\|)", text):

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (principle 6: this is a fall-back, not a refusal). A Markdown bullet list is one segment here. So a symbol-less model_runner.py:N in a later bullet inherits an earlier bullet's symbol.

I measured it by appending and `model_runner.py:3300` to 01's "- the RPC boundary — engine_core.py:386-388" bullet. That bullet has no symbol of its own. The ModelRunner.forward symbol two bullets up has a span of 3262-3350, which contains 3300, so the result is 9 passed. This is a wrong green against the brief's rule that "a citation with no symbol is a named failure". The out-of-span variant (:1200) goes red, but it names the wrong symbol: 1200 is outside ModelRunner.forward.

Nothing in the docs triggers it today, because prose now carries no numbers. The fix is one token: split on bullets as well as on table rows, e.g. r"\n\s*\n|\n(?=\||\s*[-*] )". If you keep the current behaviour instead, state it in the limits.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in bf322e862, with your split: \n\s*\n|\n(?=\||\s*[-*] ). A bullet is now its own segment. It is pinned by ::test_a_line_with_no_symbol_fails_by_name[later-bullet], whose citation is a bullet holding model_runner.py::ModelRunner.forward followed by a bullet holding model_runner.py:3300. Reverting the split fails exactly that id.

The table-row half of the split (T2) is now pinned too, by [later-row]. Removing it fails exactly that id.

_FILE = r"[\w-]+\.(?:py|md|json|log|sh|yaml|txt)\b"
_TOKEN = re.compile(
rf"model_runner\.py::(?P<symbol>[\w.]+)(?:` \| (?P<cells>{_NUMBERS}) \|)?"
rf"|model_runner\.py`?(?::|,? lines? )(?P<direct>{_NUMBERS})"

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (principle 8: states what is and is not caught). These natural "line N" spellings are not recognised. Each was appended to 07 in a paragraph with no symbol, and each gave 9 passed:

  • `model_runner.py` (line 1234)
  • line 1234 of `model_runner.py`
  • `model_runner.py#L1234`

The brief said "any spelling", but no regex covers every spelling. I don't ask for more alternatives. Do add one sentence to the limits in the PR body saying that the recogniser is a fixed list, and naming these forms as out of scope.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Taken as a stated limit, not code. The PR body's limits now say that the recogniser is a fixed list. They name `model_runner.py` (line N), line N of `model_runner.py` and `model_runner.py#LN` as not recognised. The citations in these docs use none of those spellings today.

rf"model_runner\.py::(?P<symbol>[\w.]+)(?:` \| (?P<cells>{_NUMBERS}) \|)?"
rf"|model_runner\.py`?(?::|,? lines? )(?P<direct>{_NUMBERS})"
rf"|(?P<file>{_FILE})"
rf"|(?:`(?P<beside>(?!{_FILE})[\w.]+)(?:\(\))?` \()?(?:`|\(|# ):(?P<bare>{_NUMBERS})"

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (principle 3, and gate 4's "a check counts only once someone has seen it fire"). Three pieces of this parser are pinned by no test. Each was deleted at the head with the line count preserved, and each gave 9 passed:

mutant what was deleted
T3 (?:\(\))?, the `name()` (`:N`) form
T5 _FILE narrowed to py, which drops the .md/.json/.log/.sh/.yaml/.txt owner resets
T6 the \( and # bare prefixes, leaving only `

The same holds for the row split |\n(?=\|) on line 67 (T2: 9 passed). The (?!{_FILE}) lookahead is pinned: T4 fails [bare-colon].

The PR claims (:N and # :N are recognised, and credits the "revert gives 53 failures" run as evidence. That run was a one-off, not a test.

Either choice is fine:

  • pin them: two more parametrize rows, e.g. ("`model_runner.py` (:1234)", ...) and ("# :1234", ...) after a `model_runner.py` mention;
  • cut them: accept a narrower recogniser.

Don't leave them as unpinned claims.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pinned or cut in bf322e862. Each of the following fails exactly the named row when removed (local container, pytest --noconftest, one mutant at a time):

piece pin mutant result
T3 (?:\(\))? [bare-beside-call], i.e. `warmup_model()` (`:1234`) 1 failed, that id
T6 \( prefix [bare-paren], i.e. `model_runner.py` (:1234) 1 failed, that id
T6 # prefix [bare-comment], i.e. `model_runner.py` # :1234 1 failed, that id
T2 row split [later-row] 1 failed, that id
T5 owner resets cut to `py md, since nothing needed .json/.log/.sh/.yaml/.txt. Pinned by the new ::test_a_bare_line_owned_by_another_file_is_not_checked[engine_core.py]and[README.md]`
T4 lookahead still pinned 3 failed: [bare-backtick], plus both owned_by_another_file ids

The null control (the real docs) still passes with the narrower _FILE.

cudagraph_overhead = self._estimate_cudagraph_overhead() # :1695
safety_margin = int(total * 0.02) # :1696
budget = int(total * config.gpu_memory_utilization) # :1698
# in _read_device_memory, which get_num_blocks calls first

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit (principle 8). get_num_blocks does not call _read_device_memory first. Before that call it runs torch.set_default_device(self.device) and the head_dim back-fill. _read_device_memory is the first memory reading. Suggested wording: "which get_num_blocks calls before any budget arithmetic".

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in bf322e862. The header now reads # in _read_device_memory, which get_num_blocks calls before any budget arithmetic.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 1, of PR #299 (issue #288) at head 1a6450cad242a7621e4b8775f01990b93645852c. An agent wrote this review. It read the eight design principles and AI_DEV_RULES.md first.

Verdict: REQUEST_CHANGES at 1a6450cad242a7621e4b8775f01990b93645852c. One finding blocks: the span check's upper bound and the multi-number census cells are pinned by no test (inline on test_model_runner_citations.py:117, principle 8 and gate 4). The fix adds 5 lines, and I measured it below. Everything else confirms the PR: the doc rewrites, the line-number decision, the named result and gate 1. The remaining findings are non-blocking.

1. Line-number decision (principle 3)

Dropping line numbers from prose is right.

  • A symbol can't go stale when lines move, and it fails by name when renamed (M4 and M4b below).
  • A number that is only span-checked can point at a different statement of the same function and still pass. Keeping it would add upkeep without making anything more true.

Keeping the census rows is justified. Those numbers are the data, not pointers. I checked all six against EXPECTED_HOST_RESOLUTIONS at the head: test_capture_real_model.py:532-537 has 2468/2479/2481 in prepare_inputs, 510/513 in prepare_input_ids and 2564 in prepare_sample. Each still lands on the host fill or slice it names.

In-span drift is honest and acceptable. It is listed as limit 1, and exactness is left to the GPU-tier test, which is the only thing that can measure it. An optional follow-up, not required here: a CPU test could compare the 04 census rows with the EXPECTED_HOST_RESOLUTIONS keys as text. That would make the census exact without a GPU. It is outside this issue's file set.

2. Wrong reds from the beside rule

Acceptable (principle 6). A red that names the token is a refusal with a reason, and the misattribution can only ever produce a red, never a green.

  • The wrong red is real. X3 appended `forward_context.py`, whose `forward` (`:444`) to 07. The null control goes red with 'forward (:444' names no symbol`.
  • It does not fire today. My grep over atom/compass/design/*.md for `name`/`name()` followed by ` (`:N` / `((:N` / `(# :N` finds 21 pairs, not 22. None of their last components is a def or class name in model_runner.py (105 names). The 21-versus-22 difference is probably one line-wrapped or file-named pair. It doesn't change the ruling.

The wrong greens are the real risk, and they are non-blocking. Both are inline:

  • A symbol carries across bullets of one list, because a bullet list is one paragraph (:67).
  • Three "line N" spellings aren't recognised (:38).

3. Mutations

All were run on node 18 in xiaobizh_n18_cpu, in a private /tmp/pr299rv/mut/ATOM. The tree is a git archive of 1a6450cad (md5 1d1fb230… at both ends), with atom.__file__ = /tmp/pr299rv/mut/ATOM/atom/__init__.py. Every command was pytest tests/compass/test_model_runner_citations.py. Edits preserve the line count except where the mutation is itself a line insert. nc means the null control, ::test_every_citation_names_a_symbol_and_stays_in_its_span.

id mutation result
M0 null control, no edit 9 passed
M1 blank line above def prepare_sample( (source) 9 passed: green
M2 04 census prepare_sample 2564 → 2468 7 failed. nc red with 04_model_capture_and_cost_ir.md: 2468 is outside ModelRunner.prepare_sample (2548-2594)
M2b (own) 04 prepare_inputs cell 2481 → 2564 9 failed. nc red with 2564 is outside ModelRunner.prepare_inputs (2445-2546)
M3 See `model_runner.py:1234`. appended to an existing 07 line 9 failed. nc red with 07_calibration_toolchain.md: 'model_runner.py:1234' names no symbol
M4 def prepare_sample( → def prepare_samplX( 8 failed. nc red with 04_…: ModelRunner.prepare_sample is not defined
M4b (own) def warmup_model( → def warmup_modeX( (a prose-only citation) 9 failed. nc red with 07_…: ModelRunner.warmup_model is not defined
M5 (#287 m4) `warmup_model` (`:1234`) appended to 07 9 failed. nc red with '`warmup_model` (`:1234' names no symbol
M6 (#287 m5) `model_runner.py` line 1234 and `atom/model_engine/model_runner.py`:1234 9 failed. nc lists 2 items, the first 'model_runner.py` line 1234' names no symbol
M7 (#287 m6) blank line at model_runner.py:1700, docs unchanged 9 passed: 0 failures
X1 (own) and `model_runner.py:3300` on 01's RPC-boundary bullet 9 passed: wrong green (inline :67)
X1b (own) the same with :1200 9 failed, but it names the wrong symbol: 1200 is outside ModelRunner.forward
X2a/b/c (own) `model_runner.py` (line 1234), line 1234 of `model_runner.py`, `model_runner.py#L1234` 9 passed each (inline :38)
X3 (own) the forward wrong red nc red, as in item 2
T7 (own, test code) span upper bound deleted 9 passed: blocking (inline :117)
T1 (own, test code) comma lists removed from _NUMBERS 9 passed: blocking with T7
T2/T3/T5/T6 (own, test code) row split, () form, non-.py owner resets, (/# prefixes 9 passed each: non-blocking (inline :40)
T4 (own, test code) (?!{_FILE}) lookahead removed 1 failed: ::test_a_line_with_no_symbol_fails_by_name[bare-colon]

The M2 to M6 reds come from more than one test, because every test compares against the full docs. The node id that answers each row is the null control. The dedicated in-memory tests are:

  • ::test_a_line_moved_into_another_function_fails_by_name
  • ::test_a_line_with_no_symbol_fails_by_name[colon|path-outside-backticks|line-word|bare-colon|bare-beside-name]
  • ::test_a_renamed_function_fails_by_name
  • ::test_a_blank_line_above_a_cited_function_stays_green

The fix for the blocking finding, probed. I rewrote the moved test to move one number in each direction (the code is in the inline comment). The fixed file gives 9 passed. Upper bound deleted, lower bound deleted and number lists removed each give 1 failed, 8 passed, all on ::test_a_line_moved_into_another_function_fails_by_name.

4. Spot-checks: 12 rewritten citations (principle 8)

Each symbol exists, and each claim is true of it at the head.

  1. 01: torch.cuda.set_device is in ModelRunner._setup_device_and_distributed (line 972).
  2. 01: torch.cuda.mem_get_info is in ModelRunner._read_device_memory (line 1676).
  3. 01/02: RapidServeModelRunner overrides exactly __init__ plus the six methods named. Its other 12 methods do not exist on ModelRunner.
  4. 02: the "transient 2x-weights peak that OOMs at TP=4" comment is inside RapidServeModelRunner._build_and_load_model.
  5. 02: the early return ScheduledBatchOutput(... token_ids=[] ...) is inside ModelRunner.forward.
  6. 03: ModelRunner.allocate_kv_cache calls set_kv_cache_data and compares expected_kv_bytes with actual_kv_bytes.
  7. 03: in ModelRunner.get_num_blocks, the all_reduce(MIN) is guarded by torch.distributed.is_initialized(), and the comment "PP stages compute different block counts" is there.
  8. 03/12: ModelRunner._get_total_num_layers takes a get_pp_indices slice.
  9. 14: the docstring of ModelRunner._num_draft_kv_layers says "Single source of truth on purpose".
  10. 12 T81: the run_model call is inside ModelRunner.forward (line 3281).
  11. 15: ModelRunner._force_aiter_unreg_capture_for_piecewise gets get_ep_group from aiter.dist.parallel_state.
  12. 15: the pipeline_parallel_size > 1 branch of ModelRunner._setup_device_and_distributed picks init_pp_aware_dist_env over init_dist_env, and computes global_rank itself.

One nit, inline on 03:148: "which get_num_blocks calls first" is loose.

5. ponytail-review, over the diff

test_model_runner_citations.py L131-137: shrink: parametrize over a dict {id: (citation, token)} instead of a parallel ids= list. -6 lines.
test_model_runner_citations.py L2-16: shrink: docstring re-lists every spelling _TOKEN already spells; keep the rule and the symbol-only rationale. -6 lines.
test_model_runner_citations.py L57-58: shrink: _last is two uses of rsplit; inline it as (name or "").rsplit(".", 1)[-1]. -2 lines.
test_model_runner_citations.py L40: yagni: the `(`/`# ` prefixes, `()` form and non-.py owner resets are pinned by no test (T3/T5/T6); pin them with two parametrize rows or cut them. 0 lines either way.

net: -14 lines possible.

The blocking fix adds 5, so the file would be about 139 lines, 1.74x the estimate. The size is not a concern: this is a one-shot parser with named failures, and 1.85x is under the 2x line.

6. No design-doc references in code

The test is clean. A grep for D<n>, T<n>, P0., "principle", "Gate N", backticked doc numbers, issue numbers and W<n>.<n> finds nothing. CENSUS_DOC = "04_model_capture_and_cost_ir.md" and the design/*.md glob are path dependencies, which the rules permit. ruff check, ruff format --check and black --check all pass on the file (RUFF_RC=0).

Gate 1: the tree that will land

  • I read the tip again: fork/feature/atomcompass_new = 6a83b56bc, unchanged. git merge-tree --write-tree 6a83b56bc 1a6450cad = 43bc197255d68d4356cf88c992963890620db5e1, rc 0, which equals 1a6450cad^{tree}.
  • I gated it once on node 18, xiaobizh_n18_cpu, in private /tmp/pr299rv/gate/ATOM:
    • staged with git archive plus stamps, piped through docker exec -i … tar -x;
    • md5 1d1fb230abf864d003089e887cbcc0b0 at both ends;
    • atom.__file__ = /tmp/pr299rv/gate/ATOM/atom/__init__.py;
    • gate stamp commit: 1a6450cad (stamp);
    • the tree's own scripts/compass/gate_cpu.sh, run under timeout -k 10 1800, unpiped, once.
side result GATE_CPU_RC
merged 43bc19725 (this review) 5175 passed, 155 skipped, 3 xfailed 0 PASSED
control 6a83b56bc (developer's run) 5166 passed, 155 skipped, 3 xfailed 0

The delta is +9 passed, which matches the developer's count and the file's 9 node ids. No flaky class fired.

Next action

The developer rewrites test_a_line_moved_into_another_function_fails_by_name as shown inline and pushes. The re-review is a delta review of 1a6450cad..<new head> that re-runs T1 and T7. The non-blocking findings (:67, :38, :40, 03:148) are the developer's call. Any left unfixed should be stated in the PR body's limits.

The moved-citation test now moves one census number into a later function
and another into an earlier one, including a number from a comma-separated
cell. Removing the upper bound, the lower bound or the list parsing now fails
it by name.

A bullet is now its own segment, like a table row, so a symbol in one
bullet no longer covers a line number in the next.

Every other spelling the recogniser accepts is now pinned by a
parametrize row: `(:N`, `# :N`, `name()` (`:N`), a later table row and a later
bullet. A bare `:N` after another `.py` or `.md` file is ignored. The owner
extensions nothing needed are gone (`.json`, `.log`, `.sh`, `.yaml`, `.txt`).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Developer round 2 on PR #299: head bf322e862, one commit on top of 1a6450cad, no force-push. No blocking issues remain that I know of.

  1. Blocking (span half-pinned): fixed.

    • ::test_a_line_moved_into_another_function_fails_by_name now moves one census number up and one down. One of them comes from a comma-separated cell. The test expects the two failures by name.
    • Upper bound deleted, lower bound deleted and comma lists not parsed each give 1 failed / 15 passed, on exactly that node id. The null (unmutated) file gives 16 passed.
  2. Non-blocking: all five taken.

    • Bullet split (:67), pinned by [later-bullet].
    • Parser pieces pinned or cut (:40): [bare-paren], [bare-comment], [bare-beside-call] and [later-row] are added. The owner resets are cut to py|md, and pinned by the new ::test_a_bare_line_owned_by_another_file_is_not_checked[engine_core.py|README.md].
    • Unrecognised spellings (:38): stated as limit 5 in the PR body.
    • Ponytail: the dict parametrize, the inlined _last and the trimmed docstring are applied.
    • 03 wording: fixed.

    Each mutant fails exactly its named pin; the per-mutant tables are in the inline replies.

  3. Gate 1 on the tree that will land.

    • Tip re-read: d96027e52.
    • git merge-tree --write-tree d96027e52 bf322e862 = 771238c5c1a3d56c143fe9a8dacde99738e30937, rc 0.
    • Node 18, xiaobizh_n18_cpu: control d96027e52 gives 5168 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0. Merged 771238c5c gives 5184 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0.
    • Node ids go from 5305 to 5321: +16, −0, all in test_model_runner_citations.py.
  4. Lines.

    • Production code: 0.
    • Test: 155 lines, 1.94x the estimate. That is under 2x, so this is not an escalation.
    • Round 2 changed the test by +56/−49 and the docs by +1/−1.

Next: a delta review of 1a6450cad..bf322e862.

🤖 Generated with Claude Code

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 2, of PR #299 (issue #288) at head bf322e862af441b50d2e346a067eeb16e9733ac6. This is a delta review of 1a6450cad..bf322e862, written by an agent that read the eight design principles and AI_DEV_RULES.md first.

Verdict: APPROVE at bf322e862af441b50d2e346a067eeb16e9733ac6. Nothing is blocking.

  • The cycle-1 blocking finding is fixed, and the fix fires on node 18.
  • All four non-blocking findings are addressed. Each new pin fires on its own mutant.
  • The merged tree 771238c5c gates green at +16.

Delta

Two files changed. The test goes +56/−49, and 03 changes one line.

cycle-1 finding what the delta does
blocking, :117: the upper bound and the comma-list cells were pinned by nothing The moved-line test now moves 2564 → 2468 in the prepare_sample row and 2481 → 2564 in the prepare_inputs cell, and expects both failures by name.
:67 bullet carry The split is now \n\s*\n|\n(?=||\s*[-*] ), pinned by [later-bullet].
:38 unrecognised spellings Stated as limit 5 in the PR body.
:40 unpinned parser pieces Added [bare-paren], [bare-comment], [bare-beside-call] and [later-row]. _FILE is cut to py|md, and the new ::test_a_bare_line_owned_by_another_file_is_not_checked[engine_core.py|README.md] pins it.
03:148 nit Now reads "calls before any budget arithmetic".

The ponytail shrinks from cycle 1 were applied: the dict-driven parametrize, _last inlined, and the docstring trimmed.

Mutations

I re-ran the battery on node 18, in xiaobizh_n18_cpu, rather than locally with --noconftest:

  • private root /tmp/pr299rv2/mut/ATOM, a git archive of merged tree 771238c5c (md5 cf2fd26f… at both ends);
  • atom.__file__ = /tmp/pr299rv2/mut/ATOM/atom/__init__.py;
  • each run was pytest tests/compass/test_model_runner_citations.py, and the file was restored from the tarball between runs.

Test file: tests/compass/test_model_runner_citations.py

id mutant (all line-count preserving, except P9: −2 in the test and M1: +1 in the source) result the node id that goes red
N0 null control 16 passed none
B1 span upper bound deleted 1 failed, 15 passed ::test_a_line_moved_into_another_function_fails_by_name
B2 span lower bound deleted 1 failed, 15 passed ::test_a_line_moved_into_another_function_fails_by_name
B3 comma lists removed from _NUMBERS 1 failed, 15 passed ::test_a_line_moved_into_another_function_fails_by_name
P1 table-row split removed 1 failed, 15 passed ::test_a_line_with_no_symbol_fails_by_name[later-row]
P2 bullet split removed 1 failed, 15 passed [later-bullet]
P3 (?:\(\))? removed 1 failed, 15 passed [bare-beside-call]
P4 \( bare prefix removed 1 failed, 15 passed [bare-paren]
P5 # bare prefix removed 1 failed, 15 passed [bare-comment]
P6 _FILE cut to py 1 failed, 15 passed ::test_a_bare_line_owned_by_another_file_is_not_checked[README.md]
P7 (?!{_FILE}) lookahead removed 3 failed, 13 passed [bare-backtick], plus both [engine_core.py] and [README.md]
P8 owner = m["file"] → pass 16 failed every test
P9 the beside-a-name rule removed (the condition becomes True) 2 failed, 14 passed [bare-beside-name], [bare-beside-call]
P10 the ,? lines? spelling removed 1 failed, 15 passed [line-word]
M1 blank line above def prepare_sample( (source) 16 passed none
M3 See `model_runner.py:1234`. appended to an existing 07 line 16 failed the null control and every test (full-doc comparison)
X1 cycle 1's wrong green: and `model_runner.py:3300` on 01's RPC-boundary bullet now red, 16 failed the null control and every test
X4 `pp2_27b.log` (`:126`) before any model_runner.py mention in a paragraph 16 passed none. This is correct: the number is not a model_runner.py citation.

The developer's local claims reproduce on node 18. Every mutant fails the pin it was meant to fail, and no parser piece I can delete survives.

One consequence of cutting _FILE to py|md, accepted. A .log, .json, .sh, .yaml or .txt mention no longer changes which file a following bare :N belongs to. So `x.log` (`:N`) written after a model_runner.py mention is now a wrong red, never a wrong green (principle 6). It does not fire at this tree: the null control passes.

ponytail-review (delta, and the file at 155 lines = 1.94x)

test_model_runner_citations.py L14-15: delete: `from __future__ import annotations`. requires-python is >=3.10, so `str | None`-style hints need no import. -2 lines.

net: -2 lines possible.

Everything else earns its lines. Each NO_SYMBOL row is the only pin for one parser piece; P1 to P10 above show that. The two-line owned_by_another_file test is the only pin for _FILE and the lookahead. The file is 1.94x the estimate, under the 2x escalation line. The growth since cycle 1 is pins that the review asked for, not new machinery, so I do not count the size against it.

Code hygiene. At the head, the test has no design-doc references: a grep for D<n>, T<n>, P0., "principle", "Gate N" and issue numbers returns nothing. ruff check, ruff format --check and black --check pass (RUFF_RC=0, FMT_RC=0, BLACK_RC=0).

Gate 1: the tree that will land

  • Tip re-read. fork/feature/atomcompass_new = d96027e52, moved by compass(memory): drive two spellings of one phase at both shape checks #298. compass(memory): drive two spellings of one phase at both shape checks #298 touches only tests/compass/test_memory_compare.py, which is disjoint from this PR.
  • Merged tree. git merge-tree --write-tree d96027e52 bf322e862 = 771238c5c1a3d56c143fe9a8dacde99738e30937, rc 0. It matches the developer's value.
  • Staging. I wrapped the merged tree in an unreferenced preview commit f336224a6 (parents d96027e52 and bf322e862, created with commit-tree, so no ref was moved), then:
    • git archive plus stamps (.compass-changed against d96027e52);
    • docker exec -i … tar -x into private /tmp/pr299rv2/gate/ATOM;
    • atom.__file__ = /tmp/pr299rv2/gate/ATOM/atom/__init__.py;
    • gate line commit: f336224a6 (stamp).
  • Run. The tree's own scripts/compass/gate_cpu.sh ran once, unpiped, under timeout -k 10 1800.
side result GATE_CPU_RC
merged 771238c5c (this review) 5184 passed, 155 skipped, 3 xfailed 0 PASSED
control d96027e52 (developer's run) 5168 passed, 155 skipped, 3 xfailed 0

The delta is +16, which equals the file's 16 node ids. No flaky class fired.

Next action

Land bf322e862 squashed onto feature/atomcompass_new, provided the tip is still d96027e52 or re-computes to a tree that equals 771238c5c. The ponytail -2 is optional and does not need another cycle.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Addendum to the cycle-2 review of PR #299 at head bf322e862af441b50d2e346a067eeb16e9733ac6: the tip moved while the review was being posted. The verdict is unchanged: APPROVE at bf322e862af441b50d2e346a067eeb16e9733ac6.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant