Skip to content

compass(capture): hold the step-axis backstop, the normalisation row and both sentinel bindings - #359

Merged
jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-236
Sep 23, 2026
Merged

jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-236

Conversation

@jgong5

@jgong5 jgong5 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Closes #236

What changed

One file, tests/compass/test_capture_real_model.py: +65 / −5 test lines, 0 production lines. Head 12a2ad6d4, based on be86326cc.

  1. The _step_axis backstop now has a test. test_the_step_axis_refuses_a_hint_torch_specialises calls _step_axis(FakeTensorMode(shape_env=ShapeEnv()), 1) in process, and expects an AssertionError matching not a symbol. main refuses width 1 before the backstop is reached, so until now no test reached it. The same test calls _step_axis at MIN_STEP_WIDTH, which must return an s<n> symbol. That is the control.
  2. The width-1 refusal test now checks the exit status, not only the message. It asserts returncode == 2, argparse's usage-error status, where it used to assert != 0. With main's refusal removed, the backstop exits 1, so the test now fails on the status as well as on the message. A comment in the test says why the RECORD_MARKER half cannot tell the two refusals apart: either refusal leaves the record out.
  3. The normalisation row is an exact set, like gemm and attention. The four operators were measured on node 18 at the tip, at both TP1 and TP2 with --step-symbol, and are the same at both: aiter._fused_qk_rmsnorm_group_quant_kernel.default (129), aten.mean.dim (48), aten.pow.Tensor_Scalar (48), aten.rsqrt.default (48).
  4. The apply_simulated_tp sentinel now fires, on both names. test_a_call_through_either_binding_reaches_the_sentinel runs a fresh interpreter with _declare_cuda and _declare_arch installed. It calls model_runner.apply_simulated_tp(None) and simulated_tp.apply_simulated_tp(None), and requires the line SENTINEL-CALLS 2. It takes about 8 s on node 18. If a binding is missing, or the sentinel calls through, the real function runs and raises on None.tensor_parallel_size. It could not run in process: in the CPU container, import atom.model_engine.model_runner fails on aiter's rocminfo probe. One sentence in _watch_simulated_tp's docstring is updated to match.

Named result: mutation table, node 18 (xiaobizh_n18_cpu)

Every mutation changes one line and keeps the line count. Each ran on its own copy of the staged tree (git archive of the commit), over the whole file. Tip be86326cc is 18 tests and head 12a2ad6d4 is 20.

# mutation tip be86326cc head 12a2ad6d4
null none 18 passed 20 passed
N0 comment reworded (node 18's → node 18) – 20 passed
M2b _step_axis: if not re.fullmatch(r"s\d+", str(axis)): → if False: 18 passed (blind) 1 failed / 19 passed: test_the_step_axis_refuses_a_hint_torch_specialises, :2653, Failed: DID NOT RAISE <class 'AssertionError'>, the raise half
M2 main: if args.width < MIN_STEP_WIDTH: → if False: 1 failed / 17 passed: test_the_capture_refuses_…, :2595, 'traces no symbol' in …, the string half only 1 failed / 19 passed: same test, :2633, assert 1 == 2 on .returncode, the status half, before the string is read
M10s "aten.split_with_sizes" declared into normalisation 18 passed (blind) 1 failed / 19 passed: test_the_symbol_reaches_the_work_that_decides_the_cost, :2700, Extra items in the left set: 'aten.split_with_sizes.default'
M10 "aten.as_strided" declared into normalisation (the issue's mutation) 1 failed / 17 passed: test_the_census_counts_the_symbols_the_probe_leaves, :2400, Extra items … 'normalisation' 2 failed / 18 passed: the same probe test (:2437) and test_the_symbol_reaches_… (:2700, Extra items … 'aten.as_strided.default')
M10v "aten.view" declared into normalisation 1 failed / 17 passed: test_the_census_counts_an_expression_as_a_symbol, KeyError: 'view' 2 failed / 18 passed: that test and test_the_symbol_reaches_… (:2700, 'aten.view.default')
S1 model_runner.apply_simulated_tp = sentinel → rebinds the original 18 passed (blind) 1 failed / 19 passed: test_a_call_through_either_binding_…, :2242, traceback at <string> line 11 → simulated_tp.py:44 logical = config.tensor_parallel_size
S2 simulated_tp.apply_simulated_tp = sentinel → rebinds the original 18 passed (blind) 1 failed / 19 passed: same test, <string> line 12
S3 sentinel body → raise SystemExit 18 passed 1 failed / 19 passed: same test, assert 'SENTINEL-CALLS 2' in []

One correction to #236. The issue's own mutation, aten.as_strided declared into normalisation, is already red at the tip. It fails in test_the_census_counts_the_symbols_the_probe_leaves. #261 (#238) made that probe census an exact family set after #236 was measured at 70f8cd4db. as_strided carries a symbol on the probe pass, so moving it moves a family. The blind spot is still real for an operator that carries the symbol only on the step-symbol pass. aten.split_with_sizes (80 ops) is green at the tip and red at the head, on the new exact set, and it is the row reported for item 2. aten.view is caught at the tip too, but only by a test that builds a recorder by hand from that operator's name (KeyError: 'view').

Gate 1: CPU tier, control measured on the same tip

scripts/compass/gate_cpu.sh from each tree's own copy, on xiaobizh_n18_cpu. Each tree was staged with git archive plus .compass-commit / .compass-changed stamps, and atom.__file__ resolved under each staged root. Runs were one at a time, unpiped, under timeout -k 10 7200.

tip be86326cc head 12a2ad6d4
passed 5,261 5,263 (+2)
skipped 155 155
xfailed 3 3
failed 0 0
GATE_CPU_RC 0 0
gpu not required not required (the diff touches only tests/compass/)

The node-id delta is exactly the two new tests, from --collect-only on the changed file:
test_a_call_through_either_binding_reaches_the_sentinel and test_the_step_axis_refuses_a_hint_torch_specialises. No timing-class test failed on either side.

git merge-tree --write-tree be86326cc 12a2ad6d4 = 610e72814436ea42374c3185a406511faeb668b4 = 12a2ad6d4^{tree}.

ruff format --check and ruff check pass on the file.

Review record

A correction is posted on #150 at #150 (comment): its finding 5 was closed on a basis that was only half fail-able.

Not done

  • The GPU tier was not run. gate_cpu.sh reports it as not required, and the diff is test-only.
  • The sentinel test proves which names the sentinel covers. It does not make any capture reach ATOM's call site, which stays unreachable because _build_runner overrides that method.

🤖 Generated with Claude Code

…and both sentinel bindings

The width-1 refusal test now asserts argparse's exit status 2. Before, it
asserted any non-zero status, which the backstop in _step_axis also satisfies,
so with main's refusal removed only the message match failed. A new test calls
_step_axis at a hint of 1 directly. main never lets that hint through, so until
now nothing reached the backstop.

The normalisation row's symbol-carrying operators are now an exact set, like
gemm and attention. Before, the test only checked that one operator was
present, so an operator declared into the family passed. The four operators
were measured at TP1 and TP2.

A new subprocess test installs the apply_simulated_tp sentinel and calls
through both bindings. No capture reaches ATOM's call site, so before this
nothing checked which names the sentinel covered.

Closes #236

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
assert completed.returncode != 0
# 2 is argparse's usage error. Without `main`'s refusal, `_step_axis`
# still refuses, but as an uncaught AssertionError, which exits 1.
assert completed.returncode == 2

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking. Ruling on returncode == 2: keep it. It is sound, but it also holds the mechanism, not only the refusal.

Principle 6: "Refuse rather than fall back. A declined answer with a named reason is a result." Principle 8: "Every claim carries its measurement."

Exit status 2 is a documented stdlib contract, not an accident of the implementation: ArgumentParser.error "terminates the program with a status code of 2". The driver that uses it is main(), in this same file. So the assertion cannot drift under the test; only an edit to this file can move it. Here is what the entry point prints at --tp 1 --step-symbol --width 1, measured on node 18 (xiaobizh_n18_cpu):

tree rc stderr CAPTURE-RECORD time
head (main refuses) 2 usage: … + test_capture_real_model.py: error: --width 1 traces no symbol: …, no Traceback 0 2 s
M2 (main's check → if False:; only the backstop refuses) 1 Traceback … AssertionError: the step axis came back as '1', not a symbol: … 0 12 s
R1 (parser.error( → sys.exit(; same message, same refusal) 1 --width 1 traces no symbol: …, no usage:, no Traceback 0 2 s

R1 is still a refusal that conforms to principle 6. It declines, names its reason, and prints no record. Even so, it turns this test red: 1 failed / 19 passed, test_the_capture_refuses_a_width_that_torch_would_specialise, :2633, assert 1 == 2. At the tip the same mutation gives 18 passed. So == 2 also pins how main refuses. It is not needed to tell the two refusals apart: "traces no symbol" already does that for M2.

That trade is acceptable, because the mechanism lives in the same file and is documented. If you want a second discriminator that does not depend on the mechanism, use assert "Traceback" not in completed.stderr. It reads 0 at head, 1 under M2 and 0 under R1: a clean refusal as against the backstop's uncaught raise. Either is fine. No change is required.

" from atom.model_engine import model_runner\n"
" model_runner.apply_simulated_tp(None)\n"
" simulated_tp.apply_simulated_tp(None)\n"
"print('SENTINEL-CALLS', len(calls))\n"

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking: the child never says which atom it imported.

AI_DEV_RULES, setup rule 3: "Every command sets PYTHONPATH to its own worktree and verifies it with python -c "import atom; print(atom.__file__)" before trusting a result." Principle 8: "Every claim carries its measurement."

Measured, it is hermetic by construction. With -c, sys.path[0] is '', which is cwd, which is tree_root, and PYTHONPATH is tree_root as well. The same argv, cwd and env on node 18 printed CHILD_ATOM /tmp/pr359r1/head/ATOM/atom/__init__.py sys.path0= ''. The S1 and S2 mutants' tracebacks also name the staged copy: /tmp/pr359r1/mut/head-S1/ATOM/atom/distributed/simulated_tp.py, line 44.

Still, the capture tests in this file assert record["atom_package"].startswith(tree_root) (:2133), and this child asserts nothing about it. Two cheap changes would close that, and I'd make them only if you touch this again:

  • add import atom to the probe and print atom.__file__ beside SENTINEL-CALLS;
  • assert that the printed path starts with str(tree_root).

The rest checks out:

  • Timeout: bounded, with timeout=1800, the same as every subprocess in the file.
  • Failure mode: it fails rather than skips when the import fails. Removing _declare_arch from the probe (S5) gives 1 failed / 19 passed, :2242, assert 'SENTINEL-CALLS 2' in [], with RuntimeError: Get GPU arch from rocminfo failed in the message.
  • Cost: 9.10 s in --durations, taking the file from 115 s to 124 s.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

This review is agent-authored.

Review, cycle 1: #359 (issue #236), head 12a2ad6d4965637a16de934758b4cb080000f3d8

Verdict: APPROVE, head 12a2ad6d4965637a16de934758b4cb080000f3d8. No finding is blocking. There are two non-blocking notes, both inline:

  1. The returncode == 2 ruling. Keep it.
  2. The sentinel child does not print atom.__file__. I measured it and it is hermetic anyway.

The inline notes: :2633 #359 (comment) and :2229 #359 (comment).

I read the eight principles in atom/compass/design/README.md and atom/compass/AI_DEV_RULES.md at the tip (6987b0716, then d10cb834f, whose change to both files is nil) before the diff.

1. Mutation table, reproduced on node 18 (xiaobizh_n18_cpu)

Each mutation ran on its own copy of a git archive tree, over the whole file, and changed one line with the line count preserved:

  • head: 2919 lines;
  • tip 6987b0716: 2859 lines. The test file is byte-identical to be86326cc's. The tip moved only spec/, docs and test_spec_verbs.py, and d10cb834f since moved test_spec_verbs.py alone.

atom.__file__ resolved under each staged root in every run, and the child tracebacks name the staged copy.

"Test" is the failing node id, under tests/compass/test_capture_real_model.py::.

# mutation tip 6987b0716 head 12a2ad6d4 test (head) line, assertion (head)
null none 18 passed 20 passed
N0 comment whitespace (# Both widths. + 2 spaces) 20 passed
M2b _step_axis: if not re.fullmatch(r"s\d+", str(axis)): → if False: 18 passed (blind) 1 failed / 19 passed test_the_step_axis_refuses_a_hint_torch_specialises :2653 Failed: DID NOT RAISE <class 'AssertionError'>
M2 main: if args.width < MIN_STEP_WIDTH: → if False: 1 failed / 17 passed: same test, :2595, 'traces no symbol' in … (string half only) 1 failed / 19 passed test_the_capture_refuses_a_width_that_torch_would_specialise :2633 assert 1 == 2 (status half, before the string)
R1 main: parser.error( → sys.exit( (same message) 18 passed 1 failed / 19 passed same :2633 assert 1 == 2. See the inline ruling.
M10s "aten.split_with_sizes" declared into normalisation 18 passed (blind) 1 failed / 19 passed test_the_symbol_reaches_the_work_that_decides_the_cost :2700 Extra items in the left set: 'aten.split_with_sizes.default'
M10 "aten.as_strided" declared into normalisation 1 failed / 17 passed: test_the_census_counts_the_symbols_the_probe_leaves, :2400, Extra items … 'normalisation' 2 failed / 18 passed the same probe test (:2437) and test_the_symbol_reaches_… :2700 'aten.as_strided.default'
M10v "aten.view" declared into normalisation 1 failed / 17 passed: test_the_census_counts_an_expression_as_a_symbol, :2433, KeyError: 'view' 2 failed / 18 passed that test (:2470) and test_the_symbol_reaches_… :2700 'aten.view.default'
S1 model_runner.apply_simulated_tp = sentinel → rebinds itself 18 passed (no such test) 1 failed / 19 passed test_a_call_through_either_binding_reaches_the_sentinel :2242. Child <string> line 11 → …/head-S1/ATOM/atom/distributed/simulated_tp.py, line 44, AttributeError: 'NoneType' object has no attribute 'tensor_parallel_size'
S2 simulated_tp.apply_simulated_tp = sentinel → rebinds itself 18 passed (no such test) 1 failed / 19 passed same :2242, <string> line 12, same AttributeError
S3 sentinel body → raise SystemExit 18 passed 1 failed / 19 passed same :2242 assert 'SENTINEL-CALLS 2' in []
S4 sentinel body → pass (records nothing) 1 failed / 19 passed same :2242 assert 'SENTINEL-CALLS 2' in ['SENTINEL-CALLS 0']
S5 probe drops capture._declare_arch(tmpdir) 1 failed / 19 passed (fails, does not skip) same :2242 in [], with RuntimeError: Get GPU arch from rocminfo failed

Every row the PR claims reproduces, with the same counts, node ids and line numbers. I added four rows: R1, S4, S5 and M2 at the tip. I found no green mutant on the new code. AI_DEV_RULES gate 4 says "A check counts only once someone has seen it fire." All three #236 items have now been seen to fire.

2. The returncode == 2 ruling (principle 6): keep it, non-blocking

The full ruling and the measured table are inline on :2633.

Is it a stable contract? Yes. Exit status 2 is argparse's documented behaviour: ArgumentParser.error "terminates the program with a status code of 2". The driver that uses it is this file's own main.

What each refusal prints, measured:

  • head: rc 2, usage: plus error: --width 1 traces no symbol…, no traceback, 2 s.
  • Backstop only (M2): rc 1, an AssertionError traceback, 12 s.
  • Both: no CAPTURE-RECORD.

The coupling is real but local. R1 swaps parser.error for sys.exit with the same message. That is still a principle-6 refusal, yet it reddens on assert 1 == 2. The alternative would be a message-based assertion on the argparse error ("usage:" or ": error:"), but it couples to argparse just as much. The mechanism-agnostic option is "Traceback" not in stderr, which reads 0, 1 and 0 across head, M2 and R1. It is optional.

3. The exact normalisation set: correct, and scoped to exactly what is measured

The fixture has one config. It is the vendored, sha256-checked qwen3_5_27b_config.json (the published Qwen3.8-27B). The assertion runs inside for tp in (1, 2) at the default width DECODE_SEQS = 2, and it covers no other model family and no other width. So no other supported model can reach this set.

What could move it. Only a change to ATOM's forward or to aiter's operator names can move it. That is the intended red, the same as for the gemm and attention rows beside it. An undeclared new norm operator would fail first in the unclassified bucket.

Measured on the head tree, at TP1 and TP2, and at widths 2 and 8 (the file's SECOND_WIDTH, which the assertion does not run), all four records give the identical row:
{_fused_qk_rmsnorm_group_quant_kernel: 129, pow.Tensor_Scalar: 48, mean.dim: 48, rsqrt.default: 48}.

The four operators are the whole declared normalisation family in OP_FAMILIES. So the set now says "every declared norm operator carries the symbol", and any extra declaration that carries it reddens (M10s, M10, M10v).

4. The sentinel subprocess test

  • Hermetic? Yes, by construction, but not asserted in the child. That is non-blocking and inline on :2229. Measured CHILD_ATOM /tmp/pr359r1/head/ATOM/atom/__init__.py, with sys.path[0] == '', which is cwd, which is tree_root. The mutant tracebacks name the staged simulated_tp.py.
  • Bounded? Yes: timeout=1800, as with every subprocess in the file.
  • Cost: 9.10 s. The file takes 115 s at the tip and 124 s at head, which is noise on a CPU tier of about 5,300 tests.
  • Fails rather than skips? Yes. There is no skip path. S5 and the import failure both end in assert 'SENTINEL-CALLS 2' in [], with the stderr tail attached.
  • Why not in-process? Measured: in xiaobizh_n18_cpu, a bare import atom.model_engine.model_runner raises RuntimeError: Get GPU arch from rocminfo failed. Installing _declare_arch or _declare_cuda in the pytest process would break the file's own rule that nothing mutates torch.cuda there. So the subprocess is justified.

5. Truth checks (principle 8: "Every claim carries its measurement")

New and changed comments and docstrings. All five are true as measured:

  1. "Without main's refusal, _step_axis still refuses, but as an uncaught AssertionError, which exits 1". M2: rc 1, AssertionError traceback.
  2. "Either refusal satisfies this, so it fails only with both gone". M2: 0 CAPTURE-RECORD lines.
  3. "with the check disabled the call returns a plain 1". Measured int '1' at hint 1, and SymInt 's26' at hint 2.
  4. "raises on None.tensor_parallel_size". S1/S2: AttributeError at simulated_tp.py:44.
  5. The _watch_simulated_tp docstring, "leaves every test that runs a capture passing, and fails the one that calls both bindings directly". S3: 1 failed / 19 passed, and only the sentinel test.

Design-doc references. A grep over the whole file at head for design-doc references (D<n>, T<n>, P0.x, "principle N", "Gate N", doc numbers) returns nothing.

Ruff. ruff format --check and ruff check on the file both give rc 0.

#150 correction (#150 (comment)): accurate.

  • The quote "Both halves, because the CLI is not the only entry point" is verbatim from round 2 (5772255808).
  • "18 passed" for M2b and ":2595, string only" for M2 both reproduce at the tip.
  • The fix description matches head: DID NOT RAISE at 1 failed / 19 passed, and assert 1 == 2.

#236 correction (5804216613): accurate.

Effort: +65/−5 against the 15–40 estimate. That is under 2x, so it is not an escalation under the effort rule.

6. ponytail-review over the diff (+65/−5)

I found no delete:, stdlib:, native:, yagni: or shrink: finding. What I weighed:

  • The 13-line probe string is the minimum a fresh interpreter needs: load, declare, install, call both bindings, print. Loading through sys.path would save one line and cost resolve() symmetry. Folding the one-line lines = … into the assert wraps to three lines under ruff's 88 columns.
  • The MIN_STEP_WIDTH control in the _step_axis test is a single self-check, which ponytail excludes by rule.
  • The exact set replaces a membership line in the same shape as the two rows above it.

Lean already. Ship.

7. Gate: the landing tree, merged and gated at each tip

The tip, re-read twice.

None of the three tip commits touches this PR's file. merge-tree differs from 12a2ad6d4^{tree} (610e72814…) in both cases, so the PR's own gate does not stand for the landing tree. The first tree was gated once; after the tip moved to d10cb834f, the landing tree was gated once as well:

tree that lands at 6987b0716 tree that lands at d10cb834f (current)
git merge-tree --write-tree <tip> 12a2ad6d4 bf26d601b90c9d23f049bfddd8d577a7b2f0c5ae 0fb5c6749ad478b9957037d3bbe37dc98138dcab
commit-tree stamp (no ref) f29521a63 8134adbea
passed 5,265 5,265
skipped 155 155
xfailed 3 3
failed 0 0
GATE_CPU_RC 0 PASSED 0 PASSED
gpu not required not required

How both were staged.

  • git archive of the stamp commit was piped into xiaobizh_n18_cpu, at a private path, with the md5 matched on both ends.
  • .compass-changed = tests/compass/test_capture_real_model.py.
  • The gate is the tree's own scripts/compass/gate_cpu.sh, under timeout -k 10 7200, unpiped, with one gate at a time and no other load of mine running.
  • The gate printed atom: /tmp/pr359r1/merged{,2}/ATOM/atom/__init__.py and commit: <stamp>.

The count, decomposed (principle 7: "Never report an aggregate without its decomposition"). 5,265 is the PR's head figure of 5,263 plus 2. That +2 is #346's test_spec_verbs.py: --collect-only gives 153 at the PR's base, and 155 at 6987b0716, at d10cb834f and in both merged trees. #357 changes that file without changing its count. This file collects 18 at the tip and 20 in each merged tree. Skips and xfails are unchanged from the PR's control. No timing-class test failed.

What the next task here should watch

  • The sentinel is now proven to cover both names. Still, as the PR says, no capture reaches ATOM's call site, because _build_runner overrides it. The source scan in test_apply_simulated_tp_is_called_only_where_the_capture_does_not_go remains the only guard against a second caller.
  • returncode == 2 ties the refusal test to parser.error. If main's refusal is ever re-plumbed, change the status assertion in the same commit, or switch to the Traceback-free form.

Staging under /tmp/pr359r1/ in xiaobizh_n18_cpu is removed after this review, and the ControlMaster is closed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant