Skip to content

CAP-2: make the capture symbolic on ATOM's real decode forward - #150

Merged
jgong5 merged 2 commits into
compass/cap-1from
compass/cap-2
Sep 23, 2026
Merged

jgong5 merged 2 commits into
compass/cap-1from
compass/cap-2

Conversation

@jgong5

@jgong5 jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Closes #143. Base is compass/cap-1 at 24d75742d, its round-3 head, approved at that branch's cycle 3. Unlanded and in review. Restacked onto it in round 2 with git rebase --onto 24d75742d 9fcd6c7bd compass/cap-2; the earlier base 9fcd6c7bd was this branch's round-1 base and ae8b43ae7 its first.

What this is

A decode step of the published Qwen3.8-27B whose width is a free symbol traces through ATOM's own ModelRunner, ATOM's own model classes and ATOM's own staging with ShapeEnv.replacements empty and no symbol solved anywhere — at TP1 and at TP2, on a machine with no GPU.

Named result — the census, by operator family, never as a total

ops / shape entries that are not plain integers. The control is the same file with the step's width left an integer.

family TP1 control TP1 step-symbol TP2 control TP2 step-symbol
gemm 257 / 0 257 / 514 257 / 0 257 / 514
attention 64 / 0 64 / 304 64 / 0 64 / 304
normalisation 273 / 0 273 / 802 273 / 0 273 / 802
collective — — 133 / 0 133 / 266
activation 128 / 0 128 / 256 128 / 0 128 / 256
elementwise 212 / 0 212 / 536 212 / 0 212 / 536
allocation 565 / 0 566 / 803 566 / 0 567 / 804
view 922 / 0 921 / 1736 925 / 0 924 / 1742
transfer 14 / 0 14 / 38 16 / 0 16 / 44
embedding 1 / 0 1 / 2 1 / 0 1 / 2
sampling 2 / 0 2 / 4 2 / 0 2 / 4
bookkeeping 83 / 0 83 / 21 85 / 0 85 / 23
total 2,521 / 0 of 12,544 2,521 / 5,016 of 12,544 2,662 / 0 of 13,107 2,662 / 5,297 of 13,107

The operators carrying the symbol are named, not inferred from a family total: aiter.gemm_a16w16 (256 at linear.py:974, 1 at the LM head); aiter.linear_attention_with_output_base x48 and aiter.unified_attention_with_output_base x16; aiter._fused_qk_rmsnorm_group_quant_kernel x129 plus mean.dim / pow / rsqrt x48 each. An operator matching no declared family lands in unclassified, which the tests require to be empty — a new operator is a classification somebody makes, not a bucket it falls into.

Round 2: the classifier matched by prefix, which defeated that bucket. aten.add captured aten.addmm and aten.addbmm, which are GEMMs; aten.slice captured aten.slice_scatter, which writes. Nothing on this step is one of those, so the guard had never been wrong — but it could not have fired for the case it exists for. It matches whole operator names now (the overload dropped, the rest matched exactly), which also meant enumerating the _c10d_functional.* and profiler.* entries that had been namespace prefixes and adding aten.zeros, which the probe passes dispatch and the prefix table classified nowhere at all. To be exact about what the repair does with those four: it does not move them into gemm. None of addmm, addbmm, slice_scatter or select_scatter is declared in any family, so under whole-name matching each lands in unclassified — the bucket the tests require to be empty. An ATOM that started routing a projection through addmm now turns this file red and makes somebody classify it, rather than being absorbed into elementwise and quietly understating the three cost-bearing families. That refusal, not a re-filing, is the property being claimed. The table above was re-derived after the fix and did not move — family census identical structure-for-structure on all five passes (TP1 control, TP1 step-symbol at widths 2 and 8, TP2 step-symbol at widths 2 and 8), same digests.

Both passes record the same operators: 2,521 at TP1 and 2,662 at TP2 on both sides, and the same 12,544 / 13,107 shape entries. Exactly one call differs, and it is a value rather than a shape — gdn_attn.prepare_decode writes the step's token count into a staging tensor's tail, which materialises a symbol (scalar_tensor) where it used to lift a constant (lift_fresh). It was found by comparing distinct_ops, not the digest, which covers operator names and shapes only. What changed is what the shape entries say, not which operators ran.

The three sites are symptoms of one line of torch

Round 2 of the base branch established that the sites are an order, not a set, and that closing site one relocates the bound. Site three is in another file and a later phase of the step — forward_context.py in assert_shape_contract, reached from the run_model call at model_runner.py:3254 with prepare_inputs already returned. It is not fourteen lines after aiter_attention.py:1115; an earlier revision of this branch said so and the restack drops that.

The root is that SymInt.__index__ and SymInt.__int__ are guard_int. Every host-side use of the step's width — filling a staged buffer's numpy view, slicing a Python list, checking a staged array's length — needs a number, gets the trace-time hint, and records the ask as Eq(s, hint).

Measured on ATOM's decode path, 16 ATOM lines convert the width during one traced forward — 20 conversions, all host fills and slices:

file lines conversions
aiter_attention.py in prepare_decode 1106, 1115, 1121, 1122, 1123, 1131, 1132 10
model_runner.py in prepare_inputs 2441, 2452, 2454 4
model_runner.py in prepare_input_ids 510, 513 2
model_runner.py in prepare_sample 2537 1
backends.py in _mrope_cpu_view 398, 400 2
gdn_attn.py in _attach_gdn_decode_metadata 1237 1

plus a seventeenth in forward_context._rows and an eighteenth, ScheduledBatch.__init__, which solves the width by comparison rather than conversion. Repairing them one at a time does not converge; that is what the two previous attempts were doing.

The conversion is not the defect; the recording is. A host fill genuinely needs a number, and the number it needs is the hint, which is the count ATOM computed. So the capture keeps the conversion and replaces the guard with its own log — every conversion, with the ATOM line it happened on. Round 2: that log is now asserted. It was claimed to be in three places and implemented in none; the only assertions on it were a membership test on site one and a negative one on _rows, so a conversion at a seventeenth line would have joined the log and nothing would have failed. It is now compared as a multiset of (ATOM line, conversions) against a module constant, EXPECTED_HOST_RESOLUTIONS, the way EXPECTED_GUARDS and STAGED_COPIES already were — the table above, asserted at TP1 and TP2 and at both step widths, where it is the same 20-over-16 multiset. A multiset rather than a set, so a line that converts twice where it converted once also fails.

__bool__ is untouched, so a branch on a width still installs its guard and a step whose shape decides which path ATOM takes still records that it did.

Which route, and why

Neither of the brief's two.

  • Route 1, a symbolic CpuGpuBuffer. Built and measured, and not required. Site two's -> 16384 is a buffer capacity (max_model_len // block_size) being solved against the CPU side's constant, and it exists only because that probe makes every dimension of the staged device tensor a free symbol. Capacities are engine configuration and are not a function of the step. With only the width symbolic, self.gpu[:n] and self.cpu[:n] carry the same symbol and copy_to_gpu dispatches unchanged. The alternative was built anyway: a fake, non-numpy-backed CPU side costs 38 prim.device calls ATOM never makes and leaves every family's non-numeric shape-entry count unchanged — 5,016 either way, family for family. Round 2 corrects the wording of that row: "changes not one shape entry" was loose. The total entries rise 12,544 → 12,585, because those 38 extra calls carry entries of their own; what is unchanged is the non-numeric count, per family. And the row's code is not in the tree: it was a fake CPU side written into the capture's staged allocators, measured, and removed once it had answered, so re-checking it means rebuilding it. It is the row that decided the route and it is the weakest-sourced thing in this PR; stating that is the correction, not defending it. atom/utils/__init__.py — the brief's file set — is not touched.
  • Route 2, a capture entry point below prepare_inputs. Rejected without building it: reaching the model that way means hand-constructing AttentionMetaData and Context, which re-derives the shapes ATOM's own staging derives. The design record carries a measured account of what that costs — a prior re-derivation was wrong in both directions against the real classes. It also answers the wrong question: the reason to trace ATOM's staging is to find out what it does to a symbol.
  • What was built instead: the symbol enters ATOM's staging and ATOM derives every width from it — ForwardMode.decide settles both units, prepare_inputs writes the cu_seqlens_q boundary at running_bs + 1, prepare_decode uses it as every staged bound. Nothing re-derives a width. One symbol, because a decode step is one query row per sequence; two would have to be equated later, which is the specialisation again by a longer route — and that is demonstrated rather than argued, see site three above.

Is the symbol free, or a hint in disguise?

This is the question the result rests on, and it is the load-bearing evidence — not replacements == {}, which in this pass cannot fail for __index__ and __int__ because those two are the things that were replaced (see What is not done). ATOM computes its host values from the hint — the slot mapping, the block tables, the cumulative sequence lengths are all real numbers for one concrete width. A graph built that way would still be a graph about one width if any dimension had quietly taken the hint, and nothing in the census would say so. A symbol can also be lost rather than solved, by code reading its hint and building a tensor of that size, and that records no replacement at all.

So the step is traced at two widths, 2 and 8, at TP1 and at TP2 (round 2 adds TP2; round 1 ran TP1 only, although TP2 is the named result), and the inventories compared operator for operator and shape for shape with the symbol's name set aside:

TP1 w2 TP1 w8 TP2 w2 TP2 w8
operators 2,521 2,521 2,662 2,662
shape entries 12,544 12,544 13,107 13,107
non-numeric 5,016 5,016 5,297 5,297
inventory digest a7c1ec1b6ac1ec4b… a7c1ec1b6ac1ec4b… a37d76da81cd4547… a37d76da81cd4547…
concrete dimensions, by family and value identical identical identical identical
shape_env guards 3 4 2 3

Identical except the last row, and round 2 states that last row rather than omitting it — see Guards below. Anything carrying the hint would be a 2 in one column and an 8 in the other. That also settles the other 7,528 entries, which is the second half of the census: every dimension that stayed a number is the same number at both widths, so none is a width in disguise. There are 23 distinct concrete values in the whole TP1 inventory and every one is the model's or the engine's — 5,120 (hidden) x2,959, 17,408 and 34,816 (MLP), 248,320 (vocab), 16,480 / 14,336 / 10,240 / 8,192 / 6,144 / 1,024 (projection widths), 256 (head dim), 128 and 48 (head counts), 512/513 (max_num_seqs), 16,384/16,385 (max_model_len // block_size), 49,152, and 24/4/3/1/0. Neither 2 nor 8 appears.

At TP2, 2 does appear — 73 times — and it is the group's width, not the step's. Round 1's assertion only ever ran at TP1 and so never met this. What separates the two readings is the row above: concrete_dims is identical at step widths 2 and 8, so a 2 that survives an eight-wide step is an engine constant in the way a head count is. The test now asserts the second width is absent at both TPs — the direction that can fail — and asserts the group width's presence at TP2 with the reason.

Guards — the applicability, read rather than totalled, and hint-dependent

TP1 TP2
at every width 28*N <= 8192, N < 128, N <= 512 N < 128, N <= 512
at every width but 2 …plus N + 1 > <width> …plus N + 1 > <width>

The upper bounds are staging capacities read against the step's width; none is an equality, which is the same fact as replacements being empty. TP2 installs one fewer of them, and that is a finding rather than noise: the per-token bound goes, because the width halves each rank's share of the rows it reads.

Round 2: the guard set is the one artifact of this capture that depends on the hint, and the round-1 table was the hint-2 case published as the graph's applicability. Re-measured at TP1 hints 2, 3, 8 and 16 and at TP2 hints 2 and 8: at every hint but 2 a lower bound N + 1 > <hint> appears as well. At hint 2 it is elided because a size symbol's default range is [2, ∞), which makes the inequality vacuous. It is not a specialisation — it is a __bool__ comparison installing an inequality on a live symbol, so it is the positive evidence that __bool__ is untouched. What it costs is honesty about applicability: at any hint but 2 the traced graph carries a lower bound too. expected_guards(tp, width) composes the set per width, and the two-width test asserts the whole set at both widths and the lower bound's presence and absence, rather than leaving the guards out of the comparison.

What runs differently in production

One line, and on the evidence, nothing.

-            return None if t is None else int(t.shape[0])
+            return None if t is None else t.shape[0]

ForwardMode.assert_shape_contract's _rows helper. torch.Size.__getitem__ returns a Python int for any tensor with a real size, so int() around it is the identity in every served step — same value, same type, same assertion messages. It matters only where the size is symbolic, and there converting one side of an equality to a number forces the other side to become that number.

The change is inside a def that exists only to feed four asserts in that method, and the method is shape-only by contract ("Shapes only, never values: reading a device tensor's contents on this path is a D2H sync per step"). Nothing else reads _rows. This is site three, and the only one of the three not reachable from a capture-time substitution. Both gate tiers were run because it is atom/utils/; atom/utils/forward_context.py is not in gpu_gate_triggers.txt — the CPU tier covers it — and the GPU tier was run anyway.

The pin: updated, and three ways ablated

CAP-1's assertions were re-pointed, not removed or loosened, and the four tests added in round 1 are now five:

  • test_the_inventory_is_concrete_at_both_widths — unchanged; it is the control.
  • test_the_first_two_specialisation_sites_are_where_they_were_measured — unchanged and still passing.
  • test_closing_site_one_moves_the_bound_to_a_third_site — updated. SITE_THREE moves one frame out and the test also asserts _rows is not on the path.
  • test_atom_s_own_buffer_constructor_is_what_runs — unchanged and passing; the step-symbol pass also asserts symbolic_device_allocations == 0 and constructed == 19.
  • Added: nothing specialises (replacements empty, no specialisations, the guard set by width); the census by family with its control; the three sites by what now happens at each, with the host-resolution log asserted against its declared set; the two-width digest at TP1 and TP2; and — round 2 — test_the_capture_refuses_a_width_that_torch_would_specialise, which runs the real entry point at --width 1 and asserts a non-zero exit, the message, and that no record is printed.

Ablated, on this head, all three re-run against the 14 tests:

ablation result
restore int(t.shape[0]) in _rows 5 of 14 fail
do not resolve on the host (keep guard_int) 4 of 14 fail
make the symbolic buffer's CPU side fake instead 13 pass at round 1's head — the row that established the buffer is not load-bearing, and the row whose code is not in the tree

Gates

Both tiers, re-measured as deltas against the new base 24d75742d. Round 1's figures against 9fcd6c7bd do not carry forward. Each tree staged with git archive + docker cp into a path of its own, never into the shared mount, gated with its own scripts/compass/, COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, tarball md5 and per-file md5 checked on both ends, .compass-commit read back inside each container.

CPU tier, in xiaobizh_n18_cpu

control 24d75742d branch 70f8cd4db
passed 4,566 4,571 (+5)
skipped 149 149
xfailed 3 3
failed 0 0
GATE_CPU_RC 0 0
pytest wall 83.65 s 147.48 s
shell wall 1 m 30.1 s 2 m 33 s

The +5 is exactly this file's five new tests — round 1's four plus round 2's width refusal — and the skip count is identical on both sides, so the tier's three-way flake (tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk) did not fire in either direction: no ±1 in the passed/skipped split and no non-zero GATE_CPU_RC. The gate's own stamp reads gpu: not required (.compass-changed stamp) on the branch — atom/utils/forward_context.py is not in gpu_gate_triggers.txt, because the CPU tier covers it via tests/test_forward_mode.py. The GPU tier was run anyway, because this is atom/utils/.

GPU tier, in xiaobizh_n18, judged as a delta

control 24d75742d branch 70f8cd4db
passed 5,340 5,345 (+5)
failed 5 5
errors 0 0
skipped 105 105
xfailed 3 3
GATE_GPU_RC 0 0
PREFLIGHT_RC before / after 0 / 0 0 / 0
pytest wall 124.01 s 190.39 s

The five failures are the same five, by name, on both sides and against gpu_gate_known_failures.txt — both diffs of the sorted failing node-ids are empty, control-vs-branch and branch-vs-file:

tests/test_dcp_merge_ops.py::test_row_view_matches_output_slicing_bitwise
tests/test_fused_compress_ragged.py::...[extend0-context0-cut+whole]
tests/test_fused_compress_ragged.py::...[extend1-context1-whole+cut]
tests/test_fused_compress_ragged.py::...[extend2-context2-resume+fresh]
tests/test_fused_compress_ragged.py::...[extend4-context4-tiny-then-long]

Four are one bf16 ULP (max|diff| = 0.001953125 = 2^-9) and the fifth is the bitwise torch.equal. Toolchain matched the baseline's exactly: torch 2.10.0+rocm7.2.4.git3d3aa833, hip 7.2.53211, aiter v0.1.21.dev0-49-gf4e7c7509. Run on device 3 — devices 0 and 1 were 83% resident with someone else's job, and the gate's own bookable check offered 2 3 4 5 6 7.

Both sides' pass counts are accounted for by the gate's own arithmetic rather than by a floor: 5,340 = 4,779 baseline + (610 − 49) tests/compass and 5,345 = 4,779 + (615 − 49).

Effort

Measured against 24d75742d, at head 70f8cd4db. Two AST conventions are given because round 1's figure and the review's differed by what a counter does with docstrings; production is 0 net either way.

AST statement lines (docstrings counted) AST statement lines (docstrings excluded) physical non-blank
production (atom/utils/forward_context.py) 0 net (352 → 352) 0 net (335 → 335) +6 (851 → 857)
test (tests/compass/test_capture_real_model.py) +203 (577 → 780) +189 (540 → 729) +833 (1,307 → 2,140)
design documents — — +103 / −3 (prose)

Diff against 24d75742d: +7/−1 production, +970/−45 test, +102/−2 04 and +1/−1 12.

Correction, so the number in this body is the number at head. Three figures here were first published as measured at 2b2a48370, an earlier amend of this branch, and were stale by two amends: the test numstat read +965/−45, the design numstat +101/−3, and the test's physical non-blank +828 (1,307 → 2,135). The corrected figures are the ones in the table and the line above. The AST figures did not move and are as first published — production 0 net on both conventions, test +203 / +189 — so the only thing that changed is three physical-line counts.

Comments were not trimmed and were not meant to be, and the split says so rather than leaving it to a ratio. Of the test file's +833 physical non-blank lines: comments +164, docstrings +285, everything else +384 — so 54% of what was added is explanation, which is what six of the eleven findings asked for. Of the remaining 384, only 189 are AST statements that are not docstrings (203 if the 14 new docstrings are counted as statements); the rest are continuation lines, most of them one-per-entry rows of the two declared tables, OP_FAMILIES and EXPECTED_HOST_RESOLUTIONS. The 400-production-line halt threshold was never approached, because the production change is one line — which is itself the result.

Reproducing

pytest tests/compass/test_capture_real_model.py                                  # 14 tests
python tests/compass/test_capture_real_model.py --tp 2 --step-symbol             # the TP2 record
python tests/compass/test_capture_real_model.py --tp 1 --step-symbol --width 8   # the second width
python tests/compass/test_capture_real_model.py --tp 1                           # the control
python tests/compass/test_capture_real_model.py --tp 1 --step-symbol --width 1   # refused, rc=2

14 tests, 8 subprocess captures at ~13 s each plus one subprocess that is refused, ~108 s.

Design documents

  • T81 no longer says the capture is concrete. It says what is closed — the decode step at TP1 and TP2, with the census by family — what the mechanism is, that the log is asserted against a declared set, that "a symbolic CpuGpuBuffer" was measured not to be required and that the ablation's code is not in the tree, what site three turned out to be, that the guard set is hint-dependent, and what the symbol does not survive. It stays open for what this does not cover.
  • 04 D18 gains a fourth capture discipline: the three existing ones keep the tracing symbolic and are not enough on a real engine, because the engine's host arithmetic asks the symbol for a number and guard_int records the ask. Round 2 also retitles ### The three disciplines to four, says in that section's lead that a reader who stops at the numbered list has three of four, and updates D18's decision-log row, which had not been touched.

What is not done

  • Prefill is untouched. Only a decode step is traced, and a prefill step's token count is not its sequence count, so the one-symbol argument does not carry over. It needs its own measurement.
  • The symbol does not survive ScheduledBatch.__init__. The constructor compares the staged token array's length against the count it was handed, which solves the width by comparison rather than conversion — the __bool__ boundary this PR leaves alive on purpose. So the capture builds the batch at the concrete width and rebinds its four count fields afterwards. Round 1 disclosed this only in _decode_batch's docstring while the body said the symbol "enters at the ScheduledBatch"; it enters ATOM's staging, not ATOM's batch constructor.
  • replacements == {} proves less in this pass than it looks. __index__ and __int__ are replaced by _resolve_on_the_host for the capture's duration, so for those two routes the assertion is guaranteed by construction. What it still tests is every other route into _set_replacement — a comparison, a hash, a format string — and those are live. The load-bearing evidence is the two-width digest.
  • The digest covers operator names and tensor shapes only — not scalar arguments, dtypes or strides. A value-level difference is found by comparing distinct_ops, which is how lift_fresh → scalar_tensor was found.
  • Speculative decode, ragged steps and max_seqlen_q > 1 are untraced; the batch here has max_seqlen_q == 1.
  • Three is what the instrument reaches, not a claim that three is all there are. The host-resolution log now makes a fourth site fail rather than hide, but only for this model on this step.
  • The inventories are still DIAGNOSTIC, unchanged: raw @triton.jit launches are recorded and not executed, so anything downstream of a skipped kernel read uninitialised fake memory. These counts are not a cost-model input at any width.
  • Two declared substitutions alter what is recorded — c10d's legacy collectives routed to their functional forms, and ATOM_USE_CUSTOM_ALL_GATHER=0, which is not ATOM's default path. Both are on the control as well, so the census compares like with like, but the operator lists carry them.
  • Correction to round 1's record. Round 1 claimed to have removed two design-document references from the test file's prose. That work was already done on the current base — 24d75742d removed both, with different wording, and the restack adopted the base's. This branch removes none; the identifier sweep over the changed files and over tests/compass/, atom/compass/**/*.py and scripts/compass/ is clean, and was watched firing on 9fcd6c7bd, where it returns four hits.
  • ruff format --check, ruff check and black --check are clean on both changed files, except one pre-existing BLE001 in forward_context.py:938 that is present at 24d75742d and unrelated to this change.

🤖 Generated with Claude Code

@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Review — CAP-2, round 1

Verdict: findings requiring a round 2. Eleven, numbered below. None of them challenges the result. The central claim — a decode step of the published 27B whose width is a free symbol traces through ATOM's own runner, model classes and staging with replacements empty and free symbols on the GEMM, attention and normalisation operators — reproduces, and I could not break it. The findings are about claims the instrument does not actually make, a correction this branch overwrites, and the restack.

Everything below was re-measured from scratch on node 18 against the stated base 9fcd6c7bd, with both trees staged by snapshot.sh + docker cp into paths of my own and gated with their own scripts/compass/.


What I reproduced

The two-width digest identity holds, and I could not break it. I ran the step-symbol capture at eight widths, not two:

TP widths run digest ops shape entries non-numeric replacements specialisations
1 2, 3, 5, 7, 8, 16, 64 a7c1ec1b6ac1ec4b… identical at all seven 2,521 12,544 5,016 {} 0
2 2, 5, 8 a37d76da81cd4547… identical at all three 2,662 13,107 5,297 {} 0

concrete_dims is identical at every one of them. A width that is not a power of two (3, 5, 7) changes nothing; a width an order of magnitude above the hint (64) changes nothing. The 23 distinct concrete values reproduce exactly — [0, 1, 3, 4, 24, 48, 128, 256, 512, 513, 1024, 5120, 6144, 8192, 10240, 14336, 16384, 16385, 16480, 17408, 34816, 49152, 248320], and neither 2 nor 8 nor any other width tried appears in it. The symbol is free. Width 1 is the one exception and it is finding 5 — the axis is never a symbol there at all, which is torch's documented 0/1 specialisation, not a leak.

The census reproduces row for row, at both widths, against the PR's table: TP1 gemm 257/514, attention 64/304, normalisation 273/802, total 2,521 / 5,016 of 12,544 against a control of 2,521 / 0; TP2 gemm 257/514, attention 64/304, normalisation 273/802, collective 133/266, total 2,662 / 5,297 of 13,107 against 2,662 / 0. tp_group_world_size == 2 with apply_simulated_tp_calls empty on both sides, 133 collectives dispatched. I classified all 33 distinct operators by hand; unclassified is genuinely empty today.

__bool__ is genuinely untouched. Only torch.SymInt.__index__ and __int__ are replaced; __bool__, __eq__, __hash__, __format__ are not, and ShapeEnv._set_replacement is wrapped and called through, not suppressed. There is positive evidence rather than only the absence of a patch: at every width other than 2 a guard N + 1 > <hint> appears, which is a __bool__ comparison installing a guard on the live symbol. That is finding 4, but it is also the proof that the boundary is intact.

The design doc's conversion table is exactly right. I re-derived host_resolutions independently: 20 conversions over 16 distinct ATOM lines, matching every row of 04's new table — prepare_decode 1106×2, 1115, 1121, 1122, 1123×2, 1131, 1132×2 (10); prepare_inputs 2441, 2452×2, 2454 (4); prepare_input_ids 510, 513 (2); prepare_sample 2537 (1); _mrope_cpu_view 398, 400 (2); gdn_attn 1237 (1). The table is a measurement and it is correct. Nothing in the tree keeps it correct — that is finding 1.

Both ablations reproduce exactly. Restoring int(t.shape[0]) in _rows: 5 of 13 fail. Suppressing the host resolution so guard_int runs: 4 of 13 fail. Unmodified branch: 13 passed in 92 s.

Route 1's negative result holds — I rebuilt it. I wrote a fake, non-numpy-backed CPU side into _staged_allocators from scratch and ran the step-symbol capture against it. Result: bookkeeping operators 83 → 121, i.e. exactly +38 prim.device calls, and every family's non-numeric shape-entry count unchanged (5,016 → 5,016, identical per family), same guards, replacements still empty. The negative result is real and the one-line production change is the right conclusion from it. See finding 8 for the two things wrong with how it is reported.

The one-line production change is inert. I ran an adversarial differential harness on both trees: torch.Size.__getitem__ returns a Python int for every real tensor I could construct — sizes 0/1/2/384/65536, expanded, sliced, permuted, bool, empty, and meta — and the full AssertionError text from every failing branch of assert_shape_contract is byte-identical on both trees, including a numpy-array attribute, a zero-row tensor and a 0-dim tensor (same IndexError both sides). assert_shape_contract has exactly one caller, model_runner.py:2863 inside eager run_model, outside any torch.compile or cudagraph-captured region, and the line directly above it already reads input_ids.shape[0] raw. I accept "nothing runs differently in production".

The pin does not assert less than it did. SITE_THREE drops from two frames to one (site(three, 2) → site(three, 1)), but the frame it drops is _rows, and it is replaced by a positive assertion that _rows appears nowhere on the path — strictly a re-point, not a loosening. :437 → :444 is arithmetically right: _rows grew by six comment lines and :444 is the assert slot_rows in (None, self.running_tokens) line, which is now where the comparison happens. _Specialisations was widened to cover _decode_batch as well, which strengthens it. test_the_inventory_is_concrete_at_both_widths, test_the_first_two_specialisation_sites_are_where_they_were_measured and test_atom_s_own_buffer_constructor_is_what_runs are untouched by the diff and pass.

Gates, both tiers, my own runs.

CPU control CPU branch GPU control GPU branch
passed 4,566 4,570 (+4) 5,340 5,344 (+4)
failed 0 0 5 5
skipped 149 149 105 105
xfailed 3 3 3 3
RC 0 0 0 0

Every figure in the PR is confirmed. The GPU by-name comparison is empty both ways: diff of the sorted failing node-ids control-vs-branch is empty, and branch-vs-gpu_gate_known_failures.txt is empty — the same five node-ids, not five of a different set. Toolchain matched the baseline exactly (torch 2.10.0+rocm7.2.4.git3d3aa833, hip 7.2.53211, aiter v0.1.21.dev0-49-gf4e7c7509), PREFLIGHT_RC=0 on both sides. The CPU tier's known ±1 flake did not fire — skip counts identical on both sides. atom/utils/forward_context.py is genuinely absent from gpu_gate_triggers.txt and tests/test_forward_mode.py is genuinely in the CPU tier, so "the CPU tier covers it, and the GPU tier was run anyway" is correct.

Effort, reported not adjudicated. Diff +7/−1 production, +659/−41 test, +67/−1 design — all four numstat figures confirmed. Physical non-blank confirmed exactly: production 851 → 857 (+6), test 1,303 → 1,853 (+550). My AST statement count differs in convention (352 → 352 production, 577 → 737 test, so +160 rather than +150) — the production figure is 0 net either way and the test delta differs only by what a counter does with docstrings. The 400-line halt threshold was nowhere near.

Identifier sweep: clean on this head. Zero matches for D<n>, T<n>, P0.<n>, W<n>.<n>, "principle N" or "Gate N" in atom/utils/forward_context.py and tests/compass/test_capture_real_model.py, and zero across all of tests/compass/, atom/compass/**/*.py and scripts/compass/. The two removals claimed are real. See finding 3 for why one of them was already done on the current base.


Findings

1. The host-resolution log is not asserted against a declared set, and three places say it is. The PR body says "asserted against a declared set, so a conversion at a line this capture did not expect fails the test"; 04's new section says "asserts that log against a declared set"; _resolve_on_the_host's own docstring says "a conversion at a site this capture did not expect appears here and the test fails on it". The only assertions on host_resolutions are a membership test (SITE_ONE[1][0] in resolved_at) and a negative one (no _rows frame). A conversion at a seventeenth ATOM line adds an entry and nothing fails. The instrument's headline safety property — it fails rather than hides — is stated three times and implemented nowhere. 04's 16-line / 20-conversion table, which I verified is exact, has no pin at all. Fix: assert the multiset of (file:line, count) against a module constant, the way EXPECTED_GUARDS and STAGED_COPIES already are. Mitigating fact worth stating rather than relying on: a new conversion that reaches a shape would be caught indirectly by the two-width digest; one that only reaches a host value would not.

2. This branch restates a factual error that its own base has already corrected. tests/compass/test_capture_real_model.py:60 and :1933 both say the bound is solved "fourteen lines later". CAP-1 round 3 (24d75742d) corrected exactly that sentence: site one is aiter_attention.py:1115 in prepare_decode; site three is forward_context.py:444 in assert_shape_contract, reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned — another file and a later phase of the step. I confirmed the call graph. The restack must adopt that correction, not overwrite it.

3. The base drift is not benign, and the PR's account of the identifier removals is stale. The PR declares base compass/cap-1 at 9fcd6c7bd; that branch is now 24d75742d. That commit is one commit, 1/1 in 12_open_items.md and +24/−20 in tests/compass/test_capture_real_model.py — and every hunk collides with this PR: both rewrite the same single T81 line; both rewrite the module docstring's Where the shapes specialise block; both rewrite test_closing_site_one_moves_the_bound_to_a_third_site's docstring; and both remove principle 8 from _watch_simulated_tp and T81 from test_atom_s_own_buffer_constructor_is_what_runs, with different wording. So "It removed two design-doc references" describes work the current base has already done. Restack against 24d75742d, resolve five hunks by hand, and re-state the removals against what is actually left.

4. The guard set is hint-dependent, and the applicability table is the hint-2 case only. EXPECTED_GUARDS and the PR's Guards — the applicability table give three guards at TP1 and two at TP2. Measured: at every step width other than 2 a fourth guard appears — <axis> + 1 > 3, > 5, > 7, > 8, > 16, > 64 respectively, and at TP2 as well. It is elided at hint 2 only because a size symbol's default range [2, ∞) makes it vacuous. Three consequences: the applicability statement ("covers every decode step up to the narrowest of them") acquires a lower bound at any hint but 2 and does not say so; EXPECTED_GUARDS is asserted only at the default width so the pin cannot see it; and the two-width evidence table lists five rows as "identical" while omitting the one artifact that is not identical between widths 2 and 8. The result is unaffected — the guard is a __bool__ comparison, which is the boundary the PR deliberately left alive — but it should be reported, and the two-width test should assert what it expects the guards to do rather than leave them out.

5. --width 1 produces a record that reads as a successful step-symbol capture and is fully concrete. python tests/compass/test_capture_real_model.py --tp 1 --step-symbol --width 1 exits 0 and emits "step_symbol": true, "step_axis": "1", "non_numeric_shape_entries": 0, "host_resolutions": [], "shape_env_guards": [] — a record whose family census is byte-for-byte the concrete control's. 04 D18's third discipline and _decode_batch's own docstring both know why (torch specialises a hint of 1 silently), and test_nothing_specialises… would catch it via re.fullmatch(r"s\d+", step_axis) — but only at the default width, and the CLI path has no such check. Principle 6: main() or _step_axis should refuse a width below 2 rather than emit a diagnostic that looks exactly like the failure mode this whole task exists to detect.

6. The evidence that the symbol is free is measured at TP1 only, and TP2 is the named result. test_the_symbol_is_free_across_the_step_width calls capture(1, …) twice; nothing runs the two-width comparison at TP2, although the PR's headline census is the TP2 one and TP2 is where a capture "has most to lose" by the file's own argument. I ran it: TP2 at widths 2, 5 and 8 gives one digest, a37d76da81cd4547…, with identical concrete_dims — so this is a coverage gap, not a risk. One extra capture(2, step_symbol=True, width=SECOND_WIDTH) closes it.

7. op_family's prefix matching defeats the "unclassified must be empty" guard it is built around. The comment says a new operator is "a classification somebody makes, not a bucket it falls into quietly", enforced by assert "unclassified" not in at_width. But op_family uses name.startswith(prefixes), so any operator sharing a prefix with a declared one is classified silently: ("aten.add",) under elementwise captures aten.addmm and aten.addbmm, which are GEMMs; aten.slice under view captures aten.slice_scatter, which writes; aten.select captures aten.select_scatter. Today's 33 operators are all classified correctly — I checked each one — but the guard cannot see the case it is written for, and a future ATOM change that routes a projection through addmm would land in elementwise and understate the three cost-bearing families.

8. Route 1's third ablation row has no artifact, and "not one shape entry" is loose. Two of the three ablation rows are reproducible from this diff; the third ("make the symbolic buffer's CPU side fake instead → 13 pass") is not in the tree in any form, and neither is the code that produced "38 prim.device calls". Principle 8: it is the row that decides the route, so it is the one that most needs its source. I rebuilt it and it reproduces — but that is my reconstruction, not the PR's evidence. Separately, "changes not one shape entry" is not quite what happens: total shape entries rise 12,544 → 12,585, because the 38 extra prim.device calls carry entries of their own. What is unchanged is every family's non-numeric count. Say that.

9. ScheduledBatch.__init__ is stepped around, not shown clean, and only a test docstring says so. The PR body says the symbol "enters at the ScheduledBatch and ATOM derives every width from it". In fact _decode_batch builds the batch at the concrete width and rebinds the four count fields afterwards, because __init__ compares the staged array's length against the count — the eighteenth site, the one that solves by comparison. That is the __bool__ boundary the PR says it left intact, and here it had to be avoided. It belongs in What is not done and in T81, not only in _decode_batch's docstring: a reader should know the symbol does not survive ATOM's own constructor.

10. replacements == {} is a weaker statement in this pass than its docstring claims. test_nothing_specialises_under_the_step_symbol_capture says any route that solves a symbol — "by __index__ or __int__ or a comparison or a hash or a format string" — reaches _set_replacement and appears there. In the step-symbol pass the first two have been replaced by _resolve_on_the_host, so they cannot reach that funnel by construction. What carries the claim is the two-width digest, not replacements. Re-word so the load-bearing evidence is named as the load-bearing evidence. (Related scope note, not a defect: the digest covers operator names and tensor shapes only, not scalar arguments, dtypes or strides — the single control-vs-symbol difference the PR reports, lift_fresh → scalar_tensor, was found by comparing distinct_ops, which is the right way and worth saying.)

11. 04's "three disciplines" heading was not updated, and D18's decision-log row was not touched. The new section is titled A fourth discipline but sits two subsections below ### The three disciplines, all mandatory, all silent when omitted (04:112), with the FakeTensor material in between, and D18's row in the decision log (04:1086) is unchanged. A reader who reads the numbered list and stops does not learn there is a fourth.


For the next task in this area

  • The one-symbol argument is load-bearing and correctly scoped to decode. Prefill is not a matter of repeating this: the token count is not the sequence count, so assert_shape_contract's equalities become two symbols that ATOM's contract equates — which is precisely what the site-three probe already shows happening. Expect that, and expect it to be the interesting part rather than an obstacle.
  • The guard set is the artifact that carries a hint into the record even when no shape does (finding 4). Whatever prefill does, read the guards per width before trusting a digest.
  • _resolve_on_the_host is a global monkeypatch of two torch.SymInt dunders. It is restored in a finally, and every capture is a fresh subprocess, so it is safe here — but it is not composable with anything else that wants those methods, and a future in-process capture will have to own that.

Reviewed against head 3e3c25eda, base 9fcd6c7bd. Not merged, not landed, not undrafted, no label applied.

resolved_at = {
frame for event in record["host_resolutions"] for frame in event["frames"][-1:]
}
assert SITE_ONE[1][0] in resolved_at

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 1 — the host-resolution log is not asserted against a declared set, and three places say it is.

This is the whole of the assertion on host_resolutions: a membership test, plus the negative _rows check below. Nothing compares the log to a declared set.

Three places say otherwise:

  • the PR body — "every conversion, with the ATOM line it happened on, asserted against a declared set, so a conversion at a line this capture did not expect fails the test";
  • 04's new section — "and asserts that log against a declared set";
  • _resolve_on_the_host's own docstring — "a conversion at a site this capture did not expect appears here and the test fails on it".

A conversion at a seventeenth ATOM line adds an entry to host_resolutions and nothing fails. The instrument's headline safety property is stated three times and implemented nowhere, and 04's 16-line / 20-conversion table has no pin at all.

I re-derived that table independently and it is exact — 20 conversions over 16 distinct lines, matching every row (prepare_decode 1106x2, 1115, 1121, 1122, 1123x2, 1131, 1132x2; prepare_inputs 2441, 2452x2, 2454; prepare_input_ids 510, 513; prepare_sample 2537; _mrope_cpu_view 398, 400; gdn_attn 1237). It is a good measurement with nothing keeping it true.

Fix: assert the multiset of (file:line, count) against a module constant, the way EXPECTED_GUARDS and STAGED_COPIES already are.

One mitigating fact worth stating rather than relying on: a new conversion that reaches a shape would be caught indirectly by the two-width digest; one that only reaches a host value would not.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — implemented, not withdrawn.

The log is now asserted as a multiset against a module constant, EXPECTED_HOST_RESOLUTIONS, exactly as you asked: {ATOM line: conversions}, built by host_resolutions_by_line(record) from the innermost ATOM frame of each event.

The declared set is your table, and I re-measured it rather than transcribing it — 20 conversions over 16 lines, and it is the same multiset at TP1 and TP2 and at both step widths (2 and 8), which is how it is asserted:

for tp in (1, 2):
    for width in (DECODE_SEQS, SECOND_WIDTH):
        at = capture(tp, step_symbol=True, width=width)
        assert host_resolutions_by_line(at) == EXPECTED_HOST_RESOLUTIONS, (tp, width)

Asserting the multiset rather than the set also catches the case a set would miss: a line that converts twice where it converted once. The three places that claimed the property now describe what the code does — the PR body, 04's fourth-discipline section (which also names the constant), and _resolve_on_the_host's docstring.

Your mitigating fact is stated rather than relied on, in EXPECTED_HOST_RESOLUTIONS' own comment: a new conversion that reaches a shape would also show as a digest that differs between two widths; one that only reaches a host value would not, which is why this is asserted rather than left to the digest.

Ablated: suppressing _resolve_on_the_host so guard_int runs fails 4 of 14, this test among them.

The second site is untouched by the repair, exactly as recorded. The third
is the one CAP-2 meets the moment its site-one repair lands, and it is in
a module nothing in the design record named before this test.
survives `:1115` and is solved fourteen lines later, inside ATOM's own

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 2 — this restates a factual error the base has already corrected.

"solved fourteen lines later" — and the same phrase at line 60 of the module docstring.

CAP-1 round 3 (24d75742d, now the head of compass/cap-1) corrected exactly this sentence. Site one is aiter_attention.py:1115 in prepare_decode; site three is forward_context.py:444 in assert_shape_contract, reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned — another file and a later phase of the step, not fourteen lines on. I confirmed the call graph.

The restack has to adopt that correction rather than overwrite it; this docstring and :60 are two of the five hunks that collide (see finding 3).

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in the restack. The base moved 9fcd6c7bd → 24d75742d and I rebased with git rebase --onto 24d75742d 9fcd6c7bd compass/cap-2. Five hunks collided, including both of these; each was resolved by hand so the correction is adopted, not overwritten.

  • The module docstring (:57) now reads "into an ATOM assertion helper in another file and a later phase of the step — forward_context.py, reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned", and then continues with this branch's four new paragraphs.
  • This test's docstring carries the same correction and says the negation explicitly: "… forward_context.py:444 in assert_shape_contract, reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned — not fourteen lines after aiter_attention.py:1115."

grep -rn fourteen over the tree now returns that one deliberate negation, one unrelated line in tests/entrypoints/, and 12's D-row about the fourteen torch.cuda stubs. The :444 rather than :437 is this branch's own six added comment lines in _rows, as you read it.

# implied by the per-sequence bounds that both widths install. Asserted per
# width rather than as a shared set, because the difference is a property of
# the step and worth failing on if it changes.
EXPECTED_GUARDS = {

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 4 — the guard set is hint-dependent, so this is the hint-2 case only.

EXPECTED_GUARDS is asserted only at the default step width, and the PR's Guards — the applicability table reports it as the graph's applicability. Measured at eight widths:

step width guards at TP1
2 28*N <= 8192, N < 128, N <= 512
3 …plus N + 1 > 3
5 …plus N + 1 > 5
7 / 8 / 16 / 64 …plus N + 1 > 7 / > 8 / > 16 / > 64

Same at TP2 (N + 1 > 5 at width 5, N + 1 > 8 at width 8). The fourth guard is elided at hint 2 only because a size symbol's default range [2, inf) makes it vacuous.

Three consequences:

  1. the applicability statement — "covers every decode step up to the narrowest of them" — acquires a lower bound at any hint but 2, and does not say so;
  2. EXPECTED_GUARDS is asserted only at the default width, so the pin structurally cannot see it;
  3. the two-width evidence table lists five rows as "identical" while omitting the one artifact that is not identical between widths 2 and 8.

The result is unaffected: this guard is a __bool__ comparison, which is the boundary this PR deliberately left alive — it is in fact the positive evidence that __bool__ is untouched. But it should be reported as measured, and test_the_symbol_is_free_across_the_step_width should say what it expects the guards to do rather than leave them out of the comparison.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed, and reported as measured. You are right that the published table was the hint-2 case and that the pin structurally could not see it.

I re-measured at TP1 hints 2, 3, 8 and 16 and at TP2 hints 2 and 8, and reproduce your reading exactly: the fourth guard is <axis> + 1 > <hint> at every hint but 2, at both group widths, with the digest unchanged (a7c1ec1b… at TP1, a37d76da… at TP2) and non_numeric unchanged at 5,016 / 5,297.

Three changes, one per consequence:

  1. EXPECTED_GUARDS is no longer asserted directly. expected_guards(tp, width) composes it, adding f"<axis> + 1 > {width}" for any width > MIN_STEP_WIDTH, and its docstring says why hint 2 elides it — a size symbol's default range [2, ∞) makes the inequality vacuous — and that it is a __bool__ comparison on a live symbol, i.e. the positive evidence you point out.
  2. test_the_symbol_is_free_across_the_step_width no longer leaves the guards out of the comparison. It asserts the whole set at both widths and asserts the lower bound is present in the wide capture and absent in the narrow one, so the artifact is held in both directions rather than inferred.
  3. The applicability prose in EXPECTED_GUARDS' comment, in 04 and in T81 now says the graph acquires a lower bound at any hint but 2, and the two-width table in the PR body names this as the one row that is not identical.

One more artifact turned up while doing this, which the old assertion had also been hiding: at TP2 the value 2 is present 73 times in concrete_dims — it is the group's width, not the step's. The old blanket DECODE_SEQS not in values only ever ran at TP1. What separates the two readings is the line above it: concrete_dims is identical at step widths 2 and 8, so a 2 that survives an eight-wide step is an engine constant. The test now asserts SECOND_WIDTH not in values at both TPs (the direction that can fail), DECODE_SEQS not in values at TP1, and the presence of the group width at TP2 with the reason stated.

parser.add_argument("--symbolic", action="store_true")
parser.add_argument("--repair-site-one", action="store_true")
parser.add_argument("--step-symbol", action="store_true")
parser.add_argument("--width", type=int, default=DECODE_SEQS)

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 5 — --width 1 emits a record that reads as a successful step-symbol capture and is fully concrete.

python tests/compass/test_capture_real_model.py --tp 1 --step-symbol --width 1 exits 0 and writes:

"step_symbol": true, "step_axis": "1", "step_axis_hint": 1,
"non_numeric_shape_entries": 0, "host_resolutions": [],
"shape_env_guards": [], "shape_env_replacements": {}

Its family census is byte-for-byte the concrete control's (view 922, allocation 565, bookkeeping 83). 04 D18's third discipline and _decode_batch's own docstring both know why — torch specialises a hint of 1 silently — and test_nothing_specialises… would catch it through re.fullmatch(r"s\d+", step_axis), but only at the default width. This CLI path has no such check.

Principle 6: refuse a width below 2 here (or assert in _step_axis that what came back is a symbol), rather than emit a diagnostic record that looks exactly like the failure mode this task exists to detect.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — it refuses. main() now rejects --width below MIN_STEP_WIDTH (2) before any capture runs:

$ python tests/compass/test_capture_real_model.py --tp 1 --step-symbol --width 1
test_capture_real_model.py: error: --width 1 traces no symbol: torch specialises a size
hint of 0 or 1 to a constant, so the record would be concrete and would say it was
symbolic. The narrowest traceable width is 2.
rc=2

I did both halves of what you offered, because the CLI is not the only entry point: _step_axis also refuses if what comes back is not a symbol, tested as re.fullmatch(r"s\d+", str(axis)) — the same way every record in the file is read, so the check and the assertions downstream of it agree by construction rather than by type.

And the refusal is now fail-able rather than asserted in prose: test_the_capture_refuses_a_width_that_torch_would_specialise runs the real entry point in a subprocess, asserts a non-zero exit, asserts the message, and asserts RECORD_MARKER is not in stdout — a refusal that still printed a record would be worse than the emission it replaced. That is the 14th test.

Your reading of why is what the constant's comment now says: this is torch's documented 0/1 specialisation and not a leak, but a pass reporting success while producing a concrete record is the shape of the thing this file exists to catch.

a step width in disguise. They are the model's and the engine's -- hidden
sizes, head counts, projection widths, and the staging buffers' capacities.
"""
narrow = capture(1, step_symbol=True)

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 6 — the evidence that the symbol is free is measured at TP1 only, and TP2 is the named result.

Both captures here are capture(1, …). Nothing runs the two-width comparison at TP2, although the headline census is the TP2 one and TP2 is where this file's own argument says a capture has most to lose.

I ran it: TP2 at widths 2, 5 and 8 all give one digest, a37d76da81cd4547…, with identical concrete_dims, 2,662 operators and 5,297 non-numeric entries. So this is a coverage gap rather than a risk — one extra capture(2, step_symbol=True, width=SECOND_WIDTH) closes it.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — TP2 evidence published, and it is now the test rather than a report. test_the_symbol_is_free_across_the_step_width runs for tp in (1, 2) and takes capture(tp, step_symbol=True) against capture(tp, step_symbol=True, width=SECOND_WIDTH).

My own re-measurement agrees with yours: TP2 at widths 2 and 8 gives one digest a37d76da81cd4547…, 2,662 operators, 13,107 shape entries, 5,297 non-numeric, identical concrete_dims, replacements empty. (TP1 at 2, 3, 8 and 16: one digest a7c1ec1b6ac1ec4b…, 2,521 / 12,544 / 5,016.)

It cost one extra subprocess capture — (tp=2, width=8) — which is shared with finding 1's host-resolution assertion, so the file is 14 tests in ~108 s rather than 13 in ~92.

The one thing that is not identical between TP1 and TP2 here is the concrete 2: at TP2 it appears 73 times as the group's width. Detail in my reply to finding 4; the short form is that it is identical at step widths 2 and 8, which is what says it is not the step's.

("embedding", ("aten.embedding", "aiter.masked_embedding")),
("collective", ("aiter.all_reduce_", "_c10d_functional.")),
("sampling", ("aiter.mixed_sample_outer_exponential", "aten.exponential_")),
("elementwise", ("aten.add", "aten.mul", "aten.fill_")),

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 7 — prefix matching defeats the "unclassified must be empty" guard this is built around.

OP_FAMILIES' comment says a new operator is "a classification somebody makes, not a bucket it falls into quietly", enforced by assert "unclassified" not in at_width. But op_family matches with name.startswith(prefixes), so any operator sharing a prefix with a declared one is classified silently and never reaches unclassified:

  • ("aten.add", …) here captures aten.addmm and aten.addbmm — both GEMMs;
  • aten.slice under view captures aten.slice_scatter, which writes;
  • aten.select under view captures aten.select_scatter.

Today's 33 distinct operators are all classified correctly — I checked each one — but the guard cannot see the case it is written for. An ATOM change that routes a projection through addmm would land in elementwise, and the three cost-bearing families would understate with nothing failing.

Match on the full operator name, or keep the prefixes and add an explicit overlap check.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — matching is on whole operator names now, and I re-derived the table.

def op_family(name):
    parts = name.split(".")
    base = ".".join(parts[:2]) if len(parts) == 3 else name
    for family, names in OP_FAMILIES:
        if base in names:
            return family
    return "unclassified"

str(func) is namespace.operator.overload; the overload says how an operator was called rather than what it does, so it is dropped and the rest is matched exactly. OP_FAMILIES is now a table of whole names — which meant enumerating the _c10d_functional.* and profiler.* entries that used to be namespace prefixes, and adding aten.zeros, which the probe passes dispatch and which the prefix table did not classify at all.

Whether the headline numbers move: they do not. Re-derived on node 18 against the same records, family census compared structure-for-structure before and after:

pass family census identical digest
TP1 control yes 235e44b8…
TP1 step-symbol, width 2 yes a7c1ec1b…
TP1 step-symbol, width 8 yes a7c1ec1b…
TP2 step-symbol, width 2 yes a37d76da…
TP2 step-symbol, width 8 yes a37d76da…

gemm 257/514, attention 64/304, normalisation 273/802, collective 133/266 at TP2 — every row unchanged, as your own check of the 33 operators predicted. What changed is that the bucket can now fire for the case it was written for.

Your three examples are named in the comment above OP_FAMILIES and in the test's docstring, so the next reader meets the reason rather than the rule.

way, because the symbol is the *width*, not the content.

The batch is built at the concrete width and the four count fields are
rebound afterwards, rather than the symbol being passed to the constructor.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 9 — this is the only place that says the symbol does not survive ATOM's own constructor.

The PR body says the symbol "enters at the ScheduledBatch and ATOM derives every width from it". What actually happens is here: the batch is built at the concrete width and the four count fields are rebound afterwards, because __init__ compares the staged array's length against the count — the eighteenth site, the one that solves by comparison rather than conversion.

That is the __bool__ boundary the PR says it deliberately left intact, and here it had to be stepped around rather than shown clean. The docstring is honest about it; the PR body and T81 are not. It belongs in What is not done: a reader should know that passing the symbol to the constructor specialises it, and that this is why the capture rebinds.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — it is where a reader meets it now, in three places rather than one.

  • 04's fourth-discipline section has a paragraph of its own: "It does not survive the batch's own constructor, and that is the eighteenth site." — __init__ compares the staged token array's length against the count, which solves by comparison rather than conversion; the capture builds the batch at the concrete width and rebinds the four count fields; "the symbol enters the engine's staging; it does not enter the engine's batch constructor, and a reader who takes 'the symbol enters at the ScheduledBatch' literally will look for it in the wrong place."
  • T81 carries What the symbol does not survive, and lists the constructor in Still open alongside prefill.
  • The module docstring's mechanism paragraph now ends "…with one exception, ScheduledBatch.__init__ itself, which solves the width by comparison rather than conversion, so the batch is built at the concrete width and its four count fields are rebound afterwards (see _decode_batch)."

The PR body's What is not done says the same, and the sentence you quote — "the symbol enters at the ScheduledBatch and ATOM derives every width from it" — is rewritten so it no longer claims what the constructor does not do.

through -- its own docstring says to use it rather than assigning into
`replacements` -- so an empty `replacements` at the end is not a statement
about the routes this capture happened to think of. Any route that solves a
symbol, by `__index__` or `__int__` or a comparison or a hash or a format

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 10 — replacements == {} is a weaker statement in this pass than this says.

"Any route that solves a symbol, by __index__ or __int__ or a comparison or a hash or a format string, solves it through that funnel and appears here."

In the step-symbol pass __index__ and __int__ have been replaced by _resolve_on_the_host, so those two routes cannot reach _set_replacement by construction. For them the assertion is guaranteed rather than tested.

What carries the claim is the two-width digest, which is in the test below and which I reproduced at seven widths. Re-word so the load-bearing evidence is named as the load-bearing evidence, and so this paragraph claims only what it can fail on.

Related scope note, not a defect: graph_digest covers operator names and tensor shapes only — not scalar arguments, dtypes or strides. The single control-vs-symbol difference the PR reports, lift_fresh → scalar_tensor, was found by comparing distinct_ops, which is the right instrument for a value-level difference and worth saying explicitly.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — re-worded so the paragraph claims only what it can fail on, and the load-bearing evidence is named as such.

The docstring now says, in these words: two of the routes into the funnel, __index__ and __int__, are replaced by _resolve_on_the_host for the duration of the capture, so for those two the assertion is guaranteed by construction rather than tested; what it still tests is every other route — a comparison, a hash, a format string, anything inside torch that decides it needs a number — and those are live and able to fail the line. Then: "The load-bearing evidence is the two-width digest", with the pointer to test_the_symbol_is_free_across_the_step_width and the reason (a symbol lost rather than solved records no replacement, and only the two-width comparison sees that class).

The receiving test's docstring says the same from the other end — it opens by naming itself as the load-bearing evidence — so a reader arriving at either one is sent to the other.

Your scope note is now stated in that docstring too, as a paragraph and not a parenthesis: the digest covers operator names and tensor shapes only, not scalar arguments, dtypes or strides; the one control-versus-symbol difference this file reports, lift_fresh → scalar_tensor, is a value-level difference and was found by comparing distinct_ops, which is the instrument for it. The same sentence is in 04.

3. The torch 2.10 signature is
`symbolic_context=StatelessSymbolicContext(dynamic_sizes=[...])`, not `dynamic_dims=`.

### A fourth discipline: the engine's host arithmetic asks a symbol for a number

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 11 — the numbered list this extends still says "three", and D18's decision-log row was not touched.

This section is titled A fourth discipline, but ### The three disciplines, all mandatory, all silent when omitted is at 04:112 with the FakeTensor, not bare meta material in between, and D18's row in the decision log at 04:1086 is unchanged. A reader who reads the numbered list and stops does not learn there is a fourth.

Also in this section, at line 210: "and asserts that log against a declared set" — see finding 1, it does not.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed, all three.

  • 04:112 is now "The four disciplines, all mandatory, all silent when omitted", followed by a lead paragraph saying that three of them keep the tracing symbolic and are listed below, that the fourth is about what the engine around the trace does to a symbol, that it needs the FakeTensor material to state and so is a section of its own, and — explicitly — that "a reader who stops at the end of this numbered list has three of four."
  • The decision-log row for D18 now records the fourth discipline and carries its own date: "four disciplines, not three — the fourth is that the engine's host arithmetic asks a symbol for a number, so the capture keeps the conversion, logs it with the line it happened on, and asserts that log against a declared set | 2026-09-18, fourth discipline 2026-09-22".
  • Line 210's "and asserts that log against a declared set" is now true — see finding 1. The sentence is also expanded to say why the clause is the whole of the discipline's value: a log nothing compares records a conversion at a seventeenth line and says nothing, so the instrument reads as evidence while behaving as decoration. It names EXPECTED_HOST_RESOLUTIONS and states that it is asserted at both group widths and both step widths.

A decode step of the published Qwen3.8-27B whose width is a free symbol traces
through ATOM's own ModelRunner, model classes and staging with
ShapeEnv.replacements empty and no symbol solved anywhere, at TP1 and at TP2.

The census, by operator family, never as a total. At TP1, 5,016 of 12,544 shape
entries are not plain integers: 514 on 257 aiter.gemm_a16w16, 304 on the 64
attention calls, 802 on 273 normalisation operators. At TP2, 5,297 of 13,107,
the same three families plus 266 on the 133 collectives. The control is the
same file with the width left an integer: 0 of 12,544 and 0 of 13,107. Both
passes record the same operators -- 2,521 and 2,662 -- differing in one call
that materialises a symbol where it used to lift a constant.

The three ordered sites are symptoms of one line of torch: SymInt.__index__ and
__int__ are guard_int. 16 ATOM lines convert the step's width to a number
during one traced forward -- 20 conversions, all host fills and slices in
prepare_decode, prepare_inputs, prepare_input_ids, prepare_sample,
_mrope_cpu_view and gdn_attn -- with a seventeenth in forward_context._rows and
an eighteenth, ScheduledBatch.__init__, that solves the width by comparison.
Closing them one at a time relocates rather than converges, which is what the
third site already recorded. The capture keeps the conversion, to the hint,
which is the count ATOM computed, and replaces the recording with its own log
of every conversion and the line it happened on. __bool__ is untouched, so a
branch on a width still guards.

One production line changes: assert_shape_contract's _rows returns t.shape[0]
rather than int(t.shape[0]). torch.Size.__getitem__ already returns a Python
int for a tensor with a real size, so it is the identity in a served step; it
matters only where the size is symbolic, and there converting one side of an
equality to a number forces the other to become it.

A symbolic CpuGpuBuffer is not required, and was measured not to be. The
-> 16384 recorded inside copy_to_gpu is a buffer capacity being solved against
the CPU side's constant, and exists only because that probe symbolises every
dimension of the staged device tensor. Capacities are engine configuration, not
a function of the step. With only the width symbolic the two slices carry the
same symbol and copy_to_gpu dispatches unchanged; building the CPU side fake
instead costs 38 prim.device calls ATOM never makes and changes not one shape
entry. Nothing under atom/utils/__init__.py is touched.

The symbol is shown free rather than assumed free: the step is traced at two
widths, 2 and 8, and the inventories are identical once the symbol's name is
set aside -- same operators, same entries, same digest, same concrete
dimensions by value. That is also what says every dimension that stayed a
number is not a width in disguise. Three guards at TP1 and two at TP2, all
staging capacities read against the width, none an equality.

The existing assertions are re-pointed rather than removed. The three sites are
still held, by what now happens at each: the first appears in the
host-resolution log and in nothing else, the second as the call site of ten
symbolic copies, the third nowhere. SITE_THREE moves one frame out, to
assert_shape_contract itself: _rows no longer converts, and what is left is the
assertion equating the probe's two symbols, which is ATOM's contract working
and is also why the step-symbol pass uses one symbol for the whole step.

12_open_items.md: T81 states what is closed -- the decode step at both widths,
with the census by family -- and what is not: prefill, chunked prefill,
speculative decode, and any structure where the token count is not the sequence
count. 04 gains a fourth capture discipline for the guard_int finding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Restacked onto compass/cap-1's round-3 head, which corrected the
"fourteen lines later" account of the third site; five colliding hunks
resolved so the correction is adopted rather than overwritten.

The host-resolution log is now asserted against a declared set,
EXPECTED_HOST_RESOLUTIONS -- 20 conversions over 16 ATOM lines, the same
multiset at TP1 and TP2 and at both step widths. Until now three places
said the log fails on a conversion it did not expect and nothing
implemented it.

The census matcher matched operator names by prefix, so aten.add
captured addmm and addbmm, which are GEMMs, and the "unclassified must
be empty" guard could not fire for the cases it exists for. It matches
whole names now. Re-derived: no family's numbers move.

Also: the guard set is hint-dependent -- a fourth guard <axis> + 1 >
<hint> appears at every step width but 2, at both TP1 and TP2 -- so the
guards are computed per width and the two-width test asserts the one
artifact that differs between widths. The two-width comparison now runs
at TP2 as well as TP1. --width 1 is refused rather than emitting a
concrete record labelled step_symbol. And the design record states what
the symbol does not survive: ScheduledBatch.__init__ solves the width by
comparison, so the batch is built concrete and its counts rebound.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Developer record — CAP-2, round 2

Eleven findings, all eleven fixed. None is argued rather than fixed, and one of them — finding 4 — turned out to be hiding a second artifact of the same kind, which is fixed too. The result is unchanged and was not re-litigated — the review could not break it and I did not try to strengthen it.

Head 70f8cd4db, base compass/cap-1 at 24d75742d. Everything below was measured on node 18 by me, in xiaobizh_n18_cpu and xiaobizh_n18, trees staged with snapshot.sh + docker cp into paths of my own, each gated with its own scripts/compass/.


The restack — finding 3, and it was not benign

git rebase --onto 24d75742d 9fcd6c7bd compass/cap-2. Both refs re-read first: fork/compass/cap-1 and local compass/cap-1 are both 24d75742d, and fork/compass/cap-2 was 3e3c25eda. Five hunks collided, exactly the five the review named, and each was resolved by hand rather than by taking a side wholesale:

# where collision resolved as
1 12_open_items.md:97 the same single T81 line, rewritten by both this branch's row, with the base's site-three correction folded into it
2 test_capture_real_model.py:57 — module docstring, Where the shapes specialise both rewrote the block base's corrected sentence, then this branch's four new paragraphs
3 _watch_simulated_tp's docstring both removed principle 8, different wording the base's wording, so the diff against the new base is empty here
4 test_atom_s_own_buffer_constructor_is_what_runs' docstring both removed T81, different wording the base's wording, same reason
5 test_closing_site_one_moves_the_bound_to_a_third_site — docstring and the adjacent comment (git reports it as two regions of one hunk) both rewrote it base's correction plus this branch's What the third site turned out to be

Force-pushed, for the restack only.

Finding 2 is closed by 3 and 4 above. The "fourteen lines later" sentence is gone from :57 and from the site-three test. It appears once in the tree now, as an explicit negation: "…reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned — not fourteen lines after aiter_attention.py:1115." The other two grep hits are an unrelated line in tests/entrypoints/ and 12's row about fourteen torch.cuda stubs.

And the record about the identifier removals is corrected. "It removed two design-doc references" described work the current base had already done. This branch removes none; the diff against 24d75742d contains no principle 8 or T81 line in either direction. The sweep over the two changed files and over all of tests/compass/, atom/compass/**/*.py and scripts/compass/ is clean — and I watched it fire before trusting it: the same pattern on 9fcd6c7bd:tests/compass/test_capture_real_model.py returns four hits (:581, :1387, :1529, :1546) and returns none on 24d75742d.


Finding 1 — the host-resolution log is now asserted against a declared set

Implemented, not withdrawn. EXPECTED_HOST_RESOLUTIONS is a module constant holding the multiset {ATOM line: conversions}, and host_resolutions_by_line(record) reduces a record's log to the same shape from each event's innermost ATOM frame.

I re-derived the table rather than transcribing the review's: 20 conversions over 16 lines, and — this is the part that decided how to assert it — it is the same multiset at TP1 and TP2 and at step widths 2 and 8. So it is asserted at all four:

for tp in (1, 2):
    for width in (DECODE_SEQS, SECOND_WIDTH):
        at = capture(tp, step_symbol=True, width=width)
        assert host_resolutions_by_line(at) == EXPECTED_HOST_RESOLUTIONS, (tp, width)

A multiset rather than a set, so a line converting twice where it converted once fails too. The three places that claimed the property now describe what the code does: the PR body, 04's fourth-discipline section (which names the constant and says why the clause is the whole of the discipline's value), and _resolve_on_the_host's docstring. The review's mitigating fact is stated in the constant's own comment rather than relied on.


Finding 7 — the census classifier, and whether the table moved

op_family matched with str.startswith. It matches whole operator names now — str(func) is namespace.operator.overload, the overload is dropped, the rest is matched exactly against OP_FAMILIES, and anything else is unclassified. That meant turning the _c10d_functional. and profiler. namespace prefixes into enumerated names, and adding aten.zeros, which the two probe passes dispatch and which the prefix table classified nowhere at all.

The table does not move. Re-derived on node 18, family census compared structure-for-structure between the prefix matcher and the exact one, on five passes:

pass family census identical digest
TP1 control yes 235e44b8…
TP1 step-symbol, width 2 yes a7c1ec1b…
TP1 step-symbol, width 8 yes a7c1ec1b…
TP2 step-symbol, width 2 yes a37d76da…
TP2 step-symbol, width 8 yes a37d76da…

gemm 257/514, attention 64/304, normalisation 273/802 at both TPs, collective 133/266 at TP2, totals 2,521 / 5,016 of 12,544 and 2,662 / 5,297 of 13,107 — every published row unchanged, as the review's hand-check of the 33 operators predicted. What changed is that the bucket can now fire for the case it was written for. The union of distinct operators across the eight records this file produces is 41; all 41 classify, none lands in unclassified.

To be exact about the four operators that motivated this, because the short form invites the wrong reading: the repair does not re-file addmm and addbmm into gemm. None of addmm, addbmm, slice_scatter or select_scatter is declared in any family, so under whole-name matching each lands in unclassified, which the tests require to be empty. The behaviour gained is a refusal — an ATOM that routed a projection through addmm turns this file red and makes somebody classify it — not a silent re-filing. That is the stronger property and it is the one being claimed.


Finding 4 — the guard set is hint-dependent, and it hid a second artifact

Re-measured at TP1 hints 2, 3, 8 and 16 and TP2 hints 2 and 8, and the review's reading reproduces exactly:

TP hint guards
1 2 28*<axis> <= 8192, <axis> < 128, <axis> <= 512
1 3 …plus <axis> + 1 > 3
1 8 …plus <axis> + 1 > 8
1 16 …plus <axis> + 1 > 16
2 2 <axis> < 128, <axis> <= 512
2 8 …plus <axis> + 1 > 8

Digest, operator count and non-numeric count are unchanged at every one of them. Three changes: expected_guards(tp, width) composes the set per width and its docstring says why hint 2 elides the lower bound and why the guard is positive evidence that __bool__ is intact; the two-width test asserts the full set at both widths and the lower bound's presence in the wide capture and absence in the narrow one; and the applicability prose in the test, in 04 and in T81 now says the graph acquires a lower bound at any hint but 2.

The second artifact. The old assert DECODE_SEQS not in values and SECOND_WIDTH not in values ran at TP1 only. Run at TP2 it fails — the value 2 is present 73 times in concrete_dims, across allocations, views, transfers and one device read. It is the group's width, not the step's, and what says so is the line above it: concrete_dims is identical at step widths 2 and 8, so a 2 that survives an eight-wide step is an engine constant in the way a head count is. The test now asserts SECOND_WIDTH not in values at both TPs — the direction that can fail — DECODE_SEQS not in values at TP1, and the group width's presence at TP2 with the reason written out.


Findings 5, 6, 9, 10, 11 — in brief

  • 5 — --width 1. It refuses. main() rejects any width below MIN_STEP_WIDTH, and _step_axis refuses again if what comes back is not a symbol (tested as re.fullmatch(r"s\d+", str(axis)), the way every record in the file is read, rather than by type — so the check and the assertions downstream agree by construction). Both halves, because the CLI is not the only entry point. And the refusal is fail-able: the new 14th test runs the real entry point, asserts a non-zero exit, the message, and that no record is printed.
  • 6 — TP2 two-width evidence. test_the_symbol_is_free_across_the_step_width runs for tp in (1, 2). My own measurement agrees with the review's: TP2 at widths 2 and 8 gives one digest a37d76da81cd4547…, 2,662 operators, 13,107 entries, 5,297 non-numeric, identical concrete_dims. One extra subprocess capture, shared with finding 1's assertion.
  • 9 — ScheduledBatch.__init__. Now in three places a reader meets: a paragraph of its own in 04's fourth-discipline section, a What the symbol does not survive clause in T81 with the constructor listed under Still open, and the module docstring's mechanism paragraph. The PR body's What is not done carries it, and the sentence that said the symbol "enters at the ScheduledBatch and ATOM derives every width from it" is rewritten: it enters ATOM's staging.
  • 10 — replacements == {}. Re-worded to claim only what it can fail on: for __index__ and __int__ the assertion is guaranteed by construction in this pass, and what it still tests is every other route into _set_replacement. The load-bearing evidence is named as such — the two-width digest — in both docstrings, each pointing at the other. The scope note is promoted from a parenthesis to a paragraph, in the test and in 04: the digest covers operator names and shapes only, and lift_fresh → scalar_tensor was found by comparing distinct_ops.
  • 11 — 04. ### The three disciplines is now ### The four disciplines, with a lead paragraph saying the fourth is a section of its own and that "a reader who stops at the end of this numbered list has three of four." D18's decision-log row records the fourth discipline with its own date. Line 210's claim is now true.

Finding 8 — Route 1's third ablation row

Two corrections, no new artifact.

  • "changes not one shape entry" is replaced, in the PR body and in T81, by what was actually measured: every family's non-numeric shape-entry count is unchanged (5,016 either way, family for family) while the total entries rise 12,544 → 12,585, because the 38 extra prim.device calls carry entries of their own.
  • The row's code is not in the tree, and both the PR body and T81 now say so rather than leaving a reader to discover it: it was a fake, non-numpy-backed CPU side written into the capture's staged allocators, measured, and removed once it had answered. It is the row that decided the route and the weakest-sourced thing in this PR. I did not rebuild it — the review already did, independently and from scratch, and got bookkeeping 83 → 121, exactly +38 prim.device, every family's non-numeric count unchanged. Restating the reconstruction as my evidence would be the same defect one layer down.

The pin, re-ablated against 14 tests

ablation round 1 round 2
restore int(t.shape[0]) in _rows 5 of 13 fail 5 of 14 fail
do not resolve on the host (keep guard_int) 4 of 13 fail 4 of 14 fail
unmodified 13 passed 14 passed in 107.6 s

Ablation A fails test_closing_site_one_moves_the_bound_to_a_third_site, test_nothing_specialises…, test_the_symbol_reaches_the_work_that_decides_the_cost, test_the_three_sites… and test_the_symbol_is_free_across_the_step_width. Ablation B fails the same set less the first. Neither ablation is in the tree; both are a sed on a staged copy, named in full in the gate log.


Gates — both tiers, as deltas against the new base

Re-measured; round 1's +4 / +4 against 9fcd6c7bd do not carry forward. Trees staged by git archive + docker cp into /tmp/cap2r2gate/{branch,control}, never into the shared mount; tarball md5 and per-file md5 checked on both ends; .compass-commit read back inside each container; each tree gated with its own scripts/compass/; COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new; sequential, never piped.

CPU tier, in xiaobizh_n18_cpu

control 24d75742d branch 70f8cd4db
passed 4,566 4,571 (+5)
skipped 149 149
xfailed 3 3
failed 0 0
GATE_CPU_RC 0 0
pytest wall 83.65 s 147.48 s
shell wall 1 m 30.1 s 2 m 33 s

The +5 is exactly this file's five new tests — round 1's four plus round 2's width refusal — and the skip count is identical on both sides, so the tier's three-way flake (tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk) did not fire in either direction: no ±1 in the passed/skipped split and no non-zero GATE_CPU_RC. The gate's own stamp reads gpu: not required (.compass-changed stamp) on the branch — atom/utils/forward_context.py is not in gpu_gate_triggers.txt, because the CPU tier covers it via tests/test_forward_mode.py. The GPU tier was run anyway, because this is atom/utils/.

GPU tier, in xiaobizh_n18, judged as a delta

control 24d75742d branch 70f8cd4db
passed 5,340 5,345 (+5)
failed 5 5
errors 0 0
skipped 105 105
xfailed 3 3
GATE_GPU_RC 0 0
PREFLIGHT_RC before / after 0 / 0 0 / 0
pytest wall 124.01 s 190.39 s

The five failures are the same five, by name, on both sides and against gpu_gate_known_failures.txt — both diffs of the sorted failing node-ids are empty, control-vs-branch and branch-vs-file:

tests/test_dcp_merge_ops.py::test_row_view_matches_output_slicing_bitwise
tests/test_fused_compress_ragged.py::...[extend0-context0-cut+whole]
tests/test_fused_compress_ragged.py::...[extend1-context1-whole+cut]
tests/test_fused_compress_ragged.py::...[extend2-context2-resume+fresh]
tests/test_fused_compress_ragged.py::...[extend4-context4-tiny-then-long]

Four are one bf16 ULP (max|diff| = 0.001953125 = 2^-9) and the fifth is the bitwise torch.equal. Toolchain matched the baseline's exactly: torch 2.10.0+rocm7.2.4.git3d3aa833, hip 7.2.53211, aiter v0.1.21.dev0-49-gf4e7c7509. Run on device 3 — devices 0 and 1 were 83% resident with someone else's job, and the gate's own bookable check offered 2 3 4 5 6 7.

Both sides' pass counts are accounted for by the gate's own arithmetic rather than by a floor: 5,340 = 4,779 baseline + (610 − 49) tests/compass and 5,345 = 4,779 + (615 − 49).


Effort

Measured against 24d75742d. Both AST conventions, because round 1's counter and the review's differ on docstrings; production is 0 net either way.

AST statements (docstrings counted) AST statements (docstrings excluded) physical non-blank
production atom/utils/forward_context.py 0 net (352 → 352) 0 net (335 → 335) +6 (851 → 857)
test tests/compass/test_capture_real_model.py +203 (577 → 780) +189 (540 → 729) +833 (1,307 → 2,140)

git diff --numstat 24d75742d HEAD: +7/−1 production, +970/−45 test, +102/−2 04, +1/−1 12.

Three of these were stale when this record was first posted — they were measured at 2b2a48370, two amends before head, and read +965/−45 test, +100/−2 04 and +828 (1,307 → 2,135) physical. Corrected above against head 70f8cd4db; the AST figures did not move and stand as first published. Flagged by the reviewer, and it is the right catch: the body is what a squash message is written from, so a stale count would have outlived the PR.

Comments were not trimmed — six of the eleven findings asked for more to be stated — and the split is now published rather than left to a ratio. Of the +833 physical non-blank lines: comments +164, docstrings +285, everything else +384, i.e. 54% explanation. Of that 384, 189 are AST statements that are not docstrings (203 counting the 14 new docstrings as statements); the rest are continuation lines, mostly the one-per-entry rows of OP_FAMILIES and EXPECTED_HOST_RESOLUTIONS.


What I could not do

  • The GPU tier ran on device 3. Devices 0 and 1 were 83% resident with someone else's job; the gate's own bookable check offered 2 3 4 5 6 7. Same device as round 1.
  • Three stale python -m pytest processes have been sitting in xiaobizh_n18_cpu for two days at 0% CPU. They are not live gates and I did not wait on them; they are also not mine and I did not kill them. Noting it because a reviewer's pgrep will see them.
  • Route 1's third ablation still has no in-tree artifact. See finding 8 — the wording is corrected and the absence is disclosed, which is what was asked for, but it remains a measurement that has to be rebuilt to re-check.

Not merged, not landed, not undrafted, no label applied.


Addendum — scratch collision, for whoever is next in this tree

The reviewer's scratch and mine collided at agent_scratch/cap2r2; it moved its own output to agent_scratch/cap2r2rev/ and left my files intact. Noting it so the next reader of either directory knows which is which, and because it is the second scratch collision in this task — the first, /tmp/rh.txt, is in What I could not do above.

for tp in (1, 2):
for width in (DECODE_SEQS, SECOND_WIDTH):
at = capture(tp, step_symbol=True, width=width)
assert host_resolutions_by_line(at) == EXPECTED_HOST_RESOLUTIONS, (

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 1 — closed, and I verified it can fail rather than reading it.

This was the "claimed property that does not exist" finding, so re-reading the code proves nothing. I added an undeclared seventeenth conversion to a staged copy of this branch — one line, inserted after forward_context.py:444 so SITE_THREE's line number is untouched and the failure cannot be a line-shift artifact:

        _undeclared_conversion = int(input_ids.shape[0])

pytest tests/compass/test_capture_real_model.py on that tree: 1 failed, 13 passed, and the one failure is this assertion, naming the intruder:

E   AssertionError: (1, 2)
E   Omitting 16 identical items
E   Left contains 1 more item:
E   {'atom/utils/forward_context.py:448 in assert_shape_contract': 1}

The instrument does what the three places now say it does. I also re-derived the declared set independently rather than transcribing my round-1 table: 16 lines, 20 conversions, and host_resolutions_by_line(record) == EXPECTED_HOST_RESOLUTIONS is True at TP1 and TP2 and at step widths 2 and 8 — so the loop above is asserting a property that holds on all four rather than one that happens to hold on the default.

Keying on frames[-1] is the right half to declare: it is the line that goes stale when a file is edited, and the path above it is not a property of the conversion. The multiset rather than the set is the part I would have asked for if you had not done it.

return dict(counted)


def op_family(name):

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 7 — closed. The table did not move, and I checked the matcher rather than the digest.

I rebuilt the published census from all eight records this file produces, cell by cell: gemm 257/514, attention 64/304, normalisation 273/802 at both TPs, collective 133/266 at TP2, allocation 566/803 against a 565/0 control, view 921/1736 against 922/0, bookkeeping 83/21, totals 2,521 / 5,016 of 12,544 and 2,662 / 5,297 of 13,107. Every row stands under the exact matcher.

The 41 reproduces: the union of distinct_ops across the eight records is 41, and op_family classifies all 41 — unclassified is empty. By family: view 10, elementwise 4, collective 4, allocation 4, normalisation 4, activation 3, bookkeeping 3, attention 2, embedding 2, sampling 2, transfer 2, gemm 1.

One correction to how this is being described, and it is in your favour. addmm and addbmm do not land in gemm now. Measured against this function:

aten.addmm.default          -> unclassified
aten.addbmm.default         -> unclassified
aten.slice_scatter.default  -> unclassified
aten.select_scatter.default -> unclassified
aten.add.Tensor             -> elementwise
aten.slice.Tensor           -> view

They land in the bucket that fails the test, which is stronger than being counted correctly and is exactly what the comment above asks for — a classification somebody makes. I checked because the PR body's sentence "aten.add captured aten.addmm and aten.addbmm, which are GEMMs" can be read as saying they are now counted as GEMMs. They are not counted at all until somebody declares them. One clause in the body would settle it; it is not a finding.

Worth flagging for whoever meets it first: the next ATOM change that routes a projection through addmm fails the census test with no hint that the right answer is gemm. That is the design — but the failure is a classification request, not a regression.

# a 2 that is still a 2 when the step is eight wide is not the
# step. The group width is an engine constant here in the way a
# head count is.
assert values & {DECODE_SEQS} == {DECODE_SEQS}, tp

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new artifact — I am satisfied the distinction is real, and not a rescue.

The claim that had the counterexample was round 1's "the 23 concrete values contain no width tried". Qualifying it rather than dropping it is only legitimate if the discriminator is independent of the claim, and here it is: the eight-width digest identity was measured for a different reason and settles this one.

value 2 in concrete_dims value 8 concrete_dims identical at step widths 2 and 8
TP1 0 0 yes
TP2 73 0 yes

The 73 sit in view 36, allocation 33, transfer 3 and bookkeeping 1 — the allocations, views, transfers and one device read the record claims. The whole concrete inventory is identical when the step is eight rows wide, and the value 8 appears nowhere at either TP. Had this been a rescue, the wide capture would have carried an 8 somewhere, or concrete_dims would have differed between the two widths. Neither happens, so a 2 that survives an eight-wide step is the group's width in the way a head count is.

Two details that made me more rather than less confident: TP1 has 23 distinct concrete values and TP2 has 22, and the two sets differ exactly as a TP split predicts (17,408/34,816 → 8,704, 248,320 → 124,160). And the assertion above is now pointed at the direction that can fail — SECOND_WIDTH not in values at both TPs — with the group width asserted present at TP2 and the reason written beside it rather than inferred.

That the old assertion had only ever run at TP1 is the finding-4 repair paying for itself immediately, and disclosing it rather than quietly widening the test is the right handling.

# record is the shape of the thing this file was built to catch, so it
# is the one outcome the capture will not print.
parser.error(
f"--width {args.width} traces no symbol: torch specialises a size "

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 5 — closed. It refuses, in both places, and the refusal is fail-able.

Doing both halves was the right call: main() is not the only entry point, and _step_axis's backstop testing re.fullmatch(r"s\d+", str(axis)) rather than the type means the check and the assertions downstream of it agree by construction. The 14th test runs the real entry point and asserts a non-zero exit, the message and — the part that matters — that RECORD_MARKER is absent from stdout.

One nit, for whenever this file is next touched. This refusal fires on every pass, not only --step-symbol: --tp 1 --width 1 with no other flag is refused too, with a message about tracing no symbol. Refusing is right there as well — DECODE_SEQS' own comment says a hint of 1 gives a fully constant graph in the concrete pass too — but the message names only the symbolic reason, so a reader who hits it on the control pass is told something that is true of a different pass.

| `backends.py in _mrope_cpu_view` | 398, 400 | 2 |
| `gdn_attn.py in _attach_gdn_decode_metadata` | 1237 | 1 |

A seventeenth sits in `forward_context`'s own `assert_shape_contract`, whose `_rows`

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 11 — closed. The heading is ### The four disciplines, the lead paragraph says in as many words that a reader who stops at the numbered list has three of four, and D18's decision-log row carries the fourth discipline with its own date. 04:210's "asserts that log against a declared set" is now true — I broke the assertion to confirm it, see my comment on the test.

One nit for whenever 04 is next touched, not a finding. This sentence is in the present tense — "A seventeenth sits in forward_context's own assert_shape_contract, whose _rows helper takes int(t.shape[0])" — two paragraphs above the one that says this change removes that int(). In context it reads correctly as an account of the problem followed by the fix, but 04 will outlive this PR and the sentence will then describe a line that has not existed for months. A tense fixes it.

@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Review — CAP-2, round 2

Verdict: APPROVE. All eleven round-1 findings are closed, and I verified the four you were asked to weight rather than taking the record's word for any of them. No round-1 finding survives, so there is no two-cycle halt. The central result reproduces again on my own staging, and the instrument that carries it can now fail — I made it fail. One correction to the PR body's Effort table is required before landing; it is a stale measurement, not a defect in the change, and it does not need a round 3.

Everything below was re-measured from scratch against head 70f8cd4db and base 24d75742d, with both trees built by each tree's own scripts/compass/snapshot.sh and docker cp'd into container paths of my own (tarball md5 checked on both ends, .compass-commit read back inside), gated with their own scripts/compass/, COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, sequential and never piped.


First: the paste incident. Nothing foreign survived.

You disclosed that a heredoc to a shared /tmp/rh.txt failed because another user owns that path, and that the following cat appended their file into the test file. I treated this as the first question and not as a formality, because a git checkout -- can miss a file that was staged in between.

It is clean, on five independent checks.

  1. The line counts reconcile exactly. Base 24d75742d is 1,561 lines; +970 / −45 gives 2,486; head is 2,486. 04 is 1,045 + 102 − 2 = 1,145, which is head. 12 is 282 + 1 − 1 = 282.
  2. I read the foreign file. /tmp/rh.txt on that box is owned by another user, 2,650 bytes, dated 18 Sep: it is somebody's pytest failure listing, ending 211 failed, 5743 passed, 526 skipped, 3 xfailed, 1 warning, 112 errors. I extracted every token of it longer than seven characters and grepped each one against the test file. Zero hits.
  3. The file is structurally whole. It parses; it holds 14 test functions and no orphan or duplicated block; grep -P '[^\x00-\x7F]' returns two box-drawing separator comments, one of which is at the base and the other of which matches its style.
  4. No trace of the accident's shape. No /tmp path, no home path, no other username, no email address and no heredoc marker anywhere in the file.
  5. The "staged in between" worry is answerable, and I answered it. The branch's reflog has four states for this commit — d6e2b4a26, 2b2a48370, caf0bc87e, 70f8cd4db. I read the test file at each. The only inter-amend change to it is 2b2a48370 → caf0bc87e, +23 / −11, and it is two prose edits, both of which narrow a claim: the expected_guards docstring stops citing widths 5, 7 and 64 (which were my measurements, not yours) and the _resolve_on_the_host docstring is rewritten to describe the assertion that now exists. Nothing foreign was ever committed.

Also for the record, so the next agent does not have to rediscover it: the developer's scratch and mine collided on the directory name agent_scratch/cap2r2. Nothing durable lives there, but I have moved my own output to agent_scratch/cap2r2rev/.


The four you asked me to weight

Finding 1 — the log is asserted, and the assertion can fail. I broke it.

This was the "claimed property that does not exist" finding, so a re-read of the code is not evidence. I added an undeclared seventeenth conversion to a staged copy of the branch — one line in assert_shape_contract, inserted after line 444 so that SITE_THREE's line number is unchanged and the failure cannot be a line-shift artifact:

        _undeclared_conversion = int(input_ids.shape[0])

Result on the 14 tests: 1 failed, 13 passed, and the failure is exactly the right one, with the right message:

FAILED tests/compass/test_capture_real_model.py::test_the_three_sites_are_where_they_were_and_carry_the_symbol_instead
E   AssertionError: (1, 2)
E   Omitting 16 identical items
E   Left contains 1 more item:
E   {'atom/utils/forward_context.py:448 in assert_shape_contract': 1}

The multiset is the right shape for the claim, and host_resolutions_by_line keying on the innermost ATOM frame is the right half to declare. I re-derived the table independently: 16 lines, 20 conversions, identical at TP1 and TP2 and at step widths 2 and 8 — matches_declared=True on all four. Closed.

Finding 7 — the family table did not move. Every cell.

I re-ran all eight records the file produces and rebuilt the published table from them rather than comparing digests:

family TP1 control TP1 step-symbol TP2 control TP2 step-symbol
gemm 257 / 0 257 / 514 257 / 0 257 / 514
attention 64 / 0 64 / 304 64 / 0 64 / 304
normalisation 273 / 0 273 / 802 273 / 0 273 / 802
collective — — 133 / 0 133 / 266
activation 128 / 0 128 / 256 128 / 0 128 / 256
elementwise 212 / 0 212 / 536 212 / 0 212 / 536
allocation 565 / 0 566 / 803 566 / 0 567 / 804
view 922 / 0 921 / 1736 925 / 0 924 / 1742
transfer 14 / 0 14 / 38 16 / 0 16 / 44
embedding 1 / 0 1 / 2 1 / 0 1 / 2
sampling 2 / 0 2 / 4 2 / 0 2 / 4
bookkeeping 83 / 0 83 / 21 85 / 0 85 / 23
total 2,521 / 0 of 12,544 2,521 / 5,016 of 12,544 2,662 / 0 of 13,107 2,662 / 5,297 of 13,107

Every published row stands, cell for cell, under the exact matcher. Digests: a7c1ec1b6ac1ec4b… at TP1 and a37d76da81cd4547… at TP2, each identical at widths 2 and 8; controls 235e44b884e41cd9… and 2dac49dcc79ad264….

The 41 reproduces: the union of distinct_ops across the eight records is 41, and op_family classifies all 41 — unclassified is empty. By family: view 10, elementwise 4, collective 4, allocation 4, normalisation 4, activation 3, bookkeeping 3, attention 2, embedding 2, sampling 2, transfer 2, gemm 1. The arithmetic is checkable from the constants too: 33 at TP1 plus the six TP2-only, plus aten.zeros from the probe passes and aten.scalar_tensor from the step-symbol passes.

One correction to how this is being described, and it is in your favour. addmm and addbmm do not now land in gemm. Measured:

aten.addmm.default          -> unclassified
aten.addbmm.default         -> unclassified
aten.slice_scatter.default  -> unclassified
aten.select_scatter.default -> unclassified
aten.add.Tensor             -> elementwise
aten.slice.Tensor           -> view

They land in the bucket that fails the test, which is the correct behaviour and the one the file's own rule asks for — "a new operator is a classification somebody makes, not a bucket it falls into quietly". I checked this because a reader of the PR body's sentence "aten.add captured aten.addmm and aten.addbmm, which are GEMMs" could reasonably conclude they are now counted as GEMMs. They are not counted at all until somebody declares them, which is stronger. Worth one clause in the PR body; not a finding.

Finding 4 — the guards, and the TP2 group width. The distinction is real.

The guard set reproduces exactly as re-measured:

width 2 width 8
TP1 28*<axis> <= 8192, <axis> < 128, <axis> <= 512 …plus <axis> + 1 > 8
TP2 <axis> < 128, <axis> <= 512 …plus <axis> + 1 > 8

shape_env_replacements is {} in all four. expected_guards(tp, width) composes exactly this, and the two-width test now asserts the whole set at both widths and the lower bound's presence and absence — so the artifact is held in both directions instead of being left out of the comparison.

On the new artifact: the 2 at TP2 is the group's, and this is not a rescue of a claim with a counterexample. The eight-width digest identity is the evidence, and it settles it:

value 2 in concrete_dims value 8 concrete_dims identical at step widths 2 and 8
TP1 0 0 yes
TP2 73 0 yes

The 73 sit in view 36, allocation 33, transfer 3 and bookkeeping 1 — the allocations, views, transfers and one device read the record claims. The discriminator is the right one and it is decisive: the whole concrete inventory is byte-for-byte the same when the step is eight rows wide as when it is two, and the value 8 appears nowhere at either TP. A 2 that survives an eight-wide step is not the step's width; it is the group's, in the way a head count is. Had this been a rescue, the wide capture would have carried an 8 somewhere, or concrete_dims would have differed. Neither happens. TP1 has 23 distinct concrete values and TP2 has 22, and the two sets differ exactly as a TP split predicts (17,408/34,816 → 8,704; 248,320 → 124,160, and so on).

Round 1's "the 23 concrete values contain no width tried" is correctly qualified rather than quietly dropped, and the assertion is now pointed at the direction that can fail — SECOND_WIDTH not in values at both TPs — with the group width asserted present at TP2 and the reason written out beside it. That is the right repair.

Finding 5 — it refuses, and the refusal is fail-able.

main() rejects a width below MIN_STEP_WIDTH before any capture runs, _step_axis refuses again if what comes back does not match s\d+, and the 14th test runs the real entry point in a subprocess and asserts a non-zero exit, the message text, and that RECORD_MARKER is absent from stdout. Both halves, because the CLI is not the only entry point — that was the right call. Closed.


The other seven, spot-checked

  • 2 — the "fourteen lines later" error. Gone. grep -rn fourteen over the tree returns the two deliberate negations (T81 and the site-three test's docstring), 12's unrelated row about fourteen torch.cuda stubs, and one line in tests/entrypoints/. The base's correction was adopted: the branch's only change to that sentence is dropping the now-stale :437.
  • 3 — the restack. git merge-base 24d75742d 70f8cd4db is 24d75742d; two commits sit above it; the file set is the four expected. The five collisions are resolved as described, and the account of the identifier removals is corrected to say this branch removes none.
  • 6 — TP2 two-width evidence. for tp in (1, 2) in the free-symbol test, confirmed by my own eight records above.
  • 8 — judged below.
  • 9 — ScheduledBatch.__init__. Now in 04's fourth-discipline section as a paragraph of its own, in T81 under What the symbol does not survive and Still open, in the module docstring, and in the PR body's What is not done. The sentence that claimed the symbol enters the constructor is rewritten. Closed.
  • 10 — replacements == {}. Re-worded to claim only what it can fail on, with the two-width digest named as the load-bearing evidence in both docstrings, each pointing at the other. The scope note is promoted to a paragraph in the test and in 04. Closed.
  • 11 — 04. ### The four disciplines, with a lead paragraph that says in as many words that "a reader who stops at the end of this numbered list has three of four", and D18's decision-log row updated with its own date. 04:210's claim is now true — see finding 1. Closed.

Finding 8 — declining to restate my reconstruction was the right call

You asked me to judge this, so: yes, and I would have objected to the alternative.

Restating my round-1 reconstruction as the PR's own evidence would have been provenance laundering — the same defect one layer down, as you put it, and exactly what principle 8 exists to prevent. My reconstruction is on this PR, signed, with its own method; a reader who wants it can read it there, attributed to the person who ran it. Copying it into the dev record would have made a second-hand number look first-hand.

Two things make the residual risk smaller than the bare sentence "no in-tree artifact" suggests, and they are worth stating so a successor does not over-weight it:

  1. The route decision does not rest on the absent row alone. The positive result is in the tree and asserted: the step-symbol pass runs against ATOM's unmodified CpuGpuBuffer, with symbolic_device_allocations == 0 and constructed == 19 asserted at both TPs. "A symbolic CpuGpuBuffer is not required" is carried by that. The absent ablation only quantifies what the alternative would have cost; it is not what establishes that the alternative is unnecessary.
  2. The wording correction is the one that mattered. "Changes not one shape entry" was the loose claim; "every family's non-numeric count unchanged at 5,016 while the total rises 12,544 → 12,585" is what was measured, and it now says that in both the PR body and T81, along with the fact that the code is not in the tree.

Building a fourth permanent pass into the file to hold a row that decides nothing load-bearing would have cost more than it bought. Correct handling.


Gates — my own runs, both tiers, as deltas

CPU tier, xiaobizh_n18_cpu

control 24d75742d branch 70f8cd4db
passed 4,566 4,571 (+5)
skipped 149 149
xfailed 3 3
failed 0 0
GATE_CPU_RC 0 0
pytest wall 88.37 s 144.55 s
shell wall 1 m 35.7 s 2 m 30.6 s

Every figure confirmed. The +5 is this file's 9 → 14 tests, counted at both revisions. The skip count is identical on both sides and GATE_CPU_RC is 0 on both, so the tier's three-way flake — tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk, which can pass, skip, or return a non-zero GATE_CPU_RC indistinguishable from a regression — did not fire in either direction on my runs either.

I also confirmed the blind-spot claim mechanically: atom/utils/forward_context.py is absent from gpu_gate_triggers.txt, and tests/test_forward_mode.py is absent from cpu_gate_exclude.txt, so it is in the CPU tier. "The CPU tier covers it, and the GPU tier was run anyway" is correct.

GPU tier, xiaobizh_n18, judged as a delta

control 24d75742d branch 70f8cd4db
passed 5,340 5,345 (+5)
failed 5 5
errors 0 0
skipped 105 105
xfailed 3 3
GATE_GPU_RC 0 0
PREFLIGHT_RC before / after 0 / 0 0 / 0
pytest wall 113.61 s 182.39 s

Every figure confirmed, and the by-name check with it. The gate prints its own expectation from its own arithmetic, and both sides met it exactly:

control:  expected: 5340 passed = 4779 baseline + (610 - 49) tests/compass
branch:   expected: 5345 passed = 4779 baseline + (615 - 49) tests/compass

The five FAILED node-ids are byte-identical on both sides and are the five in gpu_gate_known_failures.txt, verbatim:

tests/test_dcp_merge_ops.py::test_row_view_matches_output_slicing_bitwise
tests/test_fused_compress_ragged.py::...[extend0-context0-cut+whole]
tests/test_fused_compress_ragged.py::...[extend1-context1-whole+cut]
tests/test_fused_compress_ragged.py::...[extend2-context2-resume+fresh]
tests/test_fused_compress_ragged.py::...[extend4-context4-tiny-then-long]

Toolchain matched the baseline's exactly on both sides — torch 2.10.0+rocm7.2.4.git3d3aa833, hip 7.2.53211, aiter v0.1.21.dev0-49-gf4e7c7509 — so the delta compares like with like rather than riding the gate's documented AITER-bump gap. Run on device 3; the gate's own bookable check offered 2 3 4 5 6 7, with devices 0 and 1 at 83% VRAM with somebody else's job, same as your run. I ran this tier twice per side, and both passes ended GATE_GPU_RC=0.

Also confirmed

  • 14 tests pass on my staged branch tree: 14 passed in 110.22s, with import atom verified to resolve under that tree before I read any count.
  • The three python -m pytest processes in xiaobizh_n18_cpu are stale, as you said: ELAPSED 2 days, TIME about one minute of CPU each, 0.0%. Not live gates. (pgrep -f on a command string also matches one's own bash -lc wrapper — I read the pytest list, not the predicate.)
  • Lint. ruff format --check clean on both changed files; ruff check reports one BLE001 at forward_context.py:938, which is the same single error at the base (there at :933, moved by this branch's six comment lines). Pre-existing, as claimed.

Identifier sweep — the control fires

I did not trust a clean sweep without watching it fail first. On 9fcd6c7bd:tests/compass/test_capture_real_model.py the same pattern returns exactly four hits, at the four lines you named:

581:    cannot fail and cannot go stale -- the defect principle 8 exists for.
1387:   anywhere in this file -- so the repair route T81 names for site two, a
1529:   T81 recorded the two sites as independent and said a symbolic bound closes
1546:   # Site two, unchanged by the repair -- which is the half of T81's sentence

The same pattern returns nothing on 24d75742d, nothing on either changed file at head, and nothing across tests/compass/, atom/compass/**/*.py or scripts/compass/. Clean, and the cleanliness is measured rather than assumed.


Effort — reported, not adjudicated, with one correction

Verified at head:

AST statements (docstrings counted) AST statements (docstrings excluded) physical non-blank
production atom/utils/forward_context.py 0 net (352 → 352) ✓ 0 net (335 → 335) ✓ +6 (851 → 857) ✓
test tests/compass/test_capture_real_model.py +203 (577 → 780) ✓ +189 (540 → 729) ✓ +833 (1,307 → 2,140)

The correction, and it is the one finding I am leaving on this PR. Three of the Effort figures do not describe the head. They describe 2b2a48370, two amends earlier:

figure PR body / dev record at head 70f8cd4db
test numstat +965 / −45 +970 / −45
04 numstat +100 / −2 (record) / +101 / −3 design (body) +102 / −2, design total +103 / −3
test physical non-blank 1,307 → 2,135 (+828) 1,307 → 2,140 (+833)

The numbers are stale, not wrong-in-kind: git diff --numstat 24d75742d...2b2a48370 gives exactly 965 / 45 and 100 / 2, and that commit's test file has exactly 2,135 non-blank lines. The last two amends were prose-only — the two docstring narrowings I described in the paste-incident section, plus one line of 04 — which is also why the AST figures survived unchanged and are correct. Fix the three figures in the PR body before landing. This does not need another review cycle; it is a number that lost track of its own measurement, which is the one thing principle 8 asks a record not to do, and it is worth saying plainly on the PR that most insists on measurement.

The size, reported rather than adjudicated. +833 physical non-blank in one round is the largest prose growth of the session, and the split is:

base head delta
comment-only lines 100 264 +164
docstring lines 450 735 +285
everything else (code and multi-line data literals) 757 1,141 +384

So +449 of the +833, or 54%, is explanation — comments and docstrings. Of the remaining +384, only +203 are AST statements; the balance is continuation lines of two declared tables (OP_FAMILIES is 65 physical lines and one statement; EXPECTED_HOST_RESOLUTIONS is 20 and one) and multi-line asserts. I am not suggesting any of it be trimmed: six of the eleven findings asked for more to be stated, and the two largest new blocks are the declared sets that findings 1 and 7 required. The instrument is an open owner decision and I am recording the figure for it, not arguing it.


For the next task in this area

  • The unclassified bucket is now load-bearing and it will fire on a true positive first. addmm, addbmm, slice_scatter and select_scatter are all undeclared today. The first ATOM change that routes a projection through addmm fails the census test with no hint that the right answer is gemm — which is the design, but the next agent should know the failure is a classification request and not a regression.
  • 04's new section says a seventeenth conversion "sits in forward_context's own assert_shape_contract, whose _rows helper takes int(t.shape[0])" in the present tense, two paragraphs above the one that says this PR changed that line. It reads correctly as an account of the problem then the fix, but it will outlive the PR; a tense would fix it whenever 04 is next touched.
  • main() refuses --width 1 on every pass, not only --step-symbol, with a message about tracing no symbol. The refusal is right for all four passes — a hint of 1 is degenerate in the concrete pass too — but the message names only the symbolic reason.
  • The one-symbol argument is where prefill will be decided. Everything here turns on a decode step's token count and sequence count being one number for a structural reason. Site three is the demonstration of what happens when they are two symbols: ATOM's contract equates them, correctly. Expect prefill to meet that on purpose rather than by accident, and read the guards per width before trusting any prefill digest.

Reviewed against head 70f8cd4db, base 24d75742d. Not merged, not landed, not undrafted, no label applied.

@jgong5
jgong5 added this pull request to stack #151 September 22, 2026 11:02
@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Pin re-verification — CAP-2 (#150). Not a new review cycle.

The existing APPROVE stands. Nothing below changes it. This is the pin
re-verification the board now requires of every approved PR: for each test
presented as pinning a defect, reinstate the defect as it actually stood in
git
and record whether the test fails and on what. Reviewer instructions did
not mandate reinstatement when this PR was approved, so its approval — like
every other on the board — was given without it. Reading a pin is not seeing it
fail.

I did not re-review the design, and I re-opened no settled finding.

State, and how the base was determined

head 70f8cd4db33dcebe27e116475d3f73d7301a6ee7, unmoved since the round-2 APPROVE (head committed 05:23:59Z, approval 06:40:29Z)
base ref / sha compass/cap-1 @ 24d75742dfbc010d856796e84882a29bfb50f069
draft true (untouched)
labels none
state open, mergeable_state: clean

Base sha determined four ways, all agreeing: .base.sha;
gh api repos/jgong5/ATOM/git/ref/heads/compass/cap-1 → 24d75742d;
compare/compass/cap-1...70f8cd4db → ahead_by 2, behind_by 0,
merge_base_commit.sha = 24d75742d; and the local compass-worktrees/cap-1
worktree sitting at the same commit. #150 is not stale.

Instrument

git archive of 70f8cd4db + docker cp into xiaobizh_n18_cpu at
/tmp/xiaobizh_pr150_pinaudit/ (my own path; nobody else's staging touched, the
shared mount never written). Tarball md5 checked both ends
(a774402fd2cb9d06ad8dc567c906d6b0). Both mutated files verified byte-identical
to their git blobs at head before and after every cycle —
forward_context.py 22e6b4548f073e02…, test_capture_real_model.py
862b64358db0b68b…. atom.__file__ asserted under the staged root and printed
before any count. __pycache__ cleared between runs. Every run under
timeout -k 10 2400, never piped — captured to a file, file read.
For the record, scripts/compass is the same tree object
(ba42b43d81cc053b920f5b554611e062788001ce) at base and at head, so a
control/branch gate here is one instrument.

Every mutation was line-count-preserving, applied one at a time, with a
restore and a re-verify between each.
No two mutations were ever live
together. Published baseline and my baseline agree: 14 passed, and the
tree re-verified at 14 passed after the last mutation was reverted.

The table

P = published/baseline count on the unmutated head. All runs are
pytest tests/compass/test_capture_real_model.py, 14 tests.

# pin defect reinstated P reinstated verdict
M0 (line-drift control) one word re-cased inside the new comment in _rows, 987 → 987 lines 14 passed 14 passed control clean — the 5 failures below are semantic, not drift
M1 test_closing_site_one_moves_the_bound_to_a_third_site int(t.shape[0]) restored in _rows (comment kept, 987 → 987) 14 passed 5 failed / 9 passed bites — :2128, site(three,1) = forward_context.py:430 in _rows ≠ :444 in assert_shape_contract
M1 test_nothing_specialises_under_the_step_symbol_capture same bites — :2180, re.fullmatch(r"s\d+", '2') is None: the step axis specialised to its hint
M1 test_the_symbol_reaches_the_work_that_decides_the_cost same bites — :2263, AssertionError: ('gemm', 1) / assert 0 > 0
M1 test_the_three_sites_are_where_they_were_and_carry_the_symbol_instead same bites — :2361, multiset gains {'atom/utils/forward_context.py:430 in _rows': 2} over the declared 16 lines
M1 test_the_symbol_is_free_across_the_step_width same bites — :2450, 7d22a9fded43… != 34d00deb9d40…, the two-width digest
M1b (realistic full revert) git show 24d75742d:atom/utils/forward_context.py copied in whole, 987 → 981 14 passed 5 failed / 9 passed same five, same assertions, line numbers shifted :430 → :424 — the semantic revert is what bites, the drift only renames the line in the message
M2 test_the_capture_refuses_a_width_that_torch_would_specialise main()'s if args.width < MIN_STEP_WIDTH: → if False: 14 passed 1 failed / 13 passed bites, but only on "traces no symbol" in stderr — see below
M2b (the second half of the same refusal) _step_axis's if not re.fullmatch(r"s\d+", str(axis)): → if False: 14 passed 14 passed blind spot — the backstop is held by nothing
M3 test_the_symbol_reaches_… / the unclassified bucket "aten.mul" → "aten.mulX" in OP_FAMILIES 14 passed 1 failed / 13 passed bites — assert 'unclassified' not in …
M4 test_the_symbol_reaches_… / the gemm row aten.as_strided declared into the gemm family 14 passed 1 failed / 13 passed bites — exact-set assertion: Extra items in the left set: 'aten.as_strided.default'
M10 test_the_symbol_reaches_… / the normalisation row the same op declared into normalisation 14 passed 14 passed blind spot — see below
M8 test_the_width_is_the_group_s_and_nothing_simulated_it sentinel's calls pre-loaded with one entry 14 passed 1 failed / 13 passed bites — assert [{'frames': ['TRIPWIRE']}] == []; the list is genuinely plumbed into the record
M9 (always-zero on the headline total) shape_census's non_numeric += 1 → += 0 14 passed 14 passed by design — no test asserts the total; see note

Ten mutations. Nine pins bite. Two blind spots. No inert pin, and no pin that
names the wrong defect.

The one production line

_rows returning t.shape[0] instead of int(t.shape[0]) is the whole
production change, and it is genuinely pinned five ways, each on a different
artifact — the site's location, the axis staying a symbol, the gemm row, the
host-resolution multiset, and the two-width digest. Every one of the five names
the right defect, and none is a restatement of another.

The reinstated failures also settle a reachability question the negative
assertions in this file otherwise leave open: _rows is executed in the
step-symbol pass. With int() back, its line appears in the host-resolution log
with 2 conversions. So assert not [... "in _rows" ...] is a negative over a
line that really runs, not a needle discriminating against a sentence nothing
reaches. No tripwire was needed; the reinstatement is the stronger evidence.

Separately measured, and it supports the PR's "nothing runs differently"
claim:
tests/test_forward_mode.py — 22 tests, eight of which call
assert_shape_contract directly — is 22 passed with int() and 22 passed
without
. ATOM's own suite for that method cannot tell the two apart, which is
the claim. The corollary is worth stating: outside
tests/compass/test_capture_real_model.py, nothing in the tree distinguishes
them, so this file is the only thing holding that line.

Is the op-family decomposition real? Yes, with one soft edge.

Real (principle 7):

  • There is no total assertion at all. M9 zeroed shape_census's
    non-numeric counter — the instrument behind the published 5,016 of 12,544
    — and all 14 tests stayed green. The headline total genuinely is not what
    anything passes on; the per-family rows are. That is the property the PR
    claims, confirmed by mutation rather than by reading.
  • A total therefore cannot be passing for a peripheral-op reason, because
    no total is passed on. The three cost-bearing rows are asserted individually
    (> 0 per family, and M1 shows the gemm one firing at assert 0 > 0).
  • The unclassified bucket is live (M3), and it is what makes the breakdown a
    measurement rather than a description.
  • The gemm row cannot be inflated: M4 moved one peripheral view operator
    into it and the exact-set assertion caught it immediately. The same holds for
    attention, which is also an exact set.

The soft edge — and this is a blind spot, not a defect in any published
number:

  • The normalisation row is held by a membership test
    (assert "aiter._fused_qk_rmsnorm_group_quant_kernel.default" in norms),
    where gemm and attention are held by exact sets. M10 declared
    aten.as_strided into normalisation — the same mutation M4 catches in
    gemm — and all 14 tests passed, with 802 symbol-carrying view entries
    silently reported as normalisation work. The published numbers are unaffected
    (nothing on this step is misfiled today, and the round-2 review re-derived the
    table cell-for-cell), but the row that would understate or overstate a
    cost-bearing family is the one family whose guard does not close. Principle 7:
    the decomposition is what carries the claim, so the third row deserves the
    same exact set as the other two.

Is the refusal-demonstrating test itself inert? No — but only half of what it is credited with is fail-able.

test_the_capture_refuses_a_width_that_torch_would_specialise is the
"it-cannot-be-done-here" test here, and it is not inert: M2 removed
main()'s width refusal and the test went red. Principle 6 holds at the CLI.

What the measurement adds is how it went red. With main()'s refusal gone,
_step_axis's backstop still raised, so the subprocess still exited non-zero
and still printed no record — the returncode != 0 and RECORD_MARKER not in stdout halves both passed. The only assertion that fired was
"traces no symbol" in completed.stderr, which discriminates the two refusals
purely by their wording. And M2b, disabling the backstop while leaving main()
intact, was 14 passed: no test in this file, or anywhere in the tree,
reaches _step_axis's refusal. The round-2 review closed finding 5 on "both
halves, because the CLI is not the only entry point". The first half is
fail-able. The second is a refusal nobody has seen fail — the exact shape of an
inert pin, sitting one level below a pin that is not inert.

Neither of these is load-bearing for the result, which is why neither changes
the approval. Both are cheap to close: one exact set, and one test that calls
_step_axis(fake_mode, 1) directly and reads the message.

What the line-drift control did

M0 — one word re-cased inside the new comment in _rows, 987 lines before and
after — came back 14 passed. This file pins two ATOM line numbers
(SITE_THREE at forward_context.py:444, and sixteen lines in
EXPECTED_HOST_RESOLUTIONS), so a drift guard was the risk worth controlling
for. It did not fire, and M1b then showed what drift alone looks like: the same
five failures with :430 reading :424. The five failures under M1 are the
defect, not the diff.

What I could not reach

  • The GPU tier. Not run; the PR's GATE_GPU_RC=0 with five known-failure
    node-ids on both sides is unverified by me.
  • The full CPU tier delta. I ran the changed test file (14) and
    tests/test_forward_mode.py (22), not the ~4,571-test suite, so the
    published +5 / GATE_CPU_RC=0 is unverified by me. Nothing I ran touched
    the three-way flake in tests/entrypoints/.
  • The route-1 row (a fake, non-numpy CPU side costing 38 prim.device
    calls). Its code is not in the tree — the PR says so itself and calls it the
    weakest-sourced row in the PR — so it cannot be reinstated and I did not.
  • The sentinel firing on a real call. M8 proves the calls list reaches
    the record, so apply_simulated_tp_calls == [] is not an always-zero
    artefact of the plumbing. It does not prove the sentinel is bound over both
    names, because nothing in the file makes apply_simulated_tp actually run.

Verdict

The APPROVE stands. Five independent pins on the one production line all
bite and all name it correctly; the census is a real decomposition with no total
to hide behind; the refusal is fail-able at the entry point that exists. The two
findings above are additive work, not regressions, and neither touches a
published number.

Staging under /tmp/xiaobizh_pr150_pinaudit/ on xiaobizh_n18_cpu has been
removed and the ControlMaster closed.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Correction to the review record: finding 5 was closed on a basis that was only half fail-able

Round 2 (comment 5772255808) closed finding 5 because the refusal was held on "both halves, because the CLI is not the only entry point". Only one of the two halves was fail-able. The pin re-verification (5784085855) measured this and #236 filed it. It is re-measured here on node 18 at be86326cc, and it still holds:

  • _step_axis's backstop was held by nothing. Replacing its check with if False: left tests/compass/test_capture_real_model.py at 18 passed. main() refuses --width 1 before the backstop is reached, so the only test that ran the refusal never got to it.
  • main()'s refusal was held by the message alone. With it disabled, the backstop still refused, so returncode != 0 and RECORD_MARKER not in stdout both passed. Only "traces no symbol" in stderr failed (:2595).

Fixed in #359, on head 12a2ad6d4:

  • test_the_step_axis_refuses_a_hint_torch_specialises calls _step_axis directly. With the backstop disabled it fails on DID NOT RAISE (1 failed / 19 passed), which is the raise and not the string.
  • The refusal test asserts returncode == 2, argparse's usage-error status. With main()'s refusal disabled it now fails on assert 1 == 2 before the message is read.

Finding 5's code was right: both refusals exist and both refuse. Only the claim that both were held was wrong.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants