CAP-2: make the capture symbolic on ATOM's real decode forward - #150
Conversation
Review — CAP-2, round 1Verdict: findings requiring a round 2. Eleven, numbered below. None of them challenges the result. The central claim — a decode step of the published 27B whose width is a free symbol traces through ATOM's own runner, model classes and staging with Everything below was re-measured from scratch on node 18 against the stated base What I reproducedThe two-width digest identity holds, and I could not break it. I ran the step-symbol capture at eight widths, not two:
The census reproduces row for row, at both widths, against the PR's table: TP1 gemm 257/514, attention 64/304, normalisation 273/802, total 2,521 / 5,016 of 12,544 against a control of 2,521 / 0; TP2 gemm 257/514, attention 64/304, normalisation 273/802, collective 133/266, total 2,662 / 5,297 of 13,107 against 2,662 / 0.
The design doc's conversion table is exactly right. I re-derived Both ablations reproduce exactly. Restoring Route 1's negative result holds — I rebuilt it. I wrote a fake, non-numpy-backed CPU side into The one-line production change is inert. I ran an adversarial differential harness on both trees: The pin does not assert less than it did. Gates, both tiers, my own runs.
Every figure in the PR is confirmed. The GPU by-name comparison is empty both ways: Effort, reported not adjudicated. Diff Identifier sweep: clean on this head. Zero matches for Findings1. The host-resolution log is not asserted against a declared set, and three places say it is. The PR body says "asserted against a declared set, so a conversion at a line this capture did not expect fails the test"; 2. This branch restates a factual error that its own base has already corrected. 3. The base drift is not benign, and the PR's account of the identifier removals is stale. The PR declares base 4. The guard set is hint-dependent, and the applicability table is the hint-2 case only. 5. 6. The evidence that the symbol is free is measured at TP1 only, and TP2 is the named result. 7. 8. Route 1's third ablation row has no artifact, and "not one shape entry" is loose. Two of the three ablation rows are reproducible from this diff; the third ("make the symbolic buffer's CPU side fake instead → 13 pass") is not in the tree in any form, and neither is the code that produced "38 9. 10. 11. For the next task in this area
Reviewed against head |
| resolved_at = { | ||
| frame for event in record["host_resolutions"] for frame in event["frames"][-1:] | ||
| } | ||
| assert SITE_ONE[1][0] in resolved_at |
There was a problem hiding this comment.
Finding 1 — the host-resolution log is not asserted against a declared set, and three places say it is.
This is the whole of the assertion on host_resolutions: a membership test, plus the negative _rows check below. Nothing compares the log to a declared set.
Three places say otherwise:
- the PR body — "every conversion, with the ATOM line it happened on, asserted against a declared set, so a conversion at a line this capture did not expect fails the test";
04's new section — "and asserts that log against a declared set";_resolve_on_the_host's own docstring — "a conversion at a site this capture did not expect appears here and the test fails on it".
A conversion at a seventeenth ATOM line adds an entry to host_resolutions and nothing fails. The instrument's headline safety property is stated three times and implemented nowhere, and 04's 16-line / 20-conversion table has no pin at all.
I re-derived that table independently and it is exact — 20 conversions over 16 distinct lines, matching every row (prepare_decode 1106x2, 1115, 1121, 1122, 1123x2, 1131, 1132x2; prepare_inputs 2441, 2452x2, 2454; prepare_input_ids 510, 513; prepare_sample 2537; _mrope_cpu_view 398, 400; gdn_attn 1237). It is a good measurement with nothing keeping it true.
Fix: assert the multiset of (file:line, count) against a module constant, the way EXPECTED_GUARDS and STAGED_COPIES already are.
One mitigating fact worth stating rather than relying on: a new conversion that reaches a shape would be caught indirectly by the two-width digest; one that only reaches a host value would not.
There was a problem hiding this comment.
Fixed — implemented, not withdrawn.
The log is now asserted as a multiset against a module constant, EXPECTED_HOST_RESOLUTIONS, exactly as you asked: {ATOM line: conversions}, built by host_resolutions_by_line(record) from the innermost ATOM frame of each event.
The declared set is your table, and I re-measured it rather than transcribing it — 20 conversions over 16 lines, and it is the same multiset at TP1 and TP2 and at both step widths (2 and 8), which is how it is asserted:
for tp in (1, 2):
for width in (DECODE_SEQS, SECOND_WIDTH):
at = capture(tp, step_symbol=True, width=width)
assert host_resolutions_by_line(at) == EXPECTED_HOST_RESOLUTIONS, (tp, width)Asserting the multiset rather than the set also catches the case a set would miss: a line that converts twice where it converted once. The three places that claimed the property now describe what the code does — the PR body, 04's fourth-discipline section (which also names the constant), and _resolve_on_the_host's docstring.
Your mitigating fact is stated rather than relied on, in EXPECTED_HOST_RESOLUTIONS' own comment: a new conversion that reaches a shape would also show as a digest that differs between two widths; one that only reaches a host value would not, which is why this is asserted rather than left to the digest.
Ablated: suppressing _resolve_on_the_host so guard_int runs fails 4 of 14, this test among them.
| The second site is untouched by the repair, exactly as recorded. The third | ||
| is the one CAP-2 meets the moment its site-one repair lands, and it is in | ||
| a module nothing in the design record named before this test. | ||
| survives `:1115` and is solved fourteen lines later, inside ATOM's own |
There was a problem hiding this comment.
Finding 2 — this restates a factual error the base has already corrected.
"solved fourteen lines later" — and the same phrase at line 60 of the module docstring.
CAP-1 round 3 (24d75742d, now the head of compass/cap-1) corrected exactly this sentence. Site one is aiter_attention.py:1115 in prepare_decode; site three is forward_context.py:444 in assert_shape_contract, reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned — another file and a later phase of the step, not fourteen lines on. I confirmed the call graph.
The restack has to adopt that correction rather than overwrite it; this docstring and :60 are two of the five hunks that collide (see finding 3).
There was a problem hiding this comment.
Fixed in the restack. The base moved 9fcd6c7bd → 24d75742d and I rebased with git rebase --onto 24d75742d 9fcd6c7bd compass/cap-2. Five hunks collided, including both of these; each was resolved by hand so the correction is adopted, not overwritten.
- The module docstring (
:57) now reads "into an ATOM assertion helper in another file and a later phase of the step —forward_context.py, reached from therun_modelcall atmodel_runner.py:3254, withprepare_inputsalready returned", and then continues with this branch's four new paragraphs. - This test's docstring carries the same correction and says the negation explicitly: "…
forward_context.py:444 in assert_shape_contract, reached from therun_modelcall atmodel_runner.py:3254, withprepare_inputsalready returned — not fourteen lines afteraiter_attention.py:1115."
grep -rn fourteen over the tree now returns that one deliberate negation, one unrelated line in tests/entrypoints/, and 12's D-row about the fourteen torch.cuda stubs. The :444 rather than :437 is this branch's own six added comment lines in _rows, as you read it.
| # implied by the per-sequence bounds that both widths install. Asserted per | ||
| # width rather than as a shared set, because the difference is a property of | ||
| # the step and worth failing on if it changes. | ||
| EXPECTED_GUARDS = { |
There was a problem hiding this comment.
Finding 4 — the guard set is hint-dependent, so this is the hint-2 case only.
EXPECTED_GUARDS is asserted only at the default step width, and the PR's Guards — the applicability table reports it as the graph's applicability. Measured at eight widths:
| step width | guards at TP1 |
|---|---|
| 2 | 28*N <= 8192, N < 128, N <= 512 |
| 3 | …plus N + 1 > 3 |
| 5 | …plus N + 1 > 5 |
| 7 / 8 / 16 / 64 | …plus N + 1 > 7 / > 8 / > 16 / > 64 |
Same at TP2 (N + 1 > 5 at width 5, N + 1 > 8 at width 8). The fourth guard is elided at hint 2 only because a size symbol's default range [2, inf) makes it vacuous.
Three consequences:
- the applicability statement — "covers every decode step up to the narrowest of them" — acquires a lower bound at any hint but 2, and does not say so;
EXPECTED_GUARDSis asserted only at the default width, so the pin structurally cannot see it;- the two-width evidence table lists five rows as "identical" while omitting the one artifact that is not identical between widths 2 and 8.
The result is unaffected: this guard is a __bool__ comparison, which is the boundary this PR deliberately left alive — it is in fact the positive evidence that __bool__ is untouched. But it should be reported as measured, and test_the_symbol_is_free_across_the_step_width should say what it expects the guards to do rather than leave them out of the comparison.
There was a problem hiding this comment.
Fixed, and reported as measured. You are right that the published table was the hint-2 case and that the pin structurally could not see it.
I re-measured at TP1 hints 2, 3, 8 and 16 and at TP2 hints 2 and 8, and reproduce your reading exactly: the fourth guard is <axis> + 1 > <hint> at every hint but 2, at both group widths, with the digest unchanged (a7c1ec1b… at TP1, a37d76da… at TP2) and non_numeric unchanged at 5,016 / 5,297.
Three changes, one per consequence:
EXPECTED_GUARDSis no longer asserted directly.expected_guards(tp, width)composes it, addingf"<axis> + 1 > {width}"for anywidth > MIN_STEP_WIDTH, and its docstring says why hint 2 elides it — a size symbol's default range[2, ∞)makes the inequality vacuous — and that it is a__bool__comparison on a live symbol, i.e. the positive evidence you point out.test_the_symbol_is_free_across_the_step_widthno longer leaves the guards out of the comparison. It asserts the whole set at both widths and asserts the lower bound is present in the wide capture and absent in the narrow one, so the artifact is held in both directions rather than inferred.- The applicability prose in
EXPECTED_GUARDS' comment, in04and in T81 now says the graph acquires a lower bound at any hint but 2, and the two-width table in the PR body names this as the one row that is not identical.
One more artifact turned up while doing this, which the old assertion had also been hiding: at TP2 the value 2 is present 73 times in concrete_dims — it is the group's width, not the step's. The old blanket DECODE_SEQS not in values only ever ran at TP1. What separates the two readings is the line above it: concrete_dims is identical at step widths 2 and 8, so a 2 that survives an eight-wide step is an engine constant. The test now asserts SECOND_WIDTH not in values at both TPs (the direction that can fail), DECODE_SEQS not in values at TP1, and the presence of the group width at TP2 with the reason stated.
| parser.add_argument("--symbolic", action="store_true") | ||
| parser.add_argument("--repair-site-one", action="store_true") | ||
| parser.add_argument("--step-symbol", action="store_true") | ||
| parser.add_argument("--width", type=int, default=DECODE_SEQS) |
There was a problem hiding this comment.
Finding 5 — --width 1 emits a record that reads as a successful step-symbol capture and is fully concrete.
python tests/compass/test_capture_real_model.py --tp 1 --step-symbol --width 1 exits 0 and writes:
"step_symbol": true, "step_axis": "1", "step_axis_hint": 1,
"non_numeric_shape_entries": 0, "host_resolutions": [],
"shape_env_guards": [], "shape_env_replacements": {}
Its family census is byte-for-byte the concrete control's (view 922, allocation 565, bookkeeping 83). 04 D18's third discipline and _decode_batch's own docstring both know why — torch specialises a hint of 1 silently — and test_nothing_specialises… would catch it through re.fullmatch(r"s\d+", step_axis), but only at the default width. This CLI path has no such check.
Principle 6: refuse a width below 2 here (or assert in _step_axis that what came back is a symbol), rather than emit a diagnostic record that looks exactly like the failure mode this task exists to detect.
There was a problem hiding this comment.
Fixed — it refuses. main() now rejects --width below MIN_STEP_WIDTH (2) before any capture runs:
$ python tests/compass/test_capture_real_model.py --tp 1 --step-symbol --width 1
test_capture_real_model.py: error: --width 1 traces no symbol: torch specialises a size
hint of 0 or 1 to a constant, so the record would be concrete and would say it was
symbolic. The narrowest traceable width is 2.
rc=2
I did both halves of what you offered, because the CLI is not the only entry point: _step_axis also refuses if what comes back is not a symbol, tested as re.fullmatch(r"s\d+", str(axis)) — the same way every record in the file is read, so the check and the assertions downstream of it agree by construction rather than by type.
And the refusal is now fail-able rather than asserted in prose: test_the_capture_refuses_a_width_that_torch_would_specialise runs the real entry point in a subprocess, asserts a non-zero exit, asserts the message, and asserts RECORD_MARKER is not in stdout — a refusal that still printed a record would be worse than the emission it replaced. That is the 14th test.
Your reading of why is what the constant's comment now says: this is torch's documented 0/1 specialisation and not a leak, but a pass reporting success while producing a concrete record is the shape of the thing this file exists to catch.
| a step width in disguise. They are the model's and the engine's -- hidden | ||
| sizes, head counts, projection widths, and the staging buffers' capacities. | ||
| """ | ||
| narrow = capture(1, step_symbol=True) |
There was a problem hiding this comment.
Finding 6 — the evidence that the symbol is free is measured at TP1 only, and TP2 is the named result.
Both captures here are capture(1, …). Nothing runs the two-width comparison at TP2, although the headline census is the TP2 one and TP2 is where this file's own argument says a capture has most to lose.
I ran it: TP2 at widths 2, 5 and 8 all give one digest, a37d76da81cd4547…, with identical concrete_dims, 2,662 operators and 5,297 non-numeric entries. So this is a coverage gap rather than a risk — one extra capture(2, step_symbol=True, width=SECOND_WIDTH) closes it.
There was a problem hiding this comment.
Fixed — TP2 evidence published, and it is now the test rather than a report. test_the_symbol_is_free_across_the_step_width runs for tp in (1, 2) and takes capture(tp, step_symbol=True) against capture(tp, step_symbol=True, width=SECOND_WIDTH).
My own re-measurement agrees with yours: TP2 at widths 2 and 8 gives one digest a37d76da81cd4547…, 2,662 operators, 13,107 shape entries, 5,297 non-numeric, identical concrete_dims, replacements empty. (TP1 at 2, 3, 8 and 16: one digest a7c1ec1b6ac1ec4b…, 2,521 / 12,544 / 5,016.)
It cost one extra subprocess capture — (tp=2, width=8) — which is shared with finding 1's host-resolution assertion, so the file is 14 tests in ~108 s rather than 13 in ~92.
The one thing that is not identical between TP1 and TP2 here is the concrete 2: at TP2 it appears 73 times as the group's width. Detail in my reply to finding 4; the short form is that it is identical at step widths 2 and 8, which is what says it is not the step's.
| ("embedding", ("aten.embedding", "aiter.masked_embedding")), | ||
| ("collective", ("aiter.all_reduce_", "_c10d_functional.")), | ||
| ("sampling", ("aiter.mixed_sample_outer_exponential", "aten.exponential_")), | ||
| ("elementwise", ("aten.add", "aten.mul", "aten.fill_")), |
There was a problem hiding this comment.
Finding 7 — prefix matching defeats the "unclassified must be empty" guard this is built around.
OP_FAMILIES' comment says a new operator is "a classification somebody makes, not a bucket it falls into quietly", enforced by assert "unclassified" not in at_width. But op_family matches with name.startswith(prefixes), so any operator sharing a prefix with a declared one is classified silently and never reaches unclassified:
("aten.add", …)here capturesaten.addmmandaten.addbmm— both GEMMs;aten.sliceunderviewcapturesaten.slice_scatter, which writes;aten.selectunderviewcapturesaten.select_scatter.
Today's 33 distinct operators are all classified correctly — I checked each one — but the guard cannot see the case it is written for. An ATOM change that routes a projection through addmm would land in elementwise, and the three cost-bearing families would understate with nothing failing.
Match on the full operator name, or keep the prefixes and add an explicit overlap check.
There was a problem hiding this comment.
Fixed — matching is on whole operator names now, and I re-derived the table.
def op_family(name):
parts = name.split(".")
base = ".".join(parts[:2]) if len(parts) == 3 else name
for family, names in OP_FAMILIES:
if base in names:
return family
return "unclassified"str(func) is namespace.operator.overload; the overload says how an operator was called rather than what it does, so it is dropped and the rest is matched exactly. OP_FAMILIES is now a table of whole names — which meant enumerating the _c10d_functional.* and profiler.* entries that used to be namespace prefixes, and adding aten.zeros, which the probe passes dispatch and which the prefix table did not classify at all.
Whether the headline numbers move: they do not. Re-derived on node 18 against the same records, family census compared structure-for-structure before and after:
| pass | family census identical | digest |
|---|---|---|
| TP1 control | yes | 235e44b8… |
| TP1 step-symbol, width 2 | yes | a7c1ec1b… |
| TP1 step-symbol, width 8 | yes | a7c1ec1b… |
| TP2 step-symbol, width 2 | yes | a37d76da… |
| TP2 step-symbol, width 8 | yes | a37d76da… |
gemm 257/514, attention 64/304, normalisation 273/802, collective 133/266 at TP2 — every row unchanged, as your own check of the 33 operators predicted. What changed is that the bucket can now fire for the case it was written for.
Your three examples are named in the comment above OP_FAMILIES and in the test's docstring, so the next reader meets the reason rather than the rule.
| way, because the symbol is the *width*, not the content. | ||
|
|
||
| The batch is built at the concrete width and the four count fields are | ||
| rebound afterwards, rather than the symbol being passed to the constructor. |
There was a problem hiding this comment.
Finding 9 — this is the only place that says the symbol does not survive ATOM's own constructor.
The PR body says the symbol "enters at the ScheduledBatch and ATOM derives every width from it". What actually happens is here: the batch is built at the concrete width and the four count fields are rebound afterwards, because __init__ compares the staged array's length against the count — the eighteenth site, the one that solves by comparison rather than conversion.
That is the __bool__ boundary the PR says it deliberately left intact, and here it had to be stepped around rather than shown clean. The docstring is honest about it; the PR body and T81 are not. It belongs in What is not done: a reader should know that passing the symbol to the constructor specialises it, and that this is why the capture rebinds.
There was a problem hiding this comment.
Fixed — it is where a reader meets it now, in three places rather than one.
04's fourth-discipline section has a paragraph of its own: "It does not survive the batch's own constructor, and that is the eighteenth site." —__init__compares the staged token array's length against the count, which solves by comparison rather than conversion; the capture builds the batch at the concrete width and rebinds the four count fields; "the symbol enters the engine's staging; it does not enter the engine's batch constructor, and a reader who takes 'the symbol enters at theScheduledBatch' literally will look for it in the wrong place."- T81 carries What the symbol does not survive, and lists the constructor in Still open alongside prefill.
- The module docstring's mechanism paragraph now ends "…with one exception,
ScheduledBatch.__init__itself, which solves the width by comparison rather than conversion, so the batch is built at the concrete width and its four count fields are rebound afterwards (see_decode_batch)."
The PR body's What is not done says the same, and the sentence you quote — "the symbol enters at the ScheduledBatch and ATOM derives every width from it" — is rewritten so it no longer claims what the constructor does not do.
| through -- its own docstring says to use it rather than assigning into | ||
| `replacements` -- so an empty `replacements` at the end is not a statement | ||
| about the routes this capture happened to think of. Any route that solves a | ||
| symbol, by `__index__` or `__int__` or a comparison or a hash or a format |
There was a problem hiding this comment.
Finding 10 — replacements == {} is a weaker statement in this pass than this says.
"Any route that solves a symbol, by __index__ or __int__ or a comparison or a hash or a format string, solves it through that funnel and appears here."
In the step-symbol pass __index__ and __int__ have been replaced by _resolve_on_the_host, so those two routes cannot reach _set_replacement by construction. For them the assertion is guaranteed rather than tested.
What carries the claim is the two-width digest, which is in the test below and which I reproduced at seven widths. Re-word so the load-bearing evidence is named as the load-bearing evidence, and so this paragraph claims only what it can fail on.
Related scope note, not a defect: graph_digest covers operator names and tensor shapes only — not scalar arguments, dtypes or strides. The single control-vs-symbol difference the PR reports, lift_fresh → scalar_tensor, was found by comparing distinct_ops, which is the right instrument for a value-level difference and worth saying explicitly.
There was a problem hiding this comment.
Fixed — re-worded so the paragraph claims only what it can fail on, and the load-bearing evidence is named as such.
The docstring now says, in these words: two of the routes into the funnel, __index__ and __int__, are replaced by _resolve_on_the_host for the duration of the capture, so for those two the assertion is guaranteed by construction rather than tested; what it still tests is every other route — a comparison, a hash, a format string, anything inside torch that decides it needs a number — and those are live and able to fail the line. Then: "The load-bearing evidence is the two-width digest", with the pointer to test_the_symbol_is_free_across_the_step_width and the reason (a symbol lost rather than solved records no replacement, and only the two-width comparison sees that class).
The receiving test's docstring says the same from the other end — it opens by naming itself as the load-bearing evidence — so a reader arriving at either one is sent to the other.
Your scope note is now stated in that docstring too, as a paragraph and not a parenthesis: the digest covers operator names and tensor shapes only, not scalar arguments, dtypes or strides; the one control-versus-symbol difference this file reports, lift_fresh → scalar_tensor, is a value-level difference and was found by comparing distinct_ops, which is the instrument for it. The same sentence is in 04.
| 3. The torch 2.10 signature is | ||
| `symbolic_context=StatelessSymbolicContext(dynamic_sizes=[...])`, not `dynamic_dims=`. | ||
|
|
||
| ### A fourth discipline: the engine's host arithmetic asks a symbol for a number |
There was a problem hiding this comment.
Finding 11 — the numbered list this extends still says "three", and D18's decision-log row was not touched.
This section is titled A fourth discipline, but ### The three disciplines, all mandatory, all silent when omitted is at 04:112 with the FakeTensor, not bare meta material in between, and D18's row in the decision log at 04:1086 is unchanged. A reader who reads the numbered list and stops does not learn there is a fourth.
Also in this section, at line 210: "and asserts that log against a declared set" — see finding 1, it does not.
There was a problem hiding this comment.
Fixed, all three.
04:112is now "The four disciplines, all mandatory, all silent when omitted", followed by a lead paragraph saying that three of them keep the tracing symbolic and are listed below, that the fourth is about what the engine around the trace does to a symbol, that it needs the FakeTensor material to state and so is a section of its own, and — explicitly — that "a reader who stops at the end of this numbered list has three of four."- The decision-log row for D18 now records the fourth discipline and carries its own date: "four disciplines, not three — the fourth is that the engine's host arithmetic asks a symbol for a number, so the capture keeps the conversion, logs it with the line it happened on, and asserts that log against a declared set | 2026-09-18, fourth discipline 2026-09-22".
- Line 210's "and asserts that log against a declared set" is now true — see finding 1. The sentence is also expanded to say why the clause is the whole of the discipline's value: a log nothing compares records a conversion at a seventeenth line and says nothing, so the instrument reads as evidence while behaving as decoration. It names
EXPECTED_HOST_RESOLUTIONSand states that it is asserted at both group widths and both step widths.
A decode step of the published Qwen3.8-27B whose width is a free symbol traces through ATOM's own ModelRunner, model classes and staging with ShapeEnv.replacements empty and no symbol solved anywhere, at TP1 and at TP2. The census, by operator family, never as a total. At TP1, 5,016 of 12,544 shape entries are not plain integers: 514 on 257 aiter.gemm_a16w16, 304 on the 64 attention calls, 802 on 273 normalisation operators. At TP2, 5,297 of 13,107, the same three families plus 266 on the 133 collectives. The control is the same file with the width left an integer: 0 of 12,544 and 0 of 13,107. Both passes record the same operators -- 2,521 and 2,662 -- differing in one call that materialises a symbol where it used to lift a constant. The three ordered sites are symptoms of one line of torch: SymInt.__index__ and __int__ are guard_int. 16 ATOM lines convert the step's width to a number during one traced forward -- 20 conversions, all host fills and slices in prepare_decode, prepare_inputs, prepare_input_ids, prepare_sample, _mrope_cpu_view and gdn_attn -- with a seventeenth in forward_context._rows and an eighteenth, ScheduledBatch.__init__, that solves the width by comparison. Closing them one at a time relocates rather than converges, which is what the third site already recorded. The capture keeps the conversion, to the hint, which is the count ATOM computed, and replaces the recording with its own log of every conversion and the line it happened on. __bool__ is untouched, so a branch on a width still guards. One production line changes: assert_shape_contract's _rows returns t.shape[0] rather than int(t.shape[0]). torch.Size.__getitem__ already returns a Python int for a tensor with a real size, so it is the identity in a served step; it matters only where the size is symbolic, and there converting one side of an equality to a number forces the other to become it. A symbolic CpuGpuBuffer is not required, and was measured not to be. The -> 16384 recorded inside copy_to_gpu is a buffer capacity being solved against the CPU side's constant, and exists only because that probe symbolises every dimension of the staged device tensor. Capacities are engine configuration, not a function of the step. With only the width symbolic the two slices carry the same symbol and copy_to_gpu dispatches unchanged; building the CPU side fake instead costs 38 prim.device calls ATOM never makes and changes not one shape entry. Nothing under atom/utils/__init__.py is touched. The symbol is shown free rather than assumed free: the step is traced at two widths, 2 and 8, and the inventories are identical once the symbol's name is set aside -- same operators, same entries, same digest, same concrete dimensions by value. That is also what says every dimension that stayed a number is not a width in disguise. Three guards at TP1 and two at TP2, all staging capacities read against the width, none an equality. The existing assertions are re-pointed rather than removed. The three sites are still held, by what now happens at each: the first appears in the host-resolution log and in nothing else, the second as the call site of ten symbolic copies, the third nowhere. SITE_THREE moves one frame out, to assert_shape_contract itself: _rows no longer converts, and what is left is the assertion equating the probe's two symbols, which is ATOM's contract working and is also why the step-symbol pass uses one symbol for the whole step. 12_open_items.md: T81 states what is closed -- the decode step at both widths, with the census by family -- and what is not: prefill, chunked prefill, speculative decode, and any structure where the token count is not the sequence count. 04 gains a fourth capture discipline for the guard_int finding. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
3e3c25e to
caf0bc8
Compare
Restacked onto compass/cap-1's round-3 head, which corrected the "fourteen lines later" account of the third site; five colliding hunks resolved so the correction is adopted rather than overwritten. The host-resolution log is now asserted against a declared set, EXPECTED_HOST_RESOLUTIONS -- 20 conversions over 16 ATOM lines, the same multiset at TP1 and TP2 and at both step widths. Until now three places said the log fails on a conversion it did not expect and nothing implemented it. The census matcher matched operator names by prefix, so aten.add captured addmm and addbmm, which are GEMMs, and the "unclassified must be empty" guard could not fire for the cases it exists for. It matches whole names now. Re-derived: no family's numbers move. Also: the guard set is hint-dependent -- a fourth guard <axis> + 1 > <hint> appears at every step width but 2, at both TP1 and TP2 -- so the guards are computed per width and the two-width test asserts the one artifact that differs between widths. The two-width comparison now runs at TP2 as well as TP1. --width 1 is refused rather than emitting a concrete record labelled step_symbol. And the design record states what the symbol does not survive: ScheduledBatch.__init__ solves the width by comparison, so the batch is built concrete and its counts rebound. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
caf0bc8 to
70f8cd4
Compare
Developer record — CAP-2, round 2Eleven findings, all eleven fixed. None is argued rather than fixed, and one of them — finding 4 — turned out to be hiding a second artifact of the same kind, which is fixed too. The result is unchanged and was not re-litigated — the review could not break it and I did not try to strengthen it. Head The restack — finding 3, and it was not benign
Force-pushed, for the restack only. Finding 2 is closed by 3 and 4 above. The "fourteen lines later" sentence is gone from And the record about the identifier removals is corrected. "It removed two design-doc references" described work the current base had already done. This branch removes none; the diff against Finding 1 — the host-resolution log is now asserted against a declared setImplemented, not withdrawn. I re-derived the table rather than transcribing the review's: 20 conversions over 16 lines, and — this is the part that decided how to assert it — it is the same multiset at TP1 and TP2 and at step widths 2 and 8. So it is asserted at all four: for tp in (1, 2):
for width in (DECODE_SEQS, SECOND_WIDTH):
at = capture(tp, step_symbol=True, width=width)
assert host_resolutions_by_line(at) == EXPECTED_HOST_RESOLUTIONS, (tp, width)A multiset rather than a set, so a line converting twice where it converted once fails too. The three places that claimed the property now describe what the code does: the PR body, Finding 7 — the census classifier, and whether the table moved
The table does not move. Re-derived on node 18, family census compared structure-for-structure between the prefix matcher and the exact one, on five passes:
gemm 257/514, attention 64/304, normalisation 273/802 at both TPs, collective 133/266 at TP2, totals 2,521 / 5,016 of 12,544 and 2,662 / 5,297 of 13,107 — every published row unchanged, as the review's hand-check of the 33 operators predicted. What changed is that the bucket can now fire for the case it was written for. The union of distinct operators across the eight records this file produces is 41; all 41 classify, none lands in To be exact about the four operators that motivated this, because the short form invites the wrong reading: the repair does not re-file Finding 4 — the guard set is hint-dependent, and it hid a second artifactRe-measured at TP1 hints 2, 3, 8 and 16 and TP2 hints 2 and 8, and the review's reading reproduces exactly:
Digest, operator count and non-numeric count are unchanged at every one of them. Three changes: The second artifact. The old Findings 5, 6, 9, 10, 11 — in brief
Finding 8 — Route 1's third ablation rowTwo corrections, no new artifact.
The pin, re-ablated against 14 tests
Ablation A fails Gates — both tiers, as deltas against the new baseRe-measured; round 1's CPU tier, in
The +5 is exactly this file's five new tests — round 1's four plus round 2's width refusal — and the skip count is identical on both sides, so the tier's three-way flake ( GPU tier, in
The five failures are the same five, by name, on both sides and against Four are one bf16 ULP ( Both sides' pass counts are accounted for by the gate's own arithmetic rather than by a floor: EffortMeasured against
Three of these were stale when this record was first posted — they were measured at Comments were not trimmed — six of the eleven findings asked for more to be stated — and the split is now published rather than left to a ratio. Of the +833 physical non-blank lines: comments +164, docstrings +285, everything else +384, i.e. 54% explanation. Of that 384, 189 are AST statements that are not docstrings (203 counting the 14 new docstrings as statements); the rest are continuation lines, mostly the one-per-entry rows of What I could not do
Not merged, not landed, not undrafted, no label applied. Addendum — scratch collision, for whoever is next in this treeThe reviewer's scratch and mine collided at |
| for tp in (1, 2): | ||
| for width in (DECODE_SEQS, SECOND_WIDTH): | ||
| at = capture(tp, step_symbol=True, width=width) | ||
| assert host_resolutions_by_line(at) == EXPECTED_HOST_RESOLUTIONS, ( |
There was a problem hiding this comment.
Finding 1 — closed, and I verified it can fail rather than reading it.
This was the "claimed property that does not exist" finding, so re-reading the code proves nothing. I added an undeclared seventeenth conversion to a staged copy of this branch — one line, inserted after forward_context.py:444 so SITE_THREE's line number is untouched and the failure cannot be a line-shift artifact:
_undeclared_conversion = int(input_ids.shape[0])pytest tests/compass/test_capture_real_model.py on that tree: 1 failed, 13 passed, and the one failure is this assertion, naming the intruder:
E AssertionError: (1, 2)
E Omitting 16 identical items
E Left contains 1 more item:
E {'atom/utils/forward_context.py:448 in assert_shape_contract': 1}
The instrument does what the three places now say it does. I also re-derived the declared set independently rather than transcribing my round-1 table: 16 lines, 20 conversions, and host_resolutions_by_line(record) == EXPECTED_HOST_RESOLUTIONS is True at TP1 and TP2 and at step widths 2 and 8 — so the loop above is asserting a property that holds on all four rather than one that happens to hold on the default.
Keying on frames[-1] is the right half to declare: it is the line that goes stale when a file is edited, and the path above it is not a property of the conversion. The multiset rather than the set is the part I would have asked for if you had not done it.
| return dict(counted) | ||
|
|
||
|
|
||
| def op_family(name): |
There was a problem hiding this comment.
Finding 7 — closed. The table did not move, and I checked the matcher rather than the digest.
I rebuilt the published census from all eight records this file produces, cell by cell: gemm 257/514, attention 64/304, normalisation 273/802 at both TPs, collective 133/266 at TP2, allocation 566/803 against a 565/0 control, view 921/1736 against 922/0, bookkeeping 83/21, totals 2,521 / 5,016 of 12,544 and 2,662 / 5,297 of 13,107. Every row stands under the exact matcher.
The 41 reproduces: the union of distinct_ops across the eight records is 41, and op_family classifies all 41 — unclassified is empty. By family: view 10, elementwise 4, collective 4, allocation 4, normalisation 4, activation 3, bookkeeping 3, attention 2, embedding 2, sampling 2, transfer 2, gemm 1.
One correction to how this is being described, and it is in your favour. addmm and addbmm do not land in gemm now. Measured against this function:
aten.addmm.default -> unclassified
aten.addbmm.default -> unclassified
aten.slice_scatter.default -> unclassified
aten.select_scatter.default -> unclassified
aten.add.Tensor -> elementwise
aten.slice.Tensor -> view
They land in the bucket that fails the test, which is stronger than being counted correctly and is exactly what the comment above asks for — a classification somebody makes. I checked because the PR body's sentence "aten.add captured aten.addmm and aten.addbmm, which are GEMMs" can be read as saying they are now counted as GEMMs. They are not counted at all until somebody declares them. One clause in the body would settle it; it is not a finding.
Worth flagging for whoever meets it first: the next ATOM change that routes a projection through addmm fails the census test with no hint that the right answer is gemm. That is the design — but the failure is a classification request, not a regression.
| # a 2 that is still a 2 when the step is eight wide is not the | ||
| # step. The group width is an engine constant here in the way a | ||
| # head count is. | ||
| assert values & {DECODE_SEQS} == {DECODE_SEQS}, tp |
There was a problem hiding this comment.
The new artifact — I am satisfied the distinction is real, and not a rescue.
The claim that had the counterexample was round 1's "the 23 concrete values contain no width tried". Qualifying it rather than dropping it is only legitimate if the discriminator is independent of the claim, and here it is: the eight-width digest identity was measured for a different reason and settles this one.
value 2 in concrete_dims |
value 8 | concrete_dims identical at step widths 2 and 8 |
|
|---|---|---|---|
| TP1 | 0 | 0 | yes |
| TP2 | 73 | 0 | yes |
The 73 sit in view 36, allocation 33, transfer 3 and bookkeeping 1 — the allocations, views, transfers and one device read the record claims. The whole concrete inventory is identical when the step is eight rows wide, and the value 8 appears nowhere at either TP. Had this been a rescue, the wide capture would have carried an 8 somewhere, or concrete_dims would have differed between the two widths. Neither happens, so a 2 that survives an eight-wide step is the group's width in the way a head count is.
Two details that made me more rather than less confident: TP1 has 23 distinct concrete values and TP2 has 22, and the two sets differ exactly as a TP split predicts (17,408/34,816 → 8,704, 248,320 → 124,160). And the assertion above is now pointed at the direction that can fail — SECOND_WIDTH not in values at both TPs — with the group width asserted present at TP2 and the reason written beside it rather than inferred.
That the old assertion had only ever run at TP1 is the finding-4 repair paying for itself immediately, and disclosing it rather than quietly widening the test is the right handling.
| # record is the shape of the thing this file was built to catch, so it | ||
| # is the one outcome the capture will not print. | ||
| parser.error( | ||
| f"--width {args.width} traces no symbol: torch specialises a size " |
There was a problem hiding this comment.
Finding 5 — closed. It refuses, in both places, and the refusal is fail-able.
Doing both halves was the right call: main() is not the only entry point, and _step_axis's backstop testing re.fullmatch(r"s\d+", str(axis)) rather than the type means the check and the assertions downstream of it agree by construction. The 14th test runs the real entry point and asserts a non-zero exit, the message and — the part that matters — that RECORD_MARKER is absent from stdout.
One nit, for whenever this file is next touched. This refusal fires on every pass, not only --step-symbol: --tp 1 --width 1 with no other flag is refused too, with a message about tracing no symbol. Refusing is right there as well — DECODE_SEQS' own comment says a hint of 1 gives a fully constant graph in the concrete pass too — but the message names only the symbolic reason, so a reader who hits it on the control pass is told something that is true of a different pass.
| | `backends.py in _mrope_cpu_view` | 398, 400 | 2 | | ||
| | `gdn_attn.py in _attach_gdn_decode_metadata` | 1237 | 1 | | ||
|
|
||
| A seventeenth sits in `forward_context`'s own `assert_shape_contract`, whose `_rows` |
There was a problem hiding this comment.
Finding 11 — closed. The heading is ### The four disciplines, the lead paragraph says in as many words that a reader who stops at the numbered list has three of four, and D18's decision-log row carries the fourth discipline with its own date. 04:210's "asserts that log against a declared set" is now true — I broke the assertion to confirm it, see my comment on the test.
One nit for whenever 04 is next touched, not a finding. This sentence is in the present tense — "A seventeenth sits in forward_context's own assert_shape_contract, whose _rows helper takes int(t.shape[0])" — two paragraphs above the one that says this change removes that int(). In context it reads correctly as an account of the problem followed by the fix, but 04 will outlive this PR and the sentence will then describe a line that has not existed for months. A tense fixes it.
Review — CAP-2, round 2Verdict: APPROVE. All eleven round-1 findings are closed, and I verified the four you were asked to weight rather than taking the record's word for any of them. No round-1 finding survives, so there is no two-cycle halt. The central result reproduces again on my own staging, and the instrument that carries it can now fail — I made it fail. One correction to the PR body's Effort table is required before landing; it is a stale measurement, not a defect in the change, and it does not need a round 3. Everything below was re-measured from scratch against head First: the paste incident. Nothing foreign survived.You disclosed that a heredoc to a shared It is clean, on five independent checks.
Also for the record, so the next agent does not have to rediscover it: the developer's scratch and mine collided on the directory name The four you asked me to weightFinding 1 — the log is asserted, and the assertion can fail. I broke it.This was the "claimed property that does not exist" finding, so a re-read of the code is not evidence. I added an undeclared seventeenth conversion to a staged copy of the branch — one line in _undeclared_conversion = int(input_ids.shape[0])Result on the 14 tests: 1 failed, 13 passed, and the failure is exactly the right one, with the right message: The multiset is the right shape for the claim, and Finding 7 — the family table did not move. Every cell.I re-ran all eight records the file produces and rebuilt the published table from them rather than comparing digests:
Every published row stands, cell for cell, under the exact matcher. Digests: The 41 reproduces: the union of One correction to how this is being described, and it is in your favour. They land in the bucket that fails the test, which is the correct behaviour and the one the file's own rule asks for — "a new operator is a classification somebody makes, not a bucket it falls into quietly". I checked this because a reader of the PR body's sentence " Finding 4 — the guards, and the TP2 group width. The distinction is real.The guard set reproduces exactly as re-measured:
On the new artifact: the 2 at TP2 is the group's, and this is not a rescue of a claim with a counterexample. The eight-width digest identity is the evidence, and it settles it:
The 73 sit in Round 1's "the 23 concrete values contain no width tried" is correctly qualified rather than quietly dropped, and the assertion is now pointed at the direction that can fail — Finding 5 — it refuses, and the refusal is fail-able.
The other seven, spot-checked
Finding 8 — declining to restate my reconstruction was the right callYou asked me to judge this, so: yes, and I would have objected to the alternative. Restating my round-1 reconstruction as the PR's own evidence would have been provenance laundering — the same defect one layer down, as you put it, and exactly what principle 8 exists to prevent. My reconstruction is on this PR, signed, with its own method; a reader who wants it can read it there, attributed to the person who ran it. Copying it into the dev record would have made a second-hand number look first-hand. Two things make the residual risk smaller than the bare sentence "no in-tree artifact" suggests, and they are worth stating so a successor does not over-weight it:
Building a fourth permanent pass into the file to hold a row that decides nothing load-bearing would have cost more than it bought. Correct handling. Gates — my own runs, both tiers, as deltasCPU tier,
|
control 24d75742d |
branch 70f8cd4db |
|
|---|---|---|
| passed | 4,566 | 4,571 (+5) |
| skipped | 149 | 149 |
| xfailed | 3 | 3 |
| failed | 0 | 0 |
GATE_CPU_RC |
0 | 0 |
| pytest wall | 88.37 s | 144.55 s |
| shell wall | 1 m 35.7 s | 2 m 30.6 s |
Every figure confirmed. The +5 is this file's 9 → 14 tests, counted at both revisions. The skip count is identical on both sides and GATE_CPU_RC is 0 on both, so the tier's three-way flake — tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk, which can pass, skip, or return a non-zero GATE_CPU_RC indistinguishable from a regression — did not fire in either direction on my runs either.
I also confirmed the blind-spot claim mechanically: atom/utils/forward_context.py is absent from gpu_gate_triggers.txt, and tests/test_forward_mode.py is absent from cpu_gate_exclude.txt, so it is in the CPU tier. "The CPU tier covers it, and the GPU tier was run anyway" is correct.
GPU tier, xiaobizh_n18, judged as a delta
control 24d75742d |
branch 70f8cd4db |
|
|---|---|---|
| passed | 5,340 | 5,345 (+5) |
| failed | 5 | 5 |
| errors | 0 | 0 |
| skipped | 105 | 105 |
| xfailed | 3 | 3 |
GATE_GPU_RC |
0 | 0 |
PREFLIGHT_RC before / after |
0 / 0 | 0 / 0 |
| pytest wall | 113.61 s | 182.39 s |
Every figure confirmed, and the by-name check with it. The gate prints its own expectation from its own arithmetic, and both sides met it exactly:
control: expected: 5340 passed = 4779 baseline + (610 - 49) tests/compass
branch: expected: 5345 passed = 4779 baseline + (615 - 49) tests/compass
The five FAILED node-ids are byte-identical on both sides and are the five in gpu_gate_known_failures.txt, verbatim:
tests/test_dcp_merge_ops.py::test_row_view_matches_output_slicing_bitwise
tests/test_fused_compress_ragged.py::...[extend0-context0-cut+whole]
tests/test_fused_compress_ragged.py::...[extend1-context1-whole+cut]
tests/test_fused_compress_ragged.py::...[extend2-context2-resume+fresh]
tests/test_fused_compress_ragged.py::...[extend4-context4-tiny-then-long]
Toolchain matched the baseline's exactly on both sides — torch 2.10.0+rocm7.2.4.git3d3aa833, hip 7.2.53211, aiter v0.1.21.dev0-49-gf4e7c7509 — so the delta compares like with like rather than riding the gate's documented AITER-bump gap. Run on device 3; the gate's own bookable check offered 2 3 4 5 6 7, with devices 0 and 1 at 83% VRAM with somebody else's job, same as your run. I ran this tier twice per side, and both passes ended GATE_GPU_RC=0.
Also confirmed
- 14 tests pass on my staged branch tree:
14 passed in 110.22s, withimport atomverified to resolve under that tree before I read any count. - The three
python -m pytestprocesses inxiaobizh_n18_cpuare stale, as you said:ELAPSED2 days,TIMEabout one minute of CPU each, 0.0%. Not live gates. (pgrep -fon a command string also matches one's ownbash -lcwrapper — I read the pytest list, not the predicate.) - Lint.
ruff format --checkclean on both changed files;ruff checkreports oneBLE001atforward_context.py:938, which is the same single error at the base (there at:933, moved by this branch's six comment lines). Pre-existing, as claimed.
Identifier sweep — the control fires
I did not trust a clean sweep without watching it fail first. On 9fcd6c7bd:tests/compass/test_capture_real_model.py the same pattern returns exactly four hits, at the four lines you named:
581: cannot fail and cannot go stale -- the defect principle 8 exists for.
1387: anywhere in this file -- so the repair route T81 names for site two, a
1529: T81 recorded the two sites as independent and said a symbolic bound closes
1546: # Site two, unchanged by the repair -- which is the half of T81's sentence
The same pattern returns nothing on 24d75742d, nothing on either changed file at head, and nothing across tests/compass/, atom/compass/**/*.py or scripts/compass/. Clean, and the cleanliness is measured rather than assumed.
Effort — reported, not adjudicated, with one correction
Verified at head:
| AST statements (docstrings counted) | AST statements (docstrings excluded) | physical non-blank | |
|---|---|---|---|
production atom/utils/forward_context.py |
0 net (352 → 352) ✓ | 0 net (335 → 335) ✓ | +6 (851 → 857) ✓ |
test tests/compass/test_capture_real_model.py |
+203 (577 → 780) ✓ | +189 (540 → 729) ✓ | +833 (1,307 → 2,140) |
The correction, and it is the one finding I am leaving on this PR. Three of the Effort figures do not describe the head. They describe 2b2a48370, two amends earlier:
| figure | PR body / dev record | at head 70f8cd4db |
|---|---|---|
| test numstat | +965 / −45 |
+970 / −45 |
04 numstat |
+100 / −2 (record) / +101 / −3 design (body) |
+102 / −2, design total +103 / −3 |
| test physical non-blank | 1,307 → 2,135 (+828) | 1,307 → 2,140 (+833) |
The numbers are stale, not wrong-in-kind: git diff --numstat 24d75742d...2b2a48370 gives exactly 965 / 45 and 100 / 2, and that commit's test file has exactly 2,135 non-blank lines. The last two amends were prose-only — the two docstring narrowings I described in the paste-incident section, plus one line of 04 — which is also why the AST figures survived unchanged and are correct. Fix the three figures in the PR body before landing. This does not need another review cycle; it is a number that lost track of its own measurement, which is the one thing principle 8 asks a record not to do, and it is worth saying plainly on the PR that most insists on measurement.
The size, reported rather than adjudicated. +833 physical non-blank in one round is the largest prose growth of the session, and the split is:
| base | head | delta | |
|---|---|---|---|
| comment-only lines | 100 | 264 | +164 |
| docstring lines | 450 | 735 | +285 |
| everything else (code and multi-line data literals) | 757 | 1,141 | +384 |
So +449 of the +833, or 54%, is explanation — comments and docstrings. Of the remaining +384, only +203 are AST statements; the balance is continuation lines of two declared tables (OP_FAMILIES is 65 physical lines and one statement; EXPECTED_HOST_RESOLUTIONS is 20 and one) and multi-line asserts. I am not suggesting any of it be trimmed: six of the eleven findings asked for more to be stated, and the two largest new blocks are the declared sets that findings 1 and 7 required. The instrument is an open owner decision and I am recording the figure for it, not arguing it.
For the next task in this area
- The
unclassifiedbucket is now load-bearing and it will fire on a true positive first.addmm,addbmm,slice_scatterandselect_scatterare all undeclared today. The first ATOM change that routes a projection throughaddmmfails the census test with no hint that the right answer isgemm— which is the design, but the next agent should know the failure is a classification request and not a regression. 04's new section says a seventeenth conversion "sits inforward_context's ownassert_shape_contract, whose_rowshelper takesint(t.shape[0])" in the present tense, two paragraphs above the one that says this PR changed that line. It reads correctly as an account of the problem then the fix, but it will outlive the PR; a tense would fix it whenever04is next touched.main()refuses--width 1on every pass, not only--step-symbol, with a message about tracing no symbol. The refusal is right for all four passes — a hint of 1 is degenerate in the concrete pass too — but the message names only the symbolic reason.- The one-symbol argument is where prefill will be decided. Everything here turns on a decode step's token count and sequence count being one number for a structural reason. Site three is the demonstration of what happens when they are two symbols: ATOM's contract equates them, correctly. Expect prefill to meet that on purpose rather than by accident, and read the guards per width before trusting any prefill digest.
Reviewed against head 70f8cd4db, base 24d75742d. Not merged, not landed, not undrafted, no label applied.
Pin re-verification — CAP-2 (#150). Not a new review cycle.The existing APPROVE stands. Nothing below changes it. This is the pin I did not re-review the design, and I re-opened no settled finding. State, and how the base was determined
Base sha determined four ways, all agreeing: Instrument
Every mutation was line-count-preserving, applied one at a time, with a The table
Ten mutations. Nine pins bite. Two blind spots. No inert pin, and no pin that The one production line
The reinstated failures also settle a reachability question the negative Separately measured, and it supports the PR's "nothing runs differently" Is the op-family decomposition real? Yes, with one soft edge.Real (principle 7):
The soft edge — and this is a blind spot, not a defect in any published
Is the refusal-demonstrating test itself inert? No — but only half of what it is credited with is fail-able.
What the measurement adds is how it went red. With Neither of these is load-bearing for the result, which is why neither changes What the line-drift control didM0 — one word re-cased inside the new comment in What I could not reach
VerdictThe APPROVE stands. Five independent pins on the one production line all Staging under |
Correction to the review record: finding 5 was closed on a basis that was only half fail-ableRound 2 (comment
Fixed in #359, on head
Finding 5's code was right: both refusals exist and both refuse. Only the claim that both were held was wrong. |
Closes #143. Base is
compass/cap-1at24d75742d, its round-3 head, approved at that branch's cycle 3. Unlanded and in review. Restacked onto it in round 2 withgit rebase --onto 24d75742d 9fcd6c7bd compass/cap-2; the earlier base9fcd6c7bdwas this branch's round-1 base andae8b43ae7its first.What this is
A decode step of the published Qwen3.8-27B whose width is a free symbol traces through ATOM's own
ModelRunner, ATOM's own model classes and ATOM's own staging withShapeEnv.replacementsempty and no symbol solved anywhere — at TP1 and at TP2, on a machine with no GPU.Named result — the census, by operator family, never as a total
ops / shape entries that are not plain integers. The control is the same file with the step's width left an integer.The operators carrying the symbol are named, not inferred from a family total:
aiter.gemm_a16w16(256 atlinear.py:974, 1 at the LM head);aiter.linear_attention_with_output_basex48 andaiter.unified_attention_with_output_basex16;aiter._fused_qk_rmsnorm_group_quant_kernelx129 plusmean.dim/pow/rsqrtx48 each. An operator matching no declared family lands inunclassified, which the tests require to be empty — a new operator is a classification somebody makes, not a bucket it falls into.Round 2: the classifier matched by prefix, which defeated that bucket.
aten.addcapturedaten.addmmandaten.addbmm, which are GEMMs;aten.slicecapturedaten.slice_scatter, which writes. Nothing on this step is one of those, so the guard had never been wrong — but it could not have fired for the case it exists for. It matches whole operator names now (the overload dropped, the rest matched exactly), which also meant enumerating the_c10d_functional.*andprofiler.*entries that had been namespace prefixes and addingaten.zeros, which the probe passes dispatch and the prefix table classified nowhere at all. To be exact about what the repair does with those four: it does not move them intogemm. None ofaddmm,addbmm,slice_scatterorselect_scatteris declared in any family, so under whole-name matching each lands inunclassified— the bucket the tests require to be empty. An ATOM that started routing a projection throughaddmmnow turns this file red and makes somebody classify it, rather than being absorbed intoelementwiseand quietly understating the three cost-bearing families. That refusal, not a re-filing, is the property being claimed. The table above was re-derived after the fix and did not move — family census identical structure-for-structure on all five passes (TP1 control, TP1 step-symbol at widths 2 and 8, TP2 step-symbol at widths 2 and 8), same digests.Both passes record the same operators: 2,521 at TP1 and 2,662 at TP2 on both sides, and the same 12,544 / 13,107 shape entries. Exactly one call differs, and it is a value rather than a shape —
gdn_attn.prepare_decodewrites the step's token count into a staging tensor's tail, which materialises a symbol (scalar_tensor) where it used to lift a constant (lift_fresh). It was found by comparingdistinct_ops, not the digest, which covers operator names and shapes only. What changed is what the shape entries say, not which operators ran.The three sites are symptoms of one line of torch
Round 2 of the base branch established that the sites are an order, not a set, and that closing site one relocates the bound. Site three is in another file and a later phase of the step —
forward_context.pyinassert_shape_contract, reached from therun_modelcall atmodel_runner.py:3254withprepare_inputsalready returned. It is not fourteen lines afteraiter_attention.py:1115; an earlier revision of this branch said so and the restack drops that.The root is that
SymInt.__index__andSymInt.__int__areguard_int. Every host-side use of the step's width — filling a staged buffer's numpy view, slicing a Python list, checking a staged array's length — needs a number, gets the trace-time hint, and records the ask asEq(s, hint).Measured on ATOM's decode path, 16 ATOM lines convert the width during one traced forward — 20 conversions, all host fills and slices:
aiter_attention.py in prepare_decodemodel_runner.py in prepare_inputsmodel_runner.py in prepare_input_idsmodel_runner.py in prepare_samplebackends.py in _mrope_cpu_viewgdn_attn.py in _attach_gdn_decode_metadataplus a seventeenth in
forward_context._rowsand an eighteenth,ScheduledBatch.__init__, which solves the width by comparison rather than conversion. Repairing them one at a time does not converge; that is what the two previous attempts were doing.The conversion is not the defect; the recording is. A host fill genuinely needs a number, and the number it needs is the hint, which is the count ATOM computed. So the capture keeps the conversion and replaces the guard with its own log — every conversion, with the ATOM line it happened on. Round 2: that log is now asserted. It was claimed to be in three places and implemented in none; the only assertions on it were a membership test on site one and a negative one on
_rows, so a conversion at a seventeenth line would have joined the log and nothing would have failed. It is now compared as a multiset of(ATOM line, conversions)against a module constant,EXPECTED_HOST_RESOLUTIONS, the wayEXPECTED_GUARDSandSTAGED_COPIESalready were — the table above, asserted at TP1 and TP2 and at both step widths, where it is the same 20-over-16 multiset. A multiset rather than a set, so a line that converts twice where it converted once also fails.__bool__is untouched, so a branch on a width still installs its guard and a step whose shape decides which path ATOM takes still records that it did.Which route, and why
Neither of the brief's two.
CpuGpuBuffer. Built and measured, and not required. Site two's-> 16384is a buffer capacity (max_model_len // block_size) being solved against the CPU side's constant, and it exists only because that probe makes every dimension of the staged device tensor a free symbol. Capacities are engine configuration and are not a function of the step. With only the width symbolic,self.gpu[:n]andself.cpu[:n]carry the same symbol andcopy_to_gpudispatches unchanged. The alternative was built anyway: a fake, non-numpy-backed CPU side costs 38prim.devicecalls ATOM never makes and leaves every family's non-numeric shape-entry count unchanged — 5,016 either way, family for family. Round 2 corrects the wording of that row: "changes not one shape entry" was loose. The total entries rise 12,544 → 12,585, because those 38 extra calls carry entries of their own; what is unchanged is the non-numeric count, per family. And the row's code is not in the tree: it was a fake CPU side written into the capture's staged allocators, measured, and removed once it had answered, so re-checking it means rebuilding it. It is the row that decided the route and it is the weakest-sourced thing in this PR; stating that is the correction, not defending it.atom/utils/__init__.py— the brief's file set — is not touched.prepare_inputs. Rejected without building it: reaching the model that way means hand-constructingAttentionMetaDataandContext, which re-derives the shapes ATOM's own staging derives. The design record carries a measured account of what that costs — a prior re-derivation was wrong in both directions against the real classes. It also answers the wrong question: the reason to trace ATOM's staging is to find out what it does to a symbol.ForwardMode.decidesettles both units,prepare_inputswrites thecu_seqlens_qboundary atrunning_bs + 1,prepare_decodeuses it as every staged bound. Nothing re-derives a width. One symbol, because a decode step is one query row per sequence; two would have to be equated later, which is the specialisation again by a longer route — and that is demonstrated rather than argued, see site three above.Is the symbol free, or a hint in disguise?
This is the question the result rests on, and it is the load-bearing evidence — not
replacements == {}, which in this pass cannot fail for__index__and__int__because those two are the things that were replaced (see What is not done). ATOM computes its host values from the hint — the slot mapping, the block tables, the cumulative sequence lengths are all real numbers for one concrete width. A graph built that way would still be a graph about one width if any dimension had quietly taken the hint, and nothing in the census would say so. A symbol can also be lost rather than solved, by code reading its hint and building a tensor of that size, and that records no replacement at all.So the step is traced at two widths, 2 and 8, at TP1 and at TP2 (round 2 adds TP2; round 1 ran TP1 only, although TP2 is the named result), and the inventories compared operator for operator and shape for shape with the symbol's name set aside:
a7c1ec1b6ac1ec4b…a7c1ec1b6ac1ec4b…a37d76da81cd4547…a37d76da81cd4547…shape_envguardsIdentical except the last row, and round 2 states that last row rather than omitting it — see Guards below. Anything carrying the hint would be a 2 in one column and an 8 in the other. That also settles the other 7,528 entries, which is the second half of the census: every dimension that stayed a number is the same number at both widths, so none is a width in disguise. There are 23 distinct concrete values in the whole TP1 inventory and every one is the model's or the engine's — 5,120 (hidden) x2,959, 17,408 and 34,816 (MLP), 248,320 (vocab), 16,480 / 14,336 / 10,240 / 8,192 / 6,144 / 1,024 (projection widths), 256 (head dim), 128 and 48 (head counts), 512/513 (
max_num_seqs), 16,384/16,385 (max_model_len // block_size), 49,152, and 24/4/3/1/0. Neither 2 nor 8 appears.At TP2, 2 does appear — 73 times — and it is the group's width, not the step's. Round 1's assertion only ever ran at TP1 and so never met this. What separates the two readings is the row above:
concrete_dimsis identical at step widths 2 and 8, so a 2 that survives an eight-wide step is an engine constant in the way a head count is. The test now asserts the second width is absent at both TPs — the direction that can fail — and asserts the group width's presence at TP2 with the reason.Guards — the applicability, read rather than totalled, and hint-dependent
28*N <= 8192,N < 128,N <= 512N < 128,N <= 512N + 1 > <width>N + 1 > <width>The upper bounds are staging capacities read against the step's width; none is an equality, which is the same fact as
replacementsbeing empty. TP2 installs one fewer of them, and that is a finding rather than noise: the per-token bound goes, because the width halves each rank's share of the rows it reads.Round 2: the guard set is the one artifact of this capture that depends on the hint, and the round-1 table was the hint-2 case published as the graph's applicability. Re-measured at TP1 hints 2, 3, 8 and 16 and at TP2 hints 2 and 8: at every hint but 2 a lower bound
N + 1 > <hint>appears as well. At hint 2 it is elided because a size symbol's default range is[2, ∞), which makes the inequality vacuous. It is not a specialisation — it is a__bool__comparison installing an inequality on a live symbol, so it is the positive evidence that__bool__is untouched. What it costs is honesty about applicability: at any hint but 2 the traced graph carries a lower bound too.expected_guards(tp, width)composes the set per width, and the two-width test asserts the whole set at both widths and the lower bound's presence and absence, rather than leaving the guards out of the comparison.What runs differently in production
One line, and on the evidence, nothing.
ForwardMode.assert_shape_contract's_rowshelper.torch.Size.__getitem__returns a Pythonintfor any tensor with a real size, soint()around it is the identity in every served step — same value, same type, same assertion messages. It matters only where the size is symbolic, and there converting one side of an equality to a number forces the other side to become that number.The change is inside a
defthat exists only to feed fourasserts in that method, and the method is shape-only by contract ("Shapes only, never values: reading a device tensor's contents on this path is a D2H sync per step"). Nothing else reads_rows. This is site three, and the only one of the three not reachable from a capture-time substitution. Both gate tiers were run because it isatom/utils/;atom/utils/forward_context.pyis not ingpu_gate_triggers.txt— the CPU tier covers it — and the GPU tier was run anyway.The pin: updated, and three ways ablated
CAP-1's assertions were re-pointed, not removed or loosened, and the four tests added in round 1 are now five:
test_the_inventory_is_concrete_at_both_widths— unchanged; it is the control.test_the_first_two_specialisation_sites_are_where_they_were_measured— unchanged and still passing.test_closing_site_one_moves_the_bound_to_a_third_site— updated.SITE_THREEmoves one frame out and the test also asserts_rowsis not on the path.test_atom_s_own_buffer_constructor_is_what_runs— unchanged and passing; the step-symbol pass also assertssymbolic_device_allocations == 0andconstructed == 19.replacementsempty, no specialisations, the guard set by width); the census by family with its control; the three sites by what now happens at each, with the host-resolution log asserted against its declared set; the two-width digest at TP1 and TP2; and — round 2 —test_the_capture_refuses_a_width_that_torch_would_specialise, which runs the real entry point at--width 1and asserts a non-zero exit, the message, and that no record is printed.Ablated, on this head, all three re-run against the 14 tests:
int(t.shape[0])in_rowsguard_int)Gates
Both tiers, re-measured as deltas against the new base
24d75742d. Round 1's figures against9fcd6c7bddo not carry forward. Each tree staged withgit archive+docker cpinto a path of its own, never into the shared mount, gated with its ownscripts/compass/,COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, tarball md5 and per-file md5 checked on both ends,.compass-commitread back inside each container.CPU tier, in
xiaobizh_n18_cpu24d75742d70f8cd4dbGATE_CPU_RCThe +5 is exactly this file's five new tests — round 1's four plus round 2's width refusal — and the skip count is identical on both sides, so the tier's three-way flake (
tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk) did not fire in either direction: no ±1 in the passed/skipped split and no non-zeroGATE_CPU_RC. The gate's own stamp readsgpu: not required (.compass-changed stamp)on the branch —atom/utils/forward_context.pyis not ingpu_gate_triggers.txt, because the CPU tier covers it viatests/test_forward_mode.py. The GPU tier was run anyway, because this isatom/utils/.GPU tier, in
xiaobizh_n18, judged as a delta24d75742d70f8cd4dbGATE_GPU_RCPREFLIGHT_RCbefore / afterThe five failures are the same five, by name, on both sides and against
gpu_gate_known_failures.txt— bothdiffs of the sorted failing node-ids are empty, control-vs-branch and branch-vs-file:Four are one bf16 ULP (
max|diff| = 0.001953125 = 2^-9) and the fifth is the bitwisetorch.equal. Toolchain matched the baseline's exactly: torch 2.10.0+rocm7.2.4.git3d3aa833, hip 7.2.53211, aiter v0.1.21.dev0-49-gf4e7c7509. Run on device 3 — devices 0 and 1 were 83% resident with someone else's job, and the gate's own bookable check offered2 3 4 5 6 7.Both sides' pass counts are accounted for by the gate's own arithmetic rather than by a floor:
5,340 = 4,779 baseline + (610 − 49) tests/compassand5,345 = 4,779 + (615 − 49).Effort
Measured against
24d75742d, at head70f8cd4db. Two AST conventions are given because round 1's figure and the review's differed by what a counter does with docstrings; production is 0 net either way.atom/utils/forward_context.py)tests/compass/test_capture_real_model.py)Diff against
24d75742d:+7/−1production,+970/−45test,+102/−204and+1/−112.Correction, so the number in this body is the number at head. Three figures here were first published as measured at
2b2a48370, an earlier amend of this branch, and were stale by two amends: the test numstat read+965/−45, the design numstat+101/−3, and the test's physical non-blank+828 (1,307 → 2,135). The corrected figures are the ones in the table and the line above. The AST figures did not move and are as first published — production 0 net on both conventions, test+203/+189— so the only thing that changed is three physical-line counts.Comments were not trimmed and were not meant to be, and the split says so rather than leaving it to a ratio. Of the test file's +833 physical non-blank lines: comments +164, docstrings +285, everything else +384 — so 54% of what was added is explanation, which is what six of the eleven findings asked for. Of the remaining 384, only 189 are AST statements that are not docstrings (203 if the 14 new docstrings are counted as statements); the rest are continuation lines, most of them one-per-entry rows of the two declared tables,
OP_FAMILIESandEXPECTED_HOST_RESOLUTIONS. The 400-production-line halt threshold was never approached, because the production change is one line — which is itself the result.Reproducing
14 tests, 8 subprocess captures at ~13 s each plus one subprocess that is refused, ~108 s.
Design documents
CpuGpuBuffer" was measured not to be required and that the ablation's code is not in the tree, what site three turned out to be, that the guard set is hint-dependent, and what the symbol does not survive. It stays open for what this does not cover.04D18 gains a fourth capture discipline: the three existing ones keep the tracing symbolic and are not enough on a real engine, because the engine's host arithmetic asks the symbol for a number andguard_intrecords the ask. Round 2 also retitles### The three disciplinesto four, says in that section's lead that a reader who stops at the numbered list has three of four, and updates D18's decision-log row, which had not been touched.What is not done
ScheduledBatch.__init__. The constructor compares the staged token array's length against the count it was handed, which solves the width by comparison rather than conversion — the__bool__boundary this PR leaves alive on purpose. So the capture builds the batch at the concrete width and rebinds its four count fields afterwards. Round 1 disclosed this only in_decode_batch's docstring while the body said the symbol "enters at theScheduledBatch"; it enters ATOM's staging, not ATOM's batch constructor.replacements == {}proves less in this pass than it looks.__index__and__int__are replaced by_resolve_on_the_hostfor the capture's duration, so for those two routes the assertion is guaranteed by construction. What it still tests is every other route into_set_replacement— a comparison, a hash, a format string — and those are live. The load-bearing evidence is the two-width digest.distinct_ops, which is howlift_fresh→scalar_tensorwas found.max_seqlen_q > 1are untraced; the batch here hasmax_seqlen_q == 1.@triton.jitlaunches are recorded and not executed, so anything downstream of a skipped kernel read uninitialised fake memory. These counts are not a cost-model input at any width.c10d's legacy collectives routed to their functional forms, andATOM_USE_CUSTOM_ALL_GATHER=0, which is not ATOM's default path. Both are on the control as well, so the census compares like with like, but the operator lists carry them.24d75742dremoved both, with different wording, and the restack adopted the base's. This branch removes none; the identifier sweep over the changed files and overtests/compass/,atom/compass/**/*.pyandscripts/compass/is clean, and was watched firing on9fcd6c7bd, where it returns four hits.ruff format --check,ruff checkandblack --checkare clean on both changed files, except one pre-existingBLE001inforward_context.py:938that is present at24d75742dand unrelated to this change.🤖 Generated with Claude Code