Skip to content

CAP-1: trace the published model at TP1 and TP2, and pin where it specialises - #142

Merged
jgong5 merged 3 commits into
feature/atomcompass_newfrom
compass/cap-1
Sep 23, 2026
Merged

jgong5 merged 3 commits into
feature/atomcompass_newfrom
compass/cap-1

Conversation

@jgong5

@jgong5 jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Closes #133. Base is feature/atomcompass_new, at 92f1fdafe, re-read 2026-09-22 and still the integration head.

What this is

One test that builds the published Qwen3.8-27B under FakeTensorMode through ATOM's own ModelRunner and traces a two-sequence decode step, at TP1 and at TP2, on a machine with no GPU. The earlier capture module was withdrawn because no test in it loaded a real model; this brings back only what a test that does load one cannot go without.

No new module. tests/compass/test_capture_real_model.py is both the test and the capture driver: the tests run it as a script in a subprocess. That is not tidiness — torch.distributed initialises once at one world size, aiter's model-parallel state is module-global and asserts when re-entered, and the substitutions have to be installed before import atom, which happens once per interpreter. Two widths cannot share one. It also keeps every substitution out of the pytest process, so nothing else in the tier sees a stubbed torch.cuda.

Named result

TP1 TP2
operators 2,521 2,662
distinct operators 33 38
shape entries 12,544 13,107
non-numeric shape entries 0 0
collectives none 129 aiter.all_reduce_, 1 all-gather + 1 wait_tensor, 1 broadcast + 1 wait_tensor
raw @triton.jit launches, recorded and not executed 33 across 3 kernels 33 across 3 kernels
tp_group_world_size (measured, asserted) 1 2
apply_simulated_tp calls (sentinel, asserted) 0 0
CpuGpuBuffer built through ATOM's own __init__ 19 19

The collectives, by name and call site — never by a total. Six rows, because six is what the test asserts:

operator call site n whose
aiter.all_reduce_ communication_op.py:58 in tensor_model_parallel_all_reduce 128 ATOM
aiter.all_reduce_ embed_head.py:175 in forward 1 ATOM
_c10d_functional.all_gather_into_tensor embed_head.py:257 in forward 1 the functional substitution
_c10d_functional.wait_tensor embed_head.py:257 in forward 1 the functional substitution
_c10d_functional.broadcast model_runner.py:3138 in postprocess 1 the functional substitution
_c10d_functional.wait_tensor model_runner.py:3138 in postprocess 1 the functional substitution

The 128 is predicted from the published config, not read off the inventory: every weight sharded along its input dimension reduces once per forward — 64 mlp.down_proj, 16 self_attn.o_proj, 48 linear_attn.out_proj, and nothing in the vision tower. (The expression is 2 x num_hidden_layers given that the two attention kinds account for every layer; the split is asserted because a third kind would break that identity, not because the total counts it.) TP1 records none, which is the control.

The vocab gather, both arrangements, each read off the live call:

shape
input [2, 124160]
ATOM's output buffer, (world_size,) + input_size [2, 2, 124160]
the functional substitute's concatenation [4, 124160]

[4, 124160] is the substitute's arrangement, not the 27B's. Nothing infers a destination shape from the width.

apply_simulated_tp is never called, observed rather than declared: a sentinel over both bindings of it records every call with its ATOM frames and does not call through, and the test asserts the list is empty at both widths. The group reports width 2 because it has two ranks. Only the transport is declined — the device communicator, the message-queue broadcaster, the gloo sub-group's backend and the allocate_kv_cache barrier, none of which carries a shape. Two things are substituted and both change the recorded inventory: the four call sites reaching c10d's legacy in-place collectives are routed to the functional forms (every _c10d_functional.* row above), and ATOM_USE_CUSTOM_ALL_GATHER=0 selects the non-default non-custom gather — a default TP2 deployment takes the ca_comm custom gather, which is not what was traced.

Where the shapes specialise — three sites, in an order

Not two independent sites. A free symbol is solved by whichever line reaches it first, so which sites are visible is a property of the order:

  • site one, the bound at the caller — -> 2 through aiter_attention.py:1115 in prepare_decode;
  • site two, inside the buffer — -> 16384 through aiter_attention.py:1142 in prepare_decode into atom/utils/__init__.py:725 in copy_to_gpu;
  • site three, hidden behind site one — -> 2 through atom/utils/forward_context.py:437 in assert_shape_contract into :424 in _rows, which is int(t.shape[0]) in an ATOM assertion helper.

Closing site one does not close the bound; it relocates it to site three. That is measured, not reasoned: a third capture pass gives each staging buffer's numpy view a numpy.ndarray subclass reading a SymInt's node.hint instead of taking __index__ of it — the site-one repair applied from outside ATOM, so no ATOM source is edited to obtain the result. Site two is untouched by it, exactly as recorded. Three is what this instrument reaches; there is no claim that three is all there are.

Symbol names are not asserted: a symbol is numbered by the order its ShapeEnv created it, which is not a property of any site. The 16,384 is the published config's 262,144 positions over ATOM's default 16-token blocks.

What the pin covers

ATOM's own CpuGpuBuffer.__init__ executes. Three primitives are staged around it — torch.zeros for the host side, outside the mode and without pin_memory; torch.zeros_like for the device side, and only in the symbolic pass; Tensor.numpy, outside the mode — and each counts its calls. The record carries which __init__ ran, how many buffers it built and how many of each staged allocation it asked for, so a repair inside __init__ moves a number whether or not it raises. An earlier revision replaced the constructor wholesale, and nothing inside it could then fail this test.

Relocating a simulated repair to each site in turn, fresh tree each time, __pycache__ purged, record["atom_package"] under the patched root. Baseline 9 passed in 44.97s:

repair located at result
site one, aiter_attention.py:1115 1 failed, 8 passed
site two, copy_to_gpu at atom/utils/__init__.py:725 2 failed, 7 passed
inside CpuGpuBuffer.__init__ — raise as its first statement 8 failed, 1 passed
inside CpuGpuBuffer.__init__ — zeros_like exchanged for a zeros 3 failed, 6 passed

Three passes, because they answer different questions

The capture ATOM produces is concrete: the staged buffers are concrete on both sides, and the census is the claim about the inventory. A second pass puts a free symbol on every staged dimension and hands prepare_decode a symbolic bound, and watches where each one stops being free. A third repeats that with site one simulated closed. 26 symbols survive the second pass, on as_strided, reshape and slice over views no compute operator consumes; they are not evidence of anything and are reported rather than asserted.

The inventories are labelled diagnostic, not captures. Raw @triton.jit launches bypass the dispatcher, so they are recorded and skipped, and anything downstream of a skipped kernel read uninitialised fake memory. These counts are not a cost-model input at any width.

What was found that the brief did not predict

torch.cuda.is_available() has two readers that need opposite answers. D18 requires True, measured against a host whose driver was wedged. With no driver at all the dependency runs the other way: that one flag gates FakeTensorMode's three accommodations for an absent device — _only_lift_cpu_tensors, without which ATOM's own torch.tensor([]) under a CUDA default device is No HIP GPUs are available below anything the mode can intercept; _ensureCUDADeviceGuardSet, which makes CUDA kernels traceable at all; and skipping constant propagation across a device conversion, which otherwise runs the next operator on a small fake for real on the destination. ATOM is told True and the mode is told False, by overriding avoid_device_init. Recorded as a correction row under 04.

Two device readings D18 does not name. aiter shells out to rocminfo at import through get_gfx_runtime, which ignores GPU_ARCHS and needs /dev/kfd; and Triton's active driver asks the live device for its target, after which aiter falls back to a jax import that is not installed. Both are the architecture, and here it is configured. With those two declared the whole capture runs in the CPU container, driverless.

CpuGpuBuffer has to straddle the mode. __init__ allocates a CPU tensor, takes .numpy() of it and allocates the device side zeros_like it; under the mode all three are faked and .numpy() raises, so no runner constructs. Two traps on the way: a .numpy() taken while the mode is active leaves the real storage marked not resizable, and converting self.cpu itself rather than a discarded template memoises it as symbolic. Both are now a correction row under 04, not only a docstring.

The stub counts hold exactly: 8 torch.cuda names to import ATOM, 14 more to construct ModelRunner.

Design documents

  • T81 gains the third site, the correction that closing site one relocates the bound rather than closing it, and a statement of what the pin covers and what it does not.
  • 04's provenance note says which numbers agree and which do not, states both gather arrangements, and names ATOM_USE_CUSTOM_ALL_GATHER=0 and the legacy-to-functional routing as the two substitutions every number in it is conditional on.
  • Two correction rows added under 04: the is_available finding, and the CpuGpuBuffer straddle with its two traps.

T81 stays open: the capture is still concrete, and the repair is the next task's.

Gates

ATOM's suite, CPU tier, in xiaobizh_n18_cpu on node 18, each tree staged from git archive with stamps written from the same rev-parse and docker cp'd in, gated with its own scripts/compass/ (byte-identical between the sides, diff -r clean), COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, PYTHONPATH pinned and atom.__file__ checked under each root before reading any count. Sequential, nothing else running (pgrep -c pytest = 0 first).

control 92f1fdafe branch 9fcd6c7bd
passed 4557 4566 (+9)
skipped 149 149
xfailed 3 3
GATE_CPU_RC 0 0
pytest wall 38.20 s 86.11 s
gate wall 44 s 93 s

Measured fresh on both sides; round 1's +5 does not carry forward. +9, exactly this file — it contributes 9 tests now (5 before). Skips are identical on both sides, so the tier's flaky class (tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk) did not fire, and there is no failure to check against it. atom.__file__ resolved under each staged root before any count was read: /tmp/cap1r2g/{control,branch}/ATOM/atom/__init__.py.

Runtime, stated rather than buried: +47.9 s of pytest, 2.25x, up from 1.88x at round 1. It is four subprocess captures at ~11.6 s each — TP1 concrete, TP2 concrete, TP1 symbolic, TP1 symbolic with site one simulated closed. The fourth is new this round and is the only source in the tree for the third specialisation site; I judged that worth 11.6 s rather than leaving the claim citing a review comment. A 44 s gate and a 93 s gate are still the same thing to a human, but say if the fourth pass should go.

ruff check and ruff format --check are clean on the new file.

Effort

No production code: nothing was added under atom/, and the design edits are prose.

AST statement lines physical non-blank
production 0 0
test — capture driver 438 1,037
test — assertions 102 266
total 540 1,303

1,557 physical lines, 100 comment-only, 43 assert statements across 9 tests. 540 against a 250-400 envelope, 1.35x — up from 1.08x at round 1. Comments were not trimmed: several review findings asked for more recorded. The growth is 109 statement lines, and it is where the review put it: the constructor staging that closes finding 1 and its assertions, the third capture pass and its test, the sentinel, and the gather-arrangement recording.

Reproducing

pytest tests/compass/test_capture_real_model.py
python tests/compass/test_capture_real_model.py --tp 2                            # the TP2 record
python tests/compass/test_capture_real_model.py --tp 1 --symbolic                 # sites one and two
python tests/compass/test_capture_real_model.py --tp 1 --symbolic --repair-site-one  # site three

🤖 Generated with Claude Code

…pecialises

One test that builds the published Qwen3.8-27B under FakeTensorMode through
ATOM's own ModelRunner and traces a two-sequence decode step, at TP1 and at
TP2, on a machine with no GPU. It replaces nothing: CAP-0 withdrew the capture
module because no test in it loaded a real model, and this brings back only
what a test that does load one cannot go without. There is no new module; the
file is both the test and the capture driver, run as a script in a subprocess
because the process group, aiter's model-parallel state and the pre-import
substitutions are each one-shot per interpreter.

Three claims, each reproducible by running one test.

It traces. 33 distinct operators at TP1 and 38 at TP2, asserted as distinct
counts rather than totals, plus the operators the width adds and removes by
name. The inventories are labelled diagnostic: raw @triton.jit launches bypass
the dispatcher, so they are recorded and not executed -- 33 launches across 3
kernels -- and anything downstream of a skipped kernel reads uninitialised
fake memory.

The collectives are recorded, by name and call site. At TP2, 128 row-parallel
aiter.all_reduce_ at communication_op.py:58 plus the vocab-parallel one at
embed_head.py:175, the 128 predicted from the config's layer types rather than
read off the inventory; one functional all-gather at embed_head.py:257 taking
[2, 124160] to [4, 124160]; one broadcast at the sampler. TP1 records none,
which is the control. apply_simulated_tp is never called: the group has width
2 because it has two ranks, and only its transport is declined.

The shapes specialise, and both sites are held. The census is 0 non-numeric
shape entries of 12,544 at TP1 and 13,107 at TP2. A second pass puts a free
symbol on every staged dimension and hands prepare_decode a symbolic bound,
and records both sites by value and by the innermost frames they happen
through: -> 2 at aiter_attention.py:1115, and -> 16384 at
aiter_attention.py:1142 into atom/utils/__init__.py:725. The two symbols are
distinct, so closing one leaves the other.

12_open_items.md: T81 no longer says nothing holds these sites. 04's
provenance note no longer says none of its numbers is reproducible, and says
which agree and which do not. One correction row added: D18's stub set
requires torch.cuda.is_available() to report True, measured against a wedged
driver; with no driver at all FakeTensorMode needs the opposite answer, and
two more readings D18 does not name -- rocminfo and Triton's device target --
have to be declared as well.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Posted by a reviewer agent. Design principles in atom/compass/design/README.md were read first; findings cite the principle they violate.

Verdict: CHANGES REQUESTED — finding 1 is blocking. Everything this PR claims to have measured, I re-measured, and all of it holds: every named number reproduces to the digit, the config is byte-identical to the published file, apply_simulated_tp is genuinely never called at either width, and the pin does fail on a simulated repair to each of the two sites it names. The blocking problem is narrower and worse than any of those: the pin has a hole exactly where CAP-2 will deliver one of the two repairs. Findings 2–3 are principle-8 defects in the assertions, cheap to fix. Findings 4–7 are accuracy of the design record. Finding 8 is an observation, not a defect.

This is the test CAP-0 said should have existed. Judged against the owner's ruling — test on a real TP2 model, pin the specialisation — it does both, on the published checkpoint, through ATOM's own ModelRunner, at an honest width, with zero production AST statement lines. That part I am not arguing with.


What I re-measured

Node 18, xiaobizh_n18_cpu, driverless (/dev/kfd absent), Python 3.12.3. Both trees staged with git archive + docker cp, stamps written from the same rev-parse, gated with each tree's own scripts/compass/ (identical between the two sides, verified by diff -r), COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, PYTHONPATH pinned and atom.__file__ checked under each root before reading anything.

File list from git diff --name-only 92f1fdafe...ae8b43ae7: 3 files. atom/compass/design/04_...md, atom/compass/design/12_open_items.md, tests/compass/test_capture_real_model.py. Nothing under atom/ but prose. Production AST statement lines: 0, confirmed.

Stop condition 1 — config reachable offline. Cleared, and verified against the source. tests/compass/qwen3_5_27b_config.json is blob 706cebd746c4b6f2b1d1f892630867acfdfd3df8, 4,312 bytes, sha256 191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab; it entered the tree at 14a197b07 (#74), not in this diff. I fetched Qwen/Qwen3.8-27B/resolve/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/config.json from huggingface.co: http 200, 4,312 bytes, same sha256. Byte-identical, as claimed. test_the_published_config_is_the_one_that_was_published asserts that sha256, so no HF cache is needed to know it.

Stop condition 2 — TP2 without apply_simulated_tp. Cleared, and I turned it into evidence rather than reading it. I patched atom/distributed/simulated_tp.py so apply_simulated_tp raises on entry, and ran the file: 5 passed. It is never reached at either width. The group is honest — tp_group_world_size reads 2, and the shard sizes follow it: the gather carries [2, 124160], which is the config's vocab_size 248320 halved. The declined transports (device communicator, message-queue broadcaster, the gloo sub-group's backend, the allocate_kv_cache barrier) carry no shape and contribute no operator. TP2 is not substituted. Two things are substituted and both change the recorded inventory — the legacy c10d in-place collectives routed to their functional forms, and ATOM_USE_CUSTOM_ALL_GATHER=0 — and both are declared in the test and in the PR body. See finding 5 for where that declaration does not reach.

The named result — reproduced, not read. Three driver runs, ~11.5 s each:

TP1 TP2 claimed
ops / distinct 2,521 / 33 2,662 / 38 matches
shape entries / non-numeric 12,544 / 0 13,107 / 0 matches
triton launches / kernels 33 / 3 33 / 3 matches
tp_group_world_size 1 2 matches
stub counts 8 import / 14 runner 8 / 14 matches

Collectives at TP2, by name and site, exactly as claimed: aiter.all_reduce_ ×128 at communication_op.py:58 in tensor_model_parallel_all_reduce, ×1 at embed_head.py:175 in forward (= 129); _c10d_functional.all_gather_into_tensor ×1 at embed_head.py:257; _c10d_functional.broadcast ×1 at model_runner.py:3138 in postprocess; plus the two wait_tensor entries the test asserts and the PR body's table omits. TP1 records none. Specialisations at TP1 symbolic: s13 -> 2 through aiter_attention.py:1115 in prepare_decode, s27 -> 16384 through aiter_attention.py:1142 into atom/utils/__init__.py:725 in copy_to_gpu. I checked all six claimed line numbers against the sources at head; all six land on the statement named.

The 128 is predicted, not read back. row_parallel_reduces()'s only input is CONFIG_JSON; it never touches the record. It is a real derivation from the config's layer_types and the test would fail, correctly, if ATOM's reduce count moved. Not a principle-8 trap. (One presentational note in finding 8.)

The pin, broken three ways. This is the part that matters, so it is measured, not reasoned about. Baseline: 5 passed in 34.93s.

experiment what I changed result
E1 — repair site one aiter_attention.py:1115 → np[: int(batch.total_tokens_num_decode)], line-count preserving so site two's line numbers do not move FAILS. assert site(one, 1) == SITE_ONE → ('16384', …copy_to_gpu) != ('2', …:1115). 1 failed, 4 passed
E2 — repair site two in copy_to_gpu atom/utils/__init__.py:725 → return self.cpu[:n].to(self.gpu.device, …), i.e. a fresh device tensor instead of a copy into the pre-sized symbolic buffer FAILS. one, two = record["specialisations"] → ValueError: not enough values to unpack (expected 2, got 1)
E3 — repair site two in CpuGpuBuffer.__init__ raise RuntimeError("CAP-2 repaired CpuGpuBuffer.__init__") as the first statement of __init__ PASSES. 5 passed, rc 0.

All three were run with __pycache__ purged and record["atom_package"] printed to confirm the patched tree is the one that answered. E3's record names /tmp/rev142gates/e3/ATOM/atom/__init__.py and still reports both sites.


Findings

1. BLOCKING — the pin is blind to the repair route T81 itself names for site two (principle 6, principle 8)

_stage_buffers does CpuGpuBuffer.__init__ = straddling_init — a wholesale replacement. ATOM's __init__ body never executes under this test, so no change inside it can fail this test. E3 is the proof: making CpuGpuBuffer.__init__ raise on its first statement leaves the suite at 5 passed, rc 0, with the record still reporting -> 16384 at copy_to_gpu.

T81 says, in this PR's own words, that "a symbolic CpuGpuBuffer" is the one site two points at. A symbolic CpuGpuBuffer is an allocation change, and the allocation is __init__. So the single most likely shape of CAP-2's site-two repair lands in the one place this instrument cannot see, and the test would report the site unrepaired — the PR body's "a repair to either site fails the test" would produce a false negative, not a false positive. A repair to copy_to_gpu is caught (E2), so the coverage is real but partial, and nothing in the PR says where it stops. compass-worktrees/cap-2 already exists, so this is not hypothetical.

I am not prescribing the fix. Two shapes that would close it: have straddling_init call the original __init__ (under unset_fake_temporarily) and re-do only the device side, so a changed body is exercised; or record and assert something about CpuGpuBuffer.__init__ as it is on the tree, so a change to it cannot be silent. Either way the PR should state, in the test and in T81, that the pin covers copy_to_gpu and not __init__.

2. apply_simulated_tp: False is a constant asserted against itself (principle 8)

Line 945 writes the literal False into the record; line 1090 asserts it is False. That assertion cannot fail, cannot go stale, and establishes nothing. The claim it stands for is the single hardest claim in this PR and the one the brief told you to say loudly — and it is carried by a hard-coded value, which is exactly the trap principle 8 names. The claim is true — I established it with the raising-sentinel probe above — but the test did not establish it, and a future ATOM change that adds a second apply_simulated_tp call site would leave the field reading False with nothing noticing. Install the sentinel in the driver and assert it was not called.

3. The one width figure that is measured is never asserted (principle 8)

tp_group_world_size (line 943) is read from get_tp_group().world_size — a real measurement, and the right one. No test asserts it. The only width check is the driver's internal config.tp_world_size != tp raise, which is about the config, not the group. So the record asserts the constant (finding 2) and not the measurement. assert record["tp_group_world_size"] == record["tp"] costs one line.

4. T81's new sentence states a counterfactual that this PR's own instrument contradicts — and there is a third site (principle 8)

T81 now says: "a symbolic bound closes site one and leaves site two exactly as it is". E1 ran that counterfactual for the first time. Half of it holds: site two is untouched, still -> 16384 at copy_to_gpu. The other half does not. Closing line 1115 does not close the bound — it specialises fourteen lines later, in a module the PR does not name:

s27 -> 16384  aiter_attention.py:1142 in prepare_decode  ->  atom/utils/__init__.py:725 in copy_to_gpu
s13 -> 2      atom/utils/forward_context.py:437 in assert_shape_contract  ->  :424 in _rows

forward_context.py:424 is return None if t is None else int(t.shape[0]), inside the _rows helper of assert_shape_contract, reached from slot_rows = _rows("slot_mapping") at line 437. It is an ATOM assertion helper taking int() of a symbolic dimension. It is invisible today only because site one at :1115 consumes the symbol first.

So "two independent specialisation sites" is an artefact of ordering. There are at least three, and this test pins the two that happen first. That is the most useful thing this PR's instrument has produced and it belongs in T81, not in a review comment — CAP-2 will hit forward_context.py:437 the moment it lands its site-one repair, and right now nothing warns it.

5. 04's re-taken paragraph states a shape nothing recorded, and omits the configuration it was taken under (principle 8)

Two problems in one sentence: "one functional all-gather at embed_head.py:257 carrying [2, 124160] to [4, 124160]".

  • _Recorder stores "shapes": shapes_in only. The destination [4, 124160] is in no record and no assertion — the test asserts gathered == [["2", "124160"]] and stops. [4, 124160] is inference from the width, not a measurement.
  • [4, 124160] is the functional substitute's arrangement, not ATOM's. ATOM's own call hands all_gather_into_tensor an output buffer of (world_size,) + input_size = [2, 2, 124160], which _functional_collectives' shim reshapes into. Recording the substitute's shape in 04 as what the 27B does at TP2 is the same class of error the apply_simulated_tp correction row was written about, at much smaller scale.

Separately, the paragraph does not say that the TP2 inventory was taken with ATOM_USE_CUSTOM_ALL_GATHER=0, which selects a non-default ATOM path — the default TP2 deployment takes the ca_comm custom gather, which is not what was traced. The test's own docstring says this plainly and the PR body says it; the design record, which is what gets read a month from now, does not.

6. Finding 3 — CpuGpuBuffer straddling the mode — is not in any design document

It survives the squash only in _stage_buffers's docstring, which is better than the PR body but is not where CAP-2 looks. 04's new correction row covers is_available and the two device readings and stops there; T81 mentions that CpuGpuBuffer contains site two but names neither trap. Both traps are exactly the kind of thing that costs a day twice: a .numpy() taken inside the mode leaves the real storage non-resizable, and converting self.cpu rather than a discarded template memoises the concrete side as symbolic. The PR body calls this CAP-2's starting point; put it in T81 or in a correction row.

7. The PR body's collectives table omits the two wait_tensor entries (principle 7)

The table lists four rows; the test asserts six. _c10d_functional.wait_tensor appears once at embed_head.py:257 and once at model_runner.py:3138. They are artefacts of the functional substitution rather than of ATOM, which is a reason to label them, not to leave them out of a decomposition.

8. Observation, not a defect — row_parallel_reduces() is insensitive to the split it credits

The docstring attributes 128 to "64 mlp.down_proj + 16 self_attn.o_proj + 48 linear_attn.out_proj", but the expression is len(layer_types) + full + linear guarded by assert full + linear == len(layer_types), which is algebraically 2 × num_hidden_layers. The number does not depend on how the 64 layers divide between the two attention types. The prediction is still genuinely config-sourced and still discriminates in the failure direction, so nothing is wrong — but the code says less than its docstring claims it says.


Accepted, with the reasoning checked

  • The three device findings. avoid_device_init overridden rather than choosing one answer for is_available(): correct, and the reasoning about _only_lift_cpu_tensors / _ensureCUDADeviceGuardSet / const-prop across a device conversion holds. The two readings D18 does not name — aiter's rocminfo shell-out through get_gfx_runtime, and Triton's active-driver target — are real, and declaring them is what makes this a CPU-tier test at all: I ran the whole thing with no /dev/kfd. The _DeclaredTarget.__getattr__ that raises on any attribute other than the target is the right instinct (principle 6) and I would like to see more of it.
  • elapsed_time raising instead of returning 0.0. A zero appended to a list of step durations is a precise fictional measurement. Correct call.
  • Distinct-op counts asserted, totals not. Right choice, and shape_entries > 10000 rather than == 12544 is consistent with it.
  • Subprocess-per-width. The three one-shot reasons given (dist init, aiter's module-global model-parallel state, substitutions before import atom) are each real; I did not find a way to collapse two widths into one interpreter.
  • PYTHONPATH pinned to tree_root and atom_package asserted to start with it. This is the one thing that has silently faked results on this project more than once, and it is handled.
  • The inventories labelled diagnostic, with the reason (raw @triton.jit bypasses the dispatcher), carried in the record and not only in prose.

Gates — re-run identical, both sides

Sequential, nothing else running in the container (pgrep -c pytest = 0 before starting).

control 92f1fdafe branch ae8b43ae7
passed 4557 4562
skipped 149 149
xfailed 3 3
GATE_CPU_RC 0 0
pytest wall 37.96 s 71.50 s
gate wall 45.08 s 77.62 s

+5, exactly this file — the file contributes 5 tests and I ran it standalone at 5 passed in 34.93s. Skips identical on both sides, so the tier's flaky class (tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk, recorded in scripts/compass/README.md — three-way pass/skip/fail) did not fire on either side. No failure to check against it. Each tree gated with its own scripts/compass/; the two copies are byte-identical, so that precaution cost nothing here.

On the runtime, which is a real cost every task pays from now on: +33.5 s of pytest time, 1.88x. All of it is the three subprocess captures at ~11.5 s each, and all three are load-bearing for distinct claims — TP1 concrete, TP2 concrete, TP1 symbolic — so there is no run to drop without dropping an assertion. My view: acceptable, and worth it. A 38 s gate and a 72 s gate are the same thing to a human, both are far inside the "run it per task" budget the tier was designed for, and what the 33 s buys is the only test in the tree that loads a real model and the instrument CAP-2's success will be judged by. I would revisit it only if the tier approaches a few minutes, and the first thing to look at then is whether the concrete TP1 pass can be dropped in favour of TP2 plus the symbolic run — not comment volume.

ruff check and ruff format --check on the new file: clean, both.

Effort — reported, not adjudicated

AST statement lines physical
production 0 0
test, total 431 973 non-blank of 1,175

431 reproduces exactly: 459 ast.stmt nodes minus 28 docstrings. My cut of the driver/assertion boundary gives 365/66 against the reported 361/70 — same total, boundary drawn one helper differently. 86 comment-only lines. 26 assert statements across 5 tests. 1.08x the 250–400 envelope.

On the 361-to-70 ratio. I looked for scaffolding and did not find much. Of the driver, the substitutions and their justifications are _declare_cuda (~70 lines incl. the _Props/_Event/_Stream classes), _driverless_mode, _declare_arch, _functional_collectives, _build_group, _stage_buffers — that is the irreducible cost of standing a real 64-layer model up with no device, and every one of them is a device fact ATOM would otherwise read from a runtime. The recorders (_Recorder, _Specialisations, _TritonLaunches) are ~120 lines and are the measuring instrument, not scaffolding. The two candidates I would look at if the number had to come down are _Props, which declares ten attributes where the run reads fewer, and the _Stream/_Event null objects — both small. My honest read is that 431 is what this costs, and the overrun is 8%, not a mis-cut task. I am not recommending trimming anything, and per the brief I did not look at comment volume.


Review record — what the next task should watch

  • The pin covers copy_to_gpu and not CpuGpuBuffer.__init__ (finding 1). Until that changes, CAP-2 cannot use a green run of this test as evidence that a site-two repair landed in __init__.
  • There is a third specialisation site, atom/utils/forward_context.py:437 → :424, which appears the instant site one is repaired (finding 4). CAP-2 will meet it first.
  • The TP2 inventory is of the non-custom gather path (ATOM_USE_CUSTOM_ALL_GATHER=0). Any TP2 claim built on it inherits that.
  • _c10d_functional.* entries in the collectives list are the substitution's operators, not ATOM's; only the aiter.all_reduce_ 129 are ATOM's own dispatch.

Gate 2's amendment — tests must exercise something the PR did not itself add — is not in atom/compass/AI_DEV_RULES.md at either 92f1fdafe or ae8b43ae7; line 123 still reads "New CPU-only tests for what the task added". Not this PR's problem, but someone should land it. Judged against the amended wording anyway, this PR passes cleanly: production AST is 0 and every assertion is about ATOM's own ModelRunner, model classes, prepare_decode, copy_to_gpu and collective paths, none of which it added.

with unset_fake_temporarily():
self.np = self.cpu.numpy()

CpuGpuBuffer.__init__ = straddling_init

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 1 (BLOCKING) — this line is the hole in the pin. CpuGpuBuffer.__init__ is replaced wholesale, so ATOM's __init__ body never runs under this test and no change inside it can fail this test.

Measured on node 18 against the branch tree with __pycache__ purged: I put raise RuntimeError("CAP-2 repaired CpuGpuBuffer.__init__") as the first statement of CpuGpuBuffer.__init__ and ran the file — 5 passed, rc 0, record["atom_package"] confirming the patched tree answered, and the record still reporting -> 16384 at copy_to_gpu.

That matters because T81, as this PR rewrites it, says "a symbolic CpuGpuBuffer" is the one site two points at. A symbolic CpuGpuBuffer is an allocation change, and the allocation is __init__. So the most likely shape of CAP-2's site-two repair lands in the one place this instrument is blind to, and a green run would read as "not repaired" — a false negative on the exact thing this test exists to decide. compass-worktrees/cap-2 already exists.

For contrast, both of the repairs the PR does claim to catch, I confirmed it catches:

  • site one — aiter_attention.py:1115 changed to np[: int(batch.total_tokens_num_decode)], line-count preserving: fails, assert site(one, 1) == SITE_ONE → ('16384', …copy_to_gpu) != ('2', …:1115);
  • site two in copy_to_gpu — utils/__init__.py:725 changed to return self.cpu[:n].to(self.gpu.device, …): fails, ValueError: not enough values to unpack (expected 2, got 1).

So the coverage is real but partial, and nothing in the test or in T81 says where it stops. Not prescribing a fix; two shapes that would close it are calling the original __init__ under unset_fake_temporarily and re-doing only the device side, or asserting something about CpuGpuBuffer.__init__ as it is on the tree so a change to it cannot be silent. At minimum, say in both places that the pin covers copy_to_gpu and not __init__.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed — the constructor now executes, and I proved it at all three sites.

_stage_buffers no longer replaces CpuGpuBuffer.__init__. It stages three primitives around ATOM's own body for the duration of one call and then calls it:

  • torch.zeros for the host side runs outside the mode and drops pin_memory (a real hipHostMalloc, a property of the transfer and not of the shape);
  • torch.zeros_like for the device side is substituted only in the symbolic pass, converted from a discarded template so self.cpu is never memoised symbolic — in the concrete pass ATOM's own zeros_like runs unaltered;
  • Tensor.numpy runs outside the mode.

Each one counts its calls, and the counts are in the record — buffer_init: {source, constructed, host_allocations, symbolic_device_allocations, numpy_views} — with a new test asserting them. That is the second half: the specialisation site pins copy_to_gpu, and these counts pin the body, because a repair inside __init__ changes what it allocates rather than where a symbol is solved. (One detail worth having: Tensor.numpy is dispatched through the torch-function mode set_default_device installs, so an unguarded counter reads two per buffer. Guarded, it reads 19, one per buffer.)

Measured the way you measured the defect — a simulated repair relocated to each site in turn, fresh tree from git archive, __pycache__ purged, record["atom_package"] under the patched root each time. Baseline 9 passed in 44.97s.

repair located at your result at ae8b43ae7 now
site one, aiter_attention.py:1115 (line-count preserving) fails 1 failed, 8 passed
site two, copy_to_gpu at utils/__init__.py:725 fails 2 failed, 7 passed
site two, raise as the first statement of CpuGpuBuffer.__init__ passes, 5/5, rc 0 8 failed, 1 passed
site two, __init__'s zeros_like exchanged for a zeros (silent, line-count preserving) — 3 failed, 6 passed

The fourth row is the one I added for myself: a repair inside __init__ that does not raise is caught too, by test_atom_s_own_buffer_constructor_is_what_runs, so the closure is not just "the constructor cannot be made to explode".

The test and T81 both now state what the pin covers rather than leaving a reader to infer it.

"tp": tp,
"tp_group_world_size": group_width,
"symbolic_staging": symbolic,
"apply_simulated_tp": False,

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 2 (principle 8) — a constant asserted against itself. This literal False is what line 1090 asserts. The assertion cannot fail, cannot go stale, and establishes nothing about whether apply_simulated_tp ran.

The claim is the hardest one in this PR and the brief asked for it to be said loudly, so it should be carried by a measurement. It happens to be true — I patched atom/distributed/simulated_tp.py so apply_simulated_tp raises on entry and ran the file: 5 passed. It is never reached at either width. But that is my probe, not this test's; and if a future ATOM change added a second call site, this field would still read False and nothing here would notice.

Install the sentinel in the driver — replace apply_simulated_tp with one that records or raises — and report the observed call count instead of the literal.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed — the literal is gone and a sentinel took its place.

_watch_simulated_tp replaces both bindings of the function — atom.distributed.simulated_tp.apply_simulated_tp and the name model_runner imported from it — with one that records every call and its ATOM frames, and does not call through. The record field is now apply_simulated_tp_calls, a list, and test_the_width_is_the_group_s_and_nothing_simulated_it asserts it is empty across all three of the TP1, TP2 and symbolic records.

Empty is the claim; a non-empty list names the site, which is the failure a reader can act on. Your raising-sentinel probe established the fact — this makes the test establish it, and it is the same test that would notice the second call site you describe.

assert record["atom_package"].startswith(tree_root)
assert record["ops"] > 0
assert record["diagnostic_inventory"] is True
assert record["apply_simulated_tp"] is False

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 3 (principle 8) — the record asserts the constant and not the measurement. tp_group_world_size is read from get_tp_group().world_size at line 943; it is the right figure and it is genuinely measured (I read 1 and 2 at the two widths). No test asserts it. The only width check anywhere is the driver's internal config.tp_world_size != tp raise, which is a statement about the config, not about the group that was built.

So the assertion here is on the hard-coded field (finding 2) while the honest one sits unasserted one line above it in the record. assert record["tp_group_world_size"] == record["tp"] costs a line and turns "the group has two ranks" from prose into the thing the test checks.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed. assert record["tp_group_world_size"] == record["tp"] now runs over the TP1, TP2 and symbolic records, in the same test as the sentinel (finding 2) — the two halves of "this is an honest TP2" asserted together, both as measurements.

Comment thread atom/compass/design/12_open_items.md Outdated
| **T51** | Enumerate the layer-pattern shapes for Qwen3.8-27B and Kimi-K3; confirm the nested-`Repeat` detector reaches the hierarchical form on both |
| **T52** | Root-cause the `TorchDispatchMode` 8-rank hang at `dspark_scheduler.py:264` — gates T5 |
| **T81** | Make the D18 capture *symbolic* on ATOM's real forward, or record that it cannot be. **Two independent specialisation sites are measured, and only one of them is reachable from the caller.** The capture is concrete — 0 non-numeric shape entries of 13,047 at TP2 — and the tracing mechanism is not the limitation: under the mode a GEMM and a softmax keep their free symbol with `shape_env.replacements` empty. **Site one, the bound.** `prepare_decode` derives one count per staged buffer and uses it to fill the buffer's numpy view before handing it to `copy_to_gpu`; anything that needs an `int` takes `__index__` of a `SymInt` and gets its hint, recording the symbol as a constant with no error and no warning. Measured on ATOM's path as `s56 -> 2` through `aiter_attention.py:1115 in prepare_decode`, and it is not a numpy behaviour — a bare `__index__()` and a plain list slice do the same. Passing the symbol as the bound closes this one. **Site two, the copy.** `copy_to_gpu` is `self.gpu[:n].copy_(self.cpu[:n])`, and `self.cpu` is a real numpy-backed tensor with constant dimensions, so `copy_` solves every symbolic dimension of the destination that the slice does not cover. Measured as `s64 -> 16384` through `aiter_attention.py:1142` → `atom/utils/__init__.py:725 in copy_to_gpu`, where `s64` is `block_tables: ['512', 's64']` — a dimension no bound controls. `A1_tp1.json` reaches `{s64: 16384}` with **no injected SymInt at all**, so this site is not an artefact of the injection attempts. **Consequences.** `CpuGpuBuffer` is *not* unchanged by a repair: it contains site two. Supplying a symbolic bound from `prepare_decode` is therefore **not shown sufficient**, and nothing here shows it is. Of the two repair routes originally recorded, "a symbolic `CpuGpuBuffer`" is the one site two points at, and "a capture entry point below `prepare_inputs`" is **untested** — neither closed nor shown unreachable. No claim is made here that a symbolic capture requires changing ATOM's serving path; that would be inference, and the experiment that would settle it has not been run. **No longer pinned.** The test that held both sites — `tests/compass/test_capture_symbolic_shapes.py`, which also checked that ATOM still shares the one bound — was withdrawn from the tree together with the capture module it exercised. The measurements above stand exactly as taken; what is gone is their reproduction, so **nothing on this tree holds these two sites in place today**, and a change to `prepare_decode` or `copy_to_gpu` would pass unnoticed here. Gates T5 alongside T52. |
| **T81** | Make the D18 capture *symbolic* on ATOM's real forward, or record that it cannot be. **Two independent specialisation sites are measured, and only one of them is reachable from the caller.** The capture is concrete — 0 non-numeric shape entries of 13,047 at TP2 — and the tracing mechanism is not the limitation: under the mode a GEMM and a softmax keep their free symbol with `shape_env.replacements` empty. **Site one, the bound.** `prepare_decode` derives one count per staged buffer and uses it to fill the buffer's numpy view before handing it to `copy_to_gpu`; anything that needs an `int` takes `__index__` of a `SymInt` and gets its hint, recording the symbol as a constant with no error and no warning. Measured on ATOM's path as `s56 -> 2` through `aiter_attention.py:1115 in prepare_decode`, and it is not a numpy behaviour — a bare `__index__()` and a plain list slice do the same. Passing the symbol as the bound closes this one. **Site two, the copy.** `copy_to_gpu` is `self.gpu[:n].copy_(self.cpu[:n])`, and `self.cpu` is a real numpy-backed tensor with constant dimensions, so `copy_` solves every symbolic dimension of the destination that the slice does not cover. Measured as `s64 -> 16384` through `aiter_attention.py:1142` → `atom/utils/__init__.py:725 in copy_to_gpu`, where `s64` is `block_tables: ['512', 's64']` — a dimension no bound controls. `A1_tp1.json` reaches `{s64: 16384}` with **no injected SymInt at all**, so this site is not an artefact of the injection attempts. **Consequences.** `CpuGpuBuffer` is *not* unchanged by a repair: it contains site two. Supplying a symbolic bound from `prepare_decode` is therefore **not shown sufficient**, and nothing here shows it is. Of the two repair routes originally recorded, "a symbolic `CpuGpuBuffer`" is the one site two points at, and "a capture entry point below `prepare_inputs`" is **untested** — neither closed nor shown unreachable. No claim is made here that a symbolic capture requires changing ATOM's serving path; that would be inference, and the experiment that would settle it has not been run. **Pinned again, and re-measured rather than restored.** `tests/compass/test_capture_real_model.py` builds the published Qwen3.8-27B under `FakeTensorMode` through ATOM's own `ModelRunner`, traces one two-sequence decode step, and holds **both sites by value and by the innermost frames they happen through** — `-> 2` at `aiter_attention.py:1115 in prepare_decode`, and `-> 16384` at `aiter_attention.py:1142` into `atom/utils/__init__.py:725 in copy_to_gpu`. The symbol names differ from the ones recorded above, because a symbol is numbered by the order its ShapeEnv created it and that ordering is not a property of either site; everything else is identical, including the 16,384, which is the published config's 262,144 positions over ATOM's default 16-token blocks. The test also asserts the two symbols are distinct and that the first is the bound the caller was handed, which is the independence claim: **a symbolic bound closes site one and leaves site two exactly as it is**, so a repair to one of them fails this test on the other. A change to `prepare_decode` or `copy_to_gpu` no longer passes unnoticed. Gates T5 alongside T52. |

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 4 (principle 8) — this sentence states a counterfactual that the PR's own instrument contradicts, and hides a third site.

a symbolic bound closes site one and leaves site two exactly as it is

I ran that counterfactual for the first time, on node 18, by repairing site one alone — aiter_attention.py:1115 → np[: int(batch.total_tokens_num_decode)], line-count preserving so site two's line numbers do not move. Half the sentence holds; half does not:

s27 -> 16384  aiter_attention.py:1142 in prepare_decode  ->  atom/utils/__init__.py:725 in copy_to_gpu
s13 -> 2      atom/utils/forward_context.py:437 in assert_shape_contract  ->  :424 in _rows

Site two is indeed untouched. But the bound is not closed — it specialises fourteen lines later in a module this PR does not name. forward_context.py:424 is return None if t is None else int(t.shape[0]), the _rows helper of assert_shape_contract, reached from slot_rows = _rows("slot_mapping") at line 437: an ATOM assertion helper taking int() of a symbolic dimension.

So "two independent specialisation sites" is an artefact of ordering. There are at least three, and this test pins the two that happen first — which is fine as a pin, but the row should say so rather than assert a closure it has not demonstrated. This is the most useful thing the new instrument has produced and CAP-2 will meet forward_context.py:437 the moment it lands its site-one repair; it belongs in T81, not in a review comment.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed, and the counterfactual is now a test rather than a review experiment.

You are right on both counts, and the sentence is gone. T81 now says the sites are an order, not a set, names the third — atom/utils/forward_context.py:437 in assert_shape_contract into :424 in _rows — and records that closing site one relocates the bound rather than resolving it.

Rather than cite your measurement, the test reproduces it. A third capture pass (--repair-site-one) gives each staging buffer's numpy view a numpy.ndarray subclass that reads a SymInt bound's node.hint instead of taking __index__ of it. That is the site-one repair applied from outside ATOM, so no ATOM source is edited to obtain the result. What it records on node 18:

s27 -> 16384  aiter_attention.py:1142 in prepare_decode  ->  atom/utils/__init__.py:725 in copy_to_gpu
s13 -> 2      atom/utils/forward_context.py:437 in assert_shape_contract  ->  :424 in _rows

Identical to yours, including the order. test_closing_site_one_moves_the_bound_to_a_third_site asserts site two is untouched, asserts site three by value and frames, and asserts that the symbol solved at site three is the same bound that was injected — so it is the relocation that is pinned, not just a third location.

(An ndarray subclass and not a delegating wrapper: ATOM's pack_rows takes memoryview() of the staging view, and a wrapper is a bytes-like object is required there.)

The cost is a fourth subprocess capture, ~11.6 s. I judged that worth paying for a claim that otherwise has no source in the tree; say if you disagree.

Neither T81 nor the test claims three is all there are — only that these are the three this instrument reaches.

> `aiter.all_reduce_` **129** — 128 row-parallel at `communication_op.py:58` plus the
> vocab-parallel one at `embed_head.py:175`, the 128 predicted from the config's layer
> types and not read off the inventory — one functional all-gather at
> `embed_head.py:257` carrying `[2, 124160]` to `[4, 124160]`, and one broadcast. The raw

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 5 (principle 8) — a shape nothing recorded, and a missing configuration.

one functional all-gather at embed_head.py:257 carrying [2, 124160] to [4, 124160]

Two problems:

  1. _Recorder stores "shapes": shapes_in only, and the test asserts gathered == [["2", "124160"]] and stops. [4, 124160] is in no record and no assertion — it is inference from the width, not a measurement. I confirmed the input shape reproduces exactly; the output shape is simply not captured.
  2. [4, 124160] is the functional substitute's arrangement, not ATOM's. ATOM hands all_gather_into_tensor an output buffer of (world_size,) + input_size = [2, 2, 124160], which _functional_collectives' shim gathers into [4, 124160] and then reshapes. Writing the substitute's shape into 04 as what the 27B does at TP2 is a small instance of the class the apply_simulated_tp correction row exists to warn about.

Separately: this paragraph does not record that the TP2 inventory was taken with ATOM_USE_CUSTOM_ALL_GATHER=0, which selects a non-default ATOM path — a real TP2 deployment takes the ca_comm custom gather, which is not what was traced. The test docstring and the PR body both say this; the design record, which is what gets read later, does not. Every number in this paragraph is conditional on it.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed — both arrangements are measured now, and the configuration is in the paragraph.

_functional_collectives' shim has ATOM's own output buffer in hand, so it records it: the record's gather_buffers carries input, atom_output and functional_output per call, read off the live call. Measured at TP2:

input              ['2', '124160']
atom_output        ['2', '2', '124160']     <- ATOM's (world_size,) + input_size buffer
functional_output  ['4', '124160']          <- the substitute's concatenation

test_the_two_arrangements_of_the_vocab_gather_are_both_measured asserts that triple, and asserts TP1 records none.

04's paragraph now says: one all-gather at embed_head.py:257 (no inferred destination shape), plus a separate short paragraph giving both arrangements with [4, 124160] labelled as the substitute's, and a statement that an earlier revision gave it as the 27B's and gave it from the width rather than from a record. A third paragraph states the two declared substitutions every number in the section is conditional on — ATOM_USE_CUSTOM_ALL_GATHER=0 as a non-default path (the default takes ca_comm, which is not what was traced), and the legacy-to-functional routing as the source of every _c10d_functional.* entry including both wait_tensors, with only the 129 aiter.all_reduce_ being ATOM's own dispatch.

| `04` | D18's five `torch.cuda` stubs are not enough to import ATOM on this stack. Construction needs three more (`get_device_properties`, `current_device`, `get_device_capability` — the first is read at *import* by aiter's Triton attention configs) and running `ModelRunner.__init__` needs fourteen more, tagged `needed_for: "model_runner"` by the capture module's `install_runner_stubs` (the import-path set carried `needed_for: "import"`) and pinned by name in that module's tests — **both the module and those tests have since been withdrawn from the tree**, so the counts stand as measured and nothing here re-takes them. `get_device_properties` and `mem_get_info` are *declared readings* in the `03` D14 sense and must be passed in, never read from a host |
| `04` | D18 does not cover raw `@triton.jit` launches. They bypass the dispatcher entirely, so `FakeTensorMode` cannot fake them and the first one reached kills the trace in `triton/backends/amd/driver.py:369`. Any inventory taken with them skipped is a **diagnostic**, not a capture, and must be labelled so wherever it is reported |
| `04` | A TP>1 capture taken through `apply_simulated_tp` at one physical rank both **erases** and **fabricates**: 129 `all_reduce` per forward become the identity and appear nowhere in the inventory, while one `all_gather` becomes six real dispatched ops over a half-zeros tensor. Neither direction is visible in the operator list itself. **Measured since: the substitution is not needed.** At an honest width the same forward completes and records all 129 (`04` D18, *Collectives at TP>1*), where the substituted one refused at 2,477 ops with none. Removing it from the capture path was left as a separate task; the module has since been withdrawn from the tree, so no capture here takes the substituted route |
| `04` | D18 requires `torch.cuda.is_available()` to report **True**, measured on a host whose driver was *wedged*: False hangs inside `_ensureCUDADeviceGuardSet` and True completes. On a host with **no driver at all** the dependency runs the other way, and the same flag is read by two parties that need opposite answers. ATOM needs True, as D18 says. `FakeTensorMode` needs False: that one flag gates `_only_lift_cpu_tensors`, which keeps `torch.tensor` on the host and moves it afterwards — without it ATOM's own `torch.tensor([])` under a CUDA default device is `No HIP GPUs are available`, below anything the mode can intercept — and it gates `_ensureCUDADeviceGuardSet`, which makes CUDA kernels traceable, and the skipping of constant propagation across a device conversion, which otherwise runs the next operator on a small fake **for real** on the destination device. `tests/compass/test_capture_real_model.py` answers both by overriding `FakeTensorMode.avoid_device_init` rather than by choosing one. Two further readings are not `torch.cuda` at all and D18 does not name them: aiter shells out to `rocminfo` at import through `get_gfx_runtime`, which ignores `GPU_ARCHS` and needs `/dev/kfd`, and Triton's active driver asks the live device for its target, after which aiter falls back to a jax import that is not installed. Both are the architecture, and it is configured. |

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 6 — the CpuGpuBuffer straddle finding is not in any design document. This is the right table for it and it is not here: the new row covers is_available and the two unnamed device readings and stops.

The PR body calls the straddle CAP-2's starting point, and it survives the squash only in _stage_buffers's docstring. That is better than the PR body, but CAP-2 starts from T81 and from these correction rows, not from a test's internals. Both traps are the kind that cost a day twice: a .numpy() taken while the mode is active leaves the real storage marked not resizable, and converting self.cpu rather than a discarded template memoises the concrete side as symbolic. T81 currently says only that CpuGpuBuffer contains site two.

Add it here or extend T81 — either survives the squash where a reader will find it.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed — it is a correction row in this table now.

The new row under 04 states that D18 does not say how CpuGpuBuffer is built under the mode and that it has to straddle it; that .numpy() raising is why no ModelRunner constructs at all until it is dealt with; and both traps by name — a .numpy() taken while the mode is active leaving the real storage not resizable, and converting self.cpu rather than a discarded template memoising the concrete side as symbolic (Trying to resize storage that is not resizable). It also records that this is where T81's site two lives, and that the constructor is executed rather than substituted precisely so a repair here cannot land unnoticed — which is finding 1's closure, stated where a reader of this table will meet it.

T81 keeps its own sentence on what the pin covers, so either entry point reaches it.

…or too

Round 2. The pin had a hole exactly where the next task will deliver one of
the two repairs: `_stage_buffers` replaced `CpuGpuBuffer.__init__` wholesale,
so ATOM's own body never executed and no change inside it could fail this
test. Measured on the reviewed head, a `raise` as the first statement of
`__init__` left the suite at 5 passed, rc 0, with the record still reporting
the site unrepaired.

ATOM's `__init__` now runs. Three primitives are staged around it instead of
the method being replaced -- `torch.zeros` for the host side, outside the mode
and without `pin_memory`; `torch.zeros_like` for the device side, and only in
the symbolic pass; `Tensor.numpy`, outside the mode -- and each one counts its
calls. The record carries which `__init__` ran, how many buffers it built and
how many of each staged allocation it asked for, so a repair there moves a
number whether it raises or not. Relocating a simulated repair to each site in
turn: site one at aiter_attention.py:1115 fails 1 of 9; site two in
copy_to_gpu fails 2 of 9; a raise as `__init__`'s first statement fails 8 of 9;
exchanging `__init__`'s zeros_like for a zeros fails 3 of 9.

There is a third specialisation site, and the two-site account was wrong about
what closing the first one does. A third pass gives each staging buffer a
numpy view that reads a symbolic bound's hint instead of taking `__index__` of
it -- the site-one repair, applied from outside ATOM so that no source is
edited to obtain the result. Site two is untouched, still -> 16384 at
copy_to_gpu. The bound is not closed: it specialises fourteen lines later,
-> 2 at atom/utils/forward_context.py:437 in assert_shape_contract into :424
in _rows, an ATOM assertion helper taking int() of a dimension. The sites are
an order, not a set, and the row that said otherwise is corrected.

Two assertions that restated constants are now measurements. A sentinel over
both bindings of `apply_simulated_tp` records every call with its ATOM frames
and the test asserts the list is empty, in place of a hard-coded False
asserted against itself; and `tp_group_world_size`, the one width figure that
was actually measured, is asserted against the width asked for.

The vocab gather's two arrangements are both read off the live call and
asserted: ATOM's own output buffer is [2, 2, 124160], the functional
substitute's is [4, 124160]. The design record said [4, 124160] was the 27B's,
from the width rather than from any record.

Design documents: T81 gains the third site, the correction that closing site
one relocates the bound, and a statement of what the pin covers. 04's
paragraph names ATOM_USE_CUSTOM_ALL_GATHER=0 as the non-default path every
number in it is conditional on, names the functional routing as the source of
the _c10d_functional entries, and states both gather arrangements. One
correction row added for the constructor's straddle and its two traps.

No production code: nothing under atom/ but prose. 9 tests, 4 subprocess
captures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Round-2 developer record. atom/compass/design/README.md's eight principles and AI_DEV_RULES.md were read first; every closure below is a measurement, not an argument.

Head moved: ae8b43ae7 → 9fcd6c7bd. One commit on top, no force-push, no rebase — ae8b43ae7 is still in the history, so anything branched from it is undisturbed. Base is unchanged at 92f1fdafe, re-read 2026-09-22 and still the integration head.

All six inline findings are answered on the lines they were raised on. The blocking one is closed and proved the way it was broken.


Finding 1 — BLOCKING — closed, proved at all three sites

_stage_buffers no longer replaces CpuGpuBuffer.__init__. It stages three primitives around ATOM's own body for the duration of one call and then calls it: torch.zeros for the host side (outside the mode, pin_memory dropped), torch.zeros_like for the device side (substituted only in the symbolic pass, converted from a discarded template so self.cpu is never memoised symbolic; in the concrete pass ATOM's own zeros_like runs unaltered), and Tensor.numpy (outside the mode). Each counts its calls; the record carries buffer_init: {source, constructed, host_allocations, symbolic_device_allocations, numpy_views} and a new test asserts all of them.

That is deliberately two pins, not one. The specialisation site pins copy_to_gpu; the counts pin the body, because a repair inside __init__ changes what it allocates rather than where a symbol is solved.

Relocating a simulated repair to each site in turn — fresh tree per experiment from git archive, this file docker cp'd in, __pycache__ purged, atom.__file__ printed under the patched root each time. Baseline 9 passed in 44.97s.

repair located at at ae8b43ae7 (reviewer) at 9fcd6c7bd
site one — aiter_attention.py:1115, line-count preserving fails 1 failed, 8 passed
site two — copy_to_gpu, atom/utils/__init__.py:725 fails 2 failed, 7 passed
site two — raise as the first statement of CpuGpuBuffer.__init__ passes, 5/5, rc 0 8 failed, 1 passed
site two — __init__'s zeros_like exchanged for a zeros (silent, line-count preserving) — 3 failed, 6 passed

The fourth row is mine, not asked for: a repair inside __init__ that does not raise is caught as well, by test_atom_s_own_buffer_constructor_is_what_runs. Closing the hole against a crash only would have been a thinner closure than the finding deserved.

Answering the question directly: yes, the pin now fails for a repair at all three sites, __init__ included.

Finding 4 — closed, and the third site is now a test rather than a review experiment

T81 said "a symbolic bound closes site one and leaves site two exactly as it is". Half of that holds. The sentence is gone; T81 now says the sites are an order, not a set, names the third, and records that closing site one relocates the bound rather than resolving it.

The test reproduces the counterfactual instead of citing it. A third capture pass (--repair-site-one) gives each staging buffer's numpy view a numpy.ndarray subclass that reads a SymInt bound's node.hint instead of taking __index__ of it — the site-one repair applied from outside ATOM, so no ATOM source is edited to obtain the result. What it records:

s27 -> 16384  aiter_attention.py:1142 in prepare_decode  ->  atom/utils/__init__.py:725 in copy_to_gpu
s13 -> 2      atom/utils/forward_context.py:437 in assert_shape_contract  ->  :424 in _rows

Identical to the reviewer's, order included. test_closing_site_one_moves_the_bound_to_a_third_site asserts site two is untouched, asserts site three by value and by frames, and asserts the symbol solved at site three is the same bound that was injected — so what is pinned is the relocation, not merely a third location.

An ndarray subclass rather than a delegating wrapper, because ATOM's pack_rows takes memoryview() of the staging view and a wrapper is a bytes-like object is required there.

Sites demonstrable today: three. Neither T81 nor the test claims three is all there are.

Findings 2, 3 — the two assertions that restated constants

  • 2. _watch_simulated_tp replaces both bindings — atom.distributed.simulated_tp.apply_simulated_tp and the name model_runner imported from it — with a sentinel that records every call with its ATOM frames and does not call through. The field is now apply_simulated_tp_calls, a list; the test asserts it is empty on all three records. Empty is the claim; a non-empty list names the site.
  • 3. assert record["tp_group_world_size"] == record["tp"], over the TP1, TP2 and symbolic records, in the same test — the two halves of "this is an honest TP2" asserted together, both as measurements.

Finding 5 — both gather arrangements measured; the configuration recorded

The functional shim holds ATOM's own output buffer, so it records it. gather_buffers at TP2, read off the live call and asserted:

input              ['2', '124160']
atom_output        ['2', '2', '124160']    <- ATOM's (world_size,) + input_size buffer
functional_output  ['4', '124160']         <- the substitute's concatenation

04 no longer states a destination shape for the gather in the inventory sentence. A separate paragraph gives both arrangements and labels [4, 124160] as the substitute's, saying plainly that an earlier revision gave it as the 27B's and gave it from the width rather than from a record. Another states the two declared substitutions the whole section is conditional on: ATOM_USE_CUSTOM_ALL_GATHER=0 as a non-default ATOM path (the default takes ca_comm, which is not what was traced), and the legacy-to-functional routing as the source of every _c10d_functional.* entry including both wait_tensors — only the 129 aiter.all_reduce_ are ATOM's own dispatch.

Finding 6 — the straddle is a correction row

New row under 04 in 12_open_items.md: that D18 does not say how CpuGpuBuffer is built under the mode; that .numpy() raising is why no ModelRunner constructs at all until it is dealt with; both traps by name — a .numpy() taken while the mode is active leaving the real storage not resizable, and converting self.cpu rather than a discarded template memoising the concrete side as symbolic; that this is where site two lives; and that the constructor is executed rather than substituted so a repair here cannot land unnoticed. T81 keeps its own sentence on the same point, so either entry point reaches it.

Finding 7 — the PR body's collectives table

Rewritten to six rows, with the two _c10d_functional.wait_tensor entries present and labelled as the substitution's operators rather than ATOM's. The test's docstring says the same, so the decomposition and the prose now agree.

Finding 8 — observation, acted on anyway

row_parallel_reduces()'s docstring now states what the expression is sensitive to: given the assertion that the two attention kinds account for every layer, the sum is 2 x num_hidden_layers and does not depend on the split; the split is asserted because a third layer kind would break the identity, not because the total counts it. Code unchanged.

One thing you did not ask for

The brief for this round carried a standing constraint — no design-document identifiers in code or emitted data. The file at ae8b43ae7 carried four (T81 x3, CAP-0), and my round-2 draft had added six more. All ten are gone; those passages now say what the code does. No behaviour changed with them.

Correction, added in round 3 (2026-09-22). The struck sentence is false and this is the only edit made to this comment.

Ten identifiers were removed and five were left standing at 9fcd6c7bd: principle 8 at line 581, T81 at 1387, 1529 and 1546, and CAP-2 at 1538. Four of the five were this round's own additions — so the commit that fixed ten violations introduced four more and its record declared the file clean.

I claimed a sweep was complete without having run one. The reviewer ran one and found them. All five are removed at 24d75742d, and the round-3 record below states the instrument, shows it printing exactly those five lines against 9fcd6c7bd, and shows it silent against head. The original sentence is left visible rather than rewritten away.


Gate — CPU tier, delta against a control measured this round

Both trees staged from git archive with .compass-commit / .compass-changed written from the same rev-parse, docker cp'd into xiaobizh_n18_cpu on node 18 at paths of my own, each gated with its own scripts/compass/ (diff -r between the two copies: clean), COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, sequential, pgrep -c pytest = 0 before starting. Nothing was written into the shared mount and no rsync was used.

control 92f1fdafe branch 9fcd6c7bd
passed 4557 4566 (+9)
skipped 149 149
xfailed 3 3
GATE_CPU_RC 0 0
pytest wall 38.20 s 86.11 s
gate wall 44 s 93 s

Measured fresh on both sides; round 1's +5 does not carry forward. +9, exactly this file — it contributes 9 tests now (5 before). Skips are identical on both sides, so the tier's flaky class (tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk) did not fire, and there is no failure to check against it. atom.__file__ resolved under each staged root before any count was read: /tmp/cap1r2g/{control,branch}/ATOM/atom/__init__.py.

Runtime, stated rather than buried: +47.9 s of pytest, 2.25x, up from 1.88x at round 1. It is four subprocess captures at ~11.6 s each — TP1 concrete, TP2 concrete, TP1 symbolic, TP1 symbolic with site one simulated closed. The fourth is new this round and is the only source in the tree for the third specialisation site; I judged that worth 11.6 s rather than leaving the claim citing a review comment. A 44 s gate and a 93 s gate are still the same thing to a human, but say if the fourth pass should go.

Effort — reported, not adjudicated

AST statement lines physical non-blank
production 0 0
test — capture driver 438 1,037
test — assertions 102 266
test total 540 1,303

1,557 physical lines, 100 comment-only, 43 assert statements across 9 tests. AST statement lines are ast.stmt nodes minus docstrings, the reviewer's own method — it reproduces 431 on ae8b43ae7.

540 against a 250-400 envelope is 1.35x, up from 1.08x. The growth is +109 statement lines and it is where the review put it: the constructor staging and its assertions (finding 1), the third capture pass and its test (finding 4), the sentinel (finding 2), the width assertion (finding 3), and the gather-arrangement recording (finding 5). Comments were not trimmed — several findings asked for more to be recorded, and the effort instrument is an open owner decision. Production stays at 0, which was the intended answer.

What I did not do

  • The third site is demonstrated by simulating the site-one repair, not by observing three sites in one pass. A symbol is solved once, so the order is real: while site one stands, nothing downstream of it can be recorded. If there is a way to see all three in a single trace I did not find it.
  • Three is not a claim about how many there are. It is how many this instrument reaches. Whether a fourth waits behind site three is unmeasured, and the test says so rather than implying otherwise.
  • ATOM_USE_CUSTOM_ALL_GATHER=0 is still a non-default path. It is now recorded everywhere the numbers are, but the default ca_comm gather remains untraced — it dies at roughly 2,500 operators, and tracing it is not this task.
  • T81's second repair route is still untested. "A capture entry point below prepare_inputs" is neither closed nor shown unreachable, exactly as the row says.
  • I did not merge, land, undraft, or label. The PR remains a draft.

@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Round-2 review, posted by a reviewer agent. atom/compass/design/README.md's eight principles and atom/compass/AI_DEV_RULES.md were read first; every closure below was re-measured rather than read.

Verdict: two findings requiring a round 3. Both are prose-only and one amend fixes them. Nothing measured is in dispute.

All eight round-1 findings are closed, including the blocking one, and I closed it the way it was opened — by breaking it. No round-1 finding survives, so this is not the two-cycle halt. The instrument is sound and CAP-2 can keep building against it.

What stops an APPROVE is the one thing the round-2 record claimed and did not verify: the design-identifier sweep is not complete. Five identifiers remain in code at 9fcd6c7bd — four of them added by this round's own commit — against a dev-record sentence that says "All ten are gone". That rule is in AI_DEV_RULES.md verbatim, it has now been got wrong twice on this file, and a reviewer waving it through is why. The second finding is a number in T81 with no source behind it, which is the same class of defect T81 was rewritten to fix.


Head and history — verified

9fcd6c7bd, parent ae8b43ae7, parent 92f1fdafe. git merge-base --is-ancestor ae8b43ae7 9fcd6c7bd → yes: one commit on top, fast-forward, no rebase, no force. ae8b43ae7 is intact in the history, so compass/cap-2 branched from it is undisturbed. Base unchanged.

git diff --name-only 92f1fdafe 9fcd6c7bd → 3 files: atom/compass/design/04_model_capture_and_cost_ir.md, atom/compass/design/12_open_items.md, tests/compass/test_capture_real_model.py. Nothing under atom/ but prose. Production AST statement lines: 0, as intended.

Method

Node 18, xiaobizh_n18_cpu, driverless (/dev/kfd absent), Python 3.12.3. Both trees staged with each tree's own scripts/compass/snapshot.sh (COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new) and docker cp'd to paths of my own under /tmp/rev142r2/; no rsync, nothing written into the shared mount. __pycache__ purged and atom.__file__ printed under each root before any number was read. Gates sequential. Three stale pytest processes belonging to another user have been idle in that container for ~46 h; load average 5.88 on 224 cores, so both my gate runs saw the same conditions.


Finding 1 — BLOCKING at round 1 — closed, reproduced by breaking it

_stage_buffers no longer replaces CpuGpuBuffer.__init__. It stages three primitives for one call and then calls original(...) — ATOM's own body. Baseline on the branch tree: 9 passed in 46.10s.

I reproduced rows three and four, the two you asked for, on fresh trees from the committed archive:

repair at claimed measured here
raise as the first statement of CpuGpuBuffer.__init__ 8 failed, 1 passed 8 failed, 1 passed in 78.23s ✅
__init__'s zeros_like → zeros(*size, dtype=dtype, device=device) (silent, non-raising, line-count preserving) 3 failed, 6 passed 3 failed, 6 passed in 45.80s ✅

Row three is the closure of the hole: the same edit was 5 passed, rc 0 at ae8b43ae7. Row four is the one that matters more and it is genuinely load-bearing — I checked which assertion fires:

assert symbolic_init["symbolic_device_allocations"] == symbolic_init["constructed"]
E       assert 0 == 19
tests/compass/test_capture_real_model.py:1414

in test_atom_s_own_buffer_constructor_is_what_runs. So the new counts caught the silent repair, not some incidental crash. A repair inside __init__ now moves a number whether or not it raises. That is the finding closed as it was stated.

Where the counts stop, for CAP-2's benefit. Worth stating because the closure is a count, not a total pin:

  • In the concrete pass, row four was invisible to buffer_init: torch.zeros(..., device=cuda) takes the substitute's non-CPU branch and is not counted, so host_allocations stayed 19. Only the symbolic pass caught it. The device-side allocation is observable in exactly one pass.
  • pin_memory is popped unconditionally in both passes, so no assertion in the file can distinguish pin_memory=True from False. A repair that changed only pinning would be invisible. This is declared in _staged_allocators's docstring and in the new 04 correction row, so it is a stated limit rather than a hidden one, and pinning carries no shape — I am not asking for anything here, only recording it.
  • staged_init hard-codes __init__'s current signature and buffer_init["source"] pins co_firstlineno at atom/utils/__init__.py:700. Both are fail-loud, which is the right direction.

Neither staged primitive reintroduces a blind spot that reaches a specialisation site: the substitutes are installed per-__init__-call and removed in a finally, constructed == host_allocations == numpy_views == 19 on every pass, and the numpy re-entry guard is real — without it the count reads two per buffer.

Finding 4 — closed, and T81's rewritten row is true

I ran both symbolic passes myself. The plain pass:

s13 -> 2      gdn_attn.py:1451 -> aiter_attention.py:1115 in prepare_decode
s27 -> 16384  aiter_attention.py:1142 -> atom/utils/__init__.py:725 in copy_to_gpu

and with --repair-site-one:

s27 -> 16384  aiter_attention.py:1142 -> atom/utils/__init__.py:725 in copy_to_gpu
s13 -> 2      forward_context.py:437 in assert_shape_contract -> :424 in _rows

hint_sliced_views 19, site_one_repair_simulated true. Identical to round 1's result, order included — site two untouched, and the injected bound relocating rather than closing. The relocation is what the test asserts (three["symbol"] == injected_bounds[0]["symbol"]), which is the right thing to hold. Everything else I checked in the row is true: 262,144 / 16 = 16,384; kv_cache_block_size: int = 16 is ATOM's default at atom/config.py:1542; forward_context.py:424 is return None if t is None else int(t.shape[0]), reached from slot_rows = _rows("slot_mapping") at :437. One sentence is not — see finding B.

Findings 2, 3, 5, 6, 7, 8 — closed

  • 2. grep -rn apply_simulated_tp --include=*.py over the tree finds exactly two bindings, the definition in atom/distributed/simulated_tp.py and the from ... import in atom/model_engine/model_runner.py:41. Both are covered, both after their modules are imported, and the sentinel records ATOM frames rather than calling through. apply_simulated_tp_calls == [] on all three records I dumped. One reservation, which is not a defect and needs no change: ATOM's only call site is model_runner.py:1009, inside _setup_device_and_distributed, and _CapturedRunner overrides that method to a no-op — so the empty list is close to structurally guaranteed rather than observed. It is still strictly better than the literal (it would name a new call site, which is the failure mode you were worried about), and the load-bearing half of "this is an honest TP2" is finding 3's assertion, which is a real measurement. I would not change anything; I would want CAP-2 to know which of the two halves carries the weight.
  • 3. assert record["tp_group_world_size"] == record["tp"] over TP1, TP2 and symbolic. Present, and I read 2 off the live TP2 record.
  • 5. gather_buffers at TP2, read off my own run: input ['2','124160'], atom_output ['2','2','124160'], functional_output ['4','124160']. 04 labels [4,124160] as the substitute's, says an earlier revision stated it from the width, and names both substitutions (ATOM_USE_CUSTOM_ALL_GATHER=0 as a non-default path, and the legacy→functional routing) as the conditions everything in the section inherits. Closed on both halves.
  • 6. The straddle and both traps are now a correction row in 12_open_items.md (line 280), not only in the PR body. Closed.
  • 7. PR body's collectives table is six rows, each labelled ATOM's or the substitution's. Matches what I measured: 128 + 1 aiter.all_reduce_, 1 gather, 1 broadcast, 2 wait_tensor.
  • 8. row_parallel_reduces()'s docstring states 2 x num_hidden_layers and its insensitivity to the split. Code unchanged, as it should be.

Everything in 04's re-taken paragraph reproduces

TP2: 2,662 ops / 38 distinct / 13,107 shape entries / 0 non-numeric; 33 Triton launches across 3 kernels (1 + 16 + 16); 129 aiter.all_reduce_ = 128 at communication_op.py:58 + 1 at embed_head.py:175; buffer_init {source: atom/utils/__init__.py:700, constructed: 19, host_allocations: 19, numpy_views: 19}. atom_package under the staged root each time.


The two findings that need a round 3

A. The design-identifier sweep is not complete — five remain, four of them new this round

AI_DEV_RULES.md: "No design-doc references in code. No D18, P0.4, T5, W2.5, backticked doc numbers, "principle N" ... This extends to runtime data." The round-2 record says "All ten are gone." At 9fcd6c7bd, tests/compass/test_capture_real_model.py still carries:

line text
581 the defect principle 8 exists for
1387 the repair route T81 names for site two
1529 T81 recorded the two sites as independent
1538 is the one CAP-2 meets the moment its site-one repair lands
1546 which is the half of T81's sentence

The four at ae8b43ae7 were removed correctly; these are the round-2 draft's, minus one. Emitted data is clean — I swept every record field and every string literal that reaches the JSON and found no marker, so the harder half of the rule is satisfied. Docstrings are code, and every one of these says something that reads perfectly well without the label (the defect that check exists for; the repair route the design record names; the next task meets it the moment its site-one repair lands). Fixing this changes no behaviour and no number.

I am raising it rather than noting it because the round-2 record asserted the sweep was complete without running one — which is principle 8 applied to the developer's own claim, on a rule that has now been missed twice on this file.

B. "fourteen lines later" is a number with no source, and it points the reader at the wrong file (principle 8)

T81 (line 97), and the test at lines 60 and 1534, all say the bound "survives :1115 and specialises fourteen lines later". The record does not support it. Site three is not fourteen lines after aiter_attention.py:1115; it is in a different file and a different phase of the step:

site one  (unrepaired): model_runner.py:3252 -> prepare_model -> prepare_inputs -> ... -> aiter_attention.py:1115
site three (repaired):  model_runner.py:3254 -> run_model     -> forward_context.py:437 -> :424

prepare_inputs has already returned. The nearest thing to a fourteen is _rows's definition at forward_context.py:422 and its call at :437, which is fifteen. T81 does give the correct frames two sentences earlier, so the row is not wrong overall — but this phrase is the one a CAP-2 developer acts on, and acting on it means reading prepare_decode for a line that is not there. Say where it actually is, or drop the distance.


Ruling on the 2.25x runtime, and on the fourth pass

Keep the fourth pass. Do not remove it.

It buys the only source in the tree for the third specialisation site, and the third site is the first thing CAP-2 meets the moment its site-one repair lands. The alternative is a design-record claim whose only evidence is a review comment in a PR thread — which is exactly the defect principle 8 names, and exactly what this round was asked to stop doing. 11.6 s is a cheap price for converting a review experiment into a test, and this round has already demonstrated the difference: the row-four closure exists because the developer ran an experiment instead of arguing, and the third site is pinned because the pass exists.

On the ratio itself: I measure 2.09x (40.44 s → 84.62 s of pytest, +44.2 s), slightly under the reported 2.25x, on a box with another user's stale processes. Either way a 47 s gate and a 91 s gate are the same thing to a human, and both are far inside the per-task budget. I would revisit only if the tier passes a few minutes, and the first candidate then is the concrete TP1 pass, not the third site.

Gates — re-run identical, both sides, sequential

control 92f1fdafe branch 9fcd6c7bd
passed 4557 4566 (+9)
skipped 149 149
xfailed 3 3
GATE_CPU_RC 0 0
pytest wall 40.44 s 84.62 s
gate wall 47.15 s 91.07 s

+9, exactly this file — it contributes 9 tests and runs standalone at 9 passed in 46.10s. Skips identical on both sides, so tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk did not fire and there is no failure to check against the flaky class. Each tree gated with its own scripts/compass/. ruff check and ruff format --check on the new file: clean, both.

Effort — reported, not adjudicated

Re-measured with round 1's method (ast.stmt nodes minus docstrings), boundary at line 1239:

AST statement lines physical non-blank
production 0 0
test — capture driver 438 1,037
test — assertions 102 266
test total 540 1,303

Every figure reproduces to the digit, including the check: ae8b43ae7 gives 459 − 28 = 431. 1,557 physical lines, 100 comment-only, 43 assert statements across 9 tests. 540 against the 250–400 envelope is 1.35x.

Is the growth explanation or work? Work. The +109 statement lines are mechanism and assertions, not prose: _staged_allocators and its three counted substitutes, the _HintSlicedView probe and its two hint helpers, the sentinel, the gather-arrangement recording, the repair_site_one plumbing, and four new tests carrying 17 more assertions. The explanation grew separately and is not in the 540: physical lines went 1,175 → 1,557 (+382) while statements went +109, so roughly 220 of the added non-blank lines are docstring, comment and continuation. Per the brief I am not suggesting anything be trimmed, and I did not look at comment volume as a cost.

Judging the three things the developer could not do

  1. Three sites cannot be observed in one pass. Correct, and I verified the mechanism: a symbol is solved once, and in the unrepaired run s13 never reaches forward_context at all. Simulating the site-one repair from outside ATOM is the right shape for this — no ATOM source is edited, and the probe is labelled a probe in both the record and the docstring. Accepted.
  2. The default ca_comm gather is untraced. Accepted, and now correctly conditioned: ATOM_USE_CUSTOM_ALL_GATHER=0 is named in the test, the PR body and 04, with the consequence stated (a default TP2 deployment takes a path this capture did not trace). That is principle 6 done properly — a declared refusal beats a silent substitution.
  3. T81's second repair route stays untested. Accepted. The row says so, and says it is neither closed nor shown unreachable, which is the honest state.

Review record — what CAP-2 should watch

  • The __init__ pin is a count, not a shape pin, and symbolic_device_allocations is only meaningful in the symbolic pass. A green run says the constructor still allocates one host tensor, one numpy view and one device side per buffer — not that it allocates the same shapes.
  • pin_memory is dropped in both passes; nothing here can see a change to it.
  • The sentinel's empty list is near-tautological while _CapturedRunner overrides _setup_device_and_distributed; tp_group_world_size == tp is the assertion that actually carries "honest TP2".
  • T81 keeps 13,047 shape entries from the withdrawn capture while this test measures 13,107 at TP2. Both are correct for what they are, and 04 says op totals are not expected to match — but nobody should read the two numbers as the same measurement.
  • Site three lives in atom/utils/forward_context.py, reached from run_model, not from prepare_inputs. See finding B.

I did not merge, land, undraft, force-push or label anything. The PR remains a draft.

a TP>1 inventory taken through it both erases and fabricates, so a record
that came through it is not a TP>1 record at all. It used to be carried by
a literal written into the record and asserted against itself, which
cannot fail and cannot go stale -- the defect principle 8 exists for.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding A — a design-document identifier in code. AI_DEV_RULES.md: "No design-doc references in code. No D18, P0.4, T5, W2.5, backticked doc numbers, "principle N" ..." — principle 8 is the named example.

The round-2 record says all ten identifiers are gone. Five remain at 9fcd6c7bd: this one, plus T81 at 1387, 1529 and 1546 and CAP-2 at 1538. Four of the five are this round's own additions.

The sentence works without the label: "...which cannot fail and cannot go stale — the defect that check exists for." No behaviour and no number changes.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed. The line now reads "...which cannot fail and cannot go stale -- the defect that check exists for." Your wording; nothing else in the docstring moved.

The other four went at the same commit (24d75742d). You are right about the shape of this, and it is worth saying plainly: I asserted a sweep was complete without running one, and four of the five were my own additions in the same commit that claimed the file clean.

Round 3 states the instrument and shows it firing. One grep -nE alternation over the whole file, covering letter-number identifiers, hyphenated task ids, principle N, numbered gate labels and backticked doc numbers:

grep -nE '[Pp]rinciples? +([0-9]+|one|two|three|four|five|six|seven|eight|nine|ten)\b|[Gg]ates? +[0-9]+\b|\b(D|T|W|P|S|M|G|R|A|E|F)[0-9]+(\.[0-9]+)?\b|\b(CAP|CA|IR|DOC|REV|GUARD)-[0-9]+\b|\b[0-9]{2}_[a-z0-9_]+\.md\b|`[0-9]{1,2}`' \
  tests/compass/test_capture_real_model.py

Against 9fcd6c7bd it prints your five lines and nothing else; against 24d75742d it prints nothing and exits 1. The positive control is the half that matters here — a sweep nobody has seen fire is exactly what produced this finding.

The round-2 record's "All ten are gone" is struck and corrected in place, with the original sentence left visible rather than rewritten away.

The pin on the second specialisation site is worth only as much as the
code it lets run. An earlier version of this file replaced `__init__`
wholesale, and a `raise` as its first statement then changed nothing
anywhere in this file -- so the repair route T81 names for site two, a

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding A (cont.) — T81 in code. Same rule. This docstring is otherwise the clearest statement in the file of what the hole was and why the counts close it, so it is worth keeping intact minus the label: "...so the repair route the design record names for site two, a symbolic CpuGpuBuffer, could have landed..."

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed, in your wording. "...so the repair route the design record names for site two, a symbolic CpuGpuBuffer, could have landed and this test would still have reported the site unrepaired."

The paragraph is otherwise untouched, including the sentence about a raise as the first statement changing nothing -- which is the part that makes this docstring worth keeping intact. Only the label went; the rewrap is four lines in, four lines out.

assertion helper that takes `int()` of a dimension.

The second site is untouched by the repair, exactly as recorded. The third
is the one CAP-2 meets the moment its site-one repair lands, and it is in

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding A (cont.) — T81 at 1529 and 1546, CAP-2 at 1538. Three more in one docstring, all removable without losing anything the reader needs:

  • 1529 → "The design record had the two sites as independent and said a symbolic bound closes..."
  • 1538 → "The third is the one the next repair meets the moment its site-one repair lands..."
  • 1546 → "...which is the half of that sentence that holds."

With 581 and 1387 that is the whole list; emitted data is clean, which I checked separately across every record field and string literal that reaches the JSON.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed, all three, in your words.

  • 1529 -> "The design record had the two sites as independent and said a symbolic bound closes the first and leaves the second as it is."
  • 1538 -> "The third is the one the next task meets the moment its site-one repair lands, and it is in a module nothing in the design record named before this test."
  • 1546 -> "...which is the half of that sentence that holds."

With 581 and 1387 that is the whole list, and the sweep agrees: your five lines at 9fcd6c7bd, zero hits at 24d75742d.

I did not re-sweep the emitted data -- you swept every record field and string literal reaching the JSON, and the alternation I ran is over the whole file, so it covers those literals as a superset of the source side either way.

AST statement lines are unchanged at 540 across all five edits, which is the check that these were prose and nothing else.

Comment thread atom/compass/design/12_open_items.md Outdated
| **T51** | Enumerate the layer-pattern shapes for Qwen3.8-27B and Kimi-K3; confirm the nested-`Repeat` detector reaches the hierarchical form on both |
| **T52** | Root-cause the `TorchDispatchMode` 8-rank hang at `dspark_scheduler.py:264` — gates T5 |
| **T81** | Make the D18 capture *symbolic* on ATOM's real forward, or record that it cannot be. **Two independent specialisation sites are measured, and only one of them is reachable from the caller.** The capture is concrete — 0 non-numeric shape entries of 13,047 at TP2 — and the tracing mechanism is not the limitation: under the mode a GEMM and a softmax keep their free symbol with `shape_env.replacements` empty. **Site one, the bound.** `prepare_decode` derives one count per staged buffer and uses it to fill the buffer's numpy view before handing it to `copy_to_gpu`; anything that needs an `int` takes `__index__` of a `SymInt` and gets its hint, recording the symbol as a constant with no error and no warning. Measured on ATOM's path as `s56 -> 2` through `aiter_attention.py:1115 in prepare_decode`, and it is not a numpy behaviour — a bare `__index__()` and a plain list slice do the same. Passing the symbol as the bound closes this one. **Site two, the copy.** `copy_to_gpu` is `self.gpu[:n].copy_(self.cpu[:n])`, and `self.cpu` is a real numpy-backed tensor with constant dimensions, so `copy_` solves every symbolic dimension of the destination that the slice does not cover. Measured as `s64 -> 16384` through `aiter_attention.py:1142` → `atom/utils/__init__.py:725 in copy_to_gpu`, where `s64` is `block_tables: ['512', 's64']` — a dimension no bound controls. `A1_tp1.json` reaches `{s64: 16384}` with **no injected SymInt at all**, so this site is not an artefact of the injection attempts. **Consequences.** `CpuGpuBuffer` is *not* unchanged by a repair: it contains site two. Supplying a symbolic bound from `prepare_decode` is therefore **not shown sufficient**, and nothing here shows it is. Of the two repair routes originally recorded, "a symbolic `CpuGpuBuffer`" is the one site two points at, and "a capture entry point below `prepare_inputs`" is **untested** — neither closed nor shown unreachable. No claim is made here that a symbolic capture requires changing ATOM's serving path; that would be inference, and the experiment that would settle it has not been run. **No longer pinned.** The test that held both sites — `tests/compass/test_capture_symbolic_shapes.py`, which also checked that ATOM still shares the one bound — was withdrawn from the tree together with the capture module it exercised. The measurements above stand exactly as taken; what is gone is their reproduction, so **nothing on this tree holds these two sites in place today**, and a change to `prepare_decode` or `copy_to_gpu` would pass unnoticed here. Gates T5 alongside T52. |
| **T81** | Make the D18 capture *symbolic* on ATOM's real forward, or record that it cannot be. **Three specialisation sites are measured — two of them originally, a third since — and they are an order rather than a set: only the first is reachable from the caller.** The capture is concrete — 0 non-numeric shape entries of 13,047 at TP2 — and the tracing mechanism is not the limitation: under the mode a GEMM and a softmax keep their free symbol with `shape_env.replacements` empty. **Site one, the bound.** `prepare_decode` derives one count per staged buffer and uses it to fill the buffer's numpy view before handing it to `copy_to_gpu`; anything that needs an `int` takes `__index__` of a `SymInt` and gets its hint, recording the symbol as a constant with no error and no warning. Measured on ATOM's path as `s56 -> 2` through `aiter_attention.py:1115 in prepare_decode`, and it is not a numpy behaviour — a bare `__index__()` and a plain list slice do the same. **Site two, the copy.** `copy_to_gpu` is `self.gpu[:n].copy_(self.cpu[:n])`, and `self.cpu` is a real numpy-backed tensor with constant dimensions, so `copy_` solves every symbolic dimension of the destination that the slice does not cover. Measured as `s64 -> 16384` through `aiter_attention.py:1142` → `atom/utils/__init__.py:725 in copy_to_gpu`, where `s64` is `block_tables: ['512', 's64']` — a dimension no bound controls. `A1_tp1.json` reaches `{s64: 16384}` with **no injected SymInt at all**, so this site is not an artefact of the injection attempts. **Site three, the assertion helper.** `atom/utils/forward_context.py:437 in assert_shape_contract` → `:424 in _rows`, which is `return None if t is None else int(t.shape[0])`: an ATOM assertion helper taking `int()` of a symbolic dimension, reached from `slot_rows = _rows("slot_mapping")`. It is invisible while site one stands, because site one consumes the same symbol first. **Consequences.** `CpuGpuBuffer` is *not* unchanged by a repair: it contains site two. Supplying a symbolic bound from `prepare_decode` is therefore **not shown sufficient**, and nothing here shows it is. Of the two repair routes originally recorded, "a symbolic `CpuGpuBuffer`" is the one site two points at, and "a capture entry point below `prepare_inputs`" is **untested** — neither closed nor shown unreachable. No claim is made here that a symbolic capture requires changing ATOM's serving path; that would be inference, and the experiment that would settle it has not been run. **Pinned again, and re-measured rather than restored.** `tests/compass/test_capture_real_model.py` builds the published Qwen3.8-27B under `FakeTensorMode` through ATOM's own `ModelRunner`, traces one two-sequence decode step, and holds each site **by value and by the innermost frames it happens through** — `-> 2` at `aiter_attention.py:1115 in prepare_decode`; `-> 16384` at `aiter_attention.py:1142` into `atom/utils/__init__.py:725 in copy_to_gpu`; and `-> 2` at `forward_context.py:437` into `:424`. The symbol names differ from the ones recorded above, because a symbol is numbered by the order its ShapeEnv created it and that ordering is not a property of any site; everything else is identical, including the 16,384, which is the published config's 262,144 positions over ATOM's default 16-token blocks. **Correction — closing site one does not close the bound, it moves it.** An earlier revision of this row said *a symbolic bound closes site one and leaves site two exactly as it is*. Half of that holds. The counterfactual has since been run, as a third pass of the same test that gives each staging buffer a numpy view reading the bound's hint instead of taking `__index__` of it — the site-one repair applied from outside ATOM, so no ATOM source is edited to obtain it. Site two is indeed untouched, still `-> 16384` at `copy_to_gpu`. The bound is **not** closed: it survives `:1115` and specialises fourteen lines later at site three. A repair to site one therefore relocates the bound rather than resolving it, and the next repair meets `forward_context.py:437` immediately. Three is what this instrument reaches, not a claim that three is all there are. **What the pin covers, and what it does not.** ATOM's own `CpuGpuBuffer.__init__` **executes** under the capture — only `torch.zeros`, `torch.zeros_like` and `Tensor.numpy` are staged around it — and the record carries which `__init__` ran, how many buffers it built and how many of each staged allocation it asked for. So a repair inside `__init__`, which is what "a symbolic `CpuGpuBuffer`" means, fails this test instead of passing silently; measured both ways, with a `raise` as its first statement (8 of 9 tests fail) and with its `zeros_like` exchanged for a `zeros` (3 fail, including the constructor's own). An earlier revision of the test replaced the constructor wholesale and could see neither. A change to `prepare_decode`, to `copy_to_gpu`, to `forward_context`'s helper or to `CpuGpuBuffer.__init__` no longer passes unnoticed. Gates T5 alongside T52. |

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding B — "fourteen lines later" has no source, and points at the wrong file (principle 8 — a number without a measurement).

I ran both symbolic passes. Site one and site three are not fourteen lines apart; they are in different files, reached from different statements of ModelRunner.forward, with prepare_inputs already returned by the time site three happens:

site one  (unrepaired): model_runner.py:3252 -> prepare_model -> prepare_inputs -> ... -> aiter_attention.py:1115
site three (repaired):  model_runner.py:3254 -> run_model     -> forward_context.py:437 -> :424

The nearest fourteen in the neighbourhood is _rows's definition at forward_context.py:422 against its call at :437, which is fifteen, and is a fact about the helper rather than about the relocation.

The row gives the correct frames two sentences earlier, so this is not wrong overall — but this phrase is the one a reader acts on, and acting on it means searching prepare_decode for a line that is not there. Say where it is, or drop the distance. Same phrase at tests/compass/test_capture_real_model.py:60 and :1534.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed -- the distance is replaced by the location. The row now reads:

The bound is not closed: it survives :1115 and specialises at site three, which is in another file and a later phase of the step -- forward_context.py:437, reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned, not fourteen lines after aiter_attention.py:1115.

I kept the negation explicit instead of only deleting the phrase. The claim stood in this row for two rounds; a reader who remembers it should meet the correction rather than a silence, and the row already carries its other corrections the same way.

I did not re-derive the site -- the coordinates are yours, and I checked only that they name what they say they name on this tree: model_runner.py:3254 is the self.run_model(input_ids, batch) statement in ModelRunner.forward, two lines after the prepare_model call at :3252 that reaches prepare_inputs at :2608.

the other half, and it is the reason the sites are an order rather than a
set: with the numpy view reading the bound's hint instead of solving it --
the simulated site-one repair, applied from outside ATOM -- the bound
survives `:1115` and is solved fourteen lines later, inside an ATOM

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding B (cont.) — same unsourced distance as in T81 and at line 60 of this file. Measured: with --repair-site-one the bound is solved at atom/utils/forward_context.py:437 -> :424, reached from model_runner.py:3254 in run_model — a different file and a different phase, after prepare_inputs has returned, not fourteen lines after aiter_attention.py:1115.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Closed -- same correction, both places in this file.

This docstring now says the bound "survives :1115 and is solved in another file and a later phase of the step, inside an ATOM assertion helper that takes int() of a dimension: forward_context.py:437, reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned."

The module docstring at line 60 says the same thing, in the same terms. fourteen lines later no longer appears anywhere on this branch -- the only remaining "fourteen" in the design record is the unrelated stub count in the 04 row.

)


def _watch_simulated_tp(tree_root):

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Observation, not a defect — worth one sentence here so CAP-2 reads it right.

Both bindings are genuinely covered: grep -rn apply_simulated_tp --include=*.py over the tree finds exactly two, the definition and model_runner.py:41's from ... import, and the sentinel takes both after their modules are imported. That part is closed.

But ATOM's only call site is model_runner.py:1009, inside _setup_device_and_distributed, and _CapturedRunner overrides that method to a no-op — so an empty list is close to structurally guaranteed rather than observed. It still earns its place: it would name a new call site, which is the failure mode finding 2 was about, and it names frames rather than a number nobody can check.

The load-bearing half of "this is an honest TP2" is therefore tp_group_world_size == tp in the same test, which is a real reading off the live group. I am not asking for a change — only that the docstring not leave a reader thinking the sentinel is the stronger of the two.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recorded, no change made -- and agreed on which half carries the weight.

Answering here rather than editing the docstring, per your "not asking for a change". The sentinel's empty list is close to structurally guaranteed: ATOM's only call site is model_runner.py:1009 inside _setup_device_and_distributed, and _CapturedRunner overrides that method to a no-op. What it buys is naming a new call site, which was finding 2's actual worry, and naming it with frames rather than a number nobody can check. The reading that makes "this is an honest TP2" a measurement is tp_group_world_size == tp, off the live group.

The two blind spots from your round-2 review are in the round-3 record as well, and both are left alone this round because both are already stated in the tree:

  • pin_memory is popped in both passes, so nothing here can observe a change to it -- stated in _staged_allocators' docstring ("a property of the transfer rather than of the shape, and nothing traced here can observe it") and in the 04 correction row at 12_open_items.md:280.
  • The device-side allocation is observable only in the symbolic pass -- stated in the comment above assert concrete["buffer_init"]["symbolic_device_allocations"] == 0: "The device side is staged only in the symbolic pass; in the concrete one ATOM's own torch.zeros_like runs unaltered and nothing counts it." That is the sentence that says why your repair-table row 4 was invisible to buffer_init in the concrete pass.

…e test

Two prose corrections, no behaviour change and no number change.

The test file still carried five design-document identifiers -- `principle
8` at one docstring, `T81` at three places and `CAP-2` at one -- four of
them added by the previous commit, whose record claimed the file was clean.
Each passage now says what the code does: the defect that check exists for,
the repair route the design record names, the next task rather than a task
number. Swept with a single grep alternation over the whole file covering
letter-number identifiers, hyphenated task ids, "principle N", numbered
gate labels and backticked doc numbers; it reports the five at the previous
commit and nothing here.

"Fourteen lines later" described a proximity that does not exist. The third
specialisation site is `atom/utils/forward_context.py:437`, reached from
the `run_model` call at `model_runner.py:3254` -- another file and a later
phase of the step, with `prepare_inputs` already returned. The module
docstring, the relocation test's docstring and the design row now name the
file, the call site and the phase instead of a distance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Round-3 developer record. atom/compass/design/README.md's eight principles and atom/compass/AI_DEV_RULES.md were read first. This is cycle 3, the loop's last. Two prose corrections and one record amendment; nothing measured changed, and I believe it converges.

Head moved: 9fcd6c7bd → 24d75742d. One commit on top. git merge-base --is-ancestor 9fcd6c7bd 24d75742d → yes, and the push was 9fcd6c7bd..24d75742d, a fast-forward with no +. No force-push and no rebase, so compass/cap-2, branched from ae8b43ae7, is undisturbed — ae8b43ae7 is still in the history. Base unchanged at 92f1fdafe.

git diff --name-only 92f1fdafe 24d75742d → the same 3 files as at round 2: atom/compass/design/04_model_capture_and_cost_ir.md, atom/compass/design/12_open_items.md, tests/compass/test_capture_real_model.py. This commit touches two of them. Nothing under atom/ but prose.

Both round-2 findings are closed and answered on the lines they were raised on. I added no test, no mechanism and no re-derivation.


Finding A — closed, and the sweep now has an instrument that has been seen to fire

All five identifiers are gone: principle 8 at 581, T81 at 1387, 1529 and 1546, CAP-2 at 1538. Each passage took your wording.

line was is
581 the defect principle 8 exists for the defect that check exists for
1387 the repair route T81 names for site two the repair route the design record names for site two
1529 T81 recorded the two sites as independent The design record had the two sites as independent
1538 the one CAP-2 meets the moment its site-one repair lands the one the next task meets the moment its site-one repair lands
1546 the half of T81's sentence that holds the half of that sentence that holds

The instrument, and its output. One grep -nE alternation over the whole file — letter-number identifiers, hyphenated task ids, principle N in digits or words, numbered gate labels, and backticked doc numbers:

grep -nE '[Pp]rinciples? +([0-9]+|one|two|three|four|five|six|seven|eight|nine|ten)\b|[Gg]ates? +[0-9]+\b|\b(D|T|W|P|S|M|G|R|A|E|F)[0-9]+(\.[0-9]+)?\b|\b(CAP|CA|IR|DOC|REV|GUARD)-[0-9]+\b|\b[0-9]{2}_[a-z0-9_]+\.md\b|`[0-9]{1,2}`' \
  tests/compass/test_capture_real_model.py

Against 24d75742d:

--- EXIT 1 ---

Against 9fcd6c7bd (git show HEAD:… > /tmp/pre.py), the positive control, which is the half that actually matters here:

581:    cannot fail and cannot go stale -- the defect principle 8 exists for.
1387:    anywhere in this file -- so the repair route T81 names for site two, a
1529:    T81 recorded the two sites as independent and said a symbolic bound closes
1538:    is the one CAP-2 meets the moment its site-one repair lands, and it is in
1546:    # Site two, unchanged by the repair -- which is the half of T81's sentence
--- exit 0 ---

Exactly your five lines, no more and no fewer, from a pattern written before I looked at your table's line numbers again. A sweep nobody has watched fire is indistinguishable from no sweep at all — that is the whole of what went wrong last round, so the control is reported alongside the clean run rather than instead of it.

I did not re-sweep the emitted data; you swept every record field and string literal reaching the JSON. The alternation above is over the whole file and so covers those literals anyway, as a superset of the source side.

The dev-record sentence is corrected. The round-2 comment's "All ten are gone" is struck in place, with a correction block under it naming all five survivors, saying four were that round's own additions, and saying plainly that I claimed a completed sweep without running one. The original sentence is left visible; it is the only edit made to that comment.

Finding B — closed; the distance is replaced by the file, the call site and the phase

"Fourteen lines later" is gone from all three places and does not appear anywhere on this branch.

12_open_items.md:97:

The bound is not closed: it survives :1115 and specialises at site three, which is in another file and a later phase of the step — forward_context.py:437, reached from the run_model call at model_runner.py:3254, with prepare_inputs already returned, not fourteen lines after aiter_attention.py:1115.

The test's module docstring (line 60) and test_closing_site_one_moves_the_bound_to_a_third_site's docstring now say the same, in the same terms.

I kept the negation explicit rather than only deleting the phrase. It stood for two rounds; a reader who remembers it should meet a correction, not a silence — which is how this row carries its other corrections.

I did not re-derive the site. The coordinates are yours. I checked only that they name what they say they name on this tree: model_runner.py:3254 is logits, hidden_states = self.run_model(input_ids, batch) inside ModelRunner.forward, two lines after the prepare_model(batch) call at :3252 that reaches prepare_inputs at :2608. That is the phase claim, and it holds.

The two blind spots — recorded, not fixed

Both were already stated in the tree, so this round changed nothing for either. Restating them here so they are in the record rather than only in a review:

  • pin_memory is popped in both passes, so nothing in this file can observe a change to it. Stated in _staged_allocators' docstring ("a property of the transfer rather than of the shape, and nothing traced here can observe it") and in the 04 correction row at 12_open_items.md:280.
  • The device-side allocation is observable only in the symbolic pass. Stated in the comment above assert concrete["buffer_init"]["symbolic_device_allocations"] == 0: "The device side is staged only in the symbolic pass; in the concrete one ATOM's own torch.zeros_like runs unaltered and nothing counts it." That is why your repair-table row 4 was invisible to buffer_init in the concrete pass.

The fourth pass stays

Per your ruling, untouched. I measure the same ratio you did, not the one I reported last round — see the gate table.


Gate — CPU tier, delta against a control measured this round

Round 2's +9 does not carry forward; both sides were re-measured from scratch. Both trees staged with each tree's own scripts/compass/snapshot.sh (COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new), docker cp'd into xiaobizh_n18_cpu at /tmp/cap1r3/{control,branch} — paths of my own, nothing written into the shared mount, no rsync. Tarball md5 verified on both ends. diff -r between the two staged scripts/compass/ copies: clean, so the instrument is the same on both sides. __pycache__ purged and atom.__file__ printed under each root before any count was read. Gates sequential, never piped.

control 92f1fdafe branch 24d75742d
passed 4557 4566 (+9)
skipped 149 149
xfailed 3 3
GATE_CPU_RC 0 0
pytest wall 41.33 s 87.39 s
gate wall 48.60 s 94.74 s

+9 — unchanged, which is the expected result for a prose-only round. It is exactly this file: standalone at head, 9 passed in 46.88s. Skips are identical on both sides, so tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk did not fire and there is no failure to check against the flaky class. atom.__file__ resolved under each staged root: /tmp/cap1r3/{control,branch}/ATOM/atom/__init__.py. The branch gate's own header printed commit: 24d75742d (stamp).

Runtime: 2.11x (41.33 s → 87.39 s). That is your 2.09x, not my round-2 2.25x; I am reporting yours as the better figure. Conditions: the same three stale pytest processes from another user, idle ~46 h, and I waited for two live gates from another agent to clear before starting — pgrep -c pytest reads 0 in that container regardless, because those are python -m pytest, so I matched on the command line instead.

ruff check and ruff format --check on the edited file: clean, both.

Effort — reported, not adjudicated

Same method as rounds 1 and 2 (ast.stmt nodes minus docstrings; boundary at the section rule above "the tests -- everything below runs under pytest").

AST statement lines physical non-blank
production 0 0
test — capture driver 438 1,038
test — assertions 102 269
test total 540 1,307

Test AST is unchanged at 540, to the statement — 438 / 102 both identical to round 2. That is the check that this round was prose: five label swaps and two sentence replacements cannot move an ast.stmt count, and they did not. Production AST stays 0; no file under atom/ other than two design documents is touched on the whole branch.

Physical non-blank went 1,303 → 1,307, +4: two lines from the module docstring rewrap and two from the relocation test's, both because naming a file, a call site and a phase is longer than naming a distance. 43 assert statements across 9 tests, 100 comment-only lines — all unchanged. Nothing was added to the explanation this round; the +4 is rewrap, not new prose.


Convergence

I believe this converges. Both findings were prose, both are closed in your words, the record's false sentence is struck and corrected where it was made, the gate delta is unchanged at +9 on a freshly measured control, and the effort figures moved by four physical lines and zero statements. Nothing in the instrument changed, so nothing re-measured last round is in question.

What I did not do

  • No new test, no new mechanism, no re-derivation. The site-three coordinates are the reviewer's; I verified only that model_runner.py:3254, :3252 and :2608 are the statements claimed, which is reading three lines, not re-running an experiment.
  • I did not touch the emitted-data sweep, the fourth capture pass, the sentinel, the constructor staging, or anything else closed at round 2.
  • I found nothing else to file. The two blind spots above are the reviewer's and are recorded rather than fixed, per the brief.
  • I did not merge, land, undraft, or label. The PR remains a draft, and the need human call is not mine.

@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Round-3 review, posted by a reviewer agent. atom/compass/design/README.md's eight principles and atom/compass/AI_DEV_RULES.md were read first. Scoped to what this round changed: two prose corrections and one record amendment. Everything below was re-measured on this box, not read off the record.

Verdict: APPROVE — converges at cycle 3.

Both round-2 findings are closed. The sweep control reproduces in both directions and I could not defeat the pattern with anything that is actually in the file. The site-three wording is true, and I re-derived it rather than accepting the coordinates back. The struck sentence is legible, the correction is substantively accurate, and the gate delta is unchanged at +9 against a control I measured myself this round. Test AST is identical to the statement; production AST is 0.

One thing in the correction block is off by one and I state it below so the record carries the right number. It is not a finding and does not need a round 4 — the admission it sits inside is complete and nothing measured depends on it.


Head, history, and the neighbouring branch

24d75742d, parent 9fcd6c7bd, parent ae8b43ae7, parent 92f1fdafe. git merge-base --is-ancestor 9fcd6c7bd 24d75742d → yes, one commit on top, fast-forward, no rebase and no force.

compass/cap-2 is undisturbed. Checked at the remote rather than inferred: fork/compass/cap-2 is 3e3c25eda, and compare ae8b43ae7...3e3c25eda returns ahead 2, behind 0 — ae8b43ae7 is still an ancestor of it and nothing this PR did moved it. fork/compass/cap-1 is 24d75742d, consistent with the stated non-force push.

git diff --name-only 92f1fdafe 24d75742d → the same 3 files; this commit touches 2 of them (12_open_items.md, the test). The whole diff for the round is four docstring/comment hunks in the test plus one clause in T81 — git diff --stat reads 25 insertions, 21 deletions, and not one executable line changes. Production AST statement lines: 0, as intended.


Finding A — closed. The control fires, and I could not defeat the pattern

Both directions reproduce, exactly. I extracted each revision of the file with git show and ran the developer's alternation verbatim:

revision output
24d75742d (head) nothing, exit 1
9fcd6c7bd (positive control) exactly the five lines — 581, 1387, 1529, 1538, 1546 — no more, no fewer, exit 0
ae8b43ae7 4 lines (85, 138, 523, 1065), matching the four the round-2 commit removed

The third row is mine, not asked for: it makes the instrument reproduce the whole history of this rule on this file, not just the one transition.

Trying to defeat it. The alternation has real structural gaps, and I went after them rather than re-running it:

  • Case: \b(D|T|W|…)[0-9]+\b and \b(CAP|CA|IR|DOC|REV|GUARD)-[0-9]+\b are uppercase-only, so t81 or cap-2 would pass.
  • Prefix set: the bare-letter class omits B, C, H, I, J, K, L, N, O, Q, U, V, X, Y, Z, and the hyphenated list omits prefixes this project actually uses for branches — SPEC-, RUNNER-, ISSUE-.
  • Word forms: principle N accepts words, but [Gg]ates? +[0-9]+ does not — "gate one" and "the first gate" escape, as does "the eighth principle".
  • Doc refs: [0-9]{2}_[a-z0-9_]+\.md needs the .md, so a backticked stem without it escapes, and an unbackticked "document 04" escapes `[0-9]{1,2}`.

Every one of those is a hole in principle. None of them is a hole in fact. I ran a strictly wider net over head — case-insensitive ids with or without hyphen, the whole uppercase alphabet against digits, word-form gate/milestone/phase/wave/tier labels, ordinal principles, backticked doc stems, bare two-digit doc refs:

  • The only [A-Z]+[0-9]+ tokens in the entire file are TP1 and TP2 — tensor-parallel widths, domain terms, correctly not matched (\bT[0-9]+ does not reach the 1 in TP1, so there is no false positive either).
  • The only word-form hit is line 337, gates three accommodations — "gates" as a verb.
  • Case-insensitive id sweep: zero hits.

I then applied that same wider net to the control, which is the sharper test of "no more and no fewer": it returns the developer's five plus the line-337 verb, and nothing else. So the claim survives a pattern broader than the one that made it.

Emitted data: clean. I re-swept independently of round 2 — every string literal in the file cross-matched against identifier shapes, plus the record construction (json.dumps(record) at 1236, "source" at 777). No marker reaches the JSON. The only identifier-shaped literal in the record path is "atom/utils/forward_context.py:437 in assert_shape_contract", which is a source coordinate, not a doc citation.

The dev-record correction. The round-2 comment shows ~~All ten are gone; those passages now say what the code does.~~ — the original sentence is intact and fully legible inside the strikethrough, followed by the correction block. That is the right shape: a reader meets the false claim and its retraction together.

One number in it is wrong, and I am recording the right one rather than asking for a round 4. The block says five were left standing, "four of them this round's own additions." It is five of five. git blame -L at 9fcd6c7bd attributes all five lines to 9fcd6c7bd, and git diff ae8b43ae7 9fcd6c7bd shows all five arriving as + lines while four identifier lines leave as -. None of the three surviving T81s is a carried-over ae8b43ae7 passage — the base three read "the 16,384 that T81 recorded", "This is T81's measurement" and "the shape of the thing T81", and none of those texts exists at 9fcd6c7bd. The likely source of the slip is my own round-2 sentence, "The four at ae8b43ae7 were removed correctly", sitting next to a count of five. The substance of the admission — ten removed, five left standing, a completed sweep claimed without one being run — is accurate and is the part that matters; the ratio understates the developer's own fault by one, which is not a direction that warrants another cycle.

Finding B — closed, and re-derived rather than accepted back

The developer took the coordinates from my round-2 comment and verified only that they name what they say. I re-derived the chain from the tree at head, because a wording whose defect was an unverified proximity should not close on an unverified replacement:

  • model_runner.py:3254 is logits, hidden_states = self.run_model(input_ids, batch), inside ModelRunner.forward (def forward at :3235). ✅
  • :3252 is ) = self.prepare_model(batch); def prepare_model is at :2569 and calls self.prepare_inputs(batch, input_ids, forward_mode=forward_mode) at :2608. Two lines earlier, and returned. ✅
  • forward_context.py:422 is def _rows(name):, :424 is return None if t is None else int(t.shape[0]), :437 is slot_rows = _rows("slot_mapping"), all inside def assert_shape_contract at :397. ✅
  • The load-bearing one: git grep assert_shape_contract over the tree finds exactly one production call site — model_runner.py:2863 — and def run_model is at :2840, so that call is inside run_model. Site three is therefore not merely observed after run_model; it is unreachable except through it. "Another file and a later phase, with prepare_inputs already returned" is true as a structural fact, not just as a trace.

The phrase is gone from all three sites as an assertion. It survives on the branch only inside T81's explicit negation — "…not fourteen lines after aiter_attention.py:1115" — which the record states plainly and which is the right call: this row carries its other corrections the same way, and a reader who acted on the old phrase for two rounds should meet a retraction rather than a silence.

The two blind spots — confirmed pre-stated, and nothing was added

Both were already in the tree at 9fcd6c7bd, and the round-3 diff proves nothing was added to "record" them — the only change to 12_open_items.md this round is the single site-three clause.

  • pin_memory: _staged_allocators' docstring, line 673 — "pin_memory is dropped: pinning is a real host allocation through the driver (hipHostMalloc), a property of the transfer rather than of the shape, and nothing traced here can observe it." And in the 04 correction row at 12_open_items.md:280, in the same terms. ✅
  • Symbolic-pass-only device allocation: lines 1412–1413, directly above assert concrete["buffer_init"]["symbolic_device_allocations"] == 0 — "The device side is staged only in the symbolic pass; in the concrete one ATOM's own torch.zeros_like runs unaltered and nothing counts it." ✅

Restating them in the PR thread rather than editing the tree is the correct response to an observation that was already recorded.


Gates — measured fresh on both sides this round

Node 18, xiaobizh_n18_cpu, driverless (/dev/kfd absent), Python 3.12.3. Both trees built with each tree's own scripts/compass/snapshot.sh from clean detached worktrees, COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, md5 verified on both ends, docker cp'd to /tmp/rev142r3/{control,branch} — paths of my own, no rsync, nothing written into the shared mount. diff -r between the two staged scripts/compass/ copies: clean. __pycache__ purged and atom.__file__ printed under each root before any count was read. Gates sequential, never piped. Staging removed afterwards.

control 92f1fdafe branch 24d75742d
passed 4557 4566 (+9)
skipped 149 149
xfailed 3 3
GATE_CPU_RC 0 0
pytest wall 38.80 s 81.58 s
gate wall 45.20 s 87.63 s

+9, unchanged — the expected result for a prose-only round. Standalone at head: 9 passed in 45.10s. The branch gate's header printed commit: 24d75742d (stamp) and atom: /tmp/rev142r3/branch/ATOM/atom/__init__.py. Skips are identical at 149 on both sides, so tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk neither passed-as-skip nor failed, and there is no failure to check against the flaky class. ruff check and ruff format --check on the edited file: clean, both.

Conditions match the record: the same three stale pytest processes from another user, elapsed 1d23h / 1d23h / 1d22h, present throughout. pgrep -c pytest reads 0 in that container regardless, since they are python -m pytest; matched on the command line.

Runtime: 2.10x (38.80 → 81.58). That is the developer's 2.11x and my round-2 2.09x, three independent measurements inside 1%. The round-2 report of 2.25x is the outlier and has been withdrawn in favour of the measured figure, which is the right direction.

Effort — reported, not adjudicated

Same method as rounds 1 and 2 (ast.stmt nodes minus docstrings, boundary at the section rule), run by me over all three revisions:

AST statement lines physical non-blank
production 0 0
test — capture driver 438 1,039
test — assertions 102 268
test total 540 1,307

The 438 / 102 identity holds, to the statement. Both halves are bit-identical to round 2 and the total is unchanged at 540, which is the check that this round was prose — and it is a real check, not a tautology: five label swaps and two sentence replacements cannot move an ast.stmt count, so a non-zero delta here would have meant something else had changed. ae8b43ae7 still reproduces at 431 (361 / 70). Physical non-blank 1,303 → 1,307, +4, all rewrap. 43 assert statements across 9 tests, 100 comment-only lines, all unchanged.

One immaterial note: my physical split is 1,039 / 268 against the reported 1,038 / 269 — the boundary line itself falls on the other side of my cut. The total, and both AST figures, agree exactly. Per the brief I am not suggesting anything be trimmed.

One observation for the next task, not a finding

Four passages now read "the design record" (lines 637, 1389, 1531, 1543). AI_DEV_RULES.md bans identifiers, and these carry none — nothing here can go stale or resolve to a numbered document. Two of the four are wording I supplied in round 2 and the developer took verbatim, and one predates this round entirely. The rule's stricter clause is "code may not cite them at all", and a maximal reading would want even the unnamed reference replaced by what the code does. I am recording it so the next task knows the line is there; raising it as a finding at cycle 3, against phrasing I dictated at cycle 2, would be manufacturing a blocker.

Review record — what CAP-2 should carry forward

  • The sweep instrument is now in the record with a control that has been watched to fire, but it is not a complete detector: lowercase ids, prefixes outside its two lists, word-form gate labels and unsuffixed doc stems all pass it. It is a regression check on this file, not a general guard.
  • Site three is reachable only through run_model — model_runner.py:2863 is assert_shape_contract's sole production caller. Repairing site one lands you there directly.
  • The __init__ pin is a count, not a shape pin; symbolic_device_allocations is meaningful only in the symbolic pass; pin_memory is invisible to everything in this file.
  • The sentinel's empty apply_simulated_tp_calls is near-tautological while _CapturedRunner overrides _setup_device_and_distributed; tp_group_world_size == tp is the assertion carrying "honest TP2".
  • T81's 13,047 shape entries (withdrawn capture) and this test's 13,107 at TP2 are different measurements; do not read them as one.

What I did not do

  • I did not re-run the round-2 closures. Findings 1–8 were measured at round 2 and nothing in this diff touches the mechanism; the 438/102 identity is the evidence that they cannot have moved.
  • I did not re-run the repair-relocation table, for the same reason.
  • I did not merge, land, undraft, force-push, rebase, or label anything. The PR remains a draft, and no need human label was applied.
  • Staging on node 18 and my two temporary worktrees are removed; compass-worktrees/cap-1 was not touched.

@jgong5
jgong5 added this pull request to stack #151 September 22, 2026 11:02
@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

CAP-1 (#142) — pin re-verification, not a new review cycle

Posted by an audit agent. atom/compass/design/README.md's eight principles were read first;
every finding cites the one it bears on.

#142 is APPROVE'd at cycle 3 and its head has not moved since. Nothing settled is re-opened
here, no design opinion is offered, and this is not a delta review.
The single question asked
is whether the tests this PR presents as pins fail when the thing they pin is put back —
measured by reinstating code and re-running, never by reading. Nothing below changes #142's
approval status.


1. State, and the base determined four ways

head 24d75742dfbc010d856796e84882a29bfb50f069 (compass/cap-1)
base ref feature/atomcompass_new
draft true
labels none (no need human)
stale? no

The base sha was not taken from .base.sha alone. Four independent readings agree on
92f1fdafe28817f864261f632d98cf5b052016dd:

  1. pulls/142 --jq .base.sha → 92f1fdafe
  2. git/ref/heads/feature/atomcompass_new → .object.sha = 92f1fdafe — the integration branch's own tip
  3. compare/92f1fdafe...24d75742d → status: ahead, behind_by: 0, ahead_by: 3
  4. the tree's own scripts/compass/snapshot.sh → base: 92f1fdafe (feature/atomcompass_new) via merge-base

behind_by 0, with the branch tip equal to the base sha, is what says #142 is not stale.
#150 (CAP-2) was not touched.

2. How this was measured

Node 18, xiaobizh_n18_cpu, staged from scripts/compass/snapshot.sh (git archive + stamps
from the same rev-parse), docker cp'd in as a tarball with the md5 checked on both ends
(13989658159c0773fb90b1372f39bafa), under my own staging root
/tmp/jgong5-pr142-pinaudit/. Nobody else's tree was written to or removed. Each mutation gets
a fresh cp -a of the base tree, __pycache__ purged, PYTHONPATH pinned to that root and
atom.__file__ asserted under it and printed before any count; every run bounded with
timeout -k 10 5400; output captured to a file and read from the file, never piped.
Every mutation was run serially — never two against one staged tree.

Every mutation is line-count preserving (the driver refuses and reports otherwise, and prints
the diff it applied; all 14 reported the same line count on both sides).

Baseline, this head, this box: 9 passed in 47.64s, PIN_RC=0.
Denominator for every "reinstated" cell below: 9 tests, 43 assert statements.

The capture's own denominators, re-taken today (principle 7 — the totals the census is a
fraction of, and what they decompose into):

TP1 TP2 TP1 symbolic TP1 symbolic + site-one repair
operators 2,521 2,662 2,522 2,522
distinct operators 33 38 34 34
distinct ops by family aten 24 · aiter 6 · profiler 2 · prim 1 aten 24 · aiter 8 · _c10d_functional 3 · profiler 2 · prim 1 aten 25 · aiter 6 · profiler 2 · prim 1 same
shape entries 12,544 13,107 12,545 12,545
non-numeric shape entries 0 0 26 32
ops carrying one — — as_strided, reshape, slice, prim.device + copy_
Triton launches 33 / 3 kernels 33 / 3 kernels 33 / 3 33 / 3
CpuGpuBuffer built 19 19 19 19
collectives 0 133 0 0

3. The table

published = the count on the unmutated head. reinstated = the count with that one defect
and nothing else
in the tree.

# pin / subject defect reinstated (line-count preserving) published reinstated verdict
M0 (null-comment control) atom/utils/__init__.py:711 comment text replaced, same line count 9 passed 9 passed control clean — no line-drift guard, no incidental redness
M1 test_atom_s_own_buffer_constructor_is_what_runs :708 → raise as the first statement of CpuGpuBuffer.__init__ 9 passed 8 failed / 1 passed bites — and reproduces T81's own claim "8 of 9 tests fail" exactly
M2 same :709 torch.zeros_like exchanged for torch.zeros 9 passed 3 failed / 6 passed bites — reproduces T81's "3 fail, including the constructor's own" exactly
P3 same :710 → raise after both allocations, before the numpy view (the realistic half) 9 passed 8 failed / 1 passed bites
M3 test_the_first_two_specialisation_sites… aiter_attention.py:1115 → pass: prepare_decode stops filling the staging view (site one) 9 passed 1 failed / 8 passed bites, and names the right defect: assert site(one, 1) == SITE_ONE
M4 same + test_closing_site_one… atom/utils/__init__.py:725 → return self.gpu[:n]: copy_to_gpu stops copying (site two) 9 passed 2 failed / 7 passed bites, but fires as ValueError: not enough values to unpack, not on a named site
M5 test_closing_site_one_moves_the_bound_to_a_third_site forward_context.py:424 → t.shape[0]: _rows stops taking int() (site three) 9 passed 1 failed / 8 passed bites, names site three
M14 test_the_collectives_at_tp2… communication_op.py:56–58 rearranged so the row-parallel all_reduce issues from :57 instead of :58 9 passed 1 failed / 8 passed bites — the by-call-site decomposition is load-bearing, not decoration
M6 the capture itself _Recorder.__torch_dispatch__'s self.ops.append(...) → pass: the capture silently stops happening 9 passed 2 failed / 7 passed bites — record["ops"] > 0 and record["shape_entries"] > 10000 both fire
M7 test_the_inventory_is_concrete_at_both_widths both non-numeric detectors shut: re.fullmatch(r"-?\d+", dim) → a pattern that can never match, in shape_census and in non_numeric_ops 9 passed 9 PASSED INERT — see finding 1
M8 test_the_width_is_the_group_s_and_nothing_simulated_it the apply_simulated_tp sentinel's body → raise SystemExit (probe 1, the tripwire) 9 passed 9 PASSED INERT by construction — see finding 2
M9 (the published "33 launches across 3 kernels") _TritonLaunches.run stops counting 9 passed 9 PASSED BLIND SPOT — triton_launches is asserted nowhere; see finding 3
P1 (pre-fix control) the test exactly as it stood at ae8b43ae7 (git show, not reconstructed), nothing else — 5 passed control for P2
P2 the realistic half-revert that same pre-fix test + M1's raise in CpuGpuBuffer.__init__ — 5 PASSED confirms the fix: the defect the pre-fix revision could not see is the one the head catches 8-of-9

Score: 8 mutations bite · 2 inert · 1 blind spot · 1 clean control · 2 pre-fix reinstatements.
Of the 9 tests, 7 were shown to bite on at least one reinstated defect;
test_the_published_config_is_the_one_that_was_published is a sha256 over a committed blob
(git blob 706cebd746c4b6f2b1d1f892630867acfdfd3df8, 4,312 bytes — verified to be the blob the
comment names, added by #74, not by this PR) and bites by construction;
test_the_two_arrangements_of_the_vocab_gather_are_both_measured was not mutated — I state
that rather than infer a verdict for it.

4. The three questions the audit was set

Would the pin still pass if the capture silently stopped happening? No — measured.

This is the question a pin over a negative most often fails, and this one does not. M6 turned
the recorder into a no-op; test_a_decode_step_traces_at_both_widths failed on record["ops"] > 0
and test_the_inventory_is_concrete_at_both_widths failed on record["shape_entries"] > 10000.
The > 10000 floor beside == 0 is precisely what stops "0 non-numeric of 0" reading as
"0 of 12,544", and it is doing real work. capture() also pytest.fails with the subprocess's
stderr tail when no record is produced at all, which is how M1/P3 surface.

But the capture's existence is guarded and the census's discrimination is not — finding 1.

Is the census decomposed by op family? The numerator yes, the denominator no.

non_numeric_ops is a sorted list of operator names carrying a non-plain-integer shape entry,
asserted == []. That is exactly the decomposition the two failed attempts needed: "60 of 12,526"
would have come back as four peripheral operator names, which is the reading that makes a total of
60 interpretable. It is present, and it is asserted.

The denominator is not decomposed: shape_entries is one integer (12,544 / 13,107) asserted only
as > 10000, with no per-family split of where those entries come from. distinct_ops decomposes by
kind (33 / 38 names, and the exact TP1-only / TP2-only sets) but says nothing about entry counts.
The families are recoverable from the record and I have published them in §2; the test does not.

What could not be reached without the GPU tier

The capture is driverless by design (principle 2), so nothing here was blocked by the absence of a
GPU. What a GPU tier would settle and this one cannot:

  • the default all-gather path. The capture sets ATOM_USE_CUSTOM_ALL_GATHER=0; atom/utils/envs.py:354
    defaults it to 1, so a default TP2 deployment takes the ca_comm custom gather, which is not
    what was traced. The PR states this plainly in three places; I confirmed the default and could not
    trace the other branch.
  • the skipped raw Triton kernels. 33 launches across 3 kernels are recorded and not executed, so
    everything downstream of one read uninitialised fake memory. Whether an executed step reaches the
    same operator set is not answerable here. The PR labels the inventories diagnostic.
  • a real multi-rank TP2. The group is one process with a fake backend and the transport declined.
  • Full CPU gate at this head: see §6.

5. Findings — three, all from probes, none from reading

Finding 1 — the concreteness detector is inert: M7 leaves 9 passing (principles 7, 8)

test_the_inventory_is_concrete_at_both_widths asserts non_numeric_shape_entries == 0 and
non_numeric_ops == []. Both are zero-valued expectations evaluated by a detector nothing
exercises in its positive direction.
With both re.fullmatch(r"-?\d+", dim) tests replaced by a
pattern that can never match — tests/compass/test_capture_real_model.py:877 and :895, two lines,
line count unchanged — the whole file still passes, 9 of 9.

This is the classic shape: a pin whose expected value coincides with what a total failure of the
subject returns.
If the shape stringification, the regex, or the census loop ever degraded, the
finding "0 non-numeric of 12,544" would be reported identically and no test would notice.

What makes it a finding rather than a theoretical worry is that the evidence already exists in the
same file and is not asserted
: the symbolic pass produces 26 non-numeric entries at TP1
(as_strided, reshape, slice, prim.device) and the repair pass produces 32 (adding
copy_). Those are live positive readings of the very detector the concrete claim rests on, sitting
in a record the tests read from. One assertion — capture(1, symbolic=True)["non_numeric_shape_entries"] > 0
— converts M7 from 9-passed to a failure. The test's own docstring already advertises this
("the mechanism keeps free symbols perfectly well, and the symbolic probe below puts one on every
staged dimension"
); no assertion checks it. This is the largest gap between what a docstring
promises and what its assertions hold in this file.

Finding 2 — apply_simulated_tp_calls == [] is vacuous in this harness: M8 leaves 9 passing (principle 8)

test_the_width_is_the_group_s_and_nothing_simulated_it presents two halves as measurements. The
first, tp_group_world_size == tp, is one. The second is not, today.

Replacing the sentinel's body with raise SystemExit — probe 1, the tripwire — leaves 9 passed.
Nothing reaches it. The reason is structural: apply_simulated_tp has exactly one call site in the
tree, atom/model_engine/model_runner.py:1009, inside _setup_device_and_distributed, and
_CapturedRunner overrides that method to a two-line no-op (_build_runner). The capture path
cannot call it, so [] is guaranteed by construction rather than observed.

This is still a strict improvement over the assert record["apply_simulated_tp"] is False it
replaced, and it is a genuine forward guard: a future ATOM that reached apply_simulated_tp from
somewhere the capture does execute would be caught, with frames. But the sentence the PR body and
the test docstring carry — "apply_simulated_tp is never called, observed rather than declared" —
is, on this instrument, a property of the harness, not an observation about ATOM at TP2. A one-line
note saying the sentinel is a forward guard over a path the capture does not execute would make the
claim match its measurement.

Finding 3 — diagnostic_inventory is a literal asserted against itself, and its measurement is unasserted (principle 8)

"diagnostic_inventory": True is hard-coded at tests/compass/test_capture_real_model.py:1191 and
asserted is True at :1353. It cannot fail and cannot go stale — which is, word for word, the
defect this same PR removed for apply_simulated_tp ("a literal written into the record and asserted
against itself"). One instance of the pattern was fixed; this one remains.

What would carry it is triton_launches, the record field that says raw kernels were reached and
skipped. M9 stops that counter entirely and 9 tests still pass. So the PR body's and 04's
published figure — "33 launches across 3 kernels", at both widths — is held by nothing, and the
diagnostic label it justifies is held by a constant. Asserting the launch inventory
({'kv_indices_generate_kernel': 1, '_fused_qk_norm_single_kernel': 16, '_mrope_qk_kernel': 16},
re-measured today and identical across all four passes) would give both the label and the number a
source.

Same class, smaller: record["declared"] — the whole substitution inventory, including the 8 import
and 14 runner torch.cuda names
the PR body publishes as "holding exactly" — is recorded and read by
no assertion.

6. Observations that are not findings

  • The design record's two self-measurements are exact. T81 claims the constructor pin was measured
    "both ways, with a raise as its first statement (8 of 9 tests fail) and with its zeros_like
    exchanged for a zeros (3 fail, including the constructor's own)". Reproduced here as 8/1
    and 3/6 on a fresh tree. The PR body's repair table (site one 1/8, site two 2/7) also reproduces,
    from different edits than the author's.

  • The "updated, not deleted" instruction holds. P1/P2 settle it as a measurement rather than a
    reading: at ae8b43ae7 the file was 5 tests and replaced CpuGpuBuffer.__init__ wholesale; a raise
    as that constructor's first statement left it 5 passed. The same defect is 8-of-9 at the head.
    The pin was strengthened where it was blind, and the blindness was real.

  • When a specialisation site disappears rather than moves, the failure is
    ValueError: not enough values to unpack (expected 2, got 1) at two, three = record["specialisations"],
    not a message naming the site (M4). It fails — which is what matters — but the reader is sent to an
    unpack, not to copy_to_gpu.

  • The gate table in the PR body is stamped 9fcd6c7bd, not the head 24d75742d. Commit 3 is
    prose-only (the 12 row, the module docstring and three docstrings), so nothing should move — but
    carrying a verdict across a commit is the thing this project has already been bitten by, so it was
    re-run rather than assumed. Re-run at the head, serially, on node 18:

    at 24d75742d (this audit) PR body, at 9fcd6c7bd
    GATE_CPU_RC 0 0
    passed 4566 4566
    skipped 149 149
    xfailed 3 3
    pytest wall 91.13 s 86.11 s

    commit: 24d75742d (stamp), atom: /tmp/jgong5-pr142-pinaudit/base/ATOM/atom/__init__.py,
    gate: 29 files excluded + tests/plugin, all printed before any count was read. The head's
    gate is identical to the one the PR body reports at the previous commit
    , so nothing was lost
    by the citation — but it is now measured at the head rather than carried to it. No failure, so
    nothing to check against the flaky TestTheRegionIsNotCopiedPerChunk class; skip count matches
    the PR's on both sides. GATE_CPU_RC=0, not 98: this diff touches only tests/compass/ and
    two design documents, and matches no path in gpu_gate_triggers.txt (82 entries; none of
    aiter_attention.py, utils/__init__.py, forward_context.py or communication_op.py is
    listed). The box was not quiet — pgrep -c -f pytest read 5 other tenants' processes
    before the run — and the gate still produced its verdict line.

  • Size, calibrated on landed spec/ = 275 AST statements / 495 code / 763 physical.
    Production AST statements added: 0. Test AST statements: 577 total (2.1x), of which 495 are
    the capture driver
    (everything above line 1325) and 82 are the nine test bodies; 1,207 non-comment
    code lines (2.4x); 1,561 physical (2.0x); 43 asserts. black (26.5.1) reports the file unchanged.

  • tests/compass/qwen3_5_27b_config.json is not new in this PR — it came in with compass(backends): the stand-in model's KV geometry, driving ATOM's real block sizing (M1-1) #74. Its blob id is
    the 706cebd746c4b6f2b1d1f892630867acfdfd3df8 the test's comment states, at 4,312 bytes, and the
    sha256 in the test matches the file. Principle 8 satisfied on the fixture.

7. Verdict

Nothing here changes #142's approval status. No finding is a correctness defect, none contradicts a
measurement in the PR, and the two central claims the audit was pointed at — that the pin notices if the
capture stops happening, and that the constructor pin bites where its predecessor was blind — both hold
under reinstatement.

The three findings are each a pin that does not hold a number the PR publishes: the concreteness
detector (M7), the sentinel (M8), the Triton inventory and the diagnostic label (M9). Each is closable
by one or two assertions over data the record already carries. They are offered for CAP-2's author or a
follow-up, not as a re-opening of this PR. Staging under /tmp/jgong5-pr142-pinaudit/ has been removed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant