Skip to content

compass(spec): the refusal two readings earn, and a width table nobody entered - #146

Merged
jgong5 merged 1 commit into
compass/spec-3b-memory-refusalsfrom
compass/spec-3c-refusal-order
Sep 23, 2026
Merged

jgong5 merged 1 commit into
compass/spec-3b-memory-refusalsfrom
compass/spec-3c-refusal-order

Conversation

@jgong5

@jgong5 jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Two defects #136's round-2 review found in atom/compass/spec/, filed as issues
#144 and #145. They cannot run concurrently -- both touch the same import
graph and the same test file -- so they land together.

Stacked on compass/spec-3b-memory-refusals at dc2409574, which is APPROVE at
round 2 and unlanded. Base is that branch, not the integration branch.


Issue #145 -- a width table with no probe entry was an import-time KeyError

PROBE_TABLES is derived from FILLED_BY while validate is being imported,
and the derivation subscripted the table. A width-keyed constant written into
the schema and not into FILLED_BY -- the ordinary shape of a half-finished
change -- therefore raised while the package was being imported.

The fix is the one the review measured: FILLED_BY.get(name, ()). A table with
no entry has no None in it, so it is not a probe table and contributes
nothing; every table that has an entry derives exactly as before.

Measured, with device.runtime_constants.graph_replay_pool_bytes added to
SCHEMA as a WIDTH_TABLE and to nothing else:

at dc2409574 at this head
importing validate KeyError: 'graph_replay_pool_bytes' imports; WIDTH_TABLES has 3 entries, PROBE_TABLES 1
the two spec test files 2 errors during collection, 0 tests run collect and run; 59 failed, 186 passed (test_spec_verbs.py alone: 36 failed, 105 passed)
what the reader is shown the import that failed assert set(FILLED_BY) == {...}, Extra items in the right set: 'graph_replay_pool_bytes'
probe_for("graph_replay_pool_bytes", 1) refuses by name refuses by name, unchanged

Correction. The head-side cell of that row previously read "collects; 1
failed, 139 passed in the spec file"
. That figure is wrong and does not
reproduce: it was read off a -k-filtered run of two tests (1 failed, 1 passed, 139 deselected) and reported as if it were the whole file. The
measured figures are the ones above, and they were re-measured by the review.

The other 58 failures are the added field itself, not the derivation: the field
is required, so every fixture document in both spec files is now missing it and
is refused for that. The row's point is the one the counts still carry -- the
session collects and the failures are assertions, where before there was no
session at all. The failure that is about this change is the one the next row
quotes, in test_the_probe_table_names_every_constant_the_schema_keys_by_width.

The refusal text is the one the package gives everything else it does not know:
`graph_replay_pool_bytes` is not one of the constants measured per tensor-parallel width: [...].

The single-source property still holds. PROBE_TABLES is still derived from
FILLED_BY and is still what both reached(PROBES, PROBE_TABLES) and _probes
read. Patching FILLED_BY["allocator_retained_after_load_bytes"] to
(SINGLE_CARD, MULTI_RANK) and re-importing the module:

PROBE_TABLES = ()
validate(document, tp_widths=(1,))  -> ok=True,  PROBES in not_asked: []  in asked_in_part: []
validate(document, tp_widths=(16,)) -> ok=False, PROBES in not_asked: []  in asked_in_part: []

The question empties with no second list to edit, and the run still counts it as
asked rather than as one it could not reach. That is now a test, so the property
is pinned rather than re-derived by hand.


Issue #144 -- the checks ran rank-major; option 1, the refusal is now the readings'

I took option 1: the refusal is a function of the readings, not of rank
order.
The cost was one statement -- the per-rank loop is now two loops, the
first asking every rank whether its readings describe a card and the second
asking every rank whether free memory was binding. Option 2 would have cost a
sentence and left the instrument with the defect in it: two readings that earn
two different remedies would still be answered by whichever arrived first, and
the docstring would only have stopped promising otherwise. Nothing else in the
module reads rank order, so the traversal was not load-bearing for anything.

Measured, on the review's own case -- one rank that read numbers no card
produced, one whose cache was sized by what was free:

rank A: free 85060000000.0 of 288000000000.0, reserved 2940000000.0, budget 92100000000.0
        impossible='' free_was_binding=True
rank B: free 400000000000.0 of 288000000000.0, reserved -120000000000.0, budget 1.0
        impossible='400000000000.0 bytes free of 288000000000.0' free_was_binding=False
ranks, in this order at dc2409574 at this head
[A, B] rank 0 had 85060000000.0 free against a cache budget of ... (free-binding) rank 1 reports 400000000000.0 bytes free of 288000000000.0, which is not a reading of a card
[B, A] rank 0 reports 400000000000.0 bytes free ... not a reading of a card rank 0 reports 400000000000.0 bytes free ... not a reading of a card

Same two readings, same refusal and same remedy either way. The rank the refusal
names is still positional -- it is the first rank that failed the check that
fired -- because which reading came in where is a question about the sequence
this caller passed and has no other answer. The docstring now says that, and the
"first in that order" sentence is true as written: each check is asked of every
rank before the next check is asked.

ranks.py's own across_ranks still raises a bare ValueError to a direct
caller at tp_width=0. That was ruled SPEC-2's and outside this file set; it is
untouched here and is not what the width refusal in non_torch_across_ranks
covers.


Gates

1. ATOM's suite, as a delta against the stated base. Both sides measured on
node 18 in xiaobizh_n18_cpu, staged by git archive + docker cp, each tree
gated with its own scripts/compass/, COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new,
stamps written from the same rev-parse that produced each archive.

control dc2409574 branch 80a508dff delta
passed 4738 4741 +3
failed 0 0 0
skipped 149 149 0
xfailed 3 3 0
GATE_CPU_RC 0 0 0
wall 38.41 s 42.41 s

The +3 is exactly the three tests this adds. Both sides report gpu: not required from their own .compass-changed stamp -- none of the changed paths
is in scripts/compass/gpu_gate_triggers.txt, so the GPU tier is not owed for
this diff. Skipped is identical on both sides, so the CPU tier's flaky class
(tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk,
three-way pass/skip/fail) landed on the same outcome on both and did not need a
re-run.

2. New CPU-only tests in tests/compass/. Three, all in
tests/compass/test_spec_verbs.py, beside the sections they belong to:
test_which_refusal_a_pair_of_readings_earns_is_not_decided_by_rank_order,
test_a_width_table_no_probe_is_named_for_is_refused_and_not_an_import_error,
and test_a_probe_given_the_width_that_has_none_empties_the_probe_tables. The
last two import a second instance of validate out of the same file under a
name of its own, because the derivation under test runs once, at import, and a
test that wants to watch it run has to import the module rather than call
something in it. The instance the rest of the suite holds is untouched.

3. The named result. Both are in the tables above: the same two ranks in
either order producing the same refusal, beside a width table with no
FILLED_BY entry that is refused by name instead of taking the package down at
import.

4. Review. Dispatched by the lead.

Each fix is pinned by a reversion. Putting the old behaviour back, alone:

reverted result
the two loops back into one test_which_refusal_a_pair_of_readings_earns_is_not_decided_by_rank_order fails on its assertion -- assert 'which is not a reading of a card' in 'rank 0 had 85060000000.0 free against a cache budget of ...'; 1 failed, 140 passed
.get(name, ()) back to the subscript test_a_width_table_no_probe_is_named_for_is_refused_and_not_an_import_error fails, KeyError: 'graph_replay_pool_bytes' raised out of validate.py:159 during the module import it performs; 1 failed, 140 passed

Lint on the changed files: ruff check clean, black --check clean, no line
over 88 characters in either module.

Effort

Neither issue names an envelope. What it cost, measured as a delta against
dc2409574, statements counted as ast.stmt nodes:

AST statement lines physical non-blank
production (memory.py, validate.py) +1 (51->52, 143->143) +23 / -3
tests (test_spec_verbs.py) +42 (704->746) +98 / -1

The production side is one statement because both fixes are one expression and
one loop header; the rest of the 23 lines is the docstring paragraph that was
false and the comment saying why the loops are separate. Comments were not
trimmed to move either number.

What I could not do

  • The whole tests/compass directory does not finish in the local container
    (it was still running at 10 minutes with no output). Everything above the gate
    line was measured on the two spec test files; the gate is the whole suite on
    node 18 and is the number that counts.
  • The docstring endpoints finding from round 2 (below 1.06 / above 13.6,
    measured 1.0621 / 13.5723) is untouched -- it is a separate finding on a
    paragraph this change does not rewrite, and the review explicitly deferred it
    to the touch that reads a second machine.

Staging trees and tarballs under agent_scratch/compass_dev/spec3c_gate and
/tmp/spec3cgates on node 18 are removed; compass-worktrees/spec-3b is
untouched.

🤖 Generated with Claude Code

…y entered

The two per-rank memory checks were asked rank by rank, so a pair of readings
that earns both of them earned whichever one the rank listed first happened to
fail. The two name different things to repair -- a reading to take again,
against a card to take it on -- so a run could be sent after the neighbour when
what was actually wrong is that a rank did not read a card, fix that, and be
refused again. Each check is now asked of every rank before the next check is
asked, which is what the module already said it did: the refusal is decided by
the readings and not by the order they arrived in, and the rank it names is
still the first that failed the check that fired.

The width tables the probe question can speak for are derived from the probe
table while the module is being imported, and the derivation subscripted it. A
width-keyed constant written into the schema and not into the probe table --
the ordinary shape of a half-finished change, and the one the test holding the
two lists together exists to catch -- therefore raised at import and took the
whole package with it, including every caller with no interest in probes. The
test failed at collection, as an import error naming the test session rather
than the term nobody entered. The derivation now reads the table with a default
of no probes: a table with no entry has no hole to report, so it is not a probe
table, and the term is named where a caller asks about it, by the refusal this
package gives everything else it does not know.

The property the derivation exists for is unchanged and is now tested from both
sides: give the width that has no probe a probe and the probe tables empty, the
question has nothing left to ask about, and there is no second list to edit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Read first, in order: the eight principles in atom/compass/design/README.md, AI_DEV_RULES.md, 05_machine_spec_and_probes.md D26 with D24/D25, issues #144 and #145, and #136's round-2 review that found both.

Range. Head 80a508dff, one commit, base compass/spec-3b-memory-refusals re-read from the repository and standing at dc2409574 — the branch the PR names, APPROVE at round 2 and unlanded. File list from git diff --name-only dc2409574...80a508dff, not a glob: 3 files, +130/−4 — atom/compass/spec/memory.py, atom/compass/spec/validate.py, tests/compass/test_spec_verbs.py. The whole of what the diff removes is four lines, and I read all four: the docstring sentence that was false, the #: line that was superseded, the subscript, and one import line that was widened. No assertion anywhere in the test file was deleted, weakened or reworded. No design-doc identifier appears in any added line.

Everything below was executed, on both sides, from worktrees of my own; nothing was read out of the body. compass-worktrees/spec-3b and compass-worktrees/spec-3c were not touched.


Verdict: APPROVE

Both defects are real, both fixes do what they claim, and the stronger option on #144 is the right one. Four items are recorded below; one is a correction to the record and none asks for a code change. Numbered so the developer can answer them in a comment rather than a round 2.


#144 — option 1, verified, and it holds for all four checks

Both rows of the table reproduce exactly, same two readings, run at the parent and at the head:

ranks, in this order dc2409574 80a508dff
[free-binding, impossible] rank 0 had 85060000000.0 free against a cache budget of 92100000000.0 ... rank 1 reports 400000000000.0 bytes free of 288000000000.0, which is not a reading of a card
[impossible, free-binding] rank 0 reports 400000000000.0 bytes free of 288000000000.0 ... not a reading of a card identical to the row above, same remedy

Then I went past the pair the issue names. Two loops fixes two checks against each other; the module has four, so I built rank sets that earn three and four of them at once and ran every permutation of each set, comparing the rule, the refusal text and the remedy with the rank index normalised out:

rank set permutations distinct outcomes at dc2409574 at 80a508dff
impossible + free-binding 2 2 1
free-binding + spread 6 1 1
free-binding + absolute 6 1 1
impossible + absolute 6 1 1
spread + absolute 6 1 1
impossible + free-binding + absolute 6 2 1
impossible + free-binding + spread 6 2 1
free-binding + spread + absolute 6 1 1
all four 24 2 1
negative non_torch + free-binding 6 1 1
negative non_torch + impossible 6 1 1

Five order-dependent sets at the parent, none at the head. The two aggregate checks were never order-dependent — across_ranks reduces with min/max and its spread refusal names no rank — so the defect really was confined to the two per-rank checks, and the two loops close it. The fifth condition inside across_ranks (a non_torch the schema would refuse as a quantity) is also asked of every rank before the spread is computed, so it does not reopen the question from underneath.

The docstring's stated order now matches the code, measured rather than read: impossible → free_was_binding → the ranks against each other → the one number that survives them, which is the order the paragraph gives. The spread + absolute row above is the one that pins the last two, and it returns RANK_AGREEMENT for the spread, not the absolute refusal.

The one residue, and it is the acknowledged one. Two ranks that fail the same check differently — say 400e9 free of 288e9 and -5.0 bytes free — still produce the text of whichever arrived first. Same rule, same remedy, different sentence. The docstring covers this ("the rank a refusal names is still the first that failed the check that fired"), and it is a question about which reading came in where, exactly as the body says.

I agree with taking option 1 over option 2, and the developer's argument for it is the right one: a docstring that stopped promising order-independence would have left two readings earning two remedies answered by arrival order, which is the instrument misdirecting a repair. One production statement for that is the correct trade.


#145 — the defect was real; I reproduced the parent side before checking the fix

The demonstration was re-run from source, not monkeypatched: Field("device.runtime_constants.graph_replay_pool_bytes", Kind.WIDTH_TABLE) added to the schema and to nothing else, on each side.

At dc2409574 — the package goes down at import:

atom/compass/spec/__init__.py:73: from .validate import Validation, validate
atom/compass/spec/validate.py:152: PROBE_TABLES = tuple(
    path for path in WIDTH_TABLES if None in FILLED_BY[path.rsplit(".", 1)[-1]]
KeyError: 'graph_replay_pool_bytes'

and pytest tests/compass/test_spec_verbs.py tests/compass/test_spec_schema.py gives Interrupted: 2 errors during collection, 0 tests run — the diagnostic naming the test session, not the term nobody entered. Exactly what the issue claims, and it is what makes the change worth its lines.

At 80a508dff — the package imports, and everything the body claims about it holds:

WIDTH_TABLES: 3   PROBE_TABLES: 1 ('device.runtime_constants.allocator_retained_after_load_bytes',)
probe_for("graph_replay_pool_bytes", 1) -> REFUSED (NO_DEFAULTS):
  `graph_replay_pool_bytes` is not one of the constants measured per tensor-parallel width: [...]

and the guarding test fails on its own assertion, with the sentence the issue asked for:

assert set(FILLED_BY) == {path.rsplit(".", 1)[-1] for path in WIDTH_TABLES}
E   Extra items in the right set: 'graph_replay_pool_bytes'
tests/compass/test_spec_verbs.py:1557: AssertionError

Finding 1 — the head-side cell of that table does not reproduce (record correction, no code change)

The body's row reads `pytest tests/compass` | 2 errors during collection, 0 tests run | collects; 1 failed, 139 passed in the spec file. The left cell is exact. The right cell is not: the same source edit that produces it gives, at the head,

measured at 80a508dff
tests/compass/test_spec_verbs.py 36 failed, 105 passed
that file plus test_spec_schema.py 59 failed, 186 passed

A width table added to the schema and to nothing else makes every fixture document incomplete, so the refusals are expected and they are the right ones — but they are 36, not 1. I tried the other plausible spelling of the same edit (appending to SCHEMA after BY_PATH rather than to DECLARED) and got 35 failed / 106 passed, so it is not an artefact of where I put the field. 1 failed, 140 deselected is what the guard test gives under -k, and 1 failed, 140 passed is what the reversion gives — both reproduce exactly — so the likeliest reading is that this one cell carries a figure from a neighbouring run.

Nothing in the fix depends on it. The claim the issue asked to see — package imports, probe_for refuses by name, guard fails on its assertion — is all measured above and all true. But it is a number in the durable record whose stated source does not produce it, so please correct the cell (or say which run it came from) before this lands. Principle 8; not a round 2.


The single-source property survives .get, and I checked the case the default could have swallowed

The property, re-run rather than read. With FILLED_BY["allocator_retained_after_load_bytes"] patched to (SINGLE_CARD, MULTI_RANK) and the module re-imported:

PROBE_TABLES = ()
validate(document, tp_widths=(1,))  -> ok=True   PROBES in not_asked: []  in asked_in_part: []
validate(document, tp_widths=(16,)) -> ok=False  PROBES in not_asked: []  in asked_in_part: []

Both widths, exactly as the body states. One list edited, no second list to remember, and the run still counts the question as asked-and-silent rather than as one it could not reach — which is the inversion #136's finding 12 closed, still closed.

The mistype the brief asked about. A FILLED_BY key misspelt against a schema path that is otherwise correct: at the parent the subscript raised at import (loud, but package-down); at the head .get returns (), the table quietly leaves PROBE_TABLES, and validate(..., tp_widths=(1,)) returns ok=True with the probe question silent. So the default does swallow a mistype in the derivation. It does not swallow it in the package: measured at the head,

  • probe_for("allocator_retained_after_load_bytes", 1) refuses by name at the call site, and
  • test_the_probe_table_names_every_constant_the_schema_keys_by_width fails — the set equality catches a mistype from both directions, an extra key in FILLED_BY as readily as a missing one.

That is precisely what the new #: paragraph claims ("the naming is left to probe_for ... and the two lists are held against each other by a test"), so the comment is accurate and the derivation is not weakened in an unguarded way. Accepted, and recorded for whoever next touches this area: the guard is a test, not an instrument in the package, and it compares leaf names rather than full paths — two width tables sharing a leaf under different blocks would still pass it. That is pre-existing and not this PR's to fix.


Reversions — both reproduced, not one

reverted, alone measured
the two loops back into one test_which_refusal_a_pair_of_readings_earns_is_not_decided_by_rank_order fails on its assertion — assert 'which is not a reading of a card' in 'rank 0 had 85060000000.0 free against a cache budget of 92100000000.0 ...'; 1 failed, 140 passed
.get(name, ()) back to the subscript test_a_width_table_no_probe_is_named_for_is_refused_and_not_an_import_error fails with KeyError: 'graph_replay_pool_bytes' raised out of validate.py:159 during the module import it performs; 1 failed, 140 passed

Both figures are the body's, to the digit. Each fix is pinned by a test that fails without it.


Gate 1 — both sides re-measured on node 18

Both trees staged with their own scripts/compass/snapshot.sh, COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, tarball md5 verified on both ends, docker cp into xiaobizh_n18_cpu at a path of my own — nothing written into the shared mount, staging and worktrees removed afterwards. PYTHONPATH confirmed with import atom under each root before any count was read; each gate's own commit: stamp checked; gates run sequentially, never piped.

control dc2409574 branch 80a508dff delta
passed 4738 4741 +3
failed 0 0 0
skipped 149 149 0
xfailed 3 3 0
GATE_CPU_RC 0 0 —
stamp dc2409574 (stamp) 80a508dff (stamp) —
gpu not required not required —
wall 40.18 s 38.95 s

Every reported number reproduces. The +3 decomposes by name, not by count: collected ids diffed between the two trees give 3 only on the branch, 0 only on the control, and the three are exactly

tests/compass/test_spec_verbs.py::test_a_probe_given_the_width_that_has_none_empties_the_probe_tables
tests/compass/test_spec_verbs.py::test_a_width_table_no_probe_is_named_for_is_refused_and_not_an_import_error
tests/compass/test_spec_verbs.py::test_which_refusal_a_pair_of_readings_earns_is_not_decided_by_rank_order

gpu: not required is earned rather than asserted: I read .compass-changed (11 files, the whole five-deep stack) against gpu_gate_triggers.txt and no changed path matches a trigger, so the GPU tier is not owed for this diff.

The CPU tier's flaky class, tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk, did not fire on either side — zero failures, skips identical at 149, GATE_CPU_RC=0 both times — so nothing had to be discriminated against it and no re-run was needed.

ruff check and black --check on the three changed files: clean. Longest line at the head: memory.py 86, validate.py 87, test_spec_verbs.py 88. Nothing over 88.


Gate 2 — the three new tests

CPU-only, fixtures throughout, no device, beside the sections they belong to. reimported_validate is the right mechanism and its docstring says why: the derivation runs once, at import, so a test that wants to watch it run has to import the module. It does not register the second instance, so the one the rest of the suite holds is untouched — I confirmed that by running the full file after each of the new tests and getting the baseline counts back.

Finding 2 — the probe-table test pins one of the two widths the body reports (non-blocking, inline)

test_a_probe_given_the_width_that_has_none_empties_the_probe_tables asserts at tp_widths=(16,) only. The body reports both (1,) and (16,), and (1,) is the more interesting of the two — it is the ok=True case, where "asked and found nothing" and "could not ask" are hardest to tell apart. I measured (1,) myself and it holds, so this is coverage, not a defect.

Finding 3 — one sentence in the docstring is wider than the code (non-blocking, inline)

"each check is asked of every rank before the next check is asked" is exactly true of the first two checks and a little loose about the last two, which are asked of the ranks collectively rather than of every rank. The order claim it supports is correct; only the phrasing generalises past what the code does.


The local suite that never finished — investigated, and it is not this branch

The body says the whole tests/compass directory did not finish in the local gpu_docker container (>10 minutes, no output), so local iteration used only the two spec test files. That is real, it is environmental, and it predates this PR by days. What I found:

  • pytest tests/compass collects in 3.66 s (785 tests) and runs cleanly until it reaches tests/compass/test_capture_collectives.py::test_a_collective_custom_op_is_recorded_and_its_body_never_runs, at 11%. It then stops producing output. My run sat on that one test for the rest of its window.
  • The container's ROCm driver is wedged. timeout 30 rocminfo did not return within 120 s — a timeout cannot kill an uninterruptible process. ps inside the container shows 176 processes in D state, 87 of them rocminfo, the oldest at 13 days, beside 3102 zombies.
  • It is not this branch and not this developer: another agent's run of the same test file on an unrelated worktree has been in D state for 13 h 19 m, and there are stuck pytest tests/compass/... processes from four other worktrees, up to 10 days old.
  • On node 18's CPU container the same directory runs to completion in the gate — all 785 compass tests collect and pass there, inside a 39 s whole-suite run. tests/compass is in neither exclusion list, so the difference is the machine: no device present, so torch never blocks on the wedged one.

Conclusion: a wedged ROCm driver in jgong5_vllm, not a hang in compass/spec-3c-refusal-order. The developer's account is accurate and the decision to measure on node 18 was the right one. Recorded here because it will cost the next agent the same hour: the local container needs a teardown.sh, and until then anything under tests/compass that touches torch device init will hang there with no output and no way to kill it.


Effort — re-measured, reported and not adjudicated

Against dc2409574, AST statements counted as distinct line numbers carrying an ast.stmt, physical counts excluding blanks. Neither issue named an envelope, and which instrument the rule means is an open owner decision; I am not suggesting anything be trimmed.

claimed measured
AST statements, production +1 (51→52, 143→143) +1 (memory.py 51→52, validate.py 143→143) ✓
AST statements, tests +42 (704→746) +42 (704→746) ✓
physical non-blank, production +23 / −3 +23 / −3 ✓
physical non-blank, tests +98 / −1 +98 / −1 ✓
numstat — 14/1, 9/2, 107/1

Every effort figure reproduces. One production statement for a defect that made the instrument name the wrong remedy is, for the record, cheap.


What the next task in this area should watch

  • The guard between FILLED_BY and the schema is a test, and it compares leaf names. With .get it is now the only thing standing between a mistyped table key and a probe question that silently has nothing to ask. Two width tables sharing a leaf name under different blocks would pass it.
  • The rank a refusal names is positional, and so is the reading it quotes. Two ranks failing the same check differently still produce the first one's sentence. Bounded and documented; worth a thought if a refusal ever has to be diffed between runs.
  • ABSOLUTE_LIMIT's endpoints (below 1.06 / above 13.6 against measured 1.0621 / 13.5723) are untouched here, correctly — deferred to whoever reads a second machine.
  • ranks.py's bare ValueError at tp_width=0 to a direct caller is still there and still SPEC-2's, outside this file set.
  • The local gpu_docker container needs a teardown before anyone tries tests/compass there again.

No blocking issues. Finding 1 is a correction to the PR body and asks for no code change; findings 2 and 3 are non-blocking and inline. Nothing to escalate, and no label applied.

all, whether the cache that rank built was sized by its own budget, the ranks
against each other, and the one number that survives them. A contaminated run
can fail more than one, and the one raised is the first in that order.
can fail more than one, and **each check is asked of every rank before the next

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 3 (non-blocking): this sentence is a little wider than the code.

"each check is asked of every rank before the next check is asked" is exactly true of the first two checks — and that is the whole of the fix, so the claim it supports is right. It is loose about the last two: the spread and the absolute limit are asked of the ranks collectively, not of every rank, because across_ranks reduces with min/max before either fires.

I checked that this does not weaken the guarantee. Running every permutation of rank sets that earn three and four of the checks at once, the head gives one outcome per set where the parent gave two, and the aggregate pair was never order-dependent to begin with — the spread refusal names no rank at all. So the order the paragraph states is the order the code executes, verified end to end; only the phrase generalises past what the last two checks do.

No change asked for. Recorded so a reader who takes the sentence literally and goes looking for a per-rank loop behind checks three and four knows why there isn't one.

)
under_test = reimported_validate("validate_with_every_width_filled")
assert under_test.PROBE_TABLES == ()
checked = under_test.validate(merged().document, tp_widths=(16,))

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 2 (non-blocking): this pins one of the two widths the PR body reports.

The body states the property at both tp_widths=(1,) and tp_widths=(16,); the test asserts it at (16,) only. (1,) is the more interesting of the two — it is the ok=True case, where "asked the question and found nothing to say" and "could not ask the question" are hardest to tell apart, which is exactly the inversion this derivation exists to keep closed.

I measured it at the head rather than asking for it:

PROBE_TABLES = ()
validate(document, tp_widths=(1,))  -> ok=True   PROBES in not_asked: []  in asked_in_part: []
validate(document, tp_widths=(16,)) -> ok=False  PROBES in not_asked: []  in asked_in_part: []

Both hold, so this is coverage rather than a defect, and one extra validate(...) line would close it if you happen to be touching the file. Not worth a round 2 on its own.

#: are held against each other by a test.
PROBE_TABLES = tuple(
path for path in WIDTH_TABLES if None in FILLED_BY[path.rsplit(".", 1)[-1]]
path for path in WIDTH_TABLES if None in FILLED_BY.get(path.rsplit(".", 1)[-1], ())

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Accepted, with one measurement for the record — the default swallows a mistyped key as well as a missing one, and the test is what catches it.

The review brief for this PR asked whether .get(name, ()) could weaken the derivation by answering a misspelt FILLED_BY key the same way it answers an absent one. It can, in the derivation itself: with the key mistyped against an otherwise-correct schema path, the head imports cleanly, the table quietly leaves PROBE_TABLES, and validate(document, tp_widths=(1,)) returns ok=True with the probe question silent — where the parent's subscript raised at import.

It does not go unnoticed in the package. Measured at this head, with the key mistyped:

probe_for("allocator_retained_after_load_bytes", 1)
  -> REFUSED: `allocator_retained_after_load_bytes` is not one of the constants measured per tensor-parallel width
set(FILLED_BY) == {leaf of p for p in WIDTH_TABLES}  ->  False   (the guard fires)

So both halves of what this comment promises are true: probe_for names the term at the call site, and test_the_probe_table_names_every_constant_the_schema_keys_by_width catches the mistype from either direction — an extra key as readily as a missing one. The comment is accurate as written and I am asking for no change.

Two things for whoever next touches this area, neither introduced here: the guard is a test rather than an instrument in the package, and it compares leaf names rather than full paths, so two width tables sharing a leaf under different blocks would pass it.

@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Pin re-verification of #146 — not a new review cycle

The existing APPROVE stands. Nothing below changes it. This is the reinstatement
audit the board did not require when #146 was reviewed: every test the PR presents as
pinning a defect was checked by putting the defect back and watching the suite. All
three pins bite, and the ordering itself is pinned — not merely that some refusal
occurs.
Two recorded observations follow, neither blocking and neither asking for a
code change.

Principles referenced: 6 (refuse rather than fall back — what every pin here is
about), 7 (no aggregate without its decomposition — every count below is named by
node id), 8 (every claim carries its measurement).


Base, and whether #146 is stale

Not stale. Determined from the repository and the API, not from the body:

head of compass/spec-3c-refusal-order (API git/ref) 80a508dffcd8344fcce084ee6976246e6c918212
.head.sha on the PR the same
git rev-parse 80a508dff^ dc2409574b9f969cbef546d7a5a3c4bd75dc5b38
git merge-base 80a508dff dc2409574 dc2409574b9f969cbef546d7a5a3c4bd75dc5b38
head of compass/spec-3b-memory-refusals (API git/ref) dc2409574b9f969cbef546d7a5a3c4bd75dc5b38

So the head's parent is the base branch's current tip, and the merge-base is that
same commit: the branch sits directly on its base with nothing in between, and its head
has not moved past the approval. mergeable_state: clean, draft: true (left as it
was), no labels. Gated against dc2409574 as the control.


How it was measured

Both trees staged with their own scripts/compass/snapshot.sh,
COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new, tarball md5 verified on both
ends (e5e6d470… branch, 0371b59f… control), docker cp-style pipe into
xiaobizh_n18_cpu at /tmp/xiaobizh_pr146pin/ — nothing written into the shared mount,
nobody else's staging touched. atom.__file__ asserted under the staged root and printed
before every count; __pycache__ cleared before every run; every run under
timeout -k 10; no run piped. Mutations applied one at a time, the tree restored
from pristine copies and md5-verified between each — no concurrent mutation against
the staged tree. Every mutation except the two whole-file reverts is
line-count preserving. The reverted code was recovered with
git show dc2409574:<path>, never reconstructed from the diff.

Baselines reproduce the PR's own figures exactly:

control dc2409574 branch 80a508dff
gate 4738 passed, 0 failed, 149 skipped, 3 xfailed 4741 passed, 0 failed, 149 skipped, 3 xfailed
GATE_CPU_RC 0 0
stamp dc2409574 (stamp) 80a508dff (stamp)
gpu not required not required
tests/compass alone — 785 passed in 7.77 s

The flaky class tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk
did not appear in any run, green or red.

Line-drift control: a single null comment inserted above import copy in
tests/compass/test_spec_verbs.py, shifting all 1758 lines below it →
785 passed, PYTEST_RC=0. Nothing in this suite is holding a line number, so every
red below is the mutation and not the drift.


The table

Counts are tests/compass (785 tests) on the branch tree unless the row says gate.

# Pin (node id) Defect reinstated Published / baseline Reinstated Verdict
A1 test_which_refusal_a_pair_of_readings_earns_is_not_decided_by_rank_order memory.py restored verbatim from dc2409574 — the two per-rank checks back in one loop (247→234 lines) 785 passed · gate 4741 1 failed, 784 passed · gate 1 failed, 4740 passed, GATE_CPU_RC=1 bites
A2 same pin the ordering partial: the two loops kept, their order swapped — free_was_binding asked of every rank before impossible 785 passed 2 failed, 783 passed bites
A3 same pin the impossible refusal hard-codes rank 0 instead of the rank that failed 785 passed 1 failed, 784 passed bites (only this pin)
A4 same pin the two aggregate checks (across_ranks, the absolute ceiling) moved ahead of the two per-rank checks — the narrowing order the docstring states, inverted 785 passed 7 failed, 778 passed bites
A5 same pin the free_was_binding check made dead (if False and …) 785 passed 1 failed, 784 passed pin blind; suite catches
B1 test_a_width_table_no_probe_is_named_for_is_refused_and_not_an_import_error validate.py restored verbatim from dc2409574 — FILLED_BY[leaf] subscript back (413→406 lines) 785 passed 1 failed, 784 passed bites
B2 same pin the wrong default: .get(leaf, (None,)), so a table nobody entered looks like a table with a hole 785 passed 1 failed, 784 passed bites
B3 pins B and C PROBE_TABLES forced always empty (the counted-quantity-to-zero probe) 785 passed 4 failed, 781 passed both pins blind; suite catches
C1 test_a_probe_given_the_width_that_has_none_empties_the_probe_tables the derivation replaced by a hard-coded second list equal to today's value 785 passed 1 failed, 784 passed bites (only this pin)
C2 same pin _unreached reports an empty reads list as "could not be asked at all" — the inversion the derivation exists to prevent 785 passed 1 failed, 784 passed bites (only this pin)
D1 line-drift control a null comment shifting every line below it 785 passed 785 passed green, as required
D2 reimported_validate helper the helper registers its second instance as atom.compass.spec.validate, which its docstring says it does not do 785 passed 785 passed docstring claim held by nothing (observation 2)

Score: 3 pins, 3 bite. Zero inert, zero vacuous, zero names-the-wrong-defect. Two
blind spots found (A5, B3), both covered by pre-existing siblings, and one unguarded
docstring sentence (D2).


Named failures and the assertions they fired on

"It fails" is not a result, so here is each one.

A1 — tests/compass/test_spec_verbs.py::test_which_refusal_a_pair_of_readings_earns_is_not_decided_by_rank_order

E  assert 'which is not a reading of a card' in 'rank 0 had 85060000000.0 free against a
   cache budget of 92100000000.0, so what was free on the card set the cache siz...'

That is the identity assertion, not a bare "something was refused": the pin fires
because the pair earned the other refusal. Gate-level the same revert gives
1 failed, 4740 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=1, stamp 80a508dff.

A2 — the ordering partial, and the reason to say the ordering is pinned. Swapping the
two loops leaves the property the docstring's first clause names intact ("each check is
asked of every rank before the next"), keeps both refusals reachable, keeps the loop
count, the line count and every rule — and changes only which of the two a
contaminated pair earns. The pin reddens on the same identity assertion above, and a
pre-existing sibling goes with it:

FAILED …::test_which_refusal_a_pair_of_readings_earns_is_not_decided_by_rank_order
FAILED …::test_a_rank_whose_readings_cannot_be_one_card_is_refused_and_named[overrides1--1.0 bytes free]
E  assert '-1.0 bytes free' in 'rank 0 had -1.0 free against a cache budget of 1.0, …'

This is the answer to the question the subject raises. The pin asserts which
refusal, in what order — "which is not a reading of a card" in …what together with
"set the cache size" not in …what — not merely that a SpecRefusal was raised.

A3 — the rank a refusal names.

E  assert 'rank 1' in 'rank 0 reports 400000000000.0 bytes free of 288000000000.0, which is
   not a reading of a card (…)'

784 of 785 still pass: this pin is the only test in the suite holding the rank index in
that refusal's text.
The docstring's residual claim ("the rank a refusal names is still
the first that failed the check that fired") therefore has exactly one guard, and it is
this one.

A4 — the narrowing order across all four checks. Seven tests go red, the pin among
them, on its rule assertion rather than its text assertion:

E  assert <Rule.RANK_AGREEMENT: 'ranks of one group measure one machine'>
       is <Rule.DEVICE_WIDE: "a device-wide reading is the whole card's, not one engine's">

plus test_a_rank_whose_free_memory_was_binding_is_refused_and_named and all five
parametrisations of test_a_rank_whose_readings_cannot_be_one_card_is_refused_and_named.
So the whole stated order — per-rank before collective — is guarded, not just the pair.

A5 — what pin A does not notice. Kill the free_was_binding branch outright and pin A
passes: it asserts the impossible refusal fires first, which it still does. The
suite catches it through the pre-existing
test_a_rank_whose_free_memory_was_binding_is_refused_and_named (1 failed, 784 passed).
Correctly divided work — the new pin owns the ordering, the old one owns the check's
existence — recorded because a reader of the new test alone could over-read it.

B1 — test_a_width_table_no_probe_is_named_for_is_refused_and_not_an_import_error,
and it fails as an error raised out of the import the test itself performs, at the
right frame:

tests/compass/test_spec_verbs.py:1695: in test_…
    under_test = reimported_validate("validate_with_a_table_no_probe_is_named_for")
tests/compass/test_spec_verbs.py:1679: in reimported_validate
    loaded.loader.exec_module(module)
atom/compass/spec/validate.py:151: in <module>
>   path for path in WIDTH_TABLES if None in FILLED_BY[path.rsplit(".", 1)[-1]]
E   KeyError: 'graph_replay_pool_bytes'

(Line 151 in the reverted file; the review's validate.py:159 is the head's numbering for
the same statement.) 1 failed, 784 passed — the same 1-of-the-file the body reports.

B2 — the realistic wrong default. (None,) instead of () is the plausible typo, and
the pin catches it on the assertion written for exactly that:

E  AssertionError: assert 'device.runtime_constants.graph_replay_pool_bytes' not in
   ('device.runtime_constants.allocator_retained_after_load_bytes',
    'device.runtime_constants.graph_replay_pool_bytes')

B3 — the one both new pins miss. Force the derivation to produce () for everything:

FAILED …::test_a_missing_entry_no_probe_fills_is_two_refusals_and_not_one
FAILED …::test_the_check_does_not_agree_a_probe_exists_for_a_width_below_one
FAILED …::test_the_probe_question_reads_only_the_tables_a_probe_can_fall_short_on
FAILED …::test_the_probe_question_says_it_could_not_be_asked_when_its_table_is_gone
4 failed, 781 passed

Neither of the two new tests is in that list. Both are blind to it, and for the shape
the audit brief names: pin C's expected value is the constant the defect returns
(assert under_test.PROBE_TABLES == ()), and pin B only asks that one path be absent
from a tuple, which an always-empty tuple satisfies. Four pre-existing tests hold it, so
the suite is covered — this is a note on how far the new tests reach, not a gap in the
tree.
Observation 1 below.

C1 — the defect pin C exists for. Replace the comprehension with a literal tuple that
is correct today:

E  AssertionError: assert ('device.runt..._load_bytes',) == ()
E    Left contains one more item: 'device.runtime_constants.allocator_retained_after_load_bytes'
1 failed, 784 passed

Only this pin fires. A hard-coded second list passes every other test in the tree, so
"no second list to remember" is guarded by exactly one test, and it is this one.

C2 — the inversion, and the sharpest single result here. Make _unreached report an
empty reads list as "could not be asked at all" (if not absent and reads:), which is
the record-confusion the #: paragraph is written against:

E  assert not True
   + where True = any(<genexpr … test_a_probe_given_the_width_that_has_none_empties_the_probe_tables>)
1 failed, 784 passed

Only this pin fires. The property "a run that asked and found nothing is not recorded
as a run that could not ask" — principle 7's distinction, and #136's finding 12 — is held
in this tree by this test alone.


Observation 1 — the probe pins cannot see an always-empty derivation (non-blocking, no code change)

Measured at B3 above. Both new probe-table tests stay green when PROBE_TABLES is forced
empty unconditionally; four older tests catch it. The reviewer's Finding 2 already noted
that pin C asserts at tp_widths=(16,) only; I probed whether (1,) would have closed
this and it would not — _reach calls reached(PROBES, PROBE_TABLES) with the same
PROBE_TABLES and the same resolved document at either width, so the probe question's
not_asked/asked_in_part result is width-independent by construction and only the
incidental assert not checked.ok differs. The (1,) gap is coverage, as the review
said; the always-empty gap is a different axis and is held elsewhere.
Nothing to change.

Observation 2 — one helper docstring sentence is held by nothing (non-blocking, no code change)

reimported_validate's docstring says it "does not register it, so the instance the rest
of the suite is holding is the one it started with". Making the helper do exactly what
that sentence disclaims — registering its second instance as atom.compass.spec.validate
— leaves 785 passed. Inert today for a real reason: atom.compass.spec binds its
names with from .validate import … at package import, so a later sys.modules swap
rebinds nothing in this suite. So the sentence is accurate about the code and latent
rather than wrong. Recorded under principle 8 for whoever next adds a test that
re-imports or reloads that module: it is the docstring, not an assertion, that is keeping
the second instance out of the way.


Verdict

Three pins presented, three bite, each on a named assertion, each reproducing the
counts already in the record. The ordering is pinned by identity and by position, not by
"a refusal occurred": swapping the two checks (A2), hard-coding the rank (A3) and moving
the collective checks ahead of the per-rank ones (A4) each redden the pin, and A3 and C2
and C1 are each held by one test in the whole tree — this PR's. The line-drift control
is green. The APPROVE stands unchanged, and the one item already open against the PR
is the review's Finding 1, the body-table cell that does not reproduce — unaffected by
anything here.

Staging under /tmp/xiaobizh_pr146pin/ and the host worktrees have been removed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant