Skip to content

compass(spec): make four spec pins name the refusal they are about (#237) - #263

Merged
jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-237-spec-pins
Sep 23, 2026
Merged

jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-237-spec-pins

Conversation

@jgong5

@jgong5 jgong5 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Refs #237: the four inert pins. It does not close the issue, because F3 and the review-record corrections are not done here (see the end).

No blocking issues. Tests only: one file, tests/compass/test_spec_verbs.py. No production file changes. atom/compass/spec/machine.py and rules.py are not touched.

What changed

Each of the four pins passed with its defect reinstated because it accepted a refusal a neighbouring check also produces, or could not observe the path at all. Each now asserts which refusal fired: its rule and a piece of its what only that refusal writes (never the Rule definition string, which nothing holds, see #249).

# test change
F2 test_a_width_with_no_ranks_is_refused_rather_than_reduced asserts "tensor-parallel width {w!r} has no ranks to read a card on" in what. The cross-rank length check in ranks.py also refuses -2 with Rule.RANK_AGREEMENT and puts -2 in its message, which satisfied both old assertions.
F6 refused_by_condition() / test_each_condition_in_the_check_set_is_earned_by_a_spec each subject now carries (checked, rule, text), and the test asserts rule and text on the same refusal. The probe subject's NO_DEFAULTS was also earned by the missing-width refusal.
F1 test_readings_no_card_could_produce_are_refused_though_their_difference_is_not asserts "rank 0 reports 400000000000.0 bytes free of 288000000000.0" in what and the remedy's "take them again and find out". That reading also has a negative reserve, the next branch down, which answered when the free > total branch was removed.
F4 new test_the_probe_question_is_asked_of_the_tables_it_says_it_reads wraps validate's probe_for in a recorder and asserts the probe question asks exactly [(name, 16) for PROBE_TABLES], non-empty. Today a walk over every width table adds only calls that return a probe, so no refusal can show it; what is asked can. The fixture's precondition (width 16 missing from every width table, so any walk would reach the probe) is asserted, not assumed.

Named result: each pin red by name with its defect back in (principle 8)

Instrument: python -m pytest tests/compass -q -p no:randomly -p no:cacheprovider, node 18 xiaobizh_n18_cpu, one staged tree per side, mutants serial, each line-count preserving and restored with a byte-equality check after the run. Control = integration tip e1da5404e (my tests absent, i.e. my fix reverted); branch = head 00a4d1479. Counts are failed / passed. Line numbers are at the head.

mutant (production, one line) control e1da5404e branch 00a4d1479 failing node on branch, and the assertion
(none), before and after the battery 0 / 1088 0 / 1089
null control: one comment line in memory.py reworded, same length in lines 0 / 1088 0 / 1089 no drift guard fires
F4 _probes: for path in PROBE_TABLES: → WIDTH_TABLES: 0 / 1088 inert 1 / 1088 test_the_probe_question_is_asked_of_the_tables_it_says_it_reads, :1743 assert asked == [...]: [('driver_and...d_bytes', 16)] == [('allocator_...d_bytes', 16)] (the extra reserve-table call)
F2 memory.py: if tp_width < 1: → < -99999: 1 / 1087 ([0] only, by ValueError) 2 / 1087 test_a_width_with_no_ranks_is_refused_rather_than_reduced[-2], :1624: assert 'tensor-parallel width -2 has no ranks to read a card on' in 'non_torch has 0 reading(s) for the -2 ranks of this group'. [0] still fails with ValueError: max() iterable argument is empty
F6 validate: refusals += _probes(...) → += [] 2 / 1086 (not the named test) 4 / 1085 test_each_condition_in_the_check_set_is_earned_by_a_spec[whether a probe fills the widths this deployment uses and nobody measured], :909 (the any(rule is … and text in what) assertion). Its message lists the one refusal present, the missing-width one. Also the new F4 test, :1742 the probe question asked nothing. The two nodes that already bit still do.
F1 impossible: free > total → free > 9.9e99 * 1.0 1 / 1087 (parametrised sibling only) 2 / 1087 test_readings_no_card_could_produce_are_refused_though_their_difference_is_not, :1597: assert 'rank 0 reports 400000000000.0 bytes free of 288000000000.0' in 'rank 0 reports -120000000000.0 bytes reserved, which is not a reading of a card (...)'
F1 converse: the reserved < 0 branch disabled 1 / 1087 1 / 1088 the same [overrides3-…] node on both. The new F1 assertion does not demand the negative-reserve branch.

Every added assertion was seen to fail, and the control column is the "fix reverted" run: with my tests absent, F4 is fully inert and F2/F6/F1 fail only on nodes other than the one each finding names.

The other per-condition texts (run at addd347dc)

spec/ and the test file are byte-identical between addd347dc and e1da5404e (git diff --stat empty on those paths).

mutant control branch reading
_widths output removed 9 / 957 10 / 957 +1 is the new F4 test, whose fixture precondition counts the width refusals
strict stack refusal never raised 4 / 962 4 / 963 same nodes; the [STACK] parametrisation already bit on its rule
_transfers output removed 4 / 962 4 / 963 same nodes; the [TRANSFERS] parametrisation already bit on its rule

So for WIDTHS, STACK and TRANSFERS the new per-condition text adds no discrimination against these removals. It is there so the helper has one shape, and so a subject that also earns a neighbour's rule is caught. For MISSING and DERATES I ran no producer-removal mutant: the survey and _missing are exercised by dozens of tests, so such a mutant tells you nothing about this parametrisation.

Three runs, same pattern. The core battery was run at three tips, because the tip moved twice during the work (#163, #256, #183, #196, then #208, #172, #164; #183 and #196 touch spec/). The mutated lines are unchanged across all of them. At each tip F4 went 0→1, F2 1→2, F6 2→4 and F1 1→2, on the same nodes:

tip → head baseline control / branch
51854571d → 0dce57806 882 / 883
addd347dc → de5f0c47a 966 / 967
e1da5404e → 00a4d1479 1088 / 1089

Instrument is sound at this head (checked, not assumed)

The brief says the two test_spec_* files are the only importers of atom.compass.spec. At this head there are four: tests/compass/test_kv_simulated_connector.py:41 and tests/compass/test_memory_readings.py:50 (landed in #164) also import MachineSpec, SpecRefusal. Under every mutant above, neither file had a failing node, and both are in tests/compass, so pytest tests/compass is still a complete denominator for any spec/ mutation. (grep -rlE "compass\.spec|from \.+spec|compass import .*spec" --include=*.py atom tests scripts, excluding the package itself.)

Gate (principle 7)

scripts/compass/gate_cpu.sh from each tree's own scripts/compass, staged by git archive + docker cp into xiaobizh_n18_cpu at /tmp/issue237gates-xbz/, tarball md5 matched on all three hops, .compass-commit/.compass-changed written from the same rev-parse, atom.__file__ asserted under the staged root before any count, timeout -k 10 2400, not piped, one gate at a time. Nothing was written into the shared mount.

control e1da5404e (tip) branch 00a4d1479
commit: stamp printed by the gate e1da5404e (stamp) 00a4d1479 (stamp)
atom: /tmp/issue237gates-xbz/control/ATOM/atom/__init__.py /tmp/issue237gates-xbz/branch/ATOM/atom/__init__.py
result 5044 passed, 149 skipped, 3 xfailed 5045 passed, 149 skipped, 3 xfailed
GATE_CPU_RC 0 0
gpu: not required not required

Delta +1 passed, 0 skipped, 0 xfailed, 0 failed, compared by node id from each side's --junitxml:

  • only on the branch: tests/compass/test_spec_verbs.py::test_the_probe_question_is_asked_of_the_tables_it_says_it_reads, passed
  • only on the control: none
  • outcome changed: none

The edited tests keep their node ids, so they do not show up in the delta. The flaky class TestTheRegionIsNotCopiedPerChunk passed on both sides.

The tip moved twice while I worked, so the same pair was gated at each tip, always with the same single-node +1 and GATE_CPU_RC=0 on both sides:

tip control branch head
51854571d (the brief's figure) 4838 / 149 / 3 0dce57806: 4839 / 149 / 3
addd347dc (after #163, #256, #183, #196) 4922 / 149 / 3 de5f0c47a: 4923 / 149 / 3
e1da5404e (after #208, #172, #164) 5044 / 149 / 3 00a4d1479: 5045 / 149 / 3

The base moved again after the PR was opened (#212, #260). #212 touches spec/machine.py, spec/rules.py and test_spec_schema.py. The only behaviour it changes is the version refusal's rule, SHAPE → VERSION, and none of the texts or rules asserted here involve it. I did not rebase. Instead I gated the merge result against the new tip, building the merge commit with git merge-tree + commit-tree, so no worktree was touched:

control 815a08679 (tip) merge of tip + 00a4d1479 (8c232cd9b)
result 5056 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0 5057 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0
delta by node id + test_the_probe_question_is_asked_of_the_tables_it_says_it_reads only; nothing removed, no outcome changed

Another agent's gate_cpu.sh was running in the same container during this pair (both runs took about 3m10s, not 2m35s). The flaky timing class passed on both sides anyway.

Size

AST statements / code lines / physical lines. The counter reproduces spec/ = 275 / 495 / 763 at 669dc3f9d.

tip e1da5404e head 00a4d1479 delta
production: atom/compass/spec/ 815 / 1424 / 2602 815 / 1424 / 2602 0 / 0 / 0
test: tests/compass/test_spec_verbs.py 765 / 1247 / 1790 778 / 1301 / 1872 +13 / +54 / +82

What these assertions still would not notice

  • F4. The recorder replaces validate's module-level probe_for. A refactor that stops calling it by that name (inline FILLED_BY lookup, say) records nothing and the test fails loudly (asked nothing), so the gap does not open silently. What it does not check is a walk that filters WIDTH_TABLES inline to the tables with a hole: same calls, same result, so there is no defect for it to catch. When the width-1 hole in FILLED_BY is filled, PROBE_TABLES becomes empty and this test goes red on assert asked. That red is correct, because at that point no fixture can observe the walk, but it will need rewriting then. The per-width skip (if width in resolved[path]) is not exercised here; the audit's m28 measured it held by 9 other nodes.
  • F1. The narrative test now also pins the order of two branches: for a reading that is both free > total and reserved < 0, the first one is named. If someone reorders those branches on purpose, this test goes red. The parametrised sibling still holds each branch on a fixture that reaches only that branch.
  • F2. Only what is pinned, not the remedy. A guard that refuses under another rule is caught by the existing rule is RANK_AGREEMENT.
  • F6. Each subject is checked for its own refusal, not that it earns only that one. The two neighbours the audit named are handled like this: test_a_width_a_probe_would_fill_is_not_reported_as_a_hole_in_the_probes (a negative needle) is unchanged, but the new F4 test now asserts that the same width-16 fixture was asked of the probe and got an answer, which separates "asked and found nothing" from "never asked" there. test_a_spec_that_carries_the_hand_measured_entry_is_not_refused_for_it asserts only ok. Its claim is an absence: with the entry present, the walk skips it before any probe is consulted, so from the outside "asked and skipped" looks exactly like "never asked". I did not add an assertion that pretends to hold it.
  • Review records. compass(spec): four inert pins, and three findings closed by tests that do not close them #237 asks that the review records of findings 4 and 12 be corrected, not just the tests. That is a comment on SPEC-3b - the absolute memory refusal, and the probe table hole that stays a hole #136, and I leave it to the reviewer or lead. The table above is the evidence it would cite.

Not in scope

Principles: 6 (a refusal names its reason, so a test of a refusal should check that reason, not just that some refusal fired), 7 (the gate delta is decomposed by node id), 8 (each claimed pin carries the mutation that shows it).

🤖 Generated with Claude Code

)

Four tests in the spec package passed with the defect they were written
around put back, because each accepted a refusal a neighbouring check
also produces, or could not see the path at all:

- the no-ranks width test now asserts the width's own refusal text; at a
  negative width the cross-rank length check used to satisfy it with the
  same rule and the width in its message
- the per-condition "earned by a spec" helper now carries, per subject,
  a piece of text only that condition's refusal writes, asserted on the
  same refusal as the rule; the probe subject's rule was also earned by
  the missing-width refusal
- the impossible-readings test now asserts it was refused for free
  memory beyond the card, not for the negative reserve it also carries
- a new test records what the probe question asks, so a walk over every
  width table rather than the ones the question registers is observable
  while the extra table is one that never refuses

Tests only; no production file changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
if "was not measured at tensor-parallel width 16" in refusal.what
]
assert len(unmeasured) == len(WIDTH_TABLES) > len(PROBE_TABLES)
assert asked, "the probe question asked nothing"

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review is agent-authored.

Non-blocking (principles 3 and 8): this line is the only reason the test turns red when the width-1 hole in FILLED_BY is filled, and the PR body's reason for accepting that red does not hold.

My ruling on brittleness. The recorder observes a call to a module-level name, which makes this a white-box pin. I accept it. The property is about what the walk asks: _reach registers PROBE_TABLES as what PROBES reads. Under today's FILLED_BY, no refusal can show which tables were walked. That is F4 on the control: 0 failed / 1151 passed. So recording what is asked is the only observation available, and principle 3 does not ask for a more elaborate one.

The claim that fails. The body says that once the hole is filled, "no fixture can observe the walk". Once the hole is filled, PROBE_TABLES is (). A walk over WIDTH_TABLES still asks two questions, and the equality on the next line sees them. I measured this on the merged tree (tip 175739f87 + 00a4d1479, node 18), using line-count-preserving mutants and running tests/compass:

mutant result
this line changed to assert asked or not PROBE_TABLES, ... 0 failed / 1152 passed
the same change + _probes walks WIDTH_TABLES 1 failed: this test, :1743 [('driver_and...d_bytes', 16)] == [('allocator_...d_bytes', 16)]
the same change + hole filled ((SINGLE_CARD, MULTI_RANK)) this test passes. The 5 other reds are the tests that are about the hole.
the same change + hole filled + _probes walks WIDTH_TABLES this test fails at :1743: [('driver_and...d_bytes', 16)] == []
this line as written + hole filled this test fails at :1742: the probe question asked nothing

With assert asked or not PROBE_TABLES, the test keeps today's result and keeps catching the wrong walk after the hole is filled. It also stops going red for a change in another module that it has no quarrel with. The case the non-empty check exists for, a refactor that stops calling probe_for, only matters while PROBE_TABLES is non-empty. With PROBE_TABLES empty, "asked nothing" is the correct answer.

The cost of the current form is small. Filling the hole already reddens 4 tests at the tip, and this PR brings that to 6. Still, please either make the one-word change or correct the "still would not notice" paragraph of the body.

# be refused by the same rule on its own. What is named is the free memory
# the card cannot hold, so a check that stopped asking that is seen here
# rather than covered by its neighbour.
assert "rank 0 reports 400000000000.0 bytes free of 288000000000.0" in (

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review is agent-authored.

Non-blocking (principles 3 and 6): this assertion pins the order of impossible's checks, and the comment above it gives the wrong reason if it ever turns red.

My ruling: the order is an accident, not a contract. DeviceMemory.impossible is a chain of first-match checks. Its docstring ("What about these readings no single card could have produced") states no precedence between free > total and reserved < 0, and nothing else holds the order. I swapped those two branches with a line-count-preserving mutant:

tip 175739f87 merged tree
free > total and reserved < 0 branches swapped 0 failed / 1151 passed 1 failed: this test, :1597 ... in 'rank 0 reports -120000000000.0 bytes reserved, which is not a reading of a card (...)'

So this PR adds a new constraint on the order. The constraint cannot be avoided in this test. The test's premise is a plausible difference (non_torch = total - free - reserved > 0) beside free > total, and that forces reserved < 0. So this fixture always carries both impossibilities, and the only way it can witness the free > total branch is to fix which of the two is named. The branch itself is already witnessed on a fixture that reaches only that branch: test_a_rank_whose_readings_cannot_be_one_card_is_refused_and_named[overrides2-...] is red with the branch removed, both on the tip and on the merged tree.

I accept the order pin, because the body discloses it. The code comment does not, though. Someone who reorders the branches on purpose will read "a check that stopped asking that is seen here" and look for a check that was removed. A cheap fix is to say what is actually held. For example: "Both impossibilities are present, and which one impossible names first is fixed here on purpose; the branch on its own is held by the parametrised test above." The alternative is to drop this assertion and let the review record on #136 name [overrides2], which #237 asks for anyway.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

This review is agent-authored. It is the first review of PR #263, at head 00a4d14799eaef7187d0ed7096767dcf51083791. I read the eight design principles and AI_DEV_RULES.md before reviewing.

Verdict: APPROVE at 00a4d14799eaef7187d0ed7096767dcf51083791. No blocking issues.

Each of the four pins turns red by name with its defect back in, and fails on the assertion that names that defect. The merged tree gates at the tip's count +1, with the delta decomposed by node id. The diff has no design-doc references. There are two non-blocking inline findings (:1742 and :1597) and one body correction, listed below.

1. Each defect reinstated (principle 8)

Method. Node 18, xiaobizh_n18_cpu. I ran python -m pytest tests/compass -q -p no:randomly -p no:cacheprovider, one mutant at a time. Every mutant was line-count-preserving and restored with a byte-equality check.

The two trees:

atom.__file__ was asserted under each staged root before any count. Results are shown as failed / passed.

mutant control branch failing node on the branch, and the assertion
none (before and after) 0 / 1151 0 / 1152
null: a comment in memory.py reworded 0 / 1151 0 / 1152 no drift guard fires
_probes walks WIDTH_TABLES 0 / 1151, inert 1 / 1151 test_the_probe_question_is_asked_of_the_tables_it_says_it_reads, :1743: [('driver_and...d_bytes', 16)] == [('allocator_...d_bytes', 16)]
if tp_width < 1 → < -99999 1 / 1150 ([0] only) 2 / 1150 test_a_width_with_no_ranks_is_refused_rather_than_reduced[-2], :1624: 'tensor-parallel width -2 has no ranks to read a card on' in 'non_torch has 0 reading(s) for the -2 ranks of this group'
refusals += _probes(...) → += [] 2 / 1149 4 / 1148 test_each_condition_in_the_check_set_is_earned_by_a_spec[whether a probe fills ...], :909. The message lists only the missing-width refusal. Also the new test at :1742 (asked nothing). The two nodes that already failed on the control still fail.
free > total → free > 9.9e99 * 1.0 1 / 1150 ([overrides2] only) 2 / 1150 test_readings_no_card_could_produce_are_refused_though_their_difference_is_not, :1597: the what names -120000000000.0 bytes reserved instead
reserved < 0 removed (the converse) 1 / 1150 1 / 1151 [overrides3] on both sides. The new assertion does not demand this branch.

These match the developer's table node for node. The only differences are larger absolute counts at the newer tip, and the same deltas: 0→1, 1→2, 2→4 and 1→2. In every run, every failing node is in test_spec_verbs.py.

2. Brittleness rulings (principles 3 and 8)

  • The new test that records probe_for calls: accepted as a white-box pin. No refusal can show which tables the walk read; with the pin absent it is 0 / 1151. So recording what is asked is the only observation available. It is also the property itself, because _reach registers PROBE_TABLES as what PROBES reads.

    Two measured points on the hole-fill red:

    • With the hole filled, 4 tests already fail at the tip, and this PR makes it 6: this test and the [probe] parametrisation. The [probe] red is correct: at the tip that node stayed green only because the WIDTHS neighbour satisfied it.
    • The body's claim that after the fill "no fixture can observe the walk" is wrong. With assert asked or not PROBE_TABLES, a wrong walk is still caught after the fill ([('driver_and...', 16)] == []), and the test no longer goes red on the fill itself. Details and the four runs are in the inline comment at :1742. Non-blocking.
  • Which of two impossibilities is named first: an accident, not a contract. impossible is a chain of first-match checks with no stated precedence. Swapping the two branches gives 0 / 1151 on the tip and 1 failed on the branch, so this PR adds the constraint.

    The constraint is unavoidable in that test. A plausible difference beside free > total forces reserved < 0, so the fixture always carries both impossibilities, and the order pin is the only way it can witness that branch. [overrides2] already holds the branch on its own.

    I accept it, since the body discloses it. However, the code comment gives the wrong reason if the test goes red on a deliberate reorder. See the inline comment at :1597. Non-blocking.

  • Checking each subject for its own refusal, not for only that refusal: honest, and enough. Exclusivity is not the property, and it cannot hold for the probe subject. A missing entry that no probe fills is also, by design, a missing width: test_a_missing_entry_no_probe_fills_is_two_refusals_and_not_one asserts exactly two.

    Requiring rule and text on the same refusal closes the neighbour path. I checked each needle against the package:

    • the STACK text is raised as a refusal only by validate's strict branch. machine.check_stack emits the same words as a warning, which is not in refusals.
    • the TRANSFERS text is written only by _transfers' mismatch branch;
    • the PROBES text is written only by probe_for.

    Independent confirmation: with the hole filled, the [probe] parametrisation goes red on the branch and stays green on the tip.

3. The "still would not notice" list

  • F4 hole-fill: cheap to close here. Change one word at :1742 and delete the body's "no fixture can observe" sentence, or correct it. I measured the change: it passes today and still catches a walk over WIDTH_TABLES both before and after the fill.
  • F1 order: cheap. Either reword the comment at :1597 or drop the assertion, as set out inline.
  • F2 remedy: not worth adding. The rule check plus the width-specific what is already unique to the guard. Removing the guard makes the [-2] node fail on that text, and the remedy would add nothing a mutant can separate.
  • F6: ...hand_measured_entry_is_not_refused_for_it asserts only ok. Agreed: it cannot be held from outside without pretending. The per-width skip is held elsewhere.

4. Importers of atom.compass.spec

I grepped for compass\.spec, compass import ... spec and relative .spec imports outside the package:

All five are under tests/compass, and no production module outside spec/ imports it. None of the three non-verbs importers had a failing node under any of the nine mutants on either side. So pytest tests/compass is a complete denominator for these mutations. The body's "four" is stale after #176 but does not affect the conclusion.

5. No design-doc references

I grepped the added lines for F\d, finding, principle, #\d+, D\d+, audit and m28, and found none. The new test name and all assertion messages describe behaviour and use no review labels. The F1–F6 labels appear only in the PR body, which is correct.

Gate (principle 7): the tree that will land

Node 18, staged with git archive + docker exec -i under /tmp/pr263rev-xbz/:

  • tarball md5s matched on both ends;
  • .compass-commit and .compass-changed were written from the same rev-parse;
  • each tree ran its own scripts/compass/gate_cpu.sh, under timeout -k 10 2400, not piped, one at a time;
  • atom.__file__ was printed under each staged root;
  • nothing was written into the shared mount, and the staging has been removed.

The tip moved during the review (#264, #267). #264 touches spec/explain.py and this same test file, so I gated the merge at both tips.

tip tip + 00a4d1479
175739f87 5107 (brief's figure) merge tree 61f6f9ba2 (def75749d): 5108 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0
77d203b86 (current) 5111 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0 merge tree baad3fd25 (43f856a41): 5112 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0

At the current tip, the delta by node id from both sides' --junitxml is:

  • only on the branch: tests/compass/test_spec_verbs.py::test_the_probe_question_is_asked_of_the_tables_it_says_it_reads, passed;
  • only on the control: none;
  • outcome changed: none.

gpu: not required on both sides. merge-tree is clean against both tips, and no test name in the merged file is duplicated.

Not in scope, untouched

"Refs #237" is right. I have not touched #243, #136 or #224.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant