compass(tests): the first-refusal test tells document order from sorted order - #497
Conversation
…ed order (#494) test_the_reader_raises_the_first_refusal_the_walk_meets put two unknown keys, first_knob and last_knob, in the device block. Their sorted order matched their document order, so a spec/machine.py::_survey that walked the block in sorted key order still raised first_knob and stayed green. The keys are now zeta_knob (before the block's fields) and alpha_knob (after them), and the test asserts zeta_knob. A walk in sorted order raises alpha_knob, and so does a reader that raises the last refusal. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
This review is agent-authored. Review cycle 1 Verdict: APPROVE Head covered: Mutations (serial, each one changes one line and keeps the line count, restored with
|
| Tree | Mutant in atom/compass/spec/machine.py |
Result |
|---|---|---|
| head | none | 298 passed |
| head | null comment on _survey's loop header (control) |
298 passed |
| head | _survey walks sorted(node.items(), key=lambda kv: str(kv[0])) |
1 failed, 297 passed |
| head | _survey walks reversed(list(node.items())) |
1 failed, 297 passed |
| head | _walk raises the last refusal (list(_survey(...))[::-1]) |
1 failed, 297 passed |
| head | _walk iterates the refusals sorted by message |
1 failed, 297 passed |
| head | _survey walks in reverse-sorted key order |
298 passed (see note 1) |
| base test file | sorted walk in _survey |
143 passed (the #494 premise, reproduced) |
| base test file | last refusal in _walk |
1 failed, 142 passed |
Every red at the head is tests/compass/test_spec_schema.py::test_the_reader_raises_the_first_refusal_the_walk_meets, and the sorted mutant fails on assert 'device.zeta_knob' in 'device.alpha_knob is not a field of this schema'. So the pin goes red on the defect it names. Both base-file rows match the PR's dev record.
I threw one mutant away: _walk raising max(_survey(...), key=...) gave 144 failed. That is a ValueError on every document that has no refusal, so it went red for the wrong reason and is not a result.
Where the mutant belongs: the developer is right. At this tip, _walk is a two-line wrapper that raises the first thing _survey yields. The key iteration is _survey's for key, value in node.items():. _survey is also imported by spec/validate.py, which calls list(_survey(document, "", found)), and spec/merge.py calls _walk. So a sorted walk in _survey would also reorder what the check path reports. The test pins the reader path only, which is what the PR body says.
Findings (none blocking)
- Reverse-sorted walk survives (non-blocking). It left 298 passed on the two spec files. It also left 777 passed across the ten
tests/compassfiles that touch the spec module (test_kv_budget*.py,test_ir_data_model.py,test_memory_*.py,test_artifact_invalidation.py,test_spec_*.py,test_kv_simulated_connector.py,test_runner_rpc_surface.py). I did not run the full CPU gate. With two keys, the document-first key has to sort either first or last, so one of the two sort directions always matches document order. A third unknown key would close the gap: putmid_knobfirst andalpha_knob/zeta_knobafter the fields, and assertmid_knob. Reverse sorting is not a realistic recurrence, and compass(tests): the first-refusal test cannot tell document order from sorted order #494 namedsorted(node), so this is the author's call and not a hold. - Assertions: the test's only assertion is kept, with the key renamed.
pytest.raises(SpecRefusal)and the fixture's shape (one unknown key before the block's fields, one after) are unchanged. - docs(compass): prose is not a test subject; no line numbers or item counts; the kind of claim decides a doc-code disagreement #489 / AI_DEV_RULES: there are no prose assertions. A grep of the added lines for design-doc references, principle or gate mentions,
file.py:NNNcitations and item counts found nothing. The one added comment line says what the fixture does. Onlytests/compass/test_spec_schema.pychanges (+4/-3). The counts in the PR body are measured records with their commits, which the rules allow. - PR body claims checked:
_surveyiterates and_walkwraps it;_surveyis shared withvalidate.py;merge.pycalls_walk; the base and head mutant counts; the assertion text. All of them hold.
Merge-tree
git merge-tree --write-tree fork/feature/atomcompass_new 856205b9d gives 4d87f44afbf6343408d347a2aab59fe042493505. That is identical to the head's tree (856205b9d^{tree} = 4d87f44af...), with no conflicts.
ponytail-review
Lean already. Ship.
Closes #494
What changed
tests/compass/test_spec_schema.py::test_the_reader_raises_the_first_refusal_the_walk_meetsput two unknown keys,first_knobandlast_knob, in thedeviceblock. They sort in the same order as they appear, so the test could not tell a walk in document order from a walk in sorted order.The keys are now
zeta_knob(before the block's fields) andalpha_knob(after them). The test asserts`device.zeta_knob`. Both sorted order and "raise the last refusal" now namealpha_knob. The test's structure and its other checks are unchanged.Dev record
Step 0: the premise was re-measured at the tip
c59351778before any edit. At the tip,spec/machine.py::_walkdoes not iterate the node itself. It raises the first refusal thatspec/machine.py::_surveyyields, and_surveyis the function that iteratesnode.items(). So the sorted-order mutant goes on_survey's loop header. It changes one line and keeps the line count:for key, value in node.items():->for key, value in sorted(node.items(), key=lambda kv: str(kv[0])):tests/compass/test_spec_schema.pyc59351778c59351778_surveyc59351778_walk_survey_walkThe "last refusal" mutant is
for refusal in _survey(node, prefix, found):->for refusal in list(_survey(node, prefix, found))[::-1]:.At the head, both mutants fail the same node,
tests/compass/test_spec_schema.py::test_the_reader_raises_the_first_refusal_the_walk_meets, with the same assertion:machine.pywas restored withgit checkoutafter each run, and the worktree diff was checked afterwards.Named result (from the brief): at the PR head, the
sorted(node)mutant of the walk fails this test. It sits on_survey, for the reason given above.Surprise: the brief puts the sorted iteration in
_walk. At this tip,_walkis a two-line wrapper, and the document-order iteration lives in_survey._surveyis shared withspec/validate.py, which collects every refusal.spec/merge.pyalso calls_walk. This test covers the reader path only.Left undone: nothing in the brief.
Gates
scripts/compass/gate_cpu.shinxiaobizh_n18_cpuon node 18. Each tree was staged withgit archiveand md5-checked on both ends. The branch and the control (the tip this branch forked from) have the same counts, as expected for a test-only rename:856205b9d: 5209 passed, 155 skipped, 3 xfailed;GATE_CPU_RC=0 PASSED, rc 0c59351778: 5209 passed, 155 skipped, 3 xfailed;GATE_CPU_RC=0 PASSED, rc 0spec/machine.py::_survey/_walk, which this PR does not touch.black --checkandruff checkpass on the edited file.Principle 8: each claim above names the run that measured it.
🤖 Generated with Claude Code