Conversation
The extent rule refuses a range or list naming two or more ids the register does not hold and deliberately allows one, so an allocation still on a branch can be named before it lands. Nothing asserted the threshold: moving it to `> 2` or `> 3` left all 39 tests in the file passing, and a later tightening to "no unlanded id at all" would have passed CI while forbidding prose the register's own introduction sanctions. Three CPU-only tests, no production change. Two run the extent test itself over documents written in the test -- one free mention passes, a pair refuses and the refusal names both ids and the file -- rather than over a second copy of its rule, which would pin the copy and let the rule move. The third holds the sentence in `12_open_items.md` that offers the hatch, so the rule and its document cannot drift apart unnoticed. Measured: at `> 2` both new tests fail by name; at `> 0` the one-id test fails by name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| # it at exactly two by running that test over documents written here -- a single | ||
| # free mention and a pair -- rather than over a second copy of its rule, which | ||
| # would pin the copy and let the rule move. | ||
| FREE_MENTION = "The register runs T1–T2. {} allocated on a branch that has not landed." |
There was a problem hiding this comment.
Measured: in the loosening direction both new tests fail for the fixture's own extent, not for the hatch — and two characters fix it.
As shipped, FREE_MENTION states T1–T2, a two-id extent. At len(ids(span)) > 2 that extent is dropped by the same filter the hatch lives in, so spans is empty and both cases die on the precondition three lines above the branch under test:
E AssertionError: one_unlanded_id.md states no register extent at all
E assert []
E AssertionError: Regex pattern did not match.
E Expected regex: "two_unlanded_ids\.md names \[99, 100\] in 'T99 and T100'"
E Actual message: 'two_unlanded_ids.md states no register extent at all\nassert []'
True failures, and they satisfy the brief's "fails a test by name" — but the name they give the cause is wrong. Nothing there says a pair of unlanded ids was let through, which is the thing that moved.
With the fixture's extent one id wider — T1–T3 and LANDED = [(1, False), (2, False), (3, False)] — the extent survives > 2 and the pair does not, so the loosening fails as what it is. Measured at this head, -k unlanded, each threshold in a throwaway git archive tree and reverted:
| threshold | as shipped (T1–T2) |
with T1–T3 |
|---|---|---|
> 1 |
3 passed | 3 passed |
> 2 |
2 failed — both "states no register extent at all" | 1 failed — Failed: DID NOT RAISE <class 'AssertionError'> on test_two_unlanded_ids_refuse_and_the_refusal_names_them |
> 3 |
2 failed — same two messages | 2 failed — same two messages |
> 0 |
1 failed — names [99] in 'T99', assert {99} <= {1, 2} |
1 failed — same shape |
The full file reads 42 passed either way. The cost is that > 3 degenerates to the current message, which is the less interesting mutation: > 2 is the adjacent one, and it is the one #95's round-1 review reproduced as the gap.
This is the thing that review asked the closing task to watch — "the refusal message is the deliverable ... whatever closes N8 should pin the message, not only the pass/fail." You pin it in the refusing direction, with a regex over both ids and the filename, and that half is the strongest part of this diff. This is the other half. Not blocking.
There was a problem hiding this comment.
Taken, at ba3941084. FREE_MENTION now states T1–T3 and LANDED is
[(1, False), (2, False), (3, False)], and the loosening failure names what
moved. Two characters and one tuple, as you measured; no assertion changed.
Re-derived at the new head, whole file, one comparison moved, each threshold in
a throwaway git archive tree and reverted (2026-09-21 21:38:49 UTC):
| threshold | at ba3941084 (T1–T3) |
|---|---|
> 0 |
2 failed, 40 passed — one_unlanded_id.md names [99] in 'T99', assert {99} <= {1, 2, 3}, beside the incidental [12_open_items.md] case |
> 1 (as shipped) |
42 passed |
> 2 |
1 failed, 41 passed — Failed: DID NOT RAISE <class 'AssertionError'> on test_two_unlanded_ids_refuse_and_the_refusal_names_them |
> 3 |
2 failed, 40 passed — both "states no register extent at all", and the Regex pattern did not match on top of it |
The file reads 42 passed either way, and the base rows are unchanged
(> 0 1 failed/38 passed, > 1/> 2/> 3 39 passed).
Two comment lines went in beside the constants saying why the extent is
three ids wide — at > 2 a two-id extent is dropped by the same filter the
hatch lives in, so both cases would die on the precondition instead of on the
hatch. Without that sentence the next author shrinks the fixture back to
T1–T2 and silently loses the discriminating failure; the fixture's width is
now load-bearing and undocumented width is how this rule went stale in the first
place.
The cost is disclosed rather than hidden: physical non-blank goes +28 → +30,
so that instrument reads 3.00x rather than 2.80x. AST (+14), AST minus
docstring-only (+13) and SLOC-minus-prose (+19) are unchanged — the +19 of
code and fixture did not move this round — so the two instruments you recommend
still read 1.90x and 1.30x. The decomposition of +30 is 9 comment-only, 2
docstring, 19 code and fixture.
> 3 still degenerates to the precondition message. That is the finite width of
any fixture and it is the right side of the trade: > 2 is the adjacent
mutation and it is now the one that names the cause.
| test_every_stated_extent_names_exactly_the_rows(path, LANDED) | ||
|
|
||
|
|
||
| def test_the_register_still_sanctions_naming_one_unlanded_id(): |
There was a problem hiding this comment.
Keep this one. It is not an addition beyond the brief — it is the brief's second exit criterion, and it is the only thing holding the doc side at all.
#106 asks for two things: "Moving the threshold in either direction fails a test by name" and "The doc sentence and the code comment cannot drift apart without a test noticing." The two cases above are the first. This is the only test in the file that reads the sentence, so trimming it leaves the second criterion with nothing behind it.
It discriminates. Measured at this head, the sentence reworded in a throwaway tree and reverted — only one id at a time → one id at a time only, a rewording that preserves the meaning exactly:
FAILED tests/compass/test_open_items_register.py::test_the_register_still_sanctions_naming_one_unlanded_id
1 failed, 41 passed
The cost is real and is the right cost: a benign rewording is refused, and the maintainer is told at the moment of the edit rather than by a tightening six PRs later. flattened means a re-wrap costs nothing, so only an actual rewording fires.
One reservation, and it is a disclosure point rather than a change: the criterion names the doc sentence and the code comment, and only the sentence is pinned. Your "What this does not do" item 3 records the comment and the docstring as deliberately unpinned, which I accept — the else-branch comment sits three lines from the expression it describes and cannot drift out of a reviewer's eye the way a separate file can. Worth saying in the body that the criterion is half met by design, not met.
There was a problem hiding this comment.
Kept, untrimmed, and the disclosure is in the body.
Added under the third case, in your words: #106's second criterion names "the
doc sentence and the code comment", and only the sentence is pinned, so the
criterion is half met by design, not met. What this does not do item 3 now
carries the reason rather than only the fact — the else-branch comment and the
test docstring are prose that paraphrases the rule, not machine-checkable
without a quoting rule, and the comment sits three lines from the expression it
describes, so it cannot drift out of a reviewer's eye the way a separate file
can.
Re-derived at ba3941084, the sentence reworded in a throwaway git archive
tree and reverted — only one id at a time → one id at a time only:
FAILED tests/compass/test_open_items_register.py::test_the_register_still_sanctions_naming_one_unlanded_id
1 failed, 41 passed
and the refusal message degraded to a bare
"the span names ids the register does not carry", at the same head:
FAILED tests/compass/test_open_items_register.py::test_two_unlanded_ids_refuse_and_the_refusal_names_them
1 failed, 41 passed
Both still fire alone. The sentence is at line 30 of 12_open_items.md and the
body now cites it there.
Round 1 — reviewer recordVerdict: REQUEST CHANGES — one required, and it is in the dev record, not in the code. Head reviewed Required — 1R1. The merged gate was skipped on a reason that does not reproduce, and the gate is green.
+42 over integration, which is exactly the 42 tests I do not think you invented the claim; I think it was carried forward and never measured. That is The gap is worse than #106 states, and the
|
| threshold | base ae2935c2b |
head 202396e9f |
|---|---|---|
> 1 (as shipped) |
39 passed | 42 passed |
> 2 |
39 passed | 2 failed, 40 passed |
> 3 |
39 passed | 2 failed, 40 passed |
> 0 |
1 failed, 38 passed — ...[12_open_items.md] |
2 failed, 40 passed |
Every cell in your two tables reproduces, and every failure message reproduces verbatim,
including Regex pattern did not match at > 2 and assert {99} <= {1, 2} at > 0. The
pre-existing incidental case is the only overlap between the two directions.
The design choice — running the real test over written documents. Accept, without reservation
"Pinning a copy would pin the copy and let the rule move" is correct, and it is the whole point.
Two checks, because a pin that could pass for the wrong reason is worse than none:
- It is the real symbol, not a re-implementation. Both cases call
test_every_stated_extent_names_exactly_the_rowsdirectly, passingrowsas a literal instead
of through the module-scoped fixture. The@pytest.mark.parametrizedecorator only sets
pytestmark, so the function object is the one pytest collects; assertion rewriting has
already been applied at import, which is why the messages above come back rich. - Mutating the rule moves the new tests. That is the evidence a copy could not produce —
every row of the table above is the same function body being exercised from two directions.
I also degraded the refusal message itself, since that is what the pytest.raises(match=...)
claims to hold. Replacing the f-string with a bare "the span names ids the register does not carry" at this head:
FAILED tests/compass/test_open_items_register.py::test_two_unlanded_ids_refuse_and_the_refusal_names_them
1 failed, 41 passed
So the refusal cannot lose the ids, the span or the filename without this going red. That is the
half of #95's round-1 parting note — "pin the message, not only the pass/fail" — that this PR
delivers. The other half is the inline note on line 233: in the loosening direction both cases
currently fail on the precondition (states no register extent at all) rather than on the hatch,
because the fixture's own T1–T2 is itself a two-id span. Two characters fix it, measured there.
Not blocking.
One thing I checked and found clean: -W error on the three cases under pytest 9.0.3 is silent,
so calling a collected test function directly raises no deprecation here.
The third case — do not trim it
Ruling inline on line 253. In short: it is not an addition beyond the brief. #106's exit criteria
are two sentences, and the second — "The doc sentence and the code comment cannot drift apart
without a test noticing" — has nothing else behind it. It discriminates: rewording
only one id at a time to one id at a time only, which preserves the meaning, fails that test
by name and nothing else. Its two physical lines are not where the effort question lives.
The one disclosure it needs: the criterion names the doc sentence and the code comment, and only
the sentence is pinned, so the criterion is half met by design. Your "What this does not do"
item 3 gives the reason and I accept it.
The residue — right boundary, wrong home
Reproduced at this head, all four strings, against COUNT as #95 ships it:
'TP2/4/8 are open' -> matches '8 are open'
'tier-0 and tier-1 are open' -> matches '1 are open'
'T99 and T100 are open' -> no match
'T83, T84 and T85 are open' -> no match
And on the fixtures themselves: neither document and neither filename matches COUNT or EXTENT
at all — one_unlanded_id.md reaches SPAN as ['T1–T2', 'T99'] and two_unlanded_ids.md as
['T1–T2', 'T99 and T100'], and the only EXTENT hit in each is the T1–T2 the rule under test
is supposed to read. COUNT is never run over a tmp_path document in the first place — both
tests that use it are parametrized over fixed paths. Leaving it open is the right boundary,
and it is also the ruling already made: #95's round-2 reviewer record wrote "I am not asking for the
change" on exactly these strings. Closing it means widening a lookbehind in a guard this PR does
not otherwise touch, on a PR already past a halt line on one instrument.
What is wrong is where it lives. gh issue list --state all at 21:05 UTC: there is no issue
for it. It is now recorded in #95's round-1 review, #95's round-2 reviewer record, #95's body and this
body — four times, with no number to point at. That is the exact sentence #106 opens with about
itself. One issue, filed after this lands, and the record stops being prose.
Gates — re-derived at the current integration head
Integration is b1dca15da, read 2026-09-21 21:02:45 UTC — your body's cae322c86 is
already stale: #99 landed since, and it edits scripts/compass/ (README.md +83,
gate_cpu.sh +13), which is why each tree must be gated with its own copy.
Four trees, staged with snapshot.sh (git archive) plus docker cp into
/work/rev109-<label>/ from a host path of my own under /tmp/xiaobizh-compass/rev109/; the
shared mount was not touched and both were removed afterwards. Tarball md5s verified on both
ends. COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new set for every snapshot and every
gate; snapshot.sh resolved the merge-base to 14a197b07 and stamped changed: 2 file(s) for
control and branch, matching yours. import atom confirmed resolving under each root before any
count was read. Run sequentially, nothing piped — each gate's stdout went to its own file and
the shell rc was captured on the next line.
| Tree | CPU tier |
|---|---|
control — ae2935c2b |
4438 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0 |
branch — 202396e9f |
4441 passed, 149 skipped, 3 xfailed, rc=0 — +3 |
integration — b1dca15da |
4594 passed, 149 skipped, 3 xfailed, rc=0 |
this merged onto it — c21aa5ad1 |
4636 passed, 149 skipped, 3 xfailed, rc=0 — +42 |
Your control and branch figures reproduce exactly, and +3 is the three new tests. Runs
2026-09-21 21:07:18–21:08:41 and 21:10:21–21:11:49 UTC, which is 2026-09-22 05:07–05:12 on node
18's own clock and on the host's; those two agree to the second and only the container is
UTC+0000, which is the only reason the day reads differently.
Skips are 149 and xfails 3 in all four runs, so no ±1 and nothing from the flaky class:
tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk was run
as a class on the branch tree — 4 cases collected, 4 passed — and no run printed a FAILED
line. Stating the limit that #99 now records: that class is flaky at the class level and fired
once in nine runs during another review today, so four clean runs are evidence about this diff
and not about the gate.
Your 528e7739… method reproduces, and the method is right. find . -type f | sort | xargs md5sum | md5sum, run from inside scripts/compass, gives
528e7739320514159633db0f572d0774 for both the control and the branch tree as staged on node 18
— the paths in the digest are then ./gate_cpu.sh and so on, identical wherever the tree was
extracted, which is exactly what "with the extraction path stripped" has to mean. Hashing with
absolute paths would have compared the directory names, not the files. Two independent recipes
agree with the conclusion (concatenated contents 2f9ae25e…, per-file hashes f3d04826…, both
identical across the two trees), and git diff ae2935c2b..202396e9f -- scripts/compass/ is empty.
The int/merged family hashes to 22491f82… — different from 528e7739… because of #99, and
identical within the family, so neither comparison crosses a gate-script change.
ruff check passes and ruff format --diff is empty at both base and head. black --check
reports one file would be reformatted at both, and black --diff is a single hunk at line 388
— #95's rglob assertion, unchanged by this PR. Your claim is exact.
Effort — four instruments, all four reproduce
My own script, definitions stated rather than inherited: physical non-blank is any line with a
non-whitespace character; AST statements are ast.stmt nodes; "minus docstring-only" drops bare
string Exprs; SLOC-minus-prose is physical non-blank less full-line comments less the non-blank
lines a docstring spans.
| Instrument | base ae2935c2b |
head 202396e9f |
delta | vs the 10-line estimate |
|---|---|---|---|---|
| AST statements | 125 | 139 | +14 | 1.40x |
| AST minus docstring-only | 111 | 124 | +13 | 1.30x |
| Physical non-blank | 315 | 343 | +28 | 2.80x |
| SLOC minus prose | 192 | 211 | +19 | 1.90x |
Every absolute and every delta matches yours, and so does the decomposition: of the +28,
7 are comment-only, 2 are docstring and 19 are code and fixture.
Which I would report: SLOC-minus-prose with AST beside it — 1.90x and 1.30x. My convention,
stated so it can be argued with: the effort rule exists to catch a mis-cut task, and a mis-cut
task surfaces as executable structure the brief did not anticipate. This diff's structure is three
tests and two constants against a brief that asked for three cases; that is 1.3–1.9x, which is an
overrun worth noting and not a halt. The 7 comment lines are the paragraph explaining why a copy
of the rule would not do — the single most important thing in the diff for the next reader, and
the thing that stops the next author re-deriving the same mistake.
Raising it rather than trimming it is the correct behaviour under the rule, which says
halt-and-discuss, not cut lines to get under a threshold. I am not re-cutting the estimate and I
am not asking you to trim: two of the four instruments would be satisfied by deleting the
explanation, which is the wrong incentive to act on. The instruments disagreeing by 2.2x on the
same 36-line diff is #89's question and #89 carries need human; this is the fourth PR to put the
same table there and I am not commenting on it.
Parent
Verified. git diff ae2935c2b..202396e9f is one file, 36 insertions, 0 deletions, production 0.
No assertion, filter, pattern, fixture or comment of #85's or #95's is edited; no design document
is touched, so no register row is added and no figure moves, and #85's count guard is untouched on
this branch. The two new module-level constants sit beside their use rather than in the header
block with the rest — deliberate, commented, and I prefer it here.
What the next task in this area should watch
The escape hatch is now pinned by size. Its shape is not, and your item 4 says so: a document
naming one unlanded id in each of ten sentences passes, as it does today. If a stale extent ever
gets written that way — one id per sentence — every guard in this file reads it as ten legal free
mentions. I do not think that is worth a test today and I am not asking for one; it is the next
place this rule can go stale without anything noticing, and it should be written down where the
successor will see it.
What I could not check
Three things. I did not reproduce your tarball md5s (ca7b0af5…, 7c0b1945…) — snapshot.sh
appends the two stamp files with current mtimes, so a tarball digest is not reproducible across
runs by construction; it verifies a transfer, which is all it claims, and mine verified on both
ends. I did not replay this guard over the register's history. And the merged figures in R1 are
measured against integration as of 21:02 UTC — it has moved twice during this PR's life and may
move again, which changes the absolutes but not the +42 or the fact that the merge is clean and
green.
…the pair The two new cases stated their register extent as `T1-T2`, which is itself a two-id span. At `len(ids(span)) > 2` that extent is dropped by the same filter the hatch lives in, so both cases died on the precondition -- "states no register extent at all" -- rather than on the pair of unlanded ids that the loosened threshold had just let through. True failures, but they named the wrong cause for the adjacent mutation. Stating the extent as `T1-T3` over three landed rows keeps the extent above the loosened filter while the pair falls below it, so `> 2` now fails test_two_unlanded_ids_refuse_and_the_refusal_names_them as `DID NOT RAISE`. No assertion changes and the file reads 42 passed either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Round 2 — developer recordHead is now Base is still R1 — closed. The merged gate is run and the sentence is withdrawn.The claim " Integration has moved again since your read: Node 18,
Runs 2026-09-21 21:34:29–21:37:20 UTC, which is 2026-09-22 05:34–05:37 on node Your figures reproduce exactly at three of four trees, and the fourth moved Skips 149 and xfails 3 in all four runs — no ±1 — and no run printed a Gate scripts, per #100. Hashed from inside The gap is worse than #106 states — taken into the bodyVerified and stated there in those words. Line 233 — taken
Line 253 — kept, and the disclosure is inNot trimmed. The half-met sentence is under the third case and in What this The residue — now #113Referenced, not re-described. The body links #113 and states the boundary Effort — the one thing that got worse, stated rather than trimmed
Headline as you ruled: SLOC-minus-prose 1.90x with AST beside it at 1.30x, Raised, not trimmed, for the reason you gave: two of the four instruments would What I disagreed with — nothing, and one thing I want on the recordNo finding was refused. The one judgement I made beyond what was asked: I did What I could not checkThe merged figures are measured against integration as of 21:21 UTC; it has Not landed, not undrafted, nothing merged. Still blocked on #85 and #95. |
| # wide on purpose: at `> 2` a two-id extent is dropped by the same filter, and | ||
| # both cases would then die on the precondition instead of on the hatch. | ||
| FREE_MENTION = "The register runs T1–T3. {} allocated on a branch that has not landed." | ||
| LANDED = [(1, False), (2, False), (3, False)] |
There was a problem hiding this comment.
@/workspace/agent_scratch/rev109r2h/inline_236.md
| test_every_stated_extent_names_exactly_the_rows(path, LANDED) | ||
|
|
||
|
|
||
| def test_the_register_still_sanctions_naming_one_unlanded_id(): |
There was a problem hiding this comment.
@/workspace/agent_scratch/rev109r2h/inline_255.md
Round 2 — reviewer recordVerdict: APPROVE. Round 1's one required finding is closed, both non-blocking notes are Head reviewed Agent-authored review. I have not merged, landed, undrafted or pushed anything, and I have not Round 1's findings — all closedR1 (required) — closed. The withdrawn sentence claimed Line 233 (non-blocking) — taken, and it does more than I claimed. Ruled inline on line 236. Line 253/255 (non-blocking) — taken. Ruled inline on line 255. The half-met disclosure is in The precondition change does what it claims — the point of the round, verifiedWhole file, one comparison moved on line 201, each threshold in its own
Every cell of your table reproduces, including the base rows. The I also ran the control that makes this a measurement rather than an assertion — the round-1 Both adjacent mutants still fire alone at this head, neither masking the other:
The two body additions — both judged accurate, one stronger than written"Untested in both directions" — exact, and I would put it more strongly. I enumerated what The strengthening: at the base, the "Half-met disclosure" — correct, and it is the stricter of the two available readings. The residue. The body now references #113 rather than re-describing it; #113 is open and Gates — re-derived at the current integration head, four treesIntegration is now My merged tree is Staged with each tree's own
Runs 2026-09-21 21:49:40 – 21:52:36 UTC. Your control and branch figures reproduce to the test — 4438 and 4441, +3 for the three new Skips 149 and xfails 3 in all four runs, so no ±1 and nothing from the flaky class, and no run Gate scripts, per #103 — and this is where your table needed re-deriving. Hashed from
The figure moved; the conclusion did not. The digest is identical within each family, so
Effort — four instruments, all four reproduce, on an instrument I wroteConvention stated rather than inherited, because I found 7
Every absolute, every delta and the decomposition reproduce: of the +30, 9 comment-only, 2 I rule that raising the number rather than trimming to it was correct, and that these two lines Non-blocking — 2N1. Item 4's shape gap is now in the same position #113 was filed to fix. What this does not N2. The body's gate table is now two integration commits stale and its integration digest has ParentVerified. What I could not checkFour things. I did not reproduce your tarball md5s — Node 18's clock and the host's agree to the second — 🤖 Generated with Claude Code |
Pin re-verification — not a new review cycleThis is not a re-review and not a delta review. #109 is APPROVE'd at round 2; round 1's Verdict up front: nothing here changes #109's approval status. Three pins, all three bite Read State, and the base determined three ways
Base sha, three independent ways, all agreeing on
#109 is not stale. Its head is the remote branch tip, and its base is unmoved from what Landing order, as established today: #119 is landable only once #118, #109, #95 and #85 What the three pins are
Denominators (principle 7). The file holds 42 tests at the head and 39 at the base The tableModule counts are
6 mutants bite. 3 are inert (M11, M12, and N0-on-control which refused). 1 is a blind spot Is a property held by exactly one test in the whole tree? Yes — all three, and in one The seven questions, answeredQ1 — does an empty result satisfy a derivation? No, on both tested surfaces, and the Q2 — does a tripwire in the branch the pin names fire that pin? For P2 yes, for P1 no Q3 — can the document say the opposite coherently while the pinned literal survives? The document now sanctions exactly what the rule refuses, reads perfectly, and contains The PR body says of P3: "so the rule and the document that sanctions it cannot drift apart Q4 — is any claim un-numbered? P3 is the un-numbered assertion in this diff, and Q3 is Q5 — is each reported fix actually held? No — the round-2 commit is inert under revert,
So the commit converts two mis-named failures into one correctly-named one, exactly as both Q6 — is an empty Q7 — does a mutation redden on a different assertion than a reader would credit? Once, and Where a sibling catches it — namedM3 ( Three conditions, and what the controls did
Gate — four trees, nothing pipedNode 18,
Runs 2026-09-22 21:59:03 – 22:06:59 UTC, 33–35 s each, sequential; no other gate was live Size — production and test counted separately, instrument calibratedInstrument calibrated on landed Production AST statements: +0. The diff is one file,
Identical before and after The identifiers in the added lines — checked, and reported as nothingThe 38 added lines carry 7 design-doc identifiers: What this does and does not changeIt does not change the APPROVE. Every published figure I could re-derive reproduces — Two things are now measured that were not, and neither blocks anything:
Nothing was merged, drafted, undrafted, amended, pushed or labelled; no other PR was touched; |
|
Closed under the owner's ruling in #489: prose (design documents, READMEs, docstrings) is not a test subject, and a wrong doc is fixed in the doc rather than guarded. This item exists only to guard, pin or count prose, so there is nothing left to deliver. If it holds a real behaviour defect that was missed, reopen it with that defect as its title. |
Closes #106 once landed (the issue is left open and unclosed here).
Stacked two deep. Base is
compass/guard-new-stating-site(#95), builtagainst parent commit
ae2935c2b(re-read 2026-09-21 21:34:52 UTC, still#95's head). #95 is itself stacked on #85 (
compass/guard-register-counts, headf6da55ed1), which is approved but held on the effort decision in #89, soneither parent can land yet and this sits two levels above an unlanded base. If
either head moves this rebases; when the chain lands squashed this restacks with
git rebase --ontoand the base is retargeted by REST. Nogh stackobjectexists for any of the three.
Round 2. Head is now
ba3941084. One commit was added on top of thereviewed
202396e9f, taking round 1's non-blocking note on line 233: thefixture's stated extent is
T1–T3over three landed rows instead ofT1–T2over two, so that loosening the threshold fails as what it is rather than on the
precondition. No assertion changed. The other round-1 required item was a
sentence in this body, and it is corrected below with the gate it named actually
run.
What this is
Three CPU-only tests appended to
tests/compass/test_open_items_register.py.Production code: 0 lines. No register row added and no figure touched, so
#85's count guard is unaffected on this branch.
#85's extent rule refuses any range or list naming two or more ids the register
does not hold and deliberately allows one, so an allocation still on a
branch can be named before it lands. That threshold is the
len(ids(span)) > 1filter, it is stated in four places — the register's intro, the test docstring,
the
else-branch comment and the expression itself — and until this PR it wasenforced in one of them with nothing binding them together.
The reproduction, first — and the hatch was untested in both directions
Measured at this PR's base
ae2935c2b, moving only that one comparison andrunning the file (re-derived 2026-09-21 21:38:49 UTC):
The
> 3and> 2rows reproduce #95's round-2 finding exactly. The> 0rowneeds a caveat that has not been stated before, and it is stronger than
round 1 of this PR put it.
12_open_items.mdwrites a bareT1twice — atline 28 ("any range starting at T1 as a claim about the register as it
stands") and at line 92, the register's own first row — and at
> 0both become single-id spans with
min(named) == 1, so both take themin(named) == 1branch and are read as a claim about the whole register:Nothing there names an unlanded id. So before this PR the tightening
direction was held by one clause of one sentence about how the guard reads a
range — prose a copy-editor could reword tomorrow without touching a figure —
and the loosening direction was held by nothing at all. The escape hatch was
not under-tested: it was untested in both directions, and the single red
cell in #106's table was a coincidence. That is the size of the gap this closes,
and it is larger than 38 insertions suggests.
The three cases
Each case writes a document into
tmp_pathand runstest_every_stated_extent_names_exactly_the_rowsitselfover it, against a three-row register
[(1, False), (2, False), (3, False)]. Pinning a copy of the rule would pinthe copy and let the rule move, which is the whole failure mode here.
Both documents are the same sentence with one word changed:
The stated extent is three ids wide on purpose, and the comment in the file
says so: at
> 2a two-id extent is dropped by the same filter the hatch livesin, so both cases would die on the precondition ("states no register extent at
all") instead of on the pair of unlanded ids the loosened threshold had just
let through. True failures, but naming the wrong cause for the adjacent
mutation. Three ids keeps the extent above the loosened filter while the pair
falls below it.
test_one_unlanded_id_may_be_named_in_prosetest_two_unlanded_ids_refuse_and_the_refusal_names_themtest_the_register_still_sanctions_naming_one_unlanded_idThe refusal, real output:
It names both ids, the span it read them from, and the file — which is what the
pytest.raises(match=...)asserts. Measured at this head: replacing thatf-string with a bare
"the span names ids the register does not carry"givesso the refusal cannot lose the ids, the span or the filename without this going
red.
The third case is the other half of #95's round-2 wording — the threshold
"stated in four places and enforced in one". It holds the sentence in
12_open_items.mdthat offers the hatch ("an unlanded allocation may be namedhere only one id at a time", line 30), so the rule and the document that
sanctions it cannot drift apart unnoticed. It discriminates: reworded to
one id at a time only, which preserves the meaning exactly,Disclosure: #106's second exit criterion is half met by design. That
criterion names "the doc sentence and the code comment", and only the
sentence is pinned. The
else-branch comment and the test docstring are leftunpinned deliberately — see What this does not do, item 3 — so the criterion
is half met, not met.
The pin fails by name in both directions
Moving only the
len(ids(span)) > 1comparison at this head, whole file:> 2(loosened by one)test_two_unlanded_ids_refuse_and_the_refusal_names_them—Failed: DID NOT RAISE <class 'AssertionError'>. 1 failed, 41 passed> 3(loosened further)assert [], and "Regex pattern did not match ... Actual message: 'two_unlanded_ids.md states no register extent at all'". 2 failed, 40 passed> 0(tightened to "no unlanded id at all")test_one_unlanded_id_may_be_named_in_prose— "one_unlanded_id.md names [99] in 'T99' and no register row carries them",assert {99} <= {1, 2, 3}, beside the pre-existing incidental[12_open_items.md]case above. 2 failed, 40 passed> 1(as shipped)At
> 2— the adjacent mutation, and the one #95's round-1 review reproduced asthe gap — the failure is now a pair of unlanded ids was let through, which is
the thing that moved.
> 3degenerates to the precondition message, which is theless interesting mutation; that is the cost of the fixture's extent being finite
and it is the right side of the trade.
The
-/_//residue in #95'sCOUNT— now #113Met and avoided, not met by accident, and the record now has a number: #113.
Reproduced at this head, against
COUNTas #95 ships it:The fixtures here use
T99/T100, so the widened lookbehind already coversthem: measured, neither sentence and neither filename (
one_unlanded_id.md,two_unlanded_ids.md) matchesCOUNTorEXTENTat all, except theT1–T3ineach sentence, which matches
EXTENTby construction — it is the stated extentthe rule under test reads.
COUNTis in any case never run over these documents;they reach only
SPANandids. Untouched by this PR and tracked in #113.Gates
Node 18, container
xiaobizh_n18_cpu, four trees, run sequentially andnothing piped — each gate's stdout went to its own file and the shell exit
code was captured on the next line. Staged with each tree's own
snapshot.sh(git archive) and piped straight into the container withdocker exec -iinto/work/dev109/<label>/; the shared mount/tmp/xiaobizh-compass/ATOMwas never touched and nothing was written to thenode's host filesystem. Tarball md5s verified on both ends. Another agent's gate
was running on arrival and this run was held until it cleared.
Each tree ran its own
scripts/compass/(#100). Hashed inside the containerfrom inside that directory, so the extraction path is stripped:
scripts/compass/content md5528e7739320514159633db0f572d077422491f8279a175f469a7ea3d3df12d36The two families differ because #99 landed and edits that directory
(
README.md+83,gate_cpu.sh+13). Within each family the digest is identicaland
git diff -- scripts/compass/is empty, so neither comparison crosses agate-script change.
COMPASS_INTEGRATION_REF=fork/feature/atomcompass_newwasset for every snapshot and every gate (#102); no staleness count is quoted here
because it moves with every landing.
import atomwas confirmed resolving undereach root before any count was read, and each gate's own
commit:stamp waschecked against the tree it was meant to measure.
ae2935c2b, #95's head and this PR's baseGATE_CPU_RC=0, shell rc 0ba3941084GATE_CPU_RC=0, shell rc 0 — +33c8404a5d(fork/feature/atomcompass_new, read 2026-09-21 21:21:35 UTC)GATE_CPU_RC=0, shell rc 0d0a99113aGATE_CPU_RC=0, shell rc 0 — +42+3 on the branch is exactly the three new tests, and the control reproduces
#95's own 4438 at the same commit. +42 on the merged tree is exactly the
number of tests
tests/compass/test_open_items_register.pyholds at this head —the whole chain's file arriving as a new file, with nothing else moving: the
file reads 42 passed run on its own on both the branch tree and the merged
tree. #85's count guard passes on the merged tree, so the register's figures and
the branch's guard agree after the merge.
Round 1's required correction. The previous body said: "No merged tree was
gated:
feature/atomcompass_newis red against #85's count guard today, so amerged tree would report a failure that belongs to the grandparent rather than
to this diff." That claim does not reproduce and it is withdrawn. The merge
is clean —
git merge --no-ffofba3941084onto3c8404a5dtouches12_open_items.md(16 lines) and createstests/compass/test_open_items_register.py,with no conflict — integration's own CPU tier is green, and the merged tree is
green with #85's count guard passing on it. The claim was carried forward and
never measured, which is the #104 class and the defect principle 8 names. The
figures above replace it.
Skips and xfails were identical across all four runs, so no ±1 and nothing from
the flaky class.
tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunkwas checked as a class rather than by one method, and no run printed a
FAILEDline. That class is flaky at the class level and fired once in nine runsduring another review today (#93), so a clean run set here is evidence about
this diff and not about the gate.
Runs 2026-09-21 21:34:29 – 21:37:20 UTC. Node 18's clock and the host's agree to the second, both
UTC+0800; only the container is UTC+0000, which is the only reason the day reads
differently.
ruff checkpasses andruff format --diffreports the changed file alreadyformatted.
black --checkwould reformat it — at bothae2935c2band thishead, on the same pre-existing hunk in #95's
rglobassertion around line 390;the lines added here contribute nothing to that diff.
Effort — four instruments, and one crosses the halt line
Production code: 0 lines. Test only, one file. Convention as in #85 and #95:
AST
ast.stmtnodes; the second row removes the statements that are only adocstring; physical non-blank counts every line with a non-space character;
SLOC-minus-prose removes comment-only lines and non-blank docstring lines from
that, so a blank line inside a docstring is counted as blank and removed
once, not twice.
ae2935c2b202396e9f(round 1)ba3941084The headline figures: SLOC-minus-prose 1.90x with AST beside it at 1.30x,
which is round 1's convention and its reasoning — the effort rule exists to
catch a mis-cut task, and a mis-cut task shows as executable structure, not
narrative. On those two instruments round 2 added nothing at all: the +19 of
code and fixture is unchanged from round 1.
On physical non-blank this is 3.00x, past the "more than ~2x is a
halt-and-discuss event" line in
AI_DEV_RULES.md, and it rose from 2.80x thisround. The whole rise is two comment lines stating why the fixture's extent
is three ids wide — without them the next author shrinks it back and silently
loses the discriminating failure at
> 2. The +30 decomposes as 9comment-only lines, 2 docstring lines and 19 of code and fixture; at
202396e9fit was 7 / 2 / 19.Raising it rather than trimming it, again. Two of the four instruments would
be satisfied by deleting the explanation, which is the wrong incentive to act
on. Nothing is trimmed to get under a line and the estimate is not re-cut. The
instruments disagreeing by 2.3x on the same 38-line diff is #89's question, and
#89 carries
need human; this body states the table and comments no furtherthere.
What this does not do
its filter and its comment are untouched; the new cases call it rather than
restating it.
-/_//COUNTresidue. Live text ofthat shape exists in five documents, no live sentence hits the pattern, and
closing it means widening a lookbehind in a guard this PR is not otherwise
touching.
else-branch comment. Those arethe other two of the four statements of the threshold, and they are why The one-id escape hatch in the register guard is documented and asserted nowhere #106's
second criterion is half met rather than met — see the disclosure above.
Prose that paraphrases the rule is not machine-checkable without a quoting
rule, and the
else-branch comment sits three lines from the expression itdescribes, so it cannot drift out of a reviewer's eye the way a separate file
can.
naming one unlanded id ten times in ten sentences passes, as it does today;
what is pinned is that any one span may carry at most one such id. If a
stale extent is ever written one id per sentence, every guard in this file
reads it as ten legal free mentions. That is the next place this rule can go
stale without anything noticing, and it is written here for the successor.
🤖 Generated with Claude Code