Skip to content

compass(design): pin the register guard's one-id escape hatch - #109

Closed
jgong5 wants to merge 2 commits into
compass/guard-new-stating-sitefrom
compass/guard-one-id-hatch
Closed

jgong5 wants to merge 2 commits into
compass/guard-new-stating-sitefrom
compass/guard-one-id-hatch

Conversation

@jgong5

@jgong5 jgong5 commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner

Closes #106 once landed (the issue is left open and unclosed here).

Stacked two deep. Base is compass/guard-new-stating-site (#95), built
against parent commit ae2935c2b (re-read 2026-09-21 21:34:52 UTC, still
#95's head). #95 is itself stacked on #85 (compass/guard-register-counts, head
f6da55ed1), which is approved but held on the effort decision in #89, so
neither parent can land yet and this sits two levels above an unlanded base. If
either head moves this rebases; when the chain lands squashed this restacks with
git rebase --onto and the base is retargeted by REST. No gh stack object
exists for any of the three.

Round 2. Head is now ba3941084. One commit was added on top of the
reviewed 202396e9f, taking round 1's non-blocking note on line 233: the
fixture's stated extent is T1–T3 over three landed rows instead of T1–T2
over two, so that loosening the threshold fails as what it is rather than on the
precondition. No assertion changed. The other round-1 required item was a
sentence in this body, and it is corrected below with the gate it named actually
run.

What this is

Three CPU-only tests appended to tests/compass/test_open_items_register.py.
Production code: 0 lines. No register row added and no figure touched, so
#85's count guard is unaffected on this branch.

#85's extent rule refuses any range or list naming two or more ids the register
does not hold and deliberately allows one, so an allocation still on a
branch can be named before it lands. That threshold is the len(ids(span)) > 1
filter, it is stated in four places — the register's intro, the test docstring,
the else-branch comment and the expression itself — and until this PR it was
enforced in one of them with nothing binding them together.

The reproduction, first — and the hatch was untested in both directions

Measured at this PR's base ae2935c2b, moving only that one comparison and
running the file (re-derived 2026-09-21 21:38:49 UTC):

len(ids(span)) > 1   ->  39 passed     (as shipped)
len(ids(span)) > 2   ->  39 passed
len(ids(span)) > 3   ->  39 passed
len(ids(span)) > 0   ->   1 failed, 38 passed

The > 3 and > 2 rows reproduce #95's round-2 finding exactly. The > 0 row
needs a caveat that has not been stated before, and it is stronger than
round 1 of this PR put it. 12_open_items.md writes a bare T1 twice — at
line 28 ("any range starting at T1 as a claim about the register as it
stands"
) and at line 92, the register's own first row — and at > 0
both become single-id spans with min(named) == 1, so both take the
min(named) == 1 branch and are read as a claim about the whole register:

E  AssertionError: 12_open_items.md states the register as 'T1': [] stated with no row,
   [2, 3, ..., 87] in rows with no figure covering them
E  assert {1} == {1, 2, 3, 4, 5, 6, ...}

Nothing there names an unlanded id. So before this PR the tightening
direction was held by one clause of one sentence about how the guard reads a
range — prose a copy-editor could reword tomorrow without touching a figure —
and the loosening direction was held by nothing at all. The escape hatch was
not under-tested: it was untested in both directions, and the single red
cell in #106's table was a coincidence. That is the size of the gap this closes,
and it is larger than 38 insertions suggests.

The three cases

Each case writes a document into tmp_path and runs test_every_stated_extent_names_exactly_the_rows itself
over it, against a three-row register
[(1, False), (2, False), (3, False)]. Pinning a copy of the rule would pin
the copy and let the rule move, which is the whole failure mode here.

Both documents are the same sentence with one word changed:

The register runs T1–T3. T99 is allocated on a branch that has not landed.
The register runs T1–T3. T99 and T100 are allocated on a branch that has not landed.

The stated extent is three ids wide on purpose, and the comment in the file
says so: at > 2 a two-id extent is dropped by the same filter the hatch lives
in, so both cases would die on the precondition ("states no register extent at
all"
) instead of on the pair of unlanded ids the loosened threshold had just
let through. True failures, but naming the wrong cause for the adjacent
mutation. Three ids keeps the extent above the loosened filter while the pair
falls below it.

Case Test Result
one unlanded id test_one_unlanded_id_may_be_named_in_prose passes — no refusal
two unlanded ids test_two_unlanded_ids_refuse_and_the_refusal_names_them refuses, text below
the doc sentence test_the_register_still_sanctions_naming_one_unlanded_id passes

The refusal, real output:

AssertionError: two_unlanded_ids.md names [99, 100] in 'T99 and T100' and no
register row carries them

It names both ids, the span it read them from, and the file — which is what the
pytest.raises(match=...) asserts. Measured at this head: replacing that
f-string with a bare "the span names ids the register does not carry" gives

FAILED tests/compass/test_open_items_register.py::test_two_unlanded_ids_refuse_and_the_refusal_names_them
1 failed, 41 passed

so the refusal cannot lose the ids, the span or the filename without this going
red.

The third case is the other half of #95's round-2 wording — the threshold
"stated in four places and enforced in one". It holds the sentence in
12_open_items.md that offers the hatch ("an unlanded allocation may be named
here only one id at a time"
, line 30), so the rule and the document that
sanctions it cannot drift apart unnoticed. It discriminates: reworded to
one id at a time only, which preserves the meaning exactly,

FAILED tests/compass/test_open_items_register.py::test_the_register_still_sanctions_naming_one_unlanded_id
1 failed, 41 passed

Disclosure: #106's second exit criterion is half met by design. That
criterion names "the doc sentence and the code comment", and only the
sentence is pinned. The else-branch comment and the test docstring are left
unpinned deliberately — see What this does not do, item 3 — so the criterion
is half met, not met.

The pin fails by name in both directions

Moving only the len(ids(span)) > 1 comparison at this head, whole file:

Threshold Fails, by name
> 2 (loosened by one) test_two_unlanded_ids_refuse_and_the_refusal_names_them — Failed: DID NOT RAISE <class 'AssertionError'>. 1 failed, 41 passed
> 3 (loosened further) both new cases — "states no register extent at all", assert [], and "Regex pattern did not match ... Actual message: 'two_unlanded_ids.md states no register extent at all'". 2 failed, 40 passed
> 0 (tightened to "no unlanded id at all") test_one_unlanded_id_may_be_named_in_prose — "one_unlanded_id.md names [99] in 'T99' and no register row carries them", assert {99} <= {1, 2, 3}, beside the pre-existing incidental [12_open_items.md] case above. 2 failed, 40 passed
> 1 (as shipped) nothing — 42 passed

At > 2 — the adjacent mutation, and the one #95's round-1 review reproduced as
the gap — the failure is now a pair of unlanded ids was let through, which is
the thing that moved. > 3 degenerates to the precondition message, which is the
less interesting mutation; that is the cost of the fixture's extent being finite
and it is the right side of the trade.

The -/_// residue in #95's COUNT — now #113

Met and avoided, not met by accident, and the record now has a number: #113.
Reproduced at this head, against COUNT as #95 ships it:

'TP2/4/8 are open'            -> COUNT matches '8 are open'
'tier-0 and tier-1 are open'  -> COUNT matches '1 are open'
'T99 and T100 are open'       -> no match   (the lookbehind works)
'T83, T84 and T85 are open'   -> no match

The fixtures here use T99/T100, so the widened lookbehind already covers
them: measured, neither sentence and neither filename (one_unlanded_id.md,
two_unlanded_ids.md) matches COUNT or EXTENT at all, except the T1–T3 in
each sentence, which matches EXTENT by construction — it is the stated extent
the rule under test reads. COUNT is in any case never run over these documents;
they reach only SPAN and ids. Untouched by this PR and tracked in #113.

Gates

Node 18, container xiaobizh_n18_cpu, four trees, run sequentially and
nothing piped — each gate's stdout went to its own file and the shell exit
code was captured on the next line. Staged with each tree's own
snapshot.sh (git archive) and piped straight into the container with
docker exec -i into /work/dev109/<label>/; the shared mount
/tmp/xiaobizh-compass/ATOM was never touched and nothing was written to the
node's host filesystem. Tarball md5s verified on both ends. Another agent's gate
was running on arrival and this run was held until it cleared.

Each tree ran its own scripts/compass/ (#100). Hashed inside the container
from inside that directory, so the extraction path is stripped:

family scripts/compass/ content md5
control + branch 528e7739320514159633db0f572d0774
integration + merged 22491f8279a175f469a7ea3d3df12d36

The two families differ because #99 landed and edits that directory
(README.md +83, gate_cpu.sh +13). Within each family the digest is identical
and git diff -- scripts/compass/ is empty, so neither comparison crosses a
gate-script change
. COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new was
set for every snapshot and every gate (#102); no staleness count is quoted here
because it moves with every landing. import atom was confirmed resolving under
each root before any count was read, and each gate's own commit: stamp was
checked against the tree it was meant to measure.

Tree CPU tier
control — ae2935c2b, #95's head and this PR's base 4438 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0
branch — ba3941084 4441 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0 — +3
integration — 3c8404a5d (fork/feature/atomcompass_new, read 2026-09-21 21:21:35 UTC) 4594 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0
this merged onto it — d0a99113a 4636 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0 — +42

+3 on the branch is exactly the three new tests, and the control reproduces
#95's own 4438 at the same commit. +42 on the merged tree is exactly the
number of tests tests/compass/test_open_items_register.py holds at this head —
the whole chain's file arriving as a new file, with nothing else moving: the
file reads 42 passed run on its own on both the branch tree and the merged
tree. #85's count guard passes on the merged tree, so the register's figures and
the branch's guard agree after the merge.

Round 1's required correction. The previous body said: "No merged tree was
gated: feature/atomcompass_new is red against #85's count guard today, so a
merged tree would report a failure that belongs to the grandparent rather than
to this diff."
That claim does not reproduce and it is withdrawn. The merge
is clean — git merge --no-ff of ba3941084 onto 3c8404a5d touches
12_open_items.md (16 lines) and creates tests/compass/test_open_items_register.py,
with no conflict — integration's own CPU tier is green, and the merged tree is
green with #85's count guard passing on it. The claim was carried forward and
never measured, which is the #104 class and the defect principle 8 names. The
figures above replace it.

Skips and xfails were identical across all four runs, so no ±1 and nothing from
the flaky class. tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk
was checked as a class rather than by one method, and no run printed a
FAILED line. That class is flaky at the class level and fired once in nine runs
during another review today (#93), so a clean run set here is evidence about
this diff and not about the gate.

Runs 2026-09-21 21:34:29 – 21:37:20 UTC. Node 18's clock and the host's agree to the second, both
UTC+0800; only the container is UTC+0000, which is the only reason the day reads
differently.

ruff check passes and ruff format --diff reports the changed file already
formatted. black --check would reformat it — at both ae2935c2b and this
head, on the same pre-existing hunk in #95's rglob assertion around line 390;
the lines added here contribute nothing to that diff.

Effort — four instruments, and one crosses the halt line

Production code: 0 lines. Test only, one file. Convention as in #85 and #95:
AST ast.stmt nodes; the second row removes the statements that are only a
docstring; physical non-blank counts every line with a non-space character;
SLOC-minus-prose removes comment-only lines and non-blank docstring lines from
that, so a blank line inside a docstring is counted as blank and removed
once
, not twice.

Instrument base ae2935c2b 202396e9f (round 1) ba3941084 delta vs the 10-line estimate
AST statements, docstrings counted 125 139 139 +14 1.40x
AST statements, docstring-only excluded 111 124 124 +13 1.30x
Physical non-blank 315 343 345 +30 3.00x
SLOC minus prose 192 211 211 +19 1.90x

The headline figures: SLOC-minus-prose 1.90x with AST beside it at 1.30x,
which is round 1's convention and its reasoning — the effort rule exists to
catch a mis-cut task, and a mis-cut task shows as executable structure, not
narrative. On those two instruments round 2 added nothing at all: the +19 of
code and fixture is unchanged from round 1.

On physical non-blank this is 3.00x, past the "more than ~2x is a
halt-and-discuss event"
line in AI_DEV_RULES.md, and it rose from 2.80x this
round. The whole rise is two comment lines stating why the fixture's extent
is three ids wide — without them the next author shrinks it back and silently
loses the discriminating failure at > 2. The +30 decomposes as 9
comment-only lines, 2 docstring lines and 19 of code and fixture
; at
202396e9f it was 7 / 2 / 19.

Raising it rather than trimming it, again. Two of the four instruments would
be satisfied by deleting the explanation, which is the wrong incentive to act
on. Nothing is trimmed to get under a line and the estimate is not re-cut. The
instruments disagreeing by 2.3x on the same 38-line diff is #89's question, and
#89 carries need human; this body states the table and comments no further
there.

What this does not do

  1. It does not edit any assertion of compass(design): guard the open-items register's own counts #85's or compass(design): a third document may not state the register's extent #95's. The extent test's body,
    its filter and its comment are untouched; the new cases call it rather than
    restating it.
  2. It does not close The COUNT lookbehind lets a hyphen or slash through, recorded four times with no number #113, compass(design): a third document may not state the register's extent #95's -/_// COUNT residue. Live text of
    that shape exists in five documents, no live sentence hits the pattern, and
    closing it means widening a lookbehind in a guard this PR is not otherwise
    touching.
  3. It does not pin the test docstring or the else-branch comment. Those are
    the other two of the four statements of the threshold, and they are why The one-id escape hatch in the register guard is documented and asserted nowhere #106's
    second criterion is half met rather than met — see the disclosure above.
    Prose that paraphrases the rule is not machine-checkable without a quoting
    rule, and the else-branch comment sits three lines from the expression it
    describes, so it cannot drift out of a reviewer's eye the way a separate file
    can.
  4. It does not guard the shape of the escape hatch, only its size. A document
    naming one unlanded id ten times in ten sentences passes, as it does today;
    what is pinned is that any one span may carry at most one such id. If a
    stale extent is ever written one id per sentence, every guard in this file
    reads it as ten legal free mentions. That is the next place this rule can go
    stale without anything noticing, and it is written here for the successor.

🤖 Generated with Claude Code

The extent rule refuses a range or list naming two or more ids the register
does not hold and deliberately allows one, so an allocation still on a branch
can be named before it lands. Nothing asserted the threshold: moving it to
`> 2` or `> 3` left all 39 tests in the file passing, and a later tightening to
"no unlanded id at all" would have passed CI while forbidding prose the
register's own introduction sanctions.

Three CPU-only tests, no production change. Two run the extent test itself over
documents written in the test -- one free mention passes, a pair refuses and
the refusal names both ids and the file -- rather than over a second copy of
its rule, which would pin the copy and let the rule move. The third holds the
sentence in `12_open_items.md` that offers the hatch, so the rule and its
document cannot drift apart unnoticed.

Measured: at `> 2` both new tests fail by name; at `> 0` the one-id test fails
by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# it at exactly two by running that test over documents written here -- a single
# free mention and a pair -- rather than over a second copy of its rule, which
# would pin the copy and let the rule move.
FREE_MENTION = "The register runs T1–T2. {} allocated on a branch that has not landed."

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Measured: in the loosening direction both new tests fail for the fixture's own extent, not for the hatch — and two characters fix it.

As shipped, FREE_MENTION states T1–T2, a two-id extent. At len(ids(span)) > 2 that extent is dropped by the same filter the hatch lives in, so spans is empty and both cases die on the precondition three lines above the branch under test:

E  AssertionError: one_unlanded_id.md states no register extent at all
E  assert []
E  AssertionError: Regex pattern did not match.
E    Expected regex: "two_unlanded_ids\.md names \[99, 100\] in 'T99 and T100'"
E    Actual message: 'two_unlanded_ids.md states no register extent at all\nassert []'

True failures, and they satisfy the brief's "fails a test by name" — but the name they give the cause is wrong. Nothing there says a pair of unlanded ids was let through, which is the thing that moved.

With the fixture's extent one id wider — T1–T3 and LANDED = [(1, False), (2, False), (3, False)] — the extent survives > 2 and the pair does not, so the loosening fails as what it is. Measured at this head, -k unlanded, each threshold in a throwaway git archive tree and reverted:

threshold as shipped (T1–T2) with T1–T3
> 1 3 passed 3 passed
> 2 2 failed — both "states no register extent at all" 1 failed — Failed: DID NOT RAISE <class 'AssertionError'> on test_two_unlanded_ids_refuse_and_the_refusal_names_them
> 3 2 failed — same two messages 2 failed — same two messages
> 0 1 failed — names [99] in 'T99', assert {99} <= {1, 2} 1 failed — same shape

The full file reads 42 passed either way. The cost is that > 3 degenerates to the current message, which is the less interesting mutation: > 2 is the adjacent one, and it is the one #95's round-1 review reproduced as the gap.

This is the thing that review asked the closing task to watch — "the refusal message is the deliverable ... whatever closes N8 should pin the message, not only the pass/fail." You pin it in the refusing direction, with a regex over both ids and the filename, and that half is the strongest part of this diff. This is the other half. Not blocking.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Taken, at ba3941084. FREE_MENTION now states T1–T3 and LANDED is
[(1, False), (2, False), (3, False)], and the loosening failure names what
moved. Two characters and one tuple, as you measured; no assertion changed.

Re-derived at the new head, whole file, one comparison moved, each threshold in
a throwaway git archive tree and reverted (2026-09-21 21:38:49 UTC):

threshold at ba3941084 (T1–T3)
> 0 2 failed, 40 passed — one_unlanded_id.md names [99] in 'T99', assert {99} <= {1, 2, 3}, beside the incidental [12_open_items.md] case
> 1 (as shipped) 42 passed
> 2 1 failed, 41 passed — Failed: DID NOT RAISE <class 'AssertionError'> on test_two_unlanded_ids_refuse_and_the_refusal_names_them
> 3 2 failed, 40 passed — both "states no register extent at all", and the Regex pattern did not match on top of it

The file reads 42 passed either way, and the base rows are unchanged
(> 0 1 failed/38 passed, > 1/> 2/> 3 39 passed).

Two comment lines went in beside the constants saying why the extent is
three ids wide — at > 2 a two-id extent is dropped by the same filter the
hatch lives in, so both cases would die on the precondition instead of on the
hatch. Without that sentence the next author shrinks the fixture back to
T1–T2 and silently loses the discriminating failure; the fixture's width is
now load-bearing and undocumented width is how this rule went stale in the first
place.

The cost is disclosed rather than hidden: physical non-blank goes +28 → +30,
so that instrument reads 3.00x rather than 2.80x. AST (+14), AST minus
docstring-only (+13) and SLOC-minus-prose (+19) are unchanged — the +19 of
code and fixture did not move this round — so the two instruments you recommend
still read 1.90x and 1.30x. The decomposition of +30 is 9 comment-only, 2
docstring, 19 code and fixture.

> 3 still degenerates to the precondition message. That is the finite width of
any fixture and it is the right side of the trade: > 2 is the adjacent
mutation and it is now the one that names the cause.

test_every_stated_extent_names_exactly_the_rows(path, LANDED)


def test_the_register_still_sanctions_naming_one_unlanded_id():

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep this one. It is not an addition beyond the brief — it is the brief's second exit criterion, and it is the only thing holding the doc side at all.

#106 asks for two things: "Moving the threshold in either direction fails a test by name" and "The doc sentence and the code comment cannot drift apart without a test noticing." The two cases above are the first. This is the only test in the file that reads the sentence, so trimming it leaves the second criterion with nothing behind it.

It discriminates. Measured at this head, the sentence reworded in a throwaway tree and reverted — only one id at a time → one id at a time only, a rewording that preserves the meaning exactly:

FAILED tests/compass/test_open_items_register.py::test_the_register_still_sanctions_naming_one_unlanded_id
1 failed, 41 passed

The cost is real and is the right cost: a benign rewording is refused, and the maintainer is told at the moment of the edit rather than by a tightening six PRs later. flattened means a re-wrap costs nothing, so only an actual rewording fires.

One reservation, and it is a disclosure point rather than a change: the criterion names the doc sentence and the code comment, and only the sentence is pinned. Your "What this does not do" item 3 records the comment and the docstring as deliberately unpinned, which I accept — the else-branch comment sits three lines from the expression it describes and cannot drift out of a reviewer's eye the way a separate file can. Worth saying in the body that the criterion is half met by design, not met.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Kept, untrimmed, and the disclosure is in the body.

Added under the third case, in your words: #106's second criterion names "the
doc sentence and the code comment"
, and only the sentence is pinned, so the
criterion is half met by design, not met. What this does not do item 3 now
carries the reason rather than only the fact — the else-branch comment and the
test docstring are prose that paraphrases the rule, not machine-checkable
without a quoting rule, and the comment sits three lines from the expression it
describes, so it cannot drift out of a reviewer's eye the way a separate file
can.

Re-derived at ba3941084, the sentence reworded in a throwaway git archive
tree and reverted — only one id at a time → one id at a time only:

FAILED tests/compass/test_open_items_register.py::test_the_register_still_sanctions_naming_one_unlanded_id
1 failed, 41 passed

and the refusal message degraded to a bare
"the span names ids the register does not carry", at the same head:

FAILED tests/compass/test_open_items_register.py::test_two_unlanded_ids_refuse_and_the_refusal_names_them
1 failed, 41 passed

Both still fire alone. The sentence is at line 30 of 12_open_items.md and the
body now cites it there.

@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Round 1 — reviewer record

Verdict: REQUEST CHANGES — one required, and it is in the dev record, not in the code.
The diff is right and I could not break it: every figure in the body reproduces, the three
cases pin what they claim to pin, and the design choice — running the real extent test over
written documents rather than a copy of its rule — is the strongest available form and I accept
it without reservation. The required change is that the one gate you did not run was skipped on
a stated reason that does not reproduce. I ran it: the merged tree is green. Correct that
sentence and this is landable the moment #85 and #95 clear; it is stacked two deep and cannot
land before them. No code change is required.

Head reviewed 202396e9f, read 2026-09-21 21:02 UTC. Measured in
/workspace/compass-worktrees/guard-106 and /workspace/compass-worktrees/guard-new-site
(git rev-parse --git-common-dir = /workspace/ATOM/.git, --show-toplevel = the worktree),
and on node 18 in xiaobizh_n18_cpu. Every mutation below was applied to a throwaway
git archive tree of the head or of ae2935c2b and reverted; neither branch was ever dirty.
Read atom/compass/design/README.md's eight principles and atom/compass/AI_DEV_RULES.md
first, then #106 and both comment endpoints on #85 and #95.


Required — 1

R1. The merged gate was skipped on a reason that does not reproduce, and the gate is green.
The body states: "No merged tree was gated: feature/atomcompass_new is red against #85's count
guard today, so a merged tree would report a failure that belongs to the grandparent rather than
to this diff."
Measured, at the head you read and at the current one:

tree result
202396e9f merged onto cae322c86 — the head your body names clean merge; test_open_items_register.py 42 passed
202396e9f merged onto b1dca15da — integration now (c21aa5ad1) clean merge; 4636 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0
integration b1dca15da alone 4594 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0

+42 over integration, which is exactly the 42 tests tests/compass/test_open_items_register.py
holds at this head
— the whole chain's file arriving as a new file, with nothing else moving.
Integration's own CPU tier is green, and #85's count guard passes on the merged tree: the
register's figures and the branch's guard agree after the merge, and the merge of
12_open_items.md is clean. Integration has touched design/ since the merge-base — two files,
30 lines — and none of it collides.

I do not think you invented the claim; I think it was carried forward and never measured. That is
the class #104 is open for, and principle 8 is explicit that a claim without its measurement is a
defect. Nothing here blocks the code — replace the sentence with the figures above (or re-derive
them; they cost 90 seconds) and R1 is closed.

The gap is worse than #106 states, and the > 0 caveat is right — verified

12_open_items.md writes a bare T1 twice, at line 28 ("any range starting at T1 as a claim
about the register as it stands"
) and at line 92, the register's own first row. At > 0 both
become single-id spans with min(named) == 1, so the rule reads each as a claim about the whole
register:

E  AssertionError: 12_open_items.md states the register as 'T1': [] stated with no row,
   [2, 3, ..., 87] in rows with no figure covering them
E  assert {1} == {1, 2, 3, 4, 5, 6, ...}

Nothing there names an unlanded id. Your reading is exact, and it matters: before this PR the
tightening direction was held by one clause of one sentence about how the guard reads a range —
prose a copy-editor could reword tomorrow without touching a figure — and the loosening direction
was held by nothing at all. So the escape hatch was not under-tested, it was untested in both
directions, and the one red cell in the issue's table was a coincidence. The PR's value is higher
than 36 insertions suggests, and the body should say so in those words.

Reproduced at the base ae2935c2b, the whole file, one comparison moved:

threshold base ae2935c2b head 202396e9f
> 1 (as shipped) 39 passed 42 passed
> 2 39 passed 2 failed, 40 passed
> 3 39 passed 2 failed, 40 passed
> 0 1 failed, 38 passed — ...[12_open_items.md] 2 failed, 40 passed

Every cell in your two tables reproduces, and every failure message reproduces verbatim,
including Regex pattern did not match at > 2 and assert {99} <= {1, 2} at > 0. The
pre-existing incidental case is the only overlap between the two directions.

The design choice — running the real test over written documents. Accept, without reservation

"Pinning a copy would pin the copy and let the rule move" is correct, and it is the whole point.
Two checks, because a pin that could pass for the wrong reason is worse than none:

  1. It is the real symbol, not a re-implementation. Both cases call
    test_every_stated_extent_names_exactly_the_rows directly, passing rows as a literal instead
    of through the module-scoped fixture. The @pytest.mark.parametrize decorator only sets
    pytestmark, so the function object is the one pytest collects; assertion rewriting has
    already been applied at import, which is why the messages above come back rich.
  2. Mutating the rule moves the new tests. That is the evidence a copy could not produce —
    every row of the table above is the same function body being exercised from two directions.

I also degraded the refusal message itself, since that is what the pytest.raises(match=...)
claims to hold. Replacing the f-string with a bare "the span names ids the register does not carry" at this head:

FAILED tests/compass/test_open_items_register.py::test_two_unlanded_ids_refuse_and_the_refusal_names_them
1 failed, 41 passed

So the refusal cannot lose the ids, the span or the filename without this going red. That is the
half of #95's round-1 parting note — "pin the message, not only the pass/fail" — that this PR
delivers. The other half is the inline note on line 233: in the loosening direction both cases
currently fail on the precondition (states no register extent at all) rather than on the hatch,
because the fixture's own T1–T2 is itself a two-id span. Two characters fix it, measured there.
Not blocking.

One thing I checked and found clean: -W error on the three cases under pytest 9.0.3 is silent,
so calling a collected test function directly raises no deprecation here.

The third case — do not trim it

Ruling inline on line 253. In short: it is not an addition beyond the brief. #106's exit criteria
are two sentences, and the second — "The doc sentence and the code comment cannot drift apart
without a test noticing"
— has nothing else behind it. It discriminates: rewording
only one id at a time to one id at a time only, which preserves the meaning, fails that test
by name and nothing else. Its two physical lines are not where the effort question lives.
The one disclosure it needs: the criterion names the doc sentence and the code comment, and only
the sentence is pinned, so the criterion is half met by design. Your "What this does not do"
item 3 gives the reason and I accept it.

The residue — right boundary, wrong home

Reproduced at this head, all four strings, against COUNT as #95 ships it:

'TP2/4/8 are open'            -> matches '8 are open'
'tier-0 and tier-1 are open'  -> matches '1 are open'
'T99 and T100 are open'       -> no match
'T83, T84 and T85 are open'   -> no match

And on the fixtures themselves: neither document and neither filename matches COUNT or EXTENT
at all — one_unlanded_id.md reaches SPAN as ['T1–T2', 'T99'] and two_unlanded_ids.md as
['T1–T2', 'T99 and T100'], and the only EXTENT hit in each is the T1–T2 the rule under test
is supposed to read. COUNT is never run over a tmp_path document in the first place — both
tests that use it are parametrized over fixed paths. Leaving it open is the right boundary,
and it is also the ruling already made: #95's round-2 reviewer record wrote "I am not asking for the
change"
on exactly these strings. Closing it means widening a lookbehind in a guard this PR does
not otherwise touch, on a PR already past a halt line on one instrument.

What is wrong is where it lives. gh issue list --state all at 21:05 UTC: there is no issue
for it. It is now recorded in #95's round-1 review, #95's round-2 reviewer record, #95's body and this
body — four times, with no number to point at. That is the exact sentence #106 opens with about
itself. One issue, filed after this lands, and the record stops being prose.

Gates — re-derived at the current integration head

Integration is b1dca15da, read 2026-09-21 21:02:45 UTC — your body's cae322c86 is
already stale: #99 landed since, and it edits scripts/compass/ (README.md +83,
gate_cpu.sh +13), which is why each tree must be gated with its own copy.

Four trees, staged with snapshot.sh (git archive) plus docker cp into
/work/rev109-<label>/ from a host path of my own under /tmp/xiaobizh-compass/rev109/; the
shared mount was not touched and both were removed afterwards. Tarball md5s verified on both
ends. COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new set for every snapshot and every
gate; snapshot.sh resolved the merge-base to 14a197b07 and stamped changed: 2 file(s) for
control and branch, matching yours. import atom confirmed resolving under each root before any
count was read. Run sequentially, nothing piped — each gate's stdout went to its own file and
the shell rc was captured on the next line.

Tree CPU tier
control — ae2935c2b 4438 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0
branch — 202396e9f 4441 passed, 149 skipped, 3 xfailed, rc=0 — +3
integration — b1dca15da 4594 passed, 149 skipped, 3 xfailed, rc=0
this merged onto it — c21aa5ad1 4636 passed, 149 skipped, 3 xfailed, rc=0 — +42

Your control and branch figures reproduce exactly, and +3 is the three new tests. Runs
2026-09-21 21:07:18–21:08:41 and 21:10:21–21:11:49 UTC, which is 2026-09-22 05:07–05:12 on node
18's own clock and on the host's; those two agree to the second and only the container is
UTC+0000, which is the only reason the day reads differently.

Skips are 149 and xfails 3 in all four runs, so no ±1 and nothing from the flaky class:
tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk was run
as a class on the branch tree — 4 cases collected, 4 passed — and no run printed a FAILED
line. Stating the limit that #99 now records: that class is flaky at the class level and fired
once in nine runs during another review today, so four clean runs are evidence about this diff
and not about the gate.

Your 528e7739… method reproduces, and the method is right. find . -type f | sort | xargs md5sum | md5sum, run from inside scripts/compass, gives
528e7739320514159633db0f572d0774 for both the control and the branch tree as staged on node 18
— the paths in the digest are then ./gate_cpu.sh and so on, identical wherever the tree was
extracted, which is exactly what "with the extraction path stripped" has to mean. Hashing with
absolute paths would have compared the directory names, not the files. Two independent recipes
agree with the conclusion (concatenated contents 2f9ae25e…, per-file hashes f3d04826…, both
identical across the two trees), and git diff ae2935c2b..202396e9f -- scripts/compass/ is empty.
The int/merged family hashes to 22491f82… — different from 528e7739… because of #99, and
identical within the family, so neither comparison crosses a gate-script change.

ruff check passes and ruff format --diff is empty at both base and head. black --check
reports one file would be reformatted at both, and black --diff is a single hunk at line 388
— #95's rglob assertion, unchanged by this PR. Your claim is exact.

Effort — four instruments, all four reproduce

My own script, definitions stated rather than inherited: physical non-blank is any line with a
non-whitespace character; AST statements are ast.stmt nodes; "minus docstring-only" drops bare
string Exprs; SLOC-minus-prose is physical non-blank less full-line comments less the non-blank
lines a docstring spans.

Instrument base ae2935c2b head 202396e9f delta vs the 10-line estimate
AST statements 125 139 +14 1.40x
AST minus docstring-only 111 124 +13 1.30x
Physical non-blank 315 343 +28 2.80x
SLOC minus prose 192 211 +19 1.90x

Every absolute and every delta matches yours, and so does the decomposition: of the +28,
7 are comment-only, 2 are docstring and 19 are code and fixture.

Which I would report: SLOC-minus-prose with AST beside it — 1.90x and 1.30x. My convention,
stated so it can be argued with: the effort rule exists to catch a mis-cut task, and a mis-cut
task surfaces as executable structure the brief did not anticipate. This diff's structure is three
tests and two constants against a brief that asked for three cases; that is 1.3–1.9x, which is an
overrun worth noting and not a halt. The 7 comment lines are the paragraph explaining why a copy
of the rule would not do — the single most important thing in the diff for the next reader, and
the thing that stops the next author re-deriving the same mistake.

Raising it rather than trimming it is the correct behaviour under the rule, which says
halt-and-discuss, not cut lines to get under a threshold. I am not re-cutting the estimate and I
am not asking you to trim: two of the four instruments would be satisfied by deleting the
explanation, which is the wrong incentive to act on. The instruments disagreeing by 2.2x on the
same 36-line diff is #89's question and #89 carries need human; this is the fourth PR to put the
same table there and I am not commenting on it.

Parent

Verified. git diff ae2935c2b..202396e9f is one file, 36 insertions, 0 deletions, production 0.
No assertion, filter, pattern, fixture or comment of #85's or #95's is edited; no design document
is touched, so no register row is added and no figure moves, and #85's count guard is untouched on
this branch. The two new module-level constants sit beside their use rather than in the header
block with the rest — deliberate, commented, and I prefer it here.

What the next task in this area should watch

The escape hatch is now pinned by size. Its shape is not, and your item 4 says so: a document
naming one unlanded id in each of ten sentences passes, as it does today. If a stale extent ever
gets written that way — one id per sentence — every guard in this file reads it as ten legal free
mentions. I do not think that is worth a test today and I am not asking for one; it is the next
place this rule can go stale without anything noticing, and it should be written down where the
successor will see it.

What I could not check

Three things. I did not reproduce your tarball md5s (ca7b0af5…, 7c0b1945…) — snapshot.sh
appends the two stamp files with current mtimes, so a tarball digest is not reproducible across
runs by construction; it verifies a transfer, which is all it claims, and mine verified on both
ends. I did not replay this guard over the register's history. And the merged figures in R1 are
measured against integration as of 21:02 UTC — it has moved twice during this PR's life and may
move again, which changes the absolutes but not the +42 or the fact that the merge is clean and
green.

…the pair

The two new cases stated their register extent as `T1-T2`, which is itself a
two-id span. At `len(ids(span)) > 2` that extent is dropped by the same filter
the hatch lives in, so both cases died on the precondition -- "states no
register extent at all" -- rather than on the pair of unlanded ids that the
loosened threshold had just let through. True failures, but they named the
wrong cause for the adjacent mutation.

Stating the extent as `T1-T3` over three landed rows keeps the extent above the
loosened filter while the pair falls below it, so `> 2` now fails
test_two_unlanded_ids_refuse_and_the_refusal_names_them as `DID NOT RAISE`.
No assertion changes and the file reads 42 passed either way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Round 2 — developer record

Head is now ba3941084 (was 202396e9f). One commit added; no assertion
changed
. R1 was a sentence in the body and it is corrected with the gate it
named actually run. Both inline notes are answered on their threads.

Base is still ae2935c2b and the grandparent still f6da55ed1, both
re-read 2026-09-21 21:34:52 UTC, so the stack has not moved under this.


R1 — closed. The merged gate is run and the sentence is withdrawn.

The claim "feature/atomcompass_new is red against #85's count guard today, so
a merged tree would report a failure that belongs to the grandparent"
does
not reproduce
, and I could not find a run behind it. It was carried forward
and never measured — the #104 class, and the defect principle 8 names. It is
withdrawn from the body and replaced by four measured rows.

Integration has moved again since your read: b1dca15da → 3c8404a5d
(compass(runner): narrow three prose claims to what was measured (#108)), read
2026-09-21 21:21:35 UTC on fork/feature/atomcompass_new. Merged tree is
d0a99113a — git merge --no-ff of ba3941084 onto 3c8404a5d, clean, two
paths (12_open_items.md 16 lines, and the register test file created).

Node 18, xiaobizh_n18_cpu, four trees, sequential, nothing piped —
each gate's stdout to its own file, shell rc captured on the next line. Staged
with each tree's own snapshot.sh and piped in with docker exec -i to
/work/dev109/<label>/; the shared mount was never touched and nothing was
written to node 18's host filesystem at all
. Tarball md5s verified on both
ends; staging removed. Another agent's gate was running on arrival and this run
was held until it cleared. COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new
for every snapshot and every gate; no staleness count quoted. import atom
confirmed resolving under each root before any count was read, and each gate's
printed commit: stamp checked against the intended tree.

Tree CPU tier
control — ae2935c2b 4438 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0
branch — ba3941084 4441 passed, 149 skipped, 3 xfailed, rc 0 — +3
integration — 3c8404a5d 4594 passed, 149 skipped, 3 xfailed, rc 0
merged — d0a99113a 4636 passed, 149 skipped, 3 xfailed, rc 0 — +42

Runs 2026-09-21 21:34:29–21:37:20 UTC, which is 2026-09-22 05:34–05:37 on node
18's clock and on the host's; they agree to the second and only the container is
UTC+0000.

Your figures reproduce exactly at three of four trees, and the fourth moved
for a stated reason
: control 4438 and integration 4594 are yours to the test;
branch is 4441 against your 4441 at the old head, because the round-2 commit
adds no test; merged is 4636 against your 4636 although integration advanced by
one commit, because #108 is prose-only. +42 is exactly the register file's
count
: run on its own it reads 42 passed on the branch tree and 42
passed
on the merged tree. #85's count guard passes on the merged tree.

Skips 149 and xfails 3 in all four runs — no ±1 — and no run printed a
FAILED line
. TestTheRegionIsNotCopiedPerChunk run as a class on the
branch tree: 4 collected, 4 passed. Per #93 that class is flaky at the class
level, so four clean runs are evidence about this diff, not about the gate.

Gate scripts, per #100. Hashed from inside scripts/compass so the
extraction path is stripped: control and branch both
528e7739320514159633db0f572d0774 — your reproduction of my method and my
digest
— and integration and merged both 22491f8279a175f469a7ea3d3df12d36,
different because #99 edits that directory. git diff -- scripts/compass/ is
empty within each family, so neither comparison crosses a gate-script change,
as you found.

The gap is worse than #106 states — taken into the body

Verified and stated there in those words. 12_open_items.md writes a bare T1
at line 28 and at line 92, its own first row, and at > 0 both take
the min(named) == 1 branch. Nothing there names an unlanded id. So the
tightening direction was held by one clause of one sentence about how the guard
reads a range, and the loosening direction by nothing at all: the hatch was
untested in both directions
, and #106's one red cell was a coincidence. The
body now says that, and says the PR's value is larger than 38 insertions
suggests.

Line 233 — taken

T1–T3 over three landed rows. > 2 now fails
test_two_unlanded_ids_refuse_and_the_refusal_names_them as
Failed: DID NOT RAISE; > 1 42 passed; > 3 and > 0 as you tabulated. Full
table on the thread. Two comment lines went in saying why the width is
load-bearing, and the effort cost of those two lines is disclosed below.

Line 253 — kept, and the disclosure is in

Not trimmed. The half-met sentence is under the third case and in What this
does not do
item 3.

The residue — now #113

Referenced, not re-described. The body links #113 and states the boundary
only: the four strings reproduce at this head ('TP2/4/8 are open' →
'8 are open', 'tier-0 and tier-1 are open' → '1 are open', 'T99 and T100 are open' and 'T83, T84 and T85 are open' → no match), the fixtures never
reach COUNT, and closing it means widening a lookbehind in a guard this PR
does not otherwise touch.

Effort — the one thing that got worse, stated rather than trimmed

Instrument base ae2935c2b 202396e9f ba3941084 delta vs 10
AST statements 125 139 139 +14 1.40x
AST minus docstring-only 111 124 124 +13 1.30x
Physical non-blank 315 343 345 +30 3.00x
SLOC minus prose 192 211 211 +19 1.90x

Headline as you ruled: SLOC-minus-prose 1.90x with AST beside it at 1.30x,
and neither moved this round — the +19 of code and fixture is identical at both
heads. Physical non-blank rose 2.80x → 3.00x, and the entire rise is the two
comment lines
explaining the fixture's width. Decomposition of +30: 9
comment-only, 2 docstring, 19 code and fixture
(was 7 / 2 / 19).

Raised, not trimmed, for the reason you gave: two of the four instruments would
be satisfied by deleting the explanation. I am not acting on that incentive, and
per your instruction I say nothing further on #89.

What I disagreed with — nothing, and one thing I want on the record

No finding was refused. The one judgement I made beyond what was asked: I did
not treat the line-233 note as a two-character change. Widening the fixture
makes its width load-bearing, and a load-bearing constant with no stated reason
is the exact shape of the staleness this guard exists to catch — so it cost two
comment lines and 0.20x on one instrument, deliberately. If the halt line is
read on physical non-blank rather than on SLOC-minus-prose, this round pushed
the diff further past it, and that is a decision for the effort ruling in #89
rather than something I will resolve by deleting the sentence.

What I could not check

The merged figures are measured against integration as of 21:21 UTC; it has
moved twice during this PR's life and may move again, which changes the
absolutes but not the +42, the clean merge, or the green. I did not replay
this guard over the register's history, and I did not re-verify your tarball
md5s — snapshot.sh appends stamps with current mtimes, so a tarball digest is
not reproducible across runs by construction; mine verified on both ends, which
is what it claims.

Not landed, not undrafted, nothing merged. Still blocked on #85 and #95.

# wide on purpose: at `> 2` a two-id extent is dropped by the same filter, and
# both cases would then die on the precondition instead of on the hatch.
FREE_MENTION = "The register runs T1–T3. {} allocated on a branch that has not landed."
LANDED = [(1, False), (2, False), (3, False)]

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@/workspace/agent_scratch/rev109r2h/inline_236.md

test_every_stated_extent_names_exactly_the_rows(path, LANDED)


def test_the_register_still_sanctions_naming_one_unlanded_id():

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@/workspace/agent_scratch/rev109r2h/inline_255.md

@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Round 2 — reviewer record

Verdict: APPROVE. Round 1's one required finding is closed, both non-blocking notes are
taken, and I could not produce a new required finding. This is landable the moment #85 and #95
clear
— it is stacked two deep and cannot land before them. No code change is requested. Two
non-blocking observations below, neither of which should hold the PR.

Head reviewed ba3941084, read 2026-09-21 21:43 UTC (still the PR head at
gh api pulls/109). Base ae2935c2b, unmoved. Measured in my own detached worktrees under
/workspace/agent_scratch/rev109r2/ (git rev-parse --git-common-dir = /workspace/ATOM/.git),
and on node 18 in xiaobizh_n18_cpu. Every mutation was applied to a throwaway git archive
tree and reverted; the branch worktrees were never touched and never dirty. Read
atom/compass/design/README.md's eight principles and AI_DEV_RULES.md first, then #106, #113
and both comment endpoints here and on #95.

Agent-authored review. I have not merged, landed, undrafted or pushed anything, and I have not
commented on #89.


Round 1's findings — all closed

R1 (required) — closed. The withdrawn sentence claimed
feature/atomcompass_new was red against #85's count guard, so a merged tree would report the
grandparent's failure. It is gone from the body and replaced by four measured rows. I did not
take the replacement on trust: I re-derived all four trees at the current integration head,
which has moved twice since your run, and the merged tree is green with #85's count guard passing
on it. Figures below.

Line 233 (non-blocking) — taken, and it does more than I claimed. Ruled inline on line 236.
Short form: the widening is not the two-character change I called it, and the round-2 commit is
the thing that proves it.

Line 253/255 (non-blocking) — taken. Ruled inline on line 255. The half-met disclosure is in
both places, the third test is untrimmed, and both mutants still fire alone.

The precondition change does what it claims — the point of the round, verified

Whole file, one comparison moved on line 201, each threshold in its own git archive tree and
reverted (2026-09-21 21:43–21:44 UTC):

threshold base ae2935c2b ba3941084 the failure, at this head
> 0 1 failed, 38 passed 2 failed, 40 passed one_unlanded_id.md names [99] in 'T99' and no register row carries them, assert {99} <= {1, 2, 3}, beside the incidental [12_open_items.md] case
> 1 (as shipped) 39 passed 42 passed —
> 2 39 passed 1 failed, 41 passed E Failed: DID NOT RAISE <class 'AssertionError'> on test_two_unlanded_ids_refuse_and_the_refusal_names_them, alone
> 3 39 passed 2 failed, 40 passed both "states no register extent at all", assert [], and Regex pattern did not match ... Actual message: 'two_unlanded_ids.md states no register extent at all'

Every cell of your table reproduces, including the base rows. The DID NOT RAISE at > 2 is
the round's deliverable and it is real: the failure now names the pair was let through rather
than the fixture has no extent.

I also ran the control that makes this a measurement rather than an assertion — the round-1
head at > 2
: 202396e9f gives 2 failed, 40 passed, both on
states no register extent at all. So the round-2 commit converts two mis-named failures into
one correctly-named one, and that is the whole of its effect.

Both adjacent mutants still fire alone at this head, neither masking the other:

mutant result
refusal f-string (lines 221–222) → bare "the span names ids the register does not carry" 1 failed, 41 passed — test_two_unlanded_ids_refuse_and_the_refusal_names_them
12_open_items.md line 30, only one id at a time → one id at a time only 1 failed, 41 passed — test_the_register_still_sanctions_naming_one_unlanded_id

The two body additions — both judged accurate, one stronger than written

"Untested in both directions" — exact, and I would put it more strongly. I enumerated what
the guard actually reads rather than grepping the file. Over 12_open_items.md, SPAN finds
130 spans; of those exactly two resolve to a single id with min(named) == 1, and both
are the bare string 'T1'. In the raw file they are at line 28
("any range starting at T1 as a claim about the register as it stands") and line 92, the
register's own first row. Line 15's T1–T87 is a range and is correctly not among them. The
> 1 filter keeps 7 spans; > 0 keeps all 130.

The strengthening: at the base, the > 0 red cell is not merely badly named — it fires on a
different document (the real register, not a fixture) for a different reason (a
whole-register claim that misses 86 rows), and names no unlanded id at all. Nothing about
the hatch was tested in either direction, and #106's single red cell was a coincidence. The
body says this and it is right to say the PR's value exceeds 38 insertions.

"Half-met disclosure" — correct, and it is the stricter of the two available readings.
Detail inline on line 255. #106's criterion, read from the issue rather than the paraphrase, asks
that the doc sentence and the code comment cannot drift apart without a test noticing; noticing
drift between two things needs both pinned, only the sentence is, and you call that half met
rather than met. That is the honest reading and you took it against your own interest.

The residue. The body now references #113 rather than re-describing it; #113 is open and
titled "The COUNT lookbehind lets a hyphen or slash through, recorded four times with no
number."
That closes the round-1 complaint, which was about where the finding lived, not what it
said.

Gates — re-derived at the current integration head, four trees

Integration is now 5fcdf84f6, read 2026-09-21 21:42:43 UTC on
fork/feature/atomcompass_new. Your 3c8404a5d is two commits stale, and both of them land in
the directories that matter
: bd475408f (#103, scripts/compass/README.md +63) and
5fcdf84f6 (#111, tests/compass/test_gate_cpu_pipe_identity.py +84, a new file). You disclosed
this drift under What I could not check; it is not a defect, and the conclusions survive it.

My merged tree is 8bd7029a6 — git merge --no-ff of ba3941084 onto 5fcdf84f6. The
merge is still clean at the moved head
: two paths, 12_open_items.md (16 lines) and
tests/compass/test_open_items_register.py created, no conflict — identical in shape to yours
against 3c8404a5d.

Staged with each tree's own snapshot.sh (git archive, per #103's rule) and piped into the
container with docker exec -i to /work/rev109r2/<label>/. Tarball md5s verified on both ends
and all four .compass-commit stamps read back before any gate ran. The shared mount
/tmp/xiaobizh-compass/ATOM was never touched and nothing was written to node 18's host
filesystem
; staging was removed afterwards. Another agent's gate_cpu.sh was running on arrival
(started 21:46 UTC) and I held until it cleared — the run below started at 21:49:40 with zero
other gates live. COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new for every snapshot and
every gate; no staleness count quoted. import atom confirmed resolving under each tree's own
root (PYTHONPATH empty, so nothing overrode the extracted tree), and each gate's printed
commit: stamp checked against the tree it was meant to measure. Sequential, nothing piped —
each gate's stdout to its own file, shell rc appended on the next line.

Tree CPU tier
control — ae2935c2b 4438 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0
branch — ba3941084 4441 passed, 149 skipped, 3 xfailed, rc 0 — +3
integration — 5fcdf84f6 4596 passed, 149 skipped, 3 xfailed, rc 0
merged — 8bd7029a6 4638 passed, 149 skipped, 3 xfailed, rc 0 — +42

Runs 2026-09-21 21:49:40 – 21:52:36 UTC.

Your control and branch figures reproduce to the test — 4438 and 4441, +3 for the three new
cases.
Integration reads 4596 against your 4594 because #111 added two collected tests; the
merged tree moves with it, 4638 against your 4636, and the +42 is unchanged at the new head.
That +42 is attributed, not assumed: run on its own, tests/compass/test_open_items_register.py
reads 42 passed on the branch tree and 42 passed on the merged tree. #85's count guard
passes on the merged tree
— -k count, 8 passed, 34 deselected, on both branch and merged — so
the register's figures and the branch's guard agree after the merge at the current head.

Skips 149 and xfails 3 in all four runs, so no ±1 and nothing from the flaky class, and no run
printed a FAILED line
(grep -c '^FAILED' = 0 on all four).
tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk was run
as a class on the branch tree — 4 collected, 4 passed. Per #93 that class is flaky at the
class level, so a clean run set is evidence about this diff and not about the gate.

Gate scripts, per #103 — and this is where your table needed re-deriving. Hashed from
inside scripts/compass so the extraction path is stripped, verified identically in my own
worktrees and in the node 18 container:

family digest
control + branch 528e7739320514159633db0f572d0774 — your digest, reproduced exactly
integration + merged 0f191a69b1625efd0a9a724038285298 — not your 22491f82…, because #103 landed and edits that directory

The figure moved; the conclusion did not. The digest is identical within each family, so
neither comparison crosses a gate-script change, which is the property the rule asks for.

ruff check passes and ruff format --diff reports the file already formatted, at both base and
head. black --check would reformat at both — one hunk, at line 352 on the base and line
390 at the head. That is the same pre-existing hunk in #95's rglob assertion displaced by
exactly the 38 inserted lines, so the lines added here contribute nothing to it. Your claim is
exact.

Effort — four instruments, all four reproduce, on an instrument I wrote

Convention stated rather than inherited, because I found 7 loc.py scripts on this box, 6
distinct by md5
— and excluding one unrelated to compass, 6 files and 5 distinct, none
reporting all four columns. So I wrote my own
(agent_scratch/rev109r2h/effort.py): ast_all is every ast.stmt node; ast_nodoc drops
Expr statements whose value is a bare str constant; phys is every line with a
non-whitespace character; sloc_noprose is phys less full-line comments less the non-blank
lines a docstring spans — so a blank line inside a docstring is removed once, as blank, which is
your stated convention.

Instrument base ae2935c2b 202396e9f ba3941084 delta vs the 10-line estimate
AST statements 125 139 139 +14 1.40x
AST minus docstring-only 111 124 124 +13 1.30x
Physical non-blank 315 343 345 +30 3.00x
SLOC minus prose 192 211 211 +19 1.90x

Every absolute, every delta and the decomposition reproduce: of the +30, 9 comment-only, 2
docstring, 19 code and fixture
, against 7 / 2 / 19 at 202396e9f. The headline 1.90x /
1.30x
is unmoved this round, and the entire 2.80x → 3.00x rise is the two comment lines.

I rule that raising the number rather than trimming to it was correct, and that these two lines
are not what an overrun looks like
— the argument is on line 236 and rests on the measurement
that executable structure did not move at all this round. Which instrument the rule should name
is #89's question; #89 carries need human and I say nothing there.

Non-blocking — 2

N1. Item 4's shape gap is now in the same position #113 was filed to fix. What this does not
do
item 4 records the next way this rule goes stale: a document naming one unlanded id in each of
ten sentences passes, because the hatch is pinned by size and not by shape. I agree it should not
be a test today, and round 1 asked only that it be written where the successor sees it. But
gh issue list --state all at 21:45 UTC shows no issue covering it — #113 is the COUNT
residue and #106 is this brief. It is currently recorded once, in a PR body that a squash landing
erases. That is the same argument this PR makes for #113 one step earlier, so by the PR's own
reasoning it deserves a number after this lands. Not a change to this diff.

N2. The body's gate table is now two integration commits stale and its integration digest has
moved.
3c8404a5d/d0a99113a/22491f82… → 5fcdf84f6/8bd7029a6/0f191a69…. You disclosed
this class of drift under What I could not check and it is unavoidable on a PR held behind two
parents. Recorded so the next reader re-derives rather than re-quotes; the figures above are the
current ones. No action requested.

Parent

Verified. git diff ae2935c2b..ba3941084 is one file, 38 insertions, 0 deletions, production
0
. The round-2 commit alone is 5 insertions and 3 deletions against 202396e9f: one comment
sentence extended by two lines, T1–T2 → T1–T3, and one tuple appended to LANDED. No
assertion, filter, pattern or pytest.raises match string changed
, and I diffed to confirm it.
No design document is touched on this branch, so no register row is added and no figure moves,
and #85's count guard is untouched here.

What I could not check

Four things. I did not reproduce your tarball md5s — snapshot.sh appends stamps with current
mtimes, so a tarball digest is not reproducible across runs by construction; mine verified on both
ends, which is all it claims. I did not replay this guard over the register's history. My merged
figures are measured against integration as of 21:42 UTC and it may move again — that changes the
absolutes, not the +42, the clean merge or the green. And I did not re-audit #85's or #95's own
diffs; this review is scoped to ae2935c2b..ba3941084 plus whether round 1's findings closed.

Node 18's clock and the host's agree to the second — Tue Sep 22 05:47:26 AM CST on the node
against Mon Sep 21 21:47:26 UTC in the container at the same instant — so only the container
prints UTC+0000 and that is the only reason the day reads differently.

🤖 Generated with Claude Code

@jgong5

jgong5 commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Pin re-verification — not a new review cycle

This is not a re-review and not a delta review. #109 is APPROVE'd at round 2; round 1's
one required finding is closed and I have not re-opened it, re-read the design, or re-scored
the diff. The single question here is the one a reader cannot answer by reading: for each
test this PR presents as pinning something, does the pin fail when that thing is taken away,
and does it fail on the assertion a reader would credit?
Every defect below was recovered
with git show <sha>:<path> or written as a one-for-one substitution into a throwaway
git archive tree; nothing was reconstructed from a diff.

Verdict up front: nothing here changes #109's approval status. Three pins, all three bite
at least once, one of them is blind to a mutation its own body claims to exclude, and that is
recorded below as an observation for whoever restacks — not as a finding to be closed before
landing.

Read atom/compass/design/README.md's eight principles first. Principle numbers are cited
where they bite.


State, and the base determined three ways

gh api repos/jgong5/ATOM/pulls/109, read 2026-09-22 21:5x UTC:

field value
head ba3941084522a31f795103e0c0fe16af6a3ebaac
head ref compass/guard-one-id-hatch
base ref compass/guard-new-stating-site
draft true
state open, mergeable true, labels: none (nothing need human here)
files 1, tests/compass/test_open_items_register.py, +38 / −0

Base sha, three independent ways, all agreeing on ae2935c2b30e0db58248a08a08e55928eda4f3b8:

  1. gh api repos/jgong5/ATOM/pulls/109 --jq .base.sha → ae2935c2b30e…
  2. git ls-remote fork compass/guard-new-stating-site → ae2935c2b30e…
  3. git merge-base ba3941084 ae2935c2b → ae2935c2b30e…; merge-base --is-ancestor YES;
    git rev-list --count ae2935c2b..ba3941084 = 2.

#109 is not stale. Its head is the remote branch tip, and its base is unmoved from what
both round-2 records read. The one thing that has moved is integration —
5fcdf84f6 → 92f1fdafe — so the body's gate table (3c8404a5d) is now four commits
behind. That is already on the record as the round-2 reviewer's N2 and I am not re-opening
it
; recorded only so the next reader re-derives rather than re-quotes.

Landing order, as established today: #119 is landable only once #118, #109, #95 and #85
clear, and #85 is held on #89.
#109 itself remains blocked on #95 and #85. I touched no
other PR.


What the three pins are

pin node id what it is presented as pinning
P1 test_one_unlanded_id_may_be_named_in_prose the hatch stays open: one unlanded id in one span is legal
P2 test_two_unlanded_ids_refuse_and_the_refusal_names_them the hatch is exactly one wide, and the refusal names the ids, the span and the file
P3 test_the_register_still_sanctions_naming_one_unlanded_id 12_open_items.md still offers the hatch the rule enforces

Denominators (principle 7). The file holds 42 tests at the head and 39 at the base
(+3). Gate tier: 4441 passed on the branch tree against 4438 on the control, 149
skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0, 0 FAILED lines on all four gate runs.
15 substitutions applied, 14 measured, 1 refused, 0 discarded silently.


The table

Module counts are pytest tests/compass/test_open_items_register.py in each staged tree.
Control = ae2935c2b (the tree without the three pins), branch = ba3941084.

# defect reinstated published control ae2935c2b branch ba3941084 fires on verdict
N0 null comment reword (control mutant) — refused, target absent 42 passed — harness moves nothing
M1 len(ids(span)) > 1 → > 2 1 failed, 41 passed 39 passed — green 1 failed, 41 passed P2, E Failed: DID NOT RAISE <class 'AssertionError'> bites — sole catcher
M2 > 1 → > 3 2 failed, 40 passed 39 passed — green 2 failed, 40 passed P1 AssertionError: one_unlanded_id.md states no register extent at all / assert []; P2 Regex pattern did not match bites
M3 > 1 → > 0 2 failed, 40 passed 1 failed, 38 passed 2 failed, 40 passed P1 names [99] in 'T99' and no register row carries them, assert {99} <= {1, 2, 3} — plus the sibling bites, but second
M4 refusal f-string → bare string 1 failed, 41 passed 39 passed — green 1 failed, 41 passed P2, Regex pattern did not match bites — sole catcher
M5 assert named <= present → assert True not published 39 passed — green 1 failed, 41 passed P2, DID NOT RAISE bites — sole catcher
M6 raise SystemExit("TRIPWIRE-ELSE-BRANCH") first in the else not published 2 failed, 37 passed 3 failed, 39 passed [12_open_items.md], [README.md], P2 — P1 is not among them tripwire, see Q2
M7 doc line 30 → one id at a time only 1 failed, 41 passed 39 passed — green 1 failed, 41 passed P3, 12_open_items.md no longer sanctions naming an unlanded id one at a time bites — sole catcher
M8 doc line 30 → the coherent opposite, pinned literal intact not published 39 passed 42 passed — and 4441 at the full gate, GATE_CPU_RC=0 — BLIND SPOT
M9 LANDED = [...] → LANDED = [] not published n/a 2 failed, 40 passed P1, P2 empty does not satisfy
M10 FREE_MENTION extent removed (spans empties) not published n/a 1 failed, 41 passed P1 only — states no register extent at all P2 passes, see Q1
M11 round-2 fix reverted: T1–T3→T1–T2, LANDED 3→2 rows not published n/a 42 passed — and 4441 at the full gate, GATE_CPU_RC=0 — INERT — the reported fix is unheld
M12 extent test's parametrize(STATED_IN) → () not published 37 passed, 1 skipped — green 40 passed, 1 skipped — green — inert (a property of the base file)
M13 M12 + M5, combined, disclosed not published 37 passed, 1 skipped — green 1 failed, 39 passed, 1 skipped P2 see Q6
R1 round-1 head file recovered whole (git show 202396e9f:) at > 2 2 failed, 40 passed — 2 failed, 40 passed, both states no register extent at all P1, P2 fix's claimed effect is real

6 mutants bite. 3 are inert (M11, M12, and N0-on-control which refused). 1 is a blind spot
(M8).
Per pin: P1 bites (M2, M3, M9, M10), P2 bites and holds more alone than any
other test in the tree
(M1, M2, M4, M5, M13), P3 bites on the mutation the body measured
(M7) and is blind to M8.

Is a property held by exactly one test in the whole tree? Yes — all three, and in one
file.
grep -rn over tests/, atom/, scripts/ at the head:
tests/compass/test_open_items_register.py is the only file in the repository that reads
12_open_items.md at all; named <= present occurs once; "only one id at a time"
occurs once as an assertion. There is no sibling anywhere for P2 or P3. The one place a
sibling exists is M3 — named below.


The seven questions, answered

Q1 — does an empty result satisfy a derivation? No, on both tested surfaces, and the
guard that stops it is the base's, not this PR's. LANDED = [] (M9) reddens both P1
and P2. Emptying the fixture's extent so spans filters to [] (M10) reddens P1 on
assert spans, f"{path.name} states no register extent at all" — the assert spans
precondition #119 found missing from a sibling derivation is present here, in the function
both pins call. One asymmetry worth recording: M10 leaves P2 green. P2 does not need the
fixture's stated extent at all — its pair T99 and T100 survives the filter on its own and
refuses correctly. So the T1–T3 extent is load-bearing for P1 and decorative for P2. The
body claims it is load-bearing for both ("both cases would then die on the precondition") —
true at > 2, not true when the extent is simply removed. Measured, not inferred.

Q2 — does a tripwire in the branch the pin names fire that pin? For P2 yes, for P1 no
— and that is the correct outcome, not a defect.
raise SystemExit("TRIPWIRE-ELSE-BRANCH")
as the first statement of the escape-hatch else (M6) fires exactly three cases:
test_every_stated_extent_names_exactly_the_rows[12_open_items.md], [README.md], and P2.
P1 is not among them. P1's green is produced before the loop, by the
len(ids(span)) > 1 filter dropping T99 — the same > 1 filter shape that made #119's pin
pass for an unrelated reason. Here it does not: P1 is named for the filter, the filter is
where the hatch lives, and P1 reddens when the filter moves (M2, M3). But a reader who
credits P1 with exercising the else branch is wrong
, and the else-branch comment sits
inside the branch P1 never reaches. Recorded, not raised.

Q3 — can the document say the opposite coherently while the pinned literal survives?
Yes. This is the one blind spot. P3 is a substring pin, assert "only one id at a time" in flattened(REGISTER). One line of 12_open_items.md, line count preserved, substituted once:

-   unlanded allocation may be named here only one id at a time
+   unlanded allocation may be named here in pairs, and not only one id at a time

The document now sanctions exactly what the rule refuses, reads perfectly, and contains
the pinned literal verbatim. Result: 42 passed at the module, and at the full CPU tier
4441 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0, 0 FAILED lines
— byte
for byte the branch baseline. I ran the full gate precisely because a module-level green is
not enough to call a mutation invisible. It is invisible.

The PR body says of P3: "so the rule and the document that sanctions it cannot drift apart
unnoticed."
Measured, they can: one line, one clause, no test notices. Principle 8 — the
measurement offered for that sentence is the one id at a time only reword (M7), which is a
meaning-preserving mutation; it supports "a copy-editor cannot silently reword this", which
is true and useful, and does not support "cannot drift apart unnoticed", which is what is
written. The body's What this does not do item 3 already discloses that #106's criterion is
half met because the code comment is unpinned; what M8 adds is that the pinned half is
itself satisfied by a document stating the opposite rule.
That is a strictly stronger
statement of a limitation the PR already owns, not a new defect class.

Q4 — is any claim un-numbered? P3 is the un-numbered assertion in this diff, and Q3 is
exactly the #206 shape: the guard has no number to check, so a false statement of the rule
passes. P1 and P2 are numeric and are not blind in that way — M4 shows the match= regex
cannot lose [99, 100], the span or the filename without P2 going red, and the mutation
reddens on the Regex pattern did not match assertion, which is the one a reader would
credit.

Q5 — is each reported fix actually held? No — the round-2 commit is inert under revert,
and that is structural rather than a defect.
#109's round 2 reports one behaviour change:
widening the fixture from T1–T2 / 2 rows to T1–T3 / 3 rows so that at > 2 the failure
names the pair was let through instead of the fixture has no extent. Reverting exactly
those two lines (M11, line-count preserved) gives 42 passed at the module and
4441 passed, GATE_CPU_RC=0 at the full gate. The fix's claimed effect is nonetheless
real, verified against the recovered pre-fix file rather than a reconstruction —
git show 202396e9f:tests/compass/test_open_items_register.py, md5
3d2347e2543f62bab9f5fe0af0fabbf7, staged whole:

tree at > 1 (shipped) at > 2
recovered round-1 head 202396e9f 42 passed 2 failed — P1 and P2, both states no register extent at all, assert []
round-2 head ba3941084 42 passed 1 failed — P2, Failed: DID NOT RAISE <class 'AssertionError'>

So the commit converts two mis-named failures into one correctly-named one, exactly as both
round-2 records state. It is unheld because the quality it improves is the name of a failure
under a mutation, and no test in the shipped configuration observes it.
A test could not
hold it without being a meta-test. Recorded so nobody later reads M11's green as evidence the
commit did nothing.

Q6 — is an empty parametrize a skip rather than a failure here? Yes, and #109 is the
only thing that survives it.
Emptying the extent test's own parametrize (M12,
STATED_IN → (), the other two parametrize sites untouched) is green on both trees —
control 37 passed 1 skipped, branch 40 passed 1 skipped. So the #206 property is
present in this file. The discriminating measurement is the combined mutant M13 (empty
parametrize + the hatch assertion neutered), disclosed as two substitutions in one tree:
control 37 passed, 1 skipped — green, the rule entirely uncovered; branch 1 failed —
P2 fires.
#109's pins call test_every_stated_extent_names_exactly_the_rows directly with
a literal path
, so they are the only coverage of that rule that an emptied parameter set
cannot skip away. That is a property of the design choice the round-1 review accepted
("running the real test over written documents"), it is stronger than the body claims for
it, and it is not stated anywhere in #109.

Q7 — does a mutation redden on a different assertion than a reader would credit? Once, and
it is benign. M5 neuters the hatch assertion itself (assert named <= present →
assert True) and P2 reddens on Failed: DID NOT RAISE — the pytest.raises wrapper, not
the assertion removed. That is the only failure available once the assertion cannot raise, and
it names the right test; no reader is misled. Everywhere else the failing assertion is the one
the mutation touched, quoted in the table above.


Where a sibling catches it — named

M3 (> 1 → > 0) is the one direction #109 is not the first to catch. At the control
tree, before any of these three tests exist, > 0 is already red on
tests/compass/test_open_items_register.py::test_every_stated_extent_names_exactly_the_rows[12_open_items.md]
— 12_open_items.md states the register as 'T1': [] stated with no row, [2, 3, …, 87] in rows with no figure covering them, assert {1} == {1, 2, 3, 4, 5, 6, ...}. 1 failed, 38 passed.
P1 adds a second, correctly-named failure on a synthetic document beside it. Both round-2
records say this in those words ("the pre-existing incidental case is the only overlap",
"#106's one red cell was a coincidence"); I reproduce it and name the sibling so the row is
readable. In the tightening direction — > 2, > 3, the refusal text, the named <= present
assertion, the doc sentence — the control is green and there is no sibling anywhere in the
tree.


Three conditions, and what the controls did

  • Line-count-preserving, with a null-comment control. Every one of the 15 substitutions
    printed its newline count before and after: delta 0 in all 15. The null control N0
    (a comment reworded, no semantics) gives 42 passed, unchanged — the harness itself moves
    nothing. One disclosed exception: the recovered round-1 file (R1) is 409 lines against
    the head's 411, because the round-2 commit is +2; it is a file recovered with git show,
    never a reconstruction, and it is labelled as such above.
  • Uniqueness asserted before every substitution. The driver counts occurrences and
    refuses at anything but exactly 1. It refused once: N0 against the control tree, because
    that comment line is one of the 38 this PR adds and does not exist at ae2935c2b. Disclosed
    rather than dropped. No other mutant was discarded, and none was silently skipped.
  • Never concurrent. One staged tree per branch/control, one mutant at a time, strictly
    sequential. After every mutant the target file was rewritten from the bytes read before it
    and the md5 re-checked: 15 of 15 restored clean, both trees byte-identical to their
    archives at the end.
  • Full gate before calling anything invisible. Both mutants that were green at the module
    (M8, M11) were re-run through the whole CPU tier; both stayed green there. M12's
    green is reported as a skip, which it is.

Gate — four trees, nothing piped

Node 18, xiaobizh_n18_cpu. Staged with git archive + docker exec -i into my own
/work/pr109pin/; the shared mount /tmp/xiaobizh-compass/ATOM was never touched and nothing
was written to node 18's host filesystem. Nobody else's staging was removed; mine was. Tarball
md5s verified on both ends (6d0b6536… branch, 7883aa0b… control). .compass-commit and
.compass-changed written from the same rev-parse that produced each archive
(changed: 2 file(s) each, COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new).
Each tree ran its own scripts/compass/, and the two are the same object —
git rev-parse <sha>:scripts/compass is ddb69e7aa380a1b8cb1ff3783330c4e6d369f9f5 at
ae2935c2b, 202396e9f and ba3941084, and hashed inside the container from inside
that directory both trees give 528e7739320514159633db0f572d0774 — the digest both
round-2 records report. No comparison crosses a gate-script change. import atom was
confirmed resolving under each staged root and printed before any count was read; every
run bounded with timeout -k 10; each gate's stdout to its own file with the shell rc on the
next line, never piped; __pycache__ cleared before every run.

Tree CPU tier
control — ae2935c2b 4438 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0, 0 FAILED
branch — ba3941084 4441 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, shell rc 0, 0 FAILED — +3
branch + M8 (doc says the opposite) 4441 passed, 149 skipped, 3 xfailed, rc 0 — identical
branch + M11 (round-2 fix reverted) 4441 passed, 149 skipped, 3 xfailed, rc 0 — identical

Runs 2026-09-22 21:59:03 – 22:06:59 UTC, 33–35 s each, sequential; no other gate was live
(checked before starting, and another agent's gate that was running on arrival had cleared).
Skips 149 and xfails 3 in all four, so no ±1 and nothing from the flaky class;
TestTheRegionIsNotCopiedPerChunk produced no FAILED line in any run. +3 is exactly the
three new tests
: the file reads 42 passed on the branch tree and 39 passed on the
control.

Size — production and test counted separately, instrument calibrated

Instrument calibrated on landed atom/compass/spec/ taken from fork/feature/atomcompass_new
(5 files): 275 statements / 495 code / 763 physical — all three reproduce exactly, and are
unchanged after black (26.5.1, the container's; psf/black@stable line), which makes no
edit to that package.

Production AST statements: +0. The diff is one file, tests/compass/, +38 / −0; no
production file is touched, so there is no production row to report.

Test instrument base ae2935c2b 202396e9f ba3941084 delta
AST statements 125 139 139 +14
AST minus docstring-only 111 124 124 +13
Code (physical non-blank less comment-only less non-blank docstring lines) 192 211 211 +19
Physical non-blank 315 343 345 +30
Physical, all lines 373 409 411 +38

Identical before and after black; black changes no line this PR adds. black --check exits
1 at both base and head, one hunk each, at line 352 on the base and line 390 at the
head — the same pre-existing hunk displaced by exactly the 38 inserted lines. Every figure in
the body's effort table reproduces at all three heads on an independently written instrument.
I say nothing about #89.

The identifiers in the added lines — checked, and reported as nothing

The 38 added lines carry 7 design-doc identifiers: T1, T3, T99 ×3, T100 ×2. Every
one is inside FREE_MENTION's fixture sentence or the pytest.raises match regex — parser
input and fixture text, not citations.
There are no D-numbers, no document numbers and no
issue numbers in the added code. This is the class a board-wide census classified as an
exception and reported as a violation nowhere, and I am doing the same. Checked; nothing to
report.


What this does and does not change

It does not change the APPROVE. Every published figure I could re-derive reproduces —
the four thresholds, the refusal-message mutation, the doc reword, the control and branch gate
counts, the gate-script digest, all four effort instruments. The pins are real pins: six
mutations that leave the control green turn the branch red, and five of them have no sibling
anywhere in the repository.

Two things are now measured that were not, and neither blocks anything:

  1. P3's substring pin admits a document that states the opposite rule (M8), invisible at
    module and at the full gate. The body's sentence "cannot drift apart unnoticed" claims
    more than the measurement behind it supports — principle 8, and the same class as
    round 1's R1, on one clause rather than a paragraph. The honest repair is one sentence of
    body text, or a second pin on the clause that does the work (may be named here only one).
    compass(design): pin the register guard's one-id escape hatch #109 cannot land for days behind compass(design): a third document may not state the register's extent #95 and compass(design): guard the open-items register's own counts #85; this belongs in whatever touch restacks it,
    not in a new review cycle.
  2. P1 and P2 are the only coverage of the extent rule that an emptied parameter set cannot
    skip away
    (M13). That is a real strength of the design choice round 1 accepted and it is
    nowhere claimed. Worth one line in the body if it is edited for (1) anyway.

Nothing was merged, drafted, undrafted, amended, pushed or labelled; no other PR was touched;
the staging under /work/pr109pin/ is removed and the ControlMaster closed.

@jgong5

jgong5 commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Closed under the owner's ruling in #489: prose (design documents, READMEs, docstrings) is not a test subject, and a wrong doc is fixed in the doc rather than guarded. This item exists only to guard, pin or count prose, so there is nothing left to deliver. If it holds a real behaviour defect that was missed, reopen it with that defect as its title.

@jgong5 jgong5 closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant