compass(docs): attribute the overlapping-gate reproduction to #93 and say no record names the skip-variants method - #376
Conversation
… qualify the one-method reading The flake-rate paragraph said "Per #93, all 21 runs were of one method"; #93 names one test id and never says all 21 ran it, and comment 5765186351 names that method only for the hard failure, so the skip-variants' method has no record. Say so, and fold the attribution of the conditions, the table and the test id into one sentence. The next paragraph said the failure was reproduced with two gate loops overlapping, with no source, right after the sentence saying (per #93) that the counted hard failure ran under direct pytest. Attribute it to #93 and say that #93 names neither that run's harness nor whether it is the counted one, so the two no longer read against each other. Drop "the only record that says," which the following clause already carries. Closes #368 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| PR #79's review (issue comment 5765186351) ran `b58a48cc2` 19 times, for 17 nominal, | ||
| 1 skip-variant and the one hard failure. The other 2, 1 nominal and 1 skip-variant | ||
| by subtraction from the table, name no commit, so half the skip-variant rate rests | ||
| on #93's count alone. No record names the skip-variants' method: #93 never says |
There was a problem hiding this comment.
B1 (blocking): this sentence says no record names the skip-variants' method, but two records do. Principle 8: "Every claim carries its measurement." Here the claim is about what the records say, and it inverts them. The records attribute the skip-variants to the test by pronoun. None of them observed it.
- The CPU tier has one three-way flaky test and it is recorded nowhere #93 body. Under "The test:" it gives
…::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow, then "A wall-clock timing test with apytest.skipnoise guard", and then "It is three-way, not two-way. Measured over 21 runs on node 18 inxiaobizh_n18_cpu, on one branch:", followed by the 18 / 2 / 1 table. "It" is the named test, and its table includes the two skip-variants. The title is "The CPU tier has one three-way flaky test". The CPU tier has one three-way flaky test and it is recorded nowhere #93 never uses the words "all 21", but it does attribute all 21 to that test. - 5765186351, finding 10. It opens "The CPU tier's flake is three-way, not two-way, and I can name it", then says "Measured over 19 branch runs, it does that once (4476 passed / 150 skipped …) and fails outright once:". The node id comes after the failure clause, but both clauses have the same subject. It goes on: "The skip is the guard firing; the failure is the machine being noisy … So the two outcomes are one mechanism." So it does not name the method "only for the failure".
- The record this sentence needs is PR docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99's body, under "Not checked": "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons and refuses a caller's
-r… The signature is recorded; the attribution is not claimed." docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99's round-1 review (issue comment 5766595486) repeats it: "The ±1 skip-variant has never been attributed by name".
The brief asked for less than this sentence says. #368 N2: "which method they were rests on #93 alone. Qualify it." #366's review said the claim "Holds as a reading".
As written, this sentence fails the PR's own named result ("every claim … names its record, or says it has none"): it says "none" where two records exist. It also sits oddly against L214's "the one method #93 names".
Suggested replacement, which is one line shorter. It is also the ponytail shrink: in the verdict.
on #93's count alone. #93 attributes the skip-variants to that test, but no run
named it (PR #99). Not-nominal combined is ~1 in 7.
There was a problem hiding this comment.
Fixed in 1c5d383d2 (delta 8f9398030..1c5d383d2). L197-198 now read, in your wording:
on #93's count alone. #93 attributes the skip-variants to that test, but no run
named it (PR #99). Not-nominal combined is ~1 in 7.
That is one line shorter (−3 / +2). I re-read each record via REST before writing it:
- The CPU tier has one three-way flaky test and it is recorded nowhere #93 body. Under "The test:" it gives
…::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow, and then "It is three-way, not two-way. Measured over 21 runs on node 18 inxiaobizh_n18_cpu, on one branch:", followed by the 18 / 2 / 1 table. That is the attribution. - 5765186351, finding 10. "The CPU tier's flake is three-way, not two-way, and I can name it. … Measured over 19 branch runs, it does that once (4476 passed / 150 skipped …) and fails outright once". It is the same attribution, and the new sentence doesn't deny it.
- PR docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99 body, "Not checked". "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons and refuses a caller's
-r, which is how you would ask for them. The signature is recorded; the attribution is not claimed." This is the record behind "no run named it". docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99's round-1 review, 5766595486, repeats it: "The ±1 skip-variant has never been attributed by name, as the PR's "Not checked" says."
L213 now agrees. It reads "…not to the one method #93 names", and #93 names that one method and attributes the skip-variants to it.
| **Run gates sequentially.** The failure was reproduced with two gate loops | ||
| overlapping on node 18. Two gates on one box compete for the CPU the control arm is | ||
| **Run gates sequentially.** Per #93, the hard failure was reproduced when two gate | ||
| loops on node 18 overlapped; it names neither that run's harness nor whether it is |
There was a problem hiding this comment.
N1 (non-blocking, nit): "it names neither" reads as if "it" were the hard failure. The nearest singular noun is "the hard failure", or "node 18". "#93 names neither that run's harness nor whether it is the one in the counts" costs nothing and leaves no doubt about the referent.
The content holds, and so does the reconciliation.
- The CPU tier has one three-way flaky test and it is recorded nowhere #93 says, under "One further observation to record": "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially." That is all it says. It gives the run no harness, no commit and no count.
- I searched compass(spec): the machine-spec schema, closed against the deployment (SPEC-1) #79 (the
issuesandpullsendpoints) and docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99 (body,issuesandpullscomments) foroverlap|two gate|gate loops|concurren|parallel. No other record of that reproduction exists; docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99's hits are only reviewers saying they waited rather than overlap. - "When … overlapped" makes the overlap a condition on the box, which fits L214-215's direct-pytest attribution. Principle 8 is met.
There was a problem hiding this comment.
Fixed in 1c5d383d2. L218-219 now read "…loops on node 18 overlapped; #93 names neither that run's harness nor whether it is the one in the counts." I reflowed it over two lines and the content is unchanged.
The record behind it is still #93's "One further observation to record": "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially."
Review, cycle 1: #376 (issue #368), head
|
| PR claim | Record | Result |
|---|---|---|
L192-193, fold: "Those conditions, the table and its one test id, test_the_cost_per_byte_does_not_grow, are issue #93's." |
#93 body | Holds. "Measured over 21 runs on node 18 in xiaobizh_n18_cpu, on one branch"; "n=21, one mechanism, on a box whose load was not controlled"; the 18 / 2 / 1 table; one node id under "The test:". |
| #93 says "one mechanism", not "one method" or "all 21" | #93 body | Holds, literally. The words are "n=21, one mechanism", and "all 21" does not occur. But #93 attributes the 21 to the test by pronoun: "The test: …test_the_cost_per_byte_does_not_grow … It is three-way, not two-way. Measured over 21 runs …" (see B1). |
| L197-198, N2: "No record names the skip-variants' method: #93 never says all 21 ran that test, and 5765186351 names it only for the failure." | #93, 5765186351, #99 | Does not hold (B1). #93: "It" (the named test) "is three-way … Measured over 21 runs", with its table including the two skip-variants. 5765186351, finding 10: "… and I can name it. … it does that once (4476 passed / 150 skipped …) and fails outright once", and "The skip is the guard firing". What is absent is an observation. PR #99's body, "Not checked", says: "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons … the attribution is not claimed." 5766595486 on #99 says the same. |
| 5765186351 names the method only for the 1.89x failure | 5765186351 | Partly. The node-id block does follow the failure clause only. But the paragraph's subject is the named test for both the skip and the failure, as quoted above. |
| L213-214: "…not to the one method #93 names." | #93 body | Holds. #93 names one method. |
L214-216, drop: "Per #93, the two skip-variants both occurred under gate_cpu.sh, and the one hard failure in the counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct." |
#93; 5765186351 | Holds. #93: "the two skip-variants both occurred under gate_cpu.sh while the failure occurred under direct pytest". 5765186351: "7 gate + 12 direct". Dropping "the only record that says," loses nothing. |
| L218-220, N1: "Per #93, the hard failure was reproduced when two gate loops on node 18 overlapped; it names neither that run's harness nor whether it is the one in the counts." | #93 body | Holds. #93: "One further observation to record: the failure was reproduced when two gate loops on node 18 overlapped." It gives the run no harness, no commit and no count. #79 and #99 hold no other record of it. Searching for overlap|two gate|gate loops|concurren|parallel, #99's only hits are reviewers who waited rather than overlap. |
| "#93's body is the only source; its one comment says only that it landed in #99" | issues/93/comments |
Holds. 5768642904 reads "Landed in #99 (the flaky-test record). Closing." |
Consistency (principle 8, and the brief's named result)
- The two harness statements now agree.
- L214-215 puts the counted hard failure under direct pytest.
- L218-219 makes the overlap a condition on the box ("when … overlapped") and says The CPU tier has one three-way flaky test and it is recorded nowhere #93 names no harness for that run.
- A direct-pytest run on a box with two gate loops overlapping is consistent. N1 from compass(docs): name the record behind each flake-rate count and harness claim (#358) #366 is closed.
- Nothing else in the README contradicts them. I checked:
:44: "run sequentially and unpiped";:103-108: "outcomes of one flaky test", and the rc of each;- the table at
:165-171; - "Seen again since" at
:224-231: "The gate prints no skip reasons … so that run did not name the test". This agrees with docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99, and it is the fact B1's replacement cites.
- The one tension is inside the changed text. L197 says no record names the skip-variants' method, while L214 says The CPU tier has one three-way flaky test and it is recorded nowhere #93 names "the one method". B1's fix removes it.
Not over-hedged?
- N1's second clause is a necessary hedge, not an over-hedge. "Nor whether it is the one in the counts" is what reconciles it with L215, and The CPU tier has one three-way flaky test and it is recorded nowhere #93 does leave it open.
- N2 is the reverse. It is an over-claim of absence, not a hedge. That is B1.
ponytail-review (principle 3: "Prioritise simplicity. Add only what is necessary, and nothing more.")
scripts/compass/README.md:L197-198: shrink: two-line absence claim over #93 and 5765186351. "#93 attributes the skip-variants to that test, but no run named it (PR #99).", one line with the reflow.
net: -1 lines possible.
N1's clause at L219-220 is load-bearing (see above); the fold at L192-193 and the drop at L213 are already done as the brief asked. The only cut is the same edit that fixes B1.
Gate 1: ATOM's suite, unmodified, on node 18 (xiaobizh_n18_cpu)
Setup
- Tip. I re-read it at 2026-09-24T00:31:40Z by
git fetch fork feature/atomcompass_new:373f1fc5c0d7a7f13780cafb3f8c8b271a6398db. It has moved past the dev record's8c0a2e6ee, because compass(rules): the reviewer checks gates 1-3 as they apply per task, so approval never waits for the per-wave GPU tier (#369) #371 landed. - Merged tree.
git merge-tree --write-tree 373f1fc5c 8f9398030gives477b04347aa226edfb505d38d99f96755130fc92, with rc=0.- This is not the head's tree
68d75c1a7.git diff 8f9398030 477b04347isatom/compass/AI_DEV_RULES.mdonly (+2 / −2, compass(rules): the reviewer checks gates 1-3 as they apply per task, so approval never waits for the per-wave GPU tier (#369) #371). git diff --name-only 373f1fc5c 477b04347isscripts/compass/README.mdonly.- So the dev record's gate doesn't stand for this tree, and I gated it once.
- This is not the head's tree
- Stamp.
git commit-tree 477b04347 -p 373f1fc5c -p 8f9398030gives8b488b6531b9b2fc3be2acf136d146b252261bcd. No ref was created. - Staging.
git archiveof the stamp..compass-commitis the stamp, and.compass-changedisscripts/compass/README.md.- The tar went through
docker exec -iinto/tmp/pr376r1/stage/merged/ATOM, with md5c0b2caba…on both ends. - The tree carries its own
scripts/compass, treee21ade45e.
- Run.
- The tree's own
scripts/compass/gate_cpu.sh --junitxml, undertimeout -k 10 3000, unpiped. - It ran alone: I waited for another tenant's
gate_cpu.sh(/tmp/i353gates) to finish first. - Load average was 15.5 → 16.8.
- The tree's own
Result
| tree | printed commit: |
atom.__file__ |
passed | skipped | xfailed | junit cases | GATE_CPU_RC |
|---|---|---|---|---|---|---|---|
merged 477b04347 (stamp 8b488b653) |
8b488b653 (stamp) |
/tmp/pr376r1/stage/merged/ATOM/atom/__init__.py |
5268 | 155 | 3 | 5426 | 0 |
- Against the dev record's 5268 / 155 / 3 (5426 cases) at
8c0a2e6ee, the delta is 0. compass(rules): the reviewer checks gates 1-3 as they apply per task, so approval never waits for the per-wave GPU tier (#369) #371 is docs only, so a docs-only change moves nothing. - Timing classes. All passed, so I did not need a re-run:
TestTheRegionIsNotCopiedPerChunk: 4 of 4;TestNoSizeAtWhichACallStopsBeingOne: 12 of 12;test_freezing_twice_is_additive_and_harmless: 1 of 1.
- GPU tier:
not required. - Staging removed from node 18 and locally.
Gates 2 and 3. This is a docs-only change, so gate 2 does not apply. Gate 3's named result is "every claim in both paragraphs names its record, or says it has none". It holds for every changed sentence except L197-198, and that one is B1.
I did not undraft, merge, label or push anything.
…amed it Review cycle 1 on #376 (B1): "No record names the skip-variants' method" was false. #93 credits all 21 runs to test_the_cost_per_byte_does_not_grow through "It" after "The test:", and comment 5765186351 does the same. What is missing is an observation by name, which PR #99's "Not checked" states. Replace the absence claim with that, one line shorter. Nit: "it names neither" read as the hard failure; say "#93 names neither". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Developer round 2: #376 (issue #368), new head
|
README sentence (at 1c5d383d2) |
Record, quoted |
|---|---|
L191-193: "n=21, node 18, container xiaobizh_n18_cpu, on a box whose load was not controlled. Those conditions, the table and its one test id … are issue #93's." |
#93: "Measured over 21 runs on node 18 in xiaobizh_n18_cpu"; "n=21, one mechanism, on a box whose load was not controlled"; one node id under "The test:". |
L193-195: "19 of the 21 are pinned: PR #79's review (issue comment 5765186351) ran b58a48cc2 19 times, for 17 nominal, 1 skip-variant and the one hard failure." |
5765186351: "| Branch | b58a48cc2 | 7 gate + 12 direct | 4477 passed, 149 skipped, 3 xfailed, rc=0 in 17 of 19 |"; "Measured over 19 branch runs, it does that once (4476 passed / 150 skipped …) and fails outright once". |
| L195-197: "The other 2 … by subtraction from the table, name no commit" | #93's table has 18 / 2 / 1. Subtracting 17 / 1 / 1 leaves 1 nominal and 1 skip-variant. #93 names no commit, only "on one branch". |
| L197-198: "#93 attributes the skip-variants to that test, but no run named it (PR #99)." | See B1 above. |
| L198: "Not-nominal combined is ~1 in 7." | #93: "Not-nominal combined is ~1 in 7." |
| L199: "The finding is the third outcome, not the rate." | #93: "The finding is the third outcome, not the rate." |
L199-205: "the class holds 3 methods / 4 collected cases … of which 2 methods / 3 cases share the same timing helpers, the same noise guard and the same < 1.5 ratio assertion"; the open-region method "is deterministic" |
The test file at the tip, tests/entrypoints/test_stream_marker_properties.py, TestTheRegionIsNotCopiedPerChunk: test_the_open_region_is_never_scanned_beyond_the_window, then test_no_format_pays_more_per_byte_as_the_payload_grows with ids=["buffered-region", "kimi-incremental"], then test_the_cost_per_byte_does_not_grow. The last two each carry pytest.skip(f"machine too noisy to measure: …") and assert large / small < 1.5. The junit for this gate has 4 cases in the class. |
L205-211: the sibling's "qwen: cost per KB grew 1.73x from 32 to 128 KB, 1 failed, 4495 passed, rc=1, on the integration head 669dc3f9d", and the re-run at 4496 |
5765660762: "— "qwen: cost per KB grew 1.73x from 32 to 128 KB" — 1 failed / 4495 passed, rc=1 … Re-run of that one tree, sequentially: 4496 passed, rc=0, and 1 + 4495 = 4496 accounts for it exactly."; the head is 669dc3f9d. |
| L212-213: "…not to the one method #93 names." | #93 names one node id. |
L213-215: "Per #93, the two skip-variants both occurred under gate_cpu.sh, and the one hard failure in the counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct." |
#93: "the two skip-variants both occurred under gate_cpu.sh while the failure occurred under direct pytest". 5765186351: "7 gate + 12 direct". |
| L217-219: "Per #93, the hard failure was reproduced when two gate loops on node 18 overlapped; #93 names neither that run's harness nor whether it is the one in the counts." | #93: "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially." It says nothing more about that run. |
L219-222: "Two gates on one box compete for the CPU the control arm is measuring … check for a running gate_cpu.sh before starting one." |
This is advice, and its record is #93: "Say that gates should be run sequentially." The mechanism is 5765186351: "the failure is the machine being noisy in a way the guard does not catch". |
Gate 1: ATOM's suite, unmodified, on node 18 (xiaobizh_n18_cpu)
Setup
- Tip. It moved while I waited for the box:
13ab24dc7became8fe47dbb5b081935917950f9e8c55ee2d79145d6(compass(spec): hold the whole transfer refusals, and every sys.modules key #374,tests/compass/test_spec_verbs.pyonly). I re-read it at 2026-09-24T00:58:05Z bygit fetch forkand gated against it. - Merged tree.
git merge-tree --write-tree 8fe47dbb5 1c5d383d2givesd14621b54bd0f99b4238a2688f90ed1f66b102cb, with rc=0.- This is not the head's tree
dc7acfb8f, because the tip carries commits the branch does not. git diff --name-only 8fe47dbb5 <stamp>isscripts/compass/README.mdonly.
- This is not the head's tree
- Stamp.
git commit-tree d14621b54 -p 8fe47dbb5 -p 1c5d383d2gives8d91a93a8b2c9e36f0813606b81eb7baf15867ab. No ref was created. - Staging.
git archiveof the stamp..compass-commitis the stamp, and.compass-changedisscripts/compass/README.md.- The tar went through
docker exec -i … tar -xinto/tmp/i368r2/m2/ATOM, with md55c531f4c…on both ends.
- Run.
- The tree's own
scripts/compass/gate_cpu.sh --junitxml, withPYTHONPATHset to the staged root, undertimeout -k 10 3000, unpiped. - It ran alone: I waited for another tenant's gate (
/tmp/pr374r1) to finish first. - Load average was 11.15 at the start.
- The tree's own
Result
| tree | printed commit: |
atom.__file__ |
passed | skipped | xfailed | junit cases | GATE_CPU_RC |
|---|---|---|---|---|---|---|---|
merged d14621b54 (stamp 8d91a93a8) |
8d91a93a8 (stamp) |
/tmp/i368r2/m2/ATOM/atom/__init__.py |
5268 | 155 | 3 | 5426 | 0 |
- Against cycle 1's 5268 / 155 / 3 (5426 cases), the delta is 0. The README change is docs only, and compass(spec): hold the whole transfer refusals, and every sys.modules key #374 changes assertions without changing the case count.
- Timing classes: 17 of 17 passed, so I did not need a re-run. They are
TestTheRegionIsNotCopiedPerChunk(4),TestNoSizeAtWhichACallStopsBeingOne(12) andtest_freezing_twice_is_additive_and_harmless(1). - GPU tier:
not required. - Staging removed from node 18 and locally, via a script.
I did not undraft, merge or label anything.
| 1 skip-variant and the one hard failure. The other 2, 1 nominal and 1 skip-variant | ||
| by subtraction from the table, name no commit, so half the skip-variant rate rests | ||
| on #93's count alone. #93 attributes the skip-variants to that test, but no run | ||
| named it (PR #99). Not-nominal combined is ~1 in 7. |
There was a problem hiding this comment.
N3 (non-blocking): "no run named it" says more than PR #99 does, and read literally it is false. Principle 8: "Every claim carries its measurement. A number without a source is a defect." Here the measurement is what the cited record says, and this sentence widens it.
- "it" is "that test",
test_the_cost_per_byte_does_not_grow, and a run did name that test. 5765186351's hard-failure run printedtests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_growwithE AssertionError: cost per KB grew 1.89x from 32 KB to 128 KB. L176-178 of this README tell the reader to rely on exactly that line: "Read theFAILEDline the gate prints first … if it names this class, re-run". - PR docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99 is narrower. Its "Not checked" says "The ±1 skip-variant has not been attributed by name on any run", and 5766595486 says "The ±1 skip-variant has never been attributed by name". Both are about the skip-variant runs, not about the test.
- Why it is not blocking. The contrast with "attributes the skip-variants" steers most readers to the intended scope. No reader acts wrongly on it, and it does not contradict L171 or L212-213 in practice.
- Why it is still worth one word. This is the sentence this task exists to get exact. The wording came from cycle 1's suggested replacement, so this corrects that suggestion, not your transcription of it.
It costs one word, and the line count does not change:
on #93's count alone. #93 attributes the skip-variants to that test, but no
skip-variant run named it (PR #99). Not-nominal combined is ~1 in 7.
Review, cycle 2 (delta): #376 (issue #368), head
|
| Finding | At head | Status |
|---|---|---|
| B1, blocking (4088688626): "No record names the skip-variants' method" inverted #93 and 5765186351 | "#93 attributes the skip-variants to that test, but no run named it (PR #99)." | Resolved. The absence claim is gone. The attribution is credited to #93, and the qualification cites PR #99. Checked below. N3 is a new, narrower point about the second clause. |
| N1, non-blocking (4088688768): "it names neither" had an ambiguous referent | "…overlapped; #93 names neither that run's harness nor whether it is the one in the counts." | Resolved. The subject is now explicit, and the content is unchanged. |
Every sentence at L191-222, against its record (principle 8: "Every claim carries its measurement. A number without a source is a defect.")
I read each record myself via REST. I did not take the developer's table.
| README (head) | Record, quoted | Result |
|---|---|---|
L191-194: "n=21, node 18, container xiaobizh_n18_cpu, on a box whose load was not controlled. Those conditions, the table and its one test id, test_the_cost_per_byte_does_not_grow, are issue #93's." |
#93: "Measured over 21 runs on node 18 in xiaobizh_n18_cpu, on one branch"; "n=21, one mechanism, on a box whose load was not controlled"; the 18 / 2 / 1 table; one node id under "The test:". |
Holds. |
L194-196: "19 of the 21 are pinned: PR #79's review (issue comment 5765186351) ran b58a48cc2 19 times, for 17 nominal, 1 skip-variant and the one hard failure." |
5765186351: "| Branch | b58a48cc2 | 7 gate + 12 direct | 4477 passed … rc=0 in 17 of 19 |", and finding 10: "Measured over 19 branch runs, it does that once (4476 passed / 150 skipped …) and fails outright once", with cost per KB grew 1.89x. |
Holds for the figures. That these 19 are 19 of #93's 21 is a reading: the 1.89x failure, the container and 17/1/1 ⊂ 18/2/1 all agree, and neither record states it. The wording predates this PR, which only reflowed it. Not raised. |
| L196-197: "The other 2, 1 nominal and 1 skip-variant by subtraction from the table, name no commit, so half the skip-variant rate rests on #93's count alone." | 18 − 17 = 1, 2 − 1 = 1, 1 − 1 = 0. #93 names no sha, only "on one branch". | Holds. |
| L197-198: "#93 attributes the skip-variants to that test, but no run named it (PR #99)." | #93: "The test:" …::test_the_cost_per_byte_does_not_grow, then "It is three-way, not two-way. Measured over 21 runs …", with the skip-variant row in the table. #99 body, "Not checked": "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons and refuses a caller's -r … The signature is recorded; the attribution is not claimed." 5766595486: "The ±1 skip-variant has never been attributed by name". |
First clause holds. Second clause holds in scope but is wider than its record (N3). #99 says no skip-variant run named it. 5765186351's failing run did name the test. |
| L198: "Not-nominal combined is ~1 in 7." | #93, verbatim. 3 / 21 = 1 / 7. | Holds. |
| L199: "The finding is the third outcome, not the rate." | #93, verbatim. | Holds. |
L199-205: "3 methods / 4 collected cases", "2 methods / 3 cases share the same timing helpers, the same noise guard and the same < 1.5 ratio assertion"; the open-region method "counts scans of the open region and is deterministic" |
The test file tests/entrypoints/test_stream_marker_properties.py, byte-identical at head and at cae9607f0, TestTheRegionIsNotCopiedPerChunk. test_the_open_region_is_never_scanned_beyond_the_window calls _scans_of_an_open_region and asserts max(seen) <= _PEEK_WINDOW. test_no_format_pays_more_per_byte_as_the_payload_grows has ids=["buffered-region", "kimi-incremental"], and it and test_the_cost_per_byte_does_not_grow each call _control_ms/_stream_ms, carry if not 0.6 < control < 1.6: pytest.skip(…), and assert large / small < 1.5. |
Holds. |
L205-212: the sibling, "qwen: cost per KB grew 1.73x from 32 to 128 KB, 1 failed, 4495 passed, rc=1, on the integration head 669dc3f9d; node 18, xiaobizh_n18_cpu, 2026-09-21, recorded in the round-2 review of PR #67 (issue comment 5765660762), which hit it on its own gate run. A sequential re-run of that same tree passed at 4496" |
5765660762 (on #67, 2026-09-21T18:42:14Z, "Round 2, second reviewer"): "The integration head's first run failed with …::test_no_format_pays_more_per_byte_as_the_payload_grows[buffered-region] — "qwen: cost per KB grew 1.73x from 32 to 128 KB" — 1 failed / 4495 passed, rc=1. … Re-run of that one tree, sequentially: 4496 passed, rc=0". The head is "feature/atomcompass_new is 669dc3f9d", and the conditions are "Node 18, xiaobizh_n18_cpu". |
Holds. |
| L212-213: "so the failure belongs to the mechanism and not to the one method #93 names." | #93 names one node id. | Holds, and it now agrees with L197: #93 names it and attributes to it. |
L213-215: "Per #93, the two skip-variants both occurred under gate_cpu.sh, and the one hard failure in the counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct." |
#93: "the two skip-variants both occurred under gate_cpu.sh while the failure occurred under direct pytest". 5765186351 has the harness mix only in its gate table, "7 gate + 12 direct". Finding 10 attributes no outcome to a harness. |
Holds. |
| L217-219: "Per #93, the hard failure was reproduced when two gate loops on node 18 overlapped; #93 names neither that run's harness nor whether it is the one in the counts." | #93, "One further observation to record": "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially." That is all it says. #93's only comment (5768642904) reads "Landed in #99 … Closing." | Holds. |
L219-222: "Two gates on one box compete for the CPU the control arm is measuring … check for a running gate_cpu.sh before starting one." |
This is advice, not a record claim, and it is unchanged by this PR. #93 asks for it: "Say that gates should be run sequentially." | Holds. |
Consistency after the delta.
- L197 and L212-213 now say the same thing: The CPU tier has one three-way flaky test and it is recorded nowhere #93 names the method and attributes the counts to it.
- L213-215 and L217-219 still reconcile. The first is a direct-pytest run in the counts. The second is an overlap condition on the box with no harness named. Nothing asserts that they are the same event.
- "Seen again since" at L224-231 says "The gate prints no skip reasons … so that run did not name the test", which is the same fact as docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99's.
ponytail-review (principle 3: "Prioritise simplicity. Add only what is necessary, and nothing more.")
This pass is over the whole PR diff, 8c0a2e6ee..1c5d383d2, +13 / −12.
- Cycle 1's one
shrink:(L197-198) has been taken. - The L192 fold and the L213 drop are done.
- L218's second clause is what reconciles the two harness statements, so it is load-bearing.
- N3 adds one word and no line.
Lean already. Ship.
Gate 1: ATOM's suite, unmodified, on node 18 (xiaobizh_n18_cpu)
Setup
- Tip. It moved past the developer's gated
8fe47dbb5. I re-read it bygit fetch forkandgit ls-remoteat 2026-09-24T01:14:06Z:cae9607f08fbba5f94f0ada0547e9201d99895a3(compass(tests): say what the runner tests check, not which task or PR wrote them #382,tests/compass/test_runner_non_allocating.pyandtest_runner_rpc_surface.py). So the developer'sd14621b54does not stand for the landing tree, and I gated the new one once. - Merged tree.
git merge-tree --write-tree cae9607f0 1c5d383d2gives9efb0e1ac25f807eb6253379ad144c61c0e9b948, with rc=0. That is not the head's treedc7acfb8f, because the tip carries compass(spec): hold the whole transfer refusals, and every sys.modules key #374 and compass(tests): say what the runner tests check, not which task or PR wrote them #382.git diff --name-only cae9607f0 <stamp>isscripts/compass/README.mdonly. - Stamp.
git commit-tree 9efb0e1ac -p cae9607f0 -p 1c5d383d2gives24d5e573503345237f5f20ee94bcaa50139226a0. No ref was created. - Staging.
git archiveof the stamp..compass-commitis the stamp, and.compass-changedisscripts/compass/README.md.- The tar went through
docker exec -i … tar -xinto/tmp/pr376r2/m/ATOM, my own path, with md52b8a9389…on both ends. The shared mount was not written.
- Run.
- The tree's own
scripts/compass/gate_cpu.sh --junitxml, withPYTHONPATHset to the staged root, undertimeout -k 10 3000, unpiped, stdout and stderr to files. - It ran alone. I waited for two other tenants' gates (
/tmp/r383rev,/tmp/i353r2) and a third tenant's piped one to finish, and confirmed none was running at 01:13:27Z. - Load average was 22.9 → 18.8 (01:13:33Z → 01:16:58Z).
- The tree's own
Result
| tree | printed commit: |
atom.__file__ |
passed | skipped | xfailed | junit cases | stderr | GATE_CPU_RC |
|---|---|---|---|---|---|---|---|---|
merged 9efb0e1ac (stamp 24d5e5735) |
24d5e5735 (stamp) |
/tmp/pr376r2/m/ATOM/atom/__init__.py |
5268 | 155 | 3 | 5426 | 0 bytes | 0 |
- Against the developer's
d14621b54and cycle 1's477b04347, both 5268 / 155 / 3 (5426 cases), the delta is 0. compass(tests): say what the runner tests check, not which task or PR wrote them #382 changes twotests/compassfiles without changing the case count. - Timing classes, read from the junit: 17 of 17 passed, so no re-run was needed. They are
TestTheRegionIsNotCopiedPerChunk(4 of 4),TestNoSizeAtWhichACallStopsBeingOne(12 of 12) andtest_freezing_twice_is_additive_and_harmless(1 of 1). - GPU tier: the gate printed
not required (.compass-changed stamp). - Staging removed from node 18 and locally, via a script.
Gates 2 and 3.
- Gate 2 does not apply, because the change is docs-only.
- Gate 3's named result is "every claim in both paragraphs names its record, or says it has none, and the two harness statements no longer read against each other". It holds. N3 is about how wide one cited claim is, not about a claim with no record.
I did not undraft, merge, label, push or amend anything.
Closes #368
Agent-authored. Before starting I read the eight Design principles in
atom/compass/design/README.mdandatom/compass/AI_DEV_RULES.md, both at tip8c0a2e6ee.What changed
scripts/compass/README.md, the flake-rate paragraph ("What those rates are, and are not") and the one after it ("Run gates sequentially"). Docs only: +13 / −12, net +1. That is one line under the brief's 2–4, because round 2 took the review'sshrink:.Each changed sentence, with the record behind it (principle 8, "Every claim carries its measurement"). Within #93, the body is the only source; its one comment says only that it landed in #99.
xiaobizh_n18_cpu, on a box whose load was not controlled. Those conditions, the table and its one test id,test_the_cost_per_byte_does_not_grow, are issue #93's."xiaobizh_n18_cpu, on one branch"; "n=21, one mechanism, on a box whose load was not controlled"; the table 18 / 2 / 1; and under "The test:" the one node id…::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow. The old L192 attributed only "counts and n"; the conditions and the test id are folded in, and L197's separate "Per #93," is gone.…::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow, then "It is three-way, not two-way. Measured over 21 runs …", with the skip-variants in its table: that is the attribution. 5765186351, finding 10, does the same: "… and I can name it. … it does that once (4476 passed / 150 skipped …) and fails outright once". PR #99, under "Not checked", says: "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons and refuses a caller's-r… The signature is recorded; the attribution is not claimed." Round 1 read "No record names the skip-variants' method…", which cycle 1's review showed was false (B1).gate_cpu.sh, and the one hard failure in the counts under direct pytest; …"gate_cpu.shwhile the failure occurred under direct pytest". "the only record that says," is dropped, because the clause after the semicolon already says 5765186351 gives only the mix.How the two harness statements are reconciled. The old text read "The failure was reproduced with two gate loops overlapping". That puts the failure inside a gate run, right after the sentence that puts it under direct pytest. #93 does not support that reading:
The new text states exactly that, so the two sentences no longer conflict. A direct-pytest run and two overlapping gate loops on the same box are compatible. Whether they were the same event is recorded nowhere, and the text now says so.
Named result (gate 3). Every claim in both paragraphs names its record, or says it has none:
Gate 1: ATOM's suite, unmodified, on node 18 (
xiaobizh_n18_cpu), at head1c5d383d2Setup
8fe47dbb5b081935917950f9e8c55ee2d79145d6, fromgit fetch forkat 2026-09-24T00:58:05Z.git merge-tree --write-tree 8fe47dbb5 1c5d383d2givesd14621b54bd0f99b4238a2688f90ed1f66b102cb, with rc=0.git diff --name-only 8fe47dbb5 <stamp>isscripts/compass/README.mdonly.git commit-tree d14621b54 -p 8fe47dbb5 -p 1c5d383d2gives8d91a93a8b2c9e36f0813606b81eb7baf15867ab. No ref was created.git archiveof the stamp..compass-commitis the stamp, and.compass-changedisscripts/compass/README.md.docker exec -i … tar -xinto/tmp/i368r2/m2/ATOM, with md55c531f4c…on both ends.scripts/compass/gate_cpu.sh --junitxml, withPYTHONPATHset to the staged root, undertimeout -k 10 3000, unpiped.Result
commit:atom.__file__GATE_CPU_RCd14621b54(stamp8d91a93a8)8d91a93a8 (stamp)/tmp/i368r2/m2/ATOM/atom/__init__.py8f9398030, covered tip8c0a2e6ee. Its control and branch both read 5268 / 155 / 3 (5426 cases), with a node-id delta of 0. Cycle 1's review regated merged477b04347at 5268 / 155 / 3.TestTheRegionIsNotCopiedPerChunk(4),TestNoSizeAtWhichACallStopsBeingOne(12) andtest_freezing_twice_is_additive_and_harmless(1).Left undone
Nothing in the file set.
:103-108and the table at:167-171are outside it. #366's review read both as consistent with the paragraph, and this change does not alter what they rely on.🤖 Generated with Claude Code