Skip to content

compass(docs): attribute the overlapping-gate reproduction to #93 and say no record names the skip-variants method - #376

Merged
jgong5 merged 2 commits into
feature/atomcompass_newfrom
compass/issue-368-flake-attribution
Sep 24, 2026
Merged

jgong5 merged 2 commits into
feature/atomcompass_newfrom
compass/issue-368-flake-attribution

Conversation

@jgong5

@jgong5 jgong5 commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Closes #368

Agent-authored. Before starting I read the eight Design principles in atom/compass/design/README.md and atom/compass/AI_DEV_RULES.md, both at tip 8c0a2e6ee.

What changed

scripts/compass/README.md, the flake-rate paragraph ("What those rates are, and are not") and the one after it ("Run gates sequentially"). Docs only: +13 / −12, net +1. That is one line under the brief's 2–4, because round 2 took the review's shrink:.

Each changed sentence, with the record behind it (principle 8, "Every claim carries its measurement"). Within #93, the body is the only source; its one comment says only that it landed in #99.

# New sentence Record, quoted
Ponytail fold (L192) "n=21, node 18, container xiaobizh_n18_cpu, on a box whose load was not controlled. Those conditions, the table and its one test id, test_the_cost_per_byte_does_not_grow, are issue #93's." #93: "Measured over 21 runs on node 18 in xiaobizh_n18_cpu, on one branch"; "n=21, one mechanism, on a box whose load was not controlled"; the table 18 / 2 / 1; and under "The test:" the one node id …::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow. The old L192 attributed only "counts and n"; the conditions and the test id are folded in, and L197's separate "Per #93," is gone.
N2 (L197-198; round 2) "#93 attributes the skip-variants to that test, but no run named it (PR #99)." #93, under "The test:", gives the one node id …::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow, then "It is three-way, not two-way. Measured over 21 runs …", with the skip-variants in its table: that is the attribution. 5765186351, finding 10, does the same: "… and I can name it. … it does that once (4476 passed / 150 skipped …) and fails outright once". PR #99, under "Not checked", says: "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons and refuses a caller's -r … The signature is recorded; the attribution is not claimed." Round 1 read "No record names the skip-variants' method…", which cycle 1's review showed was false (B1).
L213 "…not to the one method #93 names." This replaces "the one method the counts came from", which restated the unqualified claim.
Ponytail drop (L213) "Per #93, the two skip-variants both occurred under gate_cpu.sh, and the one hard failure in the counts under direct pytest; …" #93: "the two skip-variants both occurred under gate_cpu.sh while the failure occurred under direct pytest". "the only record that says," is dropped, because the clause after the semicolon already says 5765186351 gives only the mix.
N1 (L217-219) "Per #93, the hard failure was reproduced when two gate loops on node 18 overlapped; #93 names neither that run's harness nor whether it is the one in the counts." #93, under "One further observation to record": "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially." That is all it says. It names no harness for the reproducing run, and gives it no count or commit. It does not say whether it is the table's one hard failure, which ran "under direct pytest". In round 2, "it names neither" became "#93 names neither", because the old wording read as though "it" were the hard failure.

How the two harness statements are reconciled. The old text read "The failure was reproduced with two gate loops overlapping". That puts the failure inside a gate run, right after the sentence that puts it under direct pytest. #93 does not support that reading:

  • Its overlap is a condition on the box ("when two gate loops … overlapped").
  • It does not name the harness of the reproducing run.

The new text states exactly that, so the two sentences no longer conflict. A direct-pytest run and two overlapping gate loops on the same box are compatible. Whether they were the same event is recorded nowhere, and the text now says so.

Named result (gate 3). Every claim in both paragraphs names its record, or says it has none:

Gate 1: ATOM's suite, unmodified, on node 18 (xiaobizh_n18_cpu), at head 1c5d383d2

Setup

  • Tip. 8fe47dbb5b081935917950f9e8c55ee2d79145d6, from git fetch fork at 2026-09-24T00:58:05Z.
  • Merged tree. git merge-tree --write-tree 8fe47dbb5 1c5d383d2 gives d14621b54bd0f99b4238a2688f90ed1f66b102cb, with rc=0.
    • git diff --name-only 8fe47dbb5 <stamp> is scripts/compass/README.md only.
  • Stamp. git commit-tree d14621b54 -p 8fe47dbb5 -p 1c5d383d2 gives 8d91a93a8b2c9e36f0813606b81eb7baf15867ab. No ref was created.
  • Staging.
    • git archive of the stamp. .compass-commit is the stamp, and .compass-changed is scripts/compass/README.md.
    • The tar went through docker exec -i … tar -x into /tmp/i368r2/m2/ATOM, with md5 5c531f4c… on both ends.
  • Run.
    • The tree's own scripts/compass/gate_cpu.sh --junitxml, with PYTHONPATH set to the staged root, under timeout -k 10 3000, unpiped.
    • It ran alone, after another tenant's gate finished.

Result

tree printed commit: atom.__file__ passed skipped xfailed junit cases GATE_CPU_RC
merged d14621b54 (stamp 8d91a93a8) 8d91a93a8 (stamp) /tmp/i368r2/m2/ATOM/atom/__init__.py 5268 155 3 5426 0
  • Round 1's gate, at head 8f9398030, covered tip 8c0a2e6ee. Its control and branch both read 5268 / 155 / 3 (5426 cases), with a node-id delta of 0. Cycle 1's review regated merged 477b04347 at 5268 / 155 / 3.
  • Timing classes: 17 of 17 passed, so I did not need a re-run. They are TestTheRegionIsNotCopiedPerChunk (4), TestNoSizeAtWhichACallStopsBeingOne (12) and test_freezing_twice_is_additive_and_harmless (1).
  • GPU tier: the gate printed "not required".

Left undone

Nothing in the file set. :103-108 and the table at :167-171 are outside it. #366's review read both as consistent with the paragraph, and this change does not alter what they rely on.

🤖 Generated with Claude Code

… qualify the one-method reading

The flake-rate paragraph said "Per #93, all 21 runs were of one method";
#93 names one test id and never says all 21 ran it, and comment
5765186351 names that method only for the hard failure, so the
skip-variants' method has no record. Say so, and fold the attribution of
the conditions, the table and the test id into one sentence.

The next paragraph said the failure was reproduced with two gate loops
overlapping, with no source, right after the sentence saying (per #93)
that the counted hard failure ran under direct pytest. Attribute it to
#93 and say that #93 names neither that run's harness nor whether it is
the counted one, so the two no longer read against each other.

Drop "the only record that says," which the following clause already
carries.

Closes #368

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread scripts/compass/README.md Outdated
PR #79's review (issue comment 5765186351) ran `b58a48cc2` 19 times, for 17 nominal,
1 skip-variant and the one hard failure. The other 2, 1 nominal and 1 skip-variant
by subtraction from the table, name no commit, so half the skip-variant rate rests
on #93's count alone. No record names the skip-variants' method: #93 never says

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

B1 (blocking): this sentence says no record names the skip-variants' method, but two records do. Principle 8: "Every claim carries its measurement." Here the claim is about what the records say, and it inverts them. The records attribute the skip-variants to the test by pronoun. None of them observed it.

  • The CPU tier has one three-way flaky test and it is recorded nowhere #93 body. Under "The test:" it gives …::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow, then "A wall-clock timing test with a pytest.skip noise guard", and then "It is three-way, not two-way. Measured over 21 runs on node 18 in xiaobizh_n18_cpu, on one branch:", followed by the 18 / 2 / 1 table. "It" is the named test, and its table includes the two skip-variants. The title is "The CPU tier has one three-way flaky test". The CPU tier has one three-way flaky test and it is recorded nowhere #93 never uses the words "all 21", but it does attribute all 21 to that test.
  • 5765186351, finding 10. It opens "The CPU tier's flake is three-way, not two-way, and I can name it", then says "Measured over 19 branch runs, it does that once (4476 passed / 150 skipped …) and fails outright once:". The node id comes after the failure clause, but both clauses have the same subject. It goes on: "The skip is the guard firing; the failure is the machine being noisy … So the two outcomes are one mechanism." So it does not name the method "only for the failure".
  • The record this sentence needs is PR docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99's body, under "Not checked": "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons and refuses a caller's -r … The signature is recorded; the attribution is not claimed." docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99's round-1 review (issue comment 5766595486) repeats it: "The ±1 skip-variant has never been attributed by name".

The brief asked for less than this sentence says. #368 N2: "which method they were rests on #93 alone. Qualify it." #366's review said the claim "Holds as a reading".

As written, this sentence fails the PR's own named result ("every claim … names its record, or says it has none"): it says "none" where two records exist. It also sits oddly against L214's "the one method #93 names".

Suggested replacement, which is one line shorter. It is also the ponytail shrink: in the verdict.

on #93's count alone. #93 attributes the skip-variants to that test, but no run
named it (PR #99). Not-nominal combined is ~1 in 7.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 1c5d383d2 (delta 8f9398030..1c5d383d2). L197-198 now read, in your wording:

on #93's count alone. #93 attributes the skip-variants to that test, but no run
named it (PR #99). Not-nominal combined is ~1 in 7.

That is one line shorter (−3 / +2). I re-read each record via REST before writing it:

  • The CPU tier has one three-way flaky test and it is recorded nowhere #93 body. Under "The test:" it gives …::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow, and then "It is three-way, not two-way. Measured over 21 runs on node 18 in xiaobizh_n18_cpu, on one branch:", followed by the 18 / 2 / 1 table. That is the attribution.
  • 5765186351, finding 10. "The CPU tier's flake is three-way, not two-way, and I can name it. … Measured over 19 branch runs, it does that once (4476 passed / 150 skipped …) and fails outright once". It is the same attribution, and the new sentence doesn't deny it.
  • PR docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99 body, "Not checked". "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons and refuses a caller's -r, which is how you would ask for them. The signature is recorded; the attribution is not claimed." This is the record behind "no run named it". docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99's round-1 review, 5766595486, repeats it: "The ±1 skip-variant has never been attributed by name, as the PR's "Not checked" says."

L213 now agrees. It reads "…not to the one method #93 names", and #93 names that one method and attributes the skip-variants to it.

Comment thread scripts/compass/README.md Outdated
**Run gates sequentially.** The failure was reproduced with two gate loops
overlapping on node 18. Two gates on one box compete for the CPU the control arm is
**Run gates sequentially.** Per #93, the hard failure was reproduced when two gate
loops on node 18 overlapped; it names neither that run's harness nor whether it is

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N1 (non-blocking, nit): "it names neither" reads as if "it" were the hard failure. The nearest singular noun is "the hard failure", or "node 18". "#93 names neither that run's harness nor whether it is the one in the counts" costs nothing and leaves no doubt about the referent.

The content holds, and so does the reconciliation.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 1c5d383d2. L218-219 now read "…loops on node 18 overlapped; #93 names neither that run's harness nor whether it is the one in the counts." I reflowed it over two lines and the content is unchanged.

The record behind it is still #93's "One further observation to record": "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially."

@jgong5

jgong5 commented Sep 24, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 1: #376 (issue #368), head 8f9398030415c99279d4855c48957c5e5d1ea5b0

Agent-authored. Before starting I read the eight Design principles in atom/compass/design/README.md and atom/compass/AI_DEV_RULES.md, both at tip 373f1fc5c.

Verdict: REQUEST_CHANGES, head 8f9398030415c99279d4855c48957c5e5d1ea5b0.

Quote verification (principle 8: "Every claim carries its measurement. A number without a source is a defect.")

I read every record via REST: #93, #79, #99 and #368 on both the issues/<n>/comments and pulls/<n>/comments endpoints, plus issues/comments/5765186351.

PR claim Record Result
L192-193, fold: "Those conditions, the table and its one test id, test_the_cost_per_byte_does_not_grow, are issue #93's." #93 body Holds. "Measured over 21 runs on node 18 in xiaobizh_n18_cpu, on one branch"; "n=21, one mechanism, on a box whose load was not controlled"; the 18 / 2 / 1 table; one node id under "The test:".
#93 says "one mechanism", not "one method" or "all 21" #93 body Holds, literally. The words are "n=21, one mechanism", and "all 21" does not occur. But #93 attributes the 21 to the test by pronoun: "The test: …test_the_cost_per_byte_does_not_grow … It is three-way, not two-way. Measured over 21 runs …" (see B1).
L197-198, N2: "No record names the skip-variants' method: #93 never says all 21 ran that test, and 5765186351 names it only for the failure." #93, 5765186351, #99 Does not hold (B1). #93: "It" (the named test) "is three-way … Measured over 21 runs", with its table including the two skip-variants. 5765186351, finding 10: "… and I can name it. … it does that once (4476 passed / 150 skipped …) and fails outright once", and "The skip is the guard firing". What is absent is an observation. PR #99's body, "Not checked", says: "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons … the attribution is not claimed." 5766595486 on #99 says the same.
5765186351 names the method only for the 1.89x failure 5765186351 Partly. The node-id block does follow the failure clause only. But the paragraph's subject is the named test for both the skip and the failure, as quoted above.
L213-214: "…not to the one method #93 names." #93 body Holds. #93 names one method.
L214-216, drop: "Per #93, the two skip-variants both occurred under gate_cpu.sh, and the one hard failure in the counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct." #93; 5765186351 Holds. #93: "the two skip-variants both occurred under gate_cpu.sh while the failure occurred under direct pytest". 5765186351: "7 gate + 12 direct". Dropping "the only record that says," loses nothing.
L218-220, N1: "Per #93, the hard failure was reproduced when two gate loops on node 18 overlapped; it names neither that run's harness nor whether it is the one in the counts." #93 body Holds. #93: "One further observation to record: the failure was reproduced when two gate loops on node 18 overlapped." It gives the run no harness, no commit and no count. #79 and #99 hold no other record of it. Searching for overlap|two gate|gate loops|concurren|parallel, #99's only hits are reviewers who waited rather than overlap.
"#93's body is the only source; its one comment says only that it landed in #99" issues/93/comments Holds. 5768642904 reads "Landed in #99 (the flaky-test record). Closing."

Consistency (principle 8, and the brief's named result)

Not over-hedged?

ponytail-review (principle 3: "Prioritise simplicity. Add only what is necessary, and nothing more.")

scripts/compass/README.md:L197-198: shrink: two-line absence claim over #93 and 5765186351. "#93 attributes the skip-variants to that test, but no run named it (PR #99).", one line with the reflow.
net: -1 lines possible.

N1's clause at L219-220 is load-bearing (see above); the fold at L192-193 and the drop at L213 are already done as the brief asked. The only cut is the same edit that fixes B1.

Gate 1: ATOM's suite, unmodified, on node 18 (xiaobizh_n18_cpu)

Setup

  • Tip. I re-read it at 2026-09-24T00:31:40Z by git fetch fork feature/atomcompass_new: 373f1fc5c0d7a7f13780cafb3f8c8b271a6398db. It has moved past the dev record's 8c0a2e6ee, because compass(rules): the reviewer checks gates 1-3 as they apply per task, so approval never waits for the per-wave GPU tier (#369) #371 landed.
  • Merged tree. git merge-tree --write-tree 373f1fc5c 8f9398030 gives 477b04347aa226edfb505d38d99f96755130fc92, with rc=0.
  • Stamp. git commit-tree 477b04347 -p 373f1fc5c -p 8f9398030 gives 8b488b6531b9b2fc3be2acf136d146b252261bcd. No ref was created.
  • Staging.
    • git archive of the stamp.
    • .compass-commit is the stamp, and .compass-changed is scripts/compass/README.md.
    • The tar went through docker exec -i into /tmp/pr376r1/stage/merged/ATOM, with md5 c0b2caba… on both ends.
    • The tree carries its own scripts/compass, tree e21ade45e.
  • Run.
    • The tree's own scripts/compass/gate_cpu.sh --junitxml, under timeout -k 10 3000, unpiped.
    • It ran alone: I waited for another tenant's gate_cpu.sh (/tmp/i353gates) to finish first.
    • Load average was 15.5 → 16.8.

Result

tree printed commit: atom.__file__ passed skipped xfailed junit cases GATE_CPU_RC
merged 477b04347 (stamp 8b488b653) 8b488b653 (stamp) /tmp/pr376r1/stage/merged/ATOM/atom/__init__.py 5268 155 3 5426 0

Gates 2 and 3. This is a docs-only change, so gate 2 does not apply. Gate 3's named result is "every claim in both paragraphs names its record, or says it has none". It holds for every changed sentence except L197-198, and that one is B1.

I did not undraft, merge, label or push anything.

…amed it

Review cycle 1 on #376 (B1): "No record names the skip-variants' method" was
false. #93 credits all 21 runs to test_the_cost_per_byte_does_not_grow through
"It" after "The test:", and comment 5765186351 does the same. What is missing
is an observation by name, which PR #99's "Not checked" states. Replace the
absence claim with that, one line shorter.

Nit: "it names neither" read as the hard failure; say "#93 names neither".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 24, 2026

Copy link
Copy Markdown
Owner Author

Developer round 2: #376 (issue #368), new head 1c5d383d2f0dd852a92bcea7605ab801cbdc833b

Agent-authored. Before starting I read the eight Design principles in atom/compass/design/README.md and atom/compass/AI_DEV_RULES.md, both at tip 13ab24dc7.

Delta: 8f9398030..1c5d383d2, one new commit. It is not an amend and nothing was force-pushed. It touches scripts/compass/README.md only, +4 / −5, net −1.

No blocking issues remain on my side. Both cycle-1 findings are answered, and each has an inline reply.

B1 (4088688626, reply 4088713268): fixed

L197-198 now read, in the reviewer's wording:

on #93's count alone. #93 attributes the skip-variants to that test, but no run
named it (PR #99). Not-nominal combined is ~1 in 7.

I re-read each record via REST before writing it. Principle 8: "Every claim carries its measurement."

N1 (4088688768, reply 4088713391): fixed

L218-219 now read "…loops on node 18 overlapped; #93 names neither that run's harness nor whether it is the one in the counts." The record behind it is #93's "One further observation to record": "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially."

Named result: every sentence in both paragraphs, against its record

README sentence (at 1c5d383d2) Record, quoted
L191-193: "n=21, node 18, container xiaobizh_n18_cpu, on a box whose load was not controlled. Those conditions, the table and its one test id … are issue #93's." #93: "Measured over 21 runs on node 18 in xiaobizh_n18_cpu"; "n=21, one mechanism, on a box whose load was not controlled"; one node id under "The test:".
L193-195: "19 of the 21 are pinned: PR #79's review (issue comment 5765186351) ran b58a48cc2 19 times, for 17 nominal, 1 skip-variant and the one hard failure." 5765186351: "| Branch | b58a48cc2 | 7 gate + 12 direct | 4477 passed, 149 skipped, 3 xfailed, rc=0 in 17 of 19 |"; "Measured over 19 branch runs, it does that once (4476 passed / 150 skipped …) and fails outright once".
L195-197: "The other 2 … by subtraction from the table, name no commit" #93's table has 18 / 2 / 1. Subtracting 17 / 1 / 1 leaves 1 nominal and 1 skip-variant. #93 names no commit, only "on one branch".
L197-198: "#93 attributes the skip-variants to that test, but no run named it (PR #99)." See B1 above.
L198: "Not-nominal combined is ~1 in 7." #93: "Not-nominal combined is ~1 in 7."
L199: "The finding is the third outcome, not the rate." #93: "The finding is the third outcome, not the rate."
L199-205: "the class holds 3 methods / 4 collected cases … of which 2 methods / 3 cases share the same timing helpers, the same noise guard and the same < 1.5 ratio assertion"; the open-region method "is deterministic" The test file at the tip, tests/entrypoints/test_stream_marker_properties.py, TestTheRegionIsNotCopiedPerChunk: test_the_open_region_is_never_scanned_beyond_the_window, then test_no_format_pays_more_per_byte_as_the_payload_grows with ids=["buffered-region", "kimi-incremental"], then test_the_cost_per_byte_does_not_grow. The last two each carry pytest.skip(f"machine too noisy to measure: …") and assert large / small < 1.5. The junit for this gate has 4 cases in the class.
L205-211: the sibling's "qwen: cost per KB grew 1.73x from 32 to 128 KB, 1 failed, 4495 passed, rc=1, on the integration head 669dc3f9d", and the re-run at 4496 5765660762: "— "qwen: cost per KB grew 1.73x from 32 to 128 KB" — 1 failed / 4495 passed, rc=1 … Re-run of that one tree, sequentially: 4496 passed, rc=0, and 1 + 4495 = 4496 accounts for it exactly."; the head is 669dc3f9d.
L212-213: "…not to the one method #93 names." #93 names one node id.
L213-215: "Per #93, the two skip-variants both occurred under gate_cpu.sh, and the one hard failure in the counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct." #93: "the two skip-variants both occurred under gate_cpu.sh while the failure occurred under direct pytest". 5765186351: "7 gate + 12 direct".
L217-219: "Per #93, the hard failure was reproduced when two gate loops on node 18 overlapped; #93 names neither that run's harness nor whether it is the one in the counts." #93: "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially." It says nothing more about that run.
L219-222: "Two gates on one box compete for the CPU the control arm is measuring … check for a running gate_cpu.sh before starting one." This is advice, and its record is #93: "Say that gates should be run sequentially." The mechanism is 5765186351: "the failure is the machine being noisy in a way the guard does not catch".

Gate 1: ATOM's suite, unmodified, on node 18 (xiaobizh_n18_cpu)

Setup

  • Tip. It moved while I waited for the box: 13ab24dc7 became 8fe47dbb5b081935917950f9e8c55ee2d79145d6 (compass(spec): hold the whole transfer refusals, and every sys.modules key #374, tests/compass/test_spec_verbs.py only). I re-read it at 2026-09-24T00:58:05Z by git fetch fork and gated against it.
  • Merged tree. git merge-tree --write-tree 8fe47dbb5 1c5d383d2 gives d14621b54bd0f99b4238a2688f90ed1f66b102cb, with rc=0.
    • This is not the head's tree dc7acfb8f, because the tip carries commits the branch does not.
    • git diff --name-only 8fe47dbb5 <stamp> is scripts/compass/README.md only.
  • Stamp. git commit-tree d14621b54 -p 8fe47dbb5 -p 1c5d383d2 gives 8d91a93a8b2c9e36f0813606b81eb7baf15867ab. No ref was created.
  • Staging.
    • git archive of the stamp. .compass-commit is the stamp, and .compass-changed is scripts/compass/README.md.
    • The tar went through docker exec -i … tar -x into /tmp/i368r2/m2/ATOM, with md5 5c531f4c… on both ends.
  • Run.
    • The tree's own scripts/compass/gate_cpu.sh --junitxml, with PYTHONPATH set to the staged root, under timeout -k 10 3000, unpiped.
    • It ran alone: I waited for another tenant's gate (/tmp/pr374r1) to finish first.
    • Load average was 11.15 at the start.

Result

tree printed commit: atom.__file__ passed skipped xfailed junit cases GATE_CPU_RC
merged d14621b54 (stamp 8d91a93a8) 8d91a93a8 (stamp) /tmp/i368r2/m2/ATOM/atom/__init__.py 5268 155 3 5426 0
  • Against cycle 1's 5268 / 155 / 3 (5426 cases), the delta is 0. The README change is docs only, and compass(spec): hold the whole transfer refusals, and every sys.modules key #374 changes assertions without changing the case count.
  • Timing classes: 17 of 17 passed, so I did not need a re-run. They are TestTheRegionIsNotCopiedPerChunk (4), TestNoSizeAtWhichACallStopsBeingOne (12) and test_freezing_twice_is_additive_and_harmless (1).
  • GPU tier: not required.
  • Staging removed from node 18 and locally, via a script.

I did not undraft, merge or label anything.

Comment thread scripts/compass/README.md
1 skip-variant and the one hard failure. The other 2, 1 nominal and 1 skip-variant
by subtraction from the table, name no commit, so half the skip-variant rate rests
on #93's count alone. #93 attributes the skip-variants to that test, but no run
named it (PR #99). Not-nominal combined is ~1 in 7.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N3 (non-blocking): "no run named it" says more than PR #99 does, and read literally it is false. Principle 8: "Every claim carries its measurement. A number without a source is a defect." Here the measurement is what the cited record says, and this sentence widens it.

  • "it" is "that test", test_the_cost_per_byte_does_not_grow, and a run did name that test. 5765186351's hard-failure run printed tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow with E AssertionError: cost per KB grew 1.89x from 32 KB to 128 KB. L176-178 of this README tell the reader to rely on exactly that line: "Read the FAILED line the gate prints first … if it names this class, re-run".
  • PR docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99 is narrower. Its "Not checked" says "The ±1 skip-variant has not been attributed by name on any run", and 5766595486 says "The ±1 skip-variant has never been attributed by name". Both are about the skip-variant runs, not about the test.
  • Why it is not blocking. The contrast with "attributes the skip-variants" steers most readers to the intended scope. No reader acts wrongly on it, and it does not contradict L171 or L212-213 in practice.
  • Why it is still worth one word. This is the sentence this task exists to get exact. The wording came from cycle 1's suggested replacement, so this corrects that suggestion, not your transcription of it.

It costs one word, and the line count does not change:

on #93's count alone. #93 attributes the skip-variants to that test, but no
skip-variant run named it (PR #99). Not-nominal combined is ~1 in 7.

@jgong5

jgong5 commented Sep 24, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 2 (delta): #376 (issue #368), head 1c5d383d2f0dd852a92bcea7605ab801cbdc833b

Agent-authored. Before starting I read the eight Design principles in atom/compass/design/README.md and atom/compass/AI_DEV_RULES.md, both at tip cae9607f0.

Verdict: APPROVE, head 1c5d383d2f0dd852a92bcea7605ab801cbdc833b.

If N3 is taken, the head moves and needs a delta review before landing. If it is not taken, it gets an issue, per "A finding not fixed in the PR that found it gets an issue".

Scope

  • Delta 8f9398030..1c5d383d2: one commit, scripts/compass/README.md only, +4 / −5. It is not an amend.
    • Hunk 1 is L197-198, the B1 sentence.
    • Hunk 2 is L217-219, the N1 referent.
  • Whole PR 8c0a2e6ee..1c5d383d2: +13 / −12, the same file. That is net +1 against the brief's "about 2–4 lines".
  • L212-213 ("…not to the one method The CPU tier has one three-way flaky test and it is recorded nowhere #93 names") is from round 1 (8f9398030) and is untouched by this delta. I re-read it against the new L197-198 anyway (below).

Cycle-1 findings

Finding At head Status
B1, blocking (4088688626): "No record names the skip-variants' method" inverted #93 and 5765186351 "#93 attributes the skip-variants to that test, but no run named it (PR #99)." Resolved. The absence claim is gone. The attribution is credited to #93, and the qualification cites PR #99. Checked below. N3 is a new, narrower point about the second clause.
N1, non-blocking (4088688768): "it names neither" had an ambiguous referent "…overlapped; #93 names neither that run's harness nor whether it is the one in the counts." Resolved. The subject is now explicit, and the content is unchanged.

Every sentence at L191-222, against its record (principle 8: "Every claim carries its measurement. A number without a source is a defect.")

I read each record myself via REST. I did not take the developer's table.

README (head) Record, quoted Result
L191-194: "n=21, node 18, container xiaobizh_n18_cpu, on a box whose load was not controlled. Those conditions, the table and its one test id, test_the_cost_per_byte_does_not_grow, are issue #93's." #93: "Measured over 21 runs on node 18 in xiaobizh_n18_cpu, on one branch"; "n=21, one mechanism, on a box whose load was not controlled"; the 18 / 2 / 1 table; one node id under "The test:". Holds.
L194-196: "19 of the 21 are pinned: PR #79's review (issue comment 5765186351) ran b58a48cc2 19 times, for 17 nominal, 1 skip-variant and the one hard failure." 5765186351: "| Branch | b58a48cc2 | 7 gate + 12 direct | 4477 passed … rc=0 in 17 of 19 |", and finding 10: "Measured over 19 branch runs, it does that once (4476 passed / 150 skipped …) and fails outright once", with cost per KB grew 1.89x. Holds for the figures. That these 19 are 19 of #93's 21 is a reading: the 1.89x failure, the container and 17/1/1 ⊂ 18/2/1 all agree, and neither record states it. The wording predates this PR, which only reflowed it. Not raised.
L196-197: "The other 2, 1 nominal and 1 skip-variant by subtraction from the table, name no commit, so half the skip-variant rate rests on #93's count alone." 18 − 17 = 1, 2 − 1 = 1, 1 − 1 = 0. #93 names no sha, only "on one branch". Holds.
L197-198: "#93 attributes the skip-variants to that test, but no run named it (PR #99)." #93: "The test:" …::test_the_cost_per_byte_does_not_grow, then "It is three-way, not two-way. Measured over 21 runs …", with the skip-variant row in the table. #99 body, "Not checked": "The ±1 skip-variant has not been attributed by name on any run: the gate prints no skip reasons and refuses a caller's -r … The signature is recorded; the attribution is not claimed." 5766595486: "The ±1 skip-variant has never been attributed by name". First clause holds. Second clause holds in scope but is wider than its record (N3). #99 says no skip-variant run named it. 5765186351's failing run did name the test.
L198: "Not-nominal combined is ~1 in 7." #93, verbatim. 3 / 21 = 1 / 7. Holds.
L199: "The finding is the third outcome, not the rate." #93, verbatim. Holds.
L199-205: "3 methods / 4 collected cases", "2 methods / 3 cases share the same timing helpers, the same noise guard and the same < 1.5 ratio assertion"; the open-region method "counts scans of the open region and is deterministic" The test file tests/entrypoints/test_stream_marker_properties.py, byte-identical at head and at cae9607f0, TestTheRegionIsNotCopiedPerChunk. test_the_open_region_is_never_scanned_beyond_the_window calls _scans_of_an_open_region and asserts max(seen) <= _PEEK_WINDOW. test_no_format_pays_more_per_byte_as_the_payload_grows has ids=["buffered-region", "kimi-incremental"], and it and test_the_cost_per_byte_does_not_grow each call _control_ms/_stream_ms, carry if not 0.6 < control < 1.6: pytest.skip(…), and assert large / small < 1.5. Holds.
L205-212: the sibling, "qwen: cost per KB grew 1.73x from 32 to 128 KB, 1 failed, 4495 passed, rc=1, on the integration head 669dc3f9d; node 18, xiaobizh_n18_cpu, 2026-09-21, recorded in the round-2 review of PR #67 (issue comment 5765660762), which hit it on its own gate run. A sequential re-run of that same tree passed at 4496" 5765660762 (on #67, 2026-09-21T18:42:14Z, "Round 2, second reviewer"): "The integration head's first run failed with …::test_no_format_pays_more_per_byte_as_the_payload_grows[buffered-region] — "qwen: cost per KB grew 1.73x from 32 to 128 KB" — 1 failed / 4495 passed, rc=1. … Re-run of that one tree, sequentially: 4496 passed, rc=0". The head is "feature/atomcompass_new is 669dc3f9d", and the conditions are "Node 18, xiaobizh_n18_cpu". Holds.
L212-213: "so the failure belongs to the mechanism and not to the one method #93 names." #93 names one node id. Holds, and it now agrees with L197: #93 names it and attributes to it.
L213-215: "Per #93, the two skip-variants both occurred under gate_cpu.sh, and the one hard failure in the counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct." #93: "the two skip-variants both occurred under gate_cpu.sh while the failure occurred under direct pytest". 5765186351 has the harness mix only in its gate table, "7 gate + 12 direct". Finding 10 attributes no outcome to a harness. Holds.
L217-219: "Per #93, the hard failure was reproduced when two gate loops on node 18 overlapped; #93 names neither that run's harness nor whether it is the one in the counts." #93, "One further observation to record": "the failure was reproduced when two gate loops on node 18 overlapped. Say that gates should be run sequentially." That is all it says. #93's only comment (5768642904) reads "Landed in #99 … Closing." Holds.
L219-222: "Two gates on one box compete for the CPU the control arm is measuring … check for a running gate_cpu.sh before starting one." This is advice, not a record claim, and it is unchanged by this PR. #93 asks for it: "Say that gates should be run sequentially." Holds.

Consistency after the delta.

ponytail-review (principle 3: "Prioritise simplicity. Add only what is necessary, and nothing more.")

This pass is over the whole PR diff, 8c0a2e6ee..1c5d383d2, +13 / −12.

  • Cycle 1's one shrink: (L197-198) has been taken.
  • The L192 fold and the L213 drop are done.
  • L218's second clause is what reconciles the two harness statements, so it is load-bearing.
  • N3 adds one word and no line.
Lean already. Ship.

Gate 1: ATOM's suite, unmodified, on node 18 (xiaobizh_n18_cpu)

Setup

  • Tip. It moved past the developer's gated 8fe47dbb5. I re-read it by git fetch fork and git ls-remote at 2026-09-24T01:14:06Z: cae9607f08fbba5f94f0ada0547e9201d99895a3 (compass(tests): say what the runner tests check, not which task or PR wrote them #382, tests/compass/test_runner_non_allocating.py and test_runner_rpc_surface.py). So the developer's d14621b54 does not stand for the landing tree, and I gated the new one once.
  • Merged tree. git merge-tree --write-tree cae9607f0 1c5d383d2 gives 9efb0e1ac25f807eb6253379ad144c61c0e9b948, with rc=0. That is not the head's tree dc7acfb8f, because the tip carries compass(spec): hold the whole transfer refusals, and every sys.modules key #374 and compass(tests): say what the runner tests check, not which task or PR wrote them #382. git diff --name-only cae9607f0 <stamp> is scripts/compass/README.md only.
  • Stamp. git commit-tree 9efb0e1ac -p cae9607f0 -p 1c5d383d2 gives 24d5e573503345237f5f20ee94bcaa50139226a0. No ref was created.
  • Staging.
    • git archive of the stamp. .compass-commit is the stamp, and .compass-changed is scripts/compass/README.md.
    • The tar went through docker exec -i … tar -x into /tmp/pr376r2/m/ATOM, my own path, with md5 2b8a9389… on both ends. The shared mount was not written.
  • Run.
    • The tree's own scripts/compass/gate_cpu.sh --junitxml, with PYTHONPATH set to the staged root, under timeout -k 10 3000, unpiped, stdout and stderr to files.
    • It ran alone. I waited for two other tenants' gates (/tmp/r383rev, /tmp/i353r2) and a third tenant's piped one to finish, and confirmed none was running at 01:13:27Z.
    • Load average was 22.9 → 18.8 (01:13:33Z → 01:16:58Z).

Result

tree printed commit: atom.__file__ passed skipped xfailed junit cases stderr GATE_CPU_RC
merged 9efb0e1ac (stamp 24d5e5735) 24d5e5735 (stamp) /tmp/pr376r2/m/ATOM/atom/__init__.py 5268 155 3 5426 0 bytes 0
  • Against the developer's d14621b54 and cycle 1's 477b04347, both 5268 / 155 / 3 (5426 cases), the delta is 0. compass(tests): say what the runner tests check, not which task or PR wrote them #382 changes two tests/compass files without changing the case count.
  • Timing classes, read from the junit: 17 of 17 passed, so no re-run was needed. They are TestTheRegionIsNotCopiedPerChunk (4 of 4), TestNoSizeAtWhichACallStopsBeingOne (12 of 12) and test_freezing_twice_is_additive_and_harmless (1 of 1).
  • GPU tier: the gate printed not required (.compass-changed stamp).
  • Staging removed from node 18 and locally, via a script.

Gates 2 and 3.

  • Gate 2 does not apply, because the change is docs-only.
  • Gate 3's named result is "every claim in both paragraphs names its record, or says it has none, and the two harness statements no longer read against each other". It holds. N3 is about how wide one cited claim is, not about a claim with no record.

I did not undraft, merge, label, push or amend anything.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant