From 8f9398030415c99279d4855c48957c5e5d1ea5b0 Mon Sep 17 00:00:00 2001 From: Jiong Gong Date: Thu, 24 Sep 2026 00:05:34 +0000 Subject: [PATCH 1/2] compass(docs): attribute the overlapping-gate reproduction to #93 and qualify the one-method reading The flake-rate paragraph said "Per #93, all 21 runs were of one method"; #93 names one test id and never says all 21 ran it, and comment 5765186351 names that method only for the hard failure, so the skip-variants' method has no record. Say so, and fold the attribution of the conditions, the table and the test id into one sentence. The next paragraph said the failure was reproduced with two gate loops overlapping, with no source, right after the sentence saying (per #93) that the counted hard failure ran under direct pytest. Attribute it to #93 and say that #93 names neither that run's harness nor whether it is the counted one, so the two no longer read against each other. Drop "the only record that says," which the following clause already carries. Closes #368 Co-Authored-By: Claude Opus 5.5 (1M context) --- scripts/compass/README.md | 26 ++++++++++++++------------ 1 file changed, 14 insertions(+), 12 deletions(-) diff --git a/scripts/compass/README.md b/scripts/compass/README.md index e06007e793..56dc3edb41 100644 --- a/scripts/compass/README.md +++ b/scripts/compass/README.md @@ -189,13 +189,14 @@ does print `GATE_CPU_RC=` on stdout on every path, so the number survives in the *text* of an untruncated pipe — but only an unpiped run puts it in `$?`. **What those rates are, and are not.** n=21, node 18, container `xiaobizh_n18_cpu`, -on a box whose load was not controlled. The table's counts and n are issue #93's. -19 of the 21 are pinned: PR #79's review -(issue comment 5765186351) ran `b58a48cc2` 19 times, for 17 nominal, 1 skip-variant -and the one hard failure. The other 2, 1 nominal and 1 skip-variant by subtraction -from the table, name no commit, so half the skip-variant rate rests on #93's count -alone. Per #93, all 21 runs were of **one -method**, `test_the_cost_per_byte_does_not_grow`. Not-nominal combined is ~1 in 7. +on a box whose load was not controlled. Those conditions, the table and its one test +id, `test_the_cost_per_byte_does_not_grow`, are issue #93's. 19 of the 21 are pinned: +PR #79's review (issue comment 5765186351) ran `b58a48cc2` 19 times, for 17 nominal, +1 skip-variant and the one hard failure. The other 2, 1 nominal and 1 skip-variant +by subtraction from the table, name no commit, so half the skip-variant rate rests +on #93's count alone. No record names the skip-variants' method: #93 never says +all 21 ran that test, and 5765186351 names it only for the failure. Not-nominal +combined is ~1 in 7. The finding is the third outcome, not the rate. The rates are a **lower bound on the class**, not a measurement of it: the class holds **3 methods / 4 collected cases** (one is parametrised `buffered-region` and `kimi-incremental`), of which @@ -210,12 +211,13 @@ observed failing the same way — `qwen: cost per KB grew 1.73x from 32 to 128 K comment 5765660762), which hit it on its own gate run. A sequential re-run of that same tree passed at **4496**, and 1 + 4495 = 4496 accounts for it exactly. Inherited, not measured here — so the failure belongs to the mechanism and not to the one -method the counts came from. Per #93, the only record that says, the two -skip-variants both occurred under `gate_cpu.sh`, and the one hard failure in the -counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct. +method #93 names. Per #93, the two skip-variants both occurred under `gate_cpu.sh`, +and the one hard failure in the counts under direct pytest; comment 5765186351 +gives only its mix, 7 gate + 12 direct. -**Run gates sequentially.** The failure was reproduced with two gate loops -overlapping on node 18. Two gates on one box compete for the CPU the control arm is +**Run gates sequentially.** Per #93, the hard failure was reproduced when two gate +loops on node 18 overlapped; it names neither that run's harness nor whether it is +the one in the counts. Two gates on one box compete for the CPU the control arm is measuring, which is the condition this test is least able to survive — check for a running `gate_cpu.sh` before starting one. From 1c5d383d2f0dd852a92bcea7605ab801cbdc833b Mon Sep 17 00:00:00 2001 From: Jiong Gong Date: Thu, 24 Sep 2026 00:41:13 +0000 Subject: [PATCH 2/2] compass(docs): #93 attributes the skip-variants to the test; no run named it Review cycle 1 on #376 (B1): "No record names the skip-variants' method" was false. #93 credits all 21 runs to test_the_cost_per_byte_does_not_grow through "It" after "The test:", and comment 5765186351 does the same. What is missing is an observation by name, which PR #99's "Not checked" states. Replace the absence claim with that, one line shorter. Nit: "it names neither" read as the hard failure; say "#93 names neither". Co-Authored-By: Claude Opus 5.5 (1M context) --- scripts/compass/README.md | 9 ++++----- 1 file changed, 4 insertions(+), 5 deletions(-) diff --git a/scripts/compass/README.md b/scripts/compass/README.md index 56dc3edb41..f97fec04e8 100644 --- a/scripts/compass/README.md +++ b/scripts/compass/README.md @@ -194,9 +194,8 @@ id, `test_the_cost_per_byte_does_not_grow`, are issue #93's. 19 of the 21 are pi PR #79's review (issue comment 5765186351) ran `b58a48cc2` 19 times, for 17 nominal, 1 skip-variant and the one hard failure. The other 2, 1 nominal and 1 skip-variant by subtraction from the table, name no commit, so half the skip-variant rate rests -on #93's count alone. No record names the skip-variants' method: #93 never says -all 21 ran that test, and 5765186351 names it only for the failure. Not-nominal -combined is ~1 in 7. +on #93's count alone. #93 attributes the skip-variants to that test, but no run +named it (PR #99). Not-nominal combined is ~1 in 7. The finding is the third outcome, not the rate. The rates are a **lower bound on the class**, not a measurement of it: the class holds **3 methods / 4 collected cases** (one is parametrised `buffered-region` and `kimi-incremental`), of which @@ -216,8 +215,8 @@ and the one hard failure in the counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct. **Run gates sequentially.** Per #93, the hard failure was reproduced when two gate -loops on node 18 overlapped; it names neither that run's harness nor whether it is -the one in the counts. Two gates on one box compete for the CPU the control arm is +loops on node 18 overlapped; #93 names neither that run's harness nor whether it +is the one in the counts. Two gates on one box compete for the CPU the control arm is measuring, which is the condition this test is least able to survive — check for a running `gate_cpu.sh` before starting one.