Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 13 additions & 12 deletions scripts/compass/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -189,13 +189,13 @@ does print `GATE_CPU_RC=` on stdout on every path, so the number survives in the
*text* of an untruncated pipe — but only an unpiped run puts it in `$?`.

**What those rates are, and are not.** n=21, node 18, container `xiaobizh_n18_cpu`,
on a box whose load was not controlled. The table's counts and n are issue #93's.
19 of the 21 are pinned: PR #79's review
(issue comment 5765186351) ran `b58a48cc2` 19 times, for 17 nominal, 1 skip-variant
and the one hard failure. The other 2, 1 nominal and 1 skip-variant by subtraction
from the table, name no commit, so half the skip-variant rate rests on #93's count
alone. Per #93, all 21 runs were of **one
method**, `test_the_cost_per_byte_does_not_grow`. Not-nominal combined is ~1 in 7.
on a box whose load was not controlled. Those conditions, the table and its one test
id, `test_the_cost_per_byte_does_not_grow`, are issue #93's. 19 of the 21 are pinned:
PR #79's review (issue comment 5765186351) ran `b58a48cc2` 19 times, for 17 nominal,
1 skip-variant and the one hard failure. The other 2, 1 nominal and 1 skip-variant
by subtraction from the table, name no commit, so half the skip-variant rate rests
on #93's count alone. #93 attributes the skip-variants to that test, but no run
named it (PR #99). Not-nominal combined is ~1 in 7.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N3 (non-blocking): "no run named it" says more than PR #99 does, and read literally it is false. Principle 8: "Every claim carries its measurement. A number without a source is a defect." Here the measurement is what the cited record says, and this sentence widens it.

  • "it" is "that test", test_the_cost_per_byte_does_not_grow, and a run did name that test. 5765186351's hard-failure run printed tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow with E AssertionError: cost per KB grew 1.89x from 32 KB to 128 KB. L176-178 of this README tell the reader to rely on exactly that line: "Read the FAILED line the gate prints first … if it names this class, re-run".
  • PR docs(compass): name the CPU tier's flaky test and its three outcomes (#93) #99 is narrower. Its "Not checked" says "The ±1 skip-variant has not been attributed by name on any run", and 5766595486 says "The ±1 skip-variant has never been attributed by name". Both are about the skip-variant runs, not about the test.
  • Why it is not blocking. The contrast with "attributes the skip-variants" steers most readers to the intended scope. No reader acts wrongly on it, and it does not contradict L171 or L212-213 in practice.
  • Why it is still worth one word. This is the sentence this task exists to get exact. The wording came from cycle 1's suggested replacement, so this corrects that suggestion, not your transcription of it.

It costs one word, and the line count does not change:

on #93's count alone. #93 attributes the skip-variants to that test, but no
skip-variant run named it (PR #99). Not-nominal combined is ~1 in 7.

The finding is the third outcome, not the rate. The rates are a **lower bound on
the class**, not a measurement of it: the class holds **3 methods / 4 collected
cases** (one is parametrised `buffered-region` and `kimi-incremental`), of which
Expand All @@ -210,12 +210,13 @@ observed failing the same way — `qwen: cost per KB grew 1.73x from 32 to 128 K
comment 5765660762), which hit it on its own gate run. A sequential re-run of that
same tree passed at **4496**, and 1 + 4495 = 4496 accounts for it exactly. Inherited,
not measured here — so the failure belongs to the mechanism and not to the one
method the counts came from. Per #93, the only record that says, the two
skip-variants both occurred under `gate_cpu.sh`, and the one hard failure in the
counts under direct pytest; comment 5765186351 gives only its mix, 7 gate + 12 direct.
method #93 names. Per #93, the two skip-variants both occurred under `gate_cpu.sh`,
and the one hard failure in the counts under direct pytest; comment 5765186351
gives only its mix, 7 gate + 12 direct.

**Run gates sequentially.** The failure was reproduced with two gate loops
overlapping on node 18. Two gates on one box compete for the CPU the control arm is
**Run gates sequentially.** Per #93, the hard failure was reproduced when two gate
loops on node 18 overlapped; #93 names neither that run's harness nor whether it
is the one in the counts. Two gates on one box compete for the CPU the control arm is
measuring, which is the condition this test is least able to survive — check for a
running `gate_cpu.sh` before starting one.

Expand Down