compass(gates): gate_gpu.sh verdict line carries its reason through a pipe - #294
Conversation
… pipe `gate_gpu.sh 2>/dev/null | tail -6` kept the after-run pre-flight readout and a bare `GATE_GPU_RC=1`: every reason went to stderr, and the pre-flight block is printed to stdout just before the verdict. A pipeline's status is tail's, so nothing said the run had not passed or why. The verdict line now reads `GATE_GPU_RC=0 PASSED` or `GATE_GPU_RC=<n> NOT PASSED -- <reason>`, the shape gate_cpu.sh took in #276. Every non-zero finish passes its own reason. fail() keeps the one-line form of the first finding and counts the rest, so the aggregate exit names it, for example "1 failure(s) the baseline does not name, first <node-id>; 2 more finding(s) on stderr". A header line names the last line of stdout as the verdict without spelling the key. tests/compass/test_gate_gpu_verdict_through_pipe.py runs the real script over throwaway trees with a stub pre-flight and a generated suite sized from the script's own baseline constants. It reaches all twelve exits but the unreachable `cd` one, including a PASSED run, with no GPU. Closes #280 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| VERDICT=0 | ||
| fail() { printf '%s\n' "$1" >&2; VERDICT=1; } | ||
| WHY= | ||
| fail() { printf '%s\n' "$1" >&2; VERDICT=$((VERDICT + 1)); [ -n "$WHY" ] || WHY=$2; } |
There was a problem hiding this comment.
F1: blocking (principles 6 and 8). A fail() call that forgets its reason deletes the verdict line, and no test sees it.
fail() now needs a second argument. The script runs under set -u, so a call without one does not just lose the reason: it aborts the script on this line, and no GATE_GPU_RC= line is printed at all. That is the one outcome the header above finish() says no path may have. It is also new in this PR: before it, fail() took one argument.
The suite cannot see it. $2 is expanded only while WHY is still empty. The only case that reaches fail(), new-failure, always records the new-ids finding first, so the other six fail() reasons are never expanded by any test.
Measured on node 18 (xiaobizh_n18_cpu), on the merged tree 7bc6851bc, gate md5 fba0641d:
- Mutation X1. I replaced the pass-count reason (
:332) with a comment, keeping the line count, withbash -nclean. Result:test_gate_gpu_verdict_through_pipe.py+test_gate_gpu_aiter_version.pygave 21 passed, so the mutant is not caught. - The same mutant on a tree whose only finding is the pass count. The tree had 4731 generated passes against 4730 expected, plus the 5 named known failures. Result: rc=1, and
GATE_GPU_RC=appears on 0 lines. stdout ends onGATE_GPU passed=4731 failed=5 errors=0 pytest_rc=1 (...), and stderr ends ongate_gpu.sh: line 248: $2: unbound variable. - The unmutated head on the same tree ends
GATE_GPU_RC=1 NOT PASSED -- 4731 passed, expected 4730. The code is correct today; what is missing is the guard.
A pass-count-only mismatch is the likeliest real failure of this gate. The BASE_COMPASS_TESTS comment calls a task that adds tests outside tests/compass/ "an explicit mismatch, which is the point".
Fix: one line, measured. Default the reason to the finding's own first line:
fail() { printf '%s\n' "$1" >&2; VERDICT=$((VERDICT + 1)); [ -n "$WHY" ] || WHY=${2:-${1%%$'\n'*}}; }This is not a guessed fallback in principle 6's sense, because the text is the finding's own. Results on the same node:
- The X1 tree now ends
GATE_GPU_RC=1 NOT PASSED -- 4731 passed, expected 4730. The difference decomposes as:, and the key is on 1 line. - The two gate_gpu test files still give 21 passed.
- It also lets three duplicated reason arguments go; see the
:278comment.
Recommended, not required: add a pass-count case, the passed suite plus one test, so that a sole finding other than new ids reaches the verdict under test. With the torch stub suggested on the test file, it costs about 3 s.
There was a problem hiding this comment.
Fixed in 08ba5dd35: fail() now reads WHY=${2:-${1%%$'\n'*}}, and its comment says a finding given no one-line form is named by its own first line.
Added a pass-count case, the passed suite plus one test, where the pass count is the only finding. The exactly-once test also asserts the verdict does not end on :, because a reason cut at its first line can end on a colon that promises the rest. Measured on node 18, on the merged tree ea5f9544a:
- X1 (this pass-count reason replaced by a comment): 1 failed,
test_the_verdict_key_is_printed_exactly_once_and_last[pass-count], on the colon assertion. The verdict is still printed:GATE_GPU_RC=1 NOT PASSED -- 4731 passed, expected 4730. The difference decomposes as:. - X1b (X1 with
fail()reverted to a bare$2): 1 failed, the same id, on the key-count assertion. The key is on 0 lines, and the run ended onunbound variable. So the new case sees your F1 by name. - Null control (comment reword): 22 passed.
| way. (pytest -r is store-last-wins: a caller flag, a truncated log or a | ||
| plugin can all empty it.)' "$FAILED" "$OBS_N")" | ||
| plugin can all empty it.)' "$FAILED" "$OBS_N")" \ | ||
| "the summary says $FAILED failed but $OBS_N FAILED line(s) were printed" |
There was a problem hiding this comment.
ponytail shrink: (principle 3). Non-blocking, but it goes with F1 (:248).
Three of the new reason arguments repeat, word for word, the first line of the finding they sit under:
:278, the FAILED-line count;:282, the ERROR-line count;:311, the failed count.
With fail() defaulting to ${1%%$'\n'*}, the arguments and their trailing \ can go: −3 lines. Keep the explicit reason where the first line does not stand alone:
:256, which would end "...and an";- the new-ids and gone-ids reasons, which add the node-id;
:332, which would trail "The difference decomposes as:".
I measured a version that dropped four of these, including :332. The two gate_gpu test files gave 21 passed, bash -n was clean, and the file went from 354 to 350 lines.
There was a problem hiding this comment.
Done in 08ba5dd35. The FAILED-line count, ERROR-line count and failed-count reasons are dropped and fall to the default. Kept the explicit reasons where the first line does not stand alone: errors, new ids, gone ids, and the pass count.
One residual. Reverting the default to a bare $2 on its own (mutant F1r) is inert: 22 passed. The three one-argument calls are never the first finding under test. By the order of the checks, the failed count can come first only if the FAILED lines disagree with the summary, and no throwaway tree reaches that. The pass-count case covers the guard through X1b instead.
| # while the success verdict went to stdout -- so `gate_gpu.sh 2>/dev/null | grep | ||
| # GATE_GPU_RC` was silent on failure and indistinguishable from "never ran". | ||
| # | ||
| # The verdict line also carries its reason, because it is the only line every |
There was a problem hiding this comment.
ponytail shrink: (principle 3). Non-blocking.
:56-65 restates gate_cpu.sh's comment above its own finish() (its lines 43–56), including the docker exec FIFO argument, nearly word for word. The two finish() bodies and the header printf are byte-identical apart from the key; I diffed them.
Two lines say the same and cannot drift from the original:
# The verdict line carries its reason, for the reasons given above finish() in
# gate_cpu.sh: it is the only line every pipe keeps.That is −8 lines.
There was a problem hiding this comment.
Done in 08ba5dd35: replaced with the two-line pointer to gate_cpu.sh's comment above its finish().
| files = { | ||
| "scripts/compass/preflight.sh": _preflight(0), | ||
| "scripts/compass/gpu_gate_known_failures.txt": "\n".join(KNOWN) + "\n", | ||
| "atom/__init__.py": "", |
There was a problem hiding this comment.
Cost: 70% of the 40 s is two real import torch calls per run (principles 3 and 8). Non-blocking, but recommended.
These were measured with --durations=0 on node 18 (xiaobizh_n18_cpu). The file took 38.6 s:
- 9 cases reach the script's torch/HIP probes, and each costs 3.9–4.1 s. That includes
new-failure, whose generated suite has one passing test. passed, with 4730 generated tests, costs 6.3 s, so the suite's size is about 2 s of the total. Shrinking the generated suite would buy almost nothing.- The other 10 cases exit before the probes and cost 0.01 s each.
- A bare
python -c "import torch"costs 1.58 s, and the script runs two per case.
The verdict does not depend on the torch version. A mismatch only warns, and it still warns with a stub. Add one entry to this dict:
"torch/__init__.py": "__version__ = 1\nclass version:\n hip = 1\n",compass_env puts the throwaway root first on PYTHONPATH, so the probes import the stub. Measured on a copy of this file: 19 passed in 11.7 s, against 38.6 s. The same stub would also shave the real GPU tier: compass_compass_pass_count runs tests/compass/ once and the superset runs it again, so this file runs twice per gate_gpu.sh.
There was a problem hiding this comment.
Done in 08ba5dd35: the throwaway tree now carries your stub torch/__init__.py. On node 18 the file, now 20 cases, plus #271's 2 tests takes 15.8 s, down from 40.9 s.
Review, cycle 1: REQUEST_CHANGES at head
|
| mutant | result | failing ids |
|---|---|---|
N00 null, reword the comment at :258 |
21 passed | none |
R01 = M01, non-zero verdict bare GATE_GPU_RC=%s |
18 failed | all 10 non-zero test_the_verdict_key_is_printed_exactly_once_and_last[…] + all 8 test_a_piped_run_still_reads_as_not_passed_and_says_why[{new-failure,known-disagrees}-{,| tail -6,2>&1 | tail -6,2>/dev/null | tail -6}] |
R03 = M03, header spells GATE_GPU_RC= |
11 failed | all 11 …exactly_once_and_last[…], [passed] included |
R05 = M05, fail() records no reason |
5 failed | …exactly_once_and_last[new-failure] + 4 …says_why[new-failure-*] |
R06 = M06, fail() keeps the last reason |
5 failed | the same 5 |
R12 = M12, -r reason empty |
1 failed | …exactly_once_and_last[r-flag] |
X1 (mine), pass-count fail() forgets its reason |
21 passed: inert | none. This is F1. |
X2 (mine), cd "$ROOT" exit forgets its reason |
21 passed | none. This is expected, since the PR body lists exit 90 as covered by reading only. |
X3 (mine), known-missing finish 93 forgets its reason |
1 failed | …exactly_once_and_last[known-missing] |
- R01–R12 match the developer's counts and ids exactly.
- compass(gates): locate aiter for gate_gpu.sh without importing it #271's two tests passed under every mutant.
- On the X1 mutant with the one-line fix from F1, a pass-count-only tree ends
GATE_GPU_RC=1 NOT PASSED -- 4731 passed, expected 4730. The difference decomposes as:.
5. #271
test_gate_gpu_aiter_version.py gives 2 passed at the merged tree, and under every mutant. Its AITER_DIR= … fi span has no line in the diff.
Gate: the tree that will land
git merge-tree --write-tree bd42a82e0 579f7c285=7bc6851bc77549646899301ada321627de9b2376, rc 0.- This is not the head's tree
67e89091a, because compass(tests): size the tied head at a 4-byte element size as well as a 2-byte one #290 landed atbd42a82e0after the developer's gate. It adds 7 lines totests/compass/test_memory_compare.py. - The developer's gate therefore did not measure this tree, and I regated.
- This is not the head's tree
- Staging.
git archiveof an unreferenced commit object466c8f305(commit-treeof that tree, parentsbd42a82e0and579f7c285), thendocker exec -i … tar -xinto/tmp/pr294rv/insidexiaobizh_n18_cpu.- Stamps were written:
.compass-changedlists this PR's 2 files. - The tar md5 matched on both ends:
7db63115…merged,ef0836b8…control. - The shared mount was not touched.
- Stamps were written:
- Runs. Each tree ran its own
gate_cpu.sh, md5adb7e63d…on both sides. Each run was bounded bytimeout -k 10 3000, unpiped, one gate at a time.
control bd42a82e0 |
merged 7bc6851bc (commit 466c8f305) |
|
|---|---|---|
atom.__file__ |
/tmp/pr294rv/control/ATOM/atom/__init__.py |
/tmp/pr294rv/merged/ATOM/atom/__init__.py |
gate_gpu.sh md5 |
4e9c066b… |
fba0641d… |
| result | 5163 passed, 155 skipped, 3 xfailed | 5182 passed, 155 skipped, 3 xfailed |
| verdict | GATE_CPU_RC=0 PASSED |
GATE_CPU_RC=0 PASSED |
The delta is +19 / 0 failed. It is exactly the new file's 19 ids: 5162→5163 and 5181→5182 are #290's +1 on both sides. No failures occurred, so no flaky-class re-run was needed.
ponytail-review
scripts/compass/gate_gpu.shL56-65: shrink: a 10-line comment restatesgate_cpu.sh:43-56nearly verbatim. Two lines pointing there, −8.scripts/compass/gate_gpu.shL248, L277-278, L281-282, L310-311: shrink: three reason args repeat$1's first line word for word.WHY=${2:-${1%%$'\n'*}}infail()and drop them, −3. The same line fixes F1.tests/compass/test_gate_gpu_verdict_through_pipe.py: nothing to cut. The 11 + 8 cases are the brief's named result, and_expected()has two callers.
net: -11 lines possible.
To reach APPROVE
- Make F1's one-line change to
fail(). - Show a mutation run on the new head where a forgotten reason no longer deletes the verdict. A re-run of X1 on the pass-count-only tree is enough.
The torch stub, a pass-count case and the two shrinks are recommended, not required.
fail() read its one-line reason as a bare $2 under `set -u`, so a call that left it out aborted the gate with no GATE_GPU_RC line at all. It now defaults to the finding's own first line. The three reasons that repeated that line word for word are dropped, and the comment above finish() now points at gate_cpu.sh's instead of restating it. The test gains a pass-count case, where the pass count is the only finding, and asserts that a verdict does not end on a colon: a reason cut at its first line would. The throwaway tree carries a stub torch, since the gate only reads its version, which takes the file from ~40 s to ~12 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Developer round 2: head
|
| mutant | result | failing id, and the assertion it fails on |
|---|---|---|
| N00, null: reword a comment | 22 passed | none |
| X1, the pass-count reason replaced by a comment | 1 failed | test_the_verdict_key_is_printed_exactly_once_and_last[pass-count]. It fails on the :-ending check, and the verdict is still printed: ...4731 passed, expected 4730. The difference decomposes as: |
X1b, X1 plus fail() reverted to a bare $2 (F1 reinstated) |
1 failed | the same id. It fails on the key-count check: the key is on 0 lines, and the run ended on unbound variable. |
| F1r, only the default reverted | 22 passed: inert | none. The three one-argument calls are never the first finding in any tree the test builds (see the :278 reply). |
M05, fail() records no reason |
6 failed | [new-failure], [pass-count], and the 4 new-failure pipe ids |
Named result
The committed test file was run against both trees:
- tip
6a83b56bc(gate_gpu.sh4e9c066b): 20 failed, 2 passed. All 20 new ids fail. - merged tree (
gate_gpu.sh478bbf53): 22 passed in 15.8 s, down from 40.9 s.
Gate 1 on the merged tree
- Tip:
6a83b56bc. - Merged tree:
git merge-tree --write-tree 6a83b56bc 08ba5dd35=ea5f9544a1d049cef036ca40befd7779e576a4ee, staged as the unreferenced commit4114a4762. - Where and how: node 18,
xiaobizh_n18_cpu;git archivewith stamps, at a private path; the tar md5 matched on both ends. Each tree ran its owngate_cpu.sh, md5adb7e63don both sides, undertimeout -k 10 3000, unpiped, one gate at a time.
control 6a83b56bc |
merged ea5f9544a |
|
|---|---|---|
atom.__file__ |
/tmp/xiaobizh280/gates/control/ATOM/atom/__init__.py |
/tmp/xiaobizh280/gates/merged/ATOM/atom/__init__.py |
| result | 5166 passed, 155 skipped, 3 xfailed | 5186 passed, 155 skipped, 3 xfailed |
| verdict | GATE_CPU_RC=0 PASSED |
GATE_CPU_RC=0 PASSED |
- Node-id delta (
--collect-only, 5303 → 5323): +20 / -0, exactly the new file's 20 ids. - One control run is discarded. The first control run was contaminated: I had copied the new test file into the control tree for the named-result run, and its 20 tip failures gave
GATE_CPU_RC=1. I removed the file, checked the tree against its tarball (only__pycache__differed), and re-ran. The figures above are from the clean run.
Lint: black rc=0 and ruff rc=0 on the test; bash -n is clean. No GPU was taken.
🤖 Generated with Claude Code
| {"tests/test_suite.py": _suite(EXPECT + 1)}, | ||
| "", | ||
| ), | ||
| "no-summary": (1, "no usable pytest summary", {"tests/test_suite.py": ""}, ""), |
There was a problem hiding this comment.
F1r: blocking (principle 8, and AI_DEV_RULES gate 4: "an inert pin on a required finding blocks APPROVE"). The gap is reachable by a real caller flag, and one line pins it.
The :278 reply says no throwaway tree makes a default-relying fail() the first finding. One does, cheaply. gate_gpu.sh passes "$@" through to pytest, and only -r* is refused. --no-summary drops pytest's short summary, so no FAILED lines are printed while the counts still read 5 failed. The first finding is then the FAILED-line check at :269, which relies on the default since 08ba5dd35.
Measured on node 18 (xiaobizh_n18_cpu), merged tree ea5f9544a, gate md5 478bbf53. The tree has exactly the expected passes plus the 5 known failures, and the gate ran with --no-summary:
- head: rc=1, the key on 1 line:
GATE_GPU_RC=1 NOT PASSED -- the summary says 5 failed but 0 FAILED line(s) were printed.; 1 more finding(s) on stderr - F1r (the default reverted only): rc=1, the key on 0 lines. stderr ends
gate_gpu.sh: line 243: $2: unbound variable.
So F1r is F1's own defect, live on a pass-through argument, and the suite gives 22 passed with it in place.
Fix: one entry here.
"no-fail-lines": (1, "0 FAILED line(s) were printed", {}, " --no-summary"),I measured it on a copy of this file:
- head:
[no-fail-lines]passed in 0.87 s; - F1r: 1 failed,
test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines], on the key-count assertion at:167.
It uses the default suite, so it costs under a second.
There was a problem hiding this comment.
Fixed in 3d2d5cb6a. I added your case, using the default suite. Its needle is "0 FAILED line(s) were printed; 2 more": on that tree the FAILED-line check is followed by the gone-ids finding and the pass-count finding, so it pins the count and the period strip too.
Measured on node 18, on the merged tree 68dbaa7a3 (gate md5 5e12a3af):
- head:
[no-fail-lines]passes. The verdict line isGATE_GPU_RC=1 NOT PASSED -- the summary says 5 failed but 0 FAILED line(s) were printed; 2 more finding(s) on stderr. - F1r (only
WHY=$2restored): 1 failed, 22 passed,test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines]. It fails on the key-count assertion, and the output ends onunbound variable. - null control: 23 passed.
My :278 reply claimed no tree reaches a default-relying first finding. That was wrong, since a pass-through flag reaches one. Thank you for finding it.
Review, cycle 2: REQUEST_CHANGES at head
|
| mutant | result | failing id and assertion |
|---|---|---|
N00 null, reword the comment at :251 |
22 passed (15.8 s) | none |
| X1, pass-count reason replaced by a comment | 1 failed, 21 passed | test_the_verdict_key_is_printed_exactly_once_and_last[pass-count], on the colon check: ...4731 passed, expected 4730. The difference decomposes as: |
X1b, X1 plus WHY=$2 |
1 failed, 21 passed | the same id, on the key-count check. The key is absent from the output. |
F1r, WHY=$2 only |
22 passed: inert | none. This is the blocking finding above. |
These match the dev round's counts and ids.
3. The stub torch
- It hides nothing the verdict needs. In
gate_gpu.shand_lib.shat the head,torchis imported only by the two version probes (:159-160), each|| echo UNKNOWN, and read only for the warn-only toolchain comparison. - No other gate or
_lib.shline imports torch. - The stub lives only in the throwaway tree. The real imports were observed in cycle 1's GPU run on node 18, which printed torch
2.10.0+rocm7.2.4.git3d3aa833and hip7.2.53211, both matching the baseline. - Cost: 22 tests in 15.2–15.8 s, down from 40.9 s.
4. ponytail-review (delta only)
Nothing to cut:
- the new
fail()comment is two lines; EXPECTremoves a repeated expression;- the colon assertion is one line;
- the stub is one dict entry.
Lean already. Ship.
A cosmetic nit, which blocks nothing: a first line that ends in . gives printed.; 1 more finding(s).
5. Gate: the tree that will land
- Tip
6a83b56bc, re-read after fetch.git merge-tree --write-tree 6a83b56bc 08ba5dd35=ea5f9544a1d049cef036ca40befd7779e576a4ee, which is the same as the dev round's. - Staging. Unreferenced commit
9f5f5ad57, thengit archivewith stamps (.compass-changedlists the PR's 2 files), thendocker exec -i … tar -xinto/tmp/pr294rv/inside the container. The md5 matched on both ends (2df70d3c…). atom.__file__=/tmp/pr294rv/merged/ATOM/atom/__init__.py, and the gate printedcommit: 9f5f5ad57 (stamp).- Result, with the tree's own
gate_cpu.sh(md5adb7e63d…), undertimeout -k 10 3000, unpiped, run once: 5186 passed, 155 skipped, 3 xfailed,GATE_CPU_RC=0 PASSED. - Delta: +20. It is measured against the dev round's control of 5166 at the same tip, which was run with the same
gate_cpu.shmd5. It equals the file's 20 ids, and nothing failed.
To reach APPROVE
- Add the
no-fail-linescase. - Show F1r failing that id by name at the new head.
Nothing else is required.
A caller's --no-summary reaches pytest, which then prints no FAILED lines while its counts still say 5 failed. The FAILED-line check becomes the first finding, and it has no explicit reason, so the verdict depends on fail()'s default. A no-fail-lines case now drives that run; without the default it ends on `set -u` with no GATE_GPU_RC line. The aggregate verdict drops a trailing period from the first reason before appending "; N more finding(s)", so it no longer reads "printed.;". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Developer round 3: head
|
| mutant | result | failing id and assertion |
|---|---|---|
| N00, null: reword a comment | 23 passed | none |
F1r, WHY=$2 only |
1 failed, 22 passed | test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines], on the key-count assertion; the output ends on unbound variable |
P1, ${WHY%.} reverted to $WHY |
1 failed, 22 passed | the same id, on the reason needle |
Named result
The committed test file was run against both trees:
- tip
d96027e52(gate_gpu.sh4e9c066b): 21 failed, 2 passed.[no-fail-lines]is among them. This ran on a separate copy of the control tree, and the control gate ran before it. - merged tree: 23 passed in 16.2 s.
Gate 1 on the merged tree
- Tip:
d96027e52, re-read after fetch. - Merged tree:
git merge-tree --write-tree d96027e52 3d2d5cb6a=68dbaa7a3c31e389b6726a7ce1a101627db3d9aa, staged as the unreferenced commit0c1c36976. - Where and how: node 18,
xiaobizh_n18_cpu;git archivewith stamps; the tar md5 matched on both ends. Each tree ran its owngate_cpu.sh, md5adb7e63don both sides, undertimeout -k 10 3000, unpiped, one gate at a time.
control d96027e52 |
merged 68dbaa7a3 |
|
|---|---|---|
atom.__file__ |
/tmp/xiaobizh280/gates/control/ATOM/atom/__init__.py |
/tmp/xiaobizh280/gates/merged/ATOM/atom/__init__.py |
| result | 5168 passed, 155 skipped, 3 xfailed | 5189 passed, 155 skipped, 3 xfailed |
| verdict | GATE_CPU_RC=0 PASSED |
GATE_CPU_RC=0 PASSED |
- Node-id delta (
--collect-only, 5305 → 5326): +21 / -0, exactly the file's 21 ids. - Size against the tip:
gate_gpu.sh+37 / -17, test +188 lines. - Lint: black rc=0 and ruff rc=0 on the test;
bash -nis clean. No GPU was taken.
🤖 Generated with Claude Code
Review, cycle 3: APPROVE at head
|
| mutant | result | failing id and assertion |
|---|---|---|
| N00 null, reword a comment | 23 passed (16.8 s) | none |
F1r, fail() back to a bare WHY=$2 |
1 failed, 22 passed | test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines], on the key-count check. The output has no GATE_GPU_RC= line. |
P1, period fix reverted (${WHY%.} → $WHY) |
1 failed, 22 passed | the same id, on the reason check. The line is ...0 FAILED line(s) were printed.; 2 more finding(s) on stderr |
The two mutants fail on different assertions of the same id, so each assertion holds its own defect.
"No bare $2": checked independently
- In
gate_gpu.shat the head,$2appears only infail(), where it is defaulted, and infinish()'s non-zeroprintf(:63). - All 11 non-zero
finishcalls pass a literal reason::115,:121,:142,:147,:148,:154,:156,:190,:218,:231. - The aggregate call at
:332passes"$WHY", which is set wheneverVERDICT>0. finish 0(:346) takes the branch that never reads$2._lib.sh's$2uses arelocal …=$2in functions whose callers pass the argument; none is new in this PR.
The claim holds.
ponytail-review (delta only)
The delta is one parameter expansion and one dict entry. Lean already. Ship.
Gate: the tree that will land
- Tip
d96027e52, re-read after fetch.git merge-tree --write-tree d96027e52 3d2d5cb6a=68dbaa7a3c31e389b6726a7ce1a101627db3d9aa, the same as the dev round's. - Staging. Unreferenced commit
e8c36c52a, thengit archivewith stamps (.compass-changedlists the PR's 2 files), thendocker exec -i … tar -xinto/tmp/pr294rv/inside the container. The md5 matched on both ends (e6441939…). atom.__file__=/tmp/pr294rv/merged/ATOM/atom/__init__.py, and the gate printedcommit: e8c36c52a (stamp).- Result, with the tree's own
gate_cpu.sh(md5adb7e63d…), undertimeout -k 10 3000, unpiped, run once: 5189 passed, 155 skipped, 3 xfailed,GATE_CPU_RC=0 PASSED. - This equals the dev round's figure, and at +21 it is exactly the file's 21 ids against their control at the same tip.
Carried forward, not re-measured this cycle
- Cycle 1's real GPU run of this PR's
gate_gpu.shon node 18 endedGATE_GPU_RC=0 PASSEDexactly once. - Since then, only the
fail()reason plumbing and the${WHY%.}trim have changed. Neither touches the zero-verdict path.
Closes #280
gate_gpu.sh's last line now carries its reason. A run that exits 1 or 93 still reads as not passed, with the reason, through| tail -6,2>&1 | tail -6and2>/dev/null | tail -6. This follows #276's shape forgate_cpu.sh.Head:
579f7c285, a single commit onfeature/atomcompass_newat99fc506f9. It was restacked once before the PR opened, because the tip moved from37824643bwhile I worked.File set, as briefed:
scripts/compass/gate_gpu.sh(+48 / -20) andtests/compass/test_gate_gpu_verdict_through_pipe.py(+180 lines, 56 AST statements; the estimate was 30-60, and #276's file measures 67 by the same count).What changed
finish <rc> <reason>printsGATE_GPU_RC=0 PASSEDorGATE_GPU_RC=<n> NOT PASSED -- <reason>, asgate_cpu.shdoes since compass(gates): gate_cpu.sh verdict line carries its reason through a pipe #276. Each of the 11 non-zero call sites passes its own reason.finish 1names whatfail()recorded.fail()now takes a second argument, the finding's one-line form. It keeps the first finding's one-line form and counts the rest. The full multi-line text still goes to stderr, unchanged. The checks already run with the node-id comparisons ahead of the counts, so a new failure is named by its node-id. Measured output:GATE_GPU_RC=1 NOT PASSED -- 1 failure(s) the baseline does not name, first tests/test_suite.py::test_new; 2 more finding(s) on stderrverdict: the last line of stdout; if this run is piped, $? is the pipe's, is compass(gates): gate_cpu.sh verdict line carries its reason through a pipe #276's text verbatim. It does not spell the key.Named result: tip vs head, same run, same throwaway tree
This is
gate_gpu.shat the tip (4e9c066b) vs the head (fba0641d), on thenew-failuretree: one test that the baseline does not name fails. The unpiped rc is 1, and every piped rc is 0.The unpiped run,
| tail -6and2>&1 | tail -6end on the same line. The rc-93 case (known-disagrees) ends onGATE_GPU_RC=93 NOT PASSED -- BASE_FAILED disagrees with the known-failures list; nothing was runin all four; at the tip it ended on a bareGATE_GPU_RC=93.The test fails by name at the tip and passes at the head. I dropped the committed test file into the tip tree:
test_gate_gpu_verdict_through_pipe.py+test_gate_gpu_aiter_version.py99fc506f9579f7c285All 19 tip failures land on the verdict-content assertion. None is on the rc or the key count, which both already held at the tip.
Node ids (
tests/compass/test_gate_gpu_verdict_through_pipe.py::):test_the_verdict_key_is_printed_exactly_once_and_last[{passed,new-failure,no-summary,interrupted,known-missing,known-disagrees,compass-count,r-flag,preflight,atom,other-checkout}]: 11 ids. This test runs2>&1and asserts four things:GATE_GPU_RC=<n> NOT PASSED --and carries the case's reason;GATE_GPU_RC=0 PASSED.test_a_piped_run_still_reads_as_not_passed_and_says_why[{new-failure,known-disagrees}-{,| tail -6,2>&1 | tail -6,2>/dev/null | tail -6}]: 8 ids. The empty pipe is the unpiped control and asserts the gate's own rc; the three piped forms assert the pipe's 0.Verdict paths: tested vs covered by reading only
The test drives the real script over a throwaway tree with no GPU:
preflight.shthat prints six lines and exits 0 or 1;atom/__init__.py;BASE_*constants, parsed rather than hard-coded;passed): 4730 generated passes + 5 named known failuresfail()new-failure)no-summary): empty suiteinterrupted): KeyboardInterrupt, rc=2known-missing)BASE_FAILEDvs list countknown-disagrees)compass-count)-rargumentr-flag)preflight): the stub exits 1compass_require_treeatom): the stub atom raises on import, so 91compass_tree_rootother-checkout): the gate is run from inside a second checkout, so 99cd "$ROOT"compass_tree_roothas justcd'd into the same path.No verdict path needs a GPU. What the stubs replace, and what therefore stays tested by reading only:
preflight.sh, which runsrocminfo/rocm-smiand can hang in D state on a wedged node, so it must not run in a CPU test;The torch and HIP readings in the test are real imports, 1.6 s on
xiaobizh_n18_cpu. AITER readsUNKNOWN, so each run prints the toolchain WARNING, which is warn-only by design. The real script's output on a GPU node has not been observed. Per the brief, the GPU tier was not run and no GPU was taken.#271's test is unaffected.
tests/compass/test_gate_gpu_aiter_version.pystill passes at the head (2 passed). ItsAITER_DIR=...fimarker span is untouched.Mutation battery (line-count preserving, null control)
These ran on
gate_gpu.shatfba0641d, the md5 committed at the head, against the new test file plus the #271 file (21 tests). Each mutant checked that the edit applied, that the line count was unchanged, and thatbash -npassed.About the test file used: the battery ran against the test file before
blackand one message edit. After that, the only other change wasre.M→re.MULTILINE, so no assertion changed.GATE_GPU_RC=%sexactly_once+ all 8 pipe idsGATE_GPU_RC=0exactly_once[passed]GATE_GPU_RC=exactly_once[*-],[*-| tail -6],[*-2>/dev/null | tail -6]for both cases.2>&1keeps it, as it should.fail()never records a reasonexactly_once[new-failure]+ 4new-failurepipe idsfail()keeps the last reason, not the firstfinish 1 ""exactly_once[known-missing]exactly_once[known-disagrees]+ 4known-disagreespipe ids-rreason emptyexactly_once[r-flag]exactly_once[other-checkout]exactly_once[preflight]exactly_once[atom]exactly_once[compass-count]exactly_once[no-summary]exactly_once[interrupted]The #271 tests passed under every mutant.
Gates
Gate 1: ATOM's CPU tier, unmodified, each tree run with its own
scripts/compass/gate_cpu.sh. The runs used:xiaobizh_n18_cpu;git archivetrees staged at a private path, with.compass-commitand.compass-changedstamps;timeout -k 10 3000, unpiped, one gate at a time.99fc506f9579f7c285(= merged tree)atom.__file__/tmp/xiaobizh280/gates/control/ATOM/atom/__init__.py/tmp/xiaobizh280/gates/branch/ATOM/atom/__init__.pygate_cpu.shmd5adb7e63ddf3d494913bc7e12f60affbcadb7e63ddf3d494913bc7e12f60affbc(the same instrument)gate_gpu.shmd54e9c066b996be5407c08e430f05d7139fba0641dcda5d1616062cded70eead33GATE_CPU_RC=0 PASSEDGATE_CPU_RC=0 PASSEDgpu:gpu_gate_triggers.txtcoversscripts/ortests/.--collect-onlyover the gate's own ignore set): control 5299, branch 5318. +19 / -0, exactly the 19 ids above.git merge-tree --write-tree 99fc506f9 579f7c285=67e89091a6a43d6c7862d8bc46a33ef8c1165042, which equals the head's own tree. The branch run is therefore the merged-tree run.37824643b: control 5155 / 155 / 3, branch 5174 / 155 / 3, bothGATE_CPU_RC=0 PASSED, +19 / -0, and the test failed 19 at the tip and passed at the head.scripts/compass/gate_gpu.sh, +48 / -20. Test is one new file, +180. No ATOM test was edited.black --checkrc=0,ruff checkrc=0 on the new test. shellcheck is not installed in either container;bash -nis clean.Gate 2: the 19 CPU-only tests above.
Gate 3: the named result above.
Gate 4: the coordinator dispatches an independent reviewer.
Notes for the reviewer / successor
passedcase generatesBASE_PASSED - BASE_COMPASS_TESTStests, 4730 today. It follows the constants, so rebaselining the GPU gate needs no edit here.new-failurereason names only the first finding. Here that is the new node-id; a changed failed count and pass count follow as "2 more finding(s) on stderr". If a run has collection errors, the errors finding comes first, because that check runs first.🤖 Generated with Claude Code