Skip to content

compass(gates): gate_gpu.sh verdict line carries its reason through a pipe - #294

Merged
jgong5 merged 3 commits into
feature/atomcompass_newfrom
compass/issue-280-gate-gpu-verdict
Sep 23, 2026
Merged

jgong5 merged 3 commits into
feature/atomcompass_newfrom
compass/issue-280-gate-gpu-verdict

Conversation

@jgong5

@jgong5 jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner

Closes #280

gate_gpu.sh's last line now carries its reason. A run that exits 1 or 93 still reads as not passed, with the reason, through | tail -6, 2>&1 | tail -6 and 2>/dev/null | tail -6. This follows #276's shape for gate_cpu.sh.

Head: 579f7c285, a single commit on feature/atomcompass_new at 99fc506f9. It was restacked once before the PR opened, because the tip moved from 37824643b while I worked.
File set, as briefed: scripts/compass/gate_gpu.sh (+48 / -20) and tests/compass/test_gate_gpu_verdict_through_pipe.py (+180 lines, 56 AST statements; the estimate was 30-60, and #276's file measures 67 by the same count).

What changed

  • finish <rc> <reason> prints GATE_GPU_RC=0 PASSED or GATE_GPU_RC=<n> NOT PASSED -- <reason>, as gate_cpu.sh does since compass(gates): gate_cpu.sh verdict line carries its reason through a pipe #276. Each of the 11 non-zero call sites passes its own reason.
  • The aggregate finish 1 names what fail() recorded. fail() now takes a second argument, the finding's one-line form. It keeps the first finding's one-line form and counts the rest. The full multi-line text still goes to stderr, unchanged. The checks already run with the node-id comparisons ahead of the counts, so a new failure is named by its node-id. Measured output:
    GATE_GPU_RC=1 NOT PASSED -- 1 failure(s) the baseline does not name, first tests/test_suite.py::test_new; 2 more finding(s) on stderr
  • A header line, verdict: the last line of stdout; if this run is piped, $? is the pipe's, is compass(gates): gate_cpu.sh verdict line carries its reason through a pipe #276's text verbatim. It does not spell the key.
  • The after-run pre-flight block keeps its place, just before the verdict. With the reason on the verdict line, the block no longer hides anything, so moving it would have been a second behaviour change for no gain.

Named result: tip vs head, same run, same throwaway tree

This is gate_gpu.sh at the tip (4e9c066b) vs the head (fba0641d), on the new-failure tree: one test that the baseline does not name fails. The unpiped rc is 1, and every piped rc is 0.

tip,  2>/dev/null | tail -6          head, 2>/dev/null | tail -6
pre-flight 2                         pre-flight 2
...                                  ...
pre-flight 6                         pre-flight 6
GATE_GPU_RC=1                        GATE_GPU_RC=1 NOT PASSED -- 1 failure(s) the baseline does not name, first tests/test_suite.py::test_new; 2 more finding(s) on stderr

The unpiped run, | tail -6 and 2>&1 | tail -6 end on the same line. The rc-93 case (known-disagrees) ends on GATE_GPU_RC=93 NOT PASSED -- BASE_FAILED disagrees with the known-failures list; nothing was run in all four; at the tip it ended on a bare GATE_GPU_RC=93.

The test fails by name at the tip and passes at the head. I dropped the committed test file into the tip tree:

tree test_gate_gpu_verdict_through_pipe.py + test_gate_gpu_aiter_version.py
tip 99fc506f9 19 failed, 2 passed, rc=1
head 579f7c285 21 passed, rc=0

All 19 tip failures land on the verdict-content assertion. None is on the rc or the key count, which both already held at the tip.

Node ids (tests/compass/test_gate_gpu_verdict_through_pipe.py::):

  • test_the_verdict_key_is_printed_exactly_once_and_last[{passed,new-failure,no-summary,interrupted,known-missing,known-disagrees,compass-count,r-flag,preflight,atom,other-checkout}]: 11 ids. This test runs 2>&1 and asserts four things:
    • the rc;
    • the key appears on exactly one line, and that line is the last;
    • the line starts with GATE_GPU_RC=<n> NOT PASSED -- and carries the case's reason;
    • the PASSED case is exactly GATE_GPU_RC=0 PASSED.
  • test_a_piped_run_still_reads_as_not_passed_and_says_why[{new-failure,known-disagrees}-{,| tail -6,2>&1 | tail -6,2>/dev/null | tail -6}]: 8 ids. The empty pipe is the unpiped control and asserts the gate's own rc; the three piped forms assert the pipe's 0.

Verdict paths: tested vs covered by reading only

The test drives the real script over a throwaway tree with no GPU:

  • a stub preflight.sh that prints six lines and exits 0 or 1;
  • a stub atom/__init__.py;
  • a generated suite whose pass and fail counts come from the script's own BASE_* constants, parsed rather than hard-coded;
  • the known-failures file rewritten to name the suite's own failing ids.
exit path how
0 PASSED tested (passed): 4730 generated passes + 5 named known failures
1 aggregate fail() tested (new-failure)
1 no usable summary tested (no-summary): empty suite
1 pytest rc not 0/1 tested (interrupted): KeyboardInterrupt, rc=2
93 known-failures list missing tested (known-missing)
93 BASE_FAILED vs list count tested (known-disagrees)
93 tests/compass pass count has no source tested (compass-count)
95 -r argument tested (r-flag)
91 pre-flight before the run tested (preflight): the stub exits 1
91/92 compass_require_tree tested (atom): the stub atom raises on import, so 91
99/90 compass_tree_root tested (other-checkout): the gate is run from inside a second checkout, so 99
90 cd "$ROOT" reading only. Unreachable short of a race: compass_tree_root has just cd'd into the same path.

No verdict path needs a GPU. What the stubs replace, and what therefore stays tested by reading only:

  • the real preflight.sh, which runs rocminfo/rocm-smi and can hang in D state on a wedged node, so it must not run in a CPU test;
  • the real toolchain comparison against the baseline (torch, HIP, AITER);
  • the real 4779-test superset.

The torch and HIP readings in the test are real imports, 1.6 s on xiaobizh_n18_cpu. AITER reads UNKNOWN, so each run prints the toolchain WARNING, which is warn-only by design. The real script's output on a GPU node has not been observed. Per the brief, the GPU tier was not run and no GPU was taken.

#271's test is unaffected. tests/compass/test_gate_gpu_aiter_version.py still passes at the head (2 passed). Its AITER_DIR= ... fi marker span is untouched.

Mutation battery (line-count preserving, null control)

These ran on gate_gpu.sh at fba0641d, the md5 committed at the head, against the new test file plus the #271 file (21 tests). Each mutant checked that the edit applied, that the line count was unchanged, and that bash -n passed.

About the test file used: the battery ran against the test file before black and one message edit. After that, the only other change was re.M → re.MULTILINE, so no assertion changed.

mutation result failing ids
M00 null: reword a comment 21 passed —
M01 non-zero verdict prints bare GATE_GPU_RC=%s 18 failed all 10 non-zero exactly_once + all 8 pipe ids
M02 zero verdict prints bare GATE_GPU_RC=0 1 failed exactly_once[passed]
M03 header spells GATE_GPU_RC= 11 failed all 11 exactly_once
M04 non-zero verdict to stderr 6 failed pipe [*-], [*-| tail -6], [*-2>/dev/null | tail -6] for both cases. 2>&1 keeps it, as it should.
M05 fail() never records a reason 5 failed exactly_once[new-failure] + 4 new-failure pipe ids
M06 fail() keeps the last reason, not the first 5 failed same 5
M07 "N more finding(s)" suffix dropped 5 failed same 5
M08 aggregate finish 1 "" 5 failed same 5
M09 new-failure reason drops the node-id 5 failed same 5
M10 known-missing reason empty 1 failed exactly_once[known-missing]
M11 known-disagrees takes known-missing's reason 5 failed exactly_once[known-disagrees] + 4 known-disagrees pipe ids
M12 -r reason empty 1 failed exactly_once[r-flag]
M13 tree-root reason empty 1 failed exactly_once[other-checkout]
M14 pre-flight reason empty 1 failed exactly_once[preflight]
M15 require-tree reason empty 1 failed exactly_once[atom]
M16 compass-count reason empty 1 failed exactly_once[compass-count]
M17 no-summary takes the pytest-rc reason 1 failed exactly_once[no-summary]
M18 pytest-rc takes the no-summary reason 1 failed exactly_once[interrupted]

The #271 tests passed under every mutant.

Gates

Gate 1: ATOM's CPU tier, unmodified, each tree run with its own scripts/compass/gate_cpu.sh. The runs used:

  • node 18, xiaobizh_n18_cpu;
  • git archive trees staged at a private path, with .compass-commit and .compass-changed stamps;
  • timeout -k 10 3000, unpiped, one gate at a time.
control 99fc506f9 branch 579f7c285 (= merged tree)
atom.__file__ /tmp/xiaobizh280/gates/control/ATOM/atom/__init__.py /tmp/xiaobizh280/gates/branch/ATOM/atom/__init__.py
gate_cpu.sh md5 adb7e63ddf3d494913bc7e12f60affbc adb7e63ddf3d494913bc7e12f60affbc (the same instrument)
gate_gpu.sh md5 4e9c066b996be5407c08e430f05d7139 fba0641dcda5d1616062cded70eead33
result 5162 passed, 155 skipped, 3 xfailed 5181 passed, 155 skipped, 3 xfailed
verdict GATE_CPU_RC=0 PASSED GATE_CPU_RC=0 PASSED
gpu: not required not required. No trigger in gpu_gate_triggers.txt covers scripts/ or tests/.
  • Node-id delta (--collect-only over the gate's own ignore set): control 5299, branch 5318. +19 / -0, exactly the 19 ids above.
  • Merged tree: git merge-tree --write-tree 99fc506f9 579f7c285 = 67e89091a6a43d6c7862d8bc46a33ef8c1165042, which equals the head's own tree. The branch run is therefore the merged-tree run.
  • Before the restack, at tip 37824643b: control 5155 / 155 / 3, branch 5174 / 155 / 3, both GATE_CPU_RC=0 PASSED, +19 / -0, and the test failed 19 at the tip and passed at the head.
  • Production vs test lines: production is scripts/compass/gate_gpu.sh, +48 / -20. Test is one new file, +180. No ATOM test was edited.
  • Lint: black --check rc=0, ruff check rc=0 on the new test. shellcheck is not installed in either container; bash -n is clean.

Gate 2: the 19 CPU-only tests above.
Gate 3: the named result above.
Gate 4: the coordinator dispatches an independent reviewer.

Notes for the reviewer / successor

  • Cost: the new file adds about 40 s to the CPU tier, measured at 40.9 s for the file alone on node 18. Most of it is two real torch imports and a pytest start-up per gate run, across 19 runs.
  • Only one gate run is large. The passed case generates BASE_PASSED - BASE_COMPASS_TESTS tests, 4730 today. It follows the constants, so rebaselining the GPU gate needs no edit here.
  • The new-failure reason names only the first finding. Here that is the new node-id; a changed failed count and pass count follow as "2 more finding(s) on stderr". If a run has collection errors, the errors finding comes first, because that check runs first.

🤖 Generated with Claude Code

… pipe

`gate_gpu.sh 2>/dev/null | tail -6` kept the after-run pre-flight readout
and a bare `GATE_GPU_RC=1`: every reason went to stderr, and the pre-flight
block is printed to stdout just before the verdict. A pipeline's status is
tail's, so nothing said the run had not passed or why.

The verdict line now reads `GATE_GPU_RC=0 PASSED` or
`GATE_GPU_RC=<n> NOT PASSED -- <reason>`, the shape gate_cpu.sh took in
#276. Every non-zero finish passes its own reason. fail() keeps the
one-line form of the first finding and counts the rest, so the aggregate
exit names it, for example "1 failure(s) the baseline does not name,
first <node-id>; 2 more finding(s) on stderr". A header line names the
last line of stdout as the verdict without spelling the key.

tests/compass/test_gate_gpu_verdict_through_pipe.py runs the real script
over throwaway trees with a stub pre-flight and a generated suite sized
from the script's own baseline constants. It reaches all twelve exits but
the unreachable `cd` one, including a PASSED run, with no GPU.

Closes #280

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread scripts/compass/gate_gpu.sh Outdated
VERDICT=0
fail() { printf '%s\n' "$1" >&2; VERDICT=1; }
WHY=
fail() { printf '%s\n' "$1" >&2; VERDICT=$((VERDICT + 1)); [ -n "$WHY" ] || WHY=$2; }

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

F1: blocking (principles 6 and 8). A fail() call that forgets its reason deletes the verdict line, and no test sees it.

fail() now needs a second argument. The script runs under set -u, so a call without one does not just lose the reason: it aborts the script on this line, and no GATE_GPU_RC= line is printed at all. That is the one outcome the header above finish() says no path may have. It is also new in this PR: before it, fail() took one argument.

The suite cannot see it. $2 is expanded only while WHY is still empty. The only case that reaches fail(), new-failure, always records the new-ids finding first, so the other six fail() reasons are never expanded by any test.

Measured on node 18 (xiaobizh_n18_cpu), on the merged tree 7bc6851bc, gate md5 fba0641d:

  • Mutation X1. I replaced the pass-count reason (:332) with a comment, keeping the line count, with bash -n clean. Result: test_gate_gpu_verdict_through_pipe.py + test_gate_gpu_aiter_version.py gave 21 passed, so the mutant is not caught.
  • The same mutant on a tree whose only finding is the pass count. The tree had 4731 generated passes against 4730 expected, plus the 5 named known failures. Result: rc=1, and GATE_GPU_RC= appears on 0 lines. stdout ends on GATE_GPU passed=4731 failed=5 errors=0 pytest_rc=1 (...), and stderr ends on gate_gpu.sh: line 248: $2: unbound variable.
  • The unmutated head on the same tree ends GATE_GPU_RC=1 NOT PASSED -- 4731 passed, expected 4730. The code is correct today; what is missing is the guard.

A pass-count-only mismatch is the likeliest real failure of this gate. The BASE_COMPASS_TESTS comment calls a task that adds tests outside tests/compass/ "an explicit mismatch, which is the point".

Fix: one line, measured. Default the reason to the finding's own first line:

fail() { printf '%s\n' "$1" >&2; VERDICT=$((VERDICT + 1)); [ -n "$WHY" ] || WHY=${2:-${1%%$'\n'*}}; }

This is not a guessed fallback in principle 6's sense, because the text is the finding's own. Results on the same node:

  • The X1 tree now ends GATE_GPU_RC=1 NOT PASSED -- 4731 passed, expected 4730. The difference decomposes as:, and the key is on 1 line.
  • The two gate_gpu test files still give 21 passed.
  • It also lets three duplicated reason arguments go; see the :278 comment.

Recommended, not required: add a pass-count case, the passed suite plus one test, so that a sole finding other than new ids reaches the verdict under test. With the torch stub suggested on the test file, it costs about 3 s.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 08ba5dd35: fail() now reads WHY=${2:-${1%%$'\n'*}}, and its comment says a finding given no one-line form is named by its own first line.

Added a pass-count case, the passed suite plus one test, where the pass count is the only finding. The exactly-once test also asserts the verdict does not end on :, because a reason cut at its first line can end on a colon that promises the rest. Measured on node 18, on the merged tree ea5f9544a:

  • X1 (this pass-count reason replaced by a comment): 1 failed, test_the_verdict_key_is_printed_exactly_once_and_last[pass-count], on the colon assertion. The verdict is still printed: GATE_GPU_RC=1 NOT PASSED -- 4731 passed, expected 4730. The difference decomposes as:.
  • X1b (X1 with fail() reverted to a bare $2): 1 failed, the same id, on the key-count assertion. The key is on 0 lines, and the run ended on unbound variable. So the new case sees your F1 by name.
  • Null control (comment reword): 22 passed.

Comment thread scripts/compass/gate_gpu.sh Outdated
way. (pytest -r is store-last-wins: a caller flag, a truncated log or a
plugin can all empty it.)' "$FAILED" "$OBS_N")"
plugin can all empty it.)' "$FAILED" "$OBS_N")" \
"the summary says $FAILED failed but $OBS_N FAILED line(s) were printed"

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ponytail shrink: (principle 3). Non-blocking, but it goes with F1 (:248).

Three of the new reason arguments repeat, word for word, the first line of the finding they sit under:

  • :278, the FAILED-line count;
  • :282, the ERROR-line count;
  • :311, the failed count.

With fail() defaulting to ${1%%$'\n'*}, the arguments and their trailing \ can go: −3 lines. Keep the explicit reason where the first line does not stand alone:

  • :256, which would end "...and an";
  • the new-ids and gone-ids reasons, which add the node-id;
  • :332, which would trail "The difference decomposes as:".

I measured a version that dropped four of these, including :332. The two gate_gpu test files gave 21 passed, bash -n was clean, and the file went from 354 to 350 lines.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 08ba5dd35. The FAILED-line count, ERROR-line count and failed-count reasons are dropped and fall to the default. Kept the explicit reasons where the first line does not stand alone: errors, new ids, gone ids, and the pass count.

One residual. Reverting the default to a bare $2 on its own (mutant F1r) is inert: 22 passed. The three one-argument calls are never the first finding under test. By the order of the checks, the failed count can come first only if the FAILED lines disagree with the summary, and no throwaway tree reaches that. The pass-count case covers the guard through X1b instead.

Comment thread scripts/compass/gate_gpu.sh Outdated
# while the success verdict went to stdout -- so `gate_gpu.sh 2>/dev/null | grep
# GATE_GPU_RC` was silent on failure and indistinguishable from "never ran".
#
# The verdict line also carries its reason, because it is the only line every

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ponytail shrink: (principle 3). Non-blocking.

:56-65 restates gate_cpu.sh's comment above its own finish() (its lines 43–56), including the docker exec FIFO argument, nearly word for word. The two finish() bodies and the header printf are byte-identical apart from the key; I diffed them.

Two lines say the same and cannot drift from the original:

# The verdict line carries its reason, for the reasons given above finish() in
# gate_cpu.sh: it is the only line every pipe keeps.

That is −8 lines.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 08ba5dd35: replaced with the two-line pointer to gate_cpu.sh's comment above its finish().

files = {
"scripts/compass/preflight.sh": _preflight(0),
"scripts/compass/gpu_gate_known_failures.txt": "\n".join(KNOWN) + "\n",
"atom/__init__.py": "",

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cost: 70% of the 40 s is two real import torch calls per run (principles 3 and 8). Non-blocking, but recommended.

These were measured with --durations=0 on node 18 (xiaobizh_n18_cpu). The file took 38.6 s:

  • 9 cases reach the script's torch/HIP probes, and each costs 3.9–4.1 s. That includes new-failure, whose generated suite has one passing test.
  • passed, with 4730 generated tests, costs 6.3 s, so the suite's size is about 2 s of the total. Shrinking the generated suite would buy almost nothing.
  • The other 10 cases exit before the probes and cost 0.01 s each.
  • A bare python -c "import torch" costs 1.58 s, and the script runs two per case.

The verdict does not depend on the torch version. A mismatch only warns, and it still warns with a stub. Add one entry to this dict:

        "torch/__init__.py": "__version__ = 1\nclass version:\n    hip = 1\n",

compass_env puts the throwaway root first on PYTHONPATH, so the probes import the stub. Measured on a copy of this file: 19 passed in 11.7 s, against 38.6 s. The same stub would also shave the real GPU tier: compass_compass_pass_count runs tests/compass/ once and the superset runs it again, so this file runs twice per gate_gpu.sh.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 08ba5dd35: the throwaway tree now carries your stub torch/__init__.py. On node 18 the file, now 20 cases, plus #271's 2 tests takes 15.8 s, down from 40.9 s.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 1: REQUEST_CHANGES at head 579f7c285649d8a95e03ee59067873613c0d2248

Blocking: one finding, F1, with a one-line fix (inline on gate_gpu.sh:248, principles 6 and 8).

  • fail() now takes a required second argument under set -u. A future call that forgets it aborts the script with no verdict line at all.
  • The suite cannot see it: mutation X1 gives 21 passed.
  • On a tree whose only finding is the pass count, which is the likeliest real failure of this gate, the X1 mutant ends on line 248: $2: unbound variable with GATE_GPU_RC= on 0 lines.
  • The code at the head is correct. What is missing is the guard, and this PR introduces the hazard.

Non-blocking: three inline findings: two ponytail shrinks (:57, :278) and the test's cost (test_...:119).

Everything else checks out. The fix matches #276, the test drives the real script, the named result reproduces, and the real GPU run ends on the new line exactly once.

I read the eight design principles and AI_DEV_RULES.md (#137 revision) first. All measurements were taken on node 18.

1. Consistency with #276 (principle 3)

  • I diffed the verdict logic. finish() and the header printf are byte-identical apart from the key (GATE_CPU_RC vs GATE_GPU_RC), so there is no needless difference.
  • gate_gpu.sh adds only the aggregate: fail() records the first reason and counts the rest. gate_cpu.sh has no counterpart, because its verdict is pytest's rc.
  • Ruling: no shared _lib.sh helper, and no follow-up issue.
    • A compass_finish KEY RC REASON helper plus two one-line wrappers saves about 4 lines, and couples two scripts that are each pinned by their own tests.
    • The only real duplication is the 10-line comment. It is flagged as a ponytail shrink (:57), and pointing at gate_cpu.sh removes it.

2. Real script or copy? (principle 8)

  • It is the real script. _tree() copies the tree's own scripts/compass/, including _lib.sh, and replaces only preflight.sh, atom/__init__.py, the suite and the known-failures list.
  • Every mutation below was applied to the tree's gate_gpu.sh, and the test reddened. A copy would not.
  • The real GPU run, measured. I ran the merged tree's own gate_gpu.sh (md5 fba0641d…) in xiaobizh_n18 with HIP_VISIBLE_DEVICES=5.
    • Card 5 was checked free first: rocm-smi --showmemuse --showmeminfo vram read 298 MB, the idle floor, on cards 2–7. Cards 0 and 1 were held at 171 GB by other PIDs and were not touched.
    • It ran unpiped, with stdout and stderr to separate files, under timeout -k 10 3600.
    • atom.__file__ = /tmp/pr294rv/merged/ATOM/atom/__init__.py, and the gate printed commit: 466c8f305 (stamp).
    • Toolchain: torch 2.10.0+rocm7.2.4.git3d3aa833, hip 7.2.53211, aiter v0.1.21.dev0-49-gf4e7c7509, all equal to the baseline, so there was no WARNING.
    • pytest: 5 failed, 5962 passed, 105 skipped, 3 xfailed in 244 s. The 5 failures are exactly the 5 known AITER ids, and the expected count was 5962 = 4779 + (1232 − 49).
    • Last stdout line: GATE_GPU_RC=0 PASSED. The key is on 1 stdout line and 0 stderr lines, and stderr is empty. The 19 new tests ran nested inside the GPU superset and passed.
  • Not observed on a GPU: a real non-zero verdict. It would need an induced failure, and the stubbed paths cover the text.
  • One residual, no action needed. pytest's output is teed to stdout. If a test whose failure text contains the key fails inside the superset, for example this file or test_gate_cpu_verdict_through_pipe.py, its traceback puts GATE_GPU_RC= on stdout more than once. The verdict is still last, and the run is already NOT PASSED.

3. The 40 s cost (principle 3)

Justified in what it covers, but most of it is avoidable, and not by shrinking the suite. --durations=0 gives 38.6 s:

  • The 9 cases that reach the torch/HIP probes take 3.9–4.1 s each. That includes a 1-test suite.
  • passed, with 4730 generated tests, takes 6.3 s.
  • The 10 early exits take 0.01 s each.
  • import torch alone takes 1.58 s, and the script does it twice per case.

A stub torch/__init__.py in the throwaway tree gives 19 passed in 11.7 s (inline test_...:119). The generated suite can stay sized from BASE_*.

4. Mutations

These ran on the merged tree's gate_gpu.sh, each line-count preserving and bash -n clean. Each was run against test_gate_gpu_verdict_through_pipe.py + test_gate_gpu_aiter_version.py (21 tests). Node ids are under tests/compass/test_gate_gpu_verdict_through_pipe.py::.

mutant result failing ids
N00 null, reword the comment at :258 21 passed none
R01 = M01, non-zero verdict bare GATE_GPU_RC=%s 18 failed all 10 non-zero test_the_verdict_key_is_printed_exactly_once_and_last[…] + all 8 test_a_piped_run_still_reads_as_not_passed_and_says_why[{new-failure,known-disagrees}-{,| tail -6,2>&1 | tail -6,2>/dev/null | tail -6}]
R03 = M03, header spells GATE_GPU_RC= 11 failed all 11 …exactly_once_and_last[…], [passed] included
R05 = M05, fail() records no reason 5 failed …exactly_once_and_last[new-failure] + 4 …says_why[new-failure-*]
R06 = M06, fail() keeps the last reason 5 failed the same 5
R12 = M12, -r reason empty 1 failed …exactly_once_and_last[r-flag]
X1 (mine), pass-count fail() forgets its reason 21 passed: inert none. This is F1.
X2 (mine), cd "$ROOT" exit forgets its reason 21 passed none. This is expected, since the PR body lists exit 90 as covered by reading only.
X3 (mine), known-missing finish 93 forgets its reason 1 failed …exactly_once_and_last[known-missing]

5. #271

test_gate_gpu_aiter_version.py gives 2 passed at the merged tree, and under every mutant. Its AITER_DIR= … fi span has no line in the diff.

Gate: the tree that will land

  • git merge-tree --write-tree bd42a82e0 579f7c285 = 7bc6851bc77549646899301ada321627de9b2376, rc 0.
  • Staging. git archive of an unreferenced commit object 466c8f305 (commit-tree of that tree, parents bd42a82e0 and 579f7c285), then docker exec -i … tar -x into /tmp/pr294rv/ inside xiaobizh_n18_cpu.
    • Stamps were written: .compass-changed lists this PR's 2 files.
    • The tar md5 matched on both ends: 7db63115… merged, ef0836b8… control.
    • The shared mount was not touched.
  • Runs. Each tree ran its own gate_cpu.sh, md5 adb7e63d… on both sides. Each run was bounded by timeout -k 10 3000, unpiped, one gate at a time.
control bd42a82e0 merged 7bc6851bc (commit 466c8f305)
atom.__file__ /tmp/pr294rv/control/ATOM/atom/__init__.py /tmp/pr294rv/merged/ATOM/atom/__init__.py
gate_gpu.sh md5 4e9c066b… fba0641d…
result 5163 passed, 155 skipped, 3 xfailed 5182 passed, 155 skipped, 3 xfailed
verdict GATE_CPU_RC=0 PASSED GATE_CPU_RC=0 PASSED

The delta is +19 / 0 failed. It is exactly the new file's 19 ids: 5162→5163 and 5181→5182 are #290's +1 on both sides. No failures occurred, so no flaky-class re-run was needed.

ponytail-review

  • scripts/compass/gate_gpu.sh L56-65: shrink: a 10-line comment restates gate_cpu.sh:43-56 nearly verbatim. Two lines pointing there, −8.
  • scripts/compass/gate_gpu.sh L248, L277-278, L281-282, L310-311: shrink: three reason args repeat $1's first line word for word. WHY=${2:-${1%%$'\n'*}} in fail() and drop them, −3. The same line fixes F1.
  • tests/compass/test_gate_gpu_verdict_through_pipe.py: nothing to cut. The 11 + 8 cases are the brief's named result, and _expected() has two callers.

net: -11 lines possible.

To reach APPROVE

  1. Make F1's one-line change to fail().
  2. Show a mutation run on the new head where a forgotten reason no longer deletes the verdict. A re-run of X1 on the pass-count-only tree is enough.

The torch stub, a pass-count case and the two shrinks are recommended, not required.

fail() read its one-line reason as a bare $2 under `set -u`, so a call
that left it out aborted the gate with no GATE_GPU_RC line at all. It now
defaults to the finding's own first line. The three reasons that repeated
that line word for word are dropped, and the comment above finish() now
points at gate_cpu.sh's instead of restating it.

The test gains a pass-count case, where the pass count is the only
finding, and asserts that a verdict does not end on a colon: a reason cut
at its first line would. The throwaway tree carries a stub torch, since
the gate only reads its version, which takes the file from ~40 s to ~12 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Developer round 2: head 08ba5dd35

This commit sits on top of 579f7c285; nothing was force-pushed. It addresses review cycle 1: F1, which was blocking, plus the three recommended changes. The replies are on each inline thread.

Changes

  • F1 (gate_gpu.sh): fail() now reads WHY=${2:-${1%%$'\n'*}}, so a finding left without a one-line reason is named by its own first line.
  • Shrink (:278): the three reasons that repeated their finding's first line are dropped.
  • Shrink (:57): the 10-line comment is now a two-line pointer to gate_cpu.sh.
  • Test: a new case, [pass-count], where the pass count is the only finding.
  • Test: the exactly-once test also asserts the verdict does not end on :.
  • Test: the throwaway tree has a stub torch.
  • Size:
    • against 579f7c285: gate_gpu.sh +8 / -16, test +12 / -5;
    • against the tip: gate_gpu.sh +37 / -17, test +187 lines, 58 AST statements.

Mutations (merged tree ea5f9544a, the new file + #271's file, 22 tests)

Node ids are under tests/compass/test_gate_gpu_verdict_through_pipe.py::.

mutant result failing id, and the assertion it fails on
N00, null: reword a comment 22 passed none
X1, the pass-count reason replaced by a comment 1 failed test_the_verdict_key_is_printed_exactly_once_and_last[pass-count]. It fails on the :-ending check, and the verdict is still printed: ...4731 passed, expected 4730. The difference decomposes as:
X1b, X1 plus fail() reverted to a bare $2 (F1 reinstated) 1 failed the same id. It fails on the key-count check: the key is on 0 lines, and the run ended on unbound variable.
F1r, only the default reverted 22 passed: inert none. The three one-argument calls are never the first finding in any tree the test builds (see the :278 reply).
M05, fail() records no reason 6 failed [new-failure], [pass-count], and the 4 new-failure pipe ids

Named result

The committed test file was run against both trees:

  • tip 6a83b56bc (gate_gpu.sh 4e9c066b): 20 failed, 2 passed. All 20 new ids fail.
  • merged tree (gate_gpu.sh 478bbf53): 22 passed in 15.8 s, down from 40.9 s.

Gate 1 on the merged tree

  • Tip: 6a83b56bc.
  • Merged tree: git merge-tree --write-tree 6a83b56bc 08ba5dd35 = ea5f9544a1d049cef036ca40befd7779e576a4ee, staged as the unreferenced commit 4114a4762.
  • Where and how: node 18, xiaobizh_n18_cpu; git archive with stamps, at a private path; the tar md5 matched on both ends. Each tree ran its own gate_cpu.sh, md5 adb7e63d on both sides, under timeout -k 10 3000, unpiped, one gate at a time.
control 6a83b56bc merged ea5f9544a
atom.__file__ /tmp/xiaobizh280/gates/control/ATOM/atom/__init__.py /tmp/xiaobizh280/gates/merged/ATOM/atom/__init__.py
result 5166 passed, 155 skipped, 3 xfailed 5186 passed, 155 skipped, 3 xfailed
verdict GATE_CPU_RC=0 PASSED GATE_CPU_RC=0 PASSED
  • Node-id delta (--collect-only, 5303 → 5323): +20 / -0, exactly the new file's 20 ids.
  • One control run is discarded. The first control run was contaminated: I had copied the new test file into the control tree for the named-result run, and its 20 tip failures gave GATE_CPU_RC=1. I removed the file, checked the tree against its tarball (only __pycache__ differed), and re-ran. The figures above are from the clean run.

Lint: black rc=0 and ruff rc=0 on the test; bash -n is clean. No GPU was taken.

🤖 Generated with Claude Code

{"tests/test_suite.py": _suite(EXPECT + 1)},
"",
),
"no-summary": (1, "no usable pytest summary", {"tests/test_suite.py": ""}, ""),

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

F1r: blocking (principle 8, and AI_DEV_RULES gate 4: "an inert pin on a required finding blocks APPROVE"). The gap is reachable by a real caller flag, and one line pins it.

The :278 reply says no throwaway tree makes a default-relying fail() the first finding. One does, cheaply. gate_gpu.sh passes "$@" through to pytest, and only -r* is refused. --no-summary drops pytest's short summary, so no FAILED lines are printed while the counts still read 5 failed. The first finding is then the FAILED-line check at :269, which relies on the default since 08ba5dd35.

Measured on node 18 (xiaobizh_n18_cpu), merged tree ea5f9544a, gate md5 478bbf53. The tree has exactly the expected passes plus the 5 known failures, and the gate ran with --no-summary:

  • head: rc=1, the key on 1 line: GATE_GPU_RC=1 NOT PASSED -- the summary says 5 failed but 0 FAILED line(s) were printed.; 1 more finding(s) on stderr
  • F1r (the default reverted only): rc=1, the key on 0 lines. stderr ends gate_gpu.sh: line 243: $2: unbound variable.

So F1r is F1's own defect, live on a pass-through argument, and the suite gives 22 passed with it in place.

Fix: one entry here.

    "no-fail-lines": (1, "0 FAILED line(s) were printed", {}, " --no-summary"),

I measured it on a copy of this file:

  • head: [no-fail-lines] passed in 0.87 s;
  • F1r: 1 failed, test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines], on the key-count assertion at :167.

It uses the default suite, so it costs under a second.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3d2d5cb6a. I added your case, using the default suite. Its needle is "0 FAILED line(s) were printed; 2 more": on that tree the FAILED-line check is followed by the gone-ids finding and the pass-count finding, so it pins the count and the period strip too.

Measured on node 18, on the merged tree 68dbaa7a3 (gate md5 5e12a3af):

  • head: [no-fail-lines] passes. The verdict line is GATE_GPU_RC=1 NOT PASSED -- the summary says 5 failed but 0 FAILED line(s) were printed; 2 more finding(s) on stderr.
  • F1r (only WHY=$2 restored): 1 failed, 22 passed, test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines]. It fails on the key-count assertion, and the output ends on unbound variable.
  • null control: 23 passed.

My :278 reply claimed no tree reaches a default-relying first finding. That was wrong, since a pass-through flag reaches one. Thank you for finding it.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 2: REQUEST_CHANGES at head 08ba5dd35e863a035143663e769ac44bb0e336da

Blocking: one finding, F1r, with a one-line fix (inline on test_gate_gpu_verdict_through_pipe.py:71; principle 8 and gate 4's inert-pin rule).

  • Reverting fail()'s default alone leaves the suite at 22 passed.
  • The dev round says no tree reaches it, but a real caller flag does. gate_gpu.sh --no-summary passes through to pytest, prints no FAILED lines, and makes the FAILED-line check (which relies on the default) the first finding.
  • With F1r, that run ends on line 243: $2: unbound variable, with GATE_GPU_RC= on 0 lines. That is F1's defect, live at this head without the default.
  • The fix is one CASES entry. It passes at the head in 0.87 s and fails by name under F1r.

Everything else in the delta checks out. X1, X1b and the null control reproduce exactly. The stub torch hides nothing the verdict needs. The two shrinks are applied. The merged tree gates green at +20.

I reviewed only the delta 579f7c285..08ba5dd35: one commit, with no force-push. All measurements were taken on node 18, xiaobizh_n18_cpu, on the merged tree.

1. Ruling on F1r (principle 8)

Not acceptable as inert; one tree must make a default-relying call the first finding. The dev round treats this as unreachable, but it is reachable:

  • The path. for arg in "$@" refuses only -r*, so --no-summary reaches pytest. It removes the short summary, so OBS_N=0 while the summary still reads 5 failed. The first finding is then :269, which since this commit has no explicit reason.
  • Measured at the head. On a tree with exactly the expected passes plus the 5 known failures, run with --no-summary, the gate ends GATE_GPU_RC=1 NOT PASSED -- the summary says 5 failed but 0 FAILED line(s) were printed.; 1 more finding(s) on stderr. The key is on 1 line.
  • Measured under F1r (WHY=$2 restored, nothing else changed). The key is on 0 lines, stderr ends gate_gpu.sh: line 243: $2: unbound variable, and the suite gives 22 passed.
  • Why this blocks. The default is load-bearing in production code at this head, not a future-edit guard, and its only witness (X1b) needs a second, unrelated mutation to fire. That is an inert pin on the required finding, which AI_DEV_RULES gate 4 says blocks APPROVE.
  • The fix, measured on a copy of the file. One entry, "no-fail-lines": (1, "0 FAILED line(s) were printed", {}, " --no-summary"). At the head: 1 passed in 0.87 s. Under F1r: test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines] fails on the key-count assertion (:167).
  • Stop rule. This is the unclosed half of F1. If it is still open after the next cycle, the review-loop stop rule applies.

2. Mutations re-run

These ran on the merged tree ea5f9544a, gate md5 478bbf53. The test set was the new file plus #271's, 22 tests. Each mutant was line-count preserving and bash -n clean. Node ids are under tests/compass/test_gate_gpu_verdict_through_pipe.py::.

mutant result failing id and assertion
N00 null, reword the comment at :251 22 passed (15.8 s) none
X1, pass-count reason replaced by a comment 1 failed, 21 passed test_the_verdict_key_is_printed_exactly_once_and_last[pass-count], on the colon check: ...4731 passed, expected 4730. The difference decomposes as:
X1b, X1 plus WHY=$2 1 failed, 21 passed the same id, on the key-count check. The key is absent from the output.
F1r, WHY=$2 only 22 passed: inert none. This is the blocking finding above.

These match the dev round's counts and ids.

3. The stub torch

  • It hides nothing the verdict needs. In gate_gpu.sh and _lib.sh at the head, torch is imported only by the two version probes (:159-160), each || echo UNKNOWN, and read only for the warn-only toolchain comparison.
  • No other gate or _lib.sh line imports torch.
  • The stub lives only in the throwaway tree. The real imports were observed in cycle 1's GPU run on node 18, which printed torch 2.10.0+rocm7.2.4.git3d3aa833 and hip 7.2.53211, both matching the baseline.
  • Cost: 22 tests in 15.2–15.8 s, down from 40.9 s.

4. ponytail-review (delta only)

Nothing to cut:

  • the new fail() comment is two lines;
  • EXPECT removes a repeated expression;
  • the colon assertion is one line;
  • the stub is one dict entry.

Lean already. Ship.

A cosmetic nit, which blocks nothing: a first line that ends in . gives printed.; 1 more finding(s).

5. Gate: the tree that will land

  • Tip 6a83b56bc, re-read after fetch. git merge-tree --write-tree 6a83b56bc 08ba5dd35 = ea5f9544a1d049cef036ca40befd7779e576a4ee, which is the same as the dev round's.
  • Staging. Unreferenced commit 9f5f5ad57, then git archive with stamps (.compass-changed lists the PR's 2 files), then docker exec -i … tar -x into /tmp/pr294rv/ inside the container. The md5 matched on both ends (2df70d3c…).
  • atom.__file__ = /tmp/pr294rv/merged/ATOM/atom/__init__.py, and the gate printed commit: 9f5f5ad57 (stamp).
  • Result, with the tree's own gate_cpu.sh (md5 adb7e63d…), under timeout -k 10 3000, unpiped, run once: 5186 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0 PASSED.
  • Delta: +20. It is measured against the dev round's control of 5166 at the same tip, which was run with the same gate_cpu.sh md5. It equals the file's 20 ids, and nothing failed.

To reach APPROVE

  1. Add the no-fail-lines case.
  2. Show F1r failing that id by name at the new head.

Nothing else is required.

A caller's --no-summary reaches pytest, which then prints no FAILED lines
while its counts still say 5 failed. The FAILED-line check becomes the
first finding, and it has no explicit reason, so the verdict depends on
fail()'s default. A no-fail-lines case now drives that run; without the
default it ends on `set -u` with no GATE_GPU_RC line.

The aggregate verdict drops a trailing period from the first reason
before appending "; N more finding(s)", so it no longer reads "printed.;".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Developer round 3: head 3d2d5cb6a

This commit sits on top of 08ba5dd35; nothing was force-pushed. It closes F1r and applies the nit.

Changes

  • Test: one new CASES entry, "no-fail-lines": (1, "0 FAILED line(s) were printed; 2 more", {}, " --no-summary").
  • Script, the nit: the aggregate now appends ; N more finding(s) to ${WHY%.}, so a first reason that ends in . no longer reads printed.;. This is one existing line, edited in place. The stderr texts are unchanged.

Can any fail() caller still reach a bare $2?

No. I grepped every $2 in gate_gpu.sh at the head. It appears in only two places:

  • fail(), where it is only ${2:-${1%%$'\n'*}}.
  • finish()'s non-zero branch. All 11 non-zero finish calls pass a literal reason. The aggregate call passes "$WHY", which is set whenever VERDICT>0: either to $2 or to $1's first line, possibly empty, but never unset. finish 0 never expands $2.

Of the seven fail() calls, four pass an explicit reason (errors, new ids, gone ids, pass count). The other three rely on the default:

  • the FAILED-line count, now driven first by [no-fail-lines];
  • the ERROR-line count;
  • the failed count.

The last two cannot be the first finding in any tree the test builds, and F1r shows the default itself is now pinned.

Mutations (merged tree 68dbaa7a3, gate md5 5e12a3af, the new file + #271's file, 23 tests)

Node ids are under tests/compass/test_gate_gpu_verdict_through_pipe.py::.

mutant result failing id and assertion
N00, null: reword a comment 23 passed none
F1r, WHY=$2 only 1 failed, 22 passed test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines], on the key-count assertion; the output ends on unbound variable
P1, ${WHY%.} reverted to $WHY 1 failed, 22 passed the same id, on the reason needle

Named result

The committed test file was run against both trees:

  • tip d96027e52 (gate_gpu.sh 4e9c066b): 21 failed, 2 passed. [no-fail-lines] is among them. This ran on a separate copy of the control tree, and the control gate ran before it.
  • merged tree: 23 passed in 16.2 s.

Gate 1 on the merged tree

  • Tip: d96027e52, re-read after fetch.
  • Merged tree: git merge-tree --write-tree d96027e52 3d2d5cb6a = 68dbaa7a3c31e389b6726a7ce1a101627db3d9aa, staged as the unreferenced commit 0c1c36976.
  • Where and how: node 18, xiaobizh_n18_cpu; git archive with stamps; the tar md5 matched on both ends. Each tree ran its own gate_cpu.sh, md5 adb7e63d on both sides, under timeout -k 10 3000, unpiped, one gate at a time.
control d96027e52 merged 68dbaa7a3
atom.__file__ /tmp/xiaobizh280/gates/control/ATOM/atom/__init__.py /tmp/xiaobizh280/gates/merged/ATOM/atom/__init__.py
result 5168 passed, 155 skipped, 3 xfailed 5189 passed, 155 skipped, 3 xfailed
verdict GATE_CPU_RC=0 PASSED GATE_CPU_RC=0 PASSED
  • Node-id delta (--collect-only, 5305 → 5326): +21 / -0, exactly the file's 21 ids.
  • Size against the tip: gate_gpu.sh +37 / -17, test +188 lines.
  • Lint: black rc=0 and ruff rc=0 on the test; bash -n is clean. No GPU was taken.

🤖 Generated with Claude Code

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 3: APPROVE at head 3d2d5cb6aae3546957cda785b672933b4a46c091

No blocking issues. F1r, cycle 2's only blocking finding, is closed: reinstating it fails [no-fail-lines] by name. Nothing remains open from cycles 1 and 2. The stop rule is not triggered.

I reviewed only the delta 08ba5dd35..3d2d5cb6a: one commit, with no force-push.

  • gate_gpu.sh: one line, WHY="${WHY%.}; …".
  • The test file: one new CASES entry, no-fail-lines.

All measurements were taken on node 18, xiaobizh_n18_cpu, on the merged tree.

Mutations

These ran on the merged tree 68dbaa7a3, gate md5 5e12a3af. The test set was the new file plus #271's, 23 tests. Each mutant was line-count preserving and bash -n clean. The gate's md5 was re-checked after restore. Node ids are under tests/compass/test_gate_gpu_verdict_through_pipe.py::.

mutant result failing id and assertion
N00 null, reword a comment 23 passed (16.8 s) none
F1r, fail() back to a bare WHY=$2 1 failed, 22 passed test_the_verdict_key_is_printed_exactly_once_and_last[no-fail-lines], on the key-count check. The output has no GATE_GPU_RC= line.
P1, period fix reverted (${WHY%.} → $WHY) 1 failed, 22 passed the same id, on the reason check. The line is ...0 FAILED line(s) were printed.; 2 more finding(s) on stderr

The two mutants fail on different assertions of the same id, so each assertion holds its own defect.

"No bare $2": checked independently

  • In gate_gpu.sh at the head, $2 appears only in fail(), where it is defaulted, and in finish()'s non-zero printf (:63).
  • All 11 non-zero finish calls pass a literal reason: :115, :121, :142, :147, :148, :154, :156, :190, :218, :231.
  • The aggregate call at :332 passes "$WHY", which is set whenever VERDICT>0.
  • finish 0 (:346) takes the branch that never reads $2.
  • _lib.sh's $2 uses are local …=$2 in functions whose callers pass the argument; none is new in this PR.

The claim holds.

ponytail-review (delta only)

The delta is one parameter expansion and one dict entry. Lean already. Ship.

Gate: the tree that will land

  • Tip d96027e52, re-read after fetch. git merge-tree --write-tree d96027e52 3d2d5cb6a = 68dbaa7a3c31e389b6726a7ce1a101627db3d9aa, the same as the dev round's.
  • Staging. Unreferenced commit e8c36c52a, then git archive with stamps (.compass-changed lists the PR's 2 files), then docker exec -i … tar -x into /tmp/pr294rv/ inside the container. The md5 matched on both ends (e6441939…).
  • atom.__file__ = /tmp/pr294rv/merged/ATOM/atom/__init__.py, and the gate printed commit: e8c36c52a (stamp).
  • Result, with the tree's own gate_cpu.sh (md5 adb7e63d…), under timeout -k 10 3000, unpiped, run once: 5189 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0 PASSED.
  • This equals the dev round's figure, and at +21 it is exactly the file's 21 ids against their control at the same tip.

Carried forward, not re-measured this cycle

  • Cycle 1's real GPU run of this PR's gate_gpu.sh on node 18 ended GATE_GPU_RC=0 PASSED exactly once.
  • Since then, only the fail() reason plumbing and the ${WHY%.} trim have changed. Neither touches the zero-verdict path.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant