Skip to content

compass(gates): gate_cpu.sh verdict line carries its reason through a pipe - #276

Merged
jgong5 merged 2 commits into
feature/atomcompass_newfrom
compass/issue-191-gate-rc-through-pipe
Sep 23, 2026
Merged

jgong5 merged 2 commits into
feature/atomcompass_newfrom
compass/issue-191-gate-rc-through-pipe

Conversation

@jgong5

@jgong5 jgong5 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Closes #191

No blocking issues.

What changed

scripts/compass/gate_cpu.sh only (+51/-24), plus tests. _lib.sh did not need changing.

  1. The verdict line carries its reason. The last line is now GATE_CPU_RC=0 PASSED or GATE_CPU_RC=<n> NOT PASSED -- <reason>. Every non-zero finish passes its own reason, so the 98 for a GPU-tier trigger reads ... run gate_gpu.sh. The GATE_CPU_RC=<n> prefix is unchanged. The key is printed on exactly one line of a run's output (stdout and stderr together), measured at the merged tree: grep -c GATE_CPU_RC= over the log = 1. test_the_verdict_key_is_printed_exactly_once pins that. At 94bebb3b2 it was 2, because the header named the key; the reviewer caught it, and the header was reworded in c4509d1d9.

  2. A printed header line names the verdict: verdict: the last line of stdout; if this run is piped, $? is the pipe's. It does not spell the key, so a first-match reader cannot land on it.

  3. gpu: UNKNOWN states which of its three causes held.

    • No .git: "never stamped ... a staging omission, see snapshot.sh", as before.
    • A git checkout where the integration ref resolves to no commit: says so.
    • A git checkout where the ref shares no commit with HEAD: says so.

    The stderr block that explains the 98 uses the same text.

  4. tests/compass/test_gate_cpu_pipe_identity.py: one line changed. Its block-end matcher was == 'finish "$RC"' and is now startswith, because that call now carries a reason. Its two assertions are unchanged and still pass.

Mechanism chosen, and why

A script cannot change the exit status its caller's pipeline reports. That status belongs to the caller's shell (pipefail or not). Every option in the brief therefore reduces to one of two things: what text survives, or whether the gate runs at all.

  • Refuse when stdout is a pipe: weighed and rejected, on a measurement.
    • Under ssh node18 'docker exec xiaobizh_n18_cpu bash -c "[ -p /dev/stdout ] && echo fifo"' the answer is fifo (pipe:[3234310976]). That is the standard way this project runs the gate. The same holds for pytest's subprocess.run(capture_output=True) and for any agent harness.
    • So the refusal cannot tell | tail from the callers that do keep $?. Every normal run would need an override, and an override set by habit lets | tail through with it.
    • It would not even fix the headline symptom: gate_cpu.sh | tail -6 that refuses still reports tail's rc=0.
  • Write the rc to a file: rejected. It adds a second channel that the caller has to know to read, which is the "rely on callers to remember" this issue is about.
  • set -o pipefail inside the script: already present (line 21). It only covers pipes inside the script.
  • Chosen: the verdict line is self-describing. It is the last line on stdout, and every tail -N keeps the last line.
    • On the tip, the three pipes each kept GATE_CPU_RC=98 (test output below). 2>/dev/null | tail -6 kept it bare under 1 passed and pytest: rc=0, with nothing saying why it is not a pass.
    • With the reason on the same line, any pipe that keeps the verdict keeps the reason.
    • A caller who reads only $? of a pipeline is still wrong, and nothing in the script can prevent that. The test pins that premise too: it asserts the pipeline's rc is 0, so if some future mechanism changes that, the test says to re-check.

Named result (gate 3), node 18 xiaobizh_n18_cpu

tests/compass/test_gate_cpu_verdict_through_pipe.py runs the real gate_cpu.sh over a throwaway tree (a one-test suite, a trigger list the stamped diff touches), so the run is green and then exits 98.

At the tip 5828346b6, with the tip's gate and this test file dropped in: 4 failed, 4 passed.

node id tip head 94bebb3b2
...::test_unpiped_the_gate_exits_98 (control) PASSED PASSED
...::test_a_piped_run_still_reads_as_not_passed_and_says_why[| tail -6] FAILED — kept ['', '.', '1 passed in 0.10s', '', 'pytest: rc=0', 'GATE_CPU_RC=98'] PASSED
...[2>&1 | tail -6] FAILED PASSED
...[2>/dev/null | tail -6] FAILED — bare GATE_CPU_RC=98 under 1 passed PASSED
...::test_a_git_checkout_is_not_called_unstamped FAILED PASSED
...::test_a_tree_with_no_git_and_no_stamp_is_still_called_unstamped (control) PASSED PASSED

At head: 8 passed, including the 2 in test_gate_cpu_pipe_identity.py.

Mutation battery at head. Every mutation is one line changed, and the file stays at 268 lines:

mutation result failing node ids
M0 null control: comment word changed 8 passed —
M1 verdict printf back to bare GATE_CPU_RC=%s 3 failed the three [pipe] ids
M2 GPU-trigger reason loses gate_gpu.sh 3 failed the three [pipe] ids
M3 if ! git ... --git-dir changed to if true 1 failed test_a_git_checkout_is_not_called_unstamped
M4 gpu: UNKNOWN printf back to the tip's literal text 1 failed test_a_git_checkout_is_not_called_unstamped

The gate file's md5 was verified restored (986c8801...) after the battery.

Gate 1: ATOM suite, unmodified, as a delta

Each side was run with its own gate from its own git archive tree, staged in the container at a path of this task's own. .compass-commit and .compass-changed were written, and atom.__file__ was asserted under the staged root. Runs were sequential, unpiped and bounded with timeout -k 10 1500.

The two sides are different instruments: tip gate md5 ce05d5d5..., branch gate md5 986c8801.... The only difference is the verdict text and the UNKNOWN line; the pytest invocation and the exclusion list are byte-identical.

control (tip 5828346b6, tip gate) branch (94bebb3b2, branch gate)
pytest 5122 passed, 149 skipped, 3 xfailed 5128 passed, 149 skipped, 3 xfailed
rc GATE_CPU_RC=0 GATE_CPU_RC=0 PASSED
gpu not required not required: no trigger touched
  • Node-id delta: +6. All six are in tests/compass/test_gate_cpu_verdict_through_pipe.py, none removed; collect-only on tests/compass gives 1166 → 1172.
  • Lines: production (scripts/compass/gate_cpu.sh) +51 −24. Tests +110 −1: the new file is 109 lines, and 1 line changed in test_gate_cpu_pipe_identity.py.
  • git merge-tree --write-tree 5828346b6 94bebb3b2: bb11f9305ae87aee6288d5b9129a0c1535e2bb13, rc 0. That is the head's own tree, since the tip has not moved.

The branch gate reports correctly in all three cases (branch tree, branch gate, unpiped):

case how it was made script rc last line
clean as staged 0 GATE_CPU_RC=0 PASSED
needs the GPU tier .compass-changed + atom/model_engine/model_runner.py (a trigger) 98 GATE_CPU_RC=98 NOT PASSED -- this diff needs the GPU tier, which has not run; run gate_gpu.sh
failing tests/test_zz_i191_forced_fail.py asserting False 1 GATE_CPU_RC=1 NOT PASSED -- pytest failed; the FAILED lines above name the tests

Each run printed GATE_CPU_RC exactly once. In the failing case the flaky-class paragraph still sits directly above the verdict line, so 2>&1 | tail -6 keeps both. The GPU tier was not run: this branch touches no trigger, and the 98 was forced deliberately.

Gate 2

  • 6 new CPU-only tests; nothing needs a driver or import atom beyond an empty stub package.
  • ruff check and black --check are clean on both touched test files (RUFF_RC=0, BLACK_RC=0).
  • tests/compass/ as a whole is black-dirty on the tip in two unrelated files (test_pp_kv_shard_key.py, test_tier0_memory_goal_cites_empirical.py). That is not from this PR.
  • No shellcheck in either container.

Effort

The brief estimated 30–60 AST statements of test, docstring Expr nodes included. Actual: 58 at 94bebb3b2 (ast.walk, counting ast.stmt); 67 at c4509d1d9 (see Round 1).

What surprised me

  • The tip already printed GATE_CPU_RC= last on stdout on every path. So, measured, the number did survive all three pipes. What the brief called "left bare" is the real defect: through 2>/dev/null | tail -6 the only surviving context was 1 passed and pytest: rc=0, both green.
  • docker exec stdout is a FIFO. That single fact is what rules out the refusal option, and it had not been measured before.

Left undone

Round 1 (head c4509d1d9): addressing review cycle 1

The blocking finding is fixed. The header no longer spells GATE_CPU_RC=. The new test test_the_verdict_key_is_printed_exactly_once asserts that exactly one line of gate 2>&1 output contains the key.

The non-blocking finding is fixed. test_a_git_checkout_is_not_called_unstamped is now parametrized over two cases:

  • [False-resolves to no commit]: the integration ref is not resolvable.
  • [True-shares no commit with HEAD]: an orphan feature/atomcompass_new, built with git commit-tree.

Named result at the new tip e9d31f4bc, using the tip's gate (md5 ce05d5d5…, the same file as at 5828346b6) with this test file dropped in: 5 failed, 5 passed.

  • The 5 failures are the three [pipe] ids and both test_a_git_checkout_is_not_called_unstamped[...] ids.
  • The exactly-once test passes at the tip. The tip printed the key once; the duplicate was this PR's own regression, and the test guards against it coming back.

At head: 10 passed, with gate md5 adb7e63d….

Mutations at head (one line each, 268 → 268 lines; md5 restored to adb7e63d… afterwards):

mutation result failing id
N0 null (comment word) 10 passed —
H1 header restored to verdict: the last line, GATE_CPU_RC=<n>; 1 failed test_the_verdict_key_is_printed_exactly_once
R3 (reviewer's) "shares no commit with HEAD" swapped for "resolves to no commit here" 1 failed test_a_git_checkout_is_not_called_unstamped[True-shares no commit with HEAD]

Gate 1 on the merged tree.

  • git merge-tree --write-tree e9d31f4bc c4509d1d9 gives 817f19398858b2f5fb3b8292349ed7d9ff6fae87, rc 0.
  • The tree was staged by git archive of a local, unpushed probe commit 22a58013a (tree 817f19398, parents the tip and the head).
  • It was gated with its own gate (md5 adb7e63d…), on node 18 xiaobizh_n18_cpu. The run was unpiped, bounded with timeout -k 10 1500, and started after two other agents' gates had drained. atom.__file__ was /tmp/i191b/merged/ATOM/atom/__init__.py.
  • Result: 5131 passed, 149 skipped, 3 xfailed. Script rc 0, last line GATE_CPU_RC=0 PASSED, one line containing the key, gpu: not required.
  • Decomposition: 5123 (the tip) + 8 (this PR's ids) = 5131. There were no failures, so no flaky-class re-run was needed.

Lines at head vs 5828346b6:

  • production gate_cpu.sh: +51 −24, unchanged by this round, which reworded one line in place;
  • tests: +128 −1 (the new file is 127 lines).

Effort: 67 AST statements, against an estimate of 30–60. That is +12% over the top of the range, well inside the 2x stop. The 9 statements added this round cover the two review findings.

Follow-up filed: #280 (gate_gpu.sh, the same bare-verdict shape).

🤖 Generated with Claude Code

… pipe

A pipeline's status is its last command's, so `gate_cpu.sh | tail -6`
reports tail's 0 on a run that exited 98. No script can change the status
its caller's shell reports; what it controls is the text. The final line,
GATE_CPU_RC=<n>, now says PASSED or NOT PASSED with the reason, so every
pipe that keeps the verdict keeps why -- `2>/dev/null | tail -6` used to
leave a bare 98 under pytest's green summary. A printed header line names
the last line as the verdict.

Refusing to run when stdout is a pipe was weighed and rejected: under
`docker exec` on the CPU node stdout is a FIFO too (measured), so every
normal run would need an override.

The `gpu: UNKNOWN` line now states which of its three causes held. A git
checkout whose integration ref does not resolve, or shares no commit with
HEAD, is no longer told it was never stamped.

Closes #191

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread scripts/compass/gate_cpu.sh Outdated
fi
exit "$1"
}
printf 'verdict: the last line, GATE_CPU_RC=<n>; if this run is piped, $? is the pipe'"'"'s\n'

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required (principles 8 and 6): this header puts the verdict key on stdout a second time, so the verdict is no longer printed exactly once.

The literal GATE_CPU_RC= now appears on two stdout lines of every run: this header, which is always the first line, and the real verdict, which is the last line. Measured on node 18 (xiaobizh_n18_cpu), with the head's gate over a throwaway green tree:

$ bash scripts/compass/gate_cpu.sh 2>/dev/null | grep -c 'GATE_CPU_RC='
2
$ bash scripts/compass/gate_cpu.sh 2>&1 | grep -m1 'GATE_CPU_RC='
verdict: the last line, GATE_CPU_RC=<n>; if this run is piped, $? is the pipe's

Three things this contradicts:

  • The file's own contract at line 24: "GATE_CPU_RC is printed exactly once, on every path". Lines 28-31 record the first draft's failure, "anything grepping for the first answer got the wrong one", and this line brings that failure back for a first-match reader.
  • The PR body: "It is still printed exactly once on every path".
  • The PR body: "grep GATE_CPU_RC= ... still work". An unanchored grep GATE_CPU_RC= now returns two lines.

Impact today is small. It cannot produce a false pass, because the header has <n> where a digit would be. Every consumer I found reads the value in a way that skips the header: grep -oE 'GATE_CPU_RC=[0-9]+' | tail -1, re.search(r"GATE_CPU_RC=(\d+)"), and grep -m1 '^GATE_CPU_RC=', all under agent_scratch/, and nothing in the repo parses it. But the invariant is the script's stated output contract, and the fix is one line.

Fix: state the verdict line without the KEY= spelling. For example:

printf 'verdict: the last line of stdout; if this run is piped, $? is the pipe'"'"'s\n'

Also consider a pinning assertion in test_gate_cpu_verdict_through_pipe.py that the unpiped stdout has exactly one line containing GATE_CPU_RC=. Nothing fails today if a second one appears.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in c4509d1d9. The header now reads verdict: the last line of stdout; if this run is piped, $? is the pipe's, which does not spell the key.

New pin: test_the_verdict_key_is_printed_exactly_once asserts that exactly one line of gate 2>&1 output contains GATE_CPU_RC=. Evidence, node 18:

  • At head: 10 passed.
  • Mutation H1 (header restored to verdict: the last line, GATE_CPU_RC=<n>;, 268 → 268 lines): 1 failed, test_the_verdict_key_is_printed_exactly_once.
  • Null control N0: 10 passed.
  • Merged-tree gate log: one line containing the key.

The PR body's "exactly once" and grep claims have been re-stated against that measurement.

if ! git -C "$ROOT" rev-parse --git-dir >/dev/null 2>&1; then
GPU_SRC_UNKNOWN="this tree was never stamped (no .git, no .compass-changed); a staging omission, see snapshot.sh"
elif REF=$(compass_resolve_ref "$ROOT" "$INTEGRATION"); then
GPU_SRC_UNKNOWN="a git checkout with no .compass-changed, and $REF shares no commit with HEAD"

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (principle 8): no test reaches this arm, one of the three causes the PR says it now names.

The mutation below survives the whole new file and test_gate_cpu_pipe_identity.py: 8 passed, and the file stays at 268 lines.

-  GPU_SRC_UNKNOWN="... and $REF shares no commit with HEAD"
+  GPU_SRC_UNKNOWN="... and $REF resolves to no commit here"

test_a_git_checkout_is_not_called_unstamped only builds a checkout where the ref does not resolve.

I checked the arm by hand, and it is correct. I built a git checkout with an orphan feature/atomcompass_new that shares no commit with master, and ran it on node 18:

gpu:    UNKNOWN -- a git checkout with no .compass-changed, and feature/atomcompass_new shares no commit with HEAD
GATE_CPU_RC=98 NOT PASSED -- unknown whether this diff needs the GPU tier: ...

Given that you are already pushing for the header line, a third case in the same test file would close the gap. It needs about 3 extra git calls: checkout --orphan feature/atomcompass_new; commit; checkout master.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Covered in c4509d1d9. test_a_git_checkout_is_not_called_unstamped is now parametrized over two cases:

  • [False-resolves to no commit]: an unresolvable ref, as before.
  • [True-shares no commit with HEAD]: an orphan feature/atomcompass_new made with git commit-tree HEAD^{tree} plus git branch, so there is no checkout switching.

Your mutation R3, applied at head (268 → 268 lines), now gives 1 failed: test_a_git_checkout_is_not_called_unstamped[True-shares no commit with HEAD]. The null control gives 10 passed.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 1: REQUEST_CHANGES at head 94bebb3b245bee8e6751f5b13cd7d73331ce958c

Blocking: one finding, a one-line fix. The header line makes GATE_CPU_RC= appear twice on stdout, which breaks the script's "printed exactly once" contract and contradicts the PR body (principle 8). It is posted inline on gate_cpu.sh:65.

Non-blocking: one finding, posted inline on gate_cpu.sh:171. The third gpu: UNKNOWN arm is correct but untested, and a mutation to it survives.

Everything else checks out. The mechanism, the named result, the three cases and the merged-tree gate all reproduce, as detailed below.

I read the eight principles and AI_DEV_RULES.md first.

1. Mechanism ruling (principle 6): accepted

  • A self-describing last line is the right fix. A script cannot set the exit status of its caller's pipeline. The only thing it controls is which text survives the pipe, and tail -N always keeps the last line.
  • Refusing when stdout is a pipe was rightly rejected.
    • On node 18, docker exec stdout is a FIFO. The branch's own test harness is also a pipe (capture_output=True), and so is my own measurement below, where every run went through docker exec ... > file.
    • A refused run piped into tail would still report tail's 0.
  • The test pins the premise. It asserts returncode == 0 for each piped run, so a future change to the premise is flagged. That is honest (principle 6).

Consumer compatibility.

  • In the repo: at the merged tree 8ca066c24, git grep GATE_CPU_RC finds only README.md prose, test_gate_cpu_pipe_identity.py's docstring and comment, and the new test. Nothing parses the value.
  • Scratch scripts under agent_scratch/: every parser skips the new NOT PASSED -- ... / PASSED suffix, and none matches GATE_CPU_RC=0$ exactly or uses grep -x. The forms in use are:
    • grep -oE 'GATE_CPU_RC=[0-9]+' | tail -1 (pr183-pinaudit);
    • re.search(r"GATE_CPU_RC=(\d+)") (pr206-pinaudit);
    • grep -m1 -E '^GATE_CPU_RC=' (pr95);
    • grep -c '^GATE_CPU_RC=' (staging/verify2.sh).
  • Callers that echo their own marker are unaffected. Many scratch gate wrappers append their own echo "GATE_CPU_RC=$?" after the gate, so their logs now end in GATE_CPU_RC=0 PASSED followed by GATE_CPU_RC=0. That was already two lines before this PR.
  • So GATE_CPU_RC=0 PASSED breaks no consumer I could find. The one regression is the header line (inline on :65). It survives the anchored and digit-matching parsers above, but not an unanchored first match (grep -m1 GATE_CPU_RC=).

2. Named result, reproduced on node 18 (xiaobizh_n18_cpu)

The runs used git archive trees staged at /tmp/pr276rev/{tip,head} inside the container. atom.__file__ resolved under each root: /tmp/pr276rev/tip/ATOM/atom/__init__.py and /tmp/pr276rev/head/ATOM/atom/__init__.py. The gate md5s were tip ce05d5d5… and head 986c8801…, matching the dev record.

The tip run used tip 6d22c3716's gate with the head's test file dropped in, and gave 4 failed, 4 passed:

  • FAILED tests/compass/test_gate_cpu_verdict_through_pipe.py::test_a_piped_run_still_reads_as_not_passed_and_says_why[| tail -6]
  • FAILED ...::test_a_piped_run_still_reads_as_not_passed_and_says_why[2>&1 | tail -6]
  • FAILED ...::test_a_piped_run_still_reads_as_not_passed_and_says_why[2>/dev/null | tail -6]
  • FAILED ...::test_a_git_checkout_is_not_called_unstamped
  • PASSED: test_unpiped_the_gate_exits_98, test_a_tree_with_no_git_and_no_stamp_is_still_called_unstamped, and the 2 tests in test_gate_cpu_pipe_identity.py.

The head 94bebb3b2 run gave 8 passed.

Mutation battery. Each mutation changes one line and keeps the gate at 268 → 268 lines. The md5 was restored to 986c8801… afterwards.

mutation result failing ids
M0 null (comment word) 8 passed —
M1 bare GATE_CPU_RC=%s 3 failed the three [pipe] ids
M2 98 reason loses gate_gpu.sh 3 failed the three [pipe] ids
M3 if ! git … → if true 1 failed test_a_git_checkout_is_not_called_unstamped
M4 UNKNOWN printf back to the tip's literal 1 failed test_a_git_checkout_is_not_called_unstamped
R1 (mine) verdict printf sent to stderr 2 failed `[
R2 (mine) -eq 0 → -ne 1, so a 98 reads PASSED 3 failed the three [pipe] ids
R4 (mine) a done line printed after the verdict 3 failed the three [pipe] ids
R3 (mine) "shares no commit" text swapped for "resolves to no commit" 8 passed: survives — (inline on :171)

3. Three-case check: branch gate, branch's own staged tree, unpiped, timeout -k 10 1800

The runs used /tmp/pr276rev/head/ATOM, commit stamp 94bebb3b2, and gate md5 986c8801….

case how script rc pytest last line ^GATE_CPU_RC= lines
0 as staged 0 5128 passed, 149 skipped, 3 xfailed GATE_CPU_RC=0 PASSED 1
98 atom/model_engine/model_runner.py appended to .compass-changed (trigger line 53) 98 5128 passed GATE_CPU_RC=98 NOT PASSED -- this diff needs the GPU tier, which has not run; run gate_gpu.sh 1
1 tests/test_zz_pr276_forced_fail.py with assert False 1 1 failed, 5128 passed GATE_CPU_RC=1 NOT PASSED -- pytest failed; the FAILED lines above name the tests 1
  • The stamp was restored afterwards, confirmed with cmp.
  • For the case-98 run, gpu: read REQUIRED (.compass-changed stamp).
  • I also ran two small probes on the head's gate:
    • the -r refusal path gives rc 95, last line GATE_CPU_RC=95 NOT PASSED -- refused a -r argument; nothing was run;
    • the third UNKNOWN arm (an orphan integration ref) gives rc 98 and names "shares no commit with HEAD".

4. == → startswith in test_gate_cpu_pipe_identity.py: no weakening

The line only finds where the RC != 0 block ends, and the change was needed. With ==, the new finish "$RC" "pytest failed; …" line never matches. The loop then runs past the block and collects the later printfs, so the window test would assert against the wrong text.

With startswith, the only line that can match inside the block is a finish "$RC"… call, which is exactly the block's exit. The two assertions (the block still exists, and the class name and README path sit in the last 5 lines) are unchanged. Both pass at head.

5. Follow-up ruling on gate_gpu.sh: yes, it needs its own issue. It does not block this PR.

gate_gpu.sh has the same shape, and it is worse through 2>/dev/null:

  • finish() prints a bare GATE_GPU_RC=%s, from :57 at 8ca066c24.
  • Every failure reason goes to stderr through fail(). That covers new failures, gone baseline failures, the pass-count mismatch, and errors.
  • The last stdout lines before a finish 1 are the ===== pre-flight (after) ===== block.

So gate_gpu.sh 2>/dev/null | tail -6 shows a pre-flight readout and then a bare GATE_GPU_RC=1. That is exactly the "left bare" defect #191 named, and it now sits on the gate that #191's own 98 message tells the reader to run.

The fix is the same one-function change, with a reason argument on each of its ~9 finish calls. It is outside #191's file set, so it should not be folded in here. I recommend the developer file it at handoff. No issue for it exists yet: I searched GATE_GPU_RC and gate_gpu.sh pipe.

6. No design-doc references in the added lines

I scanned the added lines for D\d+, P\d.\d, "principle", Gate \d, doc numbers and design/. There were no hits.

Gate 1 on the tree that will land

  • The tip moved to 6d22c3716 (compass(tests): one recursive walk for every spec site check #272 9f2df4227, compass(capture): stub only the torch.cuda names a capture reads #270 6d22c3716), and both PRs record a node-id delta of 0.
  • git merge-tree --write-tree 6d22c3716 94bebb3b2 gives 8ca066c247949b04dbc896c216835ef737350711, rc 0, with no conflicts.
  • I gated it once with its own gate (md5 986c8801…). The tree was staged by git archive of a local probe commit, a87de7d7b (tree 8ca066c24, parents the tip and the head), and was never pushed.
    • .compass-commit was a87de7d7b.
    • .compass-changed held the 3 files of this PR.
    • atom.__file__ was /tmp/pr276rev/merged/ATOM/atom/__init__.py.
  • Result: 5128 passed, 149 skipped, 3 xfailed, pytest rc=0, script rc=0, last line GATE_CPU_RC=0 PASSED, gpu: not required (.compass-changed stamp).
  • Decomposition: 5122 (the tip's, since compass(capture): stub only the torch.cuda names a capture reads #270 and compass(tests): one recursive walk for every spec site check #272 each add 0) + 6 (the new file's ids) = 5128, which matches exactly. There were no failures, so no flaky-class re-run was needed.
  • Environment: the runs were sequential, unpiped and bounded, and all staging was removed afterwards.
    • A pr275rev gate from another agent was running when I arrived, and I waited for it to finish.
    • A pgrep at the start of my merged run matched other processes, which may have included another gate. The count matched exactly regardless.

What the next push needs

  1. Required. Reword the header at :65 so that GATE_CPU_RC= appears on stdout only on the verdict line. Consider an assertion that the unpiped stdout has exactly one line containing GATE_CPU_RC=.
  2. Optional, same push. Add a test for the "shares no commit with HEAD" arm.
  3. After the push, re-state the PR body's "exactly once" and grep GATE_CPU_RC= claims against a measurement.

… arm

The printed header named the verdict as `GATE_CPU_RC=<n>`, which put the
key on stdout twice and made a first-match reader land on the header. It
now says "the last line of stdout". A test asserts the key appears on
exactly one line of a run's output.

The `gpu: UNKNOWN` arm for an integration ref that shares no commit with
HEAD had no test; the git-checkout test now covers it with an orphan
branch as well as an unresolvable ref.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Developer round 1: head c4509d1d9 (one commit on top of 94bebb3b2, no force-push)

No blocking issues remain on my side.

Blocking (:65): fixed.

  • The header no longer spells GATE_CPU_RC=.
  • New pin test_the_verdict_key_is_printed_exactly_once. Mutation H1 (the old header restored) fails exactly that id, and null control N0 gives 10 passed.
  • The PR body's "exactly once" and grep claims are re-stated against a measurement: in the merged-tree gate log, 1 line contains the key.

Non-blocking (:171): fixed. A new parametrized case, test_a_git_checkout_is_not_called_unstamped[True-shares no commit with HEAD], uses an orphan ref. Your R3 mutation now fails that id.

Named result at the new tip e9d31f4bc (tip gate md5 ce05d5d5…, same file as before): 5 failed, 5 passed.

  • Failed: the three [pipe] ids and both [...] ids of the git-checkout test.
  • The exactly-once test passes at the tip, as it should: the duplicate key was this PR's own regression.

At head: 10 passed (gate md5 adb7e63d…).

Gate 1, the merged tree.

  • git merge-tree --write-tree e9d31f4bc c4509d1d9 gives 817f19398858b2f5fb3b8292349ed7d9ff6fae87, rc 0.
  • It was gated with its own gate (md5 adb7e63d…) on node 18. It was staged via an unpushed probe commit 22a58013a, run unpiped with timeout -k 10 1500 after two other gates drained, with atom.__file__ under the staged root.
  • Result: 5131 passed, 149 skipped, 3 xfailed. Script rc 0, last line GATE_CPU_RC=0 PASSED.
  • 5131 = 5123 (the tip) + 8 (this PR's ids). No failures, so no flaky-class re-run was needed.

Effort: 67 AST statements, against an estimate of 30–60. That is under the 2x stop.

Follow-up filed: #280, for gate_gpu.sh's bare GATE_GPU_RC=, its stderr-only reasons, and its pre-flight block last on stdout. It follows the brief shape and depends on this PR.

Both inline threads have replies. Staging on node 18 is removed.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review, cycle 2: APPROVE at head c4509d1d9c4f0cf8119e32691fd2da30d26cd315

No blocking issues. Both cycle-1 findings are closed and verified by measurement. I reviewed only the delta, 94bebb3b2..c4509d1d9: one commit, no force-push, with 1 line changed in gate_cpu.sh and the rest in test_gate_cpu_verdict_through_pipe.py.

All measurements were taken on node 18 in xiaobizh_n18_cpu, using a git archive tree staged at /tmp/pr276rev2/head/ATOM. atom.__file__ resolved to /tmp/pr276rev2/head/ATOM/atom/__init__.py, and the gate md5 was adb7e63d….

Finding 1 (header repeated the key): closed

The header now reads verdict: the last line of stdout; …. I counted lines containing GATE_CPU_RC= on piped runs over a throwaway tree:

path 2>/dev/null | grep -c 2>&1 | grep -c 2>&1 | grep -m1
pass (rc 0) 1 1 GATE_CPU_RC=0 PASSED
GPU trigger (rc 98) 1 1 GATE_CPU_RC=98 NOT PASSED -- … run gate_gpu.sh
-rE refusal (rc 95) — 1 —

test_the_verdict_key_is_printed_exactly_once counts both streams. It runs {gate} 2>&1, so stderr is folded into the stdout it counts. I checked that with a mutation of my own, R6: the new header stays as it is, and one line printing GATE_CPU_RC=<n> goes to stderr only. The new test fails by name. It is limited to the 98 path, which is fine, because the header line is common to every path.

Finding 2 (the "shares no commit" cause was untested): closed

The orphan is built with commit-tree HEAD^{tree} and no parent, so it is a real root commit with no history in common.

Mutations, each one line with the gate staying at 268 → 268 lines. The md5 was restored to adb7e63d… afterwards.

mutation result failing id
N0 null (comment word) 10 passed —
R3 "shares no commit" text → "resolves to no commit" 1 failed test_a_git_checkout_is_not_called_unstamped[True-shares no commit with HEAD]
R5 old header restored 1 failed test_the_verdict_key_is_printed_exactly_once
R6 the key printed on stderr only 1 failed test_the_verdict_key_is_printed_exactly_once

At head, the 8 tests in test_gate_cpu_verdict_through_pipe.py and the 2 in test_gate_cpu_pipe_identity.py give 10 passed.

Follow-up

gate_gpu.sh is filed as #280 (open). That matches the cycle-1 ruling.

Gate 1 on the tree that will land

The inline threads from cycle 1 have been answered by the developer, and the fixes above verify them.

@jgong5
jgong5 marked this pull request as ready for review September 23, 2026 11:35
@jgong5
jgong5 merged commit 019eb90 into feature/atomcompass_new Sep 23, 2026
jgong5 added a commit that referenced this pull request Sep 23, 2026
… pipe

`gate_gpu.sh 2>/dev/null | tail -6` kept the after-run pre-flight readout
and a bare `GATE_GPU_RC=1`: every reason went to stderr, and the pre-flight
block is printed to stdout just before the verdict. A pipeline's status is
tail's, so nothing said the run had not passed or why.

The verdict line now reads `GATE_GPU_RC=0 PASSED` or
`GATE_GPU_RC=<n> NOT PASSED -- <reason>`, the shape gate_cpu.sh took in
#276. Every non-zero finish passes its own reason. fail() keeps the
one-line form of the first finding and counts the rest, so the aggregate
exit names it, for example "1 failure(s) the baseline does not name,
first <node-id>; 2 more finding(s) on stderr". A header line names the
last line of stdout as the verdict without spelling the key.

tests/compass/test_gate_gpu_verdict_through_pipe.py runs the real script
over throwaway trees with a stub pre-flight and a generated suite sized
from the script's own baseline constants. It reaches all twelve exits but
the unreachable `cd` one, including a PASSED run, with no GPU.

Closes #280

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
jgong5 added a commit that referenced this pull request Sep 23, 2026
… pipe (#294)

* compass(gates): gate_gpu.sh verdict line carries its reason through a pipe

`gate_gpu.sh 2>/dev/null | tail -6` kept the after-run pre-flight readout
and a bare `GATE_GPU_RC=1`: every reason went to stderr, and the pre-flight
block is printed to stdout just before the verdict. A pipeline's status is
tail's, so nothing said the run had not passed or why.

The verdict line now reads `GATE_GPU_RC=0 PASSED` or
`GATE_GPU_RC=<n> NOT PASSED -- <reason>`, the shape gate_cpu.sh took in
#276. Every non-zero finish passes its own reason. fail() keeps the
one-line form of the first finding and counts the rest, so the aggregate
exit names it, for example "1 failure(s) the baseline does not name,
first <node-id>; 2 more finding(s) on stderr". A header line names the
last line of stdout as the verdict without spelling the key.

tests/compass/test_gate_gpu_verdict_through_pipe.py runs the real script
over throwaway trees with a stub pre-flight and a generated suite sized
from the script's own baseline constants. It reaches all twelve exits but
the unreachable `cd` one, including a PASSED run, with no GPU.

Closes #280

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* compass(gates): a fail() reason left out names the finding, not set -u

fail() read its one-line reason as a bare $2 under `set -u`, so a call
that left it out aborted the gate with no GATE_GPU_RC line at all. It now
defaults to the finding's own first line. The three reasons that repeated
that line word for word are dropped, and the comment above finish() now
points at gate_cpu.sh's instead of restating it.

The test gains a pass-count case, where the pass count is the only
finding, and asserts that a verdict does not end on a colon: a reason cut
at its first line would. The throwaway tree carries a stub torch, since
the gate only reads its version, which takes the file from ~40 s to ~12 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* compass(gates): pin fail()'s default reason with a --no-summary run

A caller's --no-summary reaches pytest, which then prints no FAILED lines
while its counts still say 5 failed. The FAILED-line check becomes the
first finding, and it has no explicit reason, so the verdict depends on
fail()'s default. A no-fail-lines case now drives that run; without the
default it ends on `set -u` with no GATE_GPU_RC line.

The aggregate verdict drops a trailing period from the first reason
before appending "; N more finding(s)", so it no longer reads "printed.;".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant