Repository navigation
compass(gates): gate_cpu.sh verdict line carries its reason through a pipe - #276
Conversation
… pipe A pipeline's status is its last command's, so `gate_cpu.sh | tail -6` reports tail's 0 on a run that exited 98. No script can change the status its caller's shell reports; what it controls is the text. The final line, GATE_CPU_RC=<n>, now says PASSED or NOT PASSED with the reason, so every pipe that keeps the verdict keeps why -- `2>/dev/null | tail -6` used to leave a bare 98 under pytest's green summary. A printed header line names the last line as the verdict. Refusing to run when stdout is a pipe was weighed and rejected: under `docker exec` on the CPU node stdout is a FIFO too (measured), so every normal run would need an override. The `gpu: UNKNOWN` line now states which of its three causes held. A git checkout whose integration ref does not resolve, or shares no commit with HEAD, is no longer told it was never stamped. Closes #191 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| fi | ||
| exit "$1" | ||
| } | ||
| printf 'verdict: the last line, GATE_CPU_RC=<n>; if this run is piped, $? is the pipe'"'"'s\n' |
There was a problem hiding this comment.
Required (principles 8 and 6): this header puts the verdict key on stdout a second time, so the verdict is no longer printed exactly once.
The literal GATE_CPU_RC= now appears on two stdout lines of every run: this header, which is always the first line, and the real verdict, which is the last line. Measured on node 18 (xiaobizh_n18_cpu), with the head's gate over a throwaway green tree:
$ bash scripts/compass/gate_cpu.sh 2>/dev/null | grep -c 'GATE_CPU_RC='
2
$ bash scripts/compass/gate_cpu.sh 2>&1 | grep -m1 'GATE_CPU_RC='
verdict: the last line, GATE_CPU_RC=<n>; if this run is piped, $? is the pipe's
Three things this contradicts:
- The file's own contract at line 24: "GATE_CPU_RC is printed exactly once, on every path". Lines 28-31 record the first draft's failure, "anything grepping for the first answer got the wrong one", and this line brings that failure back for a first-match reader.
- The PR body: "It is still printed exactly once on every path".
- The PR body: "
grep GATE_CPU_RC=... still work". An unanchoredgrep GATE_CPU_RC=now returns two lines.
Impact today is small. It cannot produce a false pass, because the header has <n> where a digit would be. Every consumer I found reads the value in a way that skips the header: grep -oE 'GATE_CPU_RC=[0-9]+' | tail -1, re.search(r"GATE_CPU_RC=(\d+)"), and grep -m1 '^GATE_CPU_RC=', all under agent_scratch/, and nothing in the repo parses it. But the invariant is the script's stated output contract, and the fix is one line.
Fix: state the verdict line without the KEY= spelling. For example:
printf 'verdict: the last line of stdout; if this run is piped, $? is the pipe'"'"'s\n'
Also consider a pinning assertion in test_gate_cpu_verdict_through_pipe.py that the unpiped stdout has exactly one line containing GATE_CPU_RC=. Nothing fails today if a second one appears.
There was a problem hiding this comment.
Fixed in c4509d1d9. The header now reads verdict: the last line of stdout; if this run is piped, $? is the pipe's, which does not spell the key.
New pin: test_the_verdict_key_is_printed_exactly_once asserts that exactly one line of gate 2>&1 output contains GATE_CPU_RC=. Evidence, node 18:
- At head: 10 passed.
- Mutation H1 (header restored to
verdict: the last line, GATE_CPU_RC=<n>;, 268 → 268 lines): 1 failed,test_the_verdict_key_is_printed_exactly_once. - Null control N0: 10 passed.
- Merged-tree gate log: one line containing the key.
The PR body's "exactly once" and grep claims have been re-stated against that measurement.
| if ! git -C "$ROOT" rev-parse --git-dir >/dev/null 2>&1; then | ||
| GPU_SRC_UNKNOWN="this tree was never stamped (no .git, no .compass-changed); a staging omission, see snapshot.sh" | ||
| elif REF=$(compass_resolve_ref "$ROOT" "$INTEGRATION"); then | ||
| GPU_SRC_UNKNOWN="a git checkout with no .compass-changed, and $REF shares no commit with HEAD" |
There was a problem hiding this comment.
Non-blocking (principle 8): no test reaches this arm, one of the three causes the PR says it now names.
The mutation below survives the whole new file and test_gate_cpu_pipe_identity.py: 8 passed, and the file stays at 268 lines.
- GPU_SRC_UNKNOWN="... and $REF shares no commit with HEAD"
+ GPU_SRC_UNKNOWN="... and $REF resolves to no commit here"
test_a_git_checkout_is_not_called_unstamped only builds a checkout where the ref does not resolve.
I checked the arm by hand, and it is correct. I built a git checkout with an orphan feature/atomcompass_new that shares no commit with master, and ran it on node 18:
gpu: UNKNOWN -- a git checkout with no .compass-changed, and feature/atomcompass_new shares no commit with HEAD
GATE_CPU_RC=98 NOT PASSED -- unknown whether this diff needs the GPU tier: ...
Given that you are already pushing for the header line, a third case in the same test file would close the gap. It needs about 3 extra git calls: checkout --orphan feature/atomcompass_new; commit; checkout master.
There was a problem hiding this comment.
Covered in c4509d1d9. test_a_git_checkout_is_not_called_unstamped is now parametrized over two cases:
[False-resolves to no commit]: an unresolvable ref, as before.[True-shares no commit with HEAD]: an orphanfeature/atomcompass_newmade withgit commit-tree HEAD^{tree}plusgit branch, so there is no checkout switching.
Your mutation R3, applied at head (268 → 268 lines), now gives 1 failed: test_a_git_checkout_is_not_called_unstamped[True-shares no commit with HEAD]. The null control gives 10 passed.
Review, cycle 1: REQUEST_CHANGES at head
|
| mutation | result | failing ids |
|---|---|---|
| M0 null (comment word) | 8 passed | — |
M1 bare GATE_CPU_RC=%s |
3 failed | the three [pipe] ids |
M2 98 reason loses gate_gpu.sh |
3 failed | the three [pipe] ids |
M3 if ! git … → if true |
1 failed | test_a_git_checkout_is_not_called_unstamped |
| M4 UNKNOWN printf back to the tip's literal | 1 failed | test_a_git_checkout_is_not_called_unstamped |
| R1 (mine) verdict printf sent to stderr | 2 failed | `[ |
R2 (mine) -eq 0 → -ne 1, so a 98 reads PASSED |
3 failed | the three [pipe] ids |
R4 (mine) a done line printed after the verdict |
3 failed | the three [pipe] ids |
| R3 (mine) "shares no commit" text swapped for "resolves to no commit" | 8 passed: survives | — (inline on :171) |
3. Three-case check: branch gate, branch's own staged tree, unpiped, timeout -k 10 1800
The runs used /tmp/pr276rev/head/ATOM, commit stamp 94bebb3b2, and gate md5 986c8801….
| case | how | script rc | pytest | last line | ^GATE_CPU_RC= lines |
|---|---|---|---|---|---|
| 0 | as staged | 0 | 5128 passed, 149 skipped, 3 xfailed | GATE_CPU_RC=0 PASSED |
1 |
| 98 | atom/model_engine/model_runner.py appended to .compass-changed (trigger line 53) |
98 | 5128 passed | GATE_CPU_RC=98 NOT PASSED -- this diff needs the GPU tier, which has not run; run gate_gpu.sh |
1 |
| 1 | tests/test_zz_pr276_forced_fail.py with assert False |
1 | 1 failed, 5128 passed | GATE_CPU_RC=1 NOT PASSED -- pytest failed; the FAILED lines above name the tests |
1 |
- The stamp was restored afterwards, confirmed with
cmp. - For the case-98 run,
gpu:readREQUIRED (.compass-changed stamp). - I also ran two small probes on the head's gate:
- the
-rrefusal path gives rc 95, last lineGATE_CPU_RC=95 NOT PASSED -- refused a -r argument; nothing was run; - the third UNKNOWN arm (an orphan integration ref) gives rc 98 and names "shares no commit with HEAD".
- the
4. == → startswith in test_gate_cpu_pipe_identity.py: no weakening
The line only finds where the RC != 0 block ends, and the change was needed. With ==, the new finish "$RC" "pytest failed; …" line never matches. The loop then runs past the block and collects the later printfs, so the window test would assert against the wrong text.
With startswith, the only line that can match inside the block is a finish "$RC"… call, which is exactly the block's exit. The two assertions (the block still exists, and the class name and README path sit in the last 5 lines) are unchanged. Both pass at head.
5. Follow-up ruling on gate_gpu.sh: yes, it needs its own issue. It does not block this PR.
gate_gpu.sh has the same shape, and it is worse through 2>/dev/null:
finish()prints a bareGATE_GPU_RC=%s, from:57at8ca066c24.- Every failure reason goes to stderr through
fail(). That covers new failures, gone baseline failures, the pass-count mismatch, and errors. - The last stdout lines before a
finish 1are the===== pre-flight (after) =====block.
So gate_gpu.sh 2>/dev/null | tail -6 shows a pre-flight readout and then a bare GATE_GPU_RC=1. That is exactly the "left bare" defect #191 named, and it now sits on the gate that #191's own 98 message tells the reader to run.
The fix is the same one-function change, with a reason argument on each of its ~9 finish calls. It is outside #191's file set, so it should not be folded in here. I recommend the developer file it at handoff. No issue for it exists yet: I searched GATE_GPU_RC and gate_gpu.sh pipe.
6. No design-doc references in the added lines
I scanned the added lines for D\d+, P\d.\d, "principle", Gate \d, doc numbers and design/. There were no hits.
Gate 1 on the tree that will land
- The tip moved to
6d22c3716(compass(tests): one recursive walk for every spec site check #2729f2df4227, compass(capture): stub only the torch.cuda names a capture reads #2706d22c3716), and both PRs record a node-id delta of 0. git merge-tree --write-tree 6d22c3716 94bebb3b2gives8ca066c247949b04dbc896c216835ef737350711, rc 0, with no conflicts.- I gated it once with its own gate (md5
986c8801…). The tree was staged bygit archiveof a local probe commit,a87de7d7b(tree8ca066c24, parents the tip and the head), and was never pushed..compass-commitwasa87de7d7b..compass-changedheld the 3 files of this PR.atom.__file__was/tmp/pr276rev/merged/ATOM/atom/__init__.py.
- Result: 5128 passed, 149 skipped, 3 xfailed, pytest rc=0, script rc=0, last line
GATE_CPU_RC=0 PASSED,gpu: not required (.compass-changed stamp). - Decomposition: 5122 (the tip's, since compass(capture): stub only the torch.cuda names a capture reads #270 and compass(tests): one recursive walk for every spec site check #272 each add 0) + 6 (the new file's ids) = 5128, which matches exactly. There were no failures, so no flaky-class re-run was needed.
- Environment: the runs were sequential, unpiped and bounded, and all staging was removed afterwards.
- A
pr275revgate from another agent was running when I arrived, and I waited for it to finish. - A
pgrepat the start of my merged run matched other processes, which may have included another gate. The count matched exactly regardless.
- A
What the next push needs
- Required. Reword the header at
:65so thatGATE_CPU_RC=appears on stdout only on the verdict line. Consider an assertion that the unpiped stdout has exactly one line containingGATE_CPU_RC=. - Optional, same push. Add a test for the "shares no commit with HEAD" arm.
- After the push, re-state the PR body's "exactly once" and
grep GATE_CPU_RC=claims against a measurement.
… arm The printed header named the verdict as `GATE_CPU_RC=<n>`, which put the key on stdout twice and made a first-match reader land on the header. It now says "the last line of stdout". A test asserts the key appears on exactly one line of a run's output. The `gpu: UNKNOWN` arm for an integration ref that shares no commit with HEAD had no test; the git-checkout test now covers it with an orphan branch as well as an unresolvable ref. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Developer round 1: head
|
Review, cycle 2: APPROVE at head
|
| path | 2>/dev/null | grep -c |
2>&1 | grep -c |
2>&1 | grep -m1 |
|---|---|---|---|
| pass (rc 0) | 1 | 1 | GATE_CPU_RC=0 PASSED |
| GPU trigger (rc 98) | 1 | 1 | GATE_CPU_RC=98 NOT PASSED -- … run gate_gpu.sh |
-rE refusal (rc 95) |
— | 1 | — |
test_the_verdict_key_is_printed_exactly_once counts both streams. It runs {gate} 2>&1, so stderr is folded into the stdout it counts. I checked that with a mutation of my own, R6: the new header stays as it is, and one line printing GATE_CPU_RC=<n> goes to stderr only. The new test fails by name. It is limited to the 98 path, which is fine, because the header line is common to every path.
Finding 2 (the "shares no commit" cause was untested): closed
The orphan is built with commit-tree HEAD^{tree} and no parent, so it is a real root commit with no history in common.
Mutations, each one line with the gate staying at 268 → 268 lines. The md5 was restored to adb7e63d… afterwards.
| mutation | result | failing id |
|---|---|---|
| N0 null (comment word) | 10 passed | — |
| R3 "shares no commit" text → "resolves to no commit" | 1 failed | test_a_git_checkout_is_not_called_unstamped[True-shares no commit with HEAD] |
| R5 old header restored | 1 failed | test_the_verdict_key_is_printed_exactly_once |
| R6 the key printed on stderr only | 1 failed | test_the_verdict_key_is_printed_exactly_once |
At head, the 8 tests in test_gate_cpu_verdict_through_pipe.py and the 2 in test_gate_cpu_pipe_identity.py give 10 passed.
Follow-up
gate_gpu.sh is filed as #280 (open). That matches the cycle-1 ruling.
Gate 1 on the tree that will land
- Tip:
faf0d208b. Sincee9d31f4bc, it adds only compass(design): the stack pin warns by default, refuses under strict=True or on a transfer or merge #278, which touches three design docs. - Merge:
git merge-tree --write-tree faf0d208b c4509d1d9givesbd3fe089bc310b900a2da77f54838dd4ee2a2114, rc 0, with no conflicts. - Staging: a
git archiveof the local probe commit0f0636a48, which has that tree and the tip and head as parents, and was never pushed..compass-commitwas0f0636a48..compass-changedheld this PR's 3 files.atom.__file__was/tmp/pr276rev2/merged/ATOM/atom/__init__.py.
- Run: the tree's own gate (md5
adb7e63d…), unpiped, withtimeout -k 10 2400. - Result: 5131 passed, 149 skipped, 3 xfailed, pytest rc=0, script rc=0, last line
GATE_CPU_RC=0 PASSED,gpu: not required (.compass-changed stamp). - Decomposition:
- The tip
5828346b6had 5122. - compass(spec): every Rule member is named by at least one refusal site #275 added +1: its dev record gives 5122 → 5123, added
test_every_rule_is_named_by_at_least_one_site. - compass(capture): stub only the torch.cuda names a capture reads #270, compass(tests): one recursive walk for every spec site check #272 and compass(docs): say which imports reach rocminfo, and describe the GPU gate instead of citing its lines #274 added 0 each, and compass(design): the stack pin warns by default, refuses under strict=True or on a transfer or merge #278 is docs only.
- This PR adds +8: the 6 from cycle 1, plus
test_the_verdict_key_is_printed_exactly_once, plus the new[True-…]parameter. - So 5122 + 1 + 8 = 5131, which matches exactly and agrees with the developer's count on
e9d31f4bc.
- The tip
- Environment: I waited for another agent's gate (
i219gates) to finish before starting. Two other gates (p279rev,i219gates) started while mine was running. There were no failures, so no flaky-class re-run was needed.
The inline threads from cycle 1 have been answered by the developer, and the fixes above verify them.
… pipe `gate_gpu.sh 2>/dev/null | tail -6` kept the after-run pre-flight readout and a bare `GATE_GPU_RC=1`: every reason went to stderr, and the pre-flight block is printed to stdout just before the verdict. A pipeline's status is tail's, so nothing said the run had not passed or why. The verdict line now reads `GATE_GPU_RC=0 PASSED` or `GATE_GPU_RC=<n> NOT PASSED -- <reason>`, the shape gate_cpu.sh took in #276. Every non-zero finish passes its own reason. fail() keeps the one-line form of the first finding and counts the rest, so the aggregate exit names it, for example "1 failure(s) the baseline does not name, first <node-id>; 2 more finding(s) on stderr". A header line names the last line of stdout as the verdict without spelling the key. tests/compass/test_gate_gpu_verdict_through_pipe.py runs the real script over throwaway trees with a stub pre-flight and a generated suite sized from the script's own baseline constants. It reaches all twelve exits but the unreachable `cd` one, including a PASSED run, with no GPU. Closes #280 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… pipe (#294) * compass(gates): gate_gpu.sh verdict line carries its reason through a pipe `gate_gpu.sh 2>/dev/null | tail -6` kept the after-run pre-flight readout and a bare `GATE_GPU_RC=1`: every reason went to stderr, and the pre-flight block is printed to stdout just before the verdict. A pipeline's status is tail's, so nothing said the run had not passed or why. The verdict line now reads `GATE_GPU_RC=0 PASSED` or `GATE_GPU_RC=<n> NOT PASSED -- <reason>`, the shape gate_cpu.sh took in #276. Every non-zero finish passes its own reason. fail() keeps the one-line form of the first finding and counts the rest, so the aggregate exit names it, for example "1 failure(s) the baseline does not name, first <node-id>; 2 more finding(s) on stderr". A header line names the last line of stdout as the verdict without spelling the key. tests/compass/test_gate_gpu_verdict_through_pipe.py runs the real script over throwaway trees with a stub pre-flight and a generated suite sized from the script's own baseline constants. It reaches all twelve exits but the unreachable `cd` one, including a PASSED run, with no GPU. Closes #280 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * compass(gates): a fail() reason left out names the finding, not set -u fail() read its one-line reason as a bare $2 under `set -u`, so a call that left it out aborted the gate with no GATE_GPU_RC line at all. It now defaults to the finding's own first line. The three reasons that repeated that line word for word are dropped, and the comment above finish() now points at gate_cpu.sh's instead of restating it. The test gains a pass-count case, where the pass count is the only finding, and asserts that a verdict does not end on a colon: a reason cut at its first line would. The throwaway tree carries a stub torch, since the gate only reads its version, which takes the file from ~40 s to ~12 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * compass(gates): pin fail()'s default reason with a --no-summary run A caller's --no-summary reaches pytest, which then prints no FAILED lines while its counts still say 5 failed. The FAILED-line check becomes the first finding, and it has no explicit reason, so the verdict depends on fail()'s default. A no-fail-lines case now drives that run; without the default it ends on `set -u` with no GATE_GPU_RC line. The aggregate verdict drops a trailing period from the first reason before appending "; N more finding(s)", so it no longer reads "printed.;". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Closes #191
No blocking issues.
What changed
scripts/compass/gate_cpu.shonly (+51/-24), plus tests._lib.shdid not need changing.The verdict line carries its reason. The last line is now
GATE_CPU_RC=0 PASSEDorGATE_CPU_RC=<n> NOT PASSED -- <reason>. Every non-zerofinishpasses its own reason, so the 98 for a GPU-tier trigger reads... run gate_gpu.sh. TheGATE_CPU_RC=<n>prefix is unchanged. The key is printed on exactly one line of a run's output (stdout and stderr together), measured at the merged tree:grep -c GATE_CPU_RC=over the log = 1.test_the_verdict_key_is_printed_exactly_oncepins that. At94bebb3b2it was 2, because the header named the key; the reviewer caught it, and the header was reworded inc4509d1d9.A printed header line names the verdict:
verdict: the last line of stdout; if this run is piped, $? is the pipe's. It does not spell the key, so a first-match reader cannot land on it.gpu: UNKNOWNstates which of its three causes held..git: "never stamped ... a staging omission, see snapshot.sh", as before.The stderr block that explains the 98 uses the same text.
tests/compass/test_gate_cpu_pipe_identity.py: one line changed. Its block-end matcher was== 'finish "$RC"'and is nowstartswith, because that call now carries a reason. Its two assertions are unchanged and still pass.Mechanism chosen, and why
A script cannot change the exit status its caller's pipeline reports. That status belongs to the caller's shell (
pipefailor not). Every option in the brief therefore reduces to one of two things: what text survives, or whether the gate runs at all.ssh node18 'docker exec xiaobizh_n18_cpu bash -c "[ -p /dev/stdout ] && echo fifo"'the answer isfifo(pipe:[3234310976]). That is the standard way this project runs the gate. The same holds for pytest'ssubprocess.run(capture_output=True)and for any agent harness.| tailfrom the callers that do keep$?. Every normal run would need an override, and an override set by habit lets| tailthrough with it.gate_cpu.sh | tail -6that refuses still reports tail'src=0.set -o pipefailinside the script: already present (line 21). It only covers pipes inside the script.tail -Nkeeps the last line.GATE_CPU_RC=98(test output below).2>/dev/null | tail -6kept it bare under1 passedandpytest: rc=0, with nothing saying why it is not a pass.$?of a pipeline is still wrong, and nothing in the script can prevent that. The test pins that premise too: it asserts the pipeline's rc is 0, so if some future mechanism changes that, the test says to re-check.Named result (gate 3), node 18
xiaobizh_n18_cputests/compass/test_gate_cpu_verdict_through_pipe.pyruns the realgate_cpu.shover a throwaway tree (a one-test suite, a trigger list the stamped diff touches), so the run is green and then exits 98.At the tip
5828346b6, with the tip's gate and this test file dropped in: 4 failed, 4 passed.94bebb3b2...::test_unpiped_the_gate_exits_98(control)...::test_a_piped_run_still_reads_as_not_passed_and_says_why[| tail -6]['', '.', '1 passed in 0.10s', '', 'pytest: rc=0', 'GATE_CPU_RC=98']...[2>&1 | tail -6]...[2>/dev/null | tail -6]GATE_CPU_RC=98under1 passed...::test_a_git_checkout_is_not_called_unstamped...::test_a_tree_with_no_git_and_no_stamp_is_still_called_unstamped(control)At head: 8 passed, including the 2 in
test_gate_cpu_pipe_identity.py.Mutation battery at head. Every mutation is one line changed, and the file stays at 268 lines:
GATE_CPU_RC=%s[pipe]idsgate_gpu.sh[pipe]idsif ! git ... --git-dirchanged toif truetest_a_git_checkout_is_not_called_unstampedgpu: UNKNOWNprintf back to the tip's literal texttest_a_git_checkout_is_not_called_unstampedThe gate file's md5 was verified restored (
986c8801...) after the battery.Gate 1: ATOM suite, unmodified, as a delta
Each side was run with its own gate from its own
git archivetree, staged in the container at a path of this task's own..compass-commitand.compass-changedwere written, andatom.__file__was asserted under the staged root. Runs were sequential, unpiped and bounded withtimeout -k 10 1500.The two sides are different instruments: tip gate md5
ce05d5d5..., branch gate md5986c8801.... The only difference is the verdict text and the UNKNOWN line; the pytest invocation and the exclusion list are byte-identical.5828346b6, tip gate)94bebb3b2, branch gate)GATE_CPU_RC=0GATE_CPU_RC=0 PASSEDtests/compass/test_gate_cpu_verdict_through_pipe.py, none removed; collect-only ontests/compassgives 1166 → 1172.scripts/compass/gate_cpu.sh) +51 −24. Tests +110 −1: the new file is 109 lines, and 1 line changed intest_gate_cpu_pipe_identity.py.git merge-tree --write-tree 5828346b6 94bebb3b2:bb11f9305ae87aee6288d5b9129a0c1535e2bb13, rc 0. That is the head's own tree, since the tip has not moved.The branch gate reports correctly in all three cases (branch tree, branch gate, unpiped):
GATE_CPU_RC=0 PASSED.compass-changed+atom/model_engine/model_runner.py(a trigger)GATE_CPU_RC=98 NOT PASSED -- this diff needs the GPU tier, which has not run; run gate_gpu.shtests/test_zz_i191_forced_fail.pyasserting FalseGATE_CPU_RC=1 NOT PASSED -- pytest failed; the FAILED lines above name the testsEach run printed
GATE_CPU_RCexactly once. In the failing case the flaky-class paragraph still sits directly above the verdict line, so2>&1 | tail -6keeps both. The GPU tier was not run: this branch touches no trigger, and the 98 was forced deliberately.Gate 2
import atombeyond an empty stub package.ruff checkandblack --checkare clean on both touched test files (RUFF_RC=0,BLACK_RC=0).tests/compass/as a whole is black-dirty on the tip in two unrelated files (test_pp_kv_shard_key.py,test_tier0_memory_goal_cites_empirical.py). That is not from this PR.shellcheckin either container.Effort
The brief estimated 30–60 AST statements of test, docstring
Exprnodes included. Actual: 58 at94bebb3b2(ast.walk, countingast.stmt); 67 atc4509d1d9(see Round 1).What surprised me
GATE_CPU_RC=last on stdout on every path. So, measured, the number did survive all three pipes. What the brief called "left bare" is the real defect: through2>/dev/null | tail -6the only surviving context was1 passedandpytest: rc=0, both green.docker execstdout is a FIFO. That single fact is what rules out the refusal option, and it had not been measured before.Left undone
gate_gpu.sh: same bare-verdict shape, filed as gate_gpu.sh: a piped run shows a pre-flight readout and a bare GATE_GPU_RC=1; the reason is lost #280.README.mdstill says "do not pipe the gate", which is still right. It is not touched here because compass(docs): five places still describe the aiter import the GPU gate no longer does #273 is editing it..compass-changedstill records no base (the existing comment names this). Out of scope.Round 1 (head
c4509d1d9): addressing review cycle 1The blocking finding is fixed. The header no longer spells
GATE_CPU_RC=. The new testtest_the_verdict_key_is_printed_exactly_onceasserts that exactly one line ofgate 2>&1output contains the key.The non-blocking finding is fixed.
test_a_git_checkout_is_not_called_unstampedis now parametrized over two cases:[False-resolves to no commit]: the integration ref is not resolvable.[True-shares no commit with HEAD]: an orphanfeature/atomcompass_new, built withgit commit-tree.Named result at the new tip
e9d31f4bc, using the tip's gate (md5ce05d5d5…, the same file as at5828346b6) with this test file dropped in: 5 failed, 5 passed.[pipe]ids and bothtest_a_git_checkout_is_not_called_unstamped[...]ids.At head: 10 passed, with gate md5
adb7e63d….Mutations at head (one line each, 268 → 268 lines; md5 restored to
adb7e63d…afterwards):verdict: the last line, GATE_CPU_RC=<n>;test_the_verdict_key_is_printed_exactly_oncetest_a_git_checkout_is_not_called_unstamped[True-shares no commit with HEAD]Gate 1 on the merged tree.
git merge-tree --write-tree e9d31f4bc c4509d1d9gives817f19398858b2f5fb3b8292349ed7d9ff6fae87, rc 0.git archiveof a local, unpushed probe commit22a58013a(tree817f19398, parents the tip and the head).adb7e63d…), on node 18xiaobizh_n18_cpu. The run was unpiped, bounded withtimeout -k 10 1500, and started after two other agents' gates had drained.atom.__file__was/tmp/i191b/merged/ATOM/atom/__init__.py.GATE_CPU_RC=0 PASSED, one line containing the key,gpu: not required.Lines at head vs
5828346b6:gate_cpu.sh: +51 −24, unchanged by this round, which reworded one line in place;Effort: 67 AST statements, against an estimate of 30–60. That is +12% over the top of the range, well inside the 2x stop. The 9 statements added this round cover the two review findings.
Follow-up filed: #280 (
gate_gpu.sh, the same bare-verdict shape).🤖 Generated with Claude Code