Repository navigation
Two classes counted as one: the action-kill ending reproduced (n=1), the O(files) growth measured, and the budget knob shown inert - #10091
Conversation
…able instead of narrowing its range gunbc.memory_stall_refusal carried two receipts bracketing GUNBC_MEMORY_BUDGET_BYTES from opposite sides -- one truthful declaration that thrashed, one overstated one that was killed -- and concluded that no good NUMBER exists on a 7 GiB runner. This adds the specimen that varies the number as its only input, and finds the knob is not connected to the quantity that ends the run at all: 2147483648 declared produced 6419416 kB resident, 1073741824 declared produced 6416112 kB. Five hundredths of one percent apart, at three and six times the declaration. That is stronger than an empty range, and it changes what a blocked lane should do. The override answers gunbc.host_budget_source's SOURCE question -- may a run begin -- and constrains nothing allocated afterwards, so setting it to clear a HostBudgetUnreadable refusal trades a loud, typed, immediate answer for an unbounded run: the escape hatch DESIGN section 5 forbids. The receipt also localizes the growth, off the floor-phase lines rather than off the ending. Preparation is bounded and invariant under both the roster flag and the budget -- the bare-reference edge-index warm completes at source_files=4475 for ~0.8 GiB in a completing run and in both dying arms alike, and every later preparation phase matches across all three. The unbounded growth is after preparation and across witness FILES, so the entry was never the variable: the same entry at the same budget with a single --function completes and returns zero. Two controls are recorded because each converts an absence into evidence. The guest OOM killer is LOUD here -- a child allocating past the cap prints Killed, the shell survives, an unconditional status echo reports 137, and the action's trailer prints -- so a log that simply stops is not an in-guest OOM. And a BuildBuddy action that completed emits a protocol terminator that a severed one does not, measured present on a genuine 255, a genuine 0 and the OOM control, and absent on a severed run; the wrapper's own status is transport-shaped and answers neither direction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6
… it is not what cuts the log Both roster arms were run to their own conclusion rather than stopped by a client-side cap, and both ended the same way: the guest OOM killer printed Killed naming the binary, the unconditional status echo reported 137, and the action's completion trailer printed. On both budgets, matching the synthetic control exactly. That refutes the explanation this investigation started from -- that the process tree is SIGKILLed and a killed process cannot report its own death. Exhausting this runner with the real workload returns a status and a complete log every time. So memory exhaustion is real on this path and is NOT the mechanism of a run whose log simply stops; the two failures are different, and a silent cut must stop being read as evidence of an out-of-memory condition. What destroys the action remains unestablished from outside the runner. The receipt now says that plainly rather than letting the exhaustion finding stand in for an answer it does not give. A third dispatch sampled /proc/PID/stat and places the specimen against this module's own verdict, which it traverses in both directions in a single run. Early: major faults flat at zero, user-CPU share 9473 bp -- correctly admitted as progress. Terminally, over an 86-second window: 391150 major faults for 272895/min against the 6000 threshold, user-CPU share collapsed to 340 bp against the 2000 floor -- both conjuncts hold, StallRefusedPageThrash. That is stronger evidence for the conjunction than either arm alone, and it vindicates the decision to measure USER cpu specifically: at the wall the process is busy, but the time is system time servicing its own faults, so a user-plus-system share would have read as healthy throughput and admitted the thrash. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6
…wo that cannot The receipt stopped at "not establishable from outside the runner", which is a legitimate terminal answer and an untracked stall if it names no next instrument (DESIGN 4b(2)). It now names one, and it is reachable today rather than aspirational: the executor's own server-side record of a SEVERED action, read for that action's terminal disposition and its worker's, set beside the client's truncated view of the same run. Every dispatch prints its invocation URL and the receipts above already cite invocations by UUID, so a severed run is addressable by identity -- nobody has simply fetched the server's side of one. That comparison is what separates the executor destroying the action from the client losing the stream: two causes, different owners, indistinguishable from the client alone. The acceptance half is stated so no lookalike retires it: the record must be for an invocation whose CLIENT log lacks the completion trailer, and must report that action's own terminal status. A record for a run that completed proves nothing about severance. Two neighbouring instruments are named only to be refused, so they are not proposed again. An in-guest supervisor cannot answer it by construction -- anything watching from inside dies with the guest, and this receipt demonstrates that limit from the other side, since the in-guest heartbeat survives a child SIGKILL well enough to establish the loudness finding and would report nothing if the action itself were destroyed. And a fast reproducer is BLOCKED rather than unbuilt: severance did not occur in any dispatch taken here, so there is no observed handle to shrink, and naming it as the next step would name work that cannot start. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6
…ipt's own detector was wrong A third run of the same subject ended a third way, and it overturns two claims this receipt made one commit ago. Both corrections are recorded rather than quietly replaced, because each was plausible and each was reached by generalising from the runs that existed at the time. FIRST: the silent signature reproduced. Two dispatches ended loudly -- guest OOM killer, Killed naming the binary, unconditional echo reporting 137, an ordinary exit-code line -- and the third ended with no exit-code line at all, "Action failed: signal: killed" in its place, the unconditional echo NEVER printing, and the wrapper returning 255. That is the signature this receipt was opened to explain. So the kill happens at two grains and only one is loud: the guest OOM killer selects a CHILD and that ending is fully reported; the executor kills the ACTION and that ending can report nothing from inside, because there is no inside left. The previous claim -- exhaustion here is loud and therefore is not what cuts a log -- generalised from the first two dispatches before the third existed. The silent cut IS memory-driven; the in-guest OOM killer is simply not what delivers it. The severing run thrashed sustainedly first: 8822792 major faults and 659 seconds of system time. SECOND: the trailer detector was wrong and is withdrawn. The completion trailer is emitted in EVERY case measured -- genuine nonzero exit, clean exit, child OOM kill, killed action alike -- so its presence proves nothing, and offering its absence as the severance signal was a mistake. The discriminator is the EXIT-CODE LINE: present and numeric when a command reached an ending, replaced by "Action failed: signal: killed" when the action was killed. A protocol-looking line can still be transport-shaped, and the way to find out is to build the failure on purpose and watch which lines survive it. The open question narrows accordingly: not what happens, but why the executor kills the action, which the guest cannot see. The server-side record instrument stands, with its acceptance half re-pinned to the corrected signature. The fast-reproducer refusal is withdrawn -- there is one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6
…an be fetched against Two additions, both aimed at the next reader rather than at the finding. The observation counts are now stated in the receipt because they are small: the loud ending is n=2 and the silent ending is n=1. That is enough to establish that the silent ending exists and to identify its grain, which is what the receipt claims. It is not enough to establish which conditions select between the two endings, so the sustained-thrash difference is written as the distinguishing feature of one specimen rather than as a rule. Every reversal recorded in this receipt came from treating a measured result as a replicated one, and the counts are the cheapest guard against the next reader repeating that. And the next-instrument trigger no longer requires reproducing the phenomenon before it can be used. The severed invocation is identified -- 94352bd4-e91b-440a-9a43-f7daa96ee1dd, whose client log carries "Action failed: signal: killed" and no exit-code line, which is exactly the acceptance half. What is missing is only access: the bb client here exposes no invocation-fetch subcommand, so the record must be read through the executor's API or web surface by someone holding that access. That is why this stays a trigger and not a task. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6
…erver, not the action A peer lane's independent n=6 shows two silent runs with NO trailer, which is the opposite of what the severing run here measured, and raised the question of whether the withdrawn trailer detector had been retired on a case that was not a killed action. Settled against the raw capture: the severing arm's log is the direct client redirect, never a fetched server view, and its trailer follows the "Action failed: signal: killed" line by five lines. The client outlived the action it was running, so the trailer genuinely prints for a killed action and the withdrawal stands. What their data does establish is that the taxonomy was one ending short. There are three: command reached an ending exit-code line PRESENT, trailer PRESENT executor killed the action no exit-code line, "Action failed: <signal>", trailer PRESENT client stream itself lost none of the three The third is real and was produced accidentally here by a caller-side timeout on the reading end, which is exactly what made an earlier draft mistake trailer-absence for action-severance. So a missing trailer means THE CLIENT did not finish -- a fact about the observer, not about the action -- while a present trailer proves nothing about whether any command ran. Only the exit-code line answers that, and it answers in both directions, which is why it is the discriminator and the trailer is not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6
…ld, the reported runs are all client-stream-lost The separating field decided it. The runs that motivated this investigation carry no "Action failed:" line, no signal or kill wording, and nothing about the runner at all -- they are the third ending, CLIENT STREAM LOST. The action-kill ending has exactly one confirmed instance and it is the arm taken here. That matters because the original signature was established over exit-code-line, marker and log-stops, and the second and third endings SHARE all three. The field that separates them was not among them, so two classes were counted as one and the merged class became the subject. The receipt now says where the merge was rather than inheriting the framing. The consequence is that the two endings must not inherit each other's support. The O(files) accumulation and the inert budget stand on instrumented arms and are independently corroborated. The action kill is real and n=1. And the ending that actually cost the reporting lanes their time is the client-stream-lost one, which this receipt does not explain and which is not a fact about this runner or this compiler at all: such a log establishes nothing about what happened on the far side, because the observation ended before the subject did. A lane holding one has measured a failed measurement, not a failed run, and the runner-side outcome is UNOBSERVED rather than bad. Recorded with it is the trap that produced the merge, because this receipt was caught by it too: a supplied cause overwrites an accurate agnostic reading. "Control never returned and the stream was cut" was correctly agnostic between the endings; "a killed process cannot report its own death" was then accepted for runs where it had never been established, having been established for a different run entirely. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6
…o a standing consequence Two corrections to the correction, both aimed at not repeating the defect that produced the merged class in the first place. The count is now exact. Client-stream-lost is n=2 CONFIRMED -- two runs actually graded on the separating field. Earlier runs from the same lane were described before that field existed and have not been re-examined against it, so they are UNCLASSIFIED CANDIDATES, not members, and joining them would mean re-grading each. The previous wording said the reported occurrences were all the third ending, which asserted a grading that had only been performed on two of them. An inflated count is what produced the merge; a correction carrying the error's own defect is not a correction. And the operative sentence is promoted from background to a standing consequence, because it is actionable now and does not wait on any explanation of why a client stream is lost -- an explanation that may never arrive. A lane holding such a log has not measured a failed run; it has measured a failed measurement, and the runner-side outcome is UNOBSERVED rather than bad. The opposite reading was in live use, with lanes treating these as runner failures and drawing conclusions about their own branches from them. Where such a conclusion has survived, the receipt now says to check WHY before keeping it: a two-arm comparison in which both arms hit this ending is still sound, since a control does not need the phenomenon named, but that soundness is a property of the comparison's design and never of the log's readability. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6
The floor red is not this PR — it reproduces on main
The floor refused with zero failures. From this run's own footer: One witness crossed the 500ms CPU line by 2ms (0.4%) and went undecided. Nothing failed. Baseline check — the same failure occurs on main, without this diff:
Each shows Why it cannot be this PR. The whole diff is The underlying condition, visible in the same log: that witness family runs at 455–465ms against a 500ms line — roughly 91% of budget — with its siblings already flagged Rerunning the failed job once the in-progress lane finishes; the red is variance, not content. — sent from proud-bear-215 |
Scope, restated — this began as two classes counted as one
The item was opened on several silent runs reported as one signature, established over exit-code line, marker, and log stops. The second and third endings below share all three; the field that separates them (
Action failed:) was not among them. Graded on it afterwards, the runs that motivated the item are all the third ending — client stream lost — with noAction failedline, no signal or kill wording, and nothing about the runner at all.So the two endings have opposite evidential standing and must not inherit each other's support:
Standing consequence — actionable now, and it does not wait on an explanation
A log that ends without an exit-code line and without an
Action failed:line is not evidence of a failed run. The observation ended before the subject did, so the runner-side outcome is unobserved, not bad. The opposite reading has been in live use — lanes treating these as runner failures and drawing conclusions about their own branches from them. Where such a conclusion has survived, check why before keeping it: a two-arm comparison in which both arms hit this ending is still sound, since a control does not need the phenomenon named to do its job — but that soundness is a property of the comparison's design, never of the log's readability.The trap that produced the merge is recorded too, because this PR was caught by it: a supplied cause overwrites an accurate agnostic reading. "Control never returned and the stream was cut" was correctly agnostic; "a killed process cannot report its own death" was then accepted for runs where it had never been established, having been established for a different run entirely.
What was asked
Explain why
.dagwitness execution dies silently on a BuildBuddy runner — exit 255, log cut mid-stream, no returned status even withRUST_BACKTRACEdemonstrably live — with a discriminating arm behind the explanation, or an honest "cannot be established from outside the runner" naming what was ruled out.Answer
The signature is reproduced, and the kill happens at two grains — only one of which is loud.
Three dispatches of the same subject, each run to its own conclusion rather than stopped by a client cap:
Action failed0)137signal: killedThe third is the signature this item was opened on: no status, log cut, wrapper 255, and
RUST_BACKTRACEwith nothing to print — because control never returned to a shell that no longer existed. The guest's OOM killer selects a child and that ending is fully reported; the executor kills the action and that ending can report nothing from inside, because there is no inside left. The severing run thrashed sustainedly first: 8,822,792 major faults, 659s of system time.So the silent cut is memory-driven — the in-guest OOM killer is simply not what delivers it.
Two corrections this PR makes against its own earlier commits
Both are kept in the receipt rather than quietly replaced, because each was plausible and each came from generalising over the runs that existed at the time.
Action failed: signal: killedby five lines — the client outlived the action. So the withdrawal stands, and their data adds the missing ending:Action failedsignal: killedThe third is real and I produced it accidentally, via a caller-side timeout on the reading end — which is exactly what made me mistake trailer-absence for action-severance. So a missing trailer is a fact about the observer, not the action; a present trailer proves nothing about whether a command ran.
The wrapper's status is transport-shaped in both directions and is never the answer:
0over an action whose payload died at 137 (correct shell semantics — a trailingechosucceeded),255over the killed action.Draw counts, stated because they are small
The loud ending is n=2; the silent ending is n=1. That establishes the silent ending exists and identifies its grain — which is what this PR claims — and does not establish which conditions select between the two endings. The sustained-thrash difference is written as the distinguishing feature of one specimen, not as a rule. Every reversal above came from treating a measured result as a replicated one.
What is still open
Not what happens but why the executor kills the action — a health check, a ceiling above the guest's own, or a supervisor timeout are all consistent with what the guest can see, and the guest cannot tell them apart. The deciding instrument is the executor's server-side record of a severed action, reachable today by the invocation URL every dispatch prints. Acceptance half: the record must be for an invocation whose client log carries
Action failed: signal: killedand no exit-code line — and the severed invocation is already identified (94352bd4-e91b-440a-9a43-f7daa96ee1dd), so the trigger does not require reproducing the phenomenon again. What is missing is only access: thebbclient here exposes no invocation-fetch subcommand, so the record must be read through the executor's API by someone holding it. That is why this stays a trigger, not a task. An in-guest supervisor is refused by construction — this PR shows that limit from both sides, since the sampler survived a child SIGKILL well enough to establish the loud ending and reported nothing at all about the action's own death.The budget knob is inert
Where the memory actually goes
Read off the
[floor-phase]lines rather than inferred from the ending. Preparation is bounded and invariant under both the roster flag and the budget:rss_growthsource_files--function(completed, rc=0)--roster-from-discovery@ 2 GiB (died)--roster-from-discovery@ 1 GiB (died)Every later preparation phase matches across all three (
modules=3155,declarers=17,declarer-closure sources=18,closure-strict-resolve sources=18). The unbounded growth is after preparation and across witness files:--roster-from-discoverywalks every discovered witness file in the corpus, and the resident set rises file after file without returning.So the entry was never the variable — the same entry at the same budget with a single
--functioncompletes and returns zero. This also reconciles the two lanes: snappy-koi-879's observation that the edge-index warm runs before any roster narrowing is correct, but the warm completes at ~0.8 GiB and is not the killer; theirHostBudgetUnreadablewas a different failure that fires in the first phase regardless of roster, so their narrowing was never reached.Two controls, each converting an absence into evidence
The guest OOM killer is loud. A child deliberately allocating past the cap prints
Killed, the shell survives, an unconditional status echo reports137, the end marker prints, and the action's trailer prints. A child SIGKILL does not sever the action — so a log that simply stops is not an in-guest OOM. The action itself is being destroyed.A completed action emits a protocol terminator; a severed one does not. BuildBuddy prints the command's exit code line followed by
Remote run completed at <ts>. Measured present on a genuine exit 255, present on a genuine exit 0, present on the OOM control, and absent on a severed run. This is the detector to use, and it needs nothing from the payload.One correction to the "the wrapper fails toward green" framing: a wrapper exit of
0is sometimes correct shell semantics, not a bug — in the OOM control the payload died at 137 and the script still exited 0 because the trailing echo succeeded. The defect is specifically the severed case, where no trailer exists and the status is arbitrary. The earlier framing would have sent someone hunting a bug in the wrong layer.The stall names its instrument
"Not establishable from outside the runner" is a legitimate terminal answer and an untracked stall if it stops there. The deciding instrument is the executor's own server-side record of a severed action — its terminal disposition and its worker's — set beside the client's truncated view of the same run. It is reachable today and simply was not taken: every dispatch prints its invocation URL, and the receipts already in this module cite invocations by UUID, so a severed run is addressable by identity. That comparison is what separates the executor destroyed the action from the client lost the stream — two causes with different owners that the client's view cannot tell apart.
Acceptance half, so no lookalike retires it: the record must be for an invocation whose client log lacks the trailer, and must report that action's own terminal status.
Two neighbours are named only to be refused. An in-guest supervisor cannot answer it by construction — anything watching from inside dies with the guest; this PR shows that limit from the other side, since the in-guest heartbeat survives a child SIGKILL well enough to establish the loudness finding and would report nothing if the action were destroyed. A fast reproducer is blocked, not unbuilt: severance did not occur in any dispatch taken here, so there is no observed handle to shrink.
Where the specimen lands against the module's own verdict
A third dispatch sampled
/proc/PID/stat. The run traverses both arms ofmemory_stall_verdictin a single execution:ProgressUnderMemoryAdmissibleStallRefusedPageThrashBoth conjuncts hold terminally, so the module's arm fires. That is stronger evidence for the conjunction than either arm alone, and it vindicates measuring user CPU specifically: at the wall the process is busy, but the time is system time servicing its own faults — a user-plus-system share would have read as healthy throughput and admitted the thrash.
I reported this the other way round mid-investigation, from the early window alone, and was wrong; the full curve inverts it.
The change
One receipt row on
gunbc.memory_stall_refusal, the module that already owns this class — as the parent asked, the specimen goes to that module rather than opening a class beside it. It records the inert-knob measurement, the localization, both controls, and the reading rule, and states plainly that setting the override to clear aHostBudgetUnreadablerefusal trades a loud typed answer for an unbounded run — the escape hatch DESIGN §5 forbids.Verification
Every number above is from an executed dispatch on tree
4e6a597de0with the runner echoingDIRTY=0; no figure is transcribed from another session. The--roster-from-discoveryarms were run to the wall rather than stopped by a client-side cap — an earlier reading of mine was truncated by my own tool timeout and is not used as evidence.Local
gunbcin this container is stale (it cannot parse//annotations and errors on files this branch never touches). The edit was checked against that: the baseline errors on the same line content, shifted by exactly the two inserted lines, so this change introduces no new refusal. CI is the real check.Follow-up, not in this PR
transport_close_read_as_completioningunbc.recurring_failure_modedeserves this specimen — with the refinement that here the protocol event exists and is simply not consulted, so the remedy is available rather than blocked. That carrier projects intodocs/design-ledgers.mdand needs a regen round, so it belongs in its own PR. The reading rule is recorded in this receipt meanwhile, where a lane debugging this will actually hit it.🤖 Generated with Claude Code
https://claude.ai/code/session_01So6KE5WgyvEYENzwMtexS6