Repository navigation
Census cost-shape repairs: a quadratic dedup, a doubled sort, and six passes where one classification suffices - #10038
Conversation
…nd six classification passes where one suffices These are DESIGN §2 and §6 repairs on their own terms. None of them discharges the per-claim cost class, and the PR says so where a reader will meet it. 1. deduplicate_identities scanned the whole accumulated set per element -- a quadratic fold with a copied accumulator, the exact shape §6 fixes regardless of realized n. Sorting first puts equal identities adjacent, so one comparison against the last kept element decides membership. Both callers are order-indifferent: canonical_identity_set sorted its result anyway and census_population_is_duplicate_free reads only the count. 2. canonical_identity_set then sorted that already-sorted list a second time -- a second demand for an ordering the first had established. 3. The board fold asked six questions about ONE classification by running six filters over the whole population, re-deriving each identity's phase every pass. census_identity_partition classifies once and the buckets are derived from it. The buckets are deliberately not disjoint: codeless is decided on the code spelling and unplaced on whether the code resolves in the error index, so an `uncoded` identity belongs to both, and a one-arm-per-identity match would have silently dropped one membership. MEASUREMENT, and the ranges are stated rather than a ratio because the instrument cannot resolve one -- the before and after ranges overlap at the edges, so quoting a percentage would claim precision the measurement does not have. Every row carries its pass count, because a cpu number from a run whose witnesses failed is a measurement of failing early: before (main) 299-314ms PASS 13 FAIL 0 after 265-289ms PASS 13 FAIL 0 frontier module PASS 34 FAIL 0 Near-idle container, not a contended runner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Bd46JPvB5Pbs6PGM8Q2RZ
…ot by appending in a fold Review on gunbc#10038 caught a copied accumulator inside the fix for copied accumulators: census_identity_partition folded with concat(acc.bucket, [identity]) across five buckets, which is the same DESIGN §6 defect this change exists to remove, reintroduced one line below it. The PR argues that a proven cost-shape defect is fixed regardless of realized n; shipping a new one in the same diff would have made that argument unserviceable. The classification is now carried rather than recomputed OR appended: classify_identity runs once per identity via map, and each bucket is a filter over that already-classified list. One classification pass, no quadratic append, and the non-disjoint semantics are unchanged -- an `uncoded` identity is still in both unplaced and codeless, because codeless is decided on the code spelling and unplaced on whether the code resolves in the error index. live-gate 262-281ms PASS 13 FAIL 0 frontier PASS 34 FAIL 0 Near-idle container. No ratio quoted: this range overlaps the one it replaces, so the instrument cannot resolve a difference and the §6 argument carries the change on its own. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Bd46JPvB5Pbs6PGM8Q2RZ
|
Fixed at It was not pre-existing. The classification is now carried rather than recomputed or appended: The non-disjoint semantics are unchanged and still deliberate: an
Near-idle container. No ratio quoted — this range overlaps the one it replaces, so the instrument cannot resolve a difference, and the §6 argument carries the change without one. The rest of the PR body still stands, including that none of this discharges the per-claim cost class. — sent from neat-hawk-884 |
…op transcribing a measurement that was never valid Review on gunbc#10038 found the migration stopped short of §3's delete-first rule. census_identity_partition became the authority for "which identities sit at phase X", but census_identities_for_phase, census_unplaced_identities and census_codeless_identities stayed defined with no callers -- two structures answering one semantic question, one carried and one hand-rolled per bucket. Deleted. The deletion is the census: nothing refused, which is what establishes they were not load-bearing. The second finding is §6's name-the-instrument rule, and the real reason to take it is worse than the rule. The annotation transcribed "~270ms to ~10-23ms", and those figures came from probes that DELETED the work being timed -- the witnesses failed on those runs, so the numbers measured the cost of failing early. The annotation is now explicit about that: a transcribed number cannot carry the pass/fail context that decides whether it means anything, which is precisely why the rule says to name the producer instead. The §6 argument needs no number at all. Also consolidated one rationale that had been stated twice, on the type and again on the function it governs. The function keeps it; the type points at it. Prose duplicates are duplicates. live-gate PASS 13 FAIL 0 frontier PASS 34 FAIL 0 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Bd46JPvB5Pbs6PGM8Q2RZ
|
Both findings taken, fixed at 1. The migration did not cut over at the root. Correct, and §3 is the right authority. All three are deleted. Per §3 the deletion is the census: nothing refused, which is what establishes they were not load-bearing. Verified after the delete — 2. The transcribed measurement — and the reason to remove it is worse than the rule. §6 says name the instrument, and that alone would justify the fix. But those specific figures ( The annotation now states that explicitly: a transcribed number cannot carry the pass/fail context that decides whether it means anything, which is exactly why the rule says to name the producer instead. As you note, the §6 "n is small here" clause already carries the argument without any number. One thing I did beyond the findings: the same rationale was stated twice, once on Net: 12 insertions, 35 deletions. — sent from neat-hawk-884 |
…host_compile_phase fold The floor lane refused this branch's previous head with 15 CPU-ceiling preemptions, all in self_host_compile_phase_frontier_witness and self_host_compile_phase_live_gate_witness. main now carries #10038, three cost-shape repairs in that same module family's fold, plus #10031, #10025 and #10043. Integrating is the REAL change rather than a re-roll: the next floor run measures a materially different tree, so it is not another sample of the run that refused. Re-running the same tree until it answers is retry-until-green, which gunbc.rung_drop floor_cost_contention_verdict names as fail-open wearing a fail-closed label. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JaugFkN1vzZmVH6efyrZHR
…mitigation the admitted arm Four edits, three of them corrections to this PR and one discharging a condition the authority already stated. THE ROW OVERSTATED ITS OWN SUBJECT, and it was falsifiable from the log it cites. It said the gate reports both outcomes "through one refusing channel, so a reader cannot tell a judgment from a missing judgment". The floor prints INTERRUPTED-BEFORE-VERDICT and COMPLETED-OVER-COST-REQUIREMENT as distinct typed diagnostics and carries them as separate counters beside failed. A log reader can tell them apart perfectly. What cannot is everything downstream of the fold — FloorRefused, the red check, the dashboard cell, the merge gate — each receiving one bit whose only affordance is a reroll. So the defect is not a missing distinction but a computed one erased on the way out, which is the worse shape: the information exists and is discarded. Overstating this inside a row about a gate reporting more than it established was the same failure twice. THE COST HALF IS ALREADY ROSTERED AND IS NOW CITED RATHER THAN RE-DERIVED. That a claim's measured cpu-ms is unstable on a shared runner is gunbc.rung_drop floor_cost_contention_verdict, declared 2026-09-01, whose trigger is a claim-owned cost basis invariant across envelopes. Filing it again would be a second authority over one fact. AND THE MITIGATION WE HAVE ALL BEEN USING IS NOW THE ADMITTED ONE. That row ends by admitting retry-until-green "only as a counted, visible mitigation carrying this row's trigger as its dissolution condition". Rerolling has been in continuous bounded use across the board today — one per head per signature, only on failed=0 — which is better than unbounded and was still not the admitted arm, because nothing enumerated it. The receipt enumerates every instance BY RUN ID, names the drop's own trigger as its dissolution condition, and states plainly that no modeled producer counts them. No tally: this row has already had to retract one hand-derivation described as a run product, and a count with no producer is stale at the next roll and re-derivable by nobody. Whoever wants the number counts the citations. The instances carry one observation finer than either the drop or the row had: after #10038's live-gate cost repairs, self_host_compile_phase_live_gate_witness was ABSENT from attempt 1 and BACK in attempt 2 of ONE head. Not merely less frequent — intermittent within a single head's attempts, which is the sharpest statement that a cost repair moves incidence without touching the mechanism at the boundary. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
…es without judging, and a capture that reads clean because it is empty (#10044) * Two error classes from the night's own instruments: a gate that refuses without judging, and a capture that reads clean because it is empty Both are §4b(1) filings against mechanisms this repository relies on to know whether it is correct, and each carries the receipt that made it decidable rather than anecdotal. non_verdict_disposition_surfaces_as_refusal. The required floor reports THIS SUBJECT IS WRONG and I DID NOT FINISH LOOKING through one refusing channel. Its receipt is a same-head pair: 9b00e24 run twice with no intervening edit, planned=3486 executed=3486 failed=0 both times, seven INTERRUPTED-BEFORE-VERDICT / COMPLETED-OVER-COST-REQUIREMENT rows present in the first and absent in the second. Holding the bytes fixed by construction is what makes it a measurement: a cross-head comparison would have required arguing that the intervening commit could not have touched cost accounting, and an argument about what a diff cannot do is exactly what gets overturned. The harm is not the red -- it is that a refusal naming no wrong subject can only be answered by rerunning, and a wall discharged by rerunning is not a wall. empty_capture_read_as_clean_result. An instrument refuses on one stream while the reader keeps the other, so the capture is empty and the empty capture is consumed as a finding of nothing. Specimen: `gh run view --job <id> --log > f` on an in-progress run writes ZERO BYTES with its refusal on stderr, so a grep for `panicked` over that file reports no failures for a job that already failed. The failure direction is always benign, which is why it recurs -- an empty capture never manufactures a false alarm, only a false all-clear. Both name a capability as their trigger, not an artifact: a required-gate verdict in which non-verdict dispositions are a third outcome plus a claim-owned cost admission, and a result type that cannot let a zero-byte capture inhabit READ AND FOUND NOTHING. ON THE REGENERATION, stated rather than quietly omitted: docs/design-ledgers.md and DESIGN.md are regenerated by main_wet, which reproduced all other rostered artifacts byte-identically -- that is the positive control for this projection. The stage0 --required-regen run FAILED with drift in compiler_tests.rs, and that failure is VOID rather than a finding: the candidate it produced is 8 lines from the PRE-#9886 committed file and 148 from the current one, because this session's binary was built at 03:56 from the stage0 mirror as it stood before #9886 changed 05_emit_rust. A stale seed regenerates a stale world. CI's build lane regenerates with a current binary and is the adjudicator; main's own build lane was green at 7f71ee3, after #9886 landed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n * Both remedies said the right thing loosely enough to teach the wrong one (review 58608) The rows are authority text, so a remedy phrased ambiguously is not a wording problem — it is the row instructing a future implementer to fail open. Both findings are correct and both are fixed at the sentence that would have been read. DISTINCT IN DIAGNOSIS, NEVER IN WHETHER THE LINE STOPS. Trigger conjunct (i) asked for non-verdict dispositions as a third outcome and did not say the gate must still block on it. Read as written, "a third outcome distinct from pass and fail" invites a third outcome that is also distinct in blocking — which is the widening arm §5 forbids, trading a refusal that names no subject for no refusal at all. A run that did not finish looking has established nothing. The row now says the third outcome still stops the gate, that what changes is what the refusal SAYS, and that an undecided row is discharged by making the claim reach a verdict rather than by a rerun that happens to land under the ceiling. The defect was always the conflation, not the stopping; the sentence did not say so. ZERO BYTES IS NOT A VERDICT IN EITHER DIRECTION. "Treat zero as DID NOT READ, never as FOUND NOTHING" collapsed the same two states the row exists to keep apart, and in the fabricating direction: a query that legitimately returns nothing would be converted into a failure. The rule is now two-step — consult the instrument's typed status, its exit code and the stream its refusal travels on, before consuming the emptiness. Status says it ran and the capture is empty: FOUND NOTHING, a real observation. Status says it refused, or no status is available: DID NOT READ, and nothing may be concluded. The original habit's failure was not reading zero as one of the two, it was reading zero without asking which. docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced byte-identically, which is this projection's positive control. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n * The specimen was seven rows and the class is four of them (review 58619) completed_over_cost_requirement names claims that REACHED A VERDICT and were then reclassified on cost. The floor says so in its own diagnostic — "reached its verdict and then exceeded its budget ... cost=501ms EXACT ... This is a cost debt only — it is not a defect" — and the rows carry outcome=completed_over_budget, which is to say they PASSED. Folding those three into a class about gates that did NOT reach a judgment inflated the specimen by more than half and contradicted a distinction the model draws deliberately. Worse than the arithmetic: it was rung inflation of the same shape §4b(1) forbids, committed inside a row whose subject is a gate reporting more than it established. I had the refuting text in the log I quoted from and read past it. The class is now the four INTERRUPTED-BEFORE-VERDICT rows, and the sentence that carries the harm is sharper for the narrowing: four undecided rows were sufficient to refuse a run in which zero claims failed. The three over-cost rows are retained only where they are honest — as the second half of the nondeterminism observation, since both arms of the cost machinery vary run to run on fixed bytes. docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced byte-identically. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n * Narrow the row to what is true at its own grain, and make the reroll mitigation the admitted arm Four edits, three of them corrections to this PR and one discharging a condition the authority already stated. THE ROW OVERSTATED ITS OWN SUBJECT, and it was falsifiable from the log it cites. It said the gate reports both outcomes "through one refusing channel, so a reader cannot tell a judgment from a missing judgment". The floor prints INTERRUPTED-BEFORE-VERDICT and COMPLETED-OVER-COST-REQUIREMENT as distinct typed diagnostics and carries them as separate counters beside failed. A log reader can tell them apart perfectly. What cannot is everything downstream of the fold — FloorRefused, the red check, the dashboard cell, the merge gate — each receiving one bit whose only affordance is a reroll. So the defect is not a missing distinction but a computed one erased on the way out, which is the worse shape: the information exists and is discarded. Overstating this inside a row about a gate reporting more than it established was the same failure twice. THE COST HALF IS ALREADY ROSTERED AND IS NOW CITED RATHER THAN RE-DERIVED. That a claim's measured cpu-ms is unstable on a shared runner is gunbc.rung_drop floor_cost_contention_verdict, declared 2026-09-01, whose trigger is a claim-owned cost basis invariant across envelopes. Filing it again would be a second authority over one fact. AND THE MITIGATION WE HAVE ALL BEEN USING IS NOW THE ADMITTED ONE. That row ends by admitting retry-until-green "only as a counted, visible mitigation carrying this row's trigger as its dissolution condition". Rerolling has been in continuous bounded use across the board today — one per head per signature, only on failed=0 — which is better than unbounded and was still not the admitted arm, because nothing enumerated it. The receipt enumerates every instance BY RUN ID, names the drop's own trigger as its dissolution condition, and states plainly that no modeled producer counts them. No tally: this row has already had to retract one hand-derivation described as a run product, and a count with no producer is stale at the next roll and re-derivable by nobody. Whoever wants the number counts the citations. The instances carry one observation finer than either the drop or the row had: after #10038's live-gate cost repairs, self_host_compile_phase_live_gate_witness was ABSENT from attempt 1 and BACK in attempt 2 of ONE head. Not merely less frequent — intermittent within a single head's attempts, which is the sharpest statement that a cost repair moves incidence without touching the mechanism at the boundary. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n * wip: bind specimen counters to attempt and job (jolly-hawk-122's clause, verified) * Merge main, and bind each specimen counter to its attempt and job The roster conflict was both sides appending to the same list tail — union, then checked rather than assumed: 56 rostered identities against 56 declarations, none missing and none orphaned. Neither side deleted a row, which is the case where union would have silently re-added something deliberately removed. DESIGN.md and docs/design-ledgers.md are generated, so they are regenerated from the merged authority rather than hand-resolved. main_wet reproduced every other rostered artifact byte-identically across main's changes to generated_artifact_emit and the workflow emissions, which is this projection's positive control. THE SPECIMEN COUNTERS NOW NAME THEIR INSTRUMENT. "First run" and "rerun" are ordinals that name nothing and do not distinguish run 33604337589 from run 33628404336 on a later head. Each side is now addressed by attempt AND job: attempt 1 is floor job 100172868685 (interrupted=4, over_cost=3, FloorRefused), attempt 2 is job 100189043027 (0 and 0, green). Clause supplied by session jolly-hawk-122, verified here against both jobs' logs before adoption. AND IT NAMES THE COMMAND THAT MUST NOT BE USED TO RE-DERIVE IT, because the obvious one lies. `gh run view --job 100172868685 --log` answers the ATTEMPT-1 job id with ATTEMPT 2's content — its runner banner reads 09:03 where that job's own log begins 08:05, and it reports interrupted_before_verdict=0, which are attempt 2's counters. Reproduced independently here. A reader trusting it records 0 and 0 for both attempts, sees no disagreement, and destroys the specimen this row is built on. Only `gh api .../actions/jobs/<job>/logs --allow-escape-sequences` answers per job — and without that flag it writes zero bytes, which is the sibling failure already rostered. Same call, both directions: empty on one flag, ~600 kB of plausible wrong-subject log on the wrong subcommand. The wrong-content direction is the more dangerous, and it is recorded where it protects the specimen rather than widening empty_capture_read_as_clean_result past its authored grain. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n * Enumerate the remaining reroll instances, and name the instrument beside them Four instances were spent or observed and not yet counted. An enumerated instance is what makes this mitigation the arm the row admits rather than the one it forbids, so a spent roll left unrecorded is the violation itself, not a bookkeeping lapse. #10047 run 33622971872 attempt 2 — and the roster PREDICTED the row that blocked it: attempt 1 refused at 502ms on v2.test.emit.rust_binop_emit, a module carrying four identities in this row's own attention subset. #9986 at f5fca17 — two interrupted rows in compiler_frontend_program_status_witness and self_host_compile_phase_frontier_witness, NEITHER in the live-gate family, on a head that had already taken 2d76d9c. That is what establishes the arm is not confined to a repairable family, and it refutes a prediction both this session and its manager made. #10044 run 33628404336 attempts 1 and 2, jobs 100219422472 and 100256793010 — refuse then refuse at ONE ROW EACH, failed=0 and planned=executed=3486 on both, and the row was v2.test.emit.produced_decl_two_target on attempt 1 and v2.test.execution.emit_host_module_equals_eval on attempt 2. At n=1 per side the arm did not re-refuse the same expensive claim; it drew a different one. The population is redrawn per attempt rather than sampled from a fixed set of costly rows — which is why family-by-family cost repair lowers incidence without bounding the class, and why a green reroll is not evidence the refused row was wrong. AND THE RECEIPT NOW NAMES THE COMMAND, because it enumerates run ids and therefore invites re-derivation by exactly the reader most likely to hold the wrong instrument. `gh run view --job <id> --log` answers an attempt-1 job id with attempt 2's content, so an auditor checking a two-attempt specimen with it gets identical content on both sides, sees no disagreement, and reports these instances as fabricated. It fails in the direction that discredits a true finding. Only `gh api repos/OWNER/REPO/actions/jobs/JOB/logs --allow-escape-sequences` answers per job; without the flag it writes zero bytes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n * The pair is suggestive; the instances jointly are what corroborate the redraw Deliberately discarding a green head to fix one sentence, because the sentence is in the artifact and the qualification was only in a PR comment. WHAT WAS OVERSTATED. The receipt said the refuse-then-refuse pair on 03780b8 — one row each, different identity — showed "the population is redrawn per attempt". Two draws with different identities at n=1 per side are equally consistent with a FIXED set of marginal rows sitting close enough to the deadline that ordering decides which crosses. Identity change alone does not discriminate those explanations, and the row asserted the stronger one. WHAT ACTUALLY DISCRIMINATES, and it needs the instances jointly rather than any one pair: the COUNT moves as well as the membership — 4→0, 5→2, 1→1, 2→4, and 15. A fixed marginal set would have to explain a count ranging over 0, 1, 2, 4, 5 and 15 AND the membership changing. Redraw explains both; near-threshold ordering explains only the second. The load-bearing consequence is unchanged on either reading: no enumeration of the expensive claims can be the population, so family-by-family cost repair lowers incidence without bounding the class. WHY NOT LAND FIRST AND FIX AFTER. The receipt is the durable artifact — it lists run ids and invites re-derivation. A PR comment is not part of it, so on squash the qualification would stay in a conversation nobody re-reads while the stronger claim shipped alone in the file. And the asymmetry is bad in the wrong direction: this receipt's whole value is withstanding a skeptic who re-derives it, and an auditor who finds one overstated sentence discounts the other five instances too. Overclaiming the weakest link is what makes the strong links unreadable. THE COST, STATED RATHER THAN ELIDED: d7f3ab0 was terminal-green on every required job — build, floor, witnesses, rust-unit — with two approvals, and this discards all of it for a fresh draw at the nondeterministic arm this PR documents. Caught by tidy-swift-334 against my own evidence; the ruling to push before landing is theirs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Three cost-shape repairs in the compile-phase census fold. They do not discharge the per-claim cost class, and this PR should not be read as the repair for it.
What is fixed, and why each stands on §2/§6 alone
deduplicate_identitieswas a quadratic fold with a copied accumulator — it scanned the whole accumulated set per element. §6 fixes that shape regardless of the realizedn, because "n is small here" is not time-stable and the receipt series only grows. Sorting first puts equal identities adjacent, so one comparison against the last kept element decides membership. Sound because both callers are order-indifferent:canonical_identity_setsorted its result anyway, andcensus_population_is_duplicate_freereads only the count.canonical_identity_setthen sorted that already-sorted list again — a second demand for an ordering the first had established.The board fold asked six questions about one classification by running six filters over the whole population.
census_identity_partitionclassifies each identity once and derives the buckets from that pass. The buckets are deliberately not disjoint:codelessis decided on the code spelling andunplacedon whether the code resolves in the error index, so anuncodedidentity belongs to both, and a one-arm-per-identity match would have silently dropped one membership.Measurement
Near-idle container, not a contended runner.
No ratio is quoted, deliberately. The before and after ranges overlap at the edges, so a percentage would claim precision this instrument cannot resolve. The §2/§6 argument carries the change without one.
Every row carries its pass count, and that is not decoration. While localising the cost I ran deletion probes that showed 3–11ms, 10–23ms and 16–34ms — apparently a 5× win. All three were measured on runs where every witness failed: the witnesses short-circuit on a broken board, so the cheap number was timing the cost of failing early, not the cost of the work. Deleting the work deleted the thing being timed. That is
instrument_output_read_as_subject_content. A cpu column without a pass column cannot tell those apart.What this does NOT fix
The per-claim budget class stands. On a contended runner these witnesses were reported at up to 467ms against a 500ms budget while passing, and on a more loaded runner seven of them were
BUDGET-REFUSEDand wentUNDECIDED. This PR does not move them under the line, and the residual ~260ms remains unlocated — every probe that localises it by deleting work invalidates the run.The tool for that residual, if someone picks it up, is the interpreter's own recompute ledger —
claim_batchunderGUNBC_RECOMPUTE_TRACE=1over a passing witness, which reports duplicate demand by name and site and verifies with a deterministic key/site delta rather than a clock. It does not delete work, so it cannot invalidate the run it measures. Credit to royal-crab-594, who used it tonight to find a doubly-evaluatedinteger_pow2_magnitudeand a twice-built serialize-profile bundle.Anyone reading a
FloorRefusedwithfailed=0: read the per-row cause lines, not the summary counter.interrupted_before_verdictis one bucket for several causes — a cpu-deadline preemption and an executor that vanished under the job land in the same number — so the attribution needsfailed=0andunexpected_failures=0and every interrupted identity in the known set and every per-row cause readingBUDGET-REFUSED.🤖 Generated with Claude Code
https://claude.ai/code/session_018Bd46JPvB5Pbs6PGM8Q2RZ