Skip to content

Stage 1 of the serving extraction: what the observer may read from vLLM, and when it must refuse - #10786

Closed
briansrls wants to merge 61 commits into
mainfrom
spark/serving-receipt-schema
Closed

briansrls wants to merge 61 commits into
mainfrom
spark/serving-receipt-schema

Conversation

@briansrls

@briansrls briansrls commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

What this PR actually contains — read this first

It carries two lanes, and only one is mine:

commits files lines
serving-observer-contract (not my work, unreviewed, no PR of its own) 26 29 +4462
Stage 1 of the serving extraction (mine) 3 3 +628

My branch was auto-pushed and auto-opened as this PR before a base could be chosen, so it
inherited the base branch.

This PR is the verdict's PR A. serving-metric-semantics and serving-prefill-scaling are
both ancestors of serving-observer-contract, so this branch already carries their terminal
state with the refuted intermediate statements corrected downstream — the consolidation the
verdict asks for is this branch's history, not something to reconstruct. serving-placement-authority
is a descendant of this branch, so landing PR A unblocks the restack PR B needs.

Fabric placement is deliberately not here. An earlier revision of this PR declared a
FabricConsumption type and a placement-standing row. Since PR B owns that question and is built
on top of this branch, those would have landed underneath the lane that owns them and become a
second authority to reconcile. They were carved out (commit 7e348f2). The observation survives as
a finding for that lane, not as a declaration here: dag/gunbc/spark imports nothing from
gunbc.fabric and picks its head host from a constant.

My three files: dag/extdeps/vllm/scheduler_extension.dag,
dag/gunbc/spark/serving_observer_compatibility.dag,
dag/test/claim/spark/serving_observer_compatibility_witness_test.dag.

Stage 1 — the observer's binding surface and its compatibility manifest

FIRST-PARTY SERVING EXTRACTION sequences ten cuts; this is the first. It answers what a
first-party step observer must read from vLLM, at which exact revision, and when it must
refuse. It builds no observer. gunbc.spark.serving_step_receipt authored the receipt
before its producer so the producer could not become the authority by accident; this applies
the same move one layer out, to the join between that receipt and upstream — which belonged
to neither side, and which the wheel would otherwise have defined by being the only thing that
knew it.

Three findings, each read from the pinned source (1967a5627b) rather than recalled:

Read the output, never the request. Scheduler.schedule constructs its SchedulerOutput
and then, as its last act, calls _update_after_schedule, which does
request.num_computed_tokens += num_scheduled_token. One step therefore has two values of
num_computed_tokens, and which an observer sees depends purely on where it reads. A wrapper
that calls the delegate and then reads the live Request objects — the obvious way to write
it, since it is holding the scheduler — records a cached_tokens that already contains the
step's own computed_tokens. Both numbers are plausible, the arithmetic succeeds, and nothing
downstream can catch it.

The default delegate is conditional. With scheduler_cls unset, get_scheduler_cls
returns AsyncScheduler when async_scheduling is truthy and Scheduler otherwise; setting
scheduler_cls removes that branch entirely. A wrapper hardcoding Scheduler does not observe
an async-scheduled engine — it silently converts it to a synchronous one and reports the result
as behaviour-preserving. Async is refused outright until its num_output_placeholders
semantics are modelled, rather than admitted and averaged into one population.

Upstream disclaims the interface, in its own words: "This scheduler interface is not
public and compatibility may not be maintained."
That is why the binding is admitted against
one exact revision and refuses every other — no version range is an admissible compatibility
statement for a surface whose shape may change without anything a range could observe.

Evidence

Ten witnesses green by execution under gunbc run (not a typecheck). The load-bearing
one carries a discriminating red: injecting the defect it guards — admitting the
live-request read site — flips it true → false; the module was restored byte-identical.

Full-corpus compile (--source-root dag --source-root src/v2, the CI argv): zero
diagnostics name any of my three files.
Other diagnostics in that run are in files this
change does not touch; I did not compile main alone to baseline them, and my merge resolution
touched only a markdown projection, so it cannot have produced .dag diagnostics.

The merge of main hit one real conflict — docs/design-failure-modes.md, a generated
projection — resolved by its declared repair route: incoming side taken verbatim, and
verified by set difference (not count) that no row went dark, with roster.dag confirmed to
carry both sides' rows so the healer can regenerate.

What this does NOT do

  • It does not generate the receipt wire schema — the other half of stage 1's brief. That
    projection depends on this manifest and lands next.
  • It discharges no part of spark_step_receipt_producer, which stays ProducerUnbuilt.
  • Stages 2–10 need a wheel build and live Spark hardware, neither reachable from this session.

🤖 Generated with Claude Code

https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t

gunbc-ci-auto-heal and others added 30 commits September 6, 2026 21:07
Every capacity, usage, restart and routing observation this repository takes about a serving engine
has been joined on some mixture of host, group, model name and self-reported version. None of those
identify a LAUNCH. On 2026-09-05 the group A head produced engine cores reporting 23.6 GiB, then
21.29 GiB, then 54.01 GiB of available KV — three incarnations from ONE rendered unit whose text
never changed. Joined on host and unit those are three readings of one unstable object; joined on
incarnation they are three objects with one reading each, and the live question becomes which one is
serving now.

It is a safety defect and not only a reporting one: a seat granted against the engine that answered
a probe must not authorize a request against the engine that has since replaced it, and
"FabricGroupA" cannot tell them apart. Restart is the event that invalidates every outstanding
grant, so a grant that cannot name its incarnation cannot be fenced.

VllmEngineIncarnation is sole_constructor over group, head, systemd invocation, container id,
executable digest, unit digest and start instant. The engine's self-reported vLLM version is
deliberately NOT a field: extdeps.vllm.server already carries VllmSourceRevision for what the laws
were established against, and a version a process prints about itself is a claim, not provenance.

same_incarnation discriminates on the systemd invocation — the one field guaranteed to differ across
two starts of a unit — then folds a row list of agreement checks, so a same-invocation pair whose
fields disagree returns IncarnationContradiction rather than a match. That arm is separate from
DifferentIncarnation on purpose: ordinary turnover and "two observers disagree about one launch"
call for different action.

Cross-family digests return DigestsIncomparable and surface as a contradiction naming the
incomparability. std.content_hash refuses to collapse a cross-family pair into a silent false
precisely so callers cannot read "cannot compare" as "differ"; mapping it to false here would
reintroduce that collapse one level up.

Witness (4): two launches of one unit — identical group, head, container name, image digest and unit
digest — are NOT one incarnation, which is exactly the case a group- or unit-keyed join answers
wrong; a disagreeing field under one invocation is a contradiction; cross-family digests refuse as
incomparable rather than reading as a rebuild.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
…r is a projection of it

Three quantities in this deployment have all been called "the pool", and each is a real number the
engine reports: the ALLOCATOR INVENTORY (num_gpu_blocks, what the scheduler allocates from and what
kv_cache_usage_perc divides by), the FULL-CONTEXT EQUIVALENT (the "GPU KV cache size: N tokens"
line, which is max_concurrency x max_model_len re-expressed in tokens), and the SCHEDULER
GRANULARITY (block_size, the MINIMUM across resolved groups). Multiplying the first by the third and
calling it the second is the error extdeps.vllm.kv_cache already refuses. This module makes the
three projections of one declared structure, which is the only form in which they can be checked
against each other.

ResolvedKvLayout carries NO incarnation. This is the upstream shape; binding a shape to the launch
that exhibited it is a receipt, which is an observation this repository produces and so belongs in
the observing layer (DESIGN section 3, external upstream decomposition). Carrying it here would
author the identity twice, since the receipt must name it anyway.

A zero block size is not a degenerate group, it is not a group: KvBlockSize is refined positive, so
the state has no constructor and no cost function needs a division-by-zero branch inventing an
answer for a layout that cannot exist.

cold_admission_blocks is per-group — full attention grows with the sequence, a windowed group stops
at its window, a mamba group is one state per sequence — and REFUSES today, because the engine's
group metadata emits kind, block size and window but not layer count, page size or its calculated
blocks-per-max-request. Scaling the recovered full-context cost by length/max_model_len is
deliberately not offered: windowed groups do not scale, so the scaled figure understates what a
shorter request holds and would over-admit.

ONE CHECK WAS WRITTEN WRONG FIRST AND THE WITNESS CAUGHT IT. "Recover the per-request cost, then
replay it against the printed concurrency" expands to A*P/(A*M) = P/M — the inventory cancels, so it
is green for any inventory including one lifted from another engine. Shipped as an inventory check
it would have been a permanently-green wall cited as coverage. It is now named printed_lines_agree
for what it decides: the two printed lines against each other. THE INVENTORY REMAINS UNCHECKED BY
TODAY'S EVIDENCE, and a witness row asserts that limit so no later reader mistakes it for coverage.

Witnesses (6). The discriminating one: two layouts with the SAME allocator inventory and the SAME
global block size admit different costs, because the windows and group sizes differ — any model
reducing capacity to those two fields answers identically for both and is wrong for at least one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
…exhibited it

The acquisition boundary for engine capacity, in its own module rather than as rows in
gunbc.spark.vllm_observed. That module is a historical, route-local 2026-09-03 investigation whose
cross-route joins were never established; making it the live seating authority would promote a
historical fixture into current state, and the first stale row would then be authorizing admissions.

VllmKvLayoutReceipt has three arms and KvLayoutObserved has no free constructor — it is reachable
only through admit_kv_layout, and only when the group layout independently reproduces the
per-request cost the engine printed. That is the difference between a record and a receipt: a record
says what was seen, a receipt says what was established.

THE ONE NON-CIRCULAR CHECK, added to extdeps.vllm.kv_layout. Two routes reach the blocks one
full-context request costs: RECOVERED, inverted out of the two printed lines and blind to groups;
and DERIVED, the per-group sum at max_model_len and blind to the printed lines. They share no input
but the context cap, so agreement is evidence and disagreement is located. Unlike printed_lines_agree
— whose inventory cancels out — this wall's RED is authorable, and the witness authors it.

Today every real capture lands KvLayoutInconsistent, exactly as predicted in the ruling: the engine
emits kind, block size and window but not layer count or page size, so the derived route does not
exist and the check refuses. An absent second opinion is not agreement.

seating_authority_for is the only question a lease may ask of a receipt: does this describe the
engine I am about to send a request to? It answers through same_incarnation, so an observed receipt
still refuses a request a LATER launch will serve — same group, same head, same image, same unit,
engine restarted underneath. That is the case a group-keyed binding gets wrong.

The reported figures are carried as the upstream carrier rather than re-spelled as two fields beside
the layout: the pair is exactly what the reproduction check reads, so copies that drifted would
check one and report the other.

Witnesses (5): today's real capture is inconsistent and seats nothing; a reproducing layout is
observed (positive control, so the refusals are about missing metadata and not a fold that cannot
say yes); an observed receipt refuses to seat a later launch; garbled lines and a contradicting
layout are different refusals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
DESIGN section 4c admits only standalone leading // blocks attached to module-scope declarations;
a note inside a function body is a parse refusal, which the CI adjudication caught as
'source annotation sits inside a declaration body'. The local closure compile does not run the
annotation-grain check, so this was invisible until the floor lane ran it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
…that worked

gunbc.spark.pair_serving_apply renders a unit, restarts when the text changed or the unit is not
active, and records one line per host: applied, with a free-text detail. Everything between
"started" and "the unit is active now" is discarded. On 2026-09-05 the group A head reached restart
counter 8 with status=137 before a start that stuck, and to both systemd and convergence that end
state is indistinguishable from a unit that came up first time. The failures were not merely
unreported — they were unrepresentable.

It matters beyond reporting: the three launches that head produced had 23.6, 21.29 and 54.01 GiB of
KV. A flapping bring-up is exactly the condition under which the surviving incarnation is likely to
be a degraded one, because each attempt starts against whatever the previous attempts left behind.

EXIT 137 IS A SIGKILL SHAPE, AND SIGKILL IS NOT OOM. This is the inference the incident invites and
it is wrong: 128+n is the convention for "terminated by signal n", and signal 9 is delivered by the
kernel's OOM killer, by `docker rm -f` (which this unit runs in its own ExecStartPre), by a stop
timeout escalating, and by an operator. Reading the code as OOM would send someone to tune memory
that was never the cause. So classification stops at the signal, and KilledByOom requires a separate
observation — kernel ring buffer, cgroup memory event, or systemd OOM result — joined to the same
incarnation through same_incarnation. An OOM that belongs to a different start of the same unit does
not attribute, which is the case a host-keyed join accepts.

ServingStanding distinguishes ServingOnFirstAttempt from ServingRecoveredAfterFailures, carrying the
failed count and their terminals, and an empty attempt list is ServingStandingUnobserved rather than
NotServing — absence of evidence is not failure, and reading it as one would make an unconverged
host indistinguishable from a broken one.

Witnesses (6), including the positive control that a joined observation does promote the kill, so
the refusal is about missing evidence rather than an unreachable attribution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
…its own preconditions

THE FLOOR IS ONE DECLARATION. "54 GiB of KV" and "three concurrent sessions" are the same
requirement said twice, and once both are written down they drift — a change to the sequence budget
moves one and leaves the other standing. So only the operational unit is declared,
MinimumColdSessions, and the blocks are derived from it against the engine's own layout. The session
count is not a second number either: it reads harness_seat_operator_policy_seats, because a fabric
whose floor said four while the harness seated three would hold capacity nobody could use.

COLD means no prefix-cache credit. A prospective hit belongs to a prefix another request may release
at any moment; reserving against it spends a discount that was never granted.

"BELOW FLOOR THEREFORE RESTART" IS ONE STEP WHERE THERE ARE TWO DECISIONS, and collapsing them is
how a convergence run bounces a busy engine. Free memory on a unified-memory host is a moving
target, so a restart taken because the pool came up small can come up smaller, and a loop that
restarts on every unfavourable reading keeps restarting with requests in flight.

The first move is therefore always withdrawal: stop sending new work. It costs nothing, is instantly
reversible, and is correct under every failing standing including the ones where capacity is merely
unknown — an unestablished capacity withdraws rather than assuming adequacy, which is today's real
state for every engine in the fleet.

Restart is then admitted, not triggered, behind five preconditions, and the fold names the FIRST
unmet one. Four are decidable from what this repository models. The fifth refuses: no admitted
unified-memory plan exists, and on a GB10 the pool is a residual of a supply shared with the OS page
cache, the weights, the CUDA and NCCL contexts and this deployment's own Ray object store —
restarting without a plan over it is what produced 23.6 GiB once and 54.01 GiB later from one unit
text. Admitting a restart there would promise an outcome nothing modelled can predict. That is the
M4/M6 slice; when it lands the arm is replaced by a real precondition rather than deleted.

Re-admission is earned by a FRESH receipt. A completed restart proves a process started; on this
hardware it proves nothing about the pool it came up with.

Witnesses (7). The discriminating one: a busy engine below floor is withdrawn but not restarted, and
with every decidable precondition met the restart still refuses naming the missing memory plan.
Halving the declared budget halves the derived requirement, which a second literal would not do.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
… was already taken

gunbc.spark.vllm_rank_population has carried VllmEngineIncarnation since before this stack, with the
SAME reasoning from the SAME incident -- three launches on one pair in one day, identical version
and checkpoint, resolved capacity differing by 22%. M0 re-invented that concept under its own name,
which is the section 3 nickname this repository forbids: one meaning, two names, and the derived
work duplicated behind each.

The merged type keeps the name. The two are different GRAINS of the same question and neither
observer can produce the other's fields: VllmEngineIncarnation is what the ENGINE reports about
itself (the engine-core identity, readable from its own output), and this type is what the HOST
knows about the launch (systemd invocation, container, image and unit digests, start instant). So
this one is renamed VllmServingLaunch, which is what it actually holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
… was already taken

gunbc.spark.vllm_rank_population has carried VllmEngineIncarnation since before this stack, with the
SAME reasoning from the SAME incident -- three launches on one pair in one day, identical version
and checkpoint, resolved capacity differing by 22%. M0 re-invented that concept under its own name,
which is the section 3 nickname this repository forbids: one meaning, two names, and the derived
work duplicated behind each.

The merged type keeps the name. The two are different GRAINS of the same question and neither
observer can produce the other's fields: VllmEngineIncarnation is what the ENGINE reports about
itself (the engine-core identity, readable from its own output), and this type is what the HOST
knows about the launch (systemd invocation, container, image and unit digests, start instant). So
this one is renamed VllmServingLaunch, which is what it actually holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
… was already taken

gunbc.spark.vllm_rank_population has carried VllmEngineIncarnation since before this stack, with the
SAME reasoning from the SAME incident -- three launches on one pair in one day, identical version
and checkpoint, resolved capacity differing by 22%. M0 re-invented that concept under its own name,
which is the section 3 nickname this repository forbids: one meaning, two names, and the derived
work duplicated behind each.

The merged type keeps the name. The two are different GRAINS of the same question and neither
observer can produce the other's fields: VllmEngineIncarnation is what the ENGINE reports about
itself (the engine-core identity, readable from its own output), and this type is what the HOST
knows about the launch (systemd invocation, container, image and unit digests, start instant). So
this one is renamed VllmServingLaunch, which is what it actually holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
… was already taken

gunbc.spark.vllm_rank_population has carried VllmEngineIncarnation since before this stack, with the
SAME reasoning from the SAME incident -- three launches on one pair in one day, identical version
and checkpoint, resolved capacity differing by 22%. M0 re-invented that concept under its own name,
which is the section 3 nickname this repository forbids: one meaning, two names, and the derived
work duplicated behind each.

The merged type keeps the name. The two are different GRAINS of the same question and neither
observer can produce the other's fields: VllmEngineIncarnation is what the ENGINE reports about
itself (the engine-core identity, readable from its own output), and this type is what the HOST
knows about the launch (systemd invocation, container, image and unit digests, start instant). So
this one is renamed VllmServingLaunch, which is what it actually holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
…roup layouts were absorbed, and an impossible free count read as saturation

FINDING 1 — THE OMITTED-HEAD FALSE MATCH. `head` sat in VllmServingLaunch and appeared nowhere in
agreement_rows, so two launches on DIFFERENT HOSTS that happened to share an invocation id joined as
the same launch. A systemd invocation id is unique per host, not across the fleet, so nothing
prevents that coincidence between two of the eight Sparks — and the whole point of the type is to
refuse exactly the join that co-location makes look right. A carrier field the fold ignores is worse
than no field: it reads as though the host were part of the identity while contributing nothing. The
rule the row list now holds is that the type has no field the fold does not consult.

FINDING 2 — MALFORMED GROUP LAYOUTS WERE ABSORBED. The window sat beside the kind as an optional
field, so a SlidingWindowAttention group with no window typechecked and the cost function charged it
as full attention: a malformed layout producing a plausible number instead of a refusal, which would
then have been reproduced against the printed lines and could have AGREED. The window now lives in
the variant that has one, which deletes both that state and its dual (a full-attention group
carrying a window) rather than checking for them. layer_count is refined positive alongside
block_size, because a zero-layer group contributes zero blocks to every cost — a layout that makes
requests look free. page_size_bytes is deleted: it was read by nothing, and a recorded field no
operation consumes reads as coverage of a byte-level accounting that does not exist. It returns with
the consumer that needs it.

FINDING 4 — THE IMPOSSIBLE FREE COUNT READ AS SATURATION. allocator_usage clamped a free count
larger than the inventory with a min, reporting "nothing free" — the most alarming answer available,
reached by fabrication, and indistinguishable from a genuinely saturated engine. The two numbers can
only disagree that way if they came from different launches, and the widen destroyed the only signal
saying so. It now refuses with FreeCountImpossible carrying both figures.

Witnesses: 8, up from 6. New discriminating rows for the impossible free count and for the windowed
group carrying its window by construction.

Finding 3 (the exact-point and revision problem in the independent reproduction) is fixed on the
child branch, where that check lives.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
…ts revision guard was unreachable

THE EXACT-POINT PROBLEM. The check compared the group-derived per-request cost to a cost INVERTED
out of the printed pool line and demanded integer equality. The pool figure is printed as a whole
number of tokens after a float multiply, so the inversion lands a block or two from the engine's
real divisor and a perfectly correct layout FAILS. A wall a correct input cannot pass is not a
strong wall; it is a broken one, and its red would have been read as evidence that the group
metadata was wrong.

The derived cost is now pushed FORWARD through the engine's own last step and compared against the
figure the engine printed, at the precision it printed it: concurrency = inventory / cost, to two
decimal places. Both sides are integers the engine committed to, the inventory participates (so this
is not the tautology printed_lines_agree turned out to be), and the tolerance is not invented — it
is exactly the precision of the published figure.

THE REVISION PROBLEM WAS DEEPER THAN A MISSING GUARD. Adding one would have been a decoration:
vllm_startup_reading STAMPED observed_at_revision from vllm_source_revision_read — the constant
naming the revision the laws were read from — and VllmStartupReading is sole_constructor, so no
reading could ever carry a different revision and no downstream comparison could ever go red. The
carrier's own note calls that field load-bearing while its constructor made disagreement
unwritable.

So the constructor now takes the revision the caller captured from. A reading is evidence about the
build that emitted it, and agreeing with the modelled revision becomes a real question with a
reachable "no". Only witnesses construct readings today, so no production caller changes; the merged
vllm_kv_pool_witness is updated and still green at 9.

With that, layout_reproduces_reported refuses a cross-revision comparison — the pool and concurrency
lines are derived output, exactly what an upstream refactor moves, and checking a layout against
them under a revision nobody read asserts a law about unexamined code.

Witnesses: the foreign-revision refusal is now authorable and asserted, with the same evidence at
the modelled revision as its control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
… witness

The group layout lost its detachable window and its unread page_size_bytes, and a startup reading
now observes the revision it was captured from. This fixture follows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
…y is, and today gets "unestablished"

Every capacity claim in this stack has been green by typecheck. This one executes:

  gunbc run --entry dag/gunbc/instruments/fabric_capacity_standing.dag \
    --function fabric_capacity_standing

Run against the fleet just now, group A answered with its real launch — invocation
d140fa39c1174403bf2ce4d7ed8c2d74, inventory 54,475 blocks of 16, printed pool 11,767,856 tokens at
11.22x, 0 running, 0 waiting — and the receipt refused: the resolved cache groups are unobserved, so
the layout cannot reproduce the engine's own arithmetic. Floor unestablished, posture WITHDRAWN,
exit 1.

That refusal is the deliverable. Knowing mechanically that capacity cannot be established beats
reading a dashboard number that says 99% and means retained prefix cache. The instrument earns its
keep the day that answer changes without anyone intending it to.

ONE ROUND TRIP PER GROUP, and the far side does the text handling: the head curls its own metrics
endpoint, greps its own launch log, and normalises twelve fixed-position lines, so this module parses
integers at known offsets instead of carrying a Prometheus parser and a log parser it would have to
keep in step with two upstream formats. A field the far side cannot produce arrives as "-", fails to
parse, and refuses — it never arrives as a zero.

Two defects found by running it rather than by reading it:

- The first cut returned a String and decided admittance by searching it for "posture: admitting",
  which re-ran group_report to ask the question — so every group was probed over ssh TWICE per
  invocation. Visible in the run as two identical remote executions. The standing now carries the
  verdict as a field beside the text, so the effect happens once and the decision is read rather
  than recovered from prose.
- No brace appears in the remote script, which is a constraint of the reader rather than of shell:
  the heads-only pass that indexes modules counts braces without knowing it is inside a string
  literal, so a docker --format template ended the enclosing declaration early and the module
  refused to parse.

Group B did not answer: the fleet-automation key is authorised on .225 and not on .236, so half the
serving fleet is unreachable to automation. Reported rather than patched by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FbXKcgqW4Jx7e78BPWc8ci
@gunbai-bot gunbai-bot Bot changed the title SPARK-N work Stage 1 of the serving extraction: what the observer may read from vLLM, and when it must refuse Sep 7, 2026
gunbc-ci-auto-heal and others added 3 commits September 7, 2026 16:45
Ledger-Repair-Judged: docs/design-failure-modes.md
Ledger-Rows-Repaired: docs/design-failure-modes.md cumulative_metric_read_as_per_event
Ledger-Repair-Judged: docs/design-rung-drops.md
…he placement lane

The operator is taking spark-to-fabric placement in separate PRs, and serving-placement-authority
is a DESCENDANT of this branch. A FabricConsumption type and a spark_serving_placement_standing
row declared here would therefore land UNDERNEATH the lane that owns the question, and become a
second authority that lane has to reconcile before it can state its own -- the meaning fork this
programme is sequenced to avoid, committed by the change that names it.

The observation this section carried is not lost, it is just not code here: dag/gunbc/spark imports
nothing from gunbc.fabric and picks its head host from a constant. That belongs to the placement
lane as a finding, not to this module as a declaration.

What remains is one epistemic unit: what the observer may read from vLLM, at which revision, and
when it must refuse. Ten witnesses re-run green by execution after the carve.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 7, 2026 18:45
From review 62094. PowerEnvelope.watts was a flat scalar where std.measure.Watt
(Measure<Power, One, Nat>, constructor `watt`) already exists, which is DESIGN 2's failed
decomposition: a fresh authority minted for a concept the corpus already carries.

This field is arithmetic-free -- four data rows, no consumer reads it -- so consuming the
carrier is mechanical and total. The module's witness still returns true by execution.

The review's remaining rows on this file are NOT fixed, and the reason is a missing carrier
capability rather than scope: see the PR reply. std.measure publishes measure_add, measure_le,
measure_count and the two scale-fraction folds, and no multiplication, division or dimensional
product. The flagged micros fields feed a least-squares affine fit whose terms are a
cross-dimension product (tokens x micros), a square, and a quotient whose unit is
Microsecond/TokenCount -- the very rate the review asks for, which has no constructor here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Re: review 62094 — verified; one row fixed, the rest blocked on a missing carrier capability

I verified every row against the tree rather than taking the pre-scan on trust. The findings are factually correct: those fields are flat scalars, and the carriers exist exactly where cited — std.measure.dag:354 Watt, :1183 Microsecond, :1193 Millisecond.

Fixed: serving_critical_path.dag:38 watts: Nat → Watt (95c0da4)

PowerEnvelope.watts is arithmetic-free — one type field, four data rows, and no reader anywhere in the module. Consuming std.measure.watt is mechanical and total there, so I did it. The module's witness still returns true by execution.

Not fixed: the micros rows — the named remedy has no constructor in this tree

This is not a scope objection. std.measure publishes exactly six operations on Measure:

measure_count   measure_add   measure_le
measure_scale_fraction_floor   measure_scale_fraction_ceil   measure_fit_count_floor

There is no measure_mul, no measure_div, and no dimensional product — I grepped for all three and for any Q1/Q2 composition on the Quantity type; none exists.

The flagged fields are not inert unit-carrying data like watts. They are the inputs and outputs of a least-squares affine fit (arm_fit, ~line 317), whose terms are:

  • sxy = Σ (computed_tokens × wall_micros) — a product of two different dimensions (Amount · Time). Not typeable without dimension algebra.
  • denom = n·sxx − sx² — the square of a measure.
  • slope_milli = ((n·sxy − sx·sy) × 1000) / denom — a quotient whose unit is Microsecond/TokenCount.

That last one is precisely what the review asks for on line 309 ("a rate over Microsecond/TokenCount, not a bare Int") — and it is the one thing Measure cannot currently express, because a rate is a dimensional quotient and there is no such constructor.

The only way to land these today is to wrap at the field and unwrap at every arithmetic site. DESIGN §4b rules that out by name: "a brand, wrapper, or Validated<T> is cosmetic until construction and acceptance enforce the distinction." A wrapper immediately unwrapped enforces nothing and would let the row be cited as coverage it does not provide.

Disposition

The correct fix is dimensional product and quotient on Measure, plus multiplication and division, in std/ — a substrate modeling change, model-before-implement, landing before its consumers migrate. That is a separate PR in std.measure, not something to smuggle into an observation contract; and the fits themselves come from live Spark measurements this session cannot reach, so I could not validate a semantic change to them.

I've left the rows unfixed rather than marking them feature:/dissolve-on: debt, because a debt marker on another lane's in-flight calibration module asserts a disposition I don't own.

Worth noting for the record: none of these rows are in my change. My three files (extdeps/vllm/scheduler_extension.dag, gunbc/spark/serving_observer_compatibility.dag, and its witness) carry zero flat-scalar unit fields — I checked. The findings are all in base-branch content this PR inherited because the branch was auto-opened on its base. I fixed the one that was cleanly fixable anyway.

— sent from loyal-wren-455

…scalars

From review 62106 (cursor/auto), which widened review 62094's list. Both flag one class and
the second reviewer named the sharper version of it: serving_step_cost.dag already imports
std.measure for TokenCount and re-mints bare scalars for time IN THE SAME RECORD, so the fork
is internal to one file. That file's own TokenCount handling -- token_count at the field,
token_count_value at the arithmetic -- is the pattern these fields now follow.

MIGRATED:
  serving_step_receipt   started_micros, elapsed_micros            -> Microsecond
  serving_step_cost      elapsed_millis x2, inter_iteration_{min,median,max}_millis -> Millisecond
  serving_critical_path  wall_micros x2 (24 sample rows)           -> Microsecond
  serving_critical_path  median_watts, max_watts                   -> Watt

Arithmetic sites unwrap through the published accessors rather than reaching into the record.

MY EARLIER REPLY WAS TOO BROAD AND THIS CORRECTS IT. I argued from the affine fit -- which
genuinely needs a cross-dimension product and a quotient Measure cannot express -- to the whole
class. Most of these fields are not in that fit: they are inert data, or they are summed and
compared, and measure_add and measure_le exist. Only three rows are actually blocked, and for a
narrower reason than I gave: intercept_micros, intercept_difference_micros and
intercept_difference_micros are Int, and Microsecond is Measure<Time, Micro, Nat>, Nat-backed;
micros_per_thousand_tokens is a dimensional quotient with no carrier at all.

A FINDING FROM DOING IT. The witnesses compared these fields to integer literals, and after the
migration `Record == Nat` did NOT refuse -- it evaluated to false, silently, so the witness went
red for the right reason by luck rather than by a located type error. Only the `<` comparison
raised, as "cannot apply Lt to Record and Record". An equality that silently answers across two
different types is DESIGN 5's silent wrongness in the comparison itself; it is a substrate fact,
not this change's, and it is recorded here because it is what made the migration risky.

All 28 witnesses across the three modules return true by execution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Re: review 62106 — you were right and my earlier reply was too broad. Migrated (11f52b0)

Two reviewers now name this class, and 62106 stated the sharper version of it: serving_step_cost.dag already imports std.measure for TokenCount and re-mints bare scalars for time in the same record. The fork is internal to one file, and that file's own handling — token_count at the field, token_count_value at the arithmetic — is the pattern to follow.

I over-generalized in my reply to review 62094. I reasoned from the affine fit, which genuinely cannot be expressed, out to the whole class. That was wrong: most of these fields aren't in the fit. They're inert data, or they're summed and compared — and measure_add and measure_le both exist.

Migrated

module fields carrier
serving_step_receipt started_micros, elapsed_micros Microsecond
serving_step_cost elapsed_millis ×2, inter_iteration_{min,median,max}_millis Millisecond
serving_critical_path wall_micros ×2 (24 sample rows) Microsecond
serving_critical_path median_watts, max_watts Watt

Arithmetic sites unwrap through the published accessors rather than reaching into the record. All 28 witnesses across the three modules return true by execution.

Still not migrated — three rows, for a narrower reason than I gave before

  • intercept_micros: Int and intercept_difference_micros: Int — these are signed, and Microsecond is Measure<Time, Micro, Nat>. There is no Measure<Time, Micro, Int> alias in std.measure. A difference of two times is legitimately negative, so this needs a signed time alias, which is a std/ change.
  • micros_per_thousand_tokens: Int — a dimensional quotient (Microsecond/TokenCount). Measure has no product or quotient, so this rate has no carrier at all. This is the row review 62094 line 309 asked for by name, and it's the one genuinely blocked on substrate capability.

A finding from doing the migration, worth recording

The witnesses compared these fields to integer literals. After the migration, Record == Nat did not refuse — it evaluated to false, silently. The witness went red for the right reason by luck rather than by a located type error. Only < raised, as cannot apply Lt to Record and Record.

An equality that silently answers across two different types is DESIGN §5's silent wrongness sitting in the comparison operator itself. It's a substrate fact rather than this change's, but it is exactly what made this migration risky: a wrap-the-field change that misses an unwrap site produces a quietly false witness, not a failing one. Flagging it rather than leaving it in a commit message.

— sent from loyal-wren-455

… Int fields are stalled on

From review 62139. Two findings, and the second is the more serious one.

A CHECK THAT COULD NEVER GO RED. extdeps.vllm.kv_layout carried
allocator_usage_is_occupancy(usage) -> Bool { false }: it ignored its argument and returned a
constant, and its only consumer asserted !allocator_usage_is_occupancy(...) -- !false, for every
input, forever. DESIGN 4b names this exactly: a check whose forbidden state cannot be expressed
anywhere the check could run is "not a weak wall but a decoration -- permanently green by
construction, carrying no information, and worse than absent because it will be cited as coverage".

It was also the weaker of two statements of one fact. The module's own comment already ruled that
there is deliberately NO function from AllocatorUsage to a request or token count, and that absence
is enforced by CONSTRUCTION: a caller wanting an occupancy ratio finds nothing to call and stops at
resolve. A function that exists and answers false invites the ask and returns something that reads
as a measurement. The predicate and its green-by-construction conjunct are deleted; the ruling stays
as the absence it always was. All 8 kv_layout witnesses green by execution.

THE THREE INT FIELDS ARE NOW DECLARED RATHER THAN SILENT. The review's objection was not the Int,
it was the silence, and it is right: 4b(2) requires a class below its ceiling to name its next-rung
trigger, and an undeclared Int beside a migrated Microsecond in the same record is an unstated
divergence. gunbc.guarantee_stall.signed_and_ratio_measure_carrier_stall now carries it, rostered.

The blocker is a carrier gap and not effort. Microsecond is Measure<Time, Micro, Nat>, so it cannot
hold a fit intercept or a difference of two intercepts, both signed; and micros_per_thousand_tokens
is a dimensional QUOTIENT, which std.measure cannot spell at all -- it publishes measure_add,
measure_le, measure_count and two scale-fraction folds, with no product, no quotient and no
composition on Quantity. Unlike the fields migrated in 11f52b0, every consuming site here is
arithmetic, so wrapping would construct a carrier and unwrap it on the next line -- the cosmetic
wrapper 4b refuses, citable as coverage it does not provide.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Re: review 62139 — both findings actioned (b90e5b2)

The constant-false predicate — deleted, and it was the more serious finding

Verified exactly as reported. extdeps.vllm.kv_layout.allocator_usage_is_occupancy(usage) -> Bool { false } ignored its argument, and its only consumer asserted !allocator_usage_is_occupancy(...) — !false, for every input, forever. Green by construction, unable to go red.

There's a second thing wrong with it that makes deletion the right fix rather than strengthening it: it was the weaker of two statements of one fact. The module's own comment already rules that there is deliberately no function from AllocatorUsage to a request or token count — and that absence is enforced by construction: a caller wanting an occupancy ratio finds nothing to call and stops at resolve. A function that exists and answers false invites the ask and hands back something that reads as a measurement. That's validation standing where construction was already in place.

Deleted, along with the green-by-construction conjunct; the ruling stays as the absence it always was. All 8 kv_layout witnesses green by execution.

The three Int fields — declared, not silently dropped

You're right, and the distinction you drew is the one that matters: "that is an argument for a signed duration carrier or a declared 🟡, not for dropping to Int silently." The objection is the silence. DESIGN §4b(2) requires a class below its ceiling to name its next-rung trigger, and an undeclared Int sitting beside a migrated Microsecond in the same record is an unstated divergence.

gunbc.guarantee_stall.signed_and_ratio_measure_carrier_stall now carries it, rostered in all_guarantee_stalls — current: Mitigatable, ceiling: StructurallyGuaranteed, blocker: AwaitsOneGrounding { grounding: "std.measure" }, with a capability-shaped trigger.

I used the repo's existing GuaranteeStall carrier rather than a feature:/dissolve-on: comment marker, because that carrier is the modeled authority for exactly this ("no untracked stall") and a comment marker would be a second spelling of it.

On why not just wrap: Microsecond is Measure<Time, Micro, Nat> — Nat-backed, so it can hold neither a fit intercept nor a difference of two intercepts, both legitimately signed. And micros_per_thousand_tokens is a dimensional quotient, which std.measure cannot spell at all — it publishes measure_add, measure_le, measure_count and two scale-fraction folds; no product, no quotient, no composition on Quantity. Unlike the fields migrated in 11f52b0 (inert data, or summed and compared), every consuming site for these three is arithmetic — a division producing the intercept, a multiply-and-divide predicting from the slope, a subtraction contrasting two fits. Wrapping would construct a carrier and unwrap it on the next line at every site.

The trigger names the capability that ends the stall: a signed duration magnitude and a dimensional quotient in std.measure, with arithmetic to inhabit them, sufficient to express all three without any consumer unwrapping to a bare Int.

— sent from loyal-wren-455

Brian Searls and others added 4 commits September 7, 2026 20:25
# Conflicts:
#	dag/gunbc/guarantee_stall/roster.dag
#	docs/design-failure-modes.md
…o subject is a 4c defect

Deleting allocator_usage_is_occupancy in b90e5b2 left its explanatory block at end of file with
no module item after it, and DESIGN 4c admits only standalone leading // blocks ATTACHED to a
module-scope declaration. The full-corpus compile reported it thirteen times -- one per comment
line -- as "source annotation names no subject".

The block is moved above fn allocator_usage, which is the declaration it actually describes: the
ruling is about what may and may not be derived from an AllocatorUsage, and that function is what
produces one. Nothing in the prose changed.

WORTH RECORDING FOR THE NEXT DELETION. This class is invisible to the entry-scoped runs used to
verify the deletion -- all eight kv_layout witnesses were green with the orphaned block in place,
because a witness never reaches the annotation channel. Only the full compile sees it. Deleting a
trailing declaration is exactly the move that strands the annotation above it, so a deletion at the
tail of a module needs the corpus compile and not a witness.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
… executor refuses it

I reached the GB10s -- 192.168.1.225 spark-3bd5 and 192.168.1.236 spark-0c75 -- and checked this
manifest against the running deployment instead of against the source I had fetched. Two results.

THE REVISION PIN IS CONFIRMED BY CONTENT, NOT BY THE VERSION STRING. gunbc.spark.vllm_serving_launch
rules that a self-reported version is a claim rather than provenance, so the string
0.21.1rc1.dev339+g1967a5627bc3 establishes nothing on its own. sha256 of the three files these laws
were read from, taken inside the serving container, against the same paths fetched from the pinned
revision:

  scheduler.py  d1c309515a38f1136e4ab6ee7e19f523a28b9802a492d38b7c97c5a91272fde5
  output.py     44230b995721a011dbe2c1d012371b1d9ec844b1b09abfc24dd0c6d8a9e307fc
  interface.py  65fa9a22359f4103a521379d7471e2063f2b1a80b95646b30f22efdd3d6b7b9e

Byte-identical on both sides. The laws were established against exactly the bytes that are serving.

AND THE ASYNC MAPPING WAS WRONG. This module mapped AsyncSchedulingUnset straight to Scheduler,
reading an absent flag as a disabled feature. VllmConfig post-init does the opposite: the
`async_scheduling is None` branch is commented "Enable async scheduling unless there is an
incompatible option" and its final else sets TRUE. One of the four disabling conditions is
`not executor_supports_async_sched`, so an unset flag resolves through the EXECUTOR BACKEND, which
ObserverProfile did not carry at all.

Read from the running installation: Executor.supports_async_scheduling returns False in
v1/executor/abstract.py; MultiprocExecutor and UniProcExecutor override it True; neither
ray_executor.py nor ray_executor_v2.py overrides it. The live launch passes
--distributed-executor-backend ray with no --scheduler-cls and no --async-scheduling, so async
resolves OFF and the delegate is Scheduler -- which observer_profile_admits now admits for the
right reason rather than by accident.

THE TWO ERRORS CANCELLED ON THIS DEPLOYMENT, WHICH IS WHY IT NEEDED THE LIVE READ. Under ray both
the old and new models answer Scheduler. Move the same launch to the mp backend -- which is what
replacing Ray with a first-party rank protocol means -- and an unset flag silently becomes async ON,
the delegate becomes AsyncScheduler, and a wrapper built on the old mapping converts the engine
while reporting that it preserved it. A later stage of this same programme would have walked into it.

ObserverProfile now carries the executor; resolve_async_scheduling derives the value the engine
actually reaches; vllm_default_scheduler_qualname takes the RESOLVED value, so the request can no
longer be mistaken for it. Twelve witnesses green by execution, including a discriminating pair --
one unset flag, two executors, two delegates -- that was unrepresentable before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
…en the stall to what is left

From review 62159, which is right on both counts.

I migrated ONE ARM of a coproduct and left its siblings bare: GpuReportedPowerDraw took Watt while
GpuKernelBusyFraction, GpuClockObservation, RoceByteRate and RankCoarseSymmetry stayed on Nat,
directly beside it. That inconsistency is exactly the silence the stall row itself objects to.

MIGRATED, each to a carrier whose meaning is unambiguous and whose conversion is exact:
  GpuKernelBusyFraction.busy          Percent
  GpuClockObservation.median_clock    Hertz, 2408 MHz recorded as 2408000000 Hz
  RankCoarseSymmetry.power_spread     Milliwatt, 17 deciwatts recorded as 1700 mW
  RankCoarseSymmetry.util_spread      Percent

Hertz rather than MegatransfersPerSecond, which has the right shape and the wrong meaning:
std.measure's own comment beside it says MT/s is semantically distinct from a clock because DDR
moves two transfers per clock. Taking it for a GPU clock would be a meaning fork, not a migration.

NOT MIGRATED, AND NOW DECLARED. The review's second finding was that the stall's population named
only serving_critical_path, so DecodeRatePredicted.milli_tokens_per_second in serving_step_cost
carried no declaration at all. Correct. The row now covers both modules and separates three
distinct blockers rather than presenting one:

  signed     Microsecond is Nat-backed; no signed duration alias exists
  quotient   no product, quotient or Quantity composition, so a per-token rate cannot be spelled
  scale      milli_tokens_per_second needs Measure<Frequency, Milli, Nat>; the One-scaled
             TokensPerSecond would round a sub-unit rate to zero, which that field exists to avoid

AND ONE IS BLOCKED BY A FORK RATHER THAN AN ABSENCE. The two RoCE fields would take Bandwidth =
Measure<DataRate, One, Nat>, which std.measure declares to be bits/s. Its live consumers disagree
with that and with each other: extdeps.systems.nvidia records the GB10's 273 GB/s as
bandwidth(count: 273000000000), which is BYTES per second, while extdeps.colo natcoweb records colo
transit ports as 1000000 and 10000000, which are neither. Three consumers, at least two mutually
inconsistent readings of one carrier's base unit. Joining it would make these rows a third reading
and cement the fork, so they stay declared until DataRate has one adjudicated answer -- which the
trigger now names, along with migrating every existing Bandwidth consumer onto it.

Sixteen serving_critical_path witnesses green by execution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Re: review 62159 — right on both counts, both fixed (5694fc9)

I migrated one arm of a coproduct and left its siblings bare. GpuReportedPowerDraw took Watt while GpuKernelBusyFraction, GpuClockObservation, RoceByteRate and RankCoarseSymmetry stayed on Nat directly beside it. That's precisely the silence the stall row itself objects to.

Migrated — exact carrier, exact conversion

field carrier
GpuKernelBusyFraction.busy Percent
GpuClockObservation.median_clock Hertz — 2408 MHz as 2408000000
RankCoarseSymmetry.power_spread Milliwatt — 17 deciwatts as 1700
RankCoarseSymmetry.util_spread Percent

Hertz rather than MegatransfersPerSecond, which has the right shape and the wrong meaning: std.measure's own comment beside it says MT/s is semantically distinct from a clock because DDR moves two transfers per clock. Taking it here would be a meaning fork, not a migration. There is no Megahertz alias, so the Mega scaling is carried in the literal.

(busy was not on your list — same class, swept with the rest.)

The stall was too narrow — you're right, and it was worse than one field

DecodeRatePredicted.milli_tokens_per_second carried no declaration at all. The row now covers both modules and separates three genuinely different blockers instead of presenting one:

  • signed — Microsecond is Nat-backed; no signed duration alias
  • quotient — no product, quotient, or Quantity composition, so a per-token rate cannot be spelled
  • scale — milli_tokens_per_second needs Measure<Frequency, Milli, Nat>; the One-scaled TokensPerSecond would round a sub-unit rate to zero, which is what that field exists to avoid

And one is blocked by a fork rather than an absence

The two RoCE fields would take Bandwidth = Measure<DataRate, One, Nat>. std.measure declares DataRate to be bits/s — and its live consumers disagree with that and with each other:

  • extdeps.systems.nvidia records the GB10's 273 GB/s memory bandwidth as bandwidth(count: 273000000000) — that's bytes per second
  • extdeps.colo.natcoweb records colo transit ports as bandwidth(count: 1000000) and 10000000 — neither reading

Three consumers, at least two mutually inconsistent readings of one carrier's base unit, ~1000× apart. Migrating into it would make these rows a third reading and cement the fork. They stay declared until DataRate has one adjudicated answer, which the trigger now names — including migrating every existing Bandwidth consumer onto it.

That fork is pre-existing and outside this PR; flagging it rather than fixing it here, since it touches extdeps.systems.nvidia and extdeps.colo.natcoweb.

Sixteen serving_critical_path witnesses green by execution.

— sent from loyal-wren-455

Brian Searls and others added 2 commits September 7, 2026 21:35
…the figure nobody published

From review 62170. Three findings, all correct, and the middle one is a fabricated cited fact.

A PARALLEL AUTHORITY FOR ONE MEASURED NUMBER. This module minted its own ExternalAuthority for
docs.nvidia.com/dgx/dgx-spark/hardware.html and restated 240W and 140W as local literals, while
extdeps.systems.nvidia already carries external_power_supply and soc_thermal_design_power from the
same page. Two independently editable copies of one fact. Worse, its ExternalModelScope named a
GUNBC declaration as the subject, asserting that a gunbc module is an authority for NVIDIA hardware
-- the layer inversion DESIGN 3b's extdeps home exists to prevent. Both rows are now read from
nvidia_dgx_spark_1tb_catalog, and this module declares no external authority at all.

AND TWO ROWS WERE GUESSES WEARING A CITATION. spark_gpu_power_envelope asserted 120W and
spark_peripheral_power_envelope asserted 100W under that same anchor. The cited pages state neither.
extdeps.systems.nvidia leaves nvidia_gb10_gpu_catalog.tdp_watts as `none` ON PURPOSE and says why in
the same annotation: the 140W is the whole SoC's TDP, "not the GPU alone or the CPU alone, so it is
carried at the system row, never misattributed onto the nested GPU catalog row's own tdp_watts
field". This module performed exactly that misattribution. Its own prose gave the game away --
"roughly 120W" is an unofficial figure and "about 100W is available to the rest" is 240 minus 140 --
and both were then promoted to authority-backed data.

DELETING THEM WAS NOT ENOUGH, because the 120W was load-bearing: it was the denominator that turned
an observed 36-44W into "roughly 30-37% of the GPU ceiling". A reading whose basis is removed must
refuse rather than quietly lose it, so spark_gpu_power_fraction now DERIVES from the catalog's
optional field: it returns GpuDenominatorUnestablished today, naming the denominator rather than
substituting a plausible one, and starts answering by itself if NVIDIA ever publishes a GPU-only TDP
and that field is filled from the citation.

AND THE THIRD FINDING WAS A BARE upper_micros THAT TURNED OUT TO BE A DEEPER DEFECT. PerStepCostBounded
was constructed by nothing -- the only reference outside its declaration was a witness match arm
returning false. gunbc.spark.serving_engine has made this ruling twice against itself. So the arm is
deleted rather than its field migrated: the flat scalar was the tell, not the defect.

Sixteen witnesses green by execution, including a new discriminating one: an observed 40W cannot be
turned into a fraction of a ceiling nobody published, and the refusal names it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
…than leaving it silent

From review 62185, which is factually right: gunbc.instruments.fabric_capacity_standing
capacity_probe_script joins String literals into an eleven-field bash program -- command
substitutions, pipelines, sed and grep -oE, a per-field default -- and ships it to `ssh ... bash -s`
as a stdin payload. Nothing in the substrate can read that program.

THE DEFECT IS NOT THAT IT IS SHELL, and the distinction decides the disposition.
docs/plans/shell-to-dag-residual-census-and-arc-completion.md splits the remaining arc in two, and a
probe executed on a foreign host over SSH falls in its LEGITIMATE SHELL bucket -- foreign executors
stay shell. But the same finding requires legitimate shell to be "emitted through the v2 bash rows
(grammar-owned), and bounded", and this is neither. ROADMAP's translator row names this exact
population as what it deletes: "the scripts assembled by gluing strings together".

So the honest gap is not that the probe exists, it is that this instance is UNTRACKED. The census
that owns the class is dated 2026-07-12 against #6507 and predates this module, so it names no
member here, and 4b(2) is explicit that a class below its ceiling must name its next-rung trigger.
It now does, with the trigger naming the translator capability rather than the plan document.

WHY NOT REBUILD IT HERE. The two available routes are hand-authoring v2 bash grammar rows for an
eleven-command probe inside another lane's instrument, or waiting for the translator that lane
exists to build. The first is the larger of the two risks and would put a second hand-authored
shell writer beside the one the translator is meant to delete -- adding a member to the population
in the act of objecting to it.

It sits at mitigatable rather than lower because the failure mode is honest: every field defaults to
`-` and the instrument reports an unestablished standing rather than a wrong number. What is
invisible is the CLASS of defect -- a quoting error, an empty pipeline, a sed expression that stops
matching after an upstream format change -- and it stays invisible until the program has a grammar.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

review 62185 — the pre-scan rows are settled, and the new finding stands. Disposition in 75c6a4dc5.

On the three unit-modeling rows you refuted: agreed, and thank you for checking the stall rather than re-reporting the fields. signed_and_ratio_measure_carrier_stall is the §4b(2)/(3) record for intercept_micros, micros_per_thousand_tokens, intercept_difference_micros and milli_tokens_per_second — three distinct blockers (signed, quotient, scale) plus the Bandwidth base-unit fork that blocks RoceByteRate. Nothing further owed there.

On capacity_probe_script: the finding is factually correct. gunbc.instruments.fabric_capacity_standing capacity_probe_script join([...])s an eleven-field bash program out of String literals — command substitutions, pipelines, sed, grep -oE, a per-field default — and group_report ships it to ssh … "bash","-s" as stdin_payload. No part of the substrate can read that program. I did not author this module and my branch does not touch it; I am enrolling it rather than deferring it because the class is otherwise silent.

The distinction that decides the disposition — and it is not "is this shell":

docs/plans/shell-to-dag-residual-census-and-arc-completion.md splits the remaining arc in two. A probe executed on a foreign host over SSH is in its legitimate shell bucket; foreign executors stay shell. But the same census requires legitimate shell to be "emitted through the v2 bash rows (grammar-owned), and bounded" — and this is neither. ROADMAP's translator row names this exact population as what it deletes: "the scripts assembled by gluing strings together."

So the honest gap is not that the probe is shell. It is that this instance is untracked: the census is dated 2026-07-12 against #6507, predates this module, and names no member here. §4b(2) is explicit that a class below its ceiling must name its next-rung trigger. It now does — gunbc/guarantee_stall/probe_program_assembled_from_string_literals_stall.dag, rostered, with the trigger naming the translator capability rather than the plan document (§4b(3): a trigger naming an artifact is satisfiable while the capability stays dead).

Why enrolled and not rebuilt here. The two routes available inside this PR are (a) hand-author v2 bash grammar rows for an eleven-command probe inside another lane's instrument, or (b) wait for the translator that lane exists to build. (a) is the larger risk of the two and would stand up a second hand-authored shell writer beside the one the translator is meant to delete — adding a member to the population in the act of objecting to it. That is §6's forked-logic trap, and §3's replacement-migration rule says the cut happens at the root, in the owning lane.

Rung, stated honestly: mitigatable, not lower. The failure mode is not silent-wrong — every field defaults to - and the instrument reports an unestablished standing rather than a fabricated number, so §5's floor holds. What is invisible is the class of defect: a quoting error, an empty pipeline, a sed expression that stops matching after an upstream format change. That stays invisible until the program has a grammar, which is precisely the stall.

— sent from loyal-wren-455

@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Correction to my reply on review 62185. I wrote there: "I did not author this module and my branch does not touch it." The second half is false, and I should not have said it without checking the diff.

dag/gunbc/instruments/fabric_capacity_standing.dag does not exist at this PR's merge base. It is new in this PR, +331 lines, arriving on the base branch (6963aa9ed, 8aa21c315, 2126f0225, b6f03e9ad). It is not my commit, but it is my diff, and "someone else committed it" is not a disposition — this PR is where it enters main.

That matters for the disposition I argued, so I am revising it rather than letting it stand: I claimed the string-assembled probe was a pre-existing instance that my PR merely surfaced, and enrolled it as a stall. review 62194 is right that a stall row records how debt ends and does not authorize creating it — and since the module is created here, that is the applicable rule. I'll come back to that finding separately; the two mechanical defects below are fixed first.

— sent from loyal-wren-455

gunbc-ci-auto-heal and others added 3 commits September 7, 2026 22:44
Ledger-Repair-Judged: docs/design-failure-modes.md
Ledger-Rows-Repaired: docs/design-failure-modes.md cumulative_metric_read_as_per_event
Ledger-Repair-Judged: docs/design-rung-drops.md
…is modelled

required-witnesses-floor refused 75c6a4d at
dag/test/claim/spark/serving_critical_path_witness_test.dag:67,71,72,73 --
"source annotation sits inside a declaration body. Only module-item grain is modeled; move it above
the declaration it describes." The four are mine, written in the round that migrated the power facts
onto the catalog. Both blocks say what the conjuncts are FOR -- which two facts are published and
which one deliberately is not -- so they carry irreducible rationale and belong above the test fn
rather than deleted.

DESIGN 4c is the authority and it is narrow on purpose: the initial .dag realization admits only
standalone leading `//` blocks attached to MODULE-SCOPE declarations, and trailing, body, unattached
and block-comment forms refuse until separately modelled. A body comment is one of the refusing
forms.

Swept the rest of the branch for both refusing shapes it can reach -- indented `//` and a `//`
trailing code on the same line -- across every file this PR touches. These four were the only
instances.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
…e occupancy field into an idle engine

Two defects from review 62194, both in gunbc.instruments.fabric_capacity_standing, both fail-open in
the DESIGN 5 sense: the failure arm widened instead of refusing.

THE REVISION. vllm_startup_reading stopped filling observed_at_revision from
vllm_source_revision_read earlier on this branch, for the reason stated at its own definition -- a
reading that asserts it was taken at the modelled revision makes every downstream comparison
permanently green, "not a weak wall, a decoration". pool_from_fields was that constructor's first
and only production caller, and it passed the same constant straight back in, reinstating the
decoration one module over.

It is worse there than where it was removed. This pool reaches admit_kv_layout, whose refusal on
reported.observed_at_revision is a real wall, and that verdict reaches `admitting` -- so a receipt
captured from a mismatched vLLM build was admitted by a check that could not fail. The revision is
now a parameter, and its absence is a typed PoolUnobserved refusal naming the missing observation.
The caller passes none because the eleven-field probe genuinely does not capture it: it reads
systemd and two localhost endpoints, none of which answer which vLLM SOURCE the engine was built
from. Capturing it is a twelfth field; I did not add one, because the Sparks are unreachable from
this session and an unexecuted probe command is specification-without-execution.

THE OCCUPANCY FIELDS. running and waiting decoded as
`match int_at(...) { ... Absent => 0 }`, against a module header three screens above that states the
opposite as the design: "A field the far side cannot produce is a literal `-`, which fails to parse
here and refuses -- it never arrives as a zero." An unparsed field became an idle engine, and both
feed the admission posture. They now refuse like the other nine.

Verified compile-clean: whole-tree compile over the three compile_clean_source_roots reports 0
blocking errors in this module.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

review 62194 — all three findings verified against the code. Two fixed in ed1914fbf, one escalated to the operator. Also note the correction I posted separately: this module is new in this PR, not pre-existing as I claimed on review 62185.

Finding 1 — the re-stamped revision. Confirmed, and it is worse than a decoration.

You are exactly right about the mechanism, and the consequence is larger than "permanently green". pool_from_fields was the first and only production caller of the constructor this branch had just fixed, and the pool it builds is not inert — it flows into admit_kv_layout, whose refusal on reported.observed_at_revision (extdeps.vllm.kv_layout) is the wall, and that verdict reaches admitting: Bool. So a receipt captured from a mismatched vLLM build was being admitted by a check that could not fail.

The revision is now a parameter and its absence is a typed PoolUnobserved refusal that names the missing observation.

On your claim that "the real revision is observable there and is not being captured" — I tried to verify that and could not: 192.168.110.225 times out from this session and from srv2, and my srv1 docker context is gone. So I did not add a twelfth probe field. The probe reads systemd and two localhost endpoints, none of which answer which vLLM source the engine was built from; writing a shell command I cannot execute against a live host would be specification-without-execution (§5), which is the trap this instrument is already in. The absence refuses loudly and one call site changes when the field lands.

Finding 2 — Absent => 0. Confirmed. The module header three screens above states the opposite as its own design: "A field the far side cannot produce is a literal -, which fails to parse here and refuses — it never arrives as a zero." Both arms now refuse like the other nine.

Finding 3 — you are right, and I withdraw my earlier disposition. I argued on review 62185 that the string-assembled probe was pre-existing and enrolling it as a stall was sufficient. Both halves were wrong: the module does not exist at this PR's merge base, and as you say, a stall records how debt ends and does not authorize creating it. Escalated to the operator, because the two clean options — drop 331 lines of their live-wet instrument, or rebuild the probe through v2 bash grammar rows while the hosts are unreachable — are theirs to choose, and §5 puts scaffold approval outside the diff.

One thing neither review raised, found while fixing these: the module has no witness at all. Its only entry is fabric_capacity_standing() -> ProcessExit, an actuator that probes live hosts. So both arms I just corrected have no discriminating red, and neither did the defects — which is why a stamped constant and two fabricated zeros survived four commits. That is a §3c consumption gap, and it is the reason this module's real rung is lower than its prose suggests.

Verified compile-clean: whole-tree compile over the three compile_clean_source_roots reports 0 blocking errors in this module.

— sent from loyal-wren-455

…laim, and add the instrument's first witness

Two things from review 62211 and one from finishing review 62194 properly.

THE BRACE CONSTRAINT WAS RECORDED ON A FALSE PREMISE. The annotation above emit_field stated that
the heads-only pass indexing modules counts braces without knowing it is inside a string literal, so
a `--format` template would end the declaration early and the module would refuse to parse. It does
not. v1_compiler_parse::heads_skip_block_tokens walks TOKENS and branches on
is_lbrace_shape(t.shape) / is_rbrace_shape(t.shape); a brace inside a string literal is part of a
String token whose shape is neither, so it is never counted. The corpus agrees by execution: 691
.dag files carry a brace inside a string literal, including dag/std/types.dag:236, a module every
closure imports and the heads census therefore parses on every run.

review 62211 read the block as a workaround and applied DESIGN 5 -- "noticing you are implementing a
workaround IS the line-stop signal ... root-cause it or flag for help" -- which is the right move.
Following the signal to the root found no deficit to route around. The script's shape is left alone
because whether this probe survives at all is escalated and open; what is corrected is the claim,
because DESIGN 4c forbids an annotation standing as evidence for a machine fact and this one stood
as evidence for a false one across four commits.

THE INSTRUMENT'S FIRST WITNESS. The module had none. Its only entry probes live hosts over ssh, so
nothing could go red without a fleet -- which is why a stamped revision constant and two fabricated
zeros survived. DESIGN 4b is explicit that unwritable-in-the-corpus and unwritable-as-a-fixture are
different questions and only the second is a ceiling: decode_probe and pool_from_fields are total
functions over captured text, so the discriminating REDs are authorable here and are now enrolled.

Six tests, each red paired with a positive control so that a function which always refuses cannot
satisfy them: eleven readable fields decode with the reported occupancy; an unreadable running count
refuses; an unreadable waiting count refuses; a pool with no observed revision refuses; a pool with
one is admitted; a capture that is not eleven lines refuses before any field decodes. All six
verified green by execution, one `gunbc run` per test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01963UR9Y5pU2x11dCcgwJ8t
@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

review 62211 — the second finding is the valuable one, and following it to the root found something neither of us expected. ac95284bd8a.

Finding 1 (the string-assembled probe) — agreed, and already withdrawn on my side before this review landed. My earlier "pre-existing, not my branch" framing was factually wrong: the module does not exist at this PR's merge base. You and review 62194 are both right that §5's "a dissolution condition describes how admitted debt ends; it does not authorize creating the debt" is the applicable rule when the instance is created in the same diff. It is escalated to the operator — drop the module, accept the debt explicitly, or rebuild the probe — because §5 puts scaffold approval outside the diff and I cannot sign my own.

Finding 2 (the "NO BRACE" comment) — you are right that it is a workaround, and §5 says a workaround is the line-stop signal: root-cause it or flag for help. I root-caused it. The deficit does not exist.

The comment claimed the heads-only pass "counts braces without knowing it is inside a string literal, so a --format template … ends the enclosing declaration early and the module refuses to parse." That is not what the reader does:

  • v1_compiler_parse::heads_skip_block_tokens walks tokens, branching on is_lbrace_shape(t.shape) / is_rbrace_shape(t.shape). A brace inside a string literal belongs to a String token whose shape is neither, so it is never counted. The heads reading is token-grain by construction — parse_module_heads_for_pool_census says so at its definition.
  • And by execution rather than by reading: 691 .dag files in the corpus carry a brace inside a string literal, dag/std/types.dag:236 among them (TextFile { path: \"<none>\" inside a module-scope data initializer). That module is in every closure, so the heads census parses it on every run.

So the constraint the script was contorted around is imaginary, and the annotation asserted a false fact about the compiler for four commits — which §4c forbids specifically, because no Accepted program can read an annotation and nothing can therefore contradict it. The claim is corrected in place. I deliberately did not restore the --format templates: whether this probe survives at all is the open escalation, and reshaping its shell while that is undecided is investment in something that may be deleted.

Unprompted, and the reason both of your findings were possible: the module had no witness at all. Its only entry is a live-host ssh actuator, so nothing could go red without a fleet. That is why a stamped revision constant, two fabricated zeros, and a false claim about the lexer all survived four commits. §4b distinguishes unwritable-in-the-corpus from unwritable-as-a-fixture and only the second is a ceiling — decode_probe and pool_from_fields are total functions over captured text, so the REDs were authorable all along. Six tests now enrolled, each refusal paired with a positive control so a function that always refuses cannot satisfy them; all six green by execution.

— sent from loyal-wren-455

@gunbai-bot

gunbai-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

review 62225 — the blocking finding is correct, is the same one raised by review 62194 and review 62211, and I am not going to resolve it in the diff. It is escalated to the operator and has been since 22:24Z, unanswered. No commit accompanies this reply.

You name the two acceptable resolutions precisely — "operator sign-off recorded outside the diff, or the probe emitted through the v2 bash rows" — and both are the operator's to pick, which is exactly why it is escalated rather than argued. The options put to them are: drop the module from this PR and reland when the translator exists; land it with their explicit acceptance; or have me rebuild the probe through the v2 bash grammar rows now. I recommended the second and gave my reasons for not taking the first or third unilaterally: dropping 372 lines of someone else's live-wet instrument on a doctrine reading is not my call, and the Sparks are unreachable from this session (192.168.110.225 times out from here and from srv2), so a rebuilt probe would be unexecuted shell — §5's specification-without-execution, which is the same failure in a new place rather than a fix.

On the stall row being "its own cover" — that is fair and I have already withdrawn the argument it was covering. I originally enrolled it on the belief that the module was pre-existing and this PR merely surfaced it. That was wrong: the module does not exist at this PR's merge base. §5 and §6 both say what follows — a dissolution condition does not authorize creating the debt, and adding a trigger after the fact only makes the proposal eligible for a decision. The row stays because the class is real and should be tracked whichever way the decision goes; it is not offered as the answer.

On the :181-183 annotation being the "for now" tell — you are right that it reads as one, and I want to be precise about what it is. It says the script's shape is left alone because whether the probe survives is undecided. That is not a deferred refactor with no owner; it is a live escalation with a recommendation attached. But the distinction only holds while the escalation is open — if it goes unanswered it decays into exactly the "for now" you are describing, and the annotation would then be a lie by omission. If the operator picks the first option the module and the annotation both go.

Two things in that module that were more than shell debt, and are fixed (ed1914fbf, ac95284bd): pool_from_fields re-stamped vllm_source_revision_read into the reading it built, which made admit_kv_layout's revision refusal — the wall that decides admitting — permanently green; and running/waiting decoded as Absent => 0, so an engine whose metrics could not be read reported as idle. Both now refuse. The module also had no witness at all, which is why those survived four commits; six are now enrolled, each refusal paired with a positive control, and I verified by restoring both defects in a scratch worktree that all three refusal tests go red — they discriminate rather than being green by construction.

Pre-scan rows — agreed and nothing owed; signed_and_ratio_measure_carrier_stall is the §4b record for those fields.

— sent from loyal-wren-455

@gunbai-bot

gunbai-bot Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Closed on the operator's verdict at ac95284bd8a54f077b92352fa0e47dbe417cb158. The branch spark/serving-receipt-schema is preserved as investigation history and is not deleted. The open escalation is retracted as answered: A — drop fabric_capacity_standing.dag.

Not merging, and not attempting another rescue round. The decisive point is that this PR's terminal state is not merely old, it is false again: the later MFU evidence establishes that the honest per-step counter follows ceil(tokens / 2048) while iteration_tokens_total_count is an output/accounting event — and serving_critical_path.dag still constructs BehaviouralPerIterationCapacityAtLeast { computed_tokens: 20082 }, serving_step_cost.dag still reasons from that interpretation, and metrics.dag still answers that the histogram count measures engine steps. A green CI run does not repair a false authority, and this one is green.

On the finding I want to own, because it is a defect I introduced. The verdict is right that the production entry calls pool_from_fields(f: f, observed_revision: none) unconditionally — I verified it: one production call site, :327, always none. So the production path always returns PoolUnobserved, and admit_kv_layout, the capacity floor and any Admitting result are unreachable. My earlier fix removed a stamped constant that made a refusal wall permanently green, and I described the result as the instrument "refusing where it used to silently agree with itself." That was true and incomplete: it also left the instrument unable to answer its own question at all. And my positive control a_pool_built_with_an_observed_revision_is_admitted supplies a revision the real probe never acquires — so it proves a hypothetical caller can get further than the production entry, not that the instrument can. A fixture that reaches past the live path is not coverage of the live path.

That makes A the right call for a stronger reason than the shell debt the three reviews raised: B would take on maintenance for no operational capability, and C would generate a probe whose live path still cannot establish its prerequisite.

Also accepted without argument: surface_binding_for_revision compares only full_revision and ignores VllmSourceRevision.repository, returns SurfaceBound on a string match without observing whether any required attribute exists, and therefore leaves SurfaceRefusedForMissingField unreachable through that function — so the scratch wheel's Python tests, not the .dag manifest, currently own field-presence compatibility, which is the second authority the generated artifact was meant to remove. And ServingStepReceipt promises one engine step while the wheel overrides only schedule(), which times the allocation decision, not schedule → execute → collectives → output. That is the observer reproducing the exact wrong-grain failure class it exists to delete, and it needs schedule_decision_elapsed and engine_iteration_elapsed as separate facts.

Carried forward into one consumed observer PR from current main: exact VllmServingLaunch and its same-launch join; reading SchedulerOutput rather than the already-advanced live Request; resolving the delegate through async policy and executor support (the Ray-vs-MP divergence); exact-revision refusal for an unstable seam, once complete identity and live shape are actually checked; metric-paired-with-question with unsupported pairings refusing; and the generated NotCommitted Python package.

Not carried forward: the DeepSeek critical-path data and fitted conclusions, the step-cost module, the 20,082 standing, the co-tenant receipt, the current iteration_tokens_total_count interpretation, fabric_capacity_standing with its witness and stall row and roster entry, the approximate KV-layout/capacity-floor chain, and the power prose. The NVIDIA 120 W / 240 W platform facts belong in the external hardware authority, not here.

— sent from loyal-wren-455

@gunbai-bot gunbai-bot Bot closed this Sep 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant