Skip to content

The observer could not see the host it existed to retire: the observation denominator made physical - #9731

Merged
briansrls merged 4 commits into
mainfrom
session/eager-pike-541-observe-roster
Aug 30, 2026
Merged

briansrls merged 4 commits into
mainfrom
session/eager-pike-541-observe-roster

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

The defect, proven by execution

spark_serving_observe_ci_wet enumerated spark_serving_desired_hosts. Correct while desired named every Spark; wrong the moment #9679 gave desired a role scope. Retirement is planned from members observed and not desired, so scoping observation by desired makes the retirement population unobservable by construction — observed − desired is empty for every host, always.

Evidence Run Result
Observer scope fleet-converge 33296910990 (spark_observe) observed_hosts=1 — srv5 only; srv6 enrolled in known-hosts, never probed
Ground truth direct ssh, srv6 unit active + enabled, listener on 11434, ollama serve pid 2349

The module already refused an empty observation list, on the stated ground that emptiness is ignorance rather than a clean fleet. The roster held {srv5}, which is not empty, so the guard never fired. empty_observation_narrow in its non-empty form.

Correction: the 0-byte plan wire is a DIFFERENT defect

An earlier version of this description cited fleet-converge run 33297525985's zero-byte spark_serving_typed_actions.wire as evidence for this defect. That chain is false in the live dataflow. fleet_converge_plan_cli.dag:170 passes observed_spark: [] and fleet_converge_plan.dag:2325 sets spark_serving_typed_actions: scheduled_rows_empty() unconditionally — the live planner reads no Spark observer at all and would have emitted an empty wire whatever this module answered.

The two defects are independent. Welding them into one story would have made a fix here look like a fix there.

What this PR does and does not do

Does: restores the physical observation denominator and supplies the exact-coverage classifier.

Does not: make live retirement planning reachable. The live plan and apply producers remain disconnected from Spark observation, and re-running mode=plan after this lands should still produce a zero-byte wire. Those are the next P0-C obligations.

Why this is a module and not a one-line edit

The obvious repair — union the two role rosters — is rejected here. Both role projections answer through SparkCellRoleUnique and deliberately drop SparkCellRoleAmbiguous and SparkCellRoleAbsent, so their concatenation silently omits a host whose role assignment is broken. An absent or contradictory role says nothing about whether a serving process runs on the metal — and a broken assignment is precisely when stale serving state most needs to stay visible. That repair fixes today's srv6 and preserves the class one input away.

Roles decide what may be done to a host. They do not decide whether it is looked at.

Coverage is a declared replacement, not a completed cutover

SparkObservationCoverage is defined and exercised by claims. serving_observe still terminates on the older cardinality guard and does not consume a coverage value. The two authorities genuinely coexist until the observer's report is built from the join; that cutover is the remaining P0-C obligation and is stated in the module rather than implied.

Terminal policy fix (defect introduced by this PR)

Widening the roster put a retired cell inside it, and a retired cell is structurally SparkServingAbsent. The workstation entry terminated on spark_serving_observation_report_succeeded, under which absent counts as failure — so the exactly-correct final state (srv5 serving, srv6 retired) would have exited non-zero. Both entries now terminate on spark_serving_report_observations_were_made ("did the probe answer").

Evidence standing, stated honestly: spark_serving_observe_witness_test:646 already asserts exactly was_made && !succeeded for an absent host — the control is authored and discriminating, and it does not execute, because that module is test.claim.spark.*, outside the required gate closure. Restating it under the executed v2.test.* prefix was attempted and abandoned on measurement: importing this module to reach the two predicates pulls its 46-import closure and the claim module is killed during resolve with no rows.

Evidence

8 claims under v2.test.*, which the required floor executes — confirmed by the gate's own output, all 8 reported standing=planned-and-passed disposition=planned outcome=passed.

Discriminating control: reinjecting the defect (scope reverted to spark_serving_desired_hosts) reds 5 of 8. The 3 that stay green are the pure coverage-classifier claims, which consult no scope producer and correctly test a different subject.

— sent from eager-pike-541

gunbc-ci-auto-heal and others added 4 commits August 30, 2026 07:39
…tion denominator made physical

srv6 ran a live, enabled serving unit and the converge plan minted zero typed
actions and called the fleet clean. Measured, not reasoned: fleet-converge run
33296910990 reported observed_hosts=1, and run 33297525985 wrote
spark_serving_typed_actions.wire at 0 bytes.

spark_serving_observe_ci_wet enumerated spark_serving_desired_hosts. That was
correct while desired named every Spark, and became wrong the moment #9679 gave
desired a role scope. Retirement is planned from members observed and NOT
desired, so scoping observation by desired makes the retirement population
unobservable by construction -- observed minus desired is empty for every host,
always. The module already refused an EMPTY observation list on the ground that
emptiness is ignorance rather than a clean fleet; the roster held {srv5}, which
is not empty, so nothing fired.

That is empty_observation_narrow in its non-empty form.

The obvious repair -- union the two role rosters -- is rejected here and the
rejection is why this is a module rather than a one-line edit. Both role
projections answer through SparkCellRoleUnique and deliberately drop
SparkCellRoleAmbiguous and SparkCellRoleAbsent, so their concatenation silently
omits a host whose role assignment is broken. An absent or contradictory ROLE
says nothing about whether a serving PROCESS runs on the metal, and a broken
assignment is precisely when stale serving state must stay visible. That repair
would have fixed today's srv6 and preserved the class one input away.

So the scope reads the procurement authority, which derives the enrolled Spark
identities from the router bindings and knows nothing about roles. Roles decide
what may be DONE to a host; they do not decide whether it is LOOKED AT.

Coverage becomes an identity join and SUBSUMES the cardinality guard rather than
standing beside it -- two authorities for one question. Observed keys must equal
expected keys, each exactly once; missing, foreign and duplicated are carried
separately because they are different defects with different causes.

Evidence: 8 claims under v2.test.*, which the required floor executes -- the
same conclusion #9711 reached when it moved a RED witness under the required
prefix rather than leaving an inert wall. They are pure over host-identity lists
and reach no fleet planner. Discriminating control run: reinjecting the defect
reds 5 of 8; the 3 that stay green are the pure coverage-classifier claims,
which consult no scope producer and correctly test a different subject.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018dHefbXnLiW1XNs2vVxhwj
Review 57545 (non-blocking): SparkObservationCoverage is defined and exercised
by claims but not consumed by serving_observe, which still refuses on the count
the join was written to subsume -- while the module comment described the
subsumption as done.

That is the overstatement this module exists to correct, one layer up: prose
asserting a migration that execution has not performed. The PR body named the
gap; the source did not, and the source is what survives.

The comment now states that the replacement has NOT happened, that the two
authorities genuinely coexist until the observer's report is built from the
join, that the narrowed-roster case is currently caught by the claims rather
than by the production path, and that the cutover is the remaining P0-C
obligation. No behaviour change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018dHefbXnLiW1XNs2vVxhwj
…hain is unwelded

Two corrections from tracing the live dataflow rather than reasoning about it.

TERMINAL POLICY (a defect this branch introduced). Widening the roster to the
physical Spark population put a RETIRED cell inside it, and a retired cell is
structurally SparkServingAbsent. The workstation entry terminated on
spark_serving_observation_report_succeeded, under which absent counts as
failure, so the exactly-correct final state -- srv5 serving, srv6 retired --
would have exited non-zero. Fail-closed, but the wrong answer for an instrument
whose job is to describe what is there. Both entries now terminate on
spark_serving_report_observations_were_made.

The row that is no longer the terminal is retained and documented rather than
deleted: its only consumers are claims in test.claim.spark.*, which is outside
the required gate closure, so editing that module would classify its witnesses
as changed-and-declined and red the required floor. It is blocked from deletion
by the same enrollment gap that owns the 59 nonterminal identities, and that is
recorded at the function rather than hidden.

CAUSAL CHAIN. The module comment cited fleet-converge run 33297525985's 0-byte
spark_serving_typed_actions.wire as evidence for the observer defect. False in
the live dataflow: fleet_converge_plan_cli passes observed_spark: [] and
fleet_converge_plan_artifact sets spark_serving_typed_actions to
scheduled_rows_empty() unconditionally, so the live planner reads NO Spark
observer and would have emitted an empty wire whatever this module answered.
Independent defects, repaired separately. The comment now says so, including why
the chain was tempting.

Evidence for the terminal swap is authored (spark_serving_observe_witness_test
asserts was_made && !succeeded) and does not execute. Restating it under
v2.test.* was attempted and abandoned on measurement: reaching the predicates
pulls this module's 46-import closure and kills the claim module during resolve.

8 scope claims re-verified 8/8 on the refreshed head.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018dHefbXnLiW1XNs2vVxhwj
@briansrls
briansrls merged commit ac0bdaf into main Aug 30, 2026
4 checks passed
@briansrls
briansrls deleted the session/eager-pike-541-observe-roster branch August 30, 2026 09:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant