Repository navigation
Restore the dashboard deploy ROUTE behind a fail-closed interlock — the capability stays withheld, in code - #9489
Conversation
The roadmap has been dead since Aug 19 because nothing deploys srv1. The belt is not broken -- it refuses every tick on revision drift, which is fail-closed behaviour working correctly. What broke is the edge: gunbc#8676 deleted deploy.yml, and the deleted workflow was the only caller of live_deploy_apply_srv1_wet. The OPERATION was never deleted, only its invocation. This adds the caller back as a fleet-converge dispatch mode, per operator direction to put the deployment carrier there. WHY THIS PLACEMENT IS THE SAFE ONE, not merely the convenient one. gunbc.live_deploy.deployed_tree_report blocks tree-only publication: moving the deployed HEAD alone flips belt_revision_admission from refusing to admitting while the installed binary is still old, so the next belt tick dispatches a mixed realization never admitted as a release. This job already builds and ships that binary in the same run, so tree and interpreter arrive together -- the pair, which is the subject gunbc.live_deploy.release_binding names. The step carries no `target` input, unlike the mutating spark converge which does. This deploy takes its host from the RUNNER: ci_deploy_srv1_access resolves the job principal from the runner it lands on and refuses when that user is not in srv1's fleet roster. So host=srv1 is what aims it, and any other host refuses at preflight rather than deploying somewhere unintended. Additive and opt-in: a new workflow_dispatch mode. No existing mode's behaviour changes, and nothing fires unless dispatched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n its carrier requires Landing the deploy as a mode-gated STEP inside the shared converge job was wrong, and the reason is a requirement the repository already states. gunbc.fleet_workflow_steps srv1_dashboard_deploy_serialization_note: blue/green selection is observe-then-mutate, so two deploys may not overlap -- both could observe blue active, then the first flips green while the second restarts what has become the active green slot, recreating the outage. The note specifies JOB-GRAIN concurrency. `concurrency` is a job key, so a step would have inherited fleet_converge_job's `concurrency: none`: the step form restored the operation WITHOUT the serialization the operation requires. That group (srv1_dashboard_deploy_concurrency_group) had zero consumers after gunbc#8676 deleted deploy.yml. This re-homes the deleted job's requirement rather than re-inventing it, with cancel_in_progress false so a queued deploy never interrupts the one holding the host mutation window. The job does NOT reuse the converge job's WIF auth or in-run fleet key. ci_deploy_srv1_access is `transport: LocalShell` -- the deploy mutates the host it runs on and reaches nothing else -- so those credentials exist for an operation this job does not perform, and granting them would widen its credential surface for nothing. Timeout gets its own tier (30m) rather than borrowing. Aux (5m) is sized for a small command, not an rsync plus unit write plus restart plus readiness wait; spark_serving (60m) is the right magnitude under a name that would be a second name for one concept (§3). gunbc_ci_fleet_job_backstop_timeout_note declares itself the single authority to update on any change to the converge job's mode set. Updated: the mode set grew and that job's 105m bound is UNCHANGED, because dashboard_deploy adds no step to it. Recorded with the distinction a future reader needs -- adding a mode and adding a step to this job are different events, and only the second moves that number. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Caught by resolve: field 'runs_on' not found in type 'Job', plus the paired missing-required-field 'runner'. The two diagnostics are one mistake seen from both sides.
Wrong variant on the first cut. QueueMax permits many pending entries, so a burst of dispatches would deploy each queued revision in turn -- installing revisions already superseded before they were installed. QueueNotMax holds at most one pending and replaces it with the newest arrival, which is exactly what srv1_dashboard_deploy_serialization_note describes: an older pending deployment gives way to the newest main revision. cancel_in_progress stays false either way, and that is the half that matters most: a queued deploy must never interrupt the one holding the host mutation window, which is the state that produces a half-applied cutover. Found by reading the emitter rather than the model -- gha_workflow's concurrency_queue_max_emission_note spells out that the two variants mean different things on the platform, after a period when they serialized identically and the distinction had no realization. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Derived from the authority via `gunbc run --entry dag/tools/generated_artifact_gate.dag --function main_wet`, never hand-edited. Only this artifact is taken. The same regeneration also rewrites .gitattributes and three docs/plans/*.md that this branch does not touch and that are byte-identical to main, so main carries pre-existing generated-artifact drift independent of this change. Sweeping it in here would mix concerns, and regenerating a drifted artifact can silently delete orphan content no authority produces -- that drift wants its own look, not a ride-along. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ot, the group's reason was not
|
Two updates from a side-chat review, one correcting this PR and one correcting the review. Fixed here ( The reason is now stated against the deployment that actually runs — two candidate transactions interleaving writes to one tree, binary, unit, process and readiness subject, so each transaction's readback can observe the other's artifacts. In-place convergence makes that stronger, not weaker. The carrier also moved from a I also weakened a claim I could not support: QueueNotMax was described as replacing a pending deploy with "the newest main revision". While the trigger is Stale in the review: "the generated workflow does not contain the change". That was measured at Not fixed here, and recorded as open rather than quietly carried — three of them are now written into the group's own annotation so a later reader cannot mistake this group for a wall it is not:
That last one carries an operational consequence I want stated plainly rather than left in a thread: merging this PR is safe, but the restored mode should not be dispatched against srv1 until the — sent from calm-ram-380 (Edited: the SHA above posted unexpanded from a quoted heredoc — a citation naming nothing, in a comment about stale citations. Corrected to the commit that carries the fix.) |
|
Measured the live consequence of the
The first column is index-vs-HEAD and the second is worktree-vs-index, so both are wrong, not just one. That is what the two This is not latent. It is the current state of the tree the belt spawns dispatch worktrees out of. One measurement that changes the available options, and it is the useful half: srv1 can fetch natively today. Credentials and network are fine — the same That also collapses two of the open items above into one. Ordering and admitted-SHA coupling were listed as separate gaps, but a git-native convergence cannot be written without naming a target commit — there is no "fetch whatever the runner happens to have", because the runner's checkout is not reachable from srv1. So the Git-safe transition and the admitted-SHA coupling are the same change, not two that must be sequenced. I have not built it, because it carries a semantics decision that is not mine to make unilaterally: a git-native convergence would make the deploy mean "converge the target to an admitted commit" rather than "push the runner's checkout", and it would collapse both rsync legs rather than just the Nothing here changes this PR. It is recorded on it because this PR is what restores the route to that operation, and the note above is the reason not to walk it yet. — sent from calm-ram-380 |
|
Replaced my own warning with a wall ( What landed.
Home chosen for acyclicity, not tidiness. It covers more ingress than the concurrency group can, and I checked the call graph rather than assuming. Both Stated precisely, because the stronger result must not be read as retiring work: this closes ingress for the admission question ("may this realization run at all"), not the serialization question ("is another transaction running now"). The host-local lease is still owed, and so is the durable actuator inhibition beside it — a process lease dies with the process, so a deploy that fails after moving the tree and before restarting would leave the next belt tick free to consume mixed state. Those are two mechanisms answering different questions, and I had them fused until review separated them. Retract deliberately does not consult this. Removing a deployment performs no repository transition, and a wall that also blocked retraction would strand an operator with a host they could neither converge nor take down. That is stuck, not fail-closed. Rung, honestly: mechanically preventable — not structural. The wall executes and refuses on the real path, but the bound realization is a hand-declared row: someone could edit it without having done the work and nothing would notice. Next-rung trigger is recorded in the annotation — derive the bound realization from the emitted tree-sync unit, so the declaration cannot disagree with what is installed. Dissolves on the Git-native transition landing, at which point the refusal arm stops being production-reachable and its control is retained as a regression probe rather than deleted (§4b(4)). Evidence — The pair is the witness, not either arm. A wall asserted only on the arm that currently refuses cannot distinguish a real decision from a function that returns the refusal unconditionally — so both arms are asserted over one function, and fusing them makes the admitted case red. The third pins the row to legacy, which means flipping that row goes red and forces the author to state that the realization actually moved; the flip cannot be silent. Also fixed on this PR: the body still carried the stale blue/green rationale after the carrier was corrected. Corrected there too, with the correction stated rather than silently swapped. Two process notes, since the first run of these witnesses looked green and was not: it failed with — sent from calm-ram-380 |
|
CI failure was inherited, and I checked my own change before saying so — because the reasoning I used on #9501 had a hole when applied here. The hole. The four admission witnesses on this PR import Closed by execution. Ran the witnesses that do reach
And the blocker is gone. — sent from calm-ram-380 |
|
Addressing Both facts check out: What changed: the PR is retitled "Restore the dashboard deploy ROUTE behind a fail-closed interlock — the capability stays withheld, in code" and the body rewritten to lead with the distinction rather than bury it. Route, serialization and interlock are delivered; the deploying capability is deliberately withheld. Why not the first remedy — landing the Git-native transition with this caller. Three reasons, and the last is the one I cannot resolve unilaterally:
So the two-increment shape is deliberate, and the review's objection was to the framing rather than to the shape. The framing is fixed. One correction to the review, on §3 grounds rather than substance: the findings cite — sent from calm-ram-380 |
# Conflicts: # .github/workflows/fleet-converge.yml # dag/gunbc/fleet_converge_workflow.dag # dag/gunbc/fleet_workflow_steps.dag # dag/test/claim/workflow_dispatch_input_witness_test.dag
The four existing witnesses all exercise repository_transition_admission, the decision function. None reaches live_deploy_wet_with_access, where the decision is consulted -- so deleting that match and calling the admitted body directly leaves every one of them green with the wall gone. That is the local-relation-promoted-to-end-to-end failure. The two added witnesses call the real guarded entry with the exact access values its two wet ingresses pass, so coverage is per-argument rather than per-call-site. They are hermetic BECAUSE the guard fires: the refusing arm returns ahead of any host effect. Their annotation records that the coming flip to GitNativeConvergence must DELETE and replace them rather than relax them, since a relaxed placement check is the decoration 4b calls worse than absent. Found by side-chat review; the two approving reviewers both read past it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
What this actually does
#8676deleteddeploy.yml, the only caller oflive_deploy_apply_srv1_wet. Since then there has been no route at all to deploy srv1 — the tree is 604 commits behind and the roadmap belt refuses every tick on revision drift.This restores that route as a
dashboard_deploymode with a dedicateddashboard-deployjob, and then refuses to walk it, in code, for a reason measured on the host.Why the route is a job and not a step
srv1_dashboard_deploy_concurrency_grouprequires job-grain concurrency, andconcurrencyis a job key — a mode-gated step inside the shared converge job would have inherited that job'sconcurrency: none. Landing the deploy as a step would have restored the operation without the serialization the operation requires.QueueNotMaxrather thanQueueMax: QueueMax permits many pending entries, so a burst would install each queued revision in turn, including revisions already superseded before they were installed.cancel_in_progress: falseis the half that matters most — a queued deploy must never interrupt the one holding the host mutation window.The timeout is its own derived tier rather than a borrowed one, and
gunbc_ci_fleet_job_backstop_timeout_noteis updated to record why this job's 105m bound did not move:dashboard_deployadds no step to the converge job at all.Why it refuses, and why that is the deliverable rather than a shortfall
The one deployment operation this route reaches transports
.gitas files —HEAD,objects,packed-refs,refscopied whileindexis excluded — alongside an rsync of the working tree. Measured on srv1 2026-08-27,/opt/gunbc/gunbc:Status column one is index-vs-HEAD, column two is worktree-vs-index — both wrong, from one execution with no concurrency involved. So this is not a race the new concurrency group addresses; one isolated run of that operation cannot produce a valid target state.
RepositoryTransitionAdmissiontherefore refusesLegacyGitFileSyncatlive_deploy_wet_with_access, beforeobserve_candidate_release, so nothing is observed let alone mutated. Retract deliberately does not consult it: a wall that also blocked retraction would strand an operator with a host they can neither converge nor take down, which is stuck rather than fail-closed.The alternative to this wall is not a working deploy — it is a deploy that corrupts the tree. The capability was already unavailable before this PR and this PR does not remove it; what changes is that the unavailability is now enforced and explained instead of resting on a sentence in a PR body. A prose warning is not a wall (§5), and I had written one.
It also closes ingress the concurrency group structurally cannot: both
live_deploy_apply_srv1_wetandlive_deploy_apply_srv1_operator_wetroute through the walled launcher, so the refusal coversgunbc.apply->DashboardSrv1— the direct local route that never traverses a GitHub concurrency group.Rung, stated honestly
Mechanically preventable, not structural. The wall executes and refuses on the real path, but the bound realization is a hand-declared row: someone could edit it without doing the work and nothing would notice. Next-rung trigger is in the annotation — derive the bound realization from the emitted tree-sync unit so the declaration cannot disagree with what is installed. It dissolves when the Git-native transition lands, at which point the refusal arm stops being production-reachable and its control is retained as a regression probe (§4b(4)).
Why the transition is not in this PR
#9506builds it: three cited git operations (compare-and-swapupdate-ref, reflog read, worktree transition), a pure convergence adjudication whose preservation checks are identity joins, and a wet composition ordered objects → CAS → reset → readback. It currently carries REQUEST_CHANGES from side-chat review, which found real defects in it (untracked residue survivingreset --hard; the primary's own branch reported as a lost ref; a partial transition left unnamed when the reset fails after a successful CAS; a reflog law that any unrelated append satisfied). Those are fixed and it now stands at 16 witnesses, but its evidence is still pure adjudication — no witness executes the wet composition — and a scratch-repository execution matrix is owed. An earlier revision of this line said it was "approved and green on nine witnesses", which was stale in both halves and is corrected here rather than quietly updated: it is the successor this PR points at, so overstating its standing overstates the case for withholding activation here.Binding it here would mean one PR that restores a route, replaces a production repository transport, deletes both rsync legs and changes what a deploy means — from push the runner's checkout to converge the target to an admitted commit. That last one is a semantics decision I have flagged for the operator and do not have confirmation on. Two reviewed increments are the correct shape; one is not.
Corrections carried in this PR
single_deployment_note). The conclusion survived the deletion; its stated reason did not. Rewritten against in-place interleaving, moved from a commentary-onlydata …: Stringto a §4c annotation, all three dangling references updated.QueueNotMaxwas described as replacing a pending deploy with "the newest main revision". Underworkflow_dispatchover a caller-selected ref, the newest arrival need not be main.dashboard-deploy'sneeds: [build]edge asserted explicitly.Still open, recorded in the group's own annotation rather than carried in prose
Serialization is not ordering (a slow build of an older revision can acquire the lease after a fast one and install backwards); the host-local lease and a durable actuator inhibition are two mechanisms, not one, because a process lease dies with the process; and the candidate is not yet proven equal to the admitted
refs/fleet/desired.