feat(bin): consume the platform decision surface atop the accumulated fleet trunk - #2018
sbracewell64 wants to merge 3 commits into
Conversation
Squash-reconcile of the fork trunk (48 commits through ed376cf, fork PRs #8-#48 plus CI mirrors) rebased onto upstream main at 74230fc, resolving 72 conflicted files. The fork trunk carries: the fm-launch-lib.sh one-owner launch refactor, the slot-base/contribution-target task base contract (fm-task-base-lib.sh), the zero-budget model registry, fleet admission control stages 0/1, the wake-outcome ledger, the LoopSpec canonical representation, the research-approved-work corpus scanner, the 70% compaction doctrine, from-firstmate steer markers, the fleet launcher menu and Windows-to-WSL bridge, forge-verified merge gates, and the fleet-view per-argument cap fix. Conflict-resolution decisions a reviewer cannot see from the diff alone: - Where the fork carried an older in-flight import of upstream work (delivery contracts kunchenguid#1563, remote secondmate homes kunchenguid#1576, trace context kunchenguid#995, and every add/add remote-secondmate file), upstream's landed and further-evolved version wins outright; nothing fork-specific lived in those copies (verified blob-by-blob against upstream history). - Upstream's --relaunch lifecycle control plane and the fork's task base contract both survive: base derivation and the brief base-contract guard are gated to fresh spawns (RELAUNCH=0) because a relaunch reuses the recorded worktree and bases, and the trunk may have moved since. - Launch commands stay one-owner in bin/fm-launch-lib.sh: upstream's Muse Code adapter is ported into the library (template arm, model flag, effort mapping with max->ultra) instead of resurrecting the inline fm-spawn template, and fm-spawn's post-template FM_PI_HARNESS prefix is dropped because the library's pi templates carry the marker themselves. - Herdr presentation spaces keep upstream's 0.8.0-floor semantics (config off/on/empty, two-argument fm_backend_herdr_presentation_enabled) across spawn, config inheritance, docs, and tests. - fm-pr-check derives its per-task lock from state/<id>.meta explicitly so upstream's kunchenguid#1568 meta locking coexists with the fork's #34 landing records (fm_meta_lock_path rejects a .landing path). - tests/lib.sh keeps both the fork's identity-verified background-process reaper and upstream's fm_fake_version_tool fixture. - Merged prose follows upstream's X-mode -> Relay rename. Pre-existing on both parents, not introduced here: tests/fm-calm-pi-extension.test.sh fails under system Node 22 (ERR_UNKNOWN_FILE_EXTENSION), verified identical on the base commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ing operational truth Firstmate reconstructed operational facts conversationally - capacity, live work, decision status - and the drift from the records was silent. The incident this addresses was a report that queued work would dispatch "as capacity frees" while nothing in the fleet was capacity-bound. Every fact needed to refute that sentence was already recorded; nothing made it get read, and nothing refused the sentence. Add bin/fm-decision-surface.sh: a read-only composer over the already-landed deterministic owners, plus three `check` verdicts that refuse a claim structured state contradicts (a capacity claim against an admitting fleet, a ruled decision reported pending, a dispatch of an identity already in flight). It adds no fact of its own; every field names the owner it was read from. An unreadable census, an undecidable admission policy, or an absent decision record is `unevaluable` - the fact may not be asserted at all - never a quiet pass. Rewire the instruction surface onto those owners and delete what they replace: the capacity reasoning at intake, which now defers to the surface and keeps only the semantic serialization judgment, and the run-step mapping restatement in Validate, which bin/fm-crew-state.sh already owns in full. Where no owner has landed, mark rather than delete. `owners` prints the durable compensation ledger: each row is either owned by a landed command or pending with the capability that must land first, and the skill maps every pending row to the instruction it deliberately keeps alive. Attempt and retry counting, a shared verifier verdict vocabulary, the pipeline invocation that replaces the keystroke handoff, and backlog time gates all remain firstmate's for now. Declare the platform seam without depending on it. The deterministic platform publishes a richer projection - why_not_now, allowed transitions, path health - and `platform-seam --probe-platform` measures its wiring rather than assuming it. Probed against platform f0da880 the launcher answers but resolves none of this home's fleet task ids, so the seam stays not-wired; consuming a projection of other identities as fleet truth would be the same silent contradiction the surface exists to prevent. config/decision-surface-platform holds one launcher path, never a command line, so a private config file cannot become a shell-execution seam. Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no live fleet, worker, or platform. Each guarantee was confirmed by breaking it and watching the suite fail; docs/verification/decision-surface.md records those seven mutations and the probe evidence.
eea9575 to
07326e4
Compare
|
Withdrawing this pull request: it was opened at the wrong venue and its head bundles work that does not belong in contribution history. Two concrete problems with this PR as it stands:
The actual change this lane produced is small and self-contained. It is being re-opened as a fork pull request cut fresh from the current fork trunk, carrying only its own net change. For the record, no content was lost in the head being withdrawn: an audit against the branch point found no missing files, and every file whose content differs is one where upstream is ahead of that branch point. |
Intent
Phase E of the captain's software-factory commission: remove deterministic compensation from FirstMate's cognition surface - the fleet-side refactor consuming the landed platform decision surface.
Governing inputs (read before the work): the commission at data/ae-factory-commission-2026-08-08/commission.md (sections 6, 7 and 25 bind this phase - FirstMate's should-not-reason-about list is the removal target); the Phase C gap report's Phase E section at data/ae-factory-phase-c-gap/report.md; and the landed FirstMateDecisionSurface on the platform trunk (commit 4d995ff, with the full D1-D8 substrate through f0da880) - the projection FirstMate consumes instead of reconstructing truth.
Scope, fleet-side:
Decisions and tradeoffs made while doing the work, which a reviewer reading only the diff would not know:
fm-decision-surface.sh owners, 21 rows) rather than a prose document, because the commission requires executable invariants and because a prose inventory drifts from the instruction surface silently. Each row is eitherownedby a landed command orpendingwith the capability that must land first.bin/fm-crew-state.shalready owns in its header and emits directly, so it was a one-owner violation as well as prohibited deterministic work; it was deleted and replaced with a pointer plus the one actionable residue (a parked state means the worker follows the gate help). Everything else that looked like a candidate had no landed owner, so it was marked, not deleted - this is the brief's explicit "no unowned deletion that would leave a gap" rule.bin/fm-decision-surface.sh checkis that host: it answers only "does structured state CONTRADICT this claim", never "is this claim true". Verdicts arecontradicted(exit 3, the claim is forbidden),not-contradicted(exit 0, no landed owner refutes it - explicitly NOT a warrant of truth), andunevaluable(exit 4, no owner could answer, so the fact may not be asserted at all). The tri-state was chosen over a boolean specifically so an unreadable census cannot render as permission to make the claim.state_snapshot_id: fixture-state-v1). So live wiring is not yet reachable for fleet work. The seam is declared with its contract, its owned fields, and the exact condition that retires the marker, and reportswiring: not-wired. Consuming a reachable-but-foreign projection as fleet truth would be the same silent contradiction the surface exists to prevent, so "reachable" and "wired" are deliberately separate fields.config/decision-surface-platformholds ONE launcher PATH and never a command line. A command line would needsh -c/eval, which would turn a private config file into a shell-execution seam; the arguments belong to the script. A test proves appending shell syntax to the configured value does not execute it, and that a path containing spaces works unquoted.Known pre-existing condition, NOT introduced by this change: tests/fm-calm-pi-extension.test.sh fails on this machine with ERR_UNKNOWN_FILE_EXTENSION ".ts" under Node 22. Verified to fail identically on the base commit through a clean git archive extract of HEAD.
What Changed
bin/fm-decision-surface.sh, a read-only composer over landed deterministic owners (capacity, decision rulings, duplicate dispatch, path health) with tri-state contradiction checks (contradicted/not-contradicted/unevaluable), a machine-readable 21-row owners inventory, and a declared but deliberately not-wired platform seam; AGENTS.md and the supervision/skill surfaces are rewired to consume it instead of restating deterministic facts, with review follow-ups for whole-token seam id matching, probe kill grace, capacity rendering, and argument parsing.fm-launch.sh,firstmate.bat,fm-wsl-entry.sh), fleet admission control, the model registry with zero-budget spawn gating and probe verification, the wake-outcome ledger, the canonical LoopSpec schema and register, the research-approved-work corpus scanner, remote secondmate provisioning/jobs/doctor with trace-context propagation, and the worktree allocation guard — each with new skills, docs, and test suites.Risk Assessment
✅ Low: The follow-up commit cleanly fixes all four previously-reported defects exactly as directed — whole-token seam matching, an enforceable kill-after probe bound, a shared reason fallback, and symmetric parser guards — each with a behavior-level regression test, fail-closed failure directions, and an accurately updated verification record, with no new issues introduced.
Testing
Ran the focused fm-decision-surface behavior suite (13/13 pass, ~8s), captured an end-to-end CLI transcript demonstrating all three section-25 focus checks with negative controls and correct tri-state exit codes, the 21-row compensation ledger, and the platform seam's not-wired/injection-safety behavior, and independently confirmed red-capability by making the suite fail under two targeted script mutations before restoring a clean tree; no failures or regressions found, and the author-documented pre-existing fm-calm-pi-extension Node-22 failure is unrelated and was not exercised.
Evidence: End-to-end CLI transcript: surface render, three section-25 checks with exit codes, owners ledger, platform seam and injection safety
Evidence: Focused test suite run (13/13 ok)
Evidence: Red-capability mutation spot-check (two mutations, matching cases fail, clean revert)
Mutation 1 (disable capacity contradiction): not ok - a capacity claim against an admitting fleet must be refused: expected exit 3, got 0 Mutation 2 (hardcode reachable=>wired): not ok - reachable but foreign must not be reported as wired Both reverted; git status clean; suite green again (13/13).Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
⏭️ **Rebase** - skipped
.agents/skills/afk/SKILL.md- branch carries 46 commit(s) that exist on your local main branch but were never pushed to origin/main; rebasing would bundle this unrelated work (229 file(s)) into the PR:Push main to origin, or rebase your branch onto origin/main, before gating.
🔧 **Review** - 4 issues found → auto-fixed ✅
bin/fm-decision-surface.sh:361- Platform-seam wiring detection uses raw substring containment (case "$probe_out" in *"$id"*) of fleet task ids against the probe's free-text output. A fleet id that appears coincidentally — inside a longer platform work id, a title describing the same program, or any prose — counts towardfleet_identities_resolvedand a single match flipswiringtowired. The skill makeswiredconsequential: the platform projection becomes authoritative and the four transition-related pending rows retire, so a false positive reintroduces exactly the silent contradiction the seam exists to prevent. Match ids as whole tokens (e.g. split the output on non-id characters, or grep -F -w per id) or parse the projection's structured id fields.bin/fm-decision-surface.sh:329- run_timed invokestimeout/gtimeoutwithout--kill-after(or-s KILL), so a platform launcher that traps or ignores SIGTERM keeps running after the 120s bound and the command-substitution read blocks indefinitely — contradicting the header's stated guarantee that "an unresponsive platform cannot wedge a read". The red-capability mutations only covered the no-bounding-tool case, not a TERM-immune child. Both GNU timeout and gtimeout support-k <grace>; add it to make the bound enforceable.bin/fm-decision-surface.sh:511- The human fleet-surface renderer printscapacity: <band> · <action> · nullfor every configured fleet: an activefm-admission.v1record carries no top-level.reasonfield (only the inactive/unconfigured record does), and the renderer interpolates.capacity.reasondirectly. check_capacity_blocked already has the.reason // controlling_rulesfallback; reuse it in the render_surface jq program so the common configured case doesn't display the literal string "null".bin/fm-decision-surface.sh:134- The argument parser lets a named subcommand silently override an earlier task-id argument:fm-decision-surface.sh <task-id> ownerssets SUBCOMMAND=surface then overwrites it with owners and drops the task scope, and<task-id> check decision-pending <id>silently reassigns TARGET — while the reverse order (owners <task-id>) correctly dies with "unexpected argument". Guard theowners/platform-seam/checkbranches with the same[ -z "$SUBCOMMAND" ]check used in the positional branch so mixed invocations are refused as usage errors.🔧 Fix: fix seam token match, probe kill grace, render, parsing
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-decision-surface.test.sh (13/13 ok, ~8s; re-run green after mutation spot-checks)Manual end-to-end CLI demo against a canned fleet home:fm-decision-surface.shfleet surface render,check capacity-blocked(exit 3 while admitting),check decision-pendingruled/open/unrecorded (exits 3/0/4),check duplicate-dispatchlive/new identity (exits 3/0)Manual ledger verification:fm-decision-surface.sh ownersandowners --json— 13 owned + 8 pending = 21 rows, no private CFVC idsManual seam verification:platform-seamunconfigured (not-wired, reachable=not_checked), config value with appended; touch markerdoes not execute and probes unreachable, reachable-but-foreign launcher stays not-wiredRed-capability mutation spot-check: disabled the capacity contradiction branch (seed-1 case failed: expected exit 3, got 0) and hardcoded reachable-implies-wired (foreign-seam case failed); both reverted withgit checkoutand cleangit status --porcelainconfirmedReviewed the AGENTS.md diff for the two intended compensation deletions (intake capacity/serialization, Validate run-step mapping) and the section 7 stub plus section 13decision-surfaceload trigger✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.