Skip to content

feat(bin): consume the platform decision surface atop the accumulated fleet trunk - #2018

Closed
sbracewell64 wants to merge 3 commits into
kunchenguid:mainfrom
sbracewell64:fm/ae-factory-phase-e-firstmate-refactor
Closed

sbracewell64 wants to merge 3 commits into
kunchenguid:mainfrom
sbracewell64:fm/ae-factory-phase-e-firstmate-refactor

Conversation

@sbracewell64

Copy link
Copy Markdown

Intent

Phase E of the captain's software-factory commission: remove deterministic compensation from FirstMate's cognition surface - the fleet-side refactor consuming the landed platform decision surface.

Governing inputs (read before the work): the commission at data/ae-factory-commission-2026-08-08/commission.md (sections 6, 7 and 25 bind this phase - FirstMate's should-not-reason-about list is the removal target); the Phase C gap report's Phase E section at data/ae-factory-phase-c-gap/report.md; and the landed FirstMateDecisionSurface on the platform trunk (commit 4d995ff, with the full D1-D8 substrate through f0da880) - the projection FirstMate consumes instead of reconstructing truth.

Scope, fleet-side:

  1. Inventory every place this repo's shared tracked surface makes FirstMate perform the commission's prohibited deterministic work: capacity counting, polling, PR-existence and check interpretation, legal-transition invention, retry counting, remembering in-flight work, LoopSpec continuation memory, known next-stage invocation. AGENTS.md, the supervision protocols, bearings, and the wake-handling contracts are the primary surfaces. Prior increments already retired several - build on them, do not re-implement: attempt counting, task outcome, duplicate detection, and structured crew state.
  2. For each compensation with a landed deterministic owner: rewire the instruction surface to consume the owner (the decision-surface projection, why_not_now, allowed transitions, path health) and delete the compensating prose. For each WITHOUT a landed owner yet: mark it with an explicit pointer to the owning future increment rather than silently keeping the prose - no unowned deletion that would leave a gap.
  3. The commission section 25 focus-test seeds become fleet-side red-capable tests where the fleet's surfaces can host them: a capacity claim contradicting free slots, a ruled decision reported pending, a duplicate dispatch - each provably contradicted from structured state.
  4. This rewrites AGENTS.md and supervision contracts - follow the firstmate-coding-guidelines skill; preserve every safety boundary; the always-loaded contract stays concise; one-owner rule throughout.
  5. EXCLUDED: platform-side changes; workflow routing (Phase F); the fleet consumption WIRING of the platform AXI decision surface where it needs live platform connectivity - define the contract seam and mark it if the live wiring is not yet reachable.

Decisions and tradeoffs made while doing the work, which a reviewer reading only the diff would not know:

  • The inventory is deliberately expressed as machine-readable data (fm-decision-surface.sh owners, 21 rows) rather than a prose document, because the commission requires executable invariants and because a prose inventory drifts from the instruction surface silently. Each row is either owned by a landed command or pending with the capability that must land first.
  • Only TWO compensations were actually deleted from AGENTS.md, on purpose. The intake capacity/serialization paragraph now defers to the decision surface and retains only the genuinely semantic serialization judgment. The Validate run-step mapping sentence was a full restatement of the mapping bin/fm-crew-state.sh already owns in its header and emits directly, so it was a one-owner violation as well as prohibited deterministic work; it was deleted and replaced with a pointer plus the one actionable residue (a parked state means the worker follows the gate help). Everything else that looked like a candidate had no landed owner, so it was marked, not deleted - this is the brief's explicit "no unowned deletion that would leave a gap" rule.
  • Pending rows are named by CAPABILITY, not by the captain's private program increment ids (CFVC-xx). Those ids live only in a gitignored private backlog, so a shared template repo referencing them would be an unresolvable pointer for every other firstmate user. The skill carries the pending-row-to-instruction mapping table so the linkage stays unambiguous.
  • The three section 25 focus tests needed an executable host, because repo policy forbids tests that assert instruction-source bytes. bin/fm-decision-surface.sh check is that host: it answers only "does structured state CONTRADICT this claim", never "is this claim true". Verdicts are contradicted (exit 3, the claim is forbidden), not-contradicted (exit 0, no landed owner refutes it - explicitly NOT a warrant of truth), and unevaluable (exit 4, no owner could answer, so the fact may not be asserted at all). The tri-state was chosen over a boolean specifically so an unreadable census cannot render as permission to make the claim.
  • The script is a COMPOSER over already-landed owners and adds no fact of its own; every field names the owner it was read from. It is read-only: no lock, no wake drain, no mutation, no spawn.
  • Every focus test is paired with a negative control that drives the same code path to the opposite verdict from opposite structured state, because a check that only ever answers "contradicted" enforces nothing. Red-capability was then confirmed empirically by applying seven separate mutations to the script and watching the matching case fail, restoring, and re-verifying green.
  • On scope item 5: the platform AXI decision surface IS reachable from the fleet side (probed live at platform f0da880), but it resolves ZERO of this home's fleet task ids and its registry is fixture-backed (state_snapshot_id: fixture-state-v1). So live wiring is not yet reachable for fleet work. The seam is declared with its contract, its owned fields, and the exact condition that retires the marker, and reports wiring: not-wired. Consuming a reachable-but-foreign projection as fleet truth would be the same silent contradiction the surface exists to prevent, so "reachable" and "wired" are deliberately separate fields.
  • config/decision-surface-platform holds ONE launcher PATH and never a command line. A command line would need sh -c/eval, which would turn a private config file into a shell-execution seam; the arguments belong to the script. A test proves appending shell syntax to the configured value does not execute it, and that a path containing spaces works unquoted.
  • The platform probe is bounded through the repo's existing timeout/gtimeout portability idiom, and when no bounding tool exists the probe does not run at all rather than running unbounded - reporting unreachable is both honest and the direction that keeps the seam not-wired.
  • Placement followed the repo's knowledge-placement tree: a six-line always-loaded stub in AGENTS.md section 7, the full procedure in a new agent-only skill with a section 13 load trigger, mechanics in the script header and --help, the config schema in docs/configuration.md, and dated empirical evidence in docs/verification/decision-surface.md. AGENTS.md grew by a net 16 lines against 4 deleted, which is deliberate given this is the commission's central Phase E requirement.

Known pre-existing condition, NOT introduced by this change: tests/fm-calm-pi-extension.test.sh fails on this machine with ERR_UNKNOWN_FILE_EXTENSION ".ts" under Node 22. Verified to fail identically on the base commit through a clean git archive extract of HEAD.

What Changed

  • Adds bin/fm-decision-surface.sh, a read-only composer over landed deterministic owners (capacity, decision rulings, duplicate dispatch, path health) with tri-state contradiction checks (contradicted/not-contradicted/unevaluable), a machine-readable 21-row owners inventory, and a declared but deliberately not-wired platform seam; AGENTS.md and the supervision/skill surfaces are rewired to consume it instead of restating deterministic facts, with review follow-ups for whole-token seam id matching, probe kill grace, capacity rendering, and argument parsing.
  • Carries the fork trunk's accumulated feature work into the branch: the fleet launcher menu with Windows→WSL bridge (fm-launch.sh, firstmate.bat, fm-wsl-entry.sh), fleet admission control, the model registry with zero-budget spawn gating and probe verification, the wake-outcome ledger, the canonical LoopSpec schema and register, the research-approved-work corpus scanner, remote secondmate provisioning/jobs/doctor with trace-context propagation, and the worktree allocation guard — each with new skills, docs, and test suites.
  • Lands the trunk's merge and supervision fixes: refusing merges without verified green checks, resolving the fork trunk as the landing target, resolving the wake sequence instead of trusting a supplied one, detecting child-process work during supervision, separating task read and contribution bases, escalating blocked-on-human transitions past status-line dedupe, and quieting wedge alarms on settled terminal tasks.

Risk Assessment

✅ Low: The follow-up commit cleanly fixes all four previously-reported defects exactly as directed — whole-token seam matching, an enforceable kill-after probe bound, a shared reason fallback, and symmetric parser guards — each with a behavior-level regression test, fail-closed failure directions, and an accurately updated verification record, with no new issues introduced.

Testing

Ran the focused fm-decision-surface behavior suite (13/13 pass, ~8s), captured an end-to-end CLI transcript demonstrating all three section-25 focus checks with negative controls and correct tri-state exit codes, the 21-row compensation ledger, and the platform seam's not-wired/injection-safety behavior, and independently confirmed red-capability by making the suite fail under two targeted script mutations before restoring a clean tree; no failures or regressions found, and the author-documented pre-existing fm-calm-pi-extension Node-22 failure is unrelated and was not exercised.

Evidence: End-to-end CLI transcript: surface render, three section-25 checks with exit codes, owners ledger, platform seam and injection safety
==============================================================
Demo fleet: 1 live task 'alpha' (working), decision 'demo-decision-open'
awaiting the captain, decision 'demo-decision-ruled' already ruled (Done).
No admission policy configured => the fleet admits another task.
==============================================================

### The decision surface firstmate reads before reasoning about work

$ fm-decision-surface.sh 
decision surface · scope: fleet
  live work:  1 (alpha)
  census:     coherent
  capacity:   preferred · admit · admission control is not configured for this home
  decisions:  demo-decision-open
  platform:   not-wired (absent)
  pending:    attempt_and_retry_counting, verifier_verdict_vocabulary, invoking_known_next_stage, deadline_and_time_gate_elapsed, transition_legality, why_not_now, path_health, deterministic_progression
[exit 0]

### Section 25 seed 1: 'work is waiting on capacity' — while the fleet admits

$ fm-decision-surface.sh check capacity-blocked
check: capacity-blocked · verdict: contradicted · admission band=preferred action=admit (absent; admission control is not configured for this home) · live work=1 · the fleet accepts another task now
[exit 3]

### Seed 2: 'the decision is still with the captain' — while its record is ruled

$ fm-decision-surface.sh check decision-pending demo-decision-ruled
check: decision-pending demo-decision-ruled · verdict: contradicted · the decision record is Done (Ruled decision) · it is ruled, not pending
[exit 3]

### Negative control: a decision that genuinely is still open

$ fm-decision-surface.sh check decision-pending demo-decision-open
check: decision-pending demo-decision-open · verdict: not-contradicted · the decision record is queued and still open
[exit 0]

### A decision with no durable record at all: unevaluable, not open

$ fm-decision-surface.sh check decision-pending never-recorded
check: decision-pending never-recorded · verdict: unevaluable · no durable decision record exists; open one with bin/fm-decision-hold.sh before reporting its status
[exit 4]

### Seed 3: dispatching an identity that is already live

$ fm-decision-surface.sh check duplicate-dispatch alpha
check: duplicate-dispatch alpha · verdict: contradicted · work already exists under this identity · state: working · source: run-step · running · steer or reconcile it rather than dispatching a second
[exit 3]

### Negative control: a genuinely new identity

$ fm-decision-surface.sh check duplicate-dispatch brand-new
check: duplicate-dispatch brand-new · verdict: not-contradicted · no live task record and no in-flight backlog row under this identity
[exit 0]

### The compensation ledger: owned deterministic work vs pending gaps

$ fm-decision-surface.sh owners
Deterministic work firstmate must not perform, and who performs it.

  owned    counting_workers
           owner: bin/fm-fleet-snapshot.sh --json .tasks
  owned    checking_capacity
           owner: bin/fm-admission.sh --json
  owned    polling
           owner: bin/fm-watch.sh through the emitted supervision protocol
  owned    known_dependencies
           owner: bin/fm-fleet-snapshot.sh --json .backlog.records[].unresolved_blocker_ids
  owned    pr_existence
           owner: bin/fm-fleet-snapshot.sh --json .tasks[].pr and bin/fm-pr-check.sh
  owned    verifier_passed
           owner: bin/fm-crew-state.sh
  owned    work_landed
           owner: bin/fm-landed-lib.sh and bin/fm-pr-merge.sh merge_verification=
  owned    terminal_state
           owner: bin/fm-wake-ledger.sh task --outcome landed|failed|abandoned
  owned    remembering_in_flight_work
           owner: bin/fm-session-start.sh digest over state/<id>.meta
  owned    reconciling_known_identifiers
           owner: bin/fm-fleet-snapshot.sh --json .main_inventory
  owned    duplicate_work
           owner: fm-decision-surface.sh check duplicate-dispatch
  owned    budgets
           owner: config/models.json concurrency cap enforced at bin/fm-spawn.sh, plus quota-axi
  owned    loopspec_continuation
           owner: bin/fm-loopspec.sh state|claim|finish

  PENDING  attempt_and_retry_counting
           needs: a durable per-work attempt and retry budget, so a retry decision reads a count instead of recalling one
           until then: until it lands, retry judgment stays with firstmate and its instruction stays in place
  PENDING  verifier_verdict_vocabulary
           needs: a shared PASS/FAIL/NO_VERIFIER_RAN verdict contract spanning every verifier, so no unobserved result can pass
           until then: bin/fm-crew-state.sh already refuses an uncorroborated checks-passed claim, but each verifier still reports in its own vocabulary
  PENDING  invoking_known_next_stage
           needs: a direct pipeline invocation replacing the harness keystroke handoff, after the shared verifier verdict lands
           until then: firstmate still triggers validation on the worker; the keystroke transition retires with its replacement, not before
  PENDING  deadline_and_time_gate_elapsed
           needs: a backlog time-gate evaluator that reports which queued work has become eligible
           until then: queued time gates are still re-read by firstmate at teardown and heartbeat
  PENDING  transition_legality
           needs: the platform allowed-transition projection resolved for a fleet work identity
           until then: see platform-seam; the projection exists but resolves platform identities only
  PENDING  why_not_now
           needs: the platform why_not_now primitive resolved for a fleet work identity
           until then: check capacity-blocked is the fleet-side subset that is answerable today
  PENDING  path_health
           needs: the platform path-health invariant model resolved for a fleet work identity
           until then: see platform-seam
  PENDING  deterministic_progression
           needs: platform transition actuation, so a deterministic transition costs no model turn
           until then: firstmate still advances lifecycle steps that CODE could advance alone

A pending row means the compensating instruction is retained on purpose.
[exit 0]

(ledger row count by status)
[{"status":"owned","rows":13},{"status":"pending","rows":8}]

### The platform seam: unconfigured => not wired, reachability never assumed

$ fm-decision-surface.sh platform-seam
platform surface: FirstMateDecisionSurface (rt.decision_surface)
  commands:   decision-surface <work_id> | decision-surface --fleet
  launcher:   not configured (config/decision-surface-platform)
  configured: absent
  reachable:  not_checked
  wiring:     not-wired
  owns:       why_not_now, available_transitions, forbidden_transitions, path_health, verification_state, review_state, certification_state, authority, reasoning_required, genuine_engineer_decisions
  retires when: the platform projection resolves this home fleet task ids, so a fleet task can be looked up by identity rather than by a platform work id
[exit 0]

### Seam safety: the config value is a PATH, never a shell command line

$ cat config/decision-surface-platform
/tmp/fm-ds-demo-home.yf84Rq/platform dir/launcher.sh; touch '/tmp/fm-ds-demo-home.yf84Rq/injected-marker'

$ fm-decision-surface.sh platform-seam --probe-platform
platform surface: FirstMateDecisionSurface (rt.decision_surface)
  commands:   decision-surface <work_id> | decision-surface --fleet
  launcher:   /tmp/fm-ds-demo-home.yf84Rq/platform dir/launcher.sh; touch '/tmp/fm-ds-demo-home.yf84Rq/injected-marker'
  configured: present
  reachable:  false
  wiring:     not-wired
  owns:       why_not_now, available_transitions, forbidden_transitions, path_health, verification_state, review_state, certification_state, authority, reasoning_required, genuine_engineer_decisions
  retires when: the platform projection resolves this home fleet task ids, so a fleet task can be looked up by identity rather than by a platform work id
[exit 0]
injection did NOT execute: no marker file created

### Reachable-but-foreign platform: reachable=true yet wiring stays not-wired

$ fm-decision-surface.sh platform-seam --probe-platform
platform surface: FirstMateDecisionSurface (rt.decision_surface)
  commands:   decision-surface <work_id> | decision-surface --fleet
  launcher:   /tmp/fm-ds-demo-home.yf84Rq/platform dir/launcher.sh
  configured: present
  reachable:  true
  wiring:     not-wired
  owns:       why_not_now, available_transitions, forbidden_transitions, path_health, verification_state, review_state, certification_state, authority, reasoning_required, genuine_engineer_decisions
  retires when: the platform projection resolves this home fleet task ids, so a fleet task can be looked up by identity rather than by a platform work id
[exit 0]
Evidence: Focused test suite run (13/13 ok)
ok - a capacity claim is refused against an admitting fleet and stands against a refusing one
ok - an undecidable admission forbids the capacity claim instead of passing it
ok - a ruled decision refutes a pending report, an open one supports it, and an absent one is unknown
ok - a live identity refuses a second dispatch and an unused one does not
ok - an incoherent inventory refuses to clear a dispatch
ok - an unreadable census makes every census-backed check unevaluable
ok - every ledger row names its owner or the capability it waits for
ok - the platform seam is wired only when it resolves this fleet's own work
ok - the seam reads one launcher path and never interprets it as a shell command line
ok - a probe that cannot be bounded does not run
ok - a launcher that ignores TERM cannot outlive the probe bound
ok - both scopes render and an unknown task is refused rather than narrated
ok - usage errors are refused before any state is read
Evidence: Red-capability mutation spot-check (two mutations, matching cases fail, clean revert)

Mutation 1 (disable capacity contradiction): not ok - a capacity claim against an admitting fleet must be refused: expected exit 3, got 0 Mutation 2 (hardcode reachable=>wired): not ok - reachable but foreign must not be reported as wired Both reverted; git status clean; suite green again (13/13).

Red-capability spot-check of tests/fm-decision-surface.test.sh (2026-08-09)

The intent claims the focus tests are red-capable (7 mutations were verified by the
author). Two of those mutation shapes were independently re-applied here to confirm
the suite actually fails when the enforcement is broken, then reverted.

Mutation 1: disable the capacity contradiction branch in check_capacity_blocked
  (bin/fm-decision-surface.sh: `[ "$action" = admit ]` -> `[ "$action" = never-admit ]`)
  Suite result:
    not ok - a capacity claim against an admitting fleet must be refused: expected exit 3, got 0
  -> the matching seed-1 case goes red.

Mutation 2: hardcode the seam wiring gate to reachable-implies-wired
  (bin/fm-decision-surface.sh: drop the fleet_identities_resolved > 0 condition)
  Suite result:
    not ok - reachable but foreign must not be reported as wired
  -> the reachable-but-foreign seam case goes red.

After `git checkout -- bin/fm-decision-surface.sh`, `git status --porcelain` is empty
and the full suite is green again (13/13 ok, ~8s).

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped

Push main to origin, or rebase your branch onto origin/main, before gating.

🔧 **Review** - 4 issues found → auto-fixed ✅
  • ⚠️ bin/fm-decision-surface.sh:361 - Platform-seam wiring detection uses raw substring containment (case &#34;$probe_out&#34; in *&#34;$id&#34;*) of fleet task ids against the probe's free-text output. A fleet id that appears coincidentally — inside a longer platform work id, a title describing the same program, or any prose — counts toward fleet_identities_resolved and a single match flips wiring to wired. The skill makes wired consequential: the platform projection becomes authoritative and the four transition-related pending rows retire, so a false positive reintroduces exactly the silent contradiction the seam exists to prevent. Match ids as whole tokens (e.g. split the output on non-id characters, or grep -F -w per id) or parse the projection's structured id fields.
  • ⚠️ bin/fm-decision-surface.sh:329 - run_timed invokes timeout/gtimeout without --kill-after (or -s KILL), so a platform launcher that traps or ignores SIGTERM keeps running after the 120s bound and the command-substitution read blocks indefinitely — contradicting the header's stated guarantee that "an unresponsive platform cannot wedge a read". The red-capability mutations only covered the no-bounding-tool case, not a TERM-immune child. Both GNU timeout and gtimeout support -k &lt;grace&gt;; add it to make the bound enforceable.
  • ℹ️ bin/fm-decision-surface.sh:511 - The human fleet-surface renderer prints capacity: &lt;band&gt; · &lt;action&gt; · null for every configured fleet: an active fm-admission.v1 record carries no top-level .reason field (only the inactive/unconfigured record does), and the renderer interpolates .capacity.reason directly. check_capacity_blocked already has the .reason // controlling_rules fallback; reuse it in the render_surface jq program so the common configured case doesn't display the literal string "null".
  • ℹ️ bin/fm-decision-surface.sh:134 - The argument parser lets a named subcommand silently override an earlier task-id argument: fm-decision-surface.sh &lt;task-id&gt; owners sets SUBCOMMAND=surface then overwrites it with owners and drops the task scope, and &lt;task-id&gt; check decision-pending &lt;id&gt; silently reassigns TARGET — while the reverse order (owners &lt;task-id&gt;) correctly dies with "unexpected argument". Guard the owners/platform-seam/check branches with the same [ -z &#34;$SUBCOMMAND&#34; ] check used in the positional branch so mixed invocations are refused as usage errors.

🔧 Fix: fix seam token match, probe kill grace, render, parsing
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-decision-surface.test.sh (13/13 ok, ~8s; re-run green after mutation spot-checks)
  • Manual end-to-end CLI demo against a canned fleet home: fm-decision-surface.sh fleet surface render, check capacity-blocked (exit 3 while admitting), check decision-pending ruled/open/unrecorded (exits 3/0/4), check duplicate-dispatch live/new identity (exits 3/0)
  • Manual ledger verification: fm-decision-surface.sh owners and owners --json — 13 owned + 8 pending = 21 rows, no private CFVC ids
  • Manual seam verification: platform-seam unconfigured (not-wired, reachable=not_checked), config value with appended ; touch marker does not execute and probes unreachable, reachable-but-foreign launcher stays not-wired
  • Red-capability mutation spot-check: disabled the capacity contradiction branch (seed-1 case failed: expected exit 3, got 0) and hardcoded reachable-implies-wired (foreign-seam case failed); both reverted with git checkout and clean git status --porcelain confirmed
  • Reviewed the AGENTS.md diff for the two intended compensation deletions (intake capacity/serialization, Validate run-step mapping) and the section 7 stub plus section 13 decision-surface load trigger
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

sbracewell64 and others added 3 commits August 9, 2026 13:27
Squash-reconcile of the fork trunk (48 commits through ed376cf, fork PRs
#8-#48 plus CI mirrors) rebased onto upstream main at 74230fc, resolving
72 conflicted files. The fork trunk carries: the fm-launch-lib.sh
one-owner launch refactor, the slot-base/contribution-target task base
contract (fm-task-base-lib.sh), the zero-budget model registry, fleet
admission control stages 0/1, the wake-outcome ledger, the LoopSpec
canonical representation, the research-approved-work corpus scanner, the
70% compaction doctrine, from-firstmate steer markers, the fleet launcher
menu and Windows-to-WSL bridge, forge-verified merge gates, and the
fleet-view per-argument cap fix.

Conflict-resolution decisions a reviewer cannot see from the diff alone:

- Where the fork carried an older in-flight import of upstream work
  (delivery contracts kunchenguid#1563, remote secondmate homes kunchenguid#1576, trace context
  kunchenguid#995, and every add/add remote-secondmate file), upstream's landed and
  further-evolved version wins outright; nothing fork-specific lived in
  those copies (verified blob-by-blob against upstream history).
- Upstream's --relaunch lifecycle control plane and the fork's task base
  contract both survive: base derivation and the brief base-contract
  guard are gated to fresh spawns (RELAUNCH=0) because a relaunch reuses
  the recorded worktree and bases, and the trunk may have moved since.
- Launch commands stay one-owner in bin/fm-launch-lib.sh: upstream's Muse
  Code adapter is ported into the library (template arm, model flag,
  effort mapping with max->ultra) instead of resurrecting the inline
  fm-spawn template, and fm-spawn's post-template FM_PI_HARNESS prefix is
  dropped because the library's pi templates carry the marker themselves.
- Herdr presentation spaces keep upstream's 0.8.0-floor semantics
  (config off/on/empty, two-argument fm_backend_herdr_presentation_enabled)
  across spawn, config inheritance, docs, and tests.
- fm-pr-check derives its per-task lock from state/<id>.meta explicitly
  so upstream's kunchenguid#1568 meta locking coexists with the fork's #34 landing
  records (fm_meta_lock_path rejects a .landing path).
- tests/lib.sh keeps both the fork's identity-verified background-process
  reaper and upstream's fm_fake_version_tool fixture.
- Merged prose follows upstream's X-mode -> Relay rename.

Pre-existing on both parents, not introduced here:
tests/fm-calm-pi-extension.test.sh fails under system Node 22
(ERR_UNKNOWN_FILE_EXTENSION), verified identical on the base commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ing operational truth

Firstmate reconstructed operational facts conversationally - capacity, live
work, decision status - and the drift from the records was silent. The incident
this addresses was a report that queued work would dispatch "as capacity frees"
while nothing in the fleet was capacity-bound. Every fact needed to refute that
sentence was already recorded; nothing made it get read, and nothing refused
the sentence.

Add bin/fm-decision-surface.sh: a read-only composer over the already-landed
deterministic owners, plus three `check` verdicts that refuse a claim structured
state contradicts (a capacity claim against an admitting fleet, a ruled decision
reported pending, a dispatch of an identity already in flight). It adds no fact
of its own; every field names the owner it was read from. An unreadable census,
an undecidable admission policy, or an absent decision record is `unevaluable`
- the fact may not be asserted at all - never a quiet pass.

Rewire the instruction surface onto those owners and delete what they replace:
the capacity reasoning at intake, which now defers to the surface and keeps only
the semantic serialization judgment, and the run-step mapping restatement in
Validate, which bin/fm-crew-state.sh already owns in full.

Where no owner has landed, mark rather than delete. `owners` prints the durable
compensation ledger: each row is either owned by a landed command or pending
with the capability that must land first, and the skill maps every pending row
to the instruction it deliberately keeps alive. Attempt and retry counting, a
shared verifier verdict vocabulary, the pipeline invocation that replaces the
keystroke handoff, and backlog time gates all remain firstmate's for now.

Declare the platform seam without depending on it. The deterministic platform
publishes a richer projection - why_not_now, allowed transitions, path health -
and `platform-seam --probe-platform` measures its wiring rather than assuming
it. Probed against platform f0da880 the launcher answers but resolves none of
this home's fleet task ids, so the seam stays not-wired; consuming a projection
of other identities as fleet truth would be the same silent contradiction the
surface exists to prevent. config/decision-surface-platform holds one launcher
path, never a command line, so a private config file cannot become a
shell-execution seam.

Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no
live fleet, worker, or platform. Each guarantee was confirmed by breaking it and
watching the suite fail; docs/verification/decision-surface.md records those
seven mutations and the probe evidence.
@sbracewell64
sbracewell64 force-pushed the fm/ae-factory-phase-e-firstmate-refactor branch from eea9575 to 07326e4 Compare August 9, 2026 17:34
@sbracewell64

Copy link
Copy Markdown
Author

Withdrawing this pull request: it was opened at the wrong venue and its head bundles work that does not belong in contribution history.

Two concrete problems with this PR as it stands:

  1. Wrong venue. This lane's landing target is the fork trunk. The branch was pushed to the fork, but the pull request was opened here against upstream.

  2. The head bundles the fork landing queue. The head commit 07326e4e is based on upstream/main and adds three commits, one of which is ef13c16 reconcile: land the accumulated fleet trunk on upstream main - 122 files, +23126/-664, squashing the entire accumulated fork landing queue into a single commit. Those commits are the fork's own landing queue and are not intended to be replayed or bundled into a contribution.

The actual change this lane produced is small and self-contained. It is being re-opened as a fork pull request cut fresh from the current fork trunk, carrying only its own net change.

For the record, no content was lost in the head being withdrawn: an audit against the branch point found no missing files, and every file whose content differs is one where upstream is ahead of that branch point.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant