feat(bin): consume a deterministic decision surface instead of reconstructing operational truth - #68
Conversation
…ing operational truth Firstmate reconstructed operational facts conversationally - capacity, live work, decision status - and the drift from the records was silent. The incident this addresses was a report that queued work would dispatch "as capacity frees" while nothing in the fleet was capacity-bound. Every fact needed to refute that sentence was already recorded; nothing made it get read, and nothing refused the sentence. Add bin/fm-decision-surface.sh: a read-only composer over the already-landed deterministic owners, plus three `check` verdicts that refuse a claim structured state contradicts (a capacity claim against an admitting fleet, a ruled decision reported pending, a dispatch of an identity already in flight). It adds no fact of its own; every field names the owner it was read from. An unreadable census, an undecidable admission policy, or an absent decision record is `unevaluable` - the fact may not be asserted at all - never a quiet pass. Rewire the instruction surface onto those owners and delete what they replace: the capacity reasoning at intake, which now defers to the surface and keeps only the semantic serialization judgment, and the run-step mapping restatement in Validate, which bin/fm-crew-state.sh already owns in full. Where no owner has landed, mark rather than delete. `owners` prints the durable compensation ledger: each row is either owned by a landed command or pending with the capability that must land first, and the skill maps every pending row to the instruction it deliberately keeps alive. Attempt and retry counting, a shared verifier verdict vocabulary, the pipeline invocation that replaces the keystroke handoff, and backlog time gates all remain firstmate's for now. Declare the platform seam without depending on it. The deterministic platform publishes a richer projection - why_not_now, allowed transitions, path health - and `platform-seam --probe-platform` measures its wiring rather than assuming it. Probed against platform f0da880 the launcher answers but resolves none of this home's fleet task ids, so the seam stays not-wired; consuming a projection of other identities as fleet truth would be the same silent contradiction the surface exists to prevent. config/decision-surface-platform holds one launcher path, never a command line, so a private config file cannot become a shell-execution seam. Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no live fleet, worker, or platform. Each guarantee was confirmed by breaking it and watching the suite fail; docs/verification/decision-surface.md records those seven mutations and the probe evidence.
… retry budget The trunk landed two changes this surface must consume rather than talk past. The identity-axis split replaced the overloaded `kind=` with role, deliverable, and stage, keeping `kind` only as a deprecated compatibility alias. The task projection read that alias, so it would have kept reading a field the census retains only for migration. It now projects the three axes, and a test pins them plus the absence of `kind` so a revert to the alias fails rather than passing quietly. The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes "should this be retried?" arithmetic over a recorded count. The compensation ledger still marked that row pending, and the skill still listed the instruction it was keeping alive. Both are wrong the moment the owner exists: a stale pending row is exactly the silent gap the ledger exists to prevent, and this file's own contract requires the row and its instruction to change in the same edit. The row is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
f7af6de to
351290d
Compare
|
Rebased onto the current trunk Reconciled with the landed vocabularyIdentity axes (#65). The task projection read Retry budget (#63). The AGENTS.md conflict was in the section 13 trigger list: the trunk added Merge screen
Fresh check state - 13 of 14 pass, not green overallThe CI workflow run concluded The single red is Awaiting the merge decision. |
…tructing operational truth (#68) * feat: consume a deterministic decision surface instead of reconstructing operational truth Firstmate reconstructed operational facts conversationally - capacity, live work, decision status - and the drift from the records was silent. The incident this addresses was a report that queued work would dispatch "as capacity frees" while nothing in the fleet was capacity-bound. Every fact needed to refute that sentence was already recorded; nothing made it get read, and nothing refused the sentence. Add bin/fm-decision-surface.sh: a read-only composer over the already-landed deterministic owners, plus three `check` verdicts that refuse a claim structured state contradicts (a capacity claim against an admitting fleet, a ruled decision reported pending, a dispatch of an identity already in flight). It adds no fact of its own; every field names the owner it was read from. An unreadable census, an undecidable admission policy, or an absent decision record is `unevaluable` - the fact may not be asserted at all - never a quiet pass. Rewire the instruction surface onto those owners and delete what they replace: the capacity reasoning at intake, which now defers to the surface and keeps only the semantic serialization judgment, and the run-step mapping restatement in Validate, which bin/fm-crew-state.sh already owns in full. Where no owner has landed, mark rather than delete. `owners` prints the durable compensation ledger: each row is either owned by a landed command or pending with the capability that must land first, and the skill maps every pending row to the instruction it deliberately keeps alive. Attempt and retry counting, a shared verifier verdict vocabulary, the pipeline invocation that replaces the keystroke handoff, and backlog time gates all remain firstmate's for now. Declare the platform seam without depending on it. The deterministic platform publishes a richer projection - why_not_now, allowed transitions, path health - and `platform-seam --probe-platform` measures its wiring rather than assuming it. Probed against platform f0da880 the launcher answers but resolves none of this home's fleet task ids, so the seam stays not-wired; consuming a projection of other identities as fleet truth would be the same silent contradiction the surface exists to prevent. config/decision-surface-platform holds one launcher path, never a command line, so a private config file cannot become a shell-execution seam. Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no live fleet, worker, or platform. Each guarantee was confirmed by breaking it and watching the suite fail; docs/verification/decision-surface.md records those seven mutations and the probe evidence. * no-mistakes(review): fix seam token match, probe kill grace, render, parsing * fix: reconcile the decision surface with the landed identity axes and retry budget The trunk landed two changes this surface must consume rather than talk past. The identity-axis split replaced the overloaded `kind=` with role, deliverable, and stage, keeping `kind` only as a deprecated compatibility alias. The task projection read that alias, so it would have kept reading a field the census retains only for migration. It now projects the three axes, and a test pins them plus the absence of `kind` so a revert to the alias fails rather than passing quietly. The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes "should this be retried?" arithmetic over a recorded count. The compensation ledger still marked that row pending, and the skill still listed the instruction it was keeping alive. Both are wrong the moment the owner exists: a stale pending row is exactly the silent gap the ledger exists to prevent, and this file's own contract requires the row and its instruction to change in the same edit. The row is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
…tructing operational truth (#68) * feat: consume a deterministic decision surface instead of reconstructing operational truth Firstmate reconstructed operational facts conversationally - capacity, live work, decision status - and the drift from the records was silent. The incident this addresses was a report that queued work would dispatch "as capacity frees" while nothing in the fleet was capacity-bound. Every fact needed to refute that sentence was already recorded; nothing made it get read, and nothing refused the sentence. Add bin/fm-decision-surface.sh: a read-only composer over the already-landed deterministic owners, plus three `check` verdicts that refuse a claim structured state contradicts (a capacity claim against an admitting fleet, a ruled decision reported pending, a dispatch of an identity already in flight). It adds no fact of its own; every field names the owner it was read from. An unreadable census, an undecidable admission policy, or an absent decision record is `unevaluable` - the fact may not be asserted at all - never a quiet pass. Rewire the instruction surface onto those owners and delete what they replace: the capacity reasoning at intake, which now defers to the surface and keeps only the semantic serialization judgment, and the run-step mapping restatement in Validate, which bin/fm-crew-state.sh already owns in full. Where no owner has landed, mark rather than delete. `owners` prints the durable compensation ledger: each row is either owned by a landed command or pending with the capability that must land first, and the skill maps every pending row to the instruction it deliberately keeps alive. Attempt and retry counting, a shared verifier verdict vocabulary, the pipeline invocation that replaces the keystroke handoff, and backlog time gates all remain firstmate's for now. Declare the platform seam without depending on it. The deterministic platform publishes a richer projection - why_not_now, allowed transitions, path health - and `platform-seam --probe-platform` measures its wiring rather than assuming it. Probed against platform f0da880 the launcher answers but resolves none of this home's fleet task ids, so the seam stays not-wired; consuming a projection of other identities as fleet truth would be the same silent contradiction the surface exists to prevent. config/decision-surface-platform holds one launcher path, never a command line, so a private config file cannot become a shell-execution seam. Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no live fleet, worker, or platform. Each guarantee was confirmed by breaking it and watching the suite fail; docs/verification/decision-surface.md records those seven mutations and the probe evidence. * no-mistakes(review): fix seam token match, probe kill grace, render, parsing * fix: reconcile the decision surface with the landed identity axes and retry budget The trunk landed two changes this surface must consume rather than talk past. The identity-axis split replaced the overloaded `kind=` with role, deliverable, and stage, keeping `kind` only as a deprecated compatibility alias. The task projection read that alias, so it would have kept reading a field the census retains only for migration. It now projects the three axes, and a test pins them plus the absence of `kind` so a revert to the alias fails rather than passing quietly. The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes "should this be retried?" arithmetic over a recorded count. The compensation ledger still marked that row pending, and the skill still listed the instruction it was keeping alive. Both are wrong the moment the owner exists: a stale pending row is exactly the silent gap the ledger exists to prevent, and this file's own contract requires the row and its instruction to change in the same edit. The row is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
…tructing operational truth (#68) * feat: consume a deterministic decision surface instead of reconstructing operational truth Firstmate reconstructed operational facts conversationally - capacity, live work, decision status - and the drift from the records was silent. The incident this addresses was a report that queued work would dispatch "as capacity frees" while nothing in the fleet was capacity-bound. Every fact needed to refute that sentence was already recorded; nothing made it get read, and nothing refused the sentence. Add bin/fm-decision-surface.sh: a read-only composer over the already-landed deterministic owners, plus three `check` verdicts that refuse a claim structured state contradicts (a capacity claim against an admitting fleet, a ruled decision reported pending, a dispatch of an identity already in flight). It adds no fact of its own; every field names the owner it was read from. An unreadable census, an undecidable admission policy, or an absent decision record is `unevaluable` - the fact may not be asserted at all - never a quiet pass. Rewire the instruction surface onto those owners and delete what they replace: the capacity reasoning at intake, which now defers to the surface and keeps only the semantic serialization judgment, and the run-step mapping restatement in Validate, which bin/fm-crew-state.sh already owns in full. Where no owner has landed, mark rather than delete. `owners` prints the durable compensation ledger: each row is either owned by a landed command or pending with the capability that must land first, and the skill maps every pending row to the instruction it deliberately keeps alive. Attempt and retry counting, a shared verifier verdict vocabulary, the pipeline invocation that replaces the keystroke handoff, and backlog time gates all remain firstmate's for now. Declare the platform seam without depending on it. The deterministic platform publishes a richer projection - why_not_now, allowed transitions, path health - and `platform-seam --probe-platform` measures its wiring rather than assuming it. Probed against platform f0da880 the launcher answers but resolves none of this home's fleet task ids, so the seam stays not-wired; consuming a projection of other identities as fleet truth would be the same silent contradiction the surface exists to prevent. config/decision-surface-platform holds one launcher path, never a command line, so a private config file cannot become a shell-execution seam. Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no live fleet, worker, or platform. Each guarantee was confirmed by breaking it and watching the suite fail; docs/verification/decision-surface.md records those seven mutations and the probe evidence. * no-mistakes(review): fix seam token match, probe kill grace, render, parsing * fix: reconcile the decision surface with the landed identity axes and retry budget The trunk landed two changes this surface must consume rather than talk past. The identity-axis split replaced the overloaded `kind=` with role, deliverable, and stage, keeping `kind` only as a deprecated compatibility alias. The task projection read that alias, so it would have kept reading a field the census retains only for migration. It now projects the three axes, and a test pins them plus the absence of `kind` so a revert to the alias fails rather than passing quietly. The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes "should this be retried?" arithmetic over a recorded count. The compensation ledger still marked that row pending, and the skill still listed the instruction it was keeping alive. Both are wrong the moment the owner exists: a stale pending row is exactly the silent gap the ledger exists to prevent, and this file's own contract requires the row and its instruction to change in the same edit. The row is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
…tructing operational truth (#68) * feat: consume a deterministic decision surface instead of reconstructing operational truth Firstmate reconstructed operational facts conversationally - capacity, live work, decision status - and the drift from the records was silent. The incident this addresses was a report that queued work would dispatch "as capacity frees" while nothing in the fleet was capacity-bound. Every fact needed to refute that sentence was already recorded; nothing made it get read, and nothing refused the sentence. Add bin/fm-decision-surface.sh: a read-only composer over the already-landed deterministic owners, plus three `check` verdicts that refuse a claim structured state contradicts (a capacity claim against an admitting fleet, a ruled decision reported pending, a dispatch of an identity already in flight). It adds no fact of its own; every field names the owner it was read from. An unreadable census, an undecidable admission policy, or an absent decision record is `unevaluable` - the fact may not be asserted at all - never a quiet pass. Rewire the instruction surface onto those owners and delete what they replace: the capacity reasoning at intake, which now defers to the surface and keeps only the semantic serialization judgment, and the run-step mapping restatement in Validate, which bin/fm-crew-state.sh already owns in full. Where no owner has landed, mark rather than delete. `owners` prints the durable compensation ledger: each row is either owned by a landed command or pending with the capability that must land first, and the skill maps every pending row to the instruction it deliberately keeps alive. Attempt and retry counting, a shared verifier verdict vocabulary, the pipeline invocation that replaces the keystroke handoff, and backlog time gates all remain firstmate's for now. Declare the platform seam without depending on it. The deterministic platform publishes a richer projection - why_not_now, allowed transitions, path health - and `platform-seam --probe-platform` measures its wiring rather than assuming it. Probed against platform f0da880 the launcher answers but resolves none of this home's fleet task ids, so the seam stays not-wired; consuming a projection of other identities as fleet truth would be the same silent contradiction the surface exists to prevent. config/decision-surface-platform holds one launcher path, never a command line, so a private config file cannot become a shell-execution seam. Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no live fleet, worker, or platform. Each guarantee was confirmed by breaking it and watching the suite fail; docs/verification/decision-surface.md records those seven mutations and the probe evidence. * no-mistakes(review): fix seam token match, probe kill grace, render, parsing * fix: reconcile the decision surface with the landed identity axes and retry budget The trunk landed two changes this surface must consume rather than talk past. The identity-axis split replaced the overloaded `kind=` with role, deliverable, and stage, keeping `kind` only as a deprecated compatibility alias. The task projection read that alias, so it would have kept reading a field the census retains only for migration. It now projects the three axes, and a test pins them plus the absence of `kind` so a revert to the alias fails rather than passing quietly. The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes "should this be retried?" arithmetic over a recorded count. The compensation ledger still marked that row pending, and the skill still listed the instruction it was keeping alive. Both are wrong the moment the owner exists: a stale pending row is exactly the silent gap the ledger exists to prevent, and this file's own contract requires the row and its instruction to change in the same edit. The row is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
Intent
Phase E of the software-factory commission: remove deterministic compensation from FirstMate's cognition surface, so operational facts are read from CODE rather than reconstructed in conversation.
The failure this addresses is not a wrong answer but a confident one. FirstMate reconstructed capacity, live work, and decision status conversationally, and the drift from the records was silent. The motivating incident was a report that queued work would dispatch "as capacity frees" while nothing in the fleet was capacity-bound. Every fact needed to refute that sentence was already recorded; nothing made it get read, and nothing refused the sentence.
What this adds
bin/fm-decision-surface.sh- a read-only composer over the already-landed deterministic owners, plus threecheckverdicts that refuse a claim structured state contradicts:check capacity-blockedcheck decision-pending <id>check duplicate-dispatch <id>It adds no fact of its own; every field names the owner it was read from. Verdicts are
contradicted(exit 3, the claim is forbidden),not-contradicted(exit 0, no landed owner refutes it, explicitly not a warrant of truth), andunevaluable(exit 4, no owner could answer, so the fact may not be asserted at all). The tri-state is deliberate: an unreadable census must not render as permission to make the claim.What this removes, and what it deliberately keeps
Two compensations had landed owners and their prose is deleted:
bin/fm-crew-state.shalready owns and emits directly - a one-owner violation as well as prohibited deterministic work. Replaced with a pointer plus the one actionable residue.Everything else that looked like a candidate has no landed owner, so it is marked rather than deleted.
fm-decision-surface.sh ownersprints the compensation ledger: each row is either owned by a landed command orpendingwith the capability that must land first. Pending rows are named by capability rather than by a private program identifier, so the pointer resolves for every reader of this repo. The new skill carries the pending-row-to-instruction mapping so the linkage stays unambiguous.Platform seam: declared, measured, not wired
The deterministic platform publishes a richer projection -
why_not_now, allowed transitions, path health.platform-seam --probe-platformmeasures its wiring rather than assuming it.Probed live: the launcher answers, but resolves zero of this home's fleet task ids, and its registry is fixture-backed. So the seam reports
wiring: not-wired. Reachable and wired are separate fields on purpose - consuming a projection of other identities as fleet truth would be the same silent contradiction the surface exists to prevent.config/decision-surface-platformholds one launcher path and never a command line: a command line would need shell interpretation, turning a private config file into an execution seam. A test proves appending shell syntax does not execute, and that a path containing spaces works unquoted.Verification
13 behavior cases run against canned
fm-fleet-snapshot.v1documents with no live fleet, worker, or platform. Each guarantee was confirmed by breaking it and watching the matching case fail;docs/verification/decision-surface.mdrecords those mutations and the probe evidence.Locally on this branch:
bin/fm-lint.shclean,bin/fm-doc-audience-check.shclean, suite green.Check state - please read
This PR has no green CI to point at, and I am not claiming otherwise.
Require no-mistakesattestation check is structurally red here. This branch was hand-cut and hand-opened, so it carries no head-bound pipeline attestation. That is a known structural red for this landing path, not a signal about the code.checks-passedoutcome was a claim no check examined. It is reported here as unverified rather than green, which is the same rule this change encodes.Merge screen: the fork trunk
faf131bis a strict ancestor, the branch adds exactly two commits, and zero landing-queue commits are replayed.