Skip to content

feat(bin): consume a deterministic decision surface instead of reconstructing operational truth - #68

Merged
sbracewell64 merged 3 commits into
mainfrom
fm/ae-factory-phase-e-landing
Aug 9, 2026
Merged

sbracewell64 merged 3 commits into
mainfrom
fm/ae-factory-phase-e-landing

Conversation

@sbracewell64

Copy link
Copy Markdown
Owner

Intent

Phase E of the software-factory commission: remove deterministic compensation from FirstMate's cognition surface, so operational facts are read from CODE rather than reconstructed in conversation.

The failure this addresses is not a wrong answer but a confident one. FirstMate reconstructed capacity, live work, and decision status conversationally, and the drift from the records was silent. The motivating incident was a report that queued work would dispatch "as capacity frees" while nothing in the fleet was capacity-bound. Every fact needed to refute that sentence was already recorded; nothing made it get read, and nothing refused the sentence.

What this adds

bin/fm-decision-surface.sh - a read-only composer over the already-landed deterministic owners, plus three check verdicts that refuse a claim structured state contradicts:

Check The claim it tests
check capacity-blocked "this is waiting on capacity"
check decision-pending <id> "that decision is still with the captain"
check duplicate-dispatch <id> "dispatch this work" when the identity may already be live

It adds no fact of its own; every field names the owner it was read from. Verdicts are contradicted (exit 3, the claim is forbidden), not-contradicted (exit 0, no landed owner refutes it, explicitly not a warrant of truth), and unevaluable (exit 4, no owner could answer, so the fact may not be asserted at all). The tri-state is deliberate: an unreadable census must not render as permission to make the claim.

What this removes, and what it deliberately keeps

Two compensations had landed owners and their prose is deleted:

  • The intake capacity and serialization paragraph now defers to the surface and keeps only the genuinely semantic serialization judgment.
  • The Validate run-step mapping sentence was a full restatement of the mapping bin/fm-crew-state.sh already owns and emits directly - a one-owner violation as well as prohibited deterministic work. Replaced with a pointer plus the one actionable residue.

Everything else that looked like a candidate has no landed owner, so it is marked rather than deleted. fm-decision-surface.sh owners prints the compensation ledger: each row is either owned by a landed command or pending with the capability that must land first. Pending rows are named by capability rather than by a private program identifier, so the pointer resolves for every reader of this repo. The new skill carries the pending-row-to-instruction mapping so the linkage stays unambiguous.

Platform seam: declared, measured, not wired

The deterministic platform publishes a richer projection - why_not_now, allowed transitions, path health. platform-seam --probe-platform measures its wiring rather than assuming it.

Probed live: the launcher answers, but resolves zero of this home's fleet task ids, and its registry is fixture-backed. So the seam reports wiring: not-wired. Reachable and wired are separate fields on purpose - consuming a projection of other identities as fleet truth would be the same silent contradiction the surface exists to prevent.

config/decision-surface-platform holds one launcher path and never a command line: a command line would need shell interpretation, turning a private config file into an execution seam. A test proves appending shell syntax does not execute, and that a path containing spaces works unquoted.

Verification

13 behavior cases run against canned fm-fleet-snapshot.v1 documents with no live fleet, worker, or platform. Each guarantee was confirmed by breaking it and watching the matching case fail; docs/verification/decision-surface.md records those mutations and the probe evidence.

Locally on this branch: bin/fm-lint.sh clean, bin/fm-doc-audience-check.sh clean, suite green.

Check state - please read

This PR has no green CI to point at, and I am not claiming otherwise.

  • The Require no-mistakes attestation check is structurally red here. This branch was hand-cut and hand-opened, so it carries no head-bound pipeline attestation. That is a known structural red for this landing path, not a signal about the code.
  • The validation pipeline did run against this exact content on a prior branch and reached its terminal state with all review findings resolved, but that run opened its PR at the wrong venue and its head bundled the fork landing queue, so it was withdrawn. This branch is cut fresh from the current fork trunk carrying only the net change.
  • That prior PR also had zero checks configured, so its checks-passed outcome was a claim no check examined. It is reported here as unverified rather than green, which is the same rule this change encodes.

Merge screen: the fork trunk faf131b is a strict ancestor, the branch adds exactly two commits, and zero landing-queue commits are replayed.

…ing operational truth

Firstmate reconstructed operational facts conversationally - capacity, live
work, decision status - and the drift from the records was silent. The incident
this addresses was a report that queued work would dispatch "as capacity frees"
while nothing in the fleet was capacity-bound. Every fact needed to refute that
sentence was already recorded; nothing made it get read, and nothing refused
the sentence.

Add bin/fm-decision-surface.sh: a read-only composer over the already-landed
deterministic owners, plus three `check` verdicts that refuse a claim structured
state contradicts (a capacity claim against an admitting fleet, a ruled decision
reported pending, a dispatch of an identity already in flight). It adds no fact
of its own; every field names the owner it was read from. An unreadable census,
an undecidable admission policy, or an absent decision record is `unevaluable`
- the fact may not be asserted at all - never a quiet pass.

Rewire the instruction surface onto those owners and delete what they replace:
the capacity reasoning at intake, which now defers to the surface and keeps only
the semantic serialization judgment, and the run-step mapping restatement in
Validate, which bin/fm-crew-state.sh already owns in full.

Where no owner has landed, mark rather than delete. `owners` prints the durable
compensation ledger: each row is either owned by a landed command or pending
with the capability that must land first, and the skill maps every pending row
to the instruction it deliberately keeps alive. Attempt and retry counting, a
shared verifier verdict vocabulary, the pipeline invocation that replaces the
keystroke handoff, and backlog time gates all remain firstmate's for now.

Declare the platform seam without depending on it. The deterministic platform
publishes a richer projection - why_not_now, allowed transitions, path health -
and `platform-seam --probe-platform` measures its wiring rather than assuming
it. Probed against platform f0da880 the launcher answers but resolves none of
this home's fleet task ids, so the seam stays not-wired; consuming a projection
of other identities as fleet truth would be the same silent contradiction the
surface exists to prevent. config/decision-surface-platform holds one launcher
path, never a command line, so a private config file cannot become a
shell-execution seam.

Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no
live fleet, worker, or platform. Each guarantee was confirmed by breaking it and
watching the suite fail; docs/verification/decision-surface.md records those
seven mutations and the probe evidence.
… retry budget

The trunk landed two changes this surface must consume rather than talk past.

The identity-axis split replaced the overloaded `kind=` with role, deliverable,
and stage, keeping `kind` only as a deprecated compatibility alias. The task
projection read that alias, so it would have kept reading a field the census
retains only for migration. It now projects the three axes, and a test pins them
plus the absence of `kind` so a revert to the alias fails rather than passing
quietly.

The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes
"should this be retried?" arithmetic over a recorded count. The compensation
ledger still marked that row pending, and the skill still listed the instruction
it was keeping alive. Both are wrong the moment the owner exists: a stale pending
row is exactly the silent gap the ledger exists to prevent, and this file's own
contract requires the row and its instruction to change in the same edit. The row
is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
@sbracewell64
sbracewell64 force-pushed the fm/ae-factory-phase-e-landing branch from f7af6de to 351290d Compare August 9, 2026 19:31
@sbracewell64

Copy link
Copy Markdown
Owner Author

Rebased onto the current trunk 825e965 to clear the conflict, and reconciled with the two changes that landed underneath this branch.

Reconciled with the landed vocabulary

Identity axes (#65). The task projection read kind, which the census now keeps only as a deprecated compatibility alias. It reads role, deliverable, and stage instead. Three assertions pin the axes and one pins the absence of kind, so reverting to the alias fails rather than passing quietly - verified by making that revert and watching the case go red.

Retry budget (#63). bin/fm-attempt.sh landed, which makes "should this be retried?" arithmetic over a recorded count. The compensation ledger still marked attempt_and_retry_counting as pending, and the skill still listed the instruction that row was keeping alive. A stale pending row is precisely the silent gap the ledger exists to prevent, and this change's own contract requires the row and its instruction to move together, so the row is now owned by bin/fm-attempt.sh and the skill's pending table drops it.

The AGENTS.md conflict was in the section 13 trigger list: the trunk added TASK_AXIS_BACKFILL: to the bootstrap-diagnostics line while this branch added the decision-surface trigger. Trunk's line kept, this branch's trigger reapplied beneath it.

Merge screen

  • current trunk 825e965 is a strict ancestor
  • 3 commits over the trunk: the feature, the pipeline's review fix, and this reconciliation
  • zero replayed landing-queue commits

Fresh check state - 13 of 14 pass, not green overall

The CI workflow run concluded success (run 31331891518): lint, repo invariants, coverage guard, Windows launcher bridge, stock macOS Bash snapshot compatibility, the Herdr lane, all four portable serial shards, both parallel shards, and the timing aggregate.

The single red is PR must be raised via no-mistakes. That workflow reads a head-bound attestation from refs/notes/no-mistakes; this branch was hand-cut and hand-opened, so no attestation exists for head 351290d. It is a structural red for this landing path, not a signal about the code - sibling fork-landing PRs #63 and #66 show the identical failure. No attestation was published to clear it, because the pipeline never validated this exact head and writing one would manufacture evidence.

Awaiting the merge decision.

@sbracewell64
sbracewell64 merged commit 3a1ab1d into main Aug 9, 2026
13 of 14 checks passed
sbracewell64 added a commit that referenced this pull request Aug 9, 2026
…tructing operational truth (#68)

* feat: consume a deterministic decision surface instead of reconstructing operational truth

Firstmate reconstructed operational facts conversationally - capacity, live
work, decision status - and the drift from the records was silent. The incident
this addresses was a report that queued work would dispatch "as capacity frees"
while nothing in the fleet was capacity-bound. Every fact needed to refute that
sentence was already recorded; nothing made it get read, and nothing refused
the sentence.

Add bin/fm-decision-surface.sh: a read-only composer over the already-landed
deterministic owners, plus three `check` verdicts that refuse a claim structured
state contradicts (a capacity claim against an admitting fleet, a ruled decision
reported pending, a dispatch of an identity already in flight). It adds no fact
of its own; every field names the owner it was read from. An unreadable census,
an undecidable admission policy, or an absent decision record is `unevaluable`
- the fact may not be asserted at all - never a quiet pass.

Rewire the instruction surface onto those owners and delete what they replace:
the capacity reasoning at intake, which now defers to the surface and keeps only
the semantic serialization judgment, and the run-step mapping restatement in
Validate, which bin/fm-crew-state.sh already owns in full.

Where no owner has landed, mark rather than delete. `owners` prints the durable
compensation ledger: each row is either owned by a landed command or pending
with the capability that must land first, and the skill maps every pending row
to the instruction it deliberately keeps alive. Attempt and retry counting, a
shared verifier verdict vocabulary, the pipeline invocation that replaces the
keystroke handoff, and backlog time gates all remain firstmate's for now.

Declare the platform seam without depending on it. The deterministic platform
publishes a richer projection - why_not_now, allowed transitions, path health -
and `platform-seam --probe-platform` measures its wiring rather than assuming
it. Probed against platform f0da880 the launcher answers but resolves none of
this home's fleet task ids, so the seam stays not-wired; consuming a projection
of other identities as fleet truth would be the same silent contradiction the
surface exists to prevent. config/decision-surface-platform holds one launcher
path, never a command line, so a private config file cannot become a
shell-execution seam.

Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no
live fleet, worker, or platform. Each guarantee was confirmed by breaking it and
watching the suite fail; docs/verification/decision-surface.md records those
seven mutations and the probe evidence.

* no-mistakes(review): fix seam token match, probe kill grace, render, parsing

* fix: reconcile the decision surface with the landed identity axes and retry budget

The trunk landed two changes this surface must consume rather than talk past.

The identity-axis split replaced the overloaded `kind=` with role, deliverable,
and stage, keeping `kind` only as a deprecated compatibility alias. The task
projection read that alias, so it would have kept reading a field the census
retains only for migration. It now projects the three axes, and a test pins them
plus the absence of `kind` so a revert to the alias fails rather than passing
quietly.

The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes
"should this be retried?" arithmetic over a recorded count. The compensation
ledger still marked that row pending, and the skill still listed the instruction
it was keeping alive. Both are wrong the moment the owner exists: a stale pending
row is exactly the silent gap the ledger exists to prevent, and this file's own
contract requires the row and its instruction to change in the same edit. The row
is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
sbracewell64 added a commit that referenced this pull request Aug 10, 2026
…tructing operational truth (#68)

* feat: consume a deterministic decision surface instead of reconstructing operational truth

Firstmate reconstructed operational facts conversationally - capacity, live
work, decision status - and the drift from the records was silent. The incident
this addresses was a report that queued work would dispatch "as capacity frees"
while nothing in the fleet was capacity-bound. Every fact needed to refute that
sentence was already recorded; nothing made it get read, and nothing refused
the sentence.

Add bin/fm-decision-surface.sh: a read-only composer over the already-landed
deterministic owners, plus three `check` verdicts that refuse a claim structured
state contradicts (a capacity claim against an admitting fleet, a ruled decision
reported pending, a dispatch of an identity already in flight). It adds no fact
of its own; every field names the owner it was read from. An unreadable census,
an undecidable admission policy, or an absent decision record is `unevaluable`
- the fact may not be asserted at all - never a quiet pass.

Rewire the instruction surface onto those owners and delete what they replace:
the capacity reasoning at intake, which now defers to the surface and keeps only
the semantic serialization judgment, and the run-step mapping restatement in
Validate, which bin/fm-crew-state.sh already owns in full.

Where no owner has landed, mark rather than delete. `owners` prints the durable
compensation ledger: each row is either owned by a landed command or pending
with the capability that must land first, and the skill maps every pending row
to the instruction it deliberately keeps alive. Attempt and retry counting, a
shared verifier verdict vocabulary, the pipeline invocation that replaces the
keystroke handoff, and backlog time gates all remain firstmate's for now.

Declare the platform seam without depending on it. The deterministic platform
publishes a richer projection - why_not_now, allowed transitions, path health -
and `platform-seam --probe-platform` measures its wiring rather than assuming
it. Probed against platform f0da880 the launcher answers but resolves none of
this home's fleet task ids, so the seam stays not-wired; consuming a projection
of other identities as fleet truth would be the same silent contradiction the
surface exists to prevent. config/decision-surface-platform holds one launcher
path, never a command line, so a private config file cannot become a
shell-execution seam.

Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no
live fleet, worker, or platform. Each guarantee was confirmed by breaking it and
watching the suite fail; docs/verification/decision-surface.md records those
seven mutations and the probe evidence.

* no-mistakes(review): fix seam token match, probe kill grace, render, parsing

* fix: reconcile the decision surface with the landed identity axes and retry budget

The trunk landed two changes this surface must consume rather than talk past.

The identity-axis split replaced the overloaded `kind=` with role, deliverable,
and stage, keeping `kind` only as a deprecated compatibility alias. The task
projection read that alias, so it would have kept reading a field the census
retains only for migration. It now projects the three axes, and a test pins them
plus the absence of `kind` so a revert to the alias fails rather than passing
quietly.

The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes
"should this be retried?" arithmetic over a recorded count. The compensation
ledger still marked that row pending, and the skill still listed the instruction
it was keeping alive. Both are wrong the moment the owner exists: a stale pending
row is exactly the silent gap the ledger exists to prevent, and this file's own
contract requires the row and its instruction to change in the same edit. The row
is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
sbracewell64 added a commit that referenced this pull request Aug 11, 2026
…tructing operational truth (#68)

* feat: consume a deterministic decision surface instead of reconstructing operational truth

Firstmate reconstructed operational facts conversationally - capacity, live
work, decision status - and the drift from the records was silent. The incident
this addresses was a report that queued work would dispatch "as capacity frees"
while nothing in the fleet was capacity-bound. Every fact needed to refute that
sentence was already recorded; nothing made it get read, and nothing refused
the sentence.

Add bin/fm-decision-surface.sh: a read-only composer over the already-landed
deterministic owners, plus three `check` verdicts that refuse a claim structured
state contradicts (a capacity claim against an admitting fleet, a ruled decision
reported pending, a dispatch of an identity already in flight). It adds no fact
of its own; every field names the owner it was read from. An unreadable census,
an undecidable admission policy, or an absent decision record is `unevaluable`
- the fact may not be asserted at all - never a quiet pass.

Rewire the instruction surface onto those owners and delete what they replace:
the capacity reasoning at intake, which now defers to the surface and keeps only
the semantic serialization judgment, and the run-step mapping restatement in
Validate, which bin/fm-crew-state.sh already owns in full.

Where no owner has landed, mark rather than delete. `owners` prints the durable
compensation ledger: each row is either owned by a landed command or pending
with the capability that must land first, and the skill maps every pending row
to the instruction it deliberately keeps alive. Attempt and retry counting, a
shared verifier verdict vocabulary, the pipeline invocation that replaces the
keystroke handoff, and backlog time gates all remain firstmate's for now.

Declare the platform seam without depending on it. The deterministic platform
publishes a richer projection - why_not_now, allowed transitions, path health -
and `platform-seam --probe-platform` measures its wiring rather than assuming
it. Probed against platform f0da880 the launcher answers but resolves none of
this home's fleet task ids, so the seam stays not-wired; consuming a projection
of other identities as fleet truth would be the same silent contradiction the
surface exists to prevent. config/decision-surface-platform holds one launcher
path, never a command line, so a private config file cannot become a
shell-execution seam.

Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no
live fleet, worker, or platform. Each guarantee was confirmed by breaking it and
watching the suite fail; docs/verification/decision-surface.md records those
seven mutations and the probe evidence.

* no-mistakes(review): fix seam token match, probe kill grace, render, parsing

* fix: reconcile the decision surface with the landed identity axes and retry budget

The trunk landed two changes this surface must consume rather than talk past.

The identity-axis split replaced the overloaded `kind=` with role, deliverable,
and stage, keeping `kind` only as a deprecated compatibility alias. The task
projection read that alias, so it would have kept reading a field the census
retains only for migration. It now projects the three axes, and a test pins them
plus the absence of `kind` so a revert to the alias fails rather than passing
quietly.

The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes
"should this be retried?" arithmetic over a recorded count. The compensation
ledger still marked that row pending, and the skill still listed the instruction
it was keeping alive. Both are wrong the moment the owner exists: a stale pending
row is exactly the silent gap the ledger exists to prevent, and this file's own
contract requires the row and its instruction to change in the same edit. The row
is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
sbracewell64 added a commit that referenced this pull request Aug 11, 2026
…tructing operational truth (#68)

* feat: consume a deterministic decision surface instead of reconstructing operational truth

Firstmate reconstructed operational facts conversationally - capacity, live
work, decision status - and the drift from the records was silent. The incident
this addresses was a report that queued work would dispatch "as capacity frees"
while nothing in the fleet was capacity-bound. Every fact needed to refute that
sentence was already recorded; nothing made it get read, and nothing refused
the sentence.

Add bin/fm-decision-surface.sh: a read-only composer over the already-landed
deterministic owners, plus three `check` verdicts that refuse a claim structured
state contradicts (a capacity claim against an admitting fleet, a ruled decision
reported pending, a dispatch of an identity already in flight). It adds no fact
of its own; every field names the owner it was read from. An unreadable census,
an undecidable admission policy, or an absent decision record is `unevaluable`
- the fact may not be asserted at all - never a quiet pass.

Rewire the instruction surface onto those owners and delete what they replace:
the capacity reasoning at intake, which now defers to the surface and keeps only
the semantic serialization judgment, and the run-step mapping restatement in
Validate, which bin/fm-crew-state.sh already owns in full.

Where no owner has landed, mark rather than delete. `owners` prints the durable
compensation ledger: each row is either owned by a landed command or pending
with the capability that must land first, and the skill maps every pending row
to the instruction it deliberately keeps alive. Attempt and retry counting, a
shared verifier verdict vocabulary, the pipeline invocation that replaces the
keystroke handoff, and backlog time gates all remain firstmate's for now.

Declare the platform seam without depending on it. The deterministic platform
publishes a richer projection - why_not_now, allowed transitions, path health -
and `platform-seam --probe-platform` measures its wiring rather than assuming
it. Probed against platform f0da880 the launcher answers but resolves none of
this home's fleet task ids, so the seam stays not-wired; consuming a projection
of other identities as fleet truth would be the same silent contradiction the
surface exists to prevent. config/decision-surface-platform holds one launcher
path, never a command line, so a private config file cannot become a
shell-execution seam.

Twelve behavior cases run against canned fm-fleet-snapshot.v1 documents with no
live fleet, worker, or platform. Each guarantee was confirmed by breaking it and
watching the suite fail; docs/verification/decision-surface.md records those
seven mutations and the probe evidence.

* no-mistakes(review): fix seam token match, probe kill grace, render, parsing

* fix: reconcile the decision surface with the landed identity axes and retry budget

The trunk landed two changes this surface must consume rather than talk past.

The identity-axis split replaced the overloaded `kind=` with role, deliverable,
and stage, keeping `kind` only as a deprecated compatibility alias. The task
projection read that alias, so it would have kept reading a field the census
retains only for migration. It now projects the three axes, and a test pins them
plus the absence of `kind` so a revert to the alias fails rather than passing
quietly.

The durable attempt-and-retry budget landed as bin/fm-attempt.sh, which makes
"should this be retried?" arithmetic over a recorded count. The compensation
ledger still marked that row pending, and the skill still listed the instruction
it was keeping alive. Both are wrong the moment the owner exists: a stale pending
row is exactly the silent gap the ledger exists to prevent, and this file's own
contract requires the row and its instruction to change in the same edit. The row
is now owned by bin/fm-attempt.sh and the skill's pending table drops it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant