feat(bin): verify the model a dispatched worker actually ran on - #29
Merged
Merged
Conversation
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
added 9 commits
August 4, 2026 07:40
`model=` in state/<id>.meta is what firstmate REQUESTED at intake, sometimes after a deliberate quota-balanced choice. Nothing verified what RAN. A worker served below its dispatched tier still reports done and the record still reads as the intended model, so later quota-aware dispatch built on that record is fiction. Placement, and why not the alternatives: - Not at spawn time: the worker has produced no turn yet, so there is nothing to compare. Verification is necessarily after the fact. - Not a PostToolUse guard: that surface exposes `resolvedModel` for the harness's OWN delegation tool, which a firstmate primary already denies (docs/subagent-guard.md), and it never observes a bin/fm-spawn.sh dispatch - the record actually at risk. - Not folded into bin/fm-crew-state.sh: that helper owns one contract, reconciling current run state. Model provenance is orthogonal to it. Shipped as bin/fm-model-verify.sh, a read-only verifier surfaced through the fleet snapshot firstmate already reviews every heartbeat. Evidence is the harness's own session transcript: the runtime writes it, the agent never authors it, and it records the model that served each assistant turn. Asking the worker instead would query the one party that cannot see the answer. It fails loudly rather than reporting compliance. `match` requires that a model was actually read and compared. Every path that cannot read the truth - no evidence adapter for the harness, evidence unlocatable or unreadable, jq absent, evidence unattributable to this dispatch - ends in `unverifiable`, never in silence. `pending` and `unpinned` are explicit no-verdict outcomes and never render as verified. bin/fm-spawn.sh now records `spawned_at=`. A worktree from a reusable pool can still carry a previous occupant's transcripts; without that anchor the two occupants cannot be told apart, which was observed live across the development fleet. The view raises a Model Routing section only when a worker did not provably run as dispatched, so correctly routed work renders exactly as before. The composition the source investigation flagged as unverified - that a PostToolUse hook on a delegation call sees `resolvedModel` - was confirmed empirically before any of this was built. It holds. That evidence, the transcript ground truth, the path encoding, and the live observation are recorded in docs/verification/model-verification.md. Also fixes an unrelated pre-existing break in the test runner's coverage guard: its lists are built with `LC_ALL=C sort`, but its `comm` calls ran under the ambient locale, which reads a C-sorted list as unsorted, exits nonzero, and lets `set -eu` abort the guard with nothing but comm's warning.
…atches Full CI caught what the pipeline's focused test step did not: the terminal model-routing refusal blocked legitimate teardown across five lanes - scouts with a report present, empty secondmate homes, zellij ghosts, Orca scouts. A harness with no evidence adapter can never produce a verdict, so refusing whenever one was absent made non-forced cleanup permanently impossible for every non-claude worker. That is a fleet-wide regression, not the boundary the refusal was meant to draw. The verdict is now ALWAYS surfaced before cleanup, so no worker's model provenance is discarded unseen. Only the refusal is conditional: a mismatch always blocks, and an absent verdict blocks only when the dispatch was verifiable in principle - a harness with an evidence adapter and a pinned model. `--force` keeps its existing discard authority, and only non-forced teardown refuses. Also fixed, all found by full CI: - The macOS lane is a pinned test-count guard, not a portability check: every test passes under stock bash 3.2. Four tests were added to the snapshot/fleet-view suite, so the pinned count moves 15 -> 19 rather than the suite being trimmed to fit a stale guard. - Three watermark cases used the bare spawn helper, so ship spawns correctly refused for a missing --mode. They now use the ship helper like their peers. - fm-teardown.sh gained a dependency on fm-model-verify.sh, which two fixtures that assemble a minimal bin/ did not provide. Both now supply it, so a genuinely missing verifier still refuses loudly rather than being papered over in the fixture. - The Herdr projection comparison masked only container ids, so per-dispatch identity fields that can never agree between two spawns read as a projection difference. Those fields are masked too, keeping the comparison on what it exists to test. Terminal-mode boundary is now pinned in both directions: no block for a harness that can never produce a verdict or a dispatch that pinned no model, always a block on a mismatch.
Rebasing onto main composed two additive test sets in the snapshot/fleet-view suite - the landed fleet-telemetry work and model verification - so the stock macOS lane's pinned count moves 19 -> 23 to match what the suite now runs.
HelloWorldSungin
force-pushed
the
fm/fm-subagent-model-routing-guard
branch
from
August 4, 2026 08:01
1af9d4a to
45f2e55
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Firstmate resolves a concrete model and effort at intake, sometimes after a deliberate quota-balanced choice, and records
model=in task metadata. That record is what was requested. Nothing verified what ran. A worker served below its dispatched tier still reports done, and the record still reads as the intended model, so later quota-aware dispatch decisions built on that record are fiction.The failure class was reproduced one level down, in the harness's own subagent routing: model resolution is explicit model, then the named subagent type's own definition frontmatter, then the parent, so a contract assuming "an omitted model inherits the parent" is wrong about the middle clause. Real transcripts on the development host showed an Opus 4.8 parent's omitted-model dispatch running on Sonnet 5, and a Fable 5 parent's running on Opus 4.8. Other dispatches in the same sample inherited correctly, which is what makes the class dangerous: it is right often enough to look right.
Where the check belongs
Three placements were considered, and the reasoning is in the commit message:
PostToolUseguard. That surface exposesresolvedModelfor the harness's own delegation tool, which a firstmate primary already denies (docs/subagent-guard.md), and it never observes abin/fm-spawn.shdispatch - the record actually at risk.bin/fm-crew-state.sh. That helper owns one contract, reconciling current run state; model provenance is orthogonal.Shipped as
bin/fm-model-verify.sh, a read-only verifier surfaced through the fleet snapshot, plus reconciliation at the shared terminal boundary so a fast mismatched worker cannot finish and be cleaned up before it is ever verified.Evidence, and the caveat that gated the design
The evidence is the harness's own session transcript: the runtime writes it, the agent never authors it, and it records the model that served each assistant turn. Asking the worker would query the one party that cannot see the answer.
The source investigation flagged one composition as unverified - that a
PostToolUsehook on a delegation call actually seesresolvedModel. That was settled empirically before anything was built, and it holds; the payload and the full evidence are recorded indocs/verification/model-verification.md. It is not the lever the shipped verifier uses, for the reason above, but it is documented as the verified mechanism if the delegation surface ever needs its own check.Fails loudly, by design
A verdict of
matchrequires that a model was actually read and compared. Every path that cannot read the truth ends inunverifiable: no evidence adapter for the harness, evidence unlocatable or unreadable,jqabsent, no working directory recorded, a malformed dispatch anchor, an absent model record, a newline-bearing store path, or evidence that cannot be attributed to this dispatch.pendingandunpinnedare explicit no-verdict outcomes that never render as verified.Only Claude has an empirically verified evidence source and is wired. Codex, OpenCode, Pi, Grok, and Kimi report
unverifiablerather than being assumed correct, following the existing rule that a harness integration is validated against the real binary before it is trusted.bin/fm-spawn.shnow records a dispatch watermark and the canonical evidence store, and pins Claude launches to that store. A worktree from a reusable pool still carries a previous occupant's transcripts, and without that binding the two occupants cannot be told apart - observed live across the development fleet.Deliberate design decision, recorded so it is not re-raised
The fleet view's Model Routing section surfaces
mismatchandunverifiableonly, and deliberately excludespending. Every freshly spawned worker is brieflypending, so including it would fire the section on healthy routine dispatch, and an indicator that alarms constantly gets ignored. The residual gap - a worker that never produces a readable model turn stayspendingand is not raised in the human view - is noted honestly indocs/model-verification.md, and the structured snapshot's per-task verdict remains where that distinction is visible.No behavior change for correctly routed work: the section renders nothing at all when every worker verifies.
Findings the pipeline fixed on top
Recorded for review, since they are corrections the original change missed:
--forceretains its existing power to discard, and only non-forced teardown refuses.findandstatfailures were silently skipped, which could collapse topendingwith exit 0 and contradict the fail-loud invariant.CLAUDE_CONFIG_DIRwhile spawn recorded only basenames, so a restart under a different store could attribute a stale matching transcript asmatch.model=record was treated asunpinned, and a malformedspawned_atsilently degraded to the legacy no-anchor path; both now reportunverifiable...lexically before resolving symlinks, solink/../cfgcould select a different physical store than the path names.Also fixed
bin/fm-test-run.sh's coverage guard was already broken on unmodifiedmain: its lists are built withLC_ALL=C sortbut itscommcalls ran under the ambient locale, which reads a C-sorted list as unsorted, exits nonzero, and letsset -euabort the guard with nothing but comm's warning on stderr.Verification
bin/fm-lint.sh,bin/fm-doc-audience-check.sh, and the test coverage guard all pass. New coverage lives intests/fm-model-verify.test.sh, with additions to the fleet-snapshot/view, spawn, and teardown suites.docs/model-verification.mdowns the contract;docs/verification/model-verification.mdholds the dated empirical evidence.Known unrelated pre-existing failure
tests/fm-calm-pi-extension.test.shfails identically on unmodifiedmainwith/calm left collapsed thinking labels in the transcript. It is untouched here and needs its own task.Pipeline
Updates from git push no-mistakes
This pull request was raised by hand on the fork. The no-mistakes pipeline ran review, test, document, and lint against exactly these commits and pushed them here, but its PR-creation step resolved the target to this fork's upstream parent and opened the pull request there by mistake; that upstream pull request has been closed. Only the PR creation was misrouted - no pipeline step was skipped, and no output below is reconstructed.