Skip to content

fix(spawn): enforce verified reasoning effort launches - #2141

Closed
coreldh wants to merge 9 commits into
kunchenguid:mainfrom
coreldh:fm/c0810n-fm-1993
Closed

coreldh wants to merge 9 commits into
kunchenguid:mainfrom
coreldh:fm/c0810n-fm-1993

Conversation

@coreldh

@coreldh coreldh commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Intent

Get upstream PR #1993 green for maintainer merge. Preserve the fail-closed invariant: fm-spawn must never record a requested reasoning effort unless the selected harness launches with that exact effort, or refuse before metadata and launch. Codex 0.147.0 must emit model_reasoning_effort=max for gpt-5.6-luna. The PR already implements that behavior; this follow-up repairs the pre-existing Kimi fixture expectation exposed by the invariant. Kimi has no verified effort launch flag, so its former --effort high success expectation is invalid. Keep normal Kimi model launch coverage without an effort, and add a regression proving high effort refuses before metadata, pane creation, launch command, or brief delivery. Update the existing PR #1993 branch only; no force-push, no second PR, no merge. Local changed-scope aggregate is NOT_VERIFIABLE because it does not terminate after Herdr backend coverage on this host; focused Kimi red/green proof is load-bearing.

What Changed

  • Fail closed before metadata or launch when a harness cannot express the requested reasoning effort, including remote secondmate endpoint reuse.
  • Pass Codex max effort through as model_reasoning_effort="max", while retaining only verified effort mappings for other harnesses.
  • Update adapter guidance and regression coverage for Kimi, Muse, remote lifecycle, and dispatch-profile effort handling.

Risk Assessment

✅ Low: The shared capability predicate now covers remote reused endpoints while the remote-only sentinel normalization preserves ordinary no-effort launches and local invalid efforts refuse early.

Testing

Confirmed the target head, ran the focused Kimi harness suite, and captured a direct CLI-flow artifact showing normal model launch and brief delivery plus high-effort refusal before metadata, pane creation, launch, or brief delivery. The known changed-scope aggregate remains NOT_VERIFIABLE on this host and was not run.

Evidence: Kimi end-to-end CLI flow

Normal Kimi model-only launch exited 0, invoked kimi --model kimi-code/k3 --auto, and delivered the brief. A requested --effort high exited 1 before metadata, launch, brief delivery, or tmux pane creation.

=== normal Kimi model-only launch ===
exit=0
output=warning: /var/folders/cq/xf4qcb9j0qzc2dbh173mflbm0000gn/T//fm-kimi-harness.d6GD6H/manual-evidence-normal/home/data/manual-kimi-normal-z1/brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
spawned manual-kimi-normal-z1 harness=kimi kind=ship mode=no-mistakes yolo=off window=firstmate:fm-manual-kimi-normal-z1 worktree=/var/folders/cq/xf4qcb9j0qzc2dbh173mflbm0000gn/T//fm-kimi-harness.d6GD6H/manual-evidence-normal/wt
launch='/var/folders/cq/xf4qcb9j0qzc2dbh173mflbm0000gn/T//fm-kimi-harness.d6GD6H/manual-evidence-normal/fake/fakebin/kimi' --model 'kimi-code/k3' --auto
brief=Read the brief at /private/var/folders/cq/xf4qcb9j0qzc2dbh173mflbm0000gn/T/fm-kimi-harness.d6GD6H/manual-evidence-normal/home/data/manual-kimi-normal-z1/brief.md and follow it exactly.
metadata=window=firstmate:fm-manual-kimi-normal-z1;endpoint_task_id=manual-kimi-normal-z1;worktree=/var/folders/cq/xf4qcb9j0qzc2dbh173mflbm0000gn/T//fm-kimi-harness.d6GD6H/manual-evidence-normal/wt;project=/var/folders/cq/xf4qcb9j0qzc2dbh173mflbm0000gn/T/fm-kimi-harness.d6GD6H/manual-evidence-normal/project;harness=kimi;kind=ship;mode=no-mistakes;yolo=off;tasktmp=/tmp/fm-manual-kimi-normal-z1;model=kimi-code/k3;effort=default;
=== requested high-effort Kimi launch ===
exit=1
output=error: refusing kimi spawn with effort 'high': no verified launch flag expresses that exact effort
metadata_exists=no
launch_bytes=       0
brief_bytes=       0
tmux_pane_creation=no
Evidence: Focused Kimi test transcript

Focused Kimi harness behavior transcript for target b4ddd853e88894b8d4e297775d35cc203520709a.

target_commit=b4ddd853e88894b8d4e297775d35cc203520709a
ok - Kimi hook install is idempotent and removal restores every foreign config byte
ok - Kimi hook removal preserves owned newline boundaries and pristine bytes
ok - Kimi hook install refuses missing, malformed, and surprising config without writing
ok - Kimi hook install refuses without jq before any config write
ok - fm-spawn: kimi launches, delivers its brief, and registers a guarded turn-end token
ok - fm-spawn: Kimi non-default effort refuses before metadata and launch
ok - fm-spawn: local Kimi effort sentinel refuses before metadata and launch
ok - Kimi hook stays silent and inert without a Firstmate registry token
ok - fm-spawn: unsafe Kimi global config refuses before pane creation
ok - fm-teardown: Kimi task pointer and registry token are removed
ok - fm-spawn: Kimi fallback expands the active HOME
ok - fm-spawn: missing Kimi executable refuses before pane creation
ok - fm-spawn: kimi treats a silent pointer drop as a failed spawn
ok - fm-spawn: kimi never sends the brief pointer before an observable ready signal
ok - fm-harness: markerless kimi is detected by ancestry after env-marker precedence
lock acquired: harness pid 5328
ok - fm-lock recognizes Kimi ancestry and live lock holders
ok - busy detection: real Kimi moon-plus-middot captures require its harness while idle labels stay idle
ok - fm-watch classifies Kimi as unknown rather than from its spinner, and Grok's fallback stays isolated
ok - composer classifier: kimi's existing bordered > shape is already safe without an override

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed (3) ✅
  • 🚨 bin/fm-spawn.sh:1371 - The changed local guard at line 1371 is bypassed by the earlier remote-secondmate return path: when a remote Kimi endpoint is already alive, fm-remote-secondmate-control.sh launch returns its route before it invokes remote fm-spawn; the parent then writes effort=high at bin/fm-spawn.sh:608. Thus fm-spawn <id> --secondmate --harness kimi --effort high can record an unlaunched, unsupported effort. This contradicts the required criterion, “fm-spawn must never record a requested reasoning effort unless the selected harness launches with that exact effort, or refuse before metadata and launch.” Put the same capability check at the parent remote-secondmate boundary (before readiness/inheritance/route reuse), using a helper defined before that branch, so a reused endpoint cannot bypass it.

🔧 Fix: Fail close remote Kimi effort reuse
1 error still open:

  • 🚨 bin/fm-spawn.sh:475 - The new shared predicate treats the remote no-effort sentinel - as an unsupported requested effort. spawn_remote_secondmate sets effort=${EFFORT:--} and calls this helper before contacting the remote host, so a normal remote secondmate spawn with no --effort now refuses instead of preserving the documented/default launch. Accept - as the remote “no requested effort” sentinel (or normalize it to empty before the shared predicate).

🔧 Fix: Normalize remote no-effort sentinel
1 error still open:

  • 🚨 bin/fm-spawn.sh:388 - The sentinel fix makes --effort - a successful local spawn for every harness: the changed && [ "$effort" != - ] || return 0 path emits no effort flag, while metadata later records effort=-. For example, a local Kimi spawn with --effort - now violates the required criterion, “fm-spawn must never record a requested reasoning effort unless the selected harness launches with that exact effort, or refuse before metadata and launch.” Keep - as a remote transport sentinel only—normalize it at the remote caller before invoking the shared predicate—so a user-supplied --effort - remains fail-closed.

🔧 Fix: Scope effort sentinel to remote transport
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-kimi-harness.test.sh
  • Manual fixture-backed Kimi CLI flow captured in kimi-user-flow-transcript.txt
  • git rev-parse HEAD confirmed b4ddd853e88894b8d4e297775d35cc203520709a
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@coreldh

coreldh commented Aug 12, 2026

Copy link
Copy Markdown
Contributor Author

Hi Kun — one consolidated status update on our four open PRs, so you can see at a glance which one is actually ready for you and which are not.

#2141 — ready for your review/merge. All 13 required checks are green at head b4ddd853, and the PR is mergeable/clean. This is the only one of the four we are asking you to look at right now.

#2137 — still shows zero CI checks. The workflows have not run on it, so there is no validation surface to judge it by. Following up on my earlier note: #2134's checks have since run, so that part of my previous message is now out of date — only #2137 is still missing them.

#2134 — not merge-ready. Its checks ran and came back with one failure and one cancelled run against 11 passing. We are repairing it on our side; please don't spend time on it yet.

#1968 — not merge-ready. Two failing checks against 12 passing. Also under repair here.

No action needed on #2134 or #1968 beyond leaving them open. For #2137, enabling or triggering the workflows would let us see whether it stands up. Thanks.

@coreldh

coreldh commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

Closing this — main now covers the Codex max case via #4497, and the maintainers settled on catalog-gated max (#3863). The broader refuse-instead-of-omit behavior also conflicts with the documented record-and-omit effort contract. Thanks!

@coreldh coreldh closed this Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant