Conversation
… both focused regressions pass
|
Hi Kun — one consolidated status update on our four open PRs, so you can see at a glance which one is actually ready for you and which are not. #2141 — ready for your review/merge. All 13 required checks are green at head #2137 — still shows zero CI checks. The workflows have not run on it, so there is no validation surface to judge it by. Following up on my earlier note: #2134's checks have since run, so that part of my previous message is now out of date — only #2137 is still missing them. #2134 — not merge-ready. Its checks ran and came back with one failure and one cancelled run against 11 passing. We are repairing it on our side; please don't spend time on it yet. #1968 — not merge-ready. Two failing checks against 12 passing. Also under repair here. No action needed on #2134 or #1968 beyond leaving them open. For #2137, enabling or triggering the workflows would let us see whether it stands up. Thanks. |
Intent
Get upstream PR #1993 green for maintainer merge. Preserve the fail-closed invariant: fm-spawn must never record a requested reasoning effort unless the selected harness launches with that exact effort, or refuse before metadata and launch. Codex 0.147.0 must emit model_reasoning_effort=max for gpt-5.6-luna. The PR already implements that behavior; this follow-up repairs the pre-existing Kimi fixture expectation exposed by the invariant. Kimi has no verified effort launch flag, so its former --effort high success expectation is invalid. Keep normal Kimi model launch coverage without an effort, and add a regression proving high effort refuses before metadata, pane creation, launch command, or brief delivery. Update the existing PR #1993 branch only; no force-push, no second PR, no merge. Local changed-scope aggregate is NOT_VERIFIABLE because it does not terminate after Herdr backend coverage on this host; focused Kimi red/green proof is load-bearing.
What Changed
maxeffort through asmodel_reasoning_effort="max", while retaining only verified effort mappings for other harnesses.Risk Assessment
✅ Low: The shared capability predicate now covers remote reused endpoints while the remote-only sentinel normalization preserves ordinary no-effort launches and local invalid efforts refuse early.
Testing
Confirmed the target head, ran the focused Kimi harness suite, and captured a direct CLI-flow artifact showing normal model launch and brief delivery plus high-effort refusal before metadata, pane creation, launch, or brief delivery. The known changed-scope aggregate remains NOT_VERIFIABLE on this host and was not run.
Evidence: Kimi end-to-end CLI flow
Normal Kimi model-only launch exited 0, invokedkimi --model kimi-code/k3 --auto, and delivered the brief. A requested--effort highexited 1 before metadata, launch, brief delivery, or tmux pane creation.Evidence: Focused Kimi test transcript
Focused Kimi harness behavior transcript for target b4ddd853e88894b8d4e297775d35cc203520709a.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed (3) ✅
bin/fm-spawn.sh:1371- The changed local guard at line 1371 is bypassed by the earlier remote-secondmate return path: when a remote Kimi endpoint is already alive,fm-remote-secondmate-control.sh launchreturns its route before it invokes remotefm-spawn; the parent then writeseffort=highatbin/fm-spawn.sh:608. Thusfm-spawn <id> --secondmate --harness kimi --effort highcan record an unlaunched, unsupported effort. This contradicts the required criterion, “fm-spawn must never record a requested reasoning effort unless the selected harness launches with that exact effort, or refuse before metadata and launch.” Put the same capability check at the parent remote-secondmate boundary (before readiness/inheritance/route reuse), using a helper defined before that branch, so a reused endpoint cannot bypass it.🔧 Fix: Fail close remote Kimi effort reuse
1 error still open:
bin/fm-spawn.sh:475- The new shared predicate treats the remote no-effort sentinel-as an unsupported requested effort.spawn_remote_secondmatesetseffort=${EFFORT:--}and calls this helper before contacting the remote host, so a normal remote secondmate spawn with no--effortnow refuses instead of preserving the documented/default launch. Accept-as the remote “no requested effort” sentinel (or normalize it to empty before the shared predicate).🔧 Fix: Normalize remote no-effort sentinel
1 error still open:
bin/fm-spawn.sh:388- The sentinel fix makes--effort -a successful local spawn for every harness: the changed&& [ "$effort" != - ] || return 0path emits no effort flag, while metadata later recordseffort=-. For example, a local Kimi spawn with--effort -now violates the required criterion, “fm-spawn must never record a requested reasoning effort unless the selected harness launches with that exact effort, or refuse before metadata and launch.” Keep-as a remote transport sentinel only—normalize it at the remote caller before invoking the shared predicate—so a user-supplied--effort -remains fail-closed.🔧 Fix: Scope effort sentinel to remote transport
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-kimi-harness.test.shManual fixture-backed Kimi CLI flow captured inkimi-user-flow-transcript.txtgit rev-parse HEADconfirmedb4ddd853e88894b8d4e297775d35cc203520709a✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.