Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
5f23f54
Add verified Kimi crewmate harness adapter
kunchenguid Jul 25, 2026
20394a8
no-mistakes(review): Scope Kimi moon detection to spinner lines
kunchenguid Jul 25, 2026
e73d18a
no-mistakes(review): Match only complete Kimi spinner rows
kunchenguid Jul 25, 2026
53ce82e
no-mistakes(review): Resolve Kimi binary portably before pane creation
kunchenguid Jul 25, 2026
787ad97
no-mistakes(document): Align Kimi adapter documentation
kunchenguid Jul 25, 2026
569e9a6
no-mistakes(lint): Suppress false-positive ShellCheck warning for sou…
kunchenguid Jul 25, 2026
1075f1c
Fix Kimi busy spinner detection
kunchenguid Jul 26, 2026
e69bb02
no-mistakes(review): Recognize Kimi session-lock ancestry and holders
kunchenguid Jul 26, 2026
1d813f3
no-mistakes(review): Scope pending-reply Kimi busy detection by harness
kunchenguid Jul 26, 2026
59dc46a
no-mistakes(document): Correct Kimi spinner capture documentation
kunchenguid Jul 26, 2026
8e06801
no-mistakes(document): Clarify optional Kimi spinner whitespace
kunchenguid Jul 26, 2026
82a9452
no-mistakes(lint): Silence intentional pending-reply test stub warnings
kunchenguid Jul 26, 2026
dbb696f
test: align rebased Kimi busy fixtures
kunchenguid Jul 26, 2026
24a1cf0
no-mistakes: apply CI fixes
kunchenguid Jul 26, 2026
ba7821a
Reconcile Kimi busy detection after per-harness scoping
kunchenguid Jul 26, 2026
2f544ca
no-mistakes(review): Clarify observed Kimi spinner whitespace contract
kunchenguid Jul 26, 2026
05ddb31
no-mistakes(document): Clarify Kimi harness documentation
kunchenguid Jul 26, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/skills/afk/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ The daemon constructs every current injection as the `away-supervisor` kind owne
The bare `FM_INJECT_MARK` form remains accepted for legacy daemon escalations during rollout.
U+2063 has no normal keyboard keystroke and survives terminal transport as UTF-8 text.
This is how firstmate tells a daemon escalation apart from a real message in the same pane.
The operational prefix travels with the message text; it does not rely on harness-level typed-vs-injected detection, which is not portable across claude, codex, opencode, pi, and grok.
The operational prefix travels with the message text; it does not rely on harness-level typed-vs-injected detection, which is not portable across claude, codex, opencode, pi, grok, and kimi.

## Busy-guard and composer guard

Expand Down
2 changes: 1 addition & 1 deletion .agents/skills/firstmate-orca/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ It does not replace `AGENTS.md`, `docs/orca-backend.md`, or `harness-adapters`.

Orca is a runtime backend, not an agent harness.
The runtime backend owns the task endpoint and, for Orca, the task worktree.
The harness is the agent process launched inside that endpoint, such as `claude`, `codex`, `opencode`, `pi`, or `grok`.
The harness is the agent process launched inside that endpoint, such as `claude`, `codex`, `opencode`, `pi`, `grok`, or `kimi`.
Load `harness-adapters` for harness-specific launch, interrupt, resume, trust-dialog, and skill-invocation facts.

Implementation details, metadata fields, teardown guarantees, and limitations live in `docs/orca-backend.md`.
Expand Down
53 changes: 49 additions & 4 deletions .agents/skills/harness-adapters/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: harness-adapters
description: Agent-only reference for firstmate harness operations. Use before spawning or recovering a crewmate or secondmate, handling a trust dialog, sending a harness-specific skill invocation, interrupting or exiting an agent, resuming an exited agent, or verifying a new harness adapter. Contains verified facts for claude, codex, opencode, pi, and grok.
description: Agent-only reference for firstmate harness operations. Use before spawning or recovering a crewmate or secondmate, handling a trust dialog, sending a harness-specific skill invocation, interrupting or exiting an agent, resuming an exited agent, or verifying a new harness adapter. Contains verified facts for claude, codex, opencode, pi, grok, and kimi.
user-invocable: false
metadata:
internal: true
Expand All @@ -25,7 +25,7 @@ If `config/crew-harness` is unset or `default`, there is no concrete value to in
Inheritance also copies the literal `config/crew-dispatch.json` file, so secondmates apply the same best-fit profile rules for their own crewmates.

Each adapter splits into mechanics and knowledge.
The per-task mechanics, including launch command, autonomy flag, and crewmate turn-end hook, live in `bin/fm-spawn.sh`.
The per-task mechanics, including launch command, autonomy flag, and any enabled crewmate turn-end hook, live in `bin/fm-spawn.sh`.
The primary-session "no turn ends blind" guard contract and harness hook installation paths live in `docs/turnend-guard.md`.
The primary-session watcher wake protocols are rendered from `docs/supervision-protocols/` by `bin/fm-supervision-instructions.sh`.
The supervision knowledge lives here: busy signature, exit command, interrupt, dialogs, resume behavior, skill invocation, and quirks.
Expand All @@ -50,16 +50,17 @@ Use that value for interrupt, exit, resume, and skill-invocation facts.

## Primary turn-end guard

Every verified primary harness has an empirically validated hook path for the "no turn ends blind" guard.
The primary integrations for `claude`, `codex`, `opencode`, `pi`, and `grok` have empirically validated hook paths for the "no turn ends blind" guard.
`claude` and `codex` block directly through Stop hooks that preserve exit status 2 and stderr from `bin/fm-turnend-guard.sh`.
`opencode`, `pi`, and `grok` expose passive lifecycle callbacks for this purpose, so their tracked primary adapters force one bounded follow-up or resume when the shared predicate blocks.
Kimi is outside the current turn-end integration scope; `docs/turnend-guard.md` owns the global-configuration boundary.
The exact hook files, commands, scoping rules, and fail-open tradeoffs are owned by `docs/turnend-guard.md`.
`docs/verification/supervision.md` "Turn-end guard" owns active validation evidence.
When changing any primary turn-end hook, validate the real harness behavior in a scratch project or throwaway home before trusting it, then update that doc and the relevant concise fact below.

## Primary pre-arm (PreToolUse) seatbelt

Every verified primary harness also has a wired PreToolUse-equivalent hook that denies a watcher-arm anti-pattern (shell `&`, truncating pipe, bundling, broad `pkill -f fm-watch`) before it runs.
The primary integrations for `claude`, `codex`, `opencode`, `pi`, and `grok` also have wired PreToolUse-equivalent hooks that deny a watcher-arm anti-pattern (shell `&`, truncating pipe, bundling, broad `pkill -f fm-watch`) before it runs.
`claude` and `codex` block directly through PreToolUse hooks; `grok` blocks the same way but requires every `$VAR` reference in its hook `command` string to carry an inline `:-default` or it fails to launch the hook entirely.
`opencode` and `pi` block by throwing from `tool.execute.before` / returning `{block: true}` from `tool_call`.
The exact hook files, commands, output-shaping quirks (Claude Code only honors the deny when stdout is empty), and validation transcripts are owned by `docs/arm-pretool-check.md`.
Expand Down Expand Up @@ -121,6 +122,7 @@ The supported launch-profile flags below are verified locally; each row records
| grok | `--model <model>` | `--reasoning-effort <low\|medium\|high>` | Verified on grok 0.2.99 (2026-07-13). `--effort` is an alias, but firstmate's profile axis is reasoning effort. As of 0.2.99 the ceiling is `high`; both `xhigh` and `max` are rejected with `use one of: high, medium, low`, so firstmate omits them. |
| pi | `--model <model>` | `--thinking <low\|medium\|high\|xhigh\|max>` | Verified 2026-07-13 on Pi 0.80.6. `pi --help` advertises `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, and `max`; `pi --print --model openai-codex/gpt-5.6-sol --thinking max 'Reply with exactly OK.'` completed successfully. |
| opencode | `--model <provider/model>` | none for firstmate's interactive launch | Verified on opencode 1.17.6. `opencode run` has `--variant`, but firstmate launches the interactive `opencode --prompt` path, which has no verified effort flag. |
| kimi | `--model <model>` | none | Verified 2026-07-25 on Kimi Code CLI 0.29.1. |

### Model support discovery

Expand All @@ -134,6 +136,7 @@ Use the discovery surface in the current authenticated environment because suppo
| opencode | Run `opencode models [provider]`, which lists available provider/model identifiers. |
| pi | Run `pi --list-models [search]`; Pi's installed `docs/models.md` owns how built-in, extension-registered, and custom provider/model entries reach that list. |
| grok | Run `grok models`, which lists the models available to the current Grok installation and account. |
| kimi | Run `kimi provider list --json`, which lists the current provider and model configuration. |

For an unfamiliar harness or model namespace, establish support and provider identity from that harness's authoritative CLI help, model listing, or current documentation rather than guessing from a name or prefix.
If those sources do not establish the relationship needed for dispatch, fail loudly and report the unresolved candidate.
Expand All @@ -151,6 +154,13 @@ Natural language is acceptable if uncertain.
- opencode: no separate verified skill invocation beyond normal slash-command behavior; use natural language if the exact skill command is uncertain.
- pi: no separate verified skill invocation beyond normal command behavior; use natural language if the exact skill command is uncertain.
- grok: `/<skill>`, for example `/no-mistakes` (same form as claude). Verified end to end: grok discovers the user-level `no-mistakes` skill, `/no-mistakes` invokes it, and grok drives a real `no-mistakes axi run`. Like codex's `$`/`/` popups, typing `/<skill>` opens grok's slash-autocomplete, so a too-fast Enter selects the popup entry instead of sending, and for an argument-taking command (like `/no-mistakes`'s optional task-first argument) that first Enter only expands the popup selection into an argument-hint placeholder rather than submitting - a genuine second Enter is required (see the grok section below for the 2026-07-03 incident and fix). `fm_tmux_submit_core`'s retried Enter (used by `fm-send` on the tmux backend) already handles this correctly by reading the cursor row; the herdr backend needed a dedicated fix (`fm_backend_herdr_composer_state`, docs/herdr-backend.md) because its prior delta-based verification false-positived on that same popup-close content change.
- kimi: `/<skill>`, for example `/no-mistakes`.

## Submission acknowledgement hazards

A send or key action reporting success is not proof that the intended action happened.
OpenCode can accept and queue an Enter while leaving text visible, Grok can consume Enter in its slash popup without submitting, and Kimi can silently drop a message sent before readiness even though the send returns success.
The shared symptom is a healthy-looking pane with no work in progress, so each adapter must verify the observable postcondition that is specific to its TUI.

## claude (VERIFIED; busy signature re-verified 2026-07-25 on Claude Code 2.1.220)

Expand Down Expand Up @@ -334,3 +344,38 @@ The adapter therefore runs the shared predicate and, when it returns 2, forces o
It does not pass `--permission-mode`, so the passive hook cannot escalate the primary session's tool permissions.
Project-local Grok hooks require folder trust, verified with launch-time `--trust`; if the primary firstmate checkout is not trusted for Grok hooks, this primary guard fails open and `fm-guard.sh` remains the next-command alarm.
Grok's primary watcher protocol is Claude-shaped background-notify around `bin/fm-watch-arm.sh`; the passive Stop hook is only a backstop for blind turn ends.

## kimi (VERIFIED 2026-07-25, kimi 0.29.1)

Kimi Code CLI launches from the absolute path resolved from `PATH`, falling back to the executable `$HOME/.kimi-code/bin/kimi`.

| Fact | Value |
|---|---|
| Binary | Executable `kimi` from `PATH`, then executable `$HOME/.kimi-code/bin/kimi`; spawning refuses if neither exists. |
| Launch | Bare interactive TUI with `--auto`, followed by readiness-gated pointer delivery; positional prompts are rejected. |
| Models | `kimi-code/kimi-for-coding` (default), `kimi-code/kimi-for-coding-highspeed`, `kimi-code/k3`, and `kimi-code/k3-256k`. |
| Busy-pane signature | A transient line with optional leading whitespace, a rotating moon-phase glyph, optional whitespace around `Β·`, and optional trailing content; the line is absent when idle. |
| Exit command | `/exit` |
| Interrupt | Single Escape, which prints `Interrupted by user`. |
| Skill invocation | `/<skill>`, for example `/no-mistakes`; firstmate skills are discovered. |
| Autonomy | `--auto`; `-y` and `--yolo` are weaker and are not used. |
| Trust dialog | None on a clean first launch in a fresh pooled worktree. |
| Slash submission | One Enter submits, with no popup swallow or settle hazard. |
| Environment marker | None; detection relies on process ancestry command name `kimi`. |
| Composer | Bordered box with a bare `>` prompt glyph and no observed ghost or placeholder text. |
| Effort | No reasoning-effort flag exists, so requested effort is recorded in task metadata but omitted from launch. |

`fm-spawn.sh` launches Kimi bare, waits for the composer box or `Welcome to Kimi Code!`, sends only `Read the brief at <absolute-path> and follow it exactly.`, and requires a cleared composer plus either the echoed `✨` submission or nonzero context before accepting delivery.
This launch-then-send shape is mandatory because Kimi rejects a positional brief as an unknown command.
Sending before readiness was reproduced as a silent drop with a zero exit status, an empty composer, `context: 0%`, no echoed user message, and a healthy-looking idle pane.
The brief path must be absolute because the brief lives outside the task worktree, and Kimi reads it there without `--add-dir`.

Observed live spinner captures included optional leading whitespace, a moon-phase glyph, whitespace around `Β·`, and rotating tip text, with the same shape observed during tool execution.
Because those prose examples illustrate the spinner shape rather than define exact bytes, the matcher permits zero whitespace around `Β·` and does not require trailing tip text.
Kimi's footer tip rotates independently and can display `ctrl+c: cancel` while completely idle, so tip text is never used as its busy signature without the leading moon-plus-middot spinner structure.
The idle status bar can contain lowercase `thinking`, which is the model's effort label rather than a busy signal.
The spinner match covers the full moon-phase glyph set rather than one frame, but it remains locale- and emoji-font-sensitive because Kimi exposes no stable ASCII busy token.

[`docs/turnend-guard.md`](../../../docs/turnend-guard.md) owns Kimi's verified global hook surface, approval boundary, and absence from the enabled integrations.
The current adapter falls back to idle detection.
That fallback is the weakest idle detection of any supported adapter because Kimi has no stable ASCII busy token, so turn completion can only be inferred from the fragile moon spinner disappearing and the pane becoming stable.
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,7 +157,7 @@ A silent bootstrap section needs no action; for any printed actionable diagnosti
## 4. Harness and runtime dispatch

Load `harness-adapters` before every spawn or recovery and before trust handling, skill invocation, interrupt, exit, resume, or adapter verification.
The verified harnesses are `claude`, `codex`, `opencode`, `pi`, and `grok`; never dispatch on an unverified adapter.
The verified harnesses are `claude`, `codex`, `opencode`, `pi`, `grok`, and `kimi`; never dispatch on an unverified adapter.
If static `config/crew-harness` or `config/secondmate-harness` names an unverified adapter, report it and fall back only to a verified adapter rather than launching it.

`docs/configuration.md` owns dispatch-profile and runtime-backend schemas, `bin/fm-harness.sh` owns static resolution, and `bin/fm-spawn.sh` owns launch flags and fail-closed validation.
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ Full detail on every feature lives in [docs/architecture.md](docs/architecture.m

### Requirements

- A verified agent harness: Claude Code, Grok, Pi, Codex, or OpenCode.
- A verified primary agent harness: Claude Code, Grok, Pi, Codex, or OpenCode.
- Git and the GitHub CLI, authenticated through `gh auth login`.
- The CLI and dependencies for your selected runtime backend; tmux is the reference default.

Expand Down
2 changes: 1 addition & 1 deletion bin/backends/tmux.sh
Original file line number Diff line number Diff line change
Expand Up @@ -181,7 +181,7 @@ fm_backend_tmux_agent_state() { # <target>
}
comm=${comm#-}
case "$comm" in
*claude*|*codex*|*opencode*|*grok*) printf 'alive' ;;
*claude*|*codex*|*opencode*|*grok*|*kimi*) printf 'alive' ;;
zsh|bash|sh|dash|ash|ksh|mksh|tcsh|csh|fish) printf 'dead' ;;
'') printf 'unreadable' ;;
*) printf 'ambiguous' ;;
Expand Down
6 changes: 3 additions & 3 deletions bin/fm-bootstrap.sh
Original file line number Diff line number Diff line change
Expand Up @@ -436,7 +436,7 @@ secondmate_liveness_sweep() {
[ -n "$target" ] || target="$window"
agent_state=$(fm_backend_agent_state "$backend" "$target" 2>/dev/null) || agent_state=unreadable
case "$harness" in
claude|codex|opencode|pi|grok) ;;
claude|codex|opencode|pi|grok|kimi) ;;
*)
case "$agent_state" in dead|missing) agent_state=unverified-harness ;; esac
;;
Expand Down Expand Up @@ -713,15 +713,15 @@ crew_dispatch_validate() {
return 0
fi
err=$(jq -r '
def verified($h): ["claude","codex","opencode","pi","grok"] | index($h);
def verified($h): ["claude","codex","opencode","pi","grok","kimi"] | index($h);
def effort_ok($h; $e):
if $e == null then true
elif ($e | type) != "string" then false
elif $h == "claude" then (["low","medium","high","xhigh","max"] | index($e))
elif $h == "codex" then (["low","medium","high","xhigh"] | index($e))
elif $h == "grok" then (["low","medium","high"] | index($e))
elif $h == "pi" then (["low","medium","high","xhigh","max"] | index($e))
elif $h == "opencode" then false
elif $h == "opencode" or $h == "kimi" then false
else true
end;
def profiles($value):
Expand Down
9 changes: 8 additions & 1 deletion bin/fm-harness.sh
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
#!/usr/bin/env bash
# Detect the agent harness this process tree runs on.
# Usage: fm-harness.sh print own harness: claude|codex|opencode|pi|grok|unknown
# Usage: fm-harness.sh print own harness: claude|codex|opencode|pi|grok|kimi|unknown
# fm-harness.sh crew print the effective CREWMATE harness
# (config/crew-harness; "default" resolves to own)
# fm-harness.sh secondmate print the harness the PRIMARY uses to launch
Expand Down Expand Up @@ -29,6 +29,12 @@ CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}"

detect_own() {
# Layer 1: environment markers for verified harnesses.
# Keep marker detection before ancestry detection as an explicit precedence rule.
# Only claude, pi, and grok set verified markers of their own; codex, opencode,
# and kimi are markerless, so a foreign marker retained in a terminal
# multiplexer's stored environment can silently misidentify one of them before
# ancestry is consulted. This is a precedence hazard, not evidence that
# CLAUDECODE inheritance into a kimi child was observed; it was not observed.
[ "${CLAUDECODE:-}" = "1" ] && { echo claude; return; }
[ "${PI_CODING_AGENT:-}" = "true" ] && { echo pi; return; }
# grok sets GROK_AGENT=1 for its child/tool processes (verified, grok 0.2.73).
Expand All @@ -44,6 +50,7 @@ detect_own() {
*codex*) echo codex; return ;;
*opencode*) echo opencode; return ;;
*grok*) echo grok; return ;;
kimi) echo kimi; return ;;
pi) echo pi; return ;;
node*|python*)
# Bare interpreter: match the harness name in its script path.
Expand Down
Loading
Loading