Skip to content

fix(bin): switch workers to declared model-matrix stand-ins when pooled OMP Codex quota is exhausted - #91

Merged
MrGTV-love merged 22 commits into
mainfrom
fm/fm-omp-codex-account-failover-v2-r3
Oct 10, 2026
Merged

MrGTV-love merged 22 commits into
mainfrom
fm/fm-omp-codex-account-failover-v2-r3

Conversation

@MrGTV-love

@MrGTV-love MrGTV-love commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

Intent

if codex is exhausted, why not switch to a different model like Deepseek flash v4, or sonnet 5.5 via tc run?

vs hold up work??!! we have a model matrix, so why not switch models accordingly to the fallback??

also confused. we have two codex credentials. the other one is completely unutilized

why did you hold up the work vs switching models?

okay. paid calls for development work do not. update so that deepseek flash replaces glm 5.3

deepseek/deepseek-v4-flash is what I think it is

Context: on 2026-10-05 the fleet sat about 13 hours paused because quota-axi reported the Codex weekly window exhausted. quota-axi reads only one Codex account (greentree promax). omp pools two openai-codex accounts (greentree promax and gmail prolite; omp usage lists both, and the gmail account also shows a banked saved reset). Workers and every no-mistakes review agent run on omp openai-codex models, so all work and all PR pipelines stopped together. config/crew-dispatch.json (the model matrix) already names stand-ins per rule: z-ai GLM 5.3 for Sonnet-class work, GLM 5.3 Flash for Luna-class work, and gpt-6.1-sol and Opus 5.5 (Claude only through the TeamClaude proxy) as equals for the strongest class. None were applied. Earlier, on 2026-09-30, a worker stalled about 9 hours on usage_limit_reached on one account while the other account had quota; a relaunch recovered it.

What Changed

  • Added bin/fm-dispatch-capacity-lib.sh and bin/fm-dispatch-capacity.sh. They read omp usage --provider openai-codex --json and report each pooled Codex account as usable, exhausted, or unknown, without account identities or credentials. fm-dispatch-resolve.sh and fm-quota-choose.sh now use this pool verdict for omp openai-codex/* candidates instead of quota-axi's single-account Codex row. Saved resets are disclosed but never spent or counted, and the pool gets no synthesized spendPriority.
  • Added optional rule fallback and top-level default_fallback lists of explicit stand-ins to the dispatch config, validated in fm-dispatch-resolve.sh. Only proven primary exhaustion activates a stand-in. fm-spawn.sh and fm-control.sh relaunch share the same lists, with new --dispatch-rule flag. The chosen rule id is recorded with the task, follows recovery, and the launch and replacement route is appended to task status. The .omp/fm-worker-overlay.yml turns on native credential rotation and usage-aware fallback. Per-task overlays carry the matched rule's model chain. A Claude stand-in requires the TeamClaude launcher.
  • Replaced the openrouter/z-ai/glm-5.3 stand-in with deepseek/deepseek-v4-flash in docs/configuration.md and docs/examples/model-index.json. Updated the dispatch and harness docs and skills. Added tests/fm-dispatch-capacity.test.sh, and extended the dispatch-resolve, quota-choose, control-relaunch, omp-harness, crew-state and session-launch-policy tests.

Risk Assessment

⚠️ Medium: I found no new defect after tracing the capacity, fallback, spawn, relaunch, session-end, status-notice and typed-resolution paths, but the change is large (24 files) and its live behavior (OMP usage timestamps, native fallback chains, auto-relaunch loops) was only checked against fixtures.

Testing

I made a disposable lab home with bin/fm-lab-home.sh. It had a private tmux socket, a lab treehouse root and a lab matrix. I ran real fm-spawn, fm-control, fm-teardown, fm-crew-state, the session-end recovery scan, fm-dispatch-capacity and fm-quota-choose against the real omp 18.8.7, using the machine's existing logins. The real Codex pool was exhausted, so every switch was driven by real evidence. For the native-fallback and recovery cases, a lab PATH wrapper made omp usage unreadable at spawn time. That kept the worker on Codex, and the real Codex error then fired. One step was staged: placing the real extension note line after each worker state line, because omp kept the fallback model for the rest of the session. The lab tmux server, lab home and /tmp/fm-labr3* task temp folders were all removed. The worktree is clean. I also ran six focused test files in the background through bin/fm-test-run.sh. fm-crew-state, fm-dispatch-capacity and fm-omp-harness passed. The runner was then stopped by a signal (exit 144) at the start of fm-teardown, after the report was sent, so fm-teardown, fm-dispatch-resolve and fm-quota-choose never finished. No test reported a failure. Remote CI owns those files. No screenshots were taken because there is no UI change. Evidence is CLI transcripts and pane captures.

  • Live validation: ✅ go - 12 of 13 scenarios driven live against the product
Scenario Result Live Evidence
Operator checks Codex capacity: fm-dispatch-capacity reads both pooled omp accounts, not quota-axi's single account ✅ pass live r3-S2-spawn-switches-to-standin.txt (header), r3-S10-quota-choose-reads-pool.txt
Unreadable omp usage is reported as unknown capacity, never as exhausted ✅ pass live r3-S1b-unreadable-usage-is-unknown.txt
Spawning a Luna-class worker while the Codex pool is exhausted launches on the rule's declared stand-in, records dispatch_rule, and reports the route ✅ pass live r3-S2-spawn-switches-to-standin.txt
An invalid fallback in an unrelated rule does not block a spawn, but an invalid fallback in the matched rule refuses before any task record exists ✅ pass live r3-S3-invalid-fallback-scoped-to-matched-rule.txt
Strongest-class work with no declared equal stand-in is refused, not moved to a weaker model ✅ pass live r3-S4-strongest-no-weak-standin.txt
A running Codex worker whose request hits usage_limit_reached is served by omp's native retry from only the declared chain ✅ pass live r3-S5-native-inline-fallback.txt
The fallback notice never hides the worker's done, needs-decision, blocked, failed or paused state ✅ pass live r3-S6-native-fallback-note-keeps-state.txt
Relaunching with the exhausted Codex profile moves to the stand-in, and a plain relaunch stays on the stand-in (no return to primary) ✅ pass live r3-S7-relaunch-switch-and-sticky.txt
A worker idle on a real usage_limit_reached is relaunched automatically onto the stand-in by one watcher tick, with uncommitted work kept ✅ pass live r3-S8-quota-exhausted-auto-recovery.txt
Tearing down an omp task removes its per-task fallback overlay together with the base's omp files ✅ pass live r3-S9-teardown-removes-fallback-overlay.txt
fm-quota-choose skips an exhausted OMP Codex pool and picks the next ordered candidate ✅ pass live r3-S10-quota-choose-reads-pool.txt
DeepSeek v4 flash replaces GLM as the documented Sonnet-class stand-in ✅ pass live r3-S11-deepseek-replaces-glm.txt
Typed dispatch resolution emits the stand-in launch flags when the matched rule's pool is exhausted ⏸️ untested no fm-dispatch-resolve.sh needs a TypeSafe API key to call api.typesafe.ai. This host has none (spawn reported 'no TypeSafe or OpenRouter key'). Signing in or adding a credential is not allowed here. The…
Evidence: Spawn on exhausted real Codex pool switches to declared stand-in

Source: Spawn on exhausted real Codex pool switches to declared stand-in

# commit under test: c6eb3237 (rebased on d69b2c5a)
$ bin/fm-dispatch-capacity.sh --harness omp --model openai-codex/gpt-6-luna   # real omp usage, both pooled accounts
capacity: exhausted
accounts[2]{status,remaining,savedResets}:
  exhausted,0,1
  exhausted,0,2
# lab matrix (config/crew-dispatch.json)
{"rules":[
 {"when":"lab luna-class work","use":{"harness":"omp","model":"openai-codex/gpt-6-luna","effort":"medium"},
  "fallback":[{"harness":"omp","model":"google-antigravity/gemini-3.7-flash","effort":"low"}]},
 {"when":"unrelated strongest work","use":{"harness":"omp","model":"openai-codex/gpt-6.1-sol","effort":"high"},
  "fallback":[{"harness":"codex","model":"gpt-6.1-sol","effort":"ultra"}]}
]}
$ fm-spawn.sh labr3a <lab>/projects/notes --mode local-only --yolo on --harness omp --model openai-codex/gpt-6-luna --effort medium
spawned labr3a harness=omp kind=ship mode=local-only yolo=on window=firstmate:fm-labr3a  (exit 0)
# state/labr3a.meta routing keys
harness=omp
model=google-antigravity/gemini-3.7-flash
effort=low
dispatch_rule=rule_1
# state/labr3a.status
working [at=1791654816]: model-matrix fallback launched omp google-antigravity/gemini-3.7-flash for rule_1
# state/labr3a.omp-fallback.yml (per-task overlay)
{
  "retry": {
    "modelFallback": true,
    "fallbackChains": {
      "default": [],
      "google-antigravity/gemini-3.7-flash": [],
      "google-antigravity/gemini-3.7-flash:low": []
    }
  }
}

# omp usage calls made by spawn (lab wrapper log; "real" = passed to real omp)
1791654797 real
1791654808 real
1791654853 real
# worker pane (lab tmux) - real worker ran on the stand-in and did the work
│   1 lab probe ok                                                             │
│   2                                                                          │
╰──────────────────────────────────────────────────────────────────────────────╯
╭──────────────────────────────────────────────────────────────────────────────╮
│ $ git add PROBE.txt && git commit -m "Add PROBE.txt with lab probe ok"       │
├─── Output ───────────────────────────────────────────────────────────────────┤
│ [detached HEAD 2f40da6] Add PROBE.txt with lab probe ok                      │
│  1 file changed, 1 insertion(+)                                              │
│  create mode 100644 PROBE.txt                                                │
│ ⟦Wall: 0.28s | Timeout: 300s⟧                                                │
╰──────────────────────────────────────────────────────────────────────────────╯
 FIRSTMATE_STATUS: done
❯
 π · ◔ Gemini 3.7 Flash 🙈 · 🌳 notes/notes · ⑂ detached · ◫ 1.0%/1M ⟲ · S0.03
Evidence: Unreadable omp usage reads as unknown; quota-axi sees one Codex account

Source: Unreadable omp usage reads as unknown; quota-axi sees one Codex account

$ omp usage made unreadable in the lab PATH wrapper; then:
$ bin/fm-dispatch-capacity.sh --harness omp --model openai-codex/gpt-6-luna
capacity: unknown
accounts[0]{status,remaining,savedResets}:
reason: omp usage failed or returned an invalid report

# contrast: quota-axi codex view (single account, no pool)
[{"provider":"codex","accounts":0}]
Evidence: Invalid fallback blocks only its own matched rule

Source: Invalid fallback blocks only its own matched rule

# rule_2 (unrelated to labr3a) has an INVALID fallback (codex harness, effort ultra); labr3a spawn above still succeeded.
$ fm-spawn.sh labr3b ... --harness omp --model openai-codex/gpt-6.1-sol --effort high   # matches rule_2 itself
fm-gate-refuse: gate agent lifecycle permitted only against lab home ~/tmp/fm-lab.74u3uJ/home
error: invalid dispatch fallback configuration: jq: error (at <stdin>:4): fallback must be an array of explicit OMP profiles or TeamClaude-required Claude profiles
exit=1
(no labr3b state files - refused before publication)
Evidence: Strongest class with no stand-in refuses instead of using a weak model

Source: Strongest class with no stand-in refuses instead of using a weak model

# rule_2 (strongest class) declares no stand-in: fallback []
{"when":"unrelated strongest work","use":{"harness":"omp","model":"openai-codex/gpt-6.1-sol","effort":"high"},"fallback":[]}
$ fm-spawn.sh labr3b ... --harness omp --model openai-codex/gpt-6.1-sol --effort high   # real pool exhausted
error: dispatch profile has exhausted capacity and no supported permitted fallback
(no labr3b state files - no weaker model was substituted)
Evidence: Native inline fallback served by omp from the per-task chain

Source: Native inline fallback served by omp from the per-task chain

# Codex pool capacity unreadable at spawn time (lab PATH wrapper) -> capacity unknown -> spawn keeps the Codex route and hands omp the rule chain
$ fm-spawn.sh labr3c <lab>/projects/notes --mode local-only --yolo on --harness omp --model openai-codex/gpt-6-luna --effort medium
spawned labr3c harness=omp kind=ship mode=local-only yolo=on window=firstmate:fm-labr3c worktree=~/tmp/fm-lab.74u3uJ/treehouse/.treehouse/notes-4486e5/2/notes
exit=0
# meta
harness=omp
model=openai-codex/gpt-6-luna
effort=medium
dispatch_rule=rule_1
# per-task overlay given to omp
{
  "retry": {
    "modelFallback": true,
    "fallbackChains": {
      "default": [],
      "openai-codex/gpt-6-luna": [
        "google-antigravity/gemini-3.7-flash:low"
      ],
      "openai-codex/gpt-6-luna:medium": [
        "google-antigravity/gemini-3.7-flash:low"
      ],
      "google-antigravity/gemini-3.7-flash": [],
      "google-antigravity/gemini-3.7-flash:low": []
    }
  }
}

# real Codex request hit usage_limit_reached; omp served the declared stand-in. Worker pane:
 Warning: Fallback: openai-codex/gpt-6-luna:medium ->
 "native fallback ok". Commit it on your task branch.
 Fallback succeeded on google-antigravity/gemini-3.7-flash:low
│   1 native fallback ok                                                       │
 Warning: Fallback: openai-codex/gpt-6.1-sol:xhigh ->
 Advisor request failed: Codex error event: The usage limit has been reached
 (code=usage_limit_reached)
│ $ git add NATIVE.txt && git commit -m "Add NATIVE.txt with native fallback   │
│ [detached HEAD c2e2df5] Add NATIVE.txt with native fallback ok               │
 Warning: Fallback: openai-codex/gpt-6-luna:medium ->
 Advisor request failed: Codex error event: The usage limit has been reached
 (code=usage_limit_reached)
# status file lines written by the generated extension
note [at=1791654915]: model-matrix fallback served google-antigravity/gemini-3.7-flash:low
note [at=1791654919]: model-matrix fallback served google-antigravity/gemini-3.7-flash:low
# worker commit on the stand-in
c2e2df5 Add NATIVE.txt with native fallback ok
Evidence: Fallback note never replaces the worker's state line

Source: Fallback note never replaces the worker's state line

# real status file of labr3c, written by the worker and the generated omp extension:
note [at=1791654915]: model-matrix fallback served google-antigravity/gemini-3.7-flash:low
note [at=1791654919]: model-matrix fallback served google-antigravity/gemini-3.7-flash:low
done: native fallback probe ready
$ fm-crew-state.sh labr3c
state: blocked · source: status-log · named head c2e2df50676850ec2a04187374ed1a426c8ff9b3 is unreachable outside the worker copy · model-drift: google-antigravity/gemini-3.7-flash live (recorded openai-codex/gpt-6-luna)

# staged order: each terminal verb, then the REAL extension note line copied from above
done then fallback note -> state: blocked · source: status-log · named head c2e2df50676850ec2a04187374ed1a426c8ff9b3 is unreachable outside the worker copy · model-drift: google-antigravity/gemini-3.7-flash live (recorded openai-codex/gpt-6-luna)
needs-decision then fallback note -> state: parked · source: status-log · lab needs-decision probe · model-drift: google-antigravity/gemini-3.7-flash live (recorded openai-codex/gpt-6-luna)
blocked then fallback note -> state: blocked · source: status-log · lab blocked probe · model-drift: google-antigravity/gemini-3.7-flash live (recorded openai-codex/gpt-6-luna)
failed then fallback note -> state: failed · source: status-log · lab failed probe · model-drift: google-antigravity/gemini-3.7-flash live (recorded openai-codex/gpt-6-luna)
paused then fallback note -> state: paused · source: status-log · lab paused probe · model-drift: google-antigravity/gemini-3.7-flash live (recorded openai-codex/gpt-6-luna)
control: worker own free-form note after done -> state: unknown · source: none · no current-state source available · model-drift: google-antigravity/gemini-3.7-flash live (recorded openai-codex/gpt-6-luna)
# live-model published by the extension
model=google-antigravity/gemini-3.7-flash

# shared reader (last_status_line in bin/fm-status-event-lib.sh) on the same staged files:
done then fallback note -> last_status_line: done [at=1]: lab done probe
needs-decision then fallback note -> last_status_line: needs-decision [at=1]: lab needs-decision probe
blocked then fallback note -> last_status_line: blocked [at=1]: lab blocked probe
failed then fallback note -> last_status_line: failed [at=1]: lab failed probe
paused then fallback note -> last_status_line: paused [at=1]: lab paused probe
# note: "done -> blocked" in fm-crew-state above is its separate done-proof check (the weak stand-in worker committed on a detached HEAD); the real file with no trailing note gives the same result, and last_status_line returns the done line.
Evidence: Relaunch switches to stand-in and stays there

Source: Relaunch switches to stand-in and stays there

# labr3a was running on the stand-in; operator asks for the Codex route again while the real pool is exhausted
$ fm-control.sh labr3a relaunch --harness omp --model openai-codex/gpt-6-luna --effort medium --note ...
relaunched labr3a harness=omp from=omp model=google-antigravity/gemini-3.7-flash effort=low backend=tmux endpoint=firstmate:fm-labr3a worktree=~/tmp/fm-lab.74u3uJ/treehouse/.treehouse/notes-4486e5/1/notes
exit=0
# status after relaunch
working [at=1791654816]: model-matrix fallback launched omp google-antigravity/gemini-3.7-flash for rule_1
working [at=1791654982]: model-matrix fallback relaunched omp google-antigravity/gemini-3.7-flash for rule_1

# sticky: plain relaunch with no profile flags (no return-to-primary machinery)
$ fm-control.sh labr3a relaunch --note "Sticky probe: end your turn."
relaunched labr3a harness=omp from=omp model=google-antigravity/gemini-3.7-flash effort=low backend=tmux endpoint=firstmate:fm-labr3a worktree=~/tmp/fm-lab.74u3uJ/treehouse/.treehouse/notes-4486e5/1/notes
exit=0
# meta after sticky relaunch
harness=omp
model=google-antigravity/gemini-3.7-flash
effort=low
dispatch_rule=rule_1
# omp usage calls during sticky relaunch: 0 (no capacity probe of the Codex primary)
# overlay after relaunch
{
  "retry": {
    "modelFallback": true,
    "fallbackChains": {
      "default": [],
      "google-antigravity/gemini-3.7-flash": [],
      "google-antigravity/gemini-3.7-flash:low": []
    }
  }
}
Evidence: Real usage_limit_reached recovers automatically onto stand-in, work kept

Source: Real usage_limit_reached recovers automatically onto stand-in, work kept

# step 1: rule_1 had NO fallback; Codex capacity unreadable at spawn -> worker launched on openai-codex/gpt-6-luna
spawned labr3d harness=omp kind=ship mode=local-only yolo=on window=firstmate:fm-labr3d worktree=~/tmp/fm-lab.74u3uJ/treehouse/.treehouse/notes-4486e5/3/notes
exit=0
# step 2: real Codex error in the worker pane
────────────────────────────────────────────────────────────────────────────────
 Error: Codex error event: The usage limit has been reached
   (code=usage_limit_reached)
 Error: Retry failed after 1 attempts: Provider requested 293515552ms wait,
 exceeds retry.maxDelayMs (300000ms). Original error: Codex error event: The
 usage limit has been reached (code=usage_limit_reached)
❯
 π · ◑ GPT-6-Luna 👁 · 🌳 notes/notes · ⑂ detached · ◫ 3.4%/272K ⟲ · (sub)
# busy record written by the generated omp extension
v1 gen=g1791655014.25490.31343 seq=3 state=idle source=omp-ext event=quota-exhausted ts=1791655018
# step 3: an uncommitted file was placed in the worker copy; operator restored the rule_1 stand-in; one watcher tick ran fm_session_end_relaunch_scan
# (tick run from `git archive HEAD` of c6eb3237 under the lab: the scan sets FM_STATE_OVERRIDE, which the gate-worktree guard refuses by design)
scan_rc=0
wake: check: labr3d auto-relaunched after quota exhaustion
# meta
harness=omp
model=google-antigravity/gemini-3.7-flash
effort=low
dispatch_rule=rule_1
# status
working [at=1791655048]: model-matrix fallback relaunched omp google-antigravity/gemini-3.7-flash for rule_1
# uncommitted work preserved
uncommitted local edit kept
# step 4: recovered worker on the stand-in
│ ⟦Wall: 0.33s | Timeout: 300s⟧                                                │
╰──────────────────────────────────────────────────────────────────────────────╯
 FIRSTMATE_OP: v1 done: Created RECOVERY.txt in worktree root and committed to
 task branch.
❯
 π · ◔ Gemini 3.7 Flash 🙈 · 🌳 …es/notes · ⑂ detached ?1 · ◫ 1.1%/1M ⟲ · S0.04
79e5ec6 Add RECOVERY.txt probe
2ea22a4 seed
Evidence: Teardown removes the per-task fallback overlay

Source: Teardown removes the per-task fallback overlay

# labr3c state files before teardown
labr3c.busy-gen
labr3c.busy-state
labr3c.git-hooks
labr3c.live-model
labr3c.meta
labr3c.omp-ext.ts
labr3c.omp-fallback.yml
labr3c.status
labr3c.turn-ended
# lab operator landed the worker commit on the lab project main (git merge --ff-only), so teardown needs no discard
$ fm-teardown.sh labr3c
A new version of treehouse is available: v2.3.1-fm.e58947de → v3.1.2
Run "treehouse update" to update

🌳 Worktree returned to pool.
teardown labr3c complete (window firstmate:fm-labr3c, worktree ~/tmp/fm-lab.74u3uJ/treehouse/.treehouse/notes-4486e5/2/notes)
Backlog: labr3c just finished (this home keeps no markdown backlog at ~/tmp/fm-lab.74u3uJ/home/data/backlog.md). Update ~/tmp/fm-lab.74u3uJ/home/data/backlog.md - move labr3c to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due.
# labr3c state files after teardown
labr3c.teardown-processes
Evidence: Quota chooser reads the OMP pool

Source: Quota chooser reads the OMP pool

$ quota-axi --json | bin/fm-quota-choose.sh --candidate omp:openai-codex/gpt-6-luna --candidate omp:openai-codex/gpt-6.1-sol   # real pool: both accounts exhausted
none
exit=1
$ quota-axi --json | bin/fm-quota-choose.sh --candidate omp:openai-codex/gpt-6-luna --candidate claude:claude-sonnet-5-5
claude claude-sonnet-5-5
exit=0
Evidence: DeepSeek flash replaces GLM stand-in in examples

Source: DeepSeek flash replaces GLM stand-in in examples

$ FM_CONFIG_OVERRIDE=<copy of docs/examples> bin/fm-model-index.sh model pi stand-in:sonnet-grade
deepseek/deepseek-v4-flash
exit=0
$ grep -rniE "glm" docs/examples docs/configuration.md
(no GLM stand-in left)
Evidence: Focused test run log (3 files passed, then interrupted)

Source: Focused test run log (3 files passed, then interrupted)

FM_TEST_BEGIN 2026-10-10T17:52:36Z tests/fm-crew-state.test.sh family=pure-contract-unit expected_gate_skip=none
ok - captured AXI replacement status replays through crew-state
ok - captured AXI parked status replays through crew-state
ok - captured AXI failed status replays through crew-state
ok - captured capped inventory replays selection, ambiguity, and unavailable lookup
ok - captured status formats reject a synthetic authority transition
ok - captured completed status yields to synthetic subsequent development
ok - captured cancelled run leaves fleet inventory unverified without a failure contradiction
ok - failed-outcome/github/open: terminal delivery uses current disposition
ok - failed-outcome/github/merged: terminal delivery uses current disposition
ok - failed-outcome/github/closed: terminal delivery uses current disposition
ok - failed-outcome/github/unreadable: terminal delivery uses current disposition
ok - failed-outcome/github/skipped: terminal delivery uses current disposition
ok - failed-outcome/github/no-identity: terminal delivery uses current disposition
ok - failed-outcome/gitlab/open: terminal delivery uses current disposition
ok - failed-outcome/gitlab/merged: terminal delivery uses current disposition
ok - failed-outcome/gitlab/closed: terminal delivery uses current disposition
ok - failed-outcome/gitlab/unreadable: terminal delivery uses current disposition
ok - failed-outcome/gitlab/skipped: terminal delivery uses current disposition
ok - failed-outcome/gitlab/no-identity: terminal delivery uses current disposition
ok - failed-outcome/gerrit/open: terminal delivery uses current disposition
ok - failed-outcome/gerrit/merged: terminal delivery uses current disposition
ok - failed-outcome/gerrit/closed: terminal delivery uses current disposition
ok - failed-outcome/gerrit/unreadable: terminal delivery uses current disposition
ok - failed-outcome/gerrit/skipped: terminal delivery uses current disposition
ok - failed-outcome/gerrit/no-identity: terminal delivery uses current disposition
ok - failed-status/github/open: terminal delivery uses current disposition
ok - failed-status/github/merged: terminal delivery uses current disposition
ok - failed-status/github/closed: terminal delivery uses current disposition
ok - failed-status/github/unreadable: terminal delivery uses current disposition
ok - failed-status/github/skipped: terminal delivery uses current disposition
ok - failed-status/github/no-identity: terminal delivery uses current disposition
ok - failed-status/gitlab/open: terminal delivery uses current disposition
ok - failed-status/gitlab/merged: terminal delivery uses current disposition
ok - failed-status/gitlab/closed: terminal delivery uses current disposition
ok - failed-status/gitlab/unreadable: terminal delivery uses current disposition
ok - failed-status/gitlab/skipped: terminal delivery uses current disposition
ok - failed-status/gitlab/no-identity: terminal delivery uses current disposition
ok - failed-status/gerrit/open: terminal delivery uses current disposition
ok - failed-status/gerrit/merged: terminal delivery uses current disposition
ok - failed-status/gerrit/closed: terminal delivery uses current disposition
ok - failed-status/gerrit/unreadable: terminal delivery uses current disposition
ok - failed-status/gerrit/skipped: terminal delivery uses current disposition
ok - failed-status/gerrit/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/github/open: terminal delivery uses current disposition
ok - cancelled-outcome/github/merged: terminal delivery uses current disposition
ok - cancelled-outcome/github/closed: terminal delivery uses current disposition
ok - cancelled-outcome/github/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/github/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/github/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/open: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/merged: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/closed: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/open: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/merged: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/closed: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/no-identity: terminal delivery uses current disposition
ok - cancelled-status/github/open: terminal delivery uses current disposition
ok - cancelled-status/github/merged: terminal delivery uses current disposition
ok - cancelled-status/github/closed: terminal delivery uses current disposition
ok - cancelled-status/github/unreadable: terminal delivery uses current disposition
ok - cancelled-status/github/skipped: terminal delivery uses current disposition
ok - cancelled-status/github/no-identity: terminal delivery uses current disposition
ok - cancelled-status/gitlab/open: terminal delivery uses current disposition
ok - cancelled-status/gitlab/merged: terminal delivery uses current disposition
ok - cancelled-status/gitlab/closed: terminal delivery uses current disposition
ok - cancelled-status/gitlab/unreadable: terminal delivery uses current disposition
ok - cancelled-status/gitlab/skipped: terminal delivery uses current disposition
ok - cancelled-status/gitlab/no-identity: terminal delivery uses current disposition
ok - cancelled-status/gerrit/open: terminal delivery uses current disposition
ok - cancelled-status/gerrit/merged: terminal delivery uses current disposition
ok - cancelled-status/gerrit/closed: terminal delivery uses current disposition
ok - cancelled-status/gerrit/unreadable: terminal delivery uses current disposition
ok - cancelled-status/gerrit/skipped: terminal delivery uses current disposition
ok - cancelled-status/gerrit/no-identity: terminal delivery uses current disposition
ok - outcome: cancellation without delivery carries no verdict
ok - status: cancellation without delivery carries no verdict
ok - selected: cancellation without delivery carries no verdict
ok - coarse: cancellation without delivery carries no verdict
ok - no-ci-log: cancellation without delivery carries no verdict
ok - red-ci: cancellation without delivery carries no verdict
ok - cancelled-test: cancellation without delivery carries no verdict
ok - skipped-test: cancellation without delivery carries no verdict
ok - synthetic cancelled run leaves fleet inventory unverified without a failure contradiction
ok - cancelled-outcome: terminal delivery reports only observed evidence
ok - cancelled-status: terminal delivery reports only observed evidence
ok - skipped-rebase: terminal delivery reports only observed evidence
ok - cancelled-skipped-rebase: terminal delivery reports only observed evidence
ok - passed: terminal delivery reports only observed evidence
ok - active run-step is authoritative
ok - stale needs-decision over active run is superseded
ok - stale blocked over active run is superseded
ok - daemon/timeout blocked claim over a live fixing run reads as run alive
ok - socket refusal or missing socket over a stale fixing run reports blocked
ok - socket refusal over a terminal attributed run reports blocked
ok - socket-down evidence outranks a live run only while it is the log's latest event
ok - broken-pipe blocker over a live run keeps the plain superseded reading
ok - genuine daemon-down blocked line still reports blocked
ok - a busy secondmate keeps its open blocker until that exact key closes
ok - the most recently opened decision supplies the reported state and detail
ok - ship and scout terminal declarations supersede stale decisions
ok - a model-matri

... [35309 bytes truncated] ...

; pid=92525 passes=18 since=1791657133 label=shard_map engine-0; pid=92525 passes=18 since=1791657133 label=shard_map engine-0; and 13 more
fm-cpu-pass: waiting 183s for 1 CPU pass(es) for fm-test-run tests/fm-quota-choose.test.sh: another request is collecting passes; pool size 18; holders: pid=92525 passes=18 since=1791657133 label=shard_map engine-0; pid=92525 passes=18 since=1791657133 label=shard_map engine-0; pid=92525 passes=18 since=1791657133 label=shard_map engine-0; pid=92525 passes=18 since=1791657133 label=shard_map engine-0; pid=92525 passes=18 since=1791657133 label=shard_map engine-0; and 13 more
fm-cpu-pass: waiting 243s for 1 CPU pass(es) for fm-test-run tests/fm-quota-choose.test.sh: all passes in use; pool size 18; holders: pid=17494 passes=18 since=1791657376 label=shard_map engine-1; pid=17494 passes=18 since=1791657376 label=shard_map engine-1; pid=17494 passes=18 since=1791657376 label=shard_map engine-1; pid=17494 passes=18 since=1791657376 label=shard_map engine-1; pid=17494 passes=18 since=1791657376 label=shard_map engine-1; and 13 more
ok - absent key is off: one stderr line, exit 0, no network call
ok - help and bare-filename startup isolate environment credentials
ok - TYPESAFE_API_KEY= in .env activates the tool; environment and config overrides work
ok - clear: one rule Choice request, key on the fd header only, spendPriority argmax over every candidate
ok - stock Bash resolver separates API, quota, local and total timing, with input-only Jev cost
ok - stock Bash timing uses the host epoch clock at millisecond precision
skip: bash on PATH has no EPOCHREALTIME
ok - missing curl transfer time is reported as unknown
ok - unavailable usage is unknown and a failed quota snapshot retains API cost
ok - quota-error model evidence accepts only safe Jev version ids
ok - successful resolution model evidence accepts only safe Jev version ids
ok - error diagnostics omit sensitive API bodies without blocking intake
ok - never-send list withholds the request on a match or a bad list, and never prints the value
ok - opt-in marked sections protect outgoing bodies and markers never send without the exact opt-in
ok - rules snapshots and shell quoting preserve the profile protocol
ok - no-rule fallback, Agy, Gemini, and documented configurations resolve
ok - ambiguous: confidence below the fixed floor hands the decision back
ok - per-rule confidence floors fall to the most probable runner-up that clears its own floor
ok - only the brief's task sections and scout tag reach the model, with a whole-brief fallback
ok - escalate: a rule declared approval: captain never yields a profile
ok - rule floor: known shortfall falls through while unavailable evidence escalates
ok - declared provider and profile floor evidence are applied in code
ok - nonnumeric spendPriority evidence is never ranked
ok - partial and missing quota evidence remain eligible but unranked
ok - provider-wide and exact quota rows combine into one limiting candidate
ok - default: no rule matched resolves among the default profiles
ok - tie: equal spendPriority never breaks by array order
ok - no rankable candidate: the tool escalates instead of guessing
ok - native Codex binds to codex-home before default, independently of Pi accounts and row order
ok - Pi native adapters bind to codex-home with existing fallbacks and schema 5 compatibility
ok - schema 6: each candidate binds to its account row; schema 5 is unchanged
ok - task-horizon runway, pooled-account, OpenRouter coverage, and minimum-version guards
ok - quota evidence comes from one quota-axi --json read, and its failure is an error outcome
ok - API, transport, and response failures are error outcomes with exit 0
ok - typed OMP dispatch preserves pooled headroom, explicit stand-ins, and uncertainty
ok - configuration errors exit 2 before any network call
ok - typed intake resolves roles before quota ranking and detects concrete duplicates
ok - typed intake checks the chosen id against the pinned worker account's catalog
ok - chosen Codex and Pi entries never use supervisor catalogs and retain authoritative export retirement checks
# all fm-dispatch-resolve tests passed
FM_TEST_END 2026-10-10T18:37:08Z tests/fm-dispatch-resolve.test.sh exit=0 duration_ms=245524 gate_skip=false
fm-cpu-pass: got 1 CPU pass(es) for fm-test-run tests/fm-quota-choose.test.sh after 274s
ok - help renders the complete header only
ok - first positive candidate wins
ok - exhausted provider is skipped
ok - no positive candidate returns none and exit 1
ok - positional candidates work
ok - specific model scope bounds generic quota
ok - OMP Codex uses its pooled accounts rather than the single-account snapshot
ok - omp without a mapped prefix fails closed
ok - default model uses provider-wide quota
ok - fractional positive quota is eligible
ok - unknown harness fails closed
ok - all candidates are validated before selection
ok - malformed provider data fails closed
ok - multiple JSON values fail closed
ok - known-empty quota fails closed
ok - semantics and availability statuses must agree
ok - exhausted runway vetoes unknown headroom
ok - unknown headroom is not positive quota
ok - partial semantics accept mixed availability
ok - stdin snapshot is accepted
ok - empty model candidate fails closed
ok - bare harness maps to default model
ok - default TOON snapshot is accepted
ok - renderer-shaped TOON snapshot is accepted
ok - malformed nonzero TOON envelopes fail closed
ok - zero-row TOON has no positive quota
ok - empty-array TOON has no positive quota
ok - inline attention TOON has no positive quota
ok - TOON attention identities fail closed
ok - truncated zero-row TOON fails closed
ok - malformed zero-row TOON fails closed
ok - zero-row TOON rejects leading garbage
ok - counted TOON rows require every declared field
ok - TOON conversion preserves unknown-headroom exhaustion
ok - trailing empty TOON fields fail closed
ok - quoted TOON scope vetoes dispatch
ok - provider-level unknown quota is not positive
ok - Muse candidate is accepted
ok - Muse uses Meta quota
ok - Agy remains resolver-only
ok - duplicate providers fail closed
ok - empty quota scopes fail closed
ok - whitespace quota identities fail closed
ok - out-of-range quota fails closed
ok - invalid runway status fails closed
ok - missing applicable quota is not positive
ok - any exhausted applicable scope vetoes dispatch
ok - named model quota matches exact identity only
ok - invalid availability status fails closed
ok - native Codex never infers an account from a Pi lane
ok - native Codex binds to codex-home before default, independently of model and row order
ok - the same path still selects from a schema 5 snapshot by provider alone
ok - schema 6 requires accountKey on every row and uniqueness on provider + accountKey
ok - schema 6 TOON with the accountKey column is accepted
ok - helper reuses the captured quota snapshot
ok - provider-table lookup writes nothing to stderr when SIGPIPE is ignored
ok - provider-table lookup reads the whole table before answering
ok - provider-table lookup prints the provider when called directly under set -e
# all fm-quota-choose tests passed
FM_TEST_END 2026-10-10T18:38:36Z tests/fm-quota-choose.test.sh exit=0 duration_ms=334095 gate_skip=false
FM_TEST_SUMMARY total=6 failed=0 skipped_gate=0 duration_ms=2761162
FM_TEST_SUMMARY_FAMILY family=pr-forge count=1 duration_ms=1340667 failed=0
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=3 duration_ms=1122070 failed=0
FM_TEST_SUMMARY_FAMILY family=standalone count=2 duration_ms=579619 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-teardown.test.sh duration_ms=1340667
FM_TEST_SLOWEST rank=2 script=tests/fm-omp-harness.test.sh duration_ms=823926
FM_TEST_SLOWEST rank=3 script=tests/fm-quota-choose.test.sh duration_ms=334095
FM_TEST_SLOWEST rank=4 script=tests/fm-crew-state.test.sh duration_ms=259343
FM_TEST_SLOWEST rank=5 script=tests/fm-dispatch-resolve.test.sh duration_ms=245524
FM_TEST_SLOWEST rank=6 script=tests/fm-dispatch-capacity.test.sh duration_ms=38801
fm-test-run: wrote timing artifact: ~/.no-mistakes/evidence/01M4J4MR2NF51QNFJBKT50MJH6/r3-targeted-tests.json
- Outcome: 🔧 2 issues found ✅ across 3 runs (1h26m38s)

Pipeline

Updates from git push no-mistakes

... (28 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

🔧 **Review** - 8 issues found → auto-fixed (9) ✅

🔧 Fix applied.
2 infos still open:

  • ℹ️ docs/configuration.md:1379 - docs/configuration.md says 'Launch and successful replacement append the selected route to task status.' When typed resolution selects the stand-in (bin/fm-dispatch-resolve.sh:201-215), it emits --harness omp --model deepseek/deepseek-v4-flash --dispatch-rule rule_N. fm-spawn then matches rule_N through its fallback list. fm_dispatch_select measures the stand-in as unknown and keeps it, so DISPATCH_SWITCHED stays false and no status line is written (bin/fm-spawn.sh:6574). Status lines appear only when spawn or control itself switches (bin/fm-control.sh:1448). The smallest correction is the docs: say that a spawn or relaunch that itself switches to a stand-in appends the route, and that a typed-resolution stand-in is disclosed in the resolver output.
  • ℹ️ docs/configuration.md:1522 - Fix round caf064f left this docs sentence behind. The round made spawn and relaunch check the fallback schema only for matched rules (bin/fm-dispatch-capacity-lib.sh:131). It updated the library header (lib:9-12). The owner doc still says only 'a retired model or unconfigured role in an unrelated rule does not block a launch or recovery'. It does not say that an invalid fallback in an unrelated rule (for example a codex stand-in, effort 'ultra', or harness 'pi') also no longer blocks spawn or relaunch. It also does not say that typed resolution still refuses the whole file for that same error (bin/fm-dispatch-resolve.sh:235-237). An operator who reads this can expect typed intake to accept a file that it refuses. Fix: extend the sentence at line 1522 to name the unrelated invalid fallback schema too. Add that typed resolution still validates every rule's fallback list. This is a docs change only.

✅ No issues found.

🔧 **Test** - 2 issues found ✅
  • ℹ️ docs/configuration.md:1458 - On this host, omp models --json has no deepseek/ provider. Only openrouter/deepseek/deepseek-v4-flash and its :free variant exist. A matrix rule whose DeepSeek stand-in is written as deepseek/deepseek-v4-flash will therefore be refused live with 'no supported permitted fallback', and the work is held again (Scenario B). The refusal does not name the cause. Before the matrix is activated, the operator must configure the direct DeepSeek provider in omp or pick a selector that is in the catalog.
  • ℹ️ bin/fm-spawn.sh:5562 - In the live lab-s4 worker, omp's advisor request first fell back 'openai-codex/gpt-6.1-sol:xhigh -> openai-codex/gpt-6-luna:high' and then followed the task chain. The per-task overlay sets no key for the advisor model, so a chain from the captain's own omp config still applied to the advisor. The task's own model followed only the declared chain.
  • Live validation: ✅ go - 9 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Operator inspects pooled Codex capacity: fm-dispatch-capacity reports every omp account (2), shows saved resets (1, 2) without counting them, and reports whole-pool exhaustion ✅ pass live scenario-A-spawn-real-exhausted-pool.txt (first block)
Whole Codex pool exhausted: fm-spawn launches the rule's declared stand-in instead of holding the work, records dispatch_rule, and writes a 'model-matrix fallback launched' status line ✅ pass live scenario-A-spawn-real-exhausted-pool.txt
Adversarial: the declared stand-in is not in omp's catalog, so spawn refuses and publishes nothing ✅ pass live scenario-B-unsupported-standin-refused.txt
Two Codex credentials, one exhausted and one healthy: spawn keeps the assigned Codex model (the unused account is not ignored) and writes a native chain that holds only the declared stand-in ✅ pass live scenario-C-healthy-sibling-keeps-primary.txt
A live omp worker hits Codex usage_limit_reached: its busy record becomes quota-exhausted and the real watcher's recovery scan picks it and calls relaunch with the quota note ✅ pass live scenario-D-quota-exhausted-detected-and-recovery-scan.txt
Recovery relaunch with the real exhausted pool moves the task to its stand-in, keeps uncommitted work, and the new worker finishes the brief ✅ pass live scenario-E-control-relaunch-switches-to-standin.txt
omp's in-session fallback: when a Codex request fails on quota, omp moves to the declared stand-in inside the same session, and task status names the serving model ✅ pass live scenario-F-native-inline-fallback-served.txt
A task already on its stand-in stays there on a later relaunch ✅ pass live scenario-H-standin-sticky-on-relaunch.txt
fm-quota-choose uses pooled omp capacity instead of quota-axi's single exhausted Codex account ✅ pass live scenario-G-quota-choose-reads-pool.txt
Typed dispatch resolution selects the declared fallback only on proven exhaustion and prints --dispatch-rule ⏸️ untested no The prior payload did not establish a live result for this scenario. Live, fm-dispatch-resolve.sh printed only 'dispatch-resolve: off' because this host has no TypeSafe or OpenRouter key. Only the stu…
  • bin/fm-dispatch-capacity.sh --harness omp --model openai-codex/gpt-6-luna against the real omp usage --provider openai-codex --json
  • bin/fm-lab-home.sh create $LAB + bin/fm-lab-home.sh tmux-dir $LAB + tmux -L fm-lab new-session (private lab socket)
  • bin/fm-spawn.sh lab-s1 $LAB/projects/demo --scout --harness omp --model openai-codex/gpt-6-luna --effort high --dispatch-rule rule_1 (real exhausted pool -> stand-in)
  • bin/fm-spawn.sh lab-s2 ... --model openai-codex/gpt-6-terra --dispatch-rule rule_2 (deepseek/deepseek-v4-flash stand-in not in this host's catalog -> refusal)
  • bin/fm-spawn.sh lab-s3 ... with a lab view of one exhausted and one usable account (keeps Codex model; writes the native chain overlay)
  • Real omp worker hit Codex usage_limit_reached; checked state/lab-s3.busy-state event=quota-exhausted
  • Real bin/fm-wake-drain.sh + bin/fm-watch.sh supervise rounds (session-end scan chose lab-s3 and invoked relaunch)
  • bin/fm-control.sh lab-s3 relaunch --note &#39;&lt;quota note&gt;&#39; with the real exhausted pool -> zai/glm-4.5-flash, unlanded.txt kept
  • bin/fm-spawn.sh lab-s4 ... -> real omp native fallback openai-codex/gpt-6-luna:high -&gt; zai/glm-4.5-flash:high with status 'model-matrix fallback served'
  • bin/fm-control.sh lab-s3 relaunch --note ... again (stays on the stand-in)
  • bin/fm-quota-choose.sh --snapshot quota-snap.json omp:openai-codex/gpt-6-luna (real pool, then lab healthy-sibling view) and codex:gpt-6-luna
  • bash tests/fm-dispatch-resolve.test.sh
  • bash tests/fm-dispatch-capacity.test.sh
  • Lab teardown: tmux -L fm-lab kill-server, bin/fm-lab-home.sh teardown, rm -rf $LAB

✅ No issues found.

  • Live validation: ✅ go - 10 of 11 scenarios driven live against the product
Scenario Result Live Evidence
Operator inspects pooled OMP Codex capacity: both accounts are read, saved resets are shown but not counted, and an unreadable report is unknown, not exhausted ✅ pass live A-pooled-capacity.txt: accounts[2] exhausted,0,1 / exhausted,0,2; the quota-axi codex row has 0 accounts and all_models only; the wrapper-induced failure prints capacity: unknown
Spawn a Luna-class omp task while the Codex pool is exhausted: it launches on the matched rule's declared stand-in instead of holding the work, and an invalid fallback in an unrelated rule does not bl… ✅ pass live C-spawn-switch.txt: exit 0; meta model=google-antigravity/gemini-3.5-flash dispatch_rule=rule_1; status working: model-matrix fallback launched ...; overlay has an empty default chain
Relaunch a task back onto exhausted Codex: relaunch switches to the declared stand-in, keeps dispatch_rule, and the stand-in finishes the work ✅ pass live F-relaunch-switch.txt: relaunched labfo1 ... model=google-antigravity/gemini-3.7-flash; status fallback relaunched; then done; commit present
A worker that starts on Codex and hits a real usage_limit_reached is served by omp's native fallback through the task-owned chain; the extension writes a note line and live-model shows the stand-in ✅ pass live E-native-fallback.txt: pane Fallback: openai-codex/gpt-6-luna:medium -&gt; google-antigravity/gemini-3.7-flash:low, Fallback succeeded; two note [at=..]: model-matrix fallback served ... lines; the…
Adversarial: a fallback note written after done, needs-decision, blocked, failed, or paused never changes the state supervision reads ✅ pass live E-native-fallback.txt staged section: fm-crew-state gives done/parked/blocked/failed/paused, and last_status_line returns the verb line every time; a worker's own note is still read as an event. The n…
Adversarial: the matched rule has an invalid fallback (codex harness, effort ultra), so spawn still refuses with the schema error and creates no task ✅ pass live D-matched-invalid-refuses.txt: error: invalid dispatch fallback configuration: ... fallback must be an array of explicit OMP profiles ..., exit=1, no labfo3 state
A live omp worker idles on a real quota error after native rotation; one watcher tick auto-relaunches it on the declared stand-in, and it reaches done ✅ pass live G-quota-exhausted-recovery.txt: busy record event=quota-exhausted; wake check: labfo4 auto-relaunched after quota exhaustion; meta moved to the stand-in; status fallback relaunched, then done,…
Adversarial: the strongest class (gpt-6.1-sol) with no equal stand-in is exhausted, so spawn refuses and does not fall back to a weaker model ✅ pass live H-strongest-no-weak-standin.txt: error: dispatch profile has exhausted capacity and no supported permitted fallback; no task launched
fm-quota-choose reads the OMP pool, not quota-axi's single account: an exhausted pool is never chosen ✅ pass live J-quota-choose.txt: none, exit 1, with both Codex candidates on the really exhausted pool
DeepSeek flash replaces GLM 5.3 as the documented Sonnet-class stand-in ✅ pass live K-deepseek-replaces-glm.txt: fm-model-index.sh resolves stand-in:sonnet-grade to deepseek/deepseek-v4-flash; no GLM left in docs/examples or docs/configuration.md
Typed dispatch resolution picks the declared stand-in when every candidate's pooled capacity is exhausted and emits --dispatch-rule ⏸️ untested no fm-dispatch-resolve.sh needs a TypeSafe System One API key, which fm_typesafe_key reads. This host has none (spawn reported no TypeSafe or OpenRouter key). Without it the tool prints `dispatch-resol…
  • bin/fm-dispatch-capacity.sh --harness omp --model openai-codex/gpt-6-luna (and --json, gpt-6.1-sol), against the real omp usage, compared with the quota-axi codex row
  • the same capacity CLI with omp usage made unreadable by a lab PATH wrapper, so capacity is unknown
  • bin/fm-spawn.sh labfo1 &lt;lab&gt;/projects/notes --mode local-only --yolo on --harness omp --model openai-codex/gpt-6-luna --effort medium in a lab FM_HOME while the Codex pool was really exhausted; an unrelated rule_2 had an invalid fallback
  • bin/fm-control.sh labfo1 relaunch --harness omp --model openai-codex/gpt-6-luna --effort medium --note ...
  • spawn labfo2 with pool usage unknown, so the worker starts on Codex; a real usage_limit_reached led to omp's native fallback through the per-task overlay chain
  • bin/fm-crew-state.sh labfo2 and last_status_line with the real fallback note line placed after each of done, needs-decision, blocked, failed, paused, plus a control with a worker's own note line
  • bin/fm-spawn.sh labfo3 ... --model openai-codex/gpt-6.1-sol --effort high, where the matched rule_2 has an invalid fallback; then again with rule_2 fallback set to []
  • spawn labfo4 on Codex with no fallback chain; real usage_limit_reached gave busy event=quota-exhausted; after adding the matrix fallback, one watcher tick of fm_session_end_relaunch_scan
  • bin/fm-quota-choose.sh --snapshot <quota-axi --json> --candidate omp:openai-codex/gpt-6-luna --candidate omp:openai-codex/gpt-6.1-sol
  • FM_CONFIG_OVERRIDE=<copy of docs/examples> bin/fm-model-index.sh model pi stand-in:sonnet-grade
  • bin/fm-test-run.sh --json ... tests/fm-dispatch-capacity.test.sh tests/fm-dispatch-resolve.test.sh tests/fm-quota-choose.test.sh tests/fm-omp-harness.test.sh tests/fm-crew-state.test.sh tests/fm-session-end-relaunch.test.sh tests/fm-control-relaunch.test.sh (all 7 passed, exit 0)

✅ No issues found.

  • Live validation: ✅ go - 12 of 13 scenarios driven live against the product
Scenario Result Live Evidence
Operator checks Codex capacity: fm-dispatch-capacity reads both pooled omp accounts, not quota-axi's single account ✅ pass live r3-S2-spawn-switches-to-standin.txt (header), r3-S10-quota-choose-reads-pool.txt
Unreadable omp usage is reported as unknown capacity, never as exhausted ✅ pass live r3-S1b-unreadable-usage-is-unknown.txt
Spawning a Luna-class worker while the Codex pool is exhausted launches on the rule's declared stand-in, records dispatch_rule, and reports the route ✅ pass live r3-S2-spawn-switches-to-standin.txt
An invalid fallback in an unrelated rule does not block a spawn, but an invalid fallback in the matched rule refuses before any task record exists ✅ pass live r3-S3-invalid-fallback-scoped-to-matched-rule.txt
Strongest-class work with no declared equal stand-in is refused, not moved to a weaker model ✅ pass live r3-S4-strongest-no-weak-standin.txt
A running Codex worker whose request hits usage_limit_reached is served by omp's native retry from only the declared chain ✅ pass live r3-S5-native-inline-fallback.txt
The fallback notice never hides the worker's done, needs-decision, blocked, failed or paused state ✅ pass live r3-S6-native-fallback-note-keeps-state.txt
Relaunching with the exhausted Codex profile moves to the stand-in, and a plain relaunch stays on the stand-in (no return to primary) ✅ pass live r3-S7-relaunch-switch-and-sticky.txt
A worker idle on a real usage_limit_reached is relaunched automatically onto the stand-in by one watcher tick, with uncommitted work kept ✅ pass live r3-S8-quota-exhausted-auto-recovery.txt
Tearing down an omp task removes its per-task fallback overlay together with the base's omp files ✅ pass live r3-S9-teardown-removes-fallback-overlay.txt
fm-quota-choose skips an exhausted OMP Codex pool and picks the next ordered candidate ✅ pass live r3-S10-quota-choose-reads-pool.txt
DeepSeek v4 flash replaces GLM as the documented Sonnet-class stand-in ✅ pass live r3-S11-deepseek-replaces-glm.txt
Typed dispatch resolution emits the stand-in launch flags when the matched rule's pool is exhausted ⏸️ untested no fm-dispatch-resolve.sh needs a TypeSafe API key to call api.typesafe.ai. This host has none (spawn reported 'no TypeSafe or OpenRouter key'). Signing in or adding a credential is not allowed here. The…
  • bin/fm-dispatch-capacity.sh --harness omp --model openai-codex/gpt-6-luna (real omp usage, 2 pooled accounts)
  • bin/fm-dispatch-capacity.sh with omp usage made unreadable (expect unknown, not exhausted)
  • fm-spawn.sh labr3a &lt;lab&gt;/projects/notes --mode local-only --yolo on --harness omp --model openai-codex/gpt-6-luna --effort medium against real exhausted pool
  • fm-spawn.sh labr3b ... --harness omp --model openai-codex/gpt-6.1-sol --effort high with invalid matched fallback (expect refusal)
  • fm-spawn.sh labr3b ... with strongest rule fallback [] (expect refusal, no weak stand-in)
  • fm-spawn.sh labr3c ... with usage unreadable → native chain overlay → real usage_limit_reached → omp served stand-in → extension note lines
  • fm-crew-state.sh labr3c and last_status_line with the real note line after done/needs-decision/blocked/failed/paused
  • fm-control.sh labr3a relaunch --harness omp --model openai-codex/gpt-6-luna --effort medium (switch to stand-in)
  • fm-control.sh labr3a relaunch with no flags (stand-in sticky, 0 usage probes)
  • fm-spawn.sh labr3d ... with no stand-in → real quota-exhausted busy event → stand-in added → fm_session_end_relaunch_scan tick (from git archive HEAD copy)
  • fm-teardown.sh labr3c after lab landing (fallback overlay removed)
  • quota-axi --json | bin/fm-quota-choose.sh --candidate omp:openai-codex/gpt-6-luna --candidate omp:openai-codex/gpt-6.1-sol and with a claude second candidate
  • FM_CONFIG_OVERRIDE=<docs/examples copy> bin/fm-model-index.sh model pi stand-in:sonnet-grade
  • bin/fm-test-run.sh tests/fm-crew-state.test.sh tests/fm-dispatch-capacity.test.sh tests/fm-omp-harness.test.sh ... (these 3 passed; fm-teardown, fm-dispatch-resolve, fm-quota-choose interrupted by signal before finishing)
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

⚠️ **Lint** - 0 issues

✅ No issues found.

✅ No issues found.

  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

@MrGTV-love
MrGTV-love force-pushed the fm/fm-omp-codex-account-failover-v2-r3 branch from a783be9 to caf064f Compare October 10, 2026 12:47
@MrGTV-love MrGTV-love changed the title fix(bin): fail over to declared model-matrix stand-ins when the pooled OMP Codex quota is exhausted fix(bin): switch workers to declared model-matrix stand-ins when the pooled OMP Codex quota is exhausted Oct 10, 2026
@MrGTV-love

Copy link
Copy Markdown
Owner Author

Follow-ups after the final validation round (head ba232b9)

The Risk Assessment above predates the last fix round.
Both warnings it names are now fixed in this PR.

  • A fallback served by the native chain now appends a note: line that status readers skip, so a worker's done, needs-decision, blocked, failed or paused stays the task state.
  • Fallback-schema validation now covers only the rules that match the launch, so an invalid unrelated rule no longer blocks a spawn or a relaunch. Typed intake still validates every rule's fallback list.

Operator follow-up (host configuration, not a code defect)

  • deepseek-standin-not-in-host-catalog: the direct selector deepseek/deepseek-v4-flash is not always resolvable in the host omp catalog.
    A read-only check on 2026-10-10 found only openrouter/deepseek/... entries, while an earlier check the same day found a direct deepseek provider.
    A rule that names the direct selector is refused with "no supported permitted fallback" when the catalog lacks it.
    Make the direct DeepSeek provider resolvable in omp, or pick a selector that is in the catalog, before activating the matrix.

Known limits and follow-up work (info findings, none changes routing correctness)

  1. spawn-reselects-on-relaunch (bin/fm-spawn.sh): on --relaunch, spawn runs the route selection again after fm-control.sh already chose and stopped the old agent. A pool that runs out between the two reads can switch the route after the stop. Fix: skip selection in spawn when relaunching and keep only the fallback-list lookup.
  2. compact-adviser-stale-after-switch (bin/fm-spawn.sh): the compact-adviser settings are computed for the original harness and are not recomputed when a stand-in switches omp to Claude. Fix: compute them after the dispatch switch.
  3. quota-recovery-note-and-double-read (bin/fm-session-end-relaunch-lib.sh): the quota recovery note always says "permitted matrix fallback", even when control kept the same route. The busy record is read twice. Fix: neutral or route-specific wording, and one read.
  4. ambiguous-match-refusal-hides-cause (bin/fm-dispatch-capacity-lib.sh): when matching rules declare different fallback lists, the refusal says "no supported permitted fallback" and does not name the colliding rules. Fix: report the rule ids in the refusal text. Routing is unchanged.
  5. skill-teamclaude-sentence-misplaced (.agents/skills/quota-array-dispatch/SKILL.md): one sentence about TeamClaude stand-ins sits in the fm-quota-choose.sh helper paragraph, which has no fallback concept. Fix: move it to the stand-in paragraph.
  6. claude-primary-uses-native-row-in-teamclaude-home (bin/fm-dispatch-capacity-lib.sh): in a TeamClaude home, a Claude primary with a declared fallback still reads capacity from the native default-account row. This is the recorded decision for fallback candidates only. Recorded, no action unless the scope is widened.
  7. advisor-uses-ambient-chain (bin/fm-spawn.sh): the per-task overlay sets no key for omp's advisor model, so a chain from the operator's own omp config can still apply to the advisor request. The task's own model follows only the declared chain.
  8. typed-stand-in-launch-not-reported (docs/configuration.md): a stand-in that typed dispatch resolution kept as unknown is launched without a route line in the task status. The docs now say so. A route line for that case is possible follow-up work.
  9. docs-unrelated-rule-scope-omits-schema (docs/configuration.md): the sentence on unrelated rules should also say that an invalid fallback schema in an unrelated rule does not block a launch, and that typed resolution still validates every rule's fallback list.

…unch-policy.test.sh, "replacement did not run omp exactly once". This PR caused it, and it is now fixed. - Rule: the fake `omp` in the test must record `launch:omp` only when a session starts. A read-only query must not count as a launch. - Cause: this PR's capacity check (bin/fm-dispatch-capacity-lib.sh, fm_omp_codex_capacity) runs `omp usage --provider openai-codex --json` on the relaunch path. It runs once from fm-control and once from fm-spawn. The fake `omp` already answered `models --json` as a query. It recorded every other call, including `usage`, as a launch. So the test saw 3 launches, not 1. I proved this by logging each call to the fake: 2 were `usage` calls and 1 was the real launch. - Fix: the fake `omp` now answers `usage` with `{"reports":[]}` and exits without recording a launch. The product reads that report as "unknown", keeps the requested profile, and launches as before. I checked every place this fake is used. The same 4-line change covers them all: the replacement case, the fresh-omp case, and the recovery cases. No product code changed. - Proof: the test failed locally before the fix. It passed on base ca6c277. After the fix it exits 0 with 149 ok lines. `bin/fm-lint.sh` on the test file exits 0. 2. tests/fm-wake-queue.test.sh, "the watcher never blocked inside the drain-ring idle capture: check: rearm-resurface". This PR did not cause it, so I did not change any code for it. - The same failure came before this PR. It hit main commit 8ce400f in CI run 38019754012 (shard 10). It also hit an unrelated branch, fm/fm-ledger-dated-fold-cost, in CI run 38027172418. It is a timing flake that was already on main. - This PR does not change any watcher or wake file. Its one change in a file the watcher loads, bin/fm-status-event-lib.sh, only skips the model-matrix fallback note. - On this macOS host, this test fails at an earlier step on both the base snapshot and the PR head, so I could not run that test case here. That local failure is a host difference, not a result of this PR. Changed file: tests/fm-session-launch-policy.test.sh only
…commits now sit on top of the base. Rule: the branch must keep both sides. The base added omp-session.json, reboot-notice, launch_proof and an omp version check. The PR added omp-fallback.yml and dispatch_rule. Each launch, cleanup and relaunch site must carry both. Conflicts, all fixed by keeping both sides: 1. bin/fm-teardown.sh, child-home cleanup: removes .omp-fallback.yml, .live-model, .omp-session.json and .reboot-notice. 2. bin/fm-teardown.sh, main-task cleanup: removes the same files. 3. bin/fm-control-lib.sh, fm_control_harness_wiring_paths omp: prints omp-ext.ts, omp-session.json and omp-fallback.yml. 4. bin/fm-spawn.sh, preserve_relaunch_meta key list: has both dispatch_rule and launch_proof. 5. tests/fm-omp-harness.test.sh: kept the base's new test, test_spawn_refuses_unsupported_omp_before_launch. 6. tests/fm-omp-harness.test.sh: the wiring-path assertion expects all 3 paths. I set it in commit fa921c6. It conflicted again in 316951d, and I kept the 3-path version. One fix after the rebase. It is in the working tree and not committed yet. - tests/fm-control-relaunch.test.sh: the base added a version check, fm_control_omp_launch_check. It needs `omp --version` to report 18.1.20 or newer. The fake `omp` in the PR's rl-pool test did not answer `--version`. So the auto-relaunch was refused with "installed omp version 'unknown'". I added `--version) printf 'omp/18.1.20\n' ;;` to that fake. The base's own fakes use this same line. No product code changed. Tests (bash). All exit 0 with 0 "not ok": 1. fm-control-relaunch: failed before the fix, at the rl-pool case. Passes after the fix with 191 ok. 2. fm-omp-harness 3. fm-teardown 4. fm-quota-choose 5. fm-dispatch-capacity 6. fm-dispatch-resolve 7. fm-crew-state 8. fm-launch-proof 9. fm-session-launch-policy Lint: bin/fm-lint.sh on the 6 changed files exits 1. It gives one warning, SC2100, on tests/fm-omp-harness.test.sh line 363: `id=omp-fallback-q2`. That line comes from the PR's own commit d79d255 and was there before the rebase. The rebase did not cause it. I did not change it, because the rules forbid unrelated edits. If lint must pass, the smallest fix is to quote it: `id='omp-fallback-q2'`
@MrGTV-love
MrGTV-love force-pushed the fm/fm-omp-codex-account-failover-v2-r3 branch from ba232b9 to 7e8628b Compare October 10, 2026 19:35
@MrGTV-love MrGTV-love changed the title fix(bin): switch workers to declared model-matrix stand-ins when the pooled OMP Codex quota is exhausted fix(bin): switch workers to declared model-matrix stand-ins when pooled OMP Codex quota is exhausted Oct 10, 2026
@MrGTV-love
MrGTV-love merged commit b2e06c4 into main Oct 10, 2026
43 of 44 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant