Repository navigation
fix(bin): switch workers to declared model-matrix stand-ins when pooled OMP Codex quota is exhausted - #91
Merged
Conversation
MrGTV-love
force-pushed
the
fm/fm-omp-codex-account-failover-v2-r3
branch
from
October 10, 2026 12:47
a783be9 to
caf064f
Compare
Owner
Author
Follow-ups after the final validation round (head ba232b9)The Risk Assessment above predates the last fix round.
Operator follow-up (host configuration, not a code defect)
Known limits and follow-up work (info findings, none changes routing correctness)
|
…rerequisite remains
… narrow TeamClaude unknown
…d-in capacity as unknown
…laude stand-in capacity
… notice keeps state
…unch-policy.test.sh, "replacement did not run omp exactly once". This PR caused it, and it is now fixed. - Rule: the fake `omp` in the test must record `launch:omp` only when a session starts. A read-only query must not count as a launch. - Cause: this PR's capacity check (bin/fm-dispatch-capacity-lib.sh, fm_omp_codex_capacity) runs `omp usage --provider openai-codex --json` on the relaunch path. It runs once from fm-control and once from fm-spawn. The fake `omp` already answered `models --json` as a query. It recorded every other call, including `usage`, as a launch. So the test saw 3 launches, not 1. I proved this by logging each call to the fake: 2 were `usage` calls and 1 was the real launch. - Fix: the fake `omp` now answers `usage` with `{"reports":[]}` and exits without recording a launch. The product reads that report as "unknown", keeps the requested profile, and launches as before. I checked every place this fake is used. The same 4-line change covers them all: the replacement case, the fresh-omp case, and the recovery cases. No product code changed. - Proof: the test failed locally before the fix. It passed on base ca6c277. After the fix it exits 0 with 149 ok lines. `bin/fm-lint.sh` on the test file exits 0. 2. tests/fm-wake-queue.test.sh, "the watcher never blocked inside the drain-ring idle capture: check: rearm-resurface". This PR did not cause it, so I did not change any code for it. - The same failure came before this PR. It hit main commit 8ce400f in CI run 38019754012 (shard 10). It also hit an unrelated branch, fm/fm-ledger-dated-fold-cost, in CI run 38027172418. It is a timing flake that was already on main. - This PR does not change any watcher or wake file. Its one change in a file the watcher loads, bin/fm-status-event-lib.sh, only skips the model-matrix fallback note. - On this macOS host, this test fails at an earlier step on both the base snapshot and the PR head, so I could not run that test case here. That local failure is a host difference, not a result of this PR. Changed file: tests/fm-session-launch-policy.test.sh only
…commits now sit on top of the base. Rule: the branch must keep both sides. The base added omp-session.json, reboot-notice, launch_proof and an omp version check. The PR added omp-fallback.yml and dispatch_rule. Each launch, cleanup and relaunch site must carry both. Conflicts, all fixed by keeping both sides: 1. bin/fm-teardown.sh, child-home cleanup: removes .omp-fallback.yml, .live-model, .omp-session.json and .reboot-notice. 2. bin/fm-teardown.sh, main-task cleanup: removes the same files. 3. bin/fm-control-lib.sh, fm_control_harness_wiring_paths omp: prints omp-ext.ts, omp-session.json and omp-fallback.yml. 4. bin/fm-spawn.sh, preserve_relaunch_meta key list: has both dispatch_rule and launch_proof. 5. tests/fm-omp-harness.test.sh: kept the base's new test, test_spawn_refuses_unsupported_omp_before_launch. 6. tests/fm-omp-harness.test.sh: the wiring-path assertion expects all 3 paths. I set it in commit fa921c6. It conflicted again in 316951d, and I kept the 3-path version. One fix after the rebase. It is in the working tree and not committed yet. - tests/fm-control-relaunch.test.sh: the base added a version check, fm_control_omp_launch_check. It needs `omp --version` to report 18.1.20 or newer. The fake `omp` in the PR's rl-pool test did not answer `--version`. So the auto-relaunch was refused with "installed omp version 'unknown'". I added `--version) printf 'omp/18.1.20\n' ;;` to that fake. The base's own fakes use this same line. No product code changed. Tests (bash). All exit 0 with 0 "not ok": 1. fm-control-relaunch: failed before the fix, at the rl-pool case. Passes after the fix with 191 ok. 2. fm-omp-harness 3. fm-teardown 4. fm-quota-choose 5. fm-dispatch-capacity 6. fm-dispatch-resolve 7. fm-crew-state 8. fm-launch-proof 9. fm-session-launch-policy Lint: bin/fm-lint.sh on the 6 changed files exits 1. It gives one warning, SC2100, on tests/fm-omp-harness.test.sh line 363: `id=omp-fallback-q2`. That line comes from the PR's own commit d79d255 and was there before the rebase. The rebase did not cause it. I did not change it, because the rules forbid unrelated edits. If lint must pass, the smallest fix is to quote it: `id='omp-fallback-q2'`
MrGTV-love
force-pushed
the
fm/fm-omp-codex-account-failover-v2-r3
branch
from
October 10, 2026 19:35
ba232b9 to
7e8628b
Compare
This was referenced Oct 10, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
if codex is exhausted, why not switch to a different model like Deepseek flash v4, or sonnet 5.5 via tc run?
vs hold up work??!! we have a model matrix, so why not switch models accordingly to the fallback??
also confused. we have two codex credentials. the other one is completely unutilized
why did you hold up the work vs switching models?
okay. paid calls for development work do not. update so that deepseek flash replaces glm 5.3
deepseek/deepseek-v4-flash is what I think it is
Context: on 2026-10-05 the fleet sat about 13 hours paused because quota-axi reported the Codex weekly window exhausted. quota-axi reads only one Codex account (greentree promax). omp pools two openai-codex accounts (greentree promax and gmail prolite;
omp usagelists both, and the gmail account also shows a banked saved reset). Workers and every no-mistakes review agent run on omp openai-codex models, so all work and all PR pipelines stopped together. config/crew-dispatch.json (the model matrix) already names stand-ins per rule: z-ai GLM 5.3 for Sonnet-class work, GLM 5.3 Flash for Luna-class work, and gpt-6.1-sol and Opus 5.5 (Claude only through the TeamClaude proxy) as equals for the strongest class. None were applied. Earlier, on 2026-09-30, a worker stalled about 9 hours on usage_limit_reached on one account while the other account had quota; a relaunch recovered it.What Changed
bin/fm-dispatch-capacity-lib.shandbin/fm-dispatch-capacity.sh. They readomp usage --provider openai-codex --jsonand report each pooled Codex account asusable,exhausted, orunknown, without account identities or credentials.fm-dispatch-resolve.shandfm-quota-choose.shnow use this pool verdict forompopenai-codex/*candidates instead of quota-axi's single-account Codex row. Saved resets are disclosed but never spent or counted, and the pool gets no synthesizedspendPriority.fallbackand top-leveldefault_fallbacklists of explicit stand-ins to the dispatch config, validated infm-dispatch-resolve.sh. Only proven primary exhaustion activates a stand-in.fm-spawn.shandfm-control.sh relaunchshare the same lists, with new--dispatch-ruleflag. The chosen rule id is recorded with the task, follows recovery, and the launch and replacement route is appended to task status. The.omp/fm-worker-overlay.ymlturns on native credential rotation and usage-aware fallback. Per-task overlays carry the matched rule's model chain. A Claude stand-in requires the TeamClaude launcher.openrouter/z-ai/glm-5.3stand-in withdeepseek/deepseek-v4-flashindocs/configuration.mdanddocs/examples/model-index.json. Updated the dispatch and harness docs and skills. Addedtests/fm-dispatch-capacity.test.sh, and extended the dispatch-resolve, quota-choose, control-relaunch, omp-harness, crew-state and session-launch-policy tests.Risk Assessment
Testing
I made a disposable lab home with bin/fm-lab-home.sh. It had a private tmux socket, a lab treehouse root and a lab matrix. I ran real fm-spawn, fm-control, fm-teardown, fm-crew-state, the session-end recovery scan, fm-dispatch-capacity and fm-quota-choose against the real omp 18.8.7, using the machine's existing logins. The real Codex pool was exhausted, so every switch was driven by real evidence. For the native-fallback and recovery cases, a lab PATH wrapper made
omp usageunreadable at spawn time. That kept the worker on Codex, and the real Codex error then fired. One step was staged: placing the real extension note line after each worker state line, because omp kept the fallback model for the rest of the session. The lab tmux server, lab home and /tmp/fm-labr3* task temp folders were all removed. The worktree is clean. I also ran six focused test files in the background through bin/fm-test-run.sh. fm-crew-state, fm-dispatch-capacity and fm-omp-harness passed. The runner was then stopped by a signal (exit 144) at the start of fm-teardown, after the report was sent, so fm-teardown, fm-dispatch-resolve and fm-quota-choose never finished. No test reported a failure. Remote CI owns those files. No screenshots were taken because there is no UI change. Evidence is CLI transcripts and pane captures.Evidence: Spawn on exhausted real Codex pool switches to declared stand-in
Source: Spawn on exhausted real Codex pool switches to declared stand-in
Evidence: Unreadable omp usage reads as unknown; quota-axi sees one Codex account
Source: Unreadable omp usage reads as unknown; quota-axi sees one Codex account
Evidence: Invalid fallback blocks only its own matched rule
Source: Invalid fallback blocks only its own matched rule
Evidence: Strongest class with no stand-in refuses instead of using a weak model
Source: Strongest class with no stand-in refuses instead of using a weak model
Evidence: Native inline fallback served by omp from the per-task chain
Source: Native inline fallback served by omp from the per-task chain
Evidence: Fallback note never replaces the worker's state line
Source: Fallback note never replaces the worker's state line
Evidence: Relaunch switches to stand-in and stays there
Source: Relaunch switches to stand-in and stays there
Evidence: Real usage_limit_reached recovers automatically onto stand-in, work kept
Source: Real usage_limit_reached recovers automatically onto stand-in, work kept
Evidence: Teardown removes the per-task fallback overlay
Source: Teardown removes the per-task fallback overlay
Evidence: Quota chooser reads the OMP pool
Source: Quota chooser reads the OMP pool
Evidence: DeepSeek flash replaces GLM stand-in in examples
Source: DeepSeek flash replaces GLM stand-in in examples
Evidence: Focused test run log (3 files passed, then interrupted)
Source: Focused test run log (3 files passed, then interrupted)
Pipeline
Updates from git push no-mistakes
... (28 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)
🔧 **Review** - 8 issues found → auto-fixed (9) ✅
🔧 Fix applied.
2 infos still open:
docs/configuration.md:1379- docs/configuration.md says 'Launch and successful replacement append the selected route to task status.' When typed resolution selects the stand-in (bin/fm-dispatch-resolve.sh:201-215), it emits--harness omp --model deepseek/deepseek-v4-flash --dispatch-rule rule_N. fm-spawn then matches rule_N through its fallback list. fm_dispatch_select measures the stand-in as unknown and keeps it, so DISPATCH_SWITCHED stays false and no status line is written (bin/fm-spawn.sh:6574). Status lines appear only when spawn or control itself switches (bin/fm-control.sh:1448). The smallest correction is the docs: say that a spawn or relaunch that itself switches to a stand-in appends the route, and that a typed-resolution stand-in is disclosed in the resolver output.docs/configuration.md:1522- Fix round caf064f left this docs sentence behind. The round made spawn and relaunch check the fallback schema only for matched rules (bin/fm-dispatch-capacity-lib.sh:131). It updated the library header (lib:9-12). The owner doc still says only 'a retired model or unconfigured role in an unrelated rule does not block a launch or recovery'. It does not say that an invalid fallback in an unrelated rule (for example a codex stand-in, effort 'ultra', or harness 'pi') also no longer blocks spawn or relaunch. It also does not say that typed resolution still refuses the whole file for that same error (bin/fm-dispatch-resolve.sh:235-237). An operator who reads this can expect typed intake to accept a file that it refuses. Fix: extend the sentence at line 1522 to name the unrelated invalid fallback schema too. Add that typed resolution still validates every rule's fallback list. This is a docs change only.✅ No issues found.
🔧 **Test** - 2 issues found ✅
docs/configuration.md:1458- On this host,omp models --jsonhas nodeepseek/provider. Onlyopenrouter/deepseek/deepseek-v4-flashand its:freevariant exist. A matrix rule whose DeepSeek stand-in is written asdeepseek/deepseek-v4-flashwill therefore be refused live with 'no supported permitted fallback', and the work is held again (Scenario B). The refusal does not name the cause. Before the matrix is activated, the operator must configure the direct DeepSeek provider in omp or pick a selector that is in the catalog.bin/fm-spawn.sh:5562- In the live lab-s4 worker, omp's advisor request first fell back 'openai-codex/gpt-6.1-sol:xhigh -> openai-codex/gpt-6-luna:high' and then followed the task chain. The per-task overlay sets no key for the advisor model, so a chain from the captain's own omp config still applied to the advisor. The task's own model followed only the declared chain.bin/fm-dispatch-capacity.sh --harness omp --model openai-codex/gpt-6-lunaagainst the realomp usage --provider openai-codex --jsonbin/fm-lab-home.sh create $LAB+bin/fm-lab-home.sh tmux-dir $LAB+tmux -L fm-lab new-session(private lab socket)bin/fm-spawn.sh lab-s1 $LAB/projects/demo --scout --harness omp --model openai-codex/gpt-6-luna --effort high --dispatch-rule rule_1(real exhausted pool -> stand-in)bin/fm-spawn.sh lab-s2 ... --model openai-codex/gpt-6-terra --dispatch-rule rule_2(deepseek/deepseek-v4-flash stand-in not in this host's catalog -> refusal)bin/fm-spawn.sh lab-s3 ...with a lab view of one exhausted and one usable account (keeps Codex model; writes the native chain overlay)Real omp worker hit Codex usage_limit_reached; checkedstate/lab-s3.busy-stateevent=quota-exhaustedRealbin/fm-wake-drain.sh+bin/fm-watch.shsupervise rounds (session-end scan chose lab-s3 and invoked relaunch)bin/fm-control.sh lab-s3 relaunch --note '<quota note>'with the real exhausted pool -> zai/glm-4.5-flash, unlanded.txt keptbin/fm-spawn.sh lab-s4 ...-> real omp native fallbackopenai-codex/gpt-6-luna:high -> zai/glm-4.5-flash:highwith status 'model-matrix fallback served'bin/fm-control.sh lab-s3 relaunch --note ...again (stays on the stand-in)bin/fm-quota-choose.sh --snapshot quota-snap.json omp:openai-codex/gpt-6-luna(real pool, then lab healthy-sibling view) andcodex:gpt-6-lunabash tests/fm-dispatch-resolve.test.shbash tests/fm-dispatch-capacity.test.shLab teardown:tmux -L fm-lab kill-server,bin/fm-lab-home.sh teardown,rm -rf $LAB✅ No issues found.
capacity: unknownworking: model-matrix fallback launched ...; overlay has an empty default chainrelaunched labfo1 ... model=google-antigravity/gemini-3.7-flash; statusfallback relaunched; thendone; commit presentFallback: openai-codex/gpt-6-luna:medium -> google-antigravity/gemini-3.7-flash:low,Fallback succeeded; twonote [at=..]: model-matrix fallback served ...lines; the…error: invalid dispatch fallback configuration: ... fallback must be an array of explicit OMP profiles ..., exit=1, no labfo3 stateevent=quota-exhausted; wakecheck: labfo4 auto-relaunched after quota exhaustion; meta moved to the stand-in; statusfallback relaunched, thendone,…error: dispatch profile has exhausted capacity and no supported permitted fallback; no task launchednone, exit 1, with both Codex candidates on the really exhausted poolno TypeSafe or OpenRouter key). Without it the tool prints `dispatch-resol…bin/fm-dispatch-capacity.sh --harness omp --model openai-codex/gpt-6-luna(and--json, gpt-6.1-sol), against the realomp usage, compared with the quota-axi codex rowthe same capacity CLI withomp usagemade unreadable by a lab PATH wrapper, so capacity is unknownbin/fm-spawn.sh labfo1 <lab>/projects/notes --mode local-only --yolo on --harness omp --model openai-codex/gpt-6-luna --effort mediumin a lab FM_HOME while the Codex pool was really exhausted; an unrelated rule_2 had an invalid fallbackbin/fm-control.sh labfo1 relaunch --harness omp --model openai-codex/gpt-6-luna --effort medium --note ...spawn labfo2 with pool usage unknown, so the worker starts on Codex; a real usage_limit_reached led to omp's native fallback through the per-task overlay chainbin/fm-crew-state.sh labfo2andlast_status_linewith the real fallback note line placed after each of done, needs-decision, blocked, failed, paused, plus a control with a worker's own note linebin/fm-spawn.sh labfo3 ... --model openai-codex/gpt-6.1-sol --effort high, where the matched rule_2 has an invalid fallback; then again with rule_2 fallback set to []spawn labfo4 on Codex with no fallback chain; real usage_limit_reached gave busy event=quota-exhausted; after adding the matrix fallback, one watcher tick offm_session_end_relaunch_scanbin/fm-quota-choose.sh --snapshot <quota-axi --json> --candidate omp:openai-codex/gpt-6-luna --candidate omp:openai-codex/gpt-6.1-solFM_CONFIG_OVERRIDE=<copy of docs/examples> bin/fm-model-index.sh model pi stand-in:sonnet-gradebin/fm-test-run.sh --json ... tests/fm-dispatch-capacity.test.sh tests/fm-dispatch-resolve.test.sh tests/fm-quota-choose.test.sh tests/fm-omp-harness.test.sh tests/fm-crew-state.test.sh tests/fm-session-end-relaunch.test.sh tests/fm-control-relaunch.test.sh(all 7 passed, exit 0)✅ No issues found.
bin/fm-dispatch-capacity.sh --harness omp --model openai-codex/gpt-6-luna(real omp usage, 2 pooled accounts)bin/fm-dispatch-capacity.shwith omp usage made unreadable (expect unknown, not exhausted)fm-spawn.sh labr3a <lab>/projects/notes --mode local-only --yolo on --harness omp --model openai-codex/gpt-6-luna --effort mediumagainst real exhausted poolfm-spawn.sh labr3b ... --harness omp --model openai-codex/gpt-6.1-sol --effort highwith invalid matched fallback (expect refusal)fm-spawn.sh labr3b ...with strongest rule fallback [] (expect refusal, no weak stand-in)fm-spawn.sh labr3c ...with usage unreadable → native chain overlay → real usage_limit_reached → omp served stand-in → extension note linesfm-crew-state.sh labr3candlast_status_linewith the real note line after done/needs-decision/blocked/failed/pausedfm-control.sh labr3a relaunch --harness omp --model openai-codex/gpt-6-luna --effort medium(switch to stand-in)fm-control.sh labr3a relaunchwith no flags (stand-in sticky, 0 usage probes)fm-spawn.sh labr3d ...with no stand-in → real quota-exhausted busy event → stand-in added →fm_session_end_relaunch_scantick (fromgit archive HEADcopy)fm-teardown.sh labr3cafter lab landing (fallback overlay removed)quota-axi --json | bin/fm-quota-choose.sh --candidate omp:openai-codex/gpt-6-luna --candidate omp:openai-codex/gpt-6.1-soland with a claude second candidateFM_CONFIG_OVERRIDE=<docs/examples copy> bin/fm-model-index.sh model pi stand-in:sonnet-gradebin/fm-test-run.sh tests/fm-crew-state.test.sh tests/fm-dispatch-capacity.test.sh tests/fm-omp-harness.test.sh ...(these 3 passed; fm-teardown, fm-dispatch-resolve, fm-quota-choose interrupted by signal before finishing)✅ **Document** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ No issues found.
✅ No issues found.
🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ **Push** - passed
✅ No issues found.
✅ No issues found.
✅ No issues found.