Skip to content

feat(bin): automate model fallback on depletion - #20

Merged
adibirzu merged 16 commits into
mainfrom
fm/fm-routing-fallback-code
Aug 24, 2026
Merged

adibirzu merged 16 commits into
mainfrom
fm/fm-routing-fallback-code

Conversation

@adibirzu

Copy link
Copy Markdown
Owner

Intent

Land PR 19's completion: automatic model-fallback as genuine code (captain ruling 2026-08-24 - real code + schema + validation, never documentation-only prose), based on fm/fm-routing-model-fallback head 184ea16, branch head 1db9aaf already carries the engine; this run must also fold in the accepted review-round fixes and two review findings below. All prior rulings remain binding.

Accepted requirements (already implemented at 1db9aaf, keep them intact):

  1. Validated per-runtime modelFallback chains (legacy alias _model_fallback) with loud malformed-config rejection; agy gemini ladder stepping past 429/RESOURCE_EXHAUSTED into its Claude tail, opencode free-tier rotation, cursor third-party ids then auto, ClinePass glm-5.3 <-> kimi-k3 error-triggered.
  2. Depletion detection via one shared subscription-vocabulary classifier (fm-dispatch-select.mjs classify-evidence): framed 429, rate limit, RESOURCE_EXHAUSTED, spending-limit/403-budget; context-window/tool-output ceilings and plain auth errors never trigger.
  3. In-place relaunch through fm-runtime-handoff.sh preserving worktree+work; effort axis reset on model switch; routing provider re-declared only where provable across spawn/handoff/control.
  4. AUTO-STEP-DOWN standing rule: strongest-class depletion still steps down automatically - availability beats escalation, routine depletion never parks on the captain; visibility via progress note + working status line naming both models and the matched signature.
  5. Chain walk then fallbackLanes lane movement; blocked routing decision only after all automatic moves are spent.

Accepted fix-round requirements from the ruled review gate (nm-review-askuser) that this run MUST implement:
A. Wire fm-model-fallback.sh apply into the supervision/recovery status-event boundary so classified depletion relaunches automatically - ship/scout tasks only, gated by the classifier (no evidence, no relaunch). Dead code is the defect class this task exists to kill.
B. Validate config shape at the consumption boundary too (plan/apply refuse duplicate-valued or malformed chains even when bootstrap has not run).
C. Scope the AGENTS.md "chain exhausted -> stop and report" sentence to dispatch-side selection so it cannot read as contradicting auto-step-down; document the universal classify-evidence trigger with its narrow signature vocabulary (telemetry-backed providers additionally park via record-failure).

Two open review findings from the failed run that this run MUST also resolve:
D. Chainless non-terminal lanes must advance through their later fallbackLanes successors instead of dying with "no modelFallback chain configured"; a lane without a chain is exhausted-in-lane for traversal purposes (or reject non-terminal chainless lanes loudly - prefer advancement, it matches availability-beats-escalation).
E. When exhausted apply records the blocked event it must consume the evidence (advance fallback_cursor) so supervision cannot reprocess the same depletion line into an infinite blocked loop.

Tests must cover each link including A's watcher wiring, B's refusal cases, D's chainless-lane advancement, and E's no-loop guarantee; docs match implemented code. Push to the captain's fork only; never merge; the pipeline raises the PR referencing #19 as the work it completes.

What Changed

  • Added validated per-harness model-fallback chains, cyclic lanes, and fallback-lane traversal with shared depletion-evidence classification.
  • Automatically applies eligible ship and scout fallback handoffs from watcher status events, preserving worktrees and recording visible switch/blocked status events.
  • Extended routing-provider support and documented the fallback configuration and supervision behavior.

Risk Assessment

✅ Low: The fallback implementation is bounded, validates configuration at use, consumes evidence safely, and the latest self-generated-status exclusion preserves fresh worker evidence.

Testing

Exercised model depletion classification, malformed-chain refusal, chainless lane traversal, exhaustion cursor consumption/no-repeat blocking, real in-place handoff preserving worktree changes, and watcher-triggered automatic apply. The captured behavior transcript demonstrates the end-user fallback flow; no targeted failures were observed.

Evidence: Model fallback end-to-end behavior transcript

Source: Model fallback end-to-end behavior transcript

ok - classify: framed 429 => depleted
ok - classify: RESOURCE_EXHAUSTED => depleted
ok - classify: 403 spending limit => depleted
ok - classify: insufficient credits => depleted
ok - classify: credit balance too low => depleted
ok - classify: explicit rate limit => depleted
ok - classify: context token ceiling => none (ordinary working state)
ok - classify: tool-output ceiling => none
ok - classify: plain authorization 403 => none
ok - classify: unframed 429 => none
ok - plan walks the chain to the entry after the current model
ok - plan starts an out-of-chain (default) model at the strongest configured entry
ok - plan reports exhaustion at the chain tail with exit 3
ok - an exhausted chain moves to the next configured lane's chain head
ok - a configured cline cycle alternates glm-5.3 and kimi-k3
ok - moving to a lane with no configured chain launches on that lane's default model
ok - a chainless non-terminal lane advances through fallbackLanes
ok - walking out the final configured lane reports exhaustion rather than wrapping
ok - _model_fallback remains a first-class spelling of the chain
ok - refuses when no crew-dispatch.json exists
ok - refuses malformed config loudly instead of selecting around it
ok - refuses duplicate chain ids at fallback execution time
ok - refuses a cycle that cannot change models
ok - a chainless terminal lane reports exhaustion without improvising a model
ok - a healthy worker is never relaunched by fallback: no evidence, no action
ok - refuses a secondmate; its recovery is a different owner
ok - apply relaunches in place with the next model, logs the downgrade, parks the provider, and consumes the evidence
ok - depletion written during handoff remains actionable on the next apply
ok - the same evidence can never trigger a second step-down
ok - new depletion evidence after the cursor continues down the chain
ok - a non-native harness records its telemetry-backed routing provider
ok - full exhaustion parks its recorded provider and consumes evidence before blocking
ok - when the relaunch fails, the evidence stays unconsumed for the retry supervisor
ok - over the real handoff path, fallback preserves the worktree and work while stepping the model down
All fm-model-fallback tests passed.
- Outcome: 🔧 1 issue found → auto-fixed ✅ across 2 runs (14m17s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 3 issues found → auto-fixed (3) ✅
  • 🚨 bin/fm-model-fallback.sh:165 - Criterion D requires “Chainless non-terminal lanes must advance through their later fallbackLanes successors,” but this unconditional refusal runs before lane traversal. A depleted cursor task with no modelFallback.cursor and fallbackLanes: [&#34;agy&#34;,&#34;cursor&#34;,&#34;opencode&#34;] exits instead of advancing to opencode; the added test covers only a chainless successor, not this required path.
  • 🚨 bin/fm-model-fallback.sh:252 - Criterion E requires exhausted apply to consume evidence when recording blocked:, but this path appends the blocked event and exits without advancing fallback_cursor. The watcher will re-read the original depletion line on the next status signal and repeatedly append blocked events. Commit the cursor after the blocked append under the existing meta lock, and add the required no-loop regression.
  • 🚨 bin/fm-model-fallback.sh:272 - The required behavior says a depleted task carrying a telemetry-backed routing provider records failure through record-failure. This derives the provider solely from the harness, ignoring valid recorded non-native routes such as harness=pi with provider=claude; those depleted tasks never create a cooldown. Exhausted tasks also return before this block. Use the validated recorded provider and record it before either an exhausted return or relaunch.

🔧 Fix: Fix fallback lane traversal and evidence consumption
2 errors still open:

  • 🚨 docs/examples/crew-dispatch.json:49 - Required criterion 1 says “ClinePass glm-5.3 <-> kimi-k3 error-triggered,” but this ordered, non-repeating chain only supports glm-5.3 -&gt; kimi-k3. A depleted task already on kimi-k3 reaches the generic exhausted path and blocks (there is no later lane), rather than returning to glm-5.3; duplicate-chain validation also prevents expressing the reverse edge. Clarify/implement the required bidirectional behavior.
  • 🚨 bin/fm-model-fallback.sh:196 - The cursor is set to the status file’s size only after the handoff and fallback logging complete. If the replacement worker appends fresh depletion evidence while that handoff is running, that new evidence is included in fallback_cursor and is never classified on the next watcher signal, leaving the newly depleted model running. Snapshot the consumed evidence end before handoff and commit that boundary (without advancing over later writes); add a regression where a post-handoff status append remains actionable.

🔧 Fix: Add cyclic cline fallback and preserve fresh evidence
1 error still open:

  • 🚨 bin/fm-model-fallback.sh:336 - The fallback’s own visibility status line embeds the matched depletion signature (for example, &#34;Error 429&#34;), but the cursor remains at the pre-handoff EVIDENCE_END. On the next watcher signal, classify-evidence reads that newly written line, matches its framed 429, and relaunches again without new worker evidence—walking the chain to exhaustion. Exclude trusted fallback-generated status entries from classification (while retaining their visibility) or otherwise consume only that self-generated entry without advancing over evidence written during handoff.

🔧 Fix: Prevent fallback visibility events retriggering model switches
✅ Re-checked - no issues remain.

🔧 **Test** - 1 issue found → auto-fixed ✅
  • ⚠️ bin/fm-model-fallback.sh:188 - On macOS, the second apply after exhausted evidence was consumed correctly avoids a repeat blocked event, but emits head: illegal byte count -- 0 before the intended no-evidence refusal. Avoid invoking BSD head -c 0 on the empty-evidence path.
  • bash tests/fm-model-fallback.test.sh
  • Focused watcher wiring: test_depletion_signal_applies_model_fallback from tests/fm-watch-triage.test.sh
  • Manual CLI/state verification of chainless-lane traversal and exhausted-evidence replay prevention

🔧 Fix: Avoid BSD head zero-byte fallback errors
✅ Re-checked - no issues remain.

  • bash tests/fm-model-fallback.test.sh
  • bash tests/fm-dispatch-select.test.sh
  • bash tests/fm-watch-triage.test.sh (the watcher’s depletion-to-apply boundary completed successfully before the tool response window)
  • bash tests/fm-runtime-handoff.test.sh (relevant handoff checks exercised before the tool response window)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

The modelFallback feature was documentation-only: bootstrap validated the
config shape, but no code read the chain when a worker actually depleted.
This lands the missing engine and its wiring.

- bin/fm-model-fallback.sh owns the depletion response mechanically:
  classify fresh status evidence past a fallback_cursor= byte guard, walk
  the harness's chain (entry-after-current, head for out-of-chain models,
  legacy alias honored), move to the next fallbackLanes entry on lane
  exhaustion, relaunch in place through fm-runtime-handoff.sh with a
  visibility note, log the downgrade as a working status line, park the
  telemetry provider via record-failure, and record blocked loudly only
  when every automatic move is spent. Auto-step-down is standing policy:
  availability beats escalation, so depletion of the strongest class still
  proceeds automatically instead of parking on the captain.
- fm-dispatch-select.mjs classify-evidence exposes the one subscription-
  vocabulary regex as a pure classifier (now covering RESOURCE_EXHAUSTED,
  spending limits, and budget-framed 403s), shared by record-failure and
  the new runner so quota vocabulary has a single owner.
- spawn reuse drops provider= on every relaunch; runtime-handoff and
  control relaunch re-declare it where provable (native target provider,
  or the recorded one when the adapter is unchanged), so a stale routing
  provider can no longer survive an adapter switch. The effort axis is
  deliberately reset on a model switch instead of silently inherited.
- bootstrap validation rejects duplicate ids inside a chain and validates
  the new optional fallbackLanes order (non-empty, verified, duplicate-free).
- tests: new fm-model-fallback suite (25 cases: signature classification,
  traversal order, lane movement, schema refusals, cursor idempotency,
  failure atomicity, real-path worktree preservation) plus handoff
  provider-axis cases and bootstrap schema rows; docs updated to match
  the implemented behavior in configuration.md, architecture.md,
  AGENTS.md section 4, the dispatch skill, and the example config.
@adibirzu
adibirzu merged commit 016ac05 into main Aug 24, 2026
13 checks passed
@adibirzu
adibirzu deleted the fm/fm-routing-fallback-code branch September 6, 2026 20:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant