Conversation
…ers into one owner
Owner
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Support agy subscription provider and automatic model fallback on depletion
What Changed
agyas a routable subscription provider:fm-dispatch-select.mjsnow includes it inPROVIDERS/NATIVE_PROVIDER(and admitsmuse/agyas verified harnesses),fm-spawn.sh--provideracceptscursorandagywith the native harness/provider match enforced, andfm-control-lib.shplusfm-runtime-handoff.shgainedcopilot/agy(andcursor/muse/cline) adapter facts — interrupt keys, exit command vs.cline'sC-ckey exit — while still refusing those adapters for--secondmate.modelFallbackconfig object (legacy alias_model_fallback) mapping a harness to its ordered model chain:fm-bootstrap.shvalidates the object shape, verified harness keys, and non-empty model-id chains asCREW_DISPATCHdiagnostics, and AGENTS.md/docs/configuration.md/docs/architecture.md/the dispatch skill document the post-dispatch relaunch-in-place path throughfm-runtime-handoff.sh --modelbefore moving to the next harness lane, plus an agy example rule and chain indocs/examples/crew-dispatch.json.RATE_LIMIT_REinfm-dispatch-select.mjsto subscription vocabulary (framed 429, explicit rate limit, named quota/credit/allowance exhaustion) so a barelimit/token/unframed429— e.g. "context token limit reached" — no longer parks a provider for a full cooldown, and consolidated the previously contradictory harness-support rosters indocs/configuration.mdand AGENTS.md into a single owner.Risk Assessment
Testing
I ran the five targeted suites covering every file the change touches (fm-dispatch-select, fm-runtime-handoff, fm-control, the crew-dispatch validation test in fm-bootstrap, and the pre-existing fm-agy-harness suite) and they all pass. Because green units alone do not show the feature, I also built and ran an end-to-end operator walkthrough that drives the real shipped scripts against fake tmux/quota-axi/harness CLIs and captures the whole depletion story as a CLI transcript: an operator declares an agy lane plus a modelFallback chain and real bootstrap accepts it while rejecting empty chains, unverified harness keys, blank model ids and the both-spellings conflict with actionable CREW_DISPATCH diagnostics; the selector prices agy on the declared Antigravity pool (fails closed on the spent claude_gpt_5h pool, selects on the healthy gemini_5h pool with native provider identity); a real 429/quota status line records a cooldown while three working-ceiling lines are correctly refused; fm-runtime-handoff.sh then relaunches the same task in the same worktree on the next chain model twice over (3.7 to 3.6 to 3.5) with HEAD, uncommitted work and PR metadata all preserved; and only after that does the cooled-down agy lane fail over to codex, with a clear restoring eligibility. The same three commands replayed against the base commit refuse agy outright and silently ignore every malformed modelFallback chain, which is the fail-before half of the contrast. No test failures, setup problems or flakiness; the two findings I filed are informational observations about fallback observability, not broken tests. The worktree is clean — all evidence lives in the evidence directory.
Evidence: End-to-end agy + model-fallback walkthrough transcript (5 acts, base-commit contrast included)
Source: End-to-end agy + model-fallback walkthrough transcript (5 acts, base-commit contrast included)
Evidence: Driver script that produced the transcript (runs the real shipped scripts against fake tmux/quota-axi fixtures)
Source: Driver script that produced the transcript (runs the real shipped scripts against fake tmux/quota-axi fixtures)
Evidence: Key excerpt: agy selected on its own pool, then the in-place model fallback
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
AGENTS.md:215- Intent requires "automatic model fallback on depletion", but that half is delivered only as prose pointing at aconfig/crew-dispatch.json_model_fallbackkey that exists nowhere else. AGENTS.md:215 ("Read model fallback chains fromconfig/crew-dispatch.json_model_fallbackwithout hardcoding a duplicate copy") and docs/configuration.md:362 are the only two references in the repo. The field is absent from the canonical schema block at docs/configuration.md:309-327 (a section that declares itself "the single owner of the canonical schema and its per-field semantics"), absent from docs/examples/crew-dispatch.json, and read by no script. It is also unvalidated:crew_dispatch_validatein bin/fm-bootstrap.sh:1046-1090 never enumerates top-level keys, so an operator-authored_model_fallbackwith a typo or malformed chain is silently ignored instead of reported, contradicting the same file's "malformed configuration must be reported and corrected rather than selected around". As shipped, the fallback behavior has no data source and cannot trigger.bin/fm-dispatch-select.mjs:104- RATE_LIMIT_RE was widened to include genericreach|exceedpaired withlimit|token|usage, plus a bare\b429\b. That gate has two consumers with real side effects: the record-failure evidence check (line 588) and the automatic telemetry-evidence cooldown at line 496. A benign status line such asworking: context token limit reached; compactingorfailed: exceeded the tool output limitnow qualifies as verified quota evidence and parks a provider with full headroom for cooldownSeconds (default 1800s), pushing dispatch onto a weaker lane. Coverage is also missing for the change: the new agy test (tests/fm-dispatch-select.test.sh:464) uses429 Too Many Requests, which the pre-change regex already matched, and there is no negative case asserting benign text is rejected. Recommend anchoring the new alternatives to quota vocabulary (e.g. require quota/credit/usage adjacency rather than barelimit) and adding one positive plus one negative case.AGENTS.md:216- The new always-loaded line says depletion detection uses pane/status-log errors "for runtimes without telemetry (such as ClinePass and Grok)". Grok does have telemetry: docs/configuration.md:359 in this same commit says "Providers exposed by quota-axi, including Claude, Codex, Grok, Cursor, and agy, require fresh telemetry", and bin/fm-dispatch-select.mjs:82 keepsgrokin PROVIDERS so it is priced against the reserve and refused on staleness. An agent following AGENTS.md would drive grok depletion purely off pane text and bypass the record-failure/telemetry contract. cline (ClinePass) is the correct example of a telemetry-free runtime; grok appears to be a mistake.tests/fm-dispatch-select.test.sh:509- The closing assertion of test_agy_record_failure_and_cooldown acceptsagyORcodex, which are the only two candidates in the profile array, so it can never fail. Ifclear --provider agyregressed to a no-op, agy would remain in cooldown, codex would be selected, and the test would still pass while claiming to prove "cleared agy cooldown did not restore candidate eligibility". Rotation is least-recently-used and agy has not been dispatched at that point, so asserting= agyis deterministic and actually exercises the clear path.bin/fm-runtime-handoff.sh:194- Beyond agy, this commit admits cursor, muse, cline, and copilot as runtime-handoff targets and wires copilot through the full control-plane fact table (bin/fm-control-lib.sh:66,92,107,119,130,148,160,176,191). The facts match .agents/skills/harness-adapters/SKILL.md, so nothing is invented, but four extra harnesses gain a new exit-then-relaunch lifecycle entry point in a change scoped to "agy subscription provider and automatic model fallback", and tests/fm-runtime-handoff.test.sh:344 only swaps them out of the refusal case rather than adding accept-path coverage. Confirm the widening belongs in this commit.docs/examples/crew-dispatch.json:28- The new agy example (and the header example at bin/fm-dispatch-select.mjs:53) usesgemini-3.7-flash-high. .agents/skills/harness-adapters/SKILL.md:664 enumeratesgemini-3.6-flash-{low,medium,high},gemini-3.5-flash-*, andgemini-3.1-pro-*for the verified Antigravity CLI 1.1.9, so a copied-verbatim 3.7 id would be rejected at launch and, per that adapter doc, surface only as an agy trust-gate timeout rather than a clean model error. The sibling cursor example carries a "confirm against quota-axi --json before copying this" caveat; the agywhyhas no equivalent for the model id.AGENTS.md:214- .agents/skills/quota-array-dispatch/SKILL.md:62 says to "stop and report that the strongest-class choice cannot proceed rather than downgrading it", while AGENTS.md:214 forbids blocking, parking, or escalating a routine depletion and AGENTS.md:218 says the strongest-class rule "still wins over conserving quota". For the common case where the depleted model IS the strongest-class one, the two always-loaded surfaces prescribe opposite actions (relaunch on a weaker chain entry vs stop and report) and nothing in the change resolves which governs.bin/fm-control-lib.sh:97- fm_control_harness_supports_kind now refuses secondmate formuse|cline|copilot|agy(line 107), but the explanatory comment above it still names only muse and cline. bin/fm-spawn.sh:729,1404 and the harness-adapters doc confirm copilot and agy are crewmate/scout-only; the comment should name them so the refusal is not later read as unexplained.🔧 Fix: define and validate modelFallback, tighten quota-evidence regex
5 issues (3 warnings, 2 infos) still open:
bin/fm-runtime-handoff.sh:355- The mandated automatic model-fallback invocation silently drops the recorded effort axis.fm-runtime-handoff.shreads onlymodeandyolofrom meta (lines 166-168) and appends--model/--effortto SPAWN_ARGS only when they were passed on its own command line (354-355).fm-spawn.sh's reuse path dropseffort=from the preserved meta (drop_re, bin/fm-spawn.sh:3130) and rewriteseffort=${EFFORT:-default}(3176), andeffort_flag_for_harnessemits nothing fordefault. So following AGENTS.md:214 / docs/configuration.md:365 verbatim (fm-runtime-handoff.sh <id> --harness <name> --model <next-model>) relaunches a task recorded ateffort=maxat the harness default — exactly the 'silently downgrading reasoning class' that AGENTS.md:217 forbids two lines later. For agy it is worse than a downgrade: a base chain entry such asgemini-3.6-flashREQUIRES--effort(.agents/skills/harness-adapters/SKILL.md:665), so the relaunch fails and, per that same doc, surfaces only as an agy trust-gate timeout. The competing owner is explicit: the sibling caller of the same spawn reuse path,bin/fm-control.sh:617-620,672,681, readsPRIOR_MODEL/PRIOR_EFFORTfrom meta and carries them forward unless overridden. Earliest shared boundary: havefm-runtime-handoff.shdefault MODEL/EFFORT from the recorded meta the way it already does for mode/yolo (or preserveeffort=in the spawn reuse path), rather than documenting the flag at every call site.bin/fm-runtime-handoff.sh:348- After the cross-harness lane move this change mandates ('move work to the next harness lane only when a harness's whole model chain is exhausted', AGENTS.md:215), the task's recorded routing provider goes stale and the cooldown contract can no longer be honoured.provider=is not in fm-spawn.sh's reuse drop_re (bin/fm-spawn.sh:3130), and handoff never passes--provider, so a task spawned asharness=claude provider=claudebecomesharness=agy provider=claudeafter the lane move.taskMetaProviderthen reads providerclaude, computesNATIVE_PROVIDER.get('agy') === 'agy', and dies with 'task <id> has mismatched native harness and provider metadata' (bin/fm-dispatch-select.mjs:598-599) — exit 2, no cooldown. The exhausted agy lane therefore stays a full-headroom dispatch candidate, defeating the fail-closed capacity contract AGENTS.md:218 says remains enforced. This commit widens reachability by adding agy/cursor to NATIVE_PROVIDER and admitting them as handoff targets, and by making the lane move a routine automatic step. Same fix site as the effort gap: carry the recorded routing identity forward (or clear/re-deriveprovider=for the new native harness) in the handoff relaunch rather than leaving the previous lane's identity on the record.tests/fm-runtime-handoff.test.sh:608- The fix round's new accept-path test relaunches on--harness cursor, but the cursor spawn path resolves a real binary:fm_cursor_resolve_binarysearches PATH forcursor-agentthenagent, then~/.local/bin/(bin/fm-cursor-lib.sh:164-185), andfm-spawn.sh:1434exits 1 when that fails. The fixture stubs a binary namedcursor(tests/fm-runtime-handoff.test.sh:115), which no lookup name matches, and PATH keeps the host's entries (line 190). On a host without Cursor Agent CLI installed the test fails with 'handoff from a key-exit harness to cursor should succeed'; on a host with it, the suite shells out to the real Cursor binary in a fixture built entirely on fakes.fm_cursor_verify_executableaccepts any executable whose basename iscursor-agent, so addingcursor-agentto the stub list at line 115 makes the test hermetic without changing what it proves.bin/fm-dispatch-select.mjs:110- The new 429 alternative accepts any of http|status|code|error|response within 16 characters of a bare429.erroris loose enough to keep a false positive alive: a status line such asfailed: error at line 429matches, and both consumers act on it —record-failure(line 602) parks the provider for the full cooldown (default 1800s) with untouched headroom, and providerReadiness (line 383) marks it as quota evidence. The new negative test only covers the unframed variant (working: applying the hunk at line 429 of the diff), so the framed line-number case is uncovered. Dropping bareerror(or requiring tight adjacency, e.g.status[ _-]?code[ _-]?429) keeps every accepted case in the new positive table matching, since those rely onstatus code 429,too many requests, and quota vocabulary.docs/examples/crew-dispatch.json:46- The shipped example chains, which docs/configuration.md:363 tells operators to copy, contain two questionable entries."claude": ["claude-sonnet-5", "sonnet", "haiku"]—sonnetis claude's alias for the same model asclaude-sonnet-5(.agents/skills/harness-adapters/SKILL.md:150), so the first fallback hop relaunches the worker on the model that just depleted and burns a full exit-and-relaunch cycle before reaching a genuinely different model."agy": ["gemini-3.7-flash-high", ...]heads with an idagy modelsdoes not list (SKILL.md:664 enumerates gemini-3.6-flash-, gemini-3.5-flash-, gemini-3.1-pro-*) while its other two entries are real ids, so the example mixes a launch-failing head with valid tail entries; per that adapter doc a bad agy id surfaces only as a trust-gate timeout. The rule profile above got a 'confirm before copying' caveat; the chain has none of its own.bin/fm-runtime-handoff.sh:367- The in-run model-fallback relaunch is not distinguishable in status reporting. bin/fm-runtime-handoff.sh:367 logs onlyworking: runtime handoff to <harness>, so an agy gemini-3.7-flash-high -> gemini-3.6-flash-high fallback and a later gemini-3.6 -> gemini-3.5 fallback produce byte-identical status lines with no model named. AGENTS.md section 4 (added by this change) requires that "Every automatic model switch must be logged and visible in status reporting rather than silently downgrading reasoning class"; the new model id is only observable by reading state/<id>.meta. Confirmed in the captured walkthrough (ACT 4).bin/fm-runtime-handoff.sh:355- Following the documented fallback procedure literally drops the recorded effort axis. AGENTS.md section 4 instructs relaunching in place "via bin/fm-runtime-handoff.sh with --model"; because bin/fm-runtime-handoff.sh:355 only forwards --effort when explicitly supplied, a task recorded at effort=high comes back as effort=default after the model switch. Observed in the captured walkthrough (ACT 4): meta goes frommodel=gemini-3.7-flash-high, effort=hightomodel=gemini-3.6-flash-high, effort=default. This is pre-existing handoff behaviour, but it now sits on the automatic depletion path the change introduces.bash tests/fm-dispatch-select.test.sh— 14 cases pass, including the three new ones (agy per-pool quotaWindow pricing, agy record-failure/cooldown/clear, depletion-evidence gate vs working ceilings)bash tests/fm-runtime-handoff.test.sh— all pass, including the new key-exit-harness -> newly-admitted-target (cline -> cursor) handoff casebash tests/fm-control.test.sh— all pass, including the copilot/agy adapter contract rows and cline/copilot/agy harness-family resolutionbash tests/fm-bootstrap.test.sh(test_crew_dispatch_validation) — all crew-dispatch rows pass, including the 9 new agy/modelFallback validation rowsbash tests/fm-agy-harness.test.sh— pre-existing agy adapter suite still green (busy regex, composer classification, project-trust gate)Manual end-to-end walkthrough driving the realbin/fm-bootstrap.sh,bin/fm-dispatch-select.mjs,bin/fm-spawn.shandbin/fm-runtime-handoff.shagainst fake tmux/quota-axi/harness CLIs:bash /var/folders/j8/ztw_4_wx3691gww5x_1n9r4c0000gn/T/no-mistakes-evidence/01M0SEQE0C1ZFKZDTZWW97YA0Y/agy-model-fallback-e2e.shFail-before/pass-after contrast inside that walkthrough: the same three commands replayed against base commit3c544d6a1958784845356e62bc09d3dd56ee8a67(extracted viagit archive)docs/configuration.md:268- docs/configuration.md "Harness support" states the same verified-harness fact twice and the two copies contradict each other. Line 262 says claude/codex/opencode/pi/pi-signed/grok/kimi/cursor are verified for crewmate and secondmate launches with cline and copilot crewmate-only; line 268 repeats the sentence with a different set (drops cursor, omits cline/copilot, names agy crewmate-only). AGENTS.md:190 carries a third variant that lists cline and copilot in the general set even though bin/fm-spawn.sh:1406-1411 refuses all of muse, agy, cline, and copilot for --secondmate. Both duplicates predate this change (line 268 was introduced by 93cc802), so this is a consolidation follow-up: collapse docs/configuration.md to one authoritative sentence covering all crewmate-only adapters and reduce AGENTS.md:190 to match it or point at it. Out of scope here because this change did not touch either line and the fix spans an always-loaded agent contract.🔧 Fix: consolidate contradictory harness-support rosters into one owner
1 info still open:
tests/fm-dispatch-select.test.sh:516- bin/fm-lint.sh exits 1 on this branch: SC2034 (warning) 'case appears unused' for the local declarationlocal home fakebin quota out rc case name status. The unusedcasevariable is new on this branch (base commit 3c544d6 has no such declaration at any of itslocal home fakebin quota out rc ...sites), so the change introduced the lint failure. Not fixable here because this phase is restricted to documentation and doc comments and this is an executable test file; the outer executor's lint phase owns the fix (dropcasefrom the declared locals).🔧 **Lint** - 1 issue found → auto-fixed ✅
🔧 Fix: drop unused local
casein dispatch-select test✅ Re-checked - no issues remain.
✅ **Push** - passed
✅ No issues found.