Repository navigation
fix: calibrate typed dispatch runway checks to task duration - #11
Merged
Merged
Conversation
MrGTV-love
force-pushed
the
fm/fm-jev-resolver-runway-guard
branch
from
October 5, 2026 16:21
24b60f9 to
9e59aa0
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
do all of hte proposed fork fixes (2026-09-29, answering the Jev value report data/fm-jev-value-scout/report.md section 5, 'Fork follow-up shortlist'). Goal stated with it: "The user is looking to maximize performance while minimizing cost. Jev seems like a massive lever." Jev is typesafe.ai's System One model; firstmate calls it only from bin/fm-dispatch-resolve.sh (typed dispatch resolution). Work lands on the fork MrGTV-love/firstmate only; never push to, open, or comment on anything in kunchenguid/firstmate.
This task covers shortlist item 3:
3. Review quota-axi minimum version/provider coverage and represent openrouter/remaining-runway constraints so "clear" cannot imply a viable profile when ranking evidence is weak.
Report evidence: a clear result picked omp openai-codex/gpt-6-luna while that account showed 6% remaining with projected exhaustion, and all 77 held-out clears in the dispatch evaluation (data/fm-jev-dispatch-eval/report.md) chose a provider with projected quota exhaustion. omp pools two Codex accounts but quota-axi sees only one, so it understates headroom; omp OpenRouter candidates have no quota row and resolve 'eligible, unranked'.
sounds like the wrong fix. if the safety check is blockign every jev pick, that suggests the safety check is miscalibrated or misconfigured
Utilizing Jev often is important. It will improve quality, accuracy, speed, while also reducing costs substantially.
What Changed
task_horizon_minutesconfiguration (default: 240). Escalate when the highest-ranked candidate has established runway shorter than the task horizon without substituting a lower-ranked profile; disclose early or unknown projections as warnings rather than vetoes.Risk Assessment
✅ Low: The changes are bounded to typed resolution, preserve the shared quota contracts, and implement the authorized horizon and pooled-account behavior without a substantiated regression or intent contradiction.
Testing
Targeted resolver regressions, controlled executions of the actual resolver CLI, and isolated read-only bootstrap checks all passed. Evidence preserves inputs, commands, exit codes, and product output. Controlled quota and model answers are explicitly distinguished from fresh live-provider readings; no fresh provider or Jev calls were made. Disposable fixtures were removed.
Evidence: Controlled resolver CLI transcripts
Source: Controlled resolver CLI transcripts
Evidence: Controlled resolver scenario results
Source: Controlled resolver scenario results
Evidence: Bootstrap task-horizon diagnostics
Source: Bootstrap task-horizon diagnostics
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-dispatch-resolve.test.shwith inherited path overrides unset and TMPDIR isolated inside the worktree; targeted resolver regressions passed.python3 .runway-validation/drive.pyexecuted the actualbin/fm-dispatch-resolve.shagainst an isolated FM_HOME, controlled quota inputs, and controlled model answers served over localhost through real curl.FM_BOOTSTRAP_NETWORK=skip FM_BOOTSTRAP_DETECT_ONLY=1 FM_HOME=<isolated> FM_ROOT_OVERRIDE=<isolated> bin/fm-bootstrap.shexercised positive and zero task horizons, environment-key and home-key opt-in, and absent-key behavior.Checked generated CLI output for exactly one unranked-candidate note and absence of quota_summary; retained transcripts and scenario inputs, stopped the disposable HTTP server, and removed the worktree fixtures.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.