Skip to content

fix: calibrate typed dispatch runway checks to task duration - #11

Merged
MrGTV-love merged 7 commits into
mainfrom
fm/fm-jev-resolver-runway-guard
Oct 5, 2026
Merged

MrGTV-love merged 7 commits into
mainfrom
fm/fm-jev-resolver-runway-guard

Conversation

@MrGTV-love

@MrGTV-love MrGTV-love commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Intent

do all of hte proposed fork fixes (2026-09-29, answering the Jev value report data/fm-jev-value-scout/report.md section 5, 'Fork follow-up shortlist'). Goal stated with it: "The user is looking to maximize performance while minimizing cost. Jev seems like a massive lever." Jev is typesafe.ai's System One model; firstmate calls it only from bin/fm-dispatch-resolve.sh (typed dispatch resolution). Work lands on the fork MrGTV-love/firstmate only; never push to, open, or comment on anything in kunchenguid/firstmate.

This task covers shortlist item 3:
3. Review quota-axi minimum version/provider coverage and represent openrouter/remaining-runway constraints so "clear" cannot imply a viable profile when ranking evidence is weak.

Report evidence: a clear result picked omp openai-codex/gpt-6-luna while that account showed 6% remaining with projected exhaustion, and all 77 held-out clears in the dispatch evaluation (data/fm-jev-dispatch-eval/report.md) chose a provider with projected quota exhaustion. omp pools two Codex accounts but quota-axi sees only one, so it understates headroom; omp OpenRouter candidates have no quota row and resolve 'eligible, unranked'.

sounds like the wrong fix. if the safety check is blockign every jev pick, that suggests the safety check is miscalibrated or misconfigured

Utilizing Jev often is important. It will improve quality, accuracy, speed, while also reducing costs substantially.

What Changed

  • Add validated task_horizon_minutes configuration (default: 240). Escalate when the highest-ranked candidate has established runway shorter than the task horizon without substituting a lower-ranked profile; disclose early or unknown projections as warnings rather than vetoes.
  • Treat omp Codex's visible account as a lower bound on its pool: exhaustion leaves the pool eligible but unranked, while declared profile floors remain binding. Document and cover absent or credit-only OpenRouter evidence as eligible but unranked.
  • Enforce the shared quota-axi minimum version before reading quota evidence, expand resolver and bootstrap regression coverage, and update architecture, configuration, and verification documentation.

Risk Assessment

✅ Low: The changes are bounded to typed resolution, preserve the shared quota contracts, and implement the authorized horizon and pooled-account behavior without a substantiated regression or intent contradiction.

Testing

Targeted resolver regressions, controlled executions of the actual resolver CLI, and isolated read-only bootstrap checks all passed. Evidence preserves inputs, commands, exit codes, and product output. Controlled quota and model answers are explicitly distinguished from fresh live-provider readings; no fresh provider or Jev calls were made. Disposable fixtures were removed.

  • Live validation: ✅ go - 17 of 17 scenarios driven live against the product
Scenario Result Live Evidence
Resolve a single account with 6% remaining and established 1800-second runway: escalate without a profile ✅ pass live Controlled resolver CLI transcripts: report-case-six-percent-1800-established
Resolve an established 3600-second runway under the default horizon: escalate without a profile ✅ pass live Controlled resolver CLI transcripts: established-3600-short
Resolve early or absent-confidence projections: clear single-account and pooled profiles with warnings ✅ pass live Controlled resolver CLI transcripts: single-low-confidence-early, single-low-confidence-None, pool-low-confidence-early, pool-low-confidence-None
Resolve established runway exactly at or beyond the 14400-second horizon: clear without warnings ✅ pass live Controlled resolver CLI transcripts: single-exact-14400-boundary, pool-exact-14400-boundary, single-covering-80796-established, pool-covering-80796-established
Configure 30-, 120-, and zero-minute task horizons: honor valid values and reject zero ✅ pass live Controlled resolver CLI transcripts: configured-30-minute-horizon, configured-120-minute-horizon, invalid-zero-horizon
Resolve an exhausted visible pooled account: keep the pool eligible but unranked with a warning ✅ pass live Controlled resolver CLI transcripts: exhausted-pool-unranked, exhausted-pool-beside-measured, schema6-exhausted-pool
Offer a lower-ranked alternative beside a short pool winner: escalate rather than fall back ✅ pass live Controlled resolver CLI transcripts: short-pool-winner-no-same-rule-fallback, schema6-short-pool-winner, healthy-pool-winner-control
Run with incompatible quota-axi: report the minimum version and perform no ranking ✅ pass live Controlled resolver CLI transcripts: outdated-quota-axi, quota-version-unparseable, quota-version-failed-read, minimum-quota-version-control
Apply an explicit captain profile floor to the pool: veto a known shortfall ✅ pass live Controlled resolver CLI transcripts: pool-explicit-floor-safe, pool-explicit-floor-exhausted
Resolve exhaustion or zero availability outside the pool: veto the single account ✅ pass live Controlled resolver CLI transcripts: single-exhausted-veto, single-zero-veto
Resolve a known zero bound on the visible pooled account: disclose uncertainty without vetoing the pool ✅ pass live Controlled resolver CLI transcripts: pool-zero-unranked
Provide healthy alternatives in another rule or default array: do not downgrade a short matched-rule winner ✅ pass live Controlled resolver CLI transcripts: short-pool-no-cross-rule-or-default-fallback, short-single-no-same-rule-fallback
Add a short exact-model bound beside a healthy limiting rank row: escalate ✅ pass live Controlled resolver CLI transcripts: short-nonlimiting-model-bound
Resolve absent or credit-only OpenRouter coverage: keep it unranked and disclose it beside a measured winner ✅ pass live Controlled resolver CLI transcripts: openrouter-absent-alone, openrouter-credit-only-alone, openrouter-absent-beside-measured, openrouter-credit-only-beside-measured
Resolve missing spendPriority evidence: escalate rather than invent a ranking ✅ pass live Controlled resolver CLI transcripts: missing-spend-priority-unranked
Resolve unknown runway or an established projection without usable seconds: clear with disclosed warnings ✅ pass live Controlled resolver CLI transcripts: pool-unknown-warning, pool-missing-seconds-warning
Run bootstrap with task-horizon configuration: validate only on the typed opt-in path ✅ pass live Bootstrap task-horizon diagnostics
Evidence: Controlled resolver CLI transcripts

Source: Controlled resolver CLI transcripts

Actual resolver CLI, controlled inputs via isolated FM_HOME. Quota and model answers are synthetic external inputs, not observed live-provider readings. Real curl communicates with a disposable localhost HTTP endpoint.

=== report-case-six-percent-1800-established ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: escalate
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  reason: highest-ranked candidate codex:gpt-6-luna has established runway shorter than the 240-minute task horizon; completion is not proven
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_models (usableRunwaySeconds=1800 projectionConfidence=established)]
exit=0
quota calls=['--version', '--json']

=== established-3600-short ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: escalate
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  reason: highest-ranked candidate codex:gpt-6-luna has established runway shorter than the 240-minute task horizon; completion is not proven
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_models (usableRunwaySeconds=3600 projectionConfidence=established)]
exit=0
quota calls=['--version', '--json']

=== single-low-confidence-early ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_models (usableRunwaySeconds=1800 projectionConfidence=early)]
  profile: --harness 'codex' --model 'gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== single-low-confidence-None ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_models (usableRunwaySeconds=1800 projectionConfidence=unknown)]
  profile: --harness 'codex' --model 'gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== single-exact-14400-boundary ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible
  profile: --harness 'codex' --model 'gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== single-covering-80796-established ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible
  profile: --harness 'codex' --model 'gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== pool-low-confidence-early ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_models (usableRunwaySeconds=1800 projectionConfidence=early)]
  profile: --harness 'omp' --model 'openai-codex/gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== pool-low-confidence-None ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_models (usableRunwaySeconds=1800 projectionConfidence=unknown)]
  profile: --harness 'omp' --model 'openai-codex/gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== pool-exact-14400-boundary ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible
  profile: --harness 'omp' --model 'openai-codex/gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== pool-covering-80796-established ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible
  profile: --harness 'omp' --model 'openai-codex/gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== configured-30-minute-horizon ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible
  profile: --harness 'codex' --model 'gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== configured-120-minute-horizon ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: escalate
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  reason: highest-ranked candidate codex:gpt-6-luna has established runway shorter than the 120-minute task horizon; completion is not proven
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_mode

... [9447 bytes truncated] ...

 note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=through_reset  bounds=all_models:6%/through_reset,model:gpt-6-luna:50%/projected_exhaustion  -> eligible [warning: projected_exhaustion at model:gpt-6-luna (usableRunwaySeconds=600 projectionConfidence=established)]
exit=0
quota calls=['--version', '--json']

=== openrouter-absent-alone ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: escalate
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  reason: no rankable eligible candidate
  note: rule matched
  candidate: omp:openrouter/provider/model  provider=openrouter  -> eligible, unranked: provider openrouter not in the quota snapshot: disclosed uncertainty
exit=0
quota calls=['--version', '--json']

=== openrouter-absent-beside-measured ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  note: 1 eligible candidate(s) unranked (openrouter)
  candidate: omp:openrouter/provider/model  provider=openrouter  -> eligible, unranked: provider openrouter not in the quota snapshot: disclosed uncertainty
  candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.7597  runway=through_reset  -> eligible
  profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'
exit=0
quota calls=['--version', '--json']

=== openrouter-credit-only-alone ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: escalate
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  reason: no rankable eligible candidate
  note: rule matched
  candidate: omp:openrouter/provider/model  provider=openrouter  -> eligible, unranked: provider openrouter unmeasured (unknown): disclosed uncertainty
exit=0
quota calls=['--version', '--json']

=== openrouter-credit-only-beside-measured ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  note: 1 eligible candidate(s) unranked (openrouter)
  candidate: omp:openrouter/provider/model  provider=openrouter  -> eligible, unranked: provider openrouter unmeasured (unknown): disclosed uncertainty
  candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.7597  runway=through_reset  -> eligible
  profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'
exit=0
quota calls=['--version', '--json']

=== missing-spend-priority-unranked ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: escalate
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  reason: no rankable eligible candidate
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=-  runway=through_reset  -> eligible, unranked: spendPriority missing or non-numeric at all_models: not rankable: disclosed uncertainty
exit=0
quota calls=['--version', '--json']

=== pool-unknown-warning ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=unknown  -> eligible [warning: unknown at all_models]
  profile: --harness 'omp' --model 'openai-codex/gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== pool-missing-seconds-warning ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_models (usableRunwaySeconds=unknown projectionConfidence=established)]
  profile: --harness 'omp' --model 'openai-codex/gpt-6-luna'
exit=0
quota calls=['--version', '--json']

=== schema6-short-pool-winner ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: escalate
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  reason: highest-ranked candidate omp:openai-codex/gpt-6-luna has established runway shorter than the 240-minute task horizon; completion is not proven
  note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=projected_exhaustion  -> eligible [warning: projected_exhaustion at all_models (usableRunwaySeconds=3600 projectionConfidence=established)]
  candidate: cursor:cursor-grok-4.6-medium  provider=cursor  scope=all_models  remaining=91%  spendPriority=0.7597  runway=through_reset  -> eligible
exit=0
quota calls=['--version', '--json']

=== schema6-exhausted-pool ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: escalate
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  reason: no rankable eligible candidate
  note: rule matched
  candidate: omp:openai-codex/gpt-6-luna  provider=codex  -> eligible, unranked: omp Codex account pool is only lower-bounded by its visible account (runway exhausted_now at all_models): disclosed uncertainty [warning: exhausted_now at all_models]
exit=0
quota calls=['--version', '--json']

=== quota-version-unparseable ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: error
  reason: quota-axi requires >= 0.1.51; installed version is older, unreadable, or unparseable
dispatch-resolve: error (quota-axi requires >= 0.1.51; installed version is older, unreadable, or unparseable)
exit=0
quota calls=['--version']

=== quota-version-failed-read ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: error
  reason: quota-axi requires >= 0.1.51; installed version is older, unreadable, or unparseable
dispatch-resolve: error (quota-axi requires >= 0.1.51; installed version is older, unreadable, or unparseable)
exit=0
quota calls=['--version']

=== minimum-quota-version-control ===
$ bin/fm-dispatch-resolve.sh .runway-validation/brief.md --project firstmate
dispatch-resolve:
  status: clear
  model: jev-controlled-input   latency_ms: 0   tokens: -/-
  rule: rule_1 (A simple bug fix with a stated root cause.)   confidence: 0.9
  probabilities: rule_1=0.97 default=0.03
  note: rule matched
  candidate: codex:gpt-6-luna  provider=codex  scope=all_models  remaining=6%  spendPriority=0.9  runway=through_reset  -> eligible
  profile: --harness 'codex' --model 'gpt-6-luna'
exit=0
quota calls=['--version', '--json']
Evidence: Controlled resolver scenario results

Source: Controlled resolver scenario results

{
  "validation_time": "2026-10-05T16:12:40.421212+00:00",
  "scope": "controlled product CLI scenarios; no fresh Jev or provider calls, no lifecycle activity",
  "results": [
    {
      "name": "report-case-six-percent-1800-established",
      "result": "pass",
      "failures": []
    },
    {
      "name": "established-3600-short",
      "result": "pass",
      "failures": []
    },
    {
      "name": "single-low-confidence-early",
      "result": "pass",
      "failures": []
    },
    {
      "name": "single-low-confidence-None",
      "result": "pass",
      "failures": []
    },
    {
      "name": "single-exact-14400-boundary",
      "result": "pass",
      "failures": []
    },
    {
      "name": "single-covering-80796-established",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-low-confidence-early",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-low-confidence-None",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-exact-14400-boundary",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-covering-80796-established",
      "result": "pass",
      "failures": []
    },
    {
      "name": "configured-30-minute-horizon",
      "result": "pass",
      "failures": []
    },
    {
      "name": "configured-120-minute-horizon",
      "result": "pass",
      "failures": []
    },
    {
      "name": "invalid-zero-horizon",
      "result": "pass",
      "failures": []
    },
    {
      "name": "exhausted-pool-unranked",
      "result": "pass",
      "failures": []
    },
    {
      "name": "exhausted-pool-beside-measured",
      "result": "pass",
      "failures": []
    },
    {
      "name": "short-pool-winner-no-same-rule-fallback",
      "result": "pass",
      "failures": []
    },
    {
      "name": "healthy-pool-winner-control",
      "result": "pass",
      "failures": []
    },
    {
      "name": "outdated-quota-axi",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-explicit-floor-safe",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-explicit-floor-exhausted",
      "result": "pass",
      "failures": []
    },
    {
      "name": "single-exhausted-veto",
      "result": "pass",
      "failures": []
    },
    {
      "name": "single-zero-veto",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-zero-unranked",
      "result": "pass",
      "failures": []
    },
    {
      "name": "short-pool-no-cross-rule-or-default-fallback",
      "result": "pass",
      "failures": []
    },
    {
      "name": "short-single-no-same-rule-fallback",
      "result": "pass",
      "failures": []
    },
    {
      "name": "short-nonlimiting-model-bound",
      "result": "pass",
      "failures": []
    },
    {
      "name": "openrouter-absent-alone",
      "result": "pass",
      "failures": []
    },
    {
      "name": "openrouter-absent-beside-measured",
      "result": "pass",
      "failures": []
    },
    {
      "name": "openrouter-credit-only-alone",
      "result": "pass",
      "failures": []
    },
    {
      "name": "openrouter-credit-only-beside-measured",
      "result": "pass",
      "failures": []
    },
    {
      "name": "missing-spend-priority-unranked",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-unknown-warning",
      "result": "pass",
      "failures": []
    },
    {
      "name": "pool-missing-seconds-warning",
      "result": "pass",
      "failures": []
    },
    {
      "name": "schema6-short-pool-winner",
      "result": "pass",
      "failures": []
    },
    {
      "name": "schema6-exhausted-pool",
      "result": "pass",
      "failures": []
    },
    {
      "name": "quota-version-unparseable",
      "result": "pass",
      "failures": []
    },
    {
      "name": "quota-version-failed-read",
      "result": "pass",
      "failures": []
    },
    {
      "name": "minimum-quota-version-control",
      "result": "pass",
      "failures": []
    }
  ]
}
Evidence: Bootstrap task-horizon diagnostics

Source: Bootstrap task-horizon diagnostics

{
  "time": "2026-10-05T16:14:30.252176+00:00",
  "results": [
    {
      "name": "positive-horizon-env-key",
      "command": "FM_BOOTSTRAP_NETWORK=skip FM_BOOTSTRAP_DETECT_ONLY=1 FM_HOME=<isolated> FM_ROOT_OVERRIDE=<isolated> bin/fm-bootstrap.sh",
      "config": {
        "task_horizon_minutes": 90,
        "rules": [
          {
            "when": "A simple bug fix with a stated root cause.",
            "use": {
              "harness": "codex",
              "model": "gpt-6-luna"
            }
          }
        ]
      },
      "optin": "env",
      "exit_code": 0,
      "stdout": "",
      "stderr": "",
      "dispatch_diagnostics": [],
      "expected_dispatch_diagnostics": [],
      "result": "pass"
    },
    {
      "name": "zero-horizon-env-key",
      "command": "FM_BOOTSTRAP_NETWORK=skip FM_BOOTSTRAP_DETECT_ONLY=1 FM_HOME=<isolated> FM_ROOT_OVERRIDE=<isolated> bin/fm-bootstrap.sh",
      "config": {
        "task_horizon_minutes": 0,
        "rules": [
          {
            "when": "A simple bug fix with a stated root cause.",
            "use": {
              "harness": "codex",
              "model": "gpt-6-luna"
            }
          }
        ]
      },
      "optin": "env",
      "exit_code": 0,
      "stdout": "CREW_DISPATCH: invalid config/crew-dispatch.json - task_horizon_minutes must be a positive number when present\n",
      "stderr": "",
      "dispatch_diagnostics": [
        "CREW_DISPATCH: invalid config/crew-dispatch.json - task_horizon_minutes must be a positive number when present"
      ],
      "expected_dispatch_diagnostics": [
        "CREW_DISPATCH: invalid config/crew-dispatch.json - task_horizon_minutes must be a positive number when present"
      ],
      "result": "pass"
    },
    {
      "name": "zero-horizon-home-key",
      "command": "FM_BOOTSTRAP_NETWORK=skip FM_BOOTSTRAP_DETECT_ONLY=1 FM_HOME=<isolated> FM_ROOT_OVERRIDE=<isolated> bin/fm-bootstrap.sh",
      "config": {
        "task_horizon_minutes": 0,
        "rules": [
          {
            "when": "A simple bug fix with a stated root cause.",
            "use": {
              "harness": "codex",
              "model": "gpt-6-luna"
            }
          }
        ]
      },
      "optin": "file",
      "exit_code": 0,
      "stdout": "CREW_DISPATCH: invalid config/crew-dispatch.json - task_horizon_minutes must be a positive number when present\n",
      "stderr": "",
      "dispatch_diagnostics": [
        "CREW_DISPATCH: invalid config/crew-dispatch.json - task_horizon_minutes must be a positive number when present"
      ],
      "expected_dispatch_diagnostics": [
        "CREW_DISPATCH: invalid config/crew-dispatch.json - task_horizon_minutes must be a positive number when present"
      ],
      "result": "pass"
    },
    {
      "name": "zero-horizon-without-optin",
      "command": "FM_BOOTSTRAP_NETWORK=skip FM_BOOTSTRAP_DETECT_ONLY=1 FM_HOME=<isolated> FM_ROOT_OVERRIDE=<isolated> bin/fm-bootstrap.sh",
      "config": {
        "task_horizon_minutes": 0,
        "rules": [
          {
            "when": "A simple bug fix with a stated root cause.",
            "use": {
              "harness": "codex",
              "model": "gpt-6-luna"
            }
          }
        ]
      },
      "optin": "absent",
      "exit_code": 0,
      "stdout": "",
      "stderr": "",
      "dispatch_diagnostics": [],
      "expected_dispatch_diagnostics": [],
      "result": "pass"
    }
  ]
}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 17 of 17 scenarios driven live against the product
Scenario Result Live Evidence
Resolve a single account with 6% remaining and established 1800-second runway: escalate without a profile ✅ pass live Controlled resolver CLI transcripts: report-case-six-percent-1800-established
Resolve an established 3600-second runway under the default horizon: escalate without a profile ✅ pass live Controlled resolver CLI transcripts: established-3600-short
Resolve early or absent-confidence projections: clear single-account and pooled profiles with warnings ✅ pass live Controlled resolver CLI transcripts: single-low-confidence-early, single-low-confidence-None, pool-low-confidence-early, pool-low-confidence-None
Resolve established runway exactly at or beyond the 14400-second horizon: clear without warnings ✅ pass live Controlled resolver CLI transcripts: single-exact-14400-boundary, pool-exact-14400-boundary, single-covering-80796-established, pool-covering-80796-established
Configure 30-, 120-, and zero-minute task horizons: honor valid values and reject zero ✅ pass live Controlled resolver CLI transcripts: configured-30-minute-horizon, configured-120-minute-horizon, invalid-zero-horizon
Resolve an exhausted visible pooled account: keep the pool eligible but unranked with a warning ✅ pass live Controlled resolver CLI transcripts: exhausted-pool-unranked, exhausted-pool-beside-measured, schema6-exhausted-pool
Offer a lower-ranked alternative beside a short pool winner: escalate rather than fall back ✅ pass live Controlled resolver CLI transcripts: short-pool-winner-no-same-rule-fallback, schema6-short-pool-winner, healthy-pool-winner-control
Run with incompatible quota-axi: report the minimum version and perform no ranking ✅ pass live Controlled resolver CLI transcripts: outdated-quota-axi, quota-version-unparseable, quota-version-failed-read, minimum-quota-version-control
Apply an explicit captain profile floor to the pool: veto a known shortfall ✅ pass live Controlled resolver CLI transcripts: pool-explicit-floor-safe, pool-explicit-floor-exhausted
Resolve exhaustion or zero availability outside the pool: veto the single account ✅ pass live Controlled resolver CLI transcripts: single-exhausted-veto, single-zero-veto
Resolve a known zero bound on the visible pooled account: disclose uncertainty without vetoing the pool ✅ pass live Controlled resolver CLI transcripts: pool-zero-unranked
Provide healthy alternatives in another rule or default array: do not downgrade a short matched-rule winner ✅ pass live Controlled resolver CLI transcripts: short-pool-no-cross-rule-or-default-fallback, short-single-no-same-rule-fallback
Add a short exact-model bound beside a healthy limiting rank row: escalate ✅ pass live Controlled resolver CLI transcripts: short-nonlimiting-model-bound
Resolve absent or credit-only OpenRouter coverage: keep it unranked and disclose it beside a measured winner ✅ pass live Controlled resolver CLI transcripts: openrouter-absent-alone, openrouter-credit-only-alone, openrouter-absent-beside-measured, openrouter-credit-only-beside-measured
Resolve missing spendPriority evidence: escalate rather than invent a ranking ✅ pass live Controlled resolver CLI transcripts: missing-spend-priority-unranked
Resolve unknown runway or an established projection without usable seconds: clear with disclosed warnings ✅ pass live Controlled resolver CLI transcripts: pool-unknown-warning, pool-missing-seconds-warning
Run bootstrap with task-horizon configuration: validate only on the typed opt-in path ✅ pass live Bootstrap task-horizon diagnostics
  • bash tests/fm-dispatch-resolve.test.sh with inherited path overrides unset and TMPDIR isolated inside the worktree; targeted resolver regressions passed.
  • python3 .runway-validation/drive.py executed the actual bin/fm-dispatch-resolve.sh against an isolated FM_HOME, controlled quota inputs, and controlled model answers served over localhost through real curl.
  • FM_BOOTSTRAP_NETWORK=skip FM_BOOTSTRAP_DETECT_ONLY=1 FM_HOME=&lt;isolated&gt; FM_ROOT_OVERRIDE=&lt;isolated&gt; bin/fm-bootstrap.sh exercised positive and zero task horizons, environment-key and home-key opt-in, and absent-key behavior.
  • Checked generated CLI output for exactly one unranked-candidate note and absence of quota_summary; retained transcripts and scenario inputs, stopped the disposable HTTP server, and removed the worktree fixtures.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@MrGTV-love
MrGTV-love force-pushed the fm/fm-jev-resolver-runway-guard branch from 24b60f9 to 9e59aa0 Compare October 5, 2026 16:21
@MrGTV-love MrGTV-love changed the title fix(bin): guard typed dispatch clears against short quota runway fix: calibrate typed dispatch runway checks to task duration Oct 5, 2026
@MrGTV-love
MrGTV-love merged commit f9e3a4d into main Oct 5, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant