Skip to content

feat: add spend ledger, cost-aware dispatch, and token ceilings - #42

Merged
yelenplays merged 3 commits into
mainfrom
fm/fm-token-governance-v1
Sep 20, 2026
Merged

yelenplays merged 3 commits into
mainfrom
fm/fm-token-governance-v1

Conversation

@yelenplays

Copy link
Copy Markdown
Owner

Intent

Captain 2026-09-20: we have to fix this for the future. We cannot afford spending so many GPT tokens, and a decision router using Jev together with quota-axi could really fix our usage issues.

Context needed to read the ask. The Codex weekly window reached 86.65 percent consumed with a projection of exhaustion days before its reset, after a week that spent about 427M Codex-plan tokens, concentrated on two days of parallel multi-agent work at high and max effort. The fleet's dispatcher already asks Jev which rule matches a brief and then ranks that rule's candidates from one quota-axi snapshot, so the router exists; what it has no knowledge of is what a task will cost, what it did cost, and nothing caps either. Effort is a static rule attribute and orchestrated pipelines spawn several role sub-agents at maximum effort.

What Changed

  • Add bin/fm-spend-ledger.py to rebuild per-task and fleet token spend from workers' Pi session logs, including nested subagent transcripts, and to predict burn for dispatch.
  • Teach bin/fm-dispatch-resolve.sh to classify a brief's effort via a second Jev Choice (a profile's declared effort is a ceiling; undeclared is xhigh) and refuse candidates whose predicted burn exceeds remaining quota percent or usable runway, showing pred=unknown when ledger evidence is missing.
  • Add bin/fm-procevent-spend.sh to enforce optional config/spend-ceilings.json budgets armed at ship/scout spawn (per-task stop-and-report through fm-control, fleet-window report-once), and make orchestrated Pi roles inherit the task's assessed effort instead of pinning thinking.

Risk Assessment

⚠️ Medium: The authorized fixes landed (unknown/partial ceilings, Pi-only arming, weekly-only calibration, per-task orchestrated effort), but the change remains a large opt-in control plane that can stop agents and refuse dispatch candidates.

Testing

Ran the four targeted public-CLI suites for the ledger, spend ceilings, dispatch resolver, and role provisioner, then drove those same binaries in an isolated home against synthetic Pi sessions plus a live quota-axi Codex snapshot. Nested task cost, empty-is-unknown polling, Codex not-armed, elapsed-budget partial undercounts, complete-scan ceiling fire with control exit, weekly-only predict, and thinking-free role provisioning all passed. Did not launch the opt-in six-role Pi e2e because the live Codex weekly window is already at 0% remaining.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A Pi ship with nested orchestrated sessions reports its real token cost ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/ledger-task-a.json
An empty or unmeasured task does not look like zero spend and is not armed ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/arm-codex.txt
A partial scan undercount cannot clear a ceiling; a complete scan fires it ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/complete-ceiling-poll.out
Cost prediction calibrates only weekly Codex windows ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/predict-live-quota.json
Crossing a task ceiling stops the agent ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/autohandle-state.txt
Orchestrated role definitions inherit task effort instead of pinning max ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/provisioned-roles.json
Dispatch assesses effort per task and refuses predicted over-quota burn ⏸️ untested no Prior payload marked this pass with live=false; that is not a live result. The only cited drive was bash tests/fm-dispatch-resolve.test.sh with the suite's fake TypeSafe curl, not the running produc…
A live orchestrated roster actually launches at the parent task's effort ⏸️ untested no Needs FM_PI_ROLE_AGENTS_LIVE=1, Pi credentials, tmux, and Codex quota. quota-axi currently reports the Codex weekly window at 0% remaining, so enabling that gate would burn the same window this change…
Evidence: Live lab results

Source: Live lab results

PASS: ledger reports nested Pi task cost (3900 tokens, 1 nested, complete)
PASS: ledger empty task is status=empty (tokens field=0, not treated as a complete zero)
PASS: empty ledger poll kept watching (no ceiling/zero fire)
PASS: arm skips unmeasured Codex harness
PASS: arm registers a measured Pi task
PASS: default ceiling poll invokes ledger task without a scan budget
PASS: ledger with an elapsed scan budget reports partial=true undercount (100 of 100099)
PASS: partial undercount did not fire the 5000-token ceiling
PASS: complete poll fires ceiling at 100099 tokens
PASS: predict calibrates weekly codex and skips daily/unmeasured providers
PASS: predict against live quota-axi snapshot calibrates the Codex weekly window (consumed=100.0)
PASS: autohandle delivers control exit and acknowledges the ceiling capture
PASS: provisioned six role definitions omit thinking (task effort is not pinned max)
fails=0
Evidence: Unmeasured Codex harness is not armed

Source: Unmeasured Codex harness is not armed

unmeasured harness codex; not arming spend-task-t-codex

unmeasured harness codex; not arming spend-task-t-codex
Evidence: Complete poll fires the task ceiling

Source: Complete poll fires the task ceiling

spend: spend-task-t-partial status: ceiling task: t-partial observed_tokens: 100099 ceiling_tokens: 5000 condition_polls: 1

spend: spend-task-t-partial
status: ceiling
task: t-partial
observed_tokens: 100099
ceiling_tokens: 5000
condition_polls: 1
Evidence: Autohandle stop-and-report

Source: Autohandle stop-and-report

control_count=1 control_args=t-stop exit failed: spend ceiling crossed - ledger total reached the task ceiling (100099 >= 5000 tokens); agent stopped via fm-control exit

control_count=1
control_args=t-stop exit
--- spend-stop ---
{
  "version": 1,
  "task": "t-stop",
  "spawnGen": "s1789430400.1.st",
  "sequence": 1,
  "observedTokens": 100099,
  "ceilingTokens": 5000,
  "action": "fm-control exit",
  "actionResult": "ok",
  "detail": "exit delivered",
  "at": "2026-09-20T13:04:56Z"
}
--- status ---
failed: spend ceiling crossed - ledger total reached the task ceiling (100099 >= 5000 tokens); agent stopped via fm-control exit
Evidence: Predict skips non-weekly windows

Source: Predict skips non-weekly windows

{
 "anyProvider": {
  "all": {
   "seconds": 300.0,
   "tasks": 1,
   "tokens": 3900
  },
  "high": {
   "seconds": 300.0,
   "tasks": 1,
   "tokens": 3900
  }
 },
 "generatedAt": "2026-09-20T13:04:55.000Z",
 "median": {
  "codex": {
   "all": {
    "seconds": 300.0,
    "tasks": 1,
    "tokens": 3900
   },
   "high": {
    "seconds": 300.0,
    "tasks": 1,
    "tokens": 3900
   }
  }
 },
 "partial": false,
 "providers": {
  "codex": {
   "percentConsumed": 60.0,
   "tokensPerPoint": 56.666666666666664,
   "windowKind": "weekly",
   "windowStart": "2026-09-14T13:04:52.000Z",
   "windowTokens": 3400
  }
 },
 "scanSeconds": 0.001,
 "status": "ok",
 "version": 1
}
Evidence: Predict against live quota-axi Codex weekly window

Source: Predict against live quota-axi Codex weekly window

{
 "anyProvider": {
  "all": {
   "seconds": 300.0,
   "tasks": 1,
   "tokens": 3900
  },
  "high": {
   "seconds": 300.0,
   "tasks": 1,
   "tokens": 3900
  }
 },
 "generatedAt": "2026-09-20T13:04:56.000Z",
 "median": {
  "codex": {
   "all": {
    "seconds": 300.0,
    "tasks": 1,
    "tokens": 3900
   },
   "high": {
    "seconds": 300.0,
    "tasks": 1,
    "tokens": 3900
   }
  }
 },
 "partial": false,
 "providers": {
  "codex": {
   "percentConsumed": 100.0,
   "windowKind": "weekly",
   "windowStart": "2026-09-19T11:27:22.000Z",
   "windowTokens": 0
  }
 },
 "scanSeconds": 0.001,
 "status": "ok",
 "version": 1
}
Evidence: Live quota-axi Codex snapshot (weekly exhausted)

Source: Live quota-axi Codex snapshot (weekly exhausted)

{
  "generatedAt": "2026-09-20T13:04:55.554Z",
  "schemaVersion": 5,
  "providers": [
    {
      "provider": "codex",
      "plan": "prolite",
      "windows": [
        {
          "id": "weekly",
          "label": "week",
          "kind": "weekly",
          "resetsAt": "2026-09-26T11:27:22.000Z",
          "percentRemaining": 0,
          "pace": {
            "status": "ahead",
            "reservePercentPoints": -84.7464,
            "burnMultiple": 6.5558
          }
        }
      ],
      "credits": {
        "remaining": 0,
        "unlimited": false,
        "unit": "credits"
      },
      "state": {
        "status": "fresh",
        "stale": false
      },
      "quotaSemantics": {
        "status": "known",
        "effectiveAvailability": [
          {
            "scope": "all_models",
            "status": "known",
            "effectivePercentRemaining": 0,
            "boundedBy": [
              "weekly"
            ],
            "limitingWindowIds": [
              "weekly"
            ],
            "pace": {
              "status": "ahead",
              "aheadWindowIds": [
                "weekly"
              ],
              "worstReservePercentPoints": -84.7464,
              "worstReserveWindowId": "weekly"
            },
            "runway": {
              "status": "exhausted_now",
              "usableRunwaySeconds": 0,
              "limitingWindowId": "weekly",
              "projectedExhaustedAt": "2026-09-20T13:04:55.554Z"
            },
            "selection": {
              "status": "known",
              "spendPriority": -6.5558
            }
          }
        ]
      }
    }
  ]
}
Evidence: Default ceiling poll has no scan budget

Source: Default ceiling poll has no scan budget

--state /var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T//fm-token-gov-live.MdugXK/home/state task t-argv
--state /var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T//fm-token-gov-live.MdugXK/home/state task t-argv
--state /var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T//fm-token-gov-live.MdugXK/home/state task t-argv
--state /var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T//fm-token-gov-live.MdugXK/home/state task t-argv
Evidence: Provisioned roles omit thinking

Source: Provisioned roles omit thinking

{
  "explorer": {
    "path": "/var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T/fm-token-gov-live.MdugXK/pi-config/agents/fm-orchestrated-explorer.md",
    "name": "fm-orchestrated-explorer",
    "model": "openai-codex/gpt-5.6-luna",
    "has_thinking": false,
    "thinking": null,
    "session-mode": "standalone"
  },
  "researcher": {
    "path": "/var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T/fm-token-gov-live.MdugXK/pi-config/agents/fm-orchestrated-researcher.md",
    "name": "fm-orchestrated-researcher",
    "model": "openai-codex/gpt-5.6-luna",
    "has_thinking": false,
    "thinking": null,
    "session-mode": "standalone"
  },
  "worker": {
    "path": "/var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T/fm-token-gov-live.MdugXK/pi-config/agents/fm-orchestrated-worker.md",
    "name": "fm-orchestrated-worker",
    "model": "openai-codex/gpt-5.6-luna",
    "has_thinking": false,
    "thinking": null,
    "session-mode": "standalone"
  },
  "tester": {
    "path": "/var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T/fm-token-gov-live.MdugXK/pi-config/agents/fm-orchestrated-tester.md",
    "name": "fm-orchestrated-tester",
    "model": "openai-codex/gpt-5.6-luna",
    "has_thinking": false,
    "thinking": null,
    "session-mode": "standalone"
  },
  "reviewer": {
    "path": "/var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T/fm-token-gov-live.MdugXK/pi-config/agents/fm-orchestrated-reviewer.md",
    "name": "fm-orchestrated-reviewer",
    "model": "openai-codex/gpt-6-astra",
    "has_thinking": false,
    "thinking": null,
    "session-mode": "standalone"
  },
  "integrator": {
    "path": "/var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T/fm-token-gov-live.MdugXK/pi-config/agents/fm-orchestrated-integrator.md",
    "name": "fm-orchestrated-integrator",
    "model": "openai-codex/gpt-6-astra",
    "has_thinking": false,
    "thinking": null,
    "session-mode": "standalone"
  }
}
Evidence: Partial undercount ledger document

Source: Partial undercount ledger document

{
 "byEffort": {
  "medium": {
   "cacheRead": 0,
   "cacheWrite": 0,
   "cost": 0.01,
   "input": 100,
   "messages": 1,
   "output": 10,
   "pricedTokens": 100,
   "reasoning": 0,
   "tokens": 100,
   "unpricedTokens": 0
  }
 },
 "byLane": {
  "codex": {
   "cacheRead": 0,
   "cacheWrite": 0,
   "cost": 0.01,
   "input": 100,
   "messages": 1,
   "output": 10,
   "pricedTokens": 100,
   "reasoning": 0,
   "tokens": 100,
   "unpricedTokens": 0
  }
 },
 "byModel": {
  "openai-codex/gpt-6-astra": {
   "cacheRead": 0,
   "cacheWrite": 0,
   "cost": 0.01,
   "input": 100,
   "messages": 1,
   "output": 10,
   "pricedTokens": 100,
   "reasoning": 0,
   "tokens": 100,
   "unpricedTokens": 0
  }
 },
 "byProvider": {
  "openai-codex": {
   "cacheRead": 0,
   "cacheWrite": 0,
   "cost": 0.01,
   "input": 100,
   "messages": 1,
   "output": 10,
   "pricedTokens": 100,
   "reasoning": 0,
   "tokens": 100,
   "unpricedTokens": 0
  }
 },
 "generatedAt": "2026-09-20T13:04:54.000Z",
 "harness": "pi",
 "partial": true,
 "sessions": [
  {
   "file": "/var/folders/9d/8w50jhgd79x63rgbq_5cyyvm0000gn/T/fm-token-gov-live.MdugXK/partial-sessions/--work-partial--/2026-09-16T11-00-00-000Z_p1111111-0000-0000-0000-000000000001.jsonl",
   "id": "p1111111-0000-0000-0000-000000000001",
   "nested": false,
   "start": "2026-09-16T11:00:00.000Z",
   "tokens": 100
  }
 ],
 "spawnEpoch": 1789430400,
 "status": "ok",
 "task": "t-partial",
 "totals": {
  "cacheRead": 0,
  "cacheWrite": 0,
  "cost": 0.01,
  "costStatus": "known",
  "input": 100,
  "messages": 1,
  "nestedSessions": 0,
  "output": 10,
  "pricedTokens": 100,
  "reasoning": 0,
  "sessions": 1,
  "tokens": 100,
  "unpricedTokens": 0
 },
 "version": 1,
 "worktree": "/work/partial"
}
Evidence: Empty task ledger document

Source: Empty task ledger document

{
 "byEffort": {},
 "byLane": {},
 "byModel": {},
 "byProvider": {},
 "generatedAt": "2026-09-20T13:04:52.000Z",
 "harness": "pi",
 "partial": false,
 "reason": "no Pi session files bound to this task yet",
 "sessions": [],
 "spawnEpoch": 1789430400,
 "status": "empty",
 "task": "task-empty",
 "totals": {
  "cacheRead": 0,
  "cacheWrite": 0,
  "cost": 0.0,
  "costStatus": "none",
  "input": 0,
  "messages": 0,
  "nestedSessions": 0,
  "output": 0,
  "pricedTokens": 0,
  "reasoning": 0,
  "sessions": 0,
  "tokens": 0,
  "unpricedTokens": 0
 },
 "version": 1,
 "worktree": "/work/empty"
}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 4 issues found → auto-fixed ✅
  • ⚠️ .agents/skills/orchestrated-delivery/agents/fm-orchestrated-worker.md:5 - Intent names the live spend cause: 'Effort is a static rule attribute and orchestrated pipelines spawn several role sub-agents at maximum effort.' Dispatch now classifies parent effort, but fm-orchestrated-worker/tester/explorer still declare thinking: max on openai-codex. Concrete path: Jev assesses medium, fm-spawn emits --effort medium, the parent then follows orchestrated-delivery and still launches those roles at max; the classifier never reaches them. Nested tokens are only counted after the fact (and only if a ceiling is configured). The omitted change is lowering or inheriting that role effort; leaving max in place keeps the named GPT-burn path.
  • ⚠️ bin/fm-procevent-spend.sh:151 - ledger_task_tokens treats status empty as 0 and arms every ship/scout regardless of harness. The ledger only reads Pi session JSONL (openai-codex -> codex). A clear dispatch to harness=codex (docs/examples/crew-dispatch.json GPT profile) therefore arms spend-task-<id>, polls 0 forever, and never delivers fm-control exit while Codex CLI tokens accrue. That is a reachable unmeasured path for the same Codex window the intent wants capped. Pi-only measurement looks deliberate; whether that is enough, or non-Pi ships must not arm / must be measured, needs authorization.
  • ⚠️ bin/fm-procevent-spend.sh:151 - The same jq accepts status ok without checking .partial, and line 149 always passes --scan-budget 30 (ledger task default is also 30 at bin/fm-spend-ledger.py:850). refresh_cache then leaves changed files unscanned when the budget elapses and still writes a numeric total. Ceiling poll compares that undercount to the cap, so a crossed bound can fail to fire while growing session files keep consuming the budget. Do not treat a partial task document as a complete total; run ceiling polls without a scan budget (or keep polling without counting a partial read as under-ceiling).
  • ⚠️ bin/fm-spend-ledger.py:747 - cmd_predict calibrates non-weekly quota windows from pace.burnMultiple, but the module header says only weekly windows are calibrated and other kinds stay unmeasured. The intent is the Codex weekly window. This extra matching path is not required; a wrong derived windowStart would also feed a bad tokensPerPoint into dispatch burn gates. Smallest honest remedy is to remove the non-weekly branch and leave those providers unmeasured, matching the stated contract.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A Pi ship with nested orchestrated sessions reports its real token cost ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/ledger-task-a.json
An empty or unmeasured task does not look like zero spend and is not armed ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/arm-codex.txt
A partial scan undercount cannot clear a ceiling; a complete scan fires it ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/complete-ceiling-poll.out
Cost prediction calibrates only weekly Codex windows ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/predict-live-quota.json
Crossing a task ceiling stops the agent ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/autohandle-state.txt
Orchestrated role definitions inherit task effort instead of pinning max ✅ pass live ~/.no-mistakes/evidence/01M2ZAZ1YHFX1Q25BX9EN05MVS/provisioned-roles.json
Dispatch assesses effort per task and refuses predicted over-quota burn ⏸️ untested no Prior payload marked this pass with live=false; that is not a live result. The only cited drive was bash tests/fm-dispatch-resolve.test.sh with the suite's fake TypeSafe curl, not the running produc…
A live orchestrated roster actually launches at the parent task's effort ⏸️ untested no Needs FM_PI_ROLE_AGENTS_LIVE=1, Pi credentials, tmux, and Codex quota. quota-axi currently reports the Codex weekly window at 0% remaining, so enabling that gate would burn the same window this change…
  • bash tests/fm-spend-ledger.test.sh
  • bash tests/fm-procevent-spend.test.sh
  • bash tests/fm-dispatch-resolve.test.sh
  • bash tests/fm-pi-role-agents.test.sh
  • quota-axi --json --provider codex
  • isolated live lab driving bin/fm-spend-ledger.py, bin/fm-procevent-spend.sh, and bin/fm-pi-role-agents.py (evidence live-lab.sh / live-lab.log)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…dispatch, ceilings

bin/fm-spend-ledger.py rebuilds per-task and fleet token spend from the
workers' own Pi session logs, including nested subagent transcripts, with
per-record day bucketing and priced/unpriced cost honesty.

bin/fm-dispatch-resolve.sh asks Jev a typed effort question over the brief:
a profile's declared effort becomes a ceiling the assessed class may not
exceed, and predicted burn from the ledger is gated against quota-axi
remaining percent and usable runway, with the prediction named in refusals.

bin/fm-procevent-spend.sh enforces config/spend-ceilings.json: a per-task
ceiling stop-and-reports through fm-control exit plus a status line, and a
fleet-window ceiling reports once per window; both are armed at spawn.

Generated with [Devin](https://devin.ai)
@yelenplays
yelenplays merged commit 4a7d9c6 into main Sep 20, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant