feat: add spend ledger, cost-aware dispatch, and token ceilings - #42
Merged
Merged
Conversation
…dispatch, ceilings bin/fm-spend-ledger.py rebuilds per-task and fleet token spend from the workers' own Pi session logs, including nested subagent transcripts, with per-record day bucketing and priced/unpriced cost honesty. bin/fm-dispatch-resolve.sh asks Jev a typed effort question over the brief: a profile's declared effort becomes a ceiling the assessed class may not exceed, and predicted burn from the ledger is gated against quota-axi remaining percent and usable runway, with the prediction named in refusals. bin/fm-procevent-spend.sh enforces config/spend-ceilings.json: a per-task ceiling stop-and-reports through fm-control exit plus a status line, and a fleet-window ceiling reports once per window; both are armed at spawn. Generated with [Devin](https://devin.ai)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Captain 2026-09-20: we have to fix this for the future. We cannot afford spending so many GPT tokens, and a decision router using Jev together with quota-axi could really fix our usage issues.
Context needed to read the ask. The Codex weekly window reached 86.65 percent consumed with a projection of exhaustion days before its reset, after a week that spent about 427M Codex-plan tokens, concentrated on two days of parallel multi-agent work at high and max effort. The fleet's dispatcher already asks Jev which rule matches a brief and then ranks that rule's candidates from one quota-axi snapshot, so the router exists; what it has no knowledge of is what a task will cost, what it did cost, and nothing caps either. Effort is a static rule attribute and orchestrated pipelines spawn several role sub-agents at maximum effort.
What Changed
bin/fm-spend-ledger.pyto rebuild per-task and fleet token spend from workers' Pi session logs, including nested subagent transcripts, and to predict burn for dispatch.bin/fm-dispatch-resolve.shto classify a brief's effort via a second Jev Choice (a profile's declared effort is a ceiling; undeclared isxhigh) and refuse candidates whose predicted burn exceeds remaining quota percent or usable runway, showingpred=unknownwhen ledger evidence is missing.bin/fm-procevent-spend.shto enforce optionalconfig/spend-ceilings.jsonbudgets armed at ship/scout spawn (per-task stop-and-report throughfm-control, fleet-window report-once), and make orchestrated Pi roles inherit the task's assessed effort instead of pinningthinking.Risk Assessment
Testing
Ran the four targeted public-CLI suites for the ledger, spend ceilings, dispatch resolver, and role provisioner, then drove those same binaries in an isolated home against synthetic Pi sessions plus a live quota-axi Codex snapshot. Nested task cost, empty-is-unknown polling, Codex not-armed, elapsed-budget partial undercounts, complete-scan ceiling fire with control exit, weekly-only predict, and thinking-free role provisioning all passed. Did not launch the opt-in six-role Pi e2e because the live Codex weekly window is already at 0% remaining.
bash tests/fm-dispatch-resolve.test.shwith the suite's fake TypeSafe curl, not the running produc…Evidence: Live lab results
Source: Live lab results
Evidence: Unmeasured Codex harness is not armed
Source: Unmeasured Codex harness is not armed
unmeasured harness codex; not arming spend-task-t-codexEvidence: Complete poll fires the task ceiling
Source: Complete poll fires the task ceiling
spend: spend-task-t-partial status: ceiling task: t-partial observed_tokens: 100099 ceiling_tokens: 5000 condition_polls: 1Evidence: Autohandle stop-and-report
Source: Autohandle stop-and-report
control_count=1 control_args=t-stop exit failed: spend ceiling crossed - ledger total reached the task ceiling (100099 >= 5000 tokens); agent stopped via fm-control exitEvidence: Predict skips non-weekly windows
Source: Predict skips non-weekly windows
Evidence: Predict against live quota-axi Codex weekly window
Source: Predict against live quota-axi Codex weekly window
Evidence: Live quota-axi Codex snapshot (weekly exhausted)
Source: Live quota-axi Codex snapshot (weekly exhausted)
Evidence: Default ceiling poll has no scan budget
Source: Default ceiling poll has no scan budget
Evidence: Provisioned roles omit thinking
Source: Provisioned roles omit thinking
Evidence: Partial undercount ledger document
Source: Partial undercount ledger document
Evidence: Empty task ledger document
Source: Empty task ledger document
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 4 issues found → auto-fixed ✅
.agents/skills/orchestrated-delivery/agents/fm-orchestrated-worker.md:5- Intent names the live spend cause: 'Effort is a static rule attribute and orchestrated pipelines spawn several role sub-agents at maximum effort.' Dispatch now classifies parent effort, but fm-orchestrated-worker/tester/explorer still declare thinking: max on openai-codex. Concrete path: Jev assesses medium, fm-spawn emits --effort medium, the parent then follows orchestrated-delivery and still launches those roles at max; the classifier never reaches them. Nested tokens are only counted after the fact (and only if a ceiling is configured). The omitted change is lowering or inheriting that role effort; leaving max in place keeps the named GPT-burn path.bin/fm-procevent-spend.sh:151- ledger_task_tokens treats status empty as 0 and arms every ship/scout regardless of harness. The ledger only reads Pi session JSONL (openai-codex -> codex). A clear dispatch to harness=codex (docs/examples/crew-dispatch.json GPT profile) therefore arms spend-task-<id>, polls 0 forever, and never delivers fm-control exit while Codex CLI tokens accrue. That is a reachable unmeasured path for the same Codex window the intent wants capped. Pi-only measurement looks deliberate; whether that is enough, or non-Pi ships must not arm / must be measured, needs authorization.bin/fm-procevent-spend.sh:151- The same jq accepts status ok without checking .partial, and line 149 always passes --scan-budget 30 (ledger task default is also 30 at bin/fm-spend-ledger.py:850). refresh_cache then leaves changed files unscanned when the budget elapses and still writes a numeric total. Ceiling poll compares that undercount to the cap, so a crossed bound can fail to fire while growing session files keep consuming the budget. Do not treat a partial task document as a complete total; run ceiling polls without a scan budget (or keep polling without counting a partial read as under-ceiling).bin/fm-spend-ledger.py:747- cmd_predict calibrates non-weekly quota windows from pace.burnMultiple, but the module header says only weekly windows are calibrated and other kinds stay unmeasured. The intent is the Codex weekly window. This extra matching path is not required; a wrong derived windowStart would also feed a bad tokensPerPoint into dispatch burn gates. Smallest honest remedy is to remove the non-weekly branch and leave those providers unmeasured, matching the stated contract.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-dispatch-resolve.test.shwith the suite's fake TypeSafe curl, not the running produc…bash tests/fm-spend-ledger.test.shbash tests/fm-procevent-spend.test.shbash tests/fm-dispatch-resolve.test.shbash tests/fm-pi-role-agents.test.shquota-axi --json --provider codexisolated live lab drivingbin/fm-spend-ledger.py,bin/fm-procevent-spend.sh, andbin/fm-pi-role-agents.py(evidencelive-lab.sh/live-lab.log)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.