Repository navigation
feat: alarm on the OpenRouter balance before it hits zero (issue #1411) - #1731
Conversation
Every paid route in deploy/litellm/config.yaml bills against one OpenRouter account, and that file declares no chat fallback on purpose, so at zero those routes refuse outright rather than degrading. That account has been running down for weeks and nothing read the balance on any schedule. The free pool health report probes only free endpoints, which the balance does not affect, and ci.yml's out of credit classifier only runs after a job has already failed. The first available signal was a paid request failing in front of an audience. Adds .github/ci/check-provider-balance.mjs, a scheduled check that reads /api/v1/credits and /api/v1/key, and .github/workflows/provider-balance-watch.yml, which runs it every six hours. The threshold is days of runway rather than a dollar amount, because a fixed dollar floor is meaningless without the spend rate. The burn rate comes from two sources and the larger one wins: the provider's own weekly usage figure, which is available on the very first run and answers what to do with no stored sample, and the account wide delta against a cached sample, which reacts to a spike within one run interval instead of diluting it across a week. A small absolute floor sits underneath as a backstop, because a burn rate measured over a quiet window says nothing about the next busy hour and at zero measured burn the runway is arithmetically infinite. Every failure to read the number exits non zero: an HTTP error, an unparseable body, a body missing the fields, an absent credential, a misconfigured threshold. A monitor that reports green when it could not reach the API is worse than no monitor, because it asserts a state nothing verified. The alarm's primary channel is a deduplicated GitHub issue rather than a red square in the Actions tab, which nobody is subscribed to. It is filed when none is open and closed on the first healthy run after it, so a stale open issue cannot swallow the next alert, which is the defect issue #1416 records.
|
Warning Review limit reachedNext included review available in 43 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Team Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (5)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Three findings from the plain adversarial pass, all in the alerting half rather than in the arithmetic. The tracking issue title claimed the balance was running out, but the same issue is filed when the balance could not be read at all, which is a different problem with a different remedy. Retitled to cover both states, and the body now says which one to read the block as and what to check when it is the second. The cache save step could turn a run red on its own. On a run that cannot reach the API the script exits before writing the sample, and with no cache to restore there is no path to save, which actions/cache/save reports as an error. That painted a red square whose cause had nothing to do with the balance, on top of the alarm that had already fired correctly. Marked continue-on-error, since losing a sample only degrades the check to the provider's weekly figure. The checker's header now records that the provider's weekly usage window is not documented as rolling or as calendar anchored, which matters in one direction: if it resets on a fixed weekday, that source under-reads for about a day afterwards. The account wide delta is what covers it, which is a second reason not to rely on the provider's figure alone once a sample exists.
Adversarial review, pipeline modeTwo streams run. Neither is skipped. Stream 1: CodeRabbit CLI (
|
The shortfall label proves the alarm fires and files its tracking issue. The close half needs proving too, and it cannot be reached by removing that label, because the real balance is genuinely under the floor today and the check would correctly stay red. A second label lowers the thresholds under the real balance instead, so the recovery run measures the same real account and demonstrates that the issue gets closed rather than left open to swallow the next alert.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Proof of delivery, on the real account and the real Actions runnerThe check itself was already proved locally against the live endpoint (
Issue #1732 is the artefact: filed by Log lines, verbatim: Three details worth naming, since each of them is a way this could have shipped broken and looked fine:
The test labels have been removed from this pull request, so no further balance runs fire on it. |
Second bug-log reconciliation batch of the day. The first batch (#1743, merged 2026-09-02T19:44:04Z) appended 197 entries from 167 PRs, taking `.wolf/buglog.jsonl` on `main` to 511 lines. This batch sweeps every PR merged after that point which carried a `## Buglog entry` heading in its body, appending 14 entries from 7 PRs: - #1733 (1 entry) - #1735 (1 entry) - #1739 (6 entries) - #1740 (1 entry) - #1748 (1 entry) - #1749 (2 entries) - #1756 (2 entries) Checked and excluded: - #1727 carries no buglog entry. It is a docs/process PR (tracking-discipline rule), not a bug fix, and its body mentions `.wolf/buglog.jsonl` only in passing prose. - #1715, #1729, #1731 and #1734 merged before this batch's window and are already present in the first batch (#1743). Verified by id/error_message lookup against the 511 lines already on `main`. Every entry was extracted from its source PR body, parsed as JSON to confirm it is well-formed, and checked for the required `error_message`, `root_cause`, `fix` and `tags` fields (all present, none reconstructed). No duplicates were found against the existing 511 lines or within this batch, checked by both `id` and exact `error_message` match. Diff is exactly one file, 14 insertions, 0 deletions. The first 511 lines byte-match `main`'s current copy (verified with `diff` against `git show origin/main:.wolf/buglog.jsonl`). This PR was not opened on a fix or feature branch, per `.claude/rules/openwolf.md`: it is the dedicated buglog-only PR, branched directly from `main`, diffing only `.wolf/buglog.jsonl`. Refs #873 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Closes #1411
What was wrong
Every paid route in
deploy/litellm/config.yamlbills against one OpenRouter account, and that file declares no chat fallback on purpose (one alias, one price), so at zero balance those routes refuse outright rather than degrading to another model. That account has been running down for weeks: 1.59 USD of 10 purchased when the issue was filed, 1.46 four days later.Nothing read the number on any schedule.
scripts/report-free-pool-health.pyprobes only free endpoints, which the balance does not affect. The out of credit classifier inci.ymlis by construction downstream of a job that has already gone red. So the first available signal was a paid request failing in front of whoever was watching.What this adds
.github/ci/check-provider-balance.mjs, run by.github/workflows/provider-balance-watch.ymlon a0 */6 * * *schedule, so every six hours on a hosted runner. It readsGET /api/v1/creditsandGET /api/v1/keywith theOPENROUTER_API_KEYrepository secret, computes remaining balance and days of runway, and exits non zero when either threshold is breached or when the number could not be read at all.Six hours is chosen against the quantity being measured. The threshold is denominated in days, so a sampling interval a quarter of a day long cannot let the account cross from above the threshold to empty between two samples unless the burn rate rises by more than an order of magnitude inside one interval, and the stored account wide delta would catch even that on the following run.
The runway calculation
Two sources, and the larger one wins, deliberately. They disagree exactly when a quiet week is followed by a busy day, and in that case the weekly average is the dangerous number: it would report a comfortable runway right through the demo that empties the account. Taking the max means the alarm can fire early and cannot fire late.
The first run, with no previous sample. Candidate B is unavailable, candidate A is not, so the check still produces a real runway from the provider's own weekly figure and says in its output which source it used and why the other was missing. It does not report green for want of a comparison, and it does not fail for want of history either. The same path covers a cache eviction and the case where the newest stored sample is under thirty minutes old, which is too close to extrapolate from.
Thresholds.
PROVIDER_BALANCE_MIN_RUNWAY_DAYSdefaults to 7, which is the primary gate: a week is enough notice for a human with a card, across a weekend.PROVIDER_BALANCE_MIN_USDdefaults to 2 and is a backstop, not the threshold, because a burn rate measured over a quiet window says nothing about the next busy hour and at zero measured burn the runway is arithmetically infinite. Stated plainly, since it matters for reading the proof below: at the burn this account has actually averaged, 0.004 USD per day, the runway test alone would not fire today. The floor is what fires. Both are needed and neither is sufficient.Persistence. The account wide sample rides in the Actions cache with a rolling key. It is best effort by design: an eviction degrades the check to the weekly figure, which is still a runway number, rather than taking the alarm offline.
Unreadable is never healthy. An HTTP error, an unparseable body, a body missing the fields, an absent credential and a misconfigured threshold all exit non zero. A monitor that reports green when it could not reach the API is worse than no monitor, because it actively asserts a state nothing verified. There is a three attempt retry ladder so one transient blip does not open a critical issue, and exhausting it is still fatal.
How anyone would actually notice it firing
This is the part the issue is really about, so it is answered rather than assumed. A red square in the Actions tab is not a notification: nobody is subscribed to a workflow they did not trigger, and this repository already has a documented case of a monitor sitting quietly red for days. So:
priority:critical,money-pathanddemo-surfaceso it lands in the same filters the demo hotfix work is already read through. GitHub emails it to everyone watching the repository. Deduplication matches on title and author, so a public repository stranger cannot suppress the channel by opening an issue with the same title.Other provider balances, probed live rather than assumed
OPENROUTER_API_KEYandGROQ_API_KEYare the only provider credentials in the live configuration. Groq is the obvious second one and the answer is that it has nothing equivalent to read:GET api.groq.com/openai/v1/modelsGET api.groq.com/openai/v1/organizationsunknown_urlGET api.groq.com/v1/creditsunknown_urlGroq publishes no balance, credit or usage endpoint on the inference API. Its spend is visible only in the console. This is worth stating rather than leaving implicit, because an alarm on one provider and silence on the rest is a false sense of coverage. The mitigating fact is that after the 2026-08-23 restructure Groq serves audio only in
config.yaml(sections 6 and 7); no chat route is on it.Can the free pool carry chat at zero balance, and is a fallback safe
Investigated, reported, and deliberately not built here, per the brief: a routing change under time pressure is how a demo breaks worse.
hive-default,hive-auto,hive-small,hive-mediumandhive-fastall resolve todots-studio/dots-3-note-preview:free, andHIVE_TOOLS_MODELdefaults tohive-free, not to a paid model. The issue's framing that "chat stops answering" is stale on that point; what stops is the paid set.litellm_settingswere deliberately removed. A customer charged the DeepSeek price and served a free model is a worse failure than a refusal. Not proposed.Proof, against the real endpoint
Committed at
docs/proof/provider-balance-alarm-2026-09-02/check-output.log. Four cases, all against the live account, no key printed.Case 3 is the delta path measured against a real historical reading, not a synthetic one: the 8.409778422 figure is the number recorded in the issue on 2026-08-29, and 0.030 USD per day is the account's genuine four day burn.
The alarm's own delivery is proved separately, on the real Actions runner rather than reasoned about, in the proof comment on this pull request: run 33662109196 filed issue #1732, run 33662109465 proved the deduplication, and run 33662230066 closed it again. Both halves of the alert channel, not just the detection logic.
Tests
node .github/ci/check-provider-balance.test.mjs, wired into the requiredlintjob inci.ymlnext to the other guards of this shape. Twenty one cases, all offline: the suite serves its own stub of the provider API. They cover the threshold boundary on both sides, a first run with no stored sample, the spike case where the stored sample must beat the weekly average, sample pruning, a future or backwards sample, the absolute floor, an HTTP 500, a 401 with no credential echoed, a malformed body, a body missing the fields, an unreachable API, an absent credential, a misconfigured threshold and a corrupt state file. The last three assert the wiring itself: that the scheduled lane still invokes the check, still persists the sample, and still both files and closes its tracking issue, so deleting any half of the alarm turns a required check red.Buglog entry
To be appended to
.wolf/buglog.jsonlon main in a separate buglog only pull request after this merges, per.claude/rules/openwolf.md.{"id":"1411-openrouter-balance-unobserved","date":"2026-09-02","title":"The OpenRouter account every paid route bills against ran down to 1.46 USD with nothing reading the balance","error_message":"No error was raised at all, which is the defect. GET https://openrouter.ai/api/v1/credits returned total_credits 10 and total_usage 8.540654392, so 1.46 USD remained, and the first signal available to anyone would have been a paid request failing with a 402 in front of a demo audience.","root_cause":"Nothing in the repository read the provider balance on any schedule. scripts/report-free-pool-health.py probes only free endpoints, whose throughput OpenRouter gates on credits purchased all time rather than on the current balance, so it is unaffected by depletion and cannot see it. The out of credit classifier in .github/workflows/ci.yml is by construction downstream of a job that has already failed. The balance was readable the whole time through a single unauthenticated-to-us GET that bills no inference, and no code path read it.","fix":"Added .github/ci/check-provider-balance.mjs and .github/workflows/provider-balance-watch.yml, a six hourly check that computes days of runway from the larger of the provider's own weekly usage figure and an account wide usage delta carried between runs in the Actions cache, fails on a runway under seven days or a balance under a two dollar backstop floor, fails loudly on any inability to read the number rather than reporting healthy, and files a deduplicated priority:critical GitHub issue that it closes itself on recovery. Regression guard .github/ci/check-provider-balance.test.mjs is wired into the required lint job and asserts the wiring as well as the arithmetic.","tags":["monitoring","money-path","openrouter","silent-absence","github-actions","demo-surface"]}