Skip to content

feat: alarm on the OpenRouter balance before it hits zero (issue #1411) - #1731

Merged
sakibsadmanshajib merged 4 commits into
mainfrom
fix/1411-provider-balance-alarm
Sep 2, 2026
Merged

sakibsadmanshajib merged 4 commits into
mainfrom
fix/1411-provider-balance-alarm

Conversation

@sakibsadmanshajib

@sakibsadmanshajib sakibsadmanshajib commented Sep 2, 2026 •

Copy link
Copy Markdown
Owner

Closes #1411

What was wrong

Every paid route in deploy/litellm/config.yaml bills against one OpenRouter account, and that file declares no chat fallback on purpose (one alias, one price), so at zero balance those routes refuse outright rather than degrading to another model. That account has been running down for weeks: 1.59 USD of 10 purchased when the issue was filed, 1.46 four days later.

Nothing read the number on any schedule. scripts/report-free-pool-health.py probes only free endpoints, which the balance does not affect. The out of credit classifier in ci.yml is by construction downstream of a job that has already gone red. So the first available signal was a paid request failing in front of whoever was watching.

What this adds

.github/ci/check-provider-balance.mjs, run by .github/workflows/provider-balance-watch.yml on a 0 */6 * * * schedule, so every six hours on a hosted runner. It reads GET /api/v1/credits and GET /api/v1/key with the OPENROUTER_API_KEY repository secret, computes remaining balance and days of runway, and exits non zero when either threshold is breached or when the number could not be read at all.

Six hours is chosen against the quantity being measured. The threshold is denominated in days, so a sampling interval a quarter of a day long cannot let the account cross from above the threshold to empty between two samples unless the burn rate rises by more than an order of magnitude inside one interval, and the stored account wide delta would catch even that on the following run.

The runway calculation

remaining = total_credits - total_usage          (from /api/v1/credits)

burn candidate A = usage_weekly / 7              (from /api/v1/key, key scoped)
burn candidate B = (total_usage_now - total_usage_oldest_sample)
                   / elapsed_days                (account wide, from the cached sample)

burn    = max(A, B)
runway  = remaining / burn        days

Two sources, and the larger one wins, deliberately. They disagree exactly when a quiet week is followed by a busy day, and in that case the weekly average is the dangerous number: it would report a comfortable runway right through the demo that empties the account. Taking the max means the alarm can fire early and cannot fire late.

The first run, with no previous sample. Candidate B is unavailable, candidate A is not, so the check still produces a real runway from the provider's own weekly figure and says in its output which source it used and why the other was missing. It does not report green for want of a comparison, and it does not fail for want of history either. The same path covers a cache eviction and the case where the newest stored sample is under thirty minutes old, which is too close to extrapolate from.

Thresholds. PROVIDER_BALANCE_MIN_RUNWAY_DAYS defaults to 7, which is the primary gate: a week is enough notice for a human with a card, across a weekend. PROVIDER_BALANCE_MIN_USD defaults to 2 and is a backstop, not the threshold, because a burn rate measured over a quiet window says nothing about the next busy hour and at zero measured burn the runway is arithmetically infinite. Stated plainly, since it matters for reading the proof below: at the burn this account has actually averaged, 0.004 USD per day, the runway test alone would not fire today. The floor is what fires. Both are needed and neither is sufficient.

Persistence. The account wide sample rides in the Actions cache with a rolling key. It is best effort by design: an eviction degrades the check to the weekly figure, which is still a runway number, rather than taking the alarm offline.

Unreadable is never healthy. An HTTP error, an unparseable body, a body missing the fields, an absent credential and a misconfigured threshold all exit non zero. A monitor that reports green when it could not reach the API is worse than no monitor, because it actively asserts a state nothing verified. There is a three attempt retry ladder so one transient blip does not open a critical issue, and exhausting it is still fatal.

How anyone would actually notice it firing

This is the part the issue is really about, so it is answered rather than assumed. A red square in the Actions tab is not a notification: nobody is subscribed to a workflow they did not trigger, and this repository already has a documented case of a monitor sitting quietly red for days. So:

  1. A deduplicated GitHub issue is the primary channel. Titled "the OpenRouter balance is low or unreadable", which covers both states it reports, since a balance that could not be READ at all is filed here too and has a different remedy. Labelled priority:critical, money-path and demo-surface so it lands in the same filters the demo hotfix work is already read through. GitHub emails it to everyone watching the repository. Deduplication matches on title and author, so a public repository stranger cannot suppress the channel by opening an issue with the same title.
  2. The issue is closed automatically on the first healthy run after the top up. Without that half, one shortfall would leave an issue open forever and the deduplication would then swallow every later alert as "already open" while the check kept running green. That is exactly the defect issue deploy-demo-box is failing on main #1416 records, where a tracking issue sat open through twenty one green runs still claiming a failure.
  3. The failing scheduled run is the secondary channel, which GitHub separately emails to whoever last touched the schedule.
  4. The balance is on the run summary page of every run, green or red, so the number is readable at a glance without opening a log.

Other provider balances, probed live rather than assumed

OPENROUTER_API_KEY and GROQ_API_KEY are the only provider credentials in the live configuration. Groq is the obvious second one and the answer is that it has nothing equivalent to read:

Probe Result
GET api.groq.com/openai/v1/models 200, so the key is live and authenticated
GET api.groq.com/openai/v1/organizations 404 unknown_url
GET api.groq.com/v1/credits 404 unknown_url

Groq publishes no balance, credit or usage endpoint on the inference API. Its spend is visible only in the console. This is worth stating rather than leaving implicit, because an alarm on one provider and silence on the rest is a false sense of coverage. The mitigating fact is that after the 2026-08-23 restructure Groq serves audio only in config.yaml (sections 6 and 7); no chat route is on it.

Can the free pool carry chat at zero balance, and is a fallback safe

Investigated, reported, and deliberately not built here, per the brief: a routing change under time pressure is how a demo breaks worse.

  • Every customer reachable chat alias except the two paid DeepSeek ones already runs on a free OpenRouter model. hive-default, hive-auto, hive-small, hive-medium and hive-fast all resolve to dots-studio/dots-3-note-preview:free, and HIVE_TOOLS_MODEL defaults to hive-free, not to a paid model. The issue's framing that "chat stops answering" is stale on that point; what stops is the paid set.
  • OpenRouter gates free model throughput on credits purchased all time, confirmed against their own limits documentation, not on the current balance. The account has purchased 10, so the 1000 requests per day free allowance survives a zero balance.
  • One caveat that argues for the floor rather than against it: the same documentation says a negative credit balance produces errors "including for free models". Zero is survivable, negative is not, which is another reason to alarm well before the balance reaches the bottom.
  • An automatic paid to free fallback would break the owner's one alias one price rule (D-032) at a layer the catalog cannot see, which is precisely why the two chat fallbacks that used to live in litellm_settings were deliberately removed. A customer charged the DeepSeek price and served a free model is a worse failure than a refusal. Not proposed.

Proof, against the real endpoint

Committed at docs/proof/provider-balance-alarm-2026-09-02/check-output.log. Four cases, all against the live account, no key printed.

=== 1. Alarm firing: real balance, default thresholds (7 days of runway, 2.00 USD floor) ===
LOW: the OpenRouter balance is 1.46 USD, under the 2.00 USD floor. The floor sits under the
runway test because a burn rate measured over a quiet window says nothing about the next busy hour.
1.46 USD remaining of 10.00 purchased (8.54 spent). Burn 0.004 USD/day from the provider's own
weekly usage figure for this key (no usable stored sample: either this is a first run, the cache
was evicted, or the newest sample is too recent to extrapolate from), runway 382.3 days.
exit=1

=== 2. Healthy branch: same real balance, thresholds moved under it (floor 1.00 USD, runway 0.1 days) ===
OK: 1.46 USD remaining of 10.00 purchased (8.54 spent). Burn 0.004 USD/day ... runway 382.3 days.
Threshold 0.1 days, floor 1.00 USD.
exit=0

=== 3. Account-wide delta path, seeded with the real reading recorded in issue #1411 ===
    (total_usage 8.409778422 at 2026-08-29T09:44Z, four days before this capture)
OK: 1.46 USD remaining of 10.00 purchased (8.54 spent). Burn 0.030 USD/day from account-wide
usage delta over the last 4.3 days, runway 48.2 days. Threshold 0.1 days, floor 1.00 USD.
exit=0

=== 4. Unreadable API must not report healthy (pointed at a loopback port nothing listens on) ===
FATAL: could not read the OpenRouter balance after 1 attempt(s). Last error: /credits could not
be reached: ECONNREFUSED
exit=1

Case 3 is the delta path measured against a real historical reading, not a synthetic one: the 8.409778422 figure is the number recorded in the issue on 2026-08-29, and 0.030 USD per day is the account's genuine four day burn.

The alarm's own delivery is proved separately, on the real Actions runner rather than reasoned about, in the proof comment on this pull request: run 33662109196 filed issue #1732, run 33662109465 proved the deduplication, and run 33662230066 closed it again. Both halves of the alert channel, not just the detection logic.

Tests

node .github/ci/check-provider-balance.test.mjs, wired into the required lint job in ci.yml next to the other guards of this shape. Twenty one cases, all offline: the suite serves its own stub of the provider API. They cover the threshold boundary on both sides, a first run with no stored sample, the spike case where the stored sample must beat the weekly average, sample pruning, a future or backwards sample, the absolute floor, an HTTP 500, a 401 with no credential echoed, a malformed body, a body missing the fields, an unreachable API, an absent credential, a misconfigured threshold and a corrupt state file. The last three assert the wiring itself: that the scheduled lane still invokes the check, still persists the sample, and still both files and closes its tracking issue, so deleting any half of the alarm turns a required check red.

Buglog entry

To be appended to .wolf/buglog.jsonl on main in a separate buglog only pull request after this merges, per .claude/rules/openwolf.md.

{"id":"1411-openrouter-balance-unobserved","date":"2026-09-02","title":"The OpenRouter account every paid route bills against ran down to 1.46 USD with nothing reading the balance","error_message":"No error was raised at all, which is the defect. GET https://openrouter.ai/api/v1/credits returned total_credits 10 and total_usage 8.540654392, so 1.46 USD remained, and the first signal available to anyone would have been a paid request failing with a 402 in front of a demo audience.","root_cause":"Nothing in the repository read the provider balance on any schedule. scripts/report-free-pool-health.py probes only free endpoints, whose throughput OpenRouter gates on credits purchased all time rather than on the current balance, so it is unaffected by depletion and cannot see it. The out of credit classifier in .github/workflows/ci.yml is by construction downstream of a job that has already failed. The balance was readable the whole time through a single unauthenticated-to-us GET that bills no inference, and no code path read it.","fix":"Added .github/ci/check-provider-balance.mjs and .github/workflows/provider-balance-watch.yml, a six hourly check that computes days of runway from the larger of the provider's own weekly usage figure and an account wide usage delta carried between runs in the Actions cache, fails on a runway under seven days or a balance under a two dollar backstop floor, fails loudly on any inability to read the number rather than reporting healthy, and files a deduplicated priority:critical GitHub issue that it closes itself on recovery. Regression guard .github/ci/check-provider-balance.test.mjs is wired into the required lint job and asserts the wiring as well as the arithmetic.","tags":["monitoring","money-path","openrouter","silent-absence","github-actions","demo-surface"]}

Every paid route in deploy/litellm/config.yaml bills against one OpenRouter
account, and that file declares no chat fallback on purpose, so at zero those
routes refuse outright rather than degrading. That account has been running
down for weeks and nothing read the balance on any schedule. The free pool
health report probes only free endpoints, which the balance does not affect,
and ci.yml's out of credit classifier only runs after a job has already failed.
The first available signal was a paid request failing in front of an audience.

Adds .github/ci/check-provider-balance.mjs, a scheduled check that reads
/api/v1/credits and /api/v1/key, and .github/workflows/provider-balance-watch.yml,
which runs it every six hours.

The threshold is days of runway rather than a dollar amount, because a fixed
dollar floor is meaningless without the spend rate. The burn rate comes from two
sources and the larger one wins: the provider's own weekly usage figure, which
is available on the very first run and answers what to do with no stored sample,
and the account wide delta against a cached sample, which reacts to a spike
within one run interval instead of diluting it across a week. A small absolute
floor sits underneath as a backstop, because a burn rate measured over a quiet
window says nothing about the next busy hour and at zero measured burn the
runway is arithmetically infinite.

Every failure to read the number exits non zero: an HTTP error, an unparseable
body, a body missing the fields, an absent credential, a misconfigured
threshold. A monitor that reports green when it could not reach the API is worse
than no monitor, because it asserts a state nothing verified.

The alarm's primary channel is a deduplicated GitHub issue rather than a red
square in the Actions tab, which nobody is subscribed to. It is filed when none
is open and closed on the first healthy run after it, so a stale open issue
cannot swallow the next alert, which is the defect issue #1416 records.
@coderabbitai

coderabbitai Bot commented Sep 2, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 43 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Team

Run ID: 604368f9-6ecb-4b3c-b422-b0ffc80301ea

📥 Commits

Reviewing files that changed from the base of the PR and between 757266b and 563b39a.

⛔ Files ignored due to path filters (1)
  • docs/proof/provider-balance-alarm-2026-09-02/check-output.log is excluded by !**/*.log
📒 Files selected for processing (5)
  • .github/ci/check-provider-balance.mjs
  • .github/ci/check-provider-balance.test.mjs
  • .github/workflows/ci.yml
  • .github/workflows/provider-balance-watch.yml
  • .gitignore

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sakibsadmanshajib sakibsadmanshajib added demo-surface Visible to the owner or a customer during the demo walk. money-path Touches billing, credits, pricing or the ledger. priority:critical Demo blocker or live outage. Drop everything. labels Sep 2, 2026
Three findings from the plain adversarial pass, all in the alerting half
rather than in the arithmetic.

The tracking issue title claimed the balance was running out, but the same
issue is filed when the balance could not be read at all, which is a different
problem with a different remedy. Retitled to cover both states, and the body
now says which one to read the block as and what to check when it is the
second.

The cache save step could turn a run red on its own. On a run that cannot
reach the API the script exits before writing the sample, and with no cache to
restore there is no path to save, which actions/cache/save reports as an error.
That painted a red square whose cause had nothing to do with the balance, on
top of the alarm that had already fired correctly. Marked continue-on-error,
since losing a sample only degrades the check to the provider's weekly figure.

The checker's header now records that the provider's weekly usage window is
not documented as rolling or as calendar anchored, which matters in one
direction: if it resets on a fixed weekday, that source under-reads for about
a day afterwards. The account wide delta is what covers it, which is a second
reason not to rely on the provider's figure alone once a sample exists.
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Adversarial review, pipeline mode

Two streams run. Neither is skipped.

Stream 1: CodeRabbit CLI (coderabbit review --agent --base main)

0 findings. Reviewed all six files in the diff: .github/ci/check-provider-balance.mjs, .github/ci/check-provider-balance.test.mjs, .github/workflows/ci.yml, .github/workflows/provider-balance-watch.yml, .gitignore, docs/proof/provider-balance-alarm-2026-09-02/check-output.log. CLI version 0.7.5, review completed cleanly rather than erroring out, so this is a real pass and not an unavailable stream.

Stream 2: plain adversarial pass

Three findings, all in the alerting half rather than the arithmetic, all fixed in 11645bb.

1. The tracking issue title lied about half the states it reports. The title read "the OpenRouter balance is running out", but the same issue is filed when the balance could not be READ at all, which is a different problem with a different remedy: a 401 or an unreachable endpoint is not fixed by topping up. An operator getting that email would have gone to the billing page over a broken credential. Fixed by retitling to "the OpenRouter balance is low or unreadable" and adding a paragraph to the body telling the reader to check whether the block begins with LOW or FATAL and what each one means. Kept as one deduplication bucket on purpose: two titles would be two buckets, and a second bucket is more moving parts than the distinction is worth when the body already carries it.

2. The cache save step could turn a run red on its own, for a reason unrelated to the balance. On a run that cannot reach the API the script exits before writing the sample, and if no cache existed to restore there is then no path to save, which actions/cache/save reports as a path validation error. The alarm had already fired correctly at that point, and this painted a second red square with an unrelated cause on top of it, which is how a monitor's output stops being read. Marked continue-on-error: true with the reasoning inline: losing a sample only degrades the check to the provider's weekly figure, so it must never be able to decide the run's colour.

3. The weekly burn source has an undocumented window and the file did not say so. OpenRouter does not document whether usage_weekly is rolling or calendar anchored, and the figures the endpoint actually returns are consistent with either reading. That matters in exactly one direction: if it resets on a fixed weekday, that source under-reads for about a day afterwards, which is the direction that makes an alarm fire late. The account wide delta covers it once a sample exists. Now recorded in the checker's header as a second reason not to rely on the provider's figure alone.

Considered and deliberately not changed

  • Both burn sources agree today, because the key's own usage equals the account's total_usage. The check compares them anyway and prints a NOTE when they diverge, rather than silently under-reading account burn if a second key is ever added.
  • The dollar floor is not the primary threshold and is not meant to be. At the burn this account has actually averaged, 0.004 USD per day, the runway test alone would not fire today and the floor is what fires. That is stated in the pull request body rather than hidden, because a reader who assumes the runway test is doing the work would draw the wrong conclusion from the proof log.
  • No automatic paid to free fallback, per the brief and per D-032. A customer charged the DeepSeek price and served a free model is a worse failure than a refusal, which is why the chat fallbacks were deliberately removed from litellm_settings in the first place.

The shortfall label proves the alarm fires and files its tracking issue. The
close half needs proving too, and it cannot be reached by removing that label,
because the real balance is genuinely under the floor today and the check would
correctly stay red. A second label lowers the thresholds under the real balance
instead, so the recovery run measures the same real account and demonstrates
that the issue gets closed rather than left open to swallow the next alert.
@sakibsadmanshajib
sakibsadmanshajib marked this pull request as ready for review September 2, 2026 17:36
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@sakibsadmanshajib sakibsadmanshajib added provider-balance-test:simulate-low Squeeze the balance runway threshold so the alarm fires, for proving it before merge run-provider-balance-watch Run provider-balance-watch.yml on this pull request provider-balance-test:simulate-healthy Prove the balance alarm closes its tracking issue on recovery and removed provider-balance-test:simulate-low Squeeze the balance runway threshold so the alarm fires, for proving it before merge provider-balance-test:simulate-healthy Prove the balance alarm closes its tracking issue on recovery run-provider-balance-watch Run provider-balance-watch.yml on this pull request labels Sep 2, 2026
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Proof of delivery, on the real account and the real Actions runner

The check itself was already proved locally against the live endpoint (docs/proof/provider-balance-alarm-2026-09-02/check-output.log, cases 1 to 4). What follows is the half that matters more, because a check whose notification never arrives is the failure this pull request exists to remove: the alert channel exercised end to end through the label gated hook, each run reading the real OpenRouter account.

Run Label Conclusion What it proves
33662109196 provider-balance-test:simulate-low failure The alarm fires on the real balance and files the tracking issue
33662109465 same failure Deduplication: tracking issue #1732 is already open, not commenting again
33662230066 provider-balance-test:simulate-healthy success Recovery closes the issue, so a stale one cannot swallow the next alert

Issue #1732 is the artefact: filed by app/github-actions at 17:36 with priority:critical, demo-surface and money-path, closed at 17:37:45Z with a recovery comment. It was opened and closed by the workflow, not by hand.

Log lines, verbatim:

Cache not found for input keys: provider-balance-state-33662109196, provider-balance-state-
simulating a shortfall: runway threshold raised to 100000 days
LOW: the OpenRouter balance is 1.46 USD, under the 2.00 USD floor.
1.46 USD remaining of 10.00 purchased (8.54 spent). Burn 0.004 USD/day from the provider's own
weekly usage figure for this key, runway 382.3 days.
Cache saved with key: provider-balance-state-33662109196
opened the tracking issue

tracking issue #1732 is already open, not commenting again

simulating recovery: thresholds lowered under the real balance, to prove the tracking issue gets closed
OK: 1.46 USD remaining of 10.00 purchased (8.54 spent). ... Threshold 0.001 days, floor 0.00 USD.

Three details worth naming, since each of them is a way this could have shipped broken and looked fine:

  1. The dedupe author match is confirmed rather than assumed. app/github-actions is what gh actually reports for an issue filed by GITHUB_TOKEN, verified here against a real bot filed issue. Had it been wrong in the tightening direction, this would file a fresh issue four times a day, which is how a monitor gets muted.
  2. The cache round trip works. The first run reports a miss and saves; the sample it stores is what a later run compares against for account wide burn. A miss is a supported state, not a failure, and the run above shows the check producing a runway anyway from the provider's weekly figure.
  3. The job gate behaves. An earlier run on this pull request was correctly SKIPPED, before the run-provider-balance-watch label existed on it. That is the if: condition working, not an accident, and it is the same gate that admits schedule on main.

The test labels have been removed from this pull request, so no further balance runs fire on it.

@sakibsadmanshajib
sakibsadmanshajib merged commit 27820db into main Sep 2, 2026
32 checks passed
@sakibsadmanshajib
sakibsadmanshajib deleted the fix/1411-provider-balance-alarm branch September 2, 2026 18:09
sakibsadmanshajib added a commit that referenced this pull request Sep 2, 2026
Second bug-log reconciliation batch of the day. The first batch (#1743,
merged 2026-09-02T19:44:04Z) appended 197 entries from 167 PRs, taking
`.wolf/buglog.jsonl` on `main` to 511 lines.

This batch sweeps every PR merged after that point which carried a `##
Buglog entry` heading in its body, appending 14 entries from 7 PRs:

- #1733 (1 entry)
- #1735 (1 entry)
- #1739 (6 entries)
- #1740 (1 entry)
- #1748 (1 entry)
- #1749 (2 entries)
- #1756 (2 entries)

Checked and excluded:
- #1727 carries no buglog entry. It is a docs/process PR
(tracking-discipline rule), not a bug fix, and its body mentions
`.wolf/buglog.jsonl` only in passing prose.
- #1715, #1729, #1731 and #1734 merged before this batch's window and
are already present in the first batch (#1743). Verified by
id/error_message lookup against the 511 lines already on `main`.

Every entry was extracted from its source PR body, parsed as JSON to
confirm it is well-formed, and checked for the required `error_message`,
`root_cause`, `fix` and `tags` fields (all present, none reconstructed).
No duplicates were found against the existing 511 lines or within this
batch, checked by both `id` and exact `error_message` match.

Diff is exactly one file, 14 insertions, 0 deletions. The first 511
lines byte-match `main`'s current copy (verified with `diff` against
`git show origin/main:.wolf/buglog.jsonl`).

This PR was not opened on a fix or feature branch, per
`.claude/rules/openwolf.md`: it is the dedicated buglog-only PR,
branched directly from `main`, diffing only `.wolf/buglog.jsonl`.

Refs #873

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

demo-surface Visible to the owner or a customer during the demo walk. money-path Touches billing, credits, pricing or the ledger. priority:critical Demo blocker or live outage. Drop everything.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

OpenRouter account is at $1.59 of $10 purchased credit, and nothing warns before it hits zero

1 participant