Repository navigation
fix: size the Hive Auto credit hold to the request, not the envelope (issue #1372) - #1378
Conversation
/console/docs picked the first catalog id starting with "hive-" for both the curl and the Python quickstart. On this deployment that is hive-auto, the variable-price Auto Router, whose up-front credit hold is sized for the worst request the bounds allow rather than for the request in hand. A developer who pasted the quickstart got a quota refusal on their first ever call, with credit still on the account. Pick a fixed-price alias that speaks chat instead, falling back through any Hive alias, then the first catalog entry, then the seeded default, so the snippets stay runnable when the catalog cannot be read at all. The capability check keeps an embeddings-only alias out of a /chat/completions snippet, which is the other way "first hive- id wins" hands someone a sample that cannot run. The pick moves into lib/quickstart-model.ts so it can be tested without standing up the whole server component. The first test pins the live catalog ordering that produced the bug.
…(issue #1372) Selecting Hive Auto in the chat model picker made every subsequent turn fail with "You exceeded your current quota, please check your plan and billing details" while the same screen showed 0.455 USD remaining. Over the API the alias answered 429 insufficient_quota for every request. hive-auto is priced upstream_actual, so there is no catalog rate to charge against and the gateway takes a credit hold up front instead. That hold was a flat 2.00 USD equivalent, the price of the LARGEST request the variable-price bounds allow: a full 256 KiB body generating a full 16,384 tokens at the route's 3.00 and 15.00 USD per million rate ceiling. Every request paid that authorization, so no account holding less than 2.00 USD could use the alias at all, whatever it was actually sending. The hold is now derived from the request in hand. Both quantities are already known at that point and both are upper bounds, not estimates: len(body) bytes is a rigorous upper bound on prompt tokens, and EnforceVariablePriceBounds has already written this request's effective completion ceiling into the body. Pricing those two at the same rate ceiling, through the same CreditsForUpstreamCost settlement uses, yields a hold that still provably covers the request and is about six times smaller for an ordinary chat turn. The catalog figure stays the upper bound and the fallback whenever no per-request bound can be computed, so this can only ever shrink a hold toward a proven number, never past it. Three things had to be true for that to be honest rather than merely smaller, and two of them were not. The session chat path took the hold but never ran the bounds at all, so a turn on a variable-price alias went upstream with no size cap and no completion ceiling. It now runs EnforceVariablePriceBounds before anything is held or dispatched, exactly as the API path does. A pass-through, and one comparison, for every fixed-price alias. route-openrouter-auto-live, the route hive-auto actually serves, carried neither usage.include nor provider.max_price on the live box. The config sync generates that entry from provider_routes and owns only model, api_base and api_key, merging every other key from deploy/litellm/config.yaml field by field, and the file had no entry for it to merge. Both absences are silent and both are expensive: without usage.include OpenRouter reports no cost and every request settles through the fail-closed path at the full hold, and without max_price the rate half of the hold's coverage proof does not exist. The entry is added here with the same three keys its beta twin carries. The rate ceilings are now also restated as Go constants, because the request path cannot read the YAML, and TestVariablePriceCeilingsMatchTheLiteLLMConfig fails when the restatement drifts from the file, for both variable-price routes by name. The refusal text is fixed too. A reservation that does not fit is not a quota overrun, and telling a customer with a positive balance to check a plan Hive does not sell was wrong twice over. The 409 branch now says the available credit does not cover this request and what to do about it; the 429 branch, which is a rate limit on the reservation call and not a credit verdict at all, says so separately. Status, type and code are unchanged, because that is what SDKs branch on. Tests. TestASolventAccountIsNotRefusedForAnEnvelopeItIsNotUsing pins the reported condition: an ordinary turn on a 0.455 USD balance is admitted, where the envelope hold is four times that balance. Mutating variablePriceRequestHold back to the old behaviour makes it fail. TestTheSizedHoldStillCoversWhatTheRequestCanCost is the property that must not break, re-derived from the rate ceiling in the LiteLLM config rather than from the Go constants; halving the computed hold makes it fail on three of its four cases. TestAnAccountThatCannotAffordTheRequestIsStillRefused holds the other direction: a request at the size cap still exceeds that balance, and nothing drops below the endpoint floor.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Warning Review limit reachedNext included review available in 8 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (17)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Adversarial review, stream 1 of 2: Antigravity (gemini-3.1-pro-high, effort high)Framed as a pre-merge review of this repository's own PR, with the seven money-path questions this diff raises stated up front. Full verdict below, then my responses. Findings1. Under-reservation and arithmetic. Checked the order of operations across 2. Call sites. All four pass the bounded body and the correct endpoint. Line-by-line confirmation of the reassignment preceding each call. 3. Hold and release balance. 4. Overflow, rounding, float64, nil. 5. Assertions that would still pass with the fix reverted. 6. 7. SDK contract and disclosure. Status, type and code unchanged in both writers. The new message discloses no provider, route, currency amount or balance figure. Verdict: LGTM. My responsesOn finding 5, rebutted, with a code change so it is not re-raised. Those two tests guard the under-reservation direction, and the old hold over-reserved, so over-reserving has to satisfy them. A guard against holding too little that went red when the hold was too large would be the wrong guard. The revert is caught by On finding 6, accepted and recorded as a deliberate consequence. Without a rate ceiling on the route Findings 1 through 4 and 7: agreed, no change. |
Antigravity, finding 5: two of the three sizing tests pass with the fix reverted. Rebutted rather than changed, with a comment on the test file so the next reader does not spend the same time on it. They guard the under-reservation direction and the old hold over-reserved, so over-reserving has to satisfy them; the revert is caught by TestASolventAccountIsNotRefusedForAnEnvelopeItIsNotUsing, and shrinking too far is caught by the coverage table. CodeRabbit, major: the Hive fallback rung in pickQuickstartAlias could still return an embeddings-only alias if the catalog listed one ahead of every chat alias. Accepted; the chat check now applies to both rungs, with a regression test for that ordering. CodeRabbit, major: a float64 conversion in a test log line. Accepted; big.Rat.FloatString instead. Nothing in this repo converts credits to money through a float, log line included. CodeRabbit, critical (duplicate test declaration) is a false positive: there is exactly one declaration of TestAFixedPriceAliasKeepsTheFlatEndpointHold and the package compiles. CodeRabbit, major (move the rate ceilings and the model slug to environment variables) is rebutted in the PR thread; the slug being a literal is a recorded decision from issue #689, and an environment variable for the ceilings would add a silent drift path where a cross-file test currently forbids one. The LiteLLM comment on usage.include is corrected to match what the live run actually showed. A hive-auto turn settled with terminal_usage_confirmed true without it, so setting it is belt and braces rather than a repair. max_price remains the load-bearing half.
Adversarial review, stream 2 of 2: CodeRabbit CLI
Critical, Functional Correctness, Rebutted, false positive. There is exactly one declaration: at line 210, in a 218-line file. The package compiles and Major, Functional Correctness, Accepted, fixed. Correct, and it is the same class of bug as the one this function exists to stop: rung 2 could hand a Major, Maintainability, Accepted, fixed. Major, Maintainability, Rebutted. The failure mode named is real, and an environment variable is the one shape that cannot prevent it. The number these constants have to agree with is Major, Maintainability, Rebutted, and this one is a recorded decision rather than a preference. From the comment on the twin route sixteen lines above the flagged one:
Both streams have now run. No stream was skipped. |
Visual proofLive demo box, chat-hive.scubed.co, signed in as qa-tester whose composer footer reads $0.455 remaining. Before: Hive Auto refuses every turn with "You exceeded your current quota, please check your plan and billing details". After, same account and same model with this branch's edge-api running on the box: "Hello. Hive auto works." The reservation for that turn held 344,853,600 credits instead of 2,000,000,000, consumed 43,526 and released 344,810,074, which balances to the credit. Full capture log and ledger row in docs/proof/hive-auto-reservation-1372-2026-08-29/. |
Review stream availabilityRecorded so an absent stream is not read as a clean pass.
No inline review threads are open on this PR ( |
Visual proofPost-deploy proof, pipeline stage 10. deploy-demo-box run 33241444493 for merge commit 027b375 completed successfully; edge-api is on the deploy's own image cafdfda15245, and the deployed /etc/litellm/config.yaml now carries max_price 3/15 and usage.include on route-openrouter-auto-live, merged in beside the sync-owned api_base exactly as intended. Same account, same model, $0.45 remaining: Hive Auto answers "hive auto post deploy ok". Reservation for that turn held 346,844,400 credits, consumed 40,110, released 346,804,290, balancing to the credit, terminal_usage_confirmed true. |
…1384) ## What this is A decision ledger correction, not code. Touches `.wolf/decisions.md` only, one appended line. ## The stale fact D-047 (2026-08-23) records that both `hive-auto` and `hive-default` were converted to fixed price at half their effective customer price. That stopped being true of `hive-auto` the very next day. `supabase/migrations/20260824_02_free_pool_router.sql:307` sets `pricing_mode = 'upstream_actual'` on `hive-auto`, with `reservation_estimate_credits = 2000000000`. The live database confirms `upstream_actual` today. The schema is authoritative on three grounds: it is the later change, the running system agrees with it, and the reservation mechanism PR #1378 repairs only exists in `upstream_actual` mode at all. `hive-default` is unaffected by this correction and remains fixed price exactly as D-047 records. This was not a paperwork problem. An agent reasoning from D-047 alone tonight concluded `hive-auto` was fixed price and missed the flat 2.00 USD hold that made the alias unusable below a two dollar balance, filed as issue #1372 and repaired in PR #1378. ## What changed Appended `D-059` to `.wolf/decisions.md`, per the ledger's own supersede-in-place protocol: D-047 stays exactly as written (it was a correct record of what was decided on 2026-08-23), and D-059 records that the `hive-auto` half of it was reversed the following day, with the schema line, the live confirmation, and the issue and PR that make the correction matter in practice. Sources cited in the entry: issue #1372, PR #1378, `supabase/migrations/20260824_02_free_pool_router.sql`. ## CI expectations `.wolf/decisions.md` matches the `*.md` arm of the changed-files allow-list in `.github/workflows/ci.yml` (the `changes` job), so the required checks should report green without running the full test matrix. This is a docs-only ledger correction, not a code or schema change, so no functional verification is applicable. ## Review No adversarial review streams requested for this PR. It is a documentation correction against verified facts (migration file, live database, linked issue and PR), not a code change; a review pipeline pass is not being skipped by omission, it is being explicitly declared not applicable here. Fixes documentation staleness surfaced while working issue #1372 / PR #1378.
## Summary This is the batched buglog follow-up for the pull requests merged to `main` on 2026-08-29. Its diff is `.wolf/buglog.jsonl` and nothing else. Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or failed build must be logged, but the line may never be appended on a fix branch. `merge=union` in `.gitattributes` resolves concurrent appends locally and is ignored by GitHub's server side merge, so two branches that both appended land in hard conflict there. An unmergeable pull request gets no `refs/pull/N/merge`, no `pull_request` run and therefore zero checks, and the required status gate then blocks the merge for a reason the page never states (issue #873). Each fix accordingly carried its entry in its own pull request body, and this pull request copies them onto `main` in one batch, which the protocol explicitly prefers over one pull request per entry. ## Scope examined Fifty nine pull requests merged to `main` on 2026-08-29. Forty eight of them carried at least one entry, for eighty two entries in total. Thirty two of those were already on `main` and are skipped, leaving fifty appended here from thirty four pull requests. The largest block of skips comes from #1342, the equivalent batch for the 2026-08-28 merges, which merged earlier the same day and already landed thirty six entries covering #1257, #1268, #1276, #1277, #1287, #1292, #1293, #1294, #1296, #1301, #1303, #1305, #1313, #1335 and #1337. ## What landed Fifty entries appended, one JSON object per line, append only. The 232 pre-existing lines are byte identical to `origin/main` (verified by hashing the first 232 lines of the result against the base file). Every line in the resulting file parses as JSON and carries `error_message`, `root_cause`, `fix` and `tags`. | Source | Entries | |---|---| | #1083 | 2 | | #1277 | 1 | | #1278 | 1 | | #1298 | 1 | | #1334 | 1 | | #1336 | 3 | | #1343 | 1 | | #1346 | 1 | | #1351 | 1 | | #1365 | 2 | | #1368 | 1 | | #1369 | 1 | | #1371 | 3 | | #1375 | 3 | | #1376 | 1 | | #1378 | 1 | | #1379 | 2 | | #1388 | 5 | | #1389 | 3 | | #1390 | 2 | | #1393 | 1 | | #1394 | 1 | | #1410 | 1 | | #1417 | 1 | | #1421 | 1 | | #1423 | 1 | | #1424 | 1 | | #1426 | 1 | | #1429 | 1 | | #1431 | 1 | | #1433 | 1 | | #1434 | 1 | | #1436 | 1 | | #1439 | 1 | Entries are copied verbatim from their source pull request bodies. Nothing was rewritten, no field was invented, and no field was added. No JSON needed repair: all eighty two extracted entries parsed on the first attempt and all four required fields were present on every one. ## Merged pull requests that carried no entry Eleven of the fifty nine. Recorded here because the gap is itself the useful signal. | Pull request | Title | Assessment | |---|---|---| | #1013 | chore(deps): bump the go-minor-patch group across 1 directory with 4 updates | Dependabot bump, no defect fixed, no entry expected | | #1015 | chore(deps): bump the go-minor-patch group across 1 directory with 6 updates | Dependabot bump, no entry expected | | #1016 | chore(deps): bump golang from 1.26-alpine to 1.27-alpine in /deploy/docker | Dependabot bump, no entry expected | | #1218 | chore(deps): bump postcss from 8.5.19 to 8.5.26 in /apps/desktop | Dependabot bump, no entry expected | | #1219 | chore(deps): bump golang.org/x/crypto from 0.41.0 to 0.52.0 in /apps/control-plane | Dependabot bump, no entry expected | | #1342 | chore: batch buglog entries for the 2026-08-28 merges | The previous batch pull request itself, correctly carries no entry of its own | | #1364 | chore: remove four dead skills and record the patterns that cost time | Protocol gap. The body records patterns that cost time, which is the shape of a buglog entry, but none was written as one | | #1383 | test: retire stale expected-failure markers, restore the ones that are true (#1381, #1382, #1324) | Protocol gap. Stale `it.fails` markers reading as red is a real defect that was fixed here and should have carried an entry | | #1384 | docs: correct D-047, hive-auto reverted to variable pricing (D-059) | Decision ledger correction, arguably a documentation defect, no entry written | | #1387 | chore(deps): bump next from 15.5.23 to 16.3.3 in /apps/agent-console | Dependabot bump, no entry expected | | #1398 | docs: rescue the 2026-08-25 parity captures and add the 2026-08-29 QA matrix evidence | Documentation and evidence rescue, no entry written | Six of the eleven are Dependabot bumps and one is the previous batch, so the genuine protocol gaps are #1364, #1383, #1384 and #1398. Of those, #1383 is the one worth a follow-up: it fixed a real defect class (a stale expected-failure marker reads as a red "Expect test to fail" and gets dismissed as pre-existing) and left no record. ## Entries skipped as already present Thirty two. Thirty of them matched an entry already on `main` on `error_message`, `id` or `fix`. Two more from #1278 are semantic duplicates that an exact match would have missed, and were skipped after reading the landed entries they duplicate: - #1278's `streaming content_block_start omits text field` entry is covered by the consolidated `bug-2026-08-28-anthropic-sdk-wire-conformance` entry landed from #1296, whose root cause names the same `omitempty` on `StreamContentBlock.Text`. - #1278's `GET /v1/models leaked an upstream provider name` entry is covered by `BUG-1284`, landed from #1300, which names the same `public.model_aliases.summary` publication path. #1278's third entry, on `top_k` forwarding producing a 400, is not covered anywhere on `main` and is appended here. #1342 recorded #1278 as fully "merged into #1296", which was accurate for two of its three entries. ## Note on entry quality One appended entry is thin: #1277's parity re-score record carries `error_message` of `n/a` and a root cause of "console had no privacy/data-policy surface at all". It is a parity gap record rather than a defect record. It is included exactly as written rather than embellished, per the protocol's preference for the author's own words. ## Test plan - [x] Branch cut fresh from `origin/main`, diff is `.wolf/buglog.jsonl` and nothing else - [x] First 232 lines byte identical to the base file (md5 match) - [x] All 282 resulting lines parse as JSON and carry `error_message`, `root_cause`, `fix` and `tags` - [x] No `.wolf/` telemetry (`anatomy.md`, `memory.md`, `token-ledger.json`, `hooks/_session.json`, `buglog.json`) in the commit - [ ] The six required checks report green via the inert path allowlist in `.github/workflows/ci.yml` --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>



Fixes #1372.
The defect
Selecting "Hive Auto" in the chat model picker made every subsequent message fail with "You exceeded your current quota, please check your plan and billing details" while the same screen showed 0.455 USD remaining. Over the API the alias answered 429
insufficient_quotafor every request.hive-autois pricedupstream_actual, so there is no catalog rate to charge against and the gateway takes a credit hold up front instead. That hold was a flat 2.00 USD equivalent (reservation_estimate_credits = 2000000000,supabase/migrations/20260824_02_free_pool_router.sql). It is the price of the LARGEST request the variable-price bounds allow: a full 256 KiB body generating a full 16,384 tokens at the route's 3.00 and 15.00 USD per million rate ceiling. Every request paid that authorization, so no account holding less than 2.00 USD could use the alias at all, whatever it was actually sending.Compounding it,
/console/docsusedhive-autoin both the curl and the Python quickstart, so a developer pasting the quickstart got a quota refusal on their first ever call.Three findings that outlive this fix
None of these were in the issue. Each is independently worth knowing after this merges, so they are stated here rather than only in the commit that happens to touch them.
1. The session chat path took a credit hold but never ran the bounds.
internal/chat/dispatch.gois the path Open WebUI chat actually uses, and it calledstartSettlementwithout ever callingEnforceVariablePriceBounds. A variable-price turn therefore went upstream with no request size cap and no completion ceiling, and the hold in front of it was covering a request nobody had bounded. That is a money hole on its own, independent of how the hold is sized: the API path has had those bounds since 2026-08-22 and this path silently did not.2. Anything not in
deploy/litellm/config.yamlsilently does not exist on a synced route.route-openrouter-auto-liveis generated fromprovider_routesby the config sync, andlitellmconfig.mergeParamsgives the generator ownership of exactly three keys:model,api_base,api_key. Every other key is merged in from the YAML. A route with aprovider_routesrow and no YAML entry therefore comes up with noextra_bodyat all, and nothing anywhere reports that. That is howhive-autoended up serving with noprovider.max_price, which is the rate half of its own credit hold's coverage proof. The general shape is worth remembering: a DB-managed route is only half-configured by the database.3.
usage.includewas belt and braces here, not a repair, and the first version of this PR body said otherwise. The claim was that without it OpenRouter reports no cost and every request settles at the full hold through the fail-closed path. The live run disproved it: a realhive-autoturn settled withterminal_usage_confirmedtrue without the key set. The claim is downgraded rather than left standing, and the flag is still set, so the charge does not depend on that continuing to be true by default.What this changes
1. The quickstart (
c8d5531, first commit, reviewable on its own). The docs page picked the first catalog id starting withhive-, which on this deployment is the variable-price router. It now prefers a fixed-price alias that speaks chat, falling back through any Hive alias, the first catalog entry, and the seeded default. The pick moved intolib/quickstart-model.tsso it is testable without standing up the server component.2. The hold (
6a35740). It is now derived from the request in hand. Both quantities are already known at that point and both are upper bounds, not estimates:len(body)bytes is a rigorous upper bound on prompt tokens, andEnforceVariablePriceBoundshas already written this request's effective completion ceiling into the body. Pricing those two at the same rate ceiling, through the sameCreditsForUpstreamCostsettlement uses, yields a hold that still provably covers the request and is about six times smaller for an ordinary chat turn (0.344 USD against 2.00 USD). The catalog figure stays the upper bound and the fallback whenever no per-request bound can be computed, so this can only ever shrink a hold toward a proven number, never past it.3. Two things that had to be true for that to be honest rather than merely smaller, and were not.
The session chat path (
internal/chat/dispatch.go) took the hold but never ran the bounds at all. A turn on a variable-price alias went upstream with no size cap and no completion ceiling, so the only thing in front of an arbitrarily large charge was a hold sized for a request nobody had checked. It now runsEnforceVariablePriceBoundsbefore anything is held or dispatched, exactly as the API path does. A pass-through, and one comparison, for every fixed-price alias.route-openrouter-auto-live, the routehive-autoactually serves, carried neitherusage.includenorprovider.max_priceon the live box (verified 2026-08-29 by reading/etc/litellm/config.yamlinside the running container). The config sync generates that entry fromprovider_routesand owns onlymodel,api_baseandapi_key, merging every other key fromdeploy/litellm/config.yamlfield by field (litellmconfig.mergeParams), and the file had no entry for it to merge.max_priceis the load-bearing absence: without it the rate half of the hold's coverage proof does not exist.usage.includeturned out to be belt and braces rather than a repair, and the PR body said otherwise until the live run corrected it: a real hive-auto turn on 2026-08-29 settled withterminal_usage_confirmedtrue without it, so a cost does reach us on this path today. Setting it explicitly means the charge does not depend on that continuing to be true by default, which is why its beta twin sets it too. The entry is added with the same three keys that twin carries; the compose entrypoint reseeds the volume when the seed checksum changes, so this reaches the box on deploy.Stated plainly, because it is a behaviour change and not only a safety net: with
max_priceset, an Auto request that resolves to a model above 3.00 USD per million prompt or 15.00 per million completion now fails rather than being served at an unbounded price. That ceiling is the one the flat 2.00 USD hold was always derived from (20260822_30_openrouter_auto_variable_pricing.sql); the live route simply never enforced it. The alias ispreviewvisibility and its beta twin has run with the identical ceiling since 2026-08-22.4. The refusal text. A reservation that does not fit is not a quota overrun, and telling a customer with a positive balance to check a plan Hive does not sell was wrong twice over. The 409 branch now says the available credit does not cover this request and what to do about it. The 429 branch, which is a rate limit on the reservation call and not a credit verdict at all, says so separately. Status, type and code are unchanged, because that is what SDKs branch on.
Which record is authoritative on hive-auto pricing
.wolf/decisions.mdD-047 records hive-auto as converted to fixed price on 2026-08-23. The schema disagrees:20260824_02_free_pool_router.sqlflipped it back toupstream_actualthe next day, and the live row confirms it.The schema is authoritative. It is the later change, the running system agrees with it, and the whole variable-price mechanism this PR touches only exists because the alias is
upstream_actual. D-047 is stale on this point and should be amended to say so. This PR does not change the pricing mode; it changes how the hold that mode requires is sized.Live proof
Captured against the demo box at
https://chat-hive.scubed.coasqa-tester@hive.test, the account the defect was reported on, whose composer footer reads$0.455 remaining. Session minted through the admin one-time-token flow; no password set, reset or rotated. Screenshots are posted as inline images on this PR; the full capture log and the ledger row are indocs/proof/hive-auto-reservation-1372-2026-08-29/.Before, the box on
main: Hive Auto selected in the picker, "Say hello in one sentence." answered withYou exceeded your current quota, please check your plan and billing details.directly aboveYou've used $0.00287 today · $0.455 remaining.After, the same box with this branch's
edge-apiimage: same account, same model,Hello. Hive auto works.and a200on/api/chat/completions.The reservation for that turn, from
public.credit_reservationsjoined topublic.request_attempts:The hold was 0.3449 USD, not 2.00. Hold and release balance to the credit: 43,526 consumed plus 344,810,074 released is exactly 344,853,600 reserved, nothing stranded. The turn really cost 43,526 credits, so the old hold demanded an authorization roughly 46,000 times the actual charge before serving it.
The box was restored to
main's image immediately afterwards (docker inspect hive-edge-api-1reports the pre-change build, container healthy) and the temporary build worktree removed. The LiteLLM half of this PR was verified by reading the running container's config, not by editing it; it reaches the box through the deploy's seed reconciliation.Money-path invariants
math/bigrationals throughCreditsForUpstreamCost, the same function settlement uses, so hold and charge carry the identical margin (7/5), credit unit (D-046), round-half-up (D-031) and one-credit floor (D-034, D-048). No float64 anywhere on the path.CreateReservationand released or finalized in full by the same paths as before. Nothing about settlement, release, or terminal-state handling moved.Tests that can fail
TestASolventAccountIsNotRefusedForAnEnvelopeItIsNotUsingpins the reported condition: an ordinary turn on a 0.455 USD balance is admitted, where the envelope hold is over four times that balance. RED confirmed: mutatingvariablePriceRequestHoldto always fall back reproduces the defect exactly,hold 2000000000 against a 455000000 balance.TestTheSizedHoldStillCoversWhatTheRequestCanCostis the property that must not break. It re-derives the bound fromprovider.max_pricein the LiteLLM config rather than from the Go constants the implementation uses, so a drift between the two fails here as well. RED confirmed: halving the computed hold fails three of its four cases, the fourth being covered by the endpoint floor.TestAnAccountThatCannotAffordTheRequestIsStillRefusedholds the other direction: a request at the size cap still exceeds that balance, and nothing drops below the endpoint floor.TestVariablePriceCeilingsMatchTheLiteLLMConfigchecks the Go rate constants against the YAML for both variable-price routes by name. Naming them is the point: the live config hadroute-openrouter-auto-livewith nomax_priceat all and nothing failed.TestTheHoldProvablyCoversTheWorstBoundedRequestis unchanged and still guards the catalog envelope against the config ceiling and the request bounds.Three existing tests asserted the old flat hold and were updated rather than deleted, each keeping its original intent: the two settlement tests now assert the derived figure as an exact magnitude (
344274000, arithmetic spelled out atautoHoldForAnEmptyBody) instead of2000000000, andTestReservationCreditsCannotUnderReservenow exercises the envelope through the unsizable-body fallback.TestSessionChatRefusesWhenTheAccountCannotPaydroppedcreditfrom its leak list, with the reasoning in the comment: it is Hive's own billing unit and the word the console uses, and forbidding it is what left OpenAI's misleading sentence in place. Provider, route name, currency and balance remain forbidden.Full
go test ./apps/edge-api/... ./apps/control-plane/... -count=1 -shortis green.Buglog entry
{"id":"bug-1372-hive-auto-flat-reservation-hold","date":"2026-08-29","title":"Hive Auto unusable below 2.00 USD: flat envelope-sized credit hold refused every request","error_message":"You exceeded your current quota, please check your plan and billing details. (429 insufficient_quota) shown while the account held 0.455 USD","root_cause":"hive-auto is priced upstream_actual, so the gateway takes an up-front credit hold instead of charging a catalog rate. reservation_estimate_credits was a flat 2000000000 (2.00 USD), the price of the largest request the variable-price bounds allow, and every request paid that authorization regardless of its actual size. control-plane's enforcePolicy refused any account whose balance was under it. Three silent enablers: the session chat path never ran EnforceVariablePriceBounds at all, so its hold covered a request with no size cap or completion ceiling; route-openrouter-auto-live carried no provider.max_price, so the rate half of the hold's coverage proof did not exist; and it carried no usage.include, so OpenRouter reported no cost and every settlement fell through to charging the full hold unconfirmed. The refusal text was OpenAI's canonical insufficient_quota sentence, which named a plan Hive does not sell.","fix":"Size the hold from the request: len(body) as a prompt-token upper bound plus the effective completion ceiling already written into the body, priced at the route's own rate ceiling through CreditsForUpstreamCost, capped by the catalog envelope and floored by the endpoint default. Run EnforceVariablePriceBounds on the session chat path. Add route-openrouter-auto-live to deploy/litellm/config.yaml with usage.include and provider.max_price. Split the refusal text by status and say what happened.","tags":["billing","reservation","upstream_actual","hive-auto","litellm-config","chat","error-message","money-path"],"files":["apps/edge-api/internal/inference/pricing.go","apps/edge-api/internal/chat/dispatch.go","apps/edge-api/internal/chat/billing.go","apps/edge-api/internal/inference/reservation_guard.go","deploy/litellm/config.yaml","apps/web-console/lib/quickstart-model.ts"],"issue":1372}🤖 Generated with Claude Code
https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1