Skip to content

fix: size the Hive Auto credit hold to the request, not the envelope (issue #1372) - #1378

Merged
sakibsadmanshajib merged 3 commits into
mainfrom
fix/1372-hive-auto-reservation
Aug 29, 2026
Merged

sakibsadmanshajib merged 3 commits into
mainfrom
fix/1372-hive-auto-reservation

Conversation

@sakibsadmanshajib

@sakibsadmanshajib sakibsadmanshajib commented Aug 29, 2026 •

Copy link
Copy Markdown
Owner

Fixes #1372.

The defect

Selecting "Hive Auto" in the chat model picker made every subsequent message fail with "You exceeded your current quota, please check your plan and billing details" while the same screen showed 0.455 USD remaining. Over the API the alias answered 429 insufficient_quota for every request.

hive-auto is priced upstream_actual, so there is no catalog rate to charge against and the gateway takes a credit hold up front instead. That hold was a flat 2.00 USD equivalent (reservation_estimate_credits = 2000000000, supabase/migrations/20260824_02_free_pool_router.sql). It is the price of the LARGEST request the variable-price bounds allow: a full 256 KiB body generating a full 16,384 tokens at the route's 3.00 and 15.00 USD per million rate ceiling. Every request paid that authorization, so no account holding less than 2.00 USD could use the alias at all, whatever it was actually sending.

Compounding it, /console/docs used hive-auto in both the curl and the Python quickstart, so a developer pasting the quickstart got a quota refusal on their first ever call.

Three findings that outlive this fix

None of these were in the issue. Each is independently worth knowing after this merges, so they are stated here rather than only in the commit that happens to touch them.

1. The session chat path took a credit hold but never ran the bounds. internal/chat/dispatch.go is the path Open WebUI chat actually uses, and it called startSettlement without ever calling EnforceVariablePriceBounds. A variable-price turn therefore went upstream with no request size cap and no completion ceiling, and the hold in front of it was covering a request nobody had bounded. That is a money hole on its own, independent of how the hold is sized: the API path has had those bounds since 2026-08-22 and this path silently did not.

2. Anything not in deploy/litellm/config.yaml silently does not exist on a synced route. route-openrouter-auto-live is generated from provider_routes by the config sync, and litellmconfig.mergeParams gives the generator ownership of exactly three keys: model, api_base, api_key. Every other key is merged in from the YAML. A route with a provider_routes row and no YAML entry therefore comes up with no extra_body at all, and nothing anywhere reports that. That is how hive-auto ended up serving with no provider.max_price, which is the rate half of its own credit hold's coverage proof. The general shape is worth remembering: a DB-managed route is only half-configured by the database.

3. usage.include was belt and braces here, not a repair, and the first version of this PR body said otherwise. The claim was that without it OpenRouter reports no cost and every request settles at the full hold through the fail-closed path. The live run disproved it: a real hive-auto turn settled with terminal_usage_confirmed true without the key set. The claim is downgraded rather than left standing, and the flag is still set, so the charge does not depend on that continuing to be true by default.

What this changes

1. The quickstart (c8d5531, first commit, reviewable on its own). The docs page picked the first catalog id starting with hive-, which on this deployment is the variable-price router. It now prefers a fixed-price alias that speaks chat, falling back through any Hive alias, the first catalog entry, and the seeded default. The pick moved into lib/quickstart-model.ts so it is testable without standing up the server component.

2. The hold (6a35740). It is now derived from the request in hand. Both quantities are already known at that point and both are upper bounds, not estimates: len(body) bytes is a rigorous upper bound on prompt tokens, and EnforceVariablePriceBounds has already written this request's effective completion ceiling into the body. Pricing those two at the same rate ceiling, through the same CreditsForUpstreamCost settlement uses, yields a hold that still provably covers the request and is about six times smaller for an ordinary chat turn (0.344 USD against 2.00 USD). The catalog figure stays the upper bound and the fallback whenever no per-request bound can be computed, so this can only ever shrink a hold toward a proven number, never past it.

3. Two things that had to be true for that to be honest rather than merely smaller, and were not.

The session chat path (internal/chat/dispatch.go) took the hold but never ran the bounds at all. A turn on a variable-price alias went upstream with no size cap and no completion ceiling, so the only thing in front of an arbitrarily large charge was a hold sized for a request nobody had checked. It now runs EnforceVariablePriceBounds before anything is held or dispatched, exactly as the API path does. A pass-through, and one comparison, for every fixed-price alias.

route-openrouter-auto-live, the route hive-auto actually serves, carried neither usage.include nor provider.max_price on the live box (verified 2026-08-29 by reading /etc/litellm/config.yaml inside the running container). The config sync generates that entry from provider_routes and owns only model, api_base and api_key, merging every other key from deploy/litellm/config.yaml field by field (litellmconfig.mergeParams), and the file had no entry for it to merge. max_price is the load-bearing absence: without it the rate half of the hold's coverage proof does not exist. usage.include turned out to be belt and braces rather than a repair, and the PR body said otherwise until the live run corrected it: a real hive-auto turn on 2026-08-29 settled with terminal_usage_confirmed true without it, so a cost does reach us on this path today. Setting it explicitly means the charge does not depend on that continuing to be true by default, which is why its beta twin sets it too. The entry is added with the same three keys that twin carries; the compose entrypoint reseeds the volume when the seed checksum changes, so this reaches the box on deploy.

Stated plainly, because it is a behaviour change and not only a safety net: with max_price set, an Auto request that resolves to a model above 3.00 USD per million prompt or 15.00 per million completion now fails rather than being served at an unbounded price. That ceiling is the one the flat 2.00 USD hold was always derived from (20260822_30_openrouter_auto_variable_pricing.sql); the live route simply never enforced it. The alias is preview visibility and its beta twin has run with the identical ceiling since 2026-08-22.

4. The refusal text. A reservation that does not fit is not a quota overrun, and telling a customer with a positive balance to check a plan Hive does not sell was wrong twice over. The 409 branch now says the available credit does not cover this request and what to do about it. The 429 branch, which is a rate limit on the reservation call and not a credit verdict at all, says so separately. Status, type and code are unchanged, because that is what SDKs branch on.

Which record is authoritative on hive-auto pricing

.wolf/decisions.md D-047 records hive-auto as converted to fixed price on 2026-08-23. The schema disagrees: 20260824_02_free_pool_router.sql flipped it back to upstream_actual the next day, and the live row confirms it.

 alias_id  |  pricing_mode   | input_price_credits | output_price_credits | reservation_estimate_credits
-----------+-----------------+---------------------+----------------------+------------------------------
 hive-auto | upstream_actual |                     |                      |                   2000000000

The schema is authoritative. It is the later change, the running system agrees with it, and the whole variable-price mechanism this PR touches only exists because the alias is upstream_actual. D-047 is stale on this point and should be amended to say so. This PR does not change the pricing mode; it changes how the hold that mode requires is sized.

Live proof

Captured against the demo box at https://chat-hive.scubed.co as qa-tester@hive.test, the account the defect was reported on, whose composer footer reads $0.455 remaining. Session minted through the admin one-time-token flow; no password set, reset or rotated. Screenshots are posted as inline images on this PR; the full capture log and the ledger row are in docs/proof/hive-auto-reservation-1372-2026-08-29/.

Before, the box on main: Hive Auto selected in the picker, "Say hello in one sentence." answered with You exceeded your current quota, please check your plan and billing details. directly above You've used $0.00287 today · $0.455 remaining.

After, the same box with this branch's edge-api image: same account, same model, Hello. Hive auto works. and a 200 on /api/chat/completions.

The reservation for that turn, from public.credit_reservations joined to public.request_attempts:

created_at                    | model_alias | reserved_credits | consumed_credits | released_credits | status    | terminal_usage_confirmed
2026-08-29 07:25:48.632339+00 | hive-auto   |        344853600 |            43526 |        344810074 | finalized | t

The hold was 0.3449 USD, not 2.00. Hold and release balance to the credit: 43,526 consumed plus 344,810,074 released is exactly 344,853,600 reserved, nothing stranded. The turn really cost 43,526 credits, so the old hold demanded an authorization roughly 46,000 times the actual charge before serving it.

The box was restored to main's image immediately afterwards (docker inspect hive-edge-api-1 reports the pre-change build, container healthy) and the temporary build worktree removed. The LiteLLM half of this PR was verified by reading the running container's config, not by editing it; it reaches the box through the deploy's seed reconciliation.

Money-path invariants

  • All hold arithmetic is math/big rationals through CreditsForUpstreamCost, the same function settlement uses, so hold and charge carry the identical margin (7/5), credit unit (D-046), round-half-up (D-031) and one-credit floor (D-034, D-048). No float64 anywhere on the path.
  • No token class starts or stops being billed, so D-055 is untouched. This changes an authorization, not a rate and not a class.
  • Hold and release still balance: the hold is a single number handed to CreateReservation and released or finalized in full by the same paths as before. Nothing about settlement, release, or terminal-state handling moved.
  • The hold can only shrink toward a number the arithmetic proves, and falls back to the catalog envelope whenever it cannot compute one. An unsizable request (over the byte cap) holds the full envelope, then gets refused by the bounds a moment later.

Tests that can fail

  • TestASolventAccountIsNotRefusedForAnEnvelopeItIsNotUsing pins the reported condition: an ordinary turn on a 0.455 USD balance is admitted, where the envelope hold is over four times that balance. RED confirmed: mutating variablePriceRequestHold to always fall back reproduces the defect exactly, hold 2000000000 against a 455000000 balance.
  • TestTheSizedHoldStillCoversWhatTheRequestCanCost is the property that must not break. It re-derives the bound from provider.max_price in the LiteLLM config rather than from the Go constants the implementation uses, so a drift between the two fails here as well. RED confirmed: halving the computed hold fails three of its four cases, the fourth being covered by the endpoint floor.
  • TestAnAccountThatCannotAffordTheRequestIsStillRefused holds the other direction: a request at the size cap still exceeds that balance, and nothing drops below the endpoint floor.
  • TestVariablePriceCeilingsMatchTheLiteLLMConfig checks the Go rate constants against the YAML for both variable-price routes by name. Naming them is the point: the live config had route-openrouter-auto-live with no max_price at all and nothing failed.
  • TestTheHoldProvablyCoversTheWorstBoundedRequest is unchanged and still guards the catalog envelope against the config ceiling and the request bounds.

Three existing tests asserted the old flat hold and were updated rather than deleted, each keeping its original intent: the two settlement tests now assert the derived figure as an exact magnitude (344274000, arithmetic spelled out at autoHoldForAnEmptyBody) instead of 2000000000, and TestReservationCreditsCannotUnderReserve now exercises the envelope through the unsizable-body fallback. TestSessionChatRefusesWhenTheAccountCannotPay dropped credit from its leak list, with the reasoning in the comment: it is Hive's own billing unit and the word the console uses, and forbidding it is what left OpenAI's misleading sentence in place. Provider, route name, currency and balance remain forbidden.

Full go test ./apps/edge-api/... ./apps/control-plane/... -count=1 -short is green.

Buglog entry

{"id":"bug-1372-hive-auto-flat-reservation-hold","date":"2026-08-29","title":"Hive Auto unusable below 2.00 USD: flat envelope-sized credit hold refused every request","error_message":"You exceeded your current quota, please check your plan and billing details. (429 insufficient_quota) shown while the account held 0.455 USD","root_cause":"hive-auto is priced upstream_actual, so the gateway takes an up-front credit hold instead of charging a catalog rate. reservation_estimate_credits was a flat 2000000000 (2.00 USD), the price of the largest request the variable-price bounds allow, and every request paid that authorization regardless of its actual size. control-plane's enforcePolicy refused any account whose balance was under it. Three silent enablers: the session chat path never ran EnforceVariablePriceBounds at all, so its hold covered a request with no size cap or completion ceiling; route-openrouter-auto-live carried no provider.max_price, so the rate half of the hold's coverage proof did not exist; and it carried no usage.include, so OpenRouter reported no cost and every settlement fell through to charging the full hold unconfirmed. The refusal text was OpenAI's canonical insufficient_quota sentence, which named a plan Hive does not sell.","fix":"Size the hold from the request: len(body) as a prompt-token upper bound plus the effective completion ceiling already written into the body, priced at the route's own rate ceiling through CreditsForUpstreamCost, capped by the catalog envelope and floored by the endpoint default. Run EnforceVariablePriceBounds on the session chat path. Add route-openrouter-auto-live to deploy/litellm/config.yaml with usage.include and provider.max_price. Split the refusal text by status and say what happened.","tags":["billing","reservation","upstream_actual","hive-auto","litellm-config","chat","error-message","money-path"],"files":["apps/edge-api/internal/inference/pricing.go","apps/edge-api/internal/chat/dispatch.go","apps/edge-api/internal/chat/billing.go","apps/edge-api/internal/inference/reservation_guard.go","deploy/litellm/config.yaml","apps/web-console/lib/quickstart-model.ts"],"issue":1372}

🤖 Generated with Claude Code

https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1

/console/docs picked the first catalog id starting with "hive-" for both
the curl and the Python quickstart. On this deployment that is hive-auto,
the variable-price Auto Router, whose up-front credit hold is sized for
the worst request the bounds allow rather than for the request in hand. A
developer who pasted the quickstart got a quota refusal on their first
ever call, with credit still on the account.

Pick a fixed-price alias that speaks chat instead, falling back through
any Hive alias, then the first catalog entry, then the seeded default, so
the snippets stay runnable when the catalog cannot be read at all. The
capability check keeps an embeddings-only alias out of a
/chat/completions snippet, which is the other way "first hive- id wins"
hands someone a sample that cannot run.

The pick moves into lib/quickstart-model.ts so it can be tested without
standing up the whole server component. The first test pins the live
catalog ordering that produced the bug.
…(issue #1372)

Selecting Hive Auto in the chat model picker made every subsequent turn
fail with "You exceeded your current quota, please check your plan and
billing details" while the same screen showed 0.455 USD remaining. Over
the API the alias answered 429 insufficient_quota for every request.

hive-auto is priced upstream_actual, so there is no catalog rate to
charge against and the gateway takes a credit hold up front instead. That
hold was a flat 2.00 USD equivalent, the price of the LARGEST request the
variable-price bounds allow: a full 256 KiB body generating a full 16,384
tokens at the route's 3.00 and 15.00 USD per million rate ceiling. Every
request paid that authorization, so no account holding less than 2.00 USD
could use the alias at all, whatever it was actually sending.

The hold is now derived from the request in hand. Both quantities are
already known at that point and both are upper bounds, not estimates:
len(body) bytes is a rigorous upper bound on prompt tokens, and
EnforceVariablePriceBounds has already written this request's effective
completion ceiling into the body. Pricing those two at the same rate
ceiling, through the same CreditsForUpstreamCost settlement uses, yields a
hold that still provably covers the request and is about six times smaller
for an ordinary chat turn. The catalog figure stays the upper bound and
the fallback whenever no per-request bound can be computed, so this can
only ever shrink a hold toward a proven number, never past it.

Three things had to be true for that to be honest rather than merely
smaller, and two of them were not.

The session chat path took the hold but never ran the bounds at all, so a
turn on a variable-price alias went upstream with no size cap and no
completion ceiling. It now runs EnforceVariablePriceBounds before anything
is held or dispatched, exactly as the API path does. A pass-through, and
one comparison, for every fixed-price alias.

route-openrouter-auto-live, the route hive-auto actually serves, carried
neither usage.include nor provider.max_price on the live box. The config
sync generates that entry from provider_routes and owns only model,
api_base and api_key, merging every other key from deploy/litellm/config.yaml
field by field, and the file had no entry for it to merge. Both absences
are silent and both are expensive: without usage.include OpenRouter
reports no cost and every request settles through the fail-closed path at
the full hold, and without max_price the rate half of the hold's coverage
proof does not exist. The entry is added here with the same three keys its
beta twin carries.

The rate ceilings are now also restated as Go constants, because the
request path cannot read the YAML, and TestVariablePriceCeilingsMatchTheLiteLLMConfig
fails when the restatement drifts from the file, for both variable-price
routes by name.

The refusal text is fixed too. A reservation that does not fit is not a
quota overrun, and telling a customer with a positive balance to check a
plan Hive does not sell was wrong twice over. The 409 branch now says the
available credit does not cover this request and what to do about it; the
429 branch, which is a rate limit on the reservation call and not a credit
verdict at all, says so separately. Status, type and code are unchanged,
because that is what SDKs branch on.

Tests. TestASolventAccountIsNotRefusedForAnEnvelopeItIsNotUsing pins the
reported condition: an ordinary turn on a 0.455 USD balance is admitted,
where the envelope hold is four times that balance. Mutating
variablePriceRequestHold back to the old behaviour makes it fail.
TestTheSizedHoldStillCoversWhatTheRequestCanCost is the property that must
not break, re-derived from the rate ceiling in the LiteLLM config rather
than from the Go constants; halving the computed hold makes it fail on
three of its four cases. TestAnAccountThatCannotAffordTheRequestIsStillRefused
holds the other direction: a request at the size cap still exceeds that
balance, and nothing drops below the endpoint floor.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 29, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 8 minutes.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 86a4a84f-0b15-46c2-b690-e36f548b6d1f

📥 Commits

Reviewing files that changed from the base of the PR and between 096c208 and 323e457.

📒 Files selected for processing (17)
  • apps/edge-api/internal/chat/billing.go
  • apps/edge-api/internal/chat/dispatch.go
  • apps/edge-api/internal/chat/settlement_test.go
  • apps/edge-api/internal/inference/orchestrator.go
  • apps/edge-api/internal/inference/pricing.go
  • apps/edge-api/internal/inference/reservation_guard.go
  • apps/edge-api/internal/inference/reservation_sizing_test.go
  • apps/edge-api/internal/inference/stream.go
  • apps/edge-api/internal/inference/stream_responses.go
  • apps/edge-api/internal/inference/upstream_cost_settlement_test.go
  • apps/edge-api/internal/inference/upstream_cost_test.go
  • apps/edge-api/internal/inference/variable_price_bounds_test.go
  • apps/web-console/__tests__/quickstart-model.test.ts
  • apps/web-console/app/console/docs/page.tsx
  • apps/web-console/lib/quickstart-model.ts
  • deploy/litellm/config.yaml
  • docs/proof/hive-auto-reservation-1372-2026-08-29/capture-log.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Adversarial review, stream 1 of 2: Antigravity (gemini-3.1-pro-high, effort high)

Framed as a pre-merge review of this repository's own PR, with the seven money-path questions this diff raises stated up front. Full verdict below, then my responses.

Findings

1. Under-reservation and arithmetic. Checked the order of operations across chat/dispatch.go, orchestrator.go, stream.go and stream_responses.go. EnforceVariablePriceBounds runs before startSettlement / ReservationCredits in all four, and the bounded body is what reaches the sizing. On the downstream rewrite: model grows by 17 bytes when the alias is replaced by the route name, but the upstream tokenises only the messages array and ignores model, so len(body) remains a robust upper bound. No way found for the new hold to be smaller than what the request can cost.

2. Call sites. All four pass the bounded body and the correct endpoint. Line-by-line confirmation of the reassignment preceding each call.

3. Hold and release balance. ReleaseReservationInput carries a reservation id and no amount, so the ledger releases exactly what was locked. The fail-closed settlement path charges heldCredits, which is now the smaller resized hold, so a missing upstream cost fails closed at the resized bound instead of at the 2.00 USD envelope.

4. Overflow, rounding, float64, nil. big.Rat throughout, no float64. CreditsForUpstreamCost uses QuoRem with explicit half-up rounding and refuses a value that does not fit in int64 before casting. ReservationEstimateCredits is nil-checked before dereference.

5. Assertions that would still pass with the fix reverted. TestAnAccountThatCannotAffordTheRequestIsStillRefused and the table assertions in TestTheSizedHoldStillCoversWhatTheRequestCanCost are satisfied by the old flat 2e9 hold.

6. provider.max_price on route-openrouter-auto-live. OpenRouter will refuse to route to a model above the 3.00 / 15.00 ceiling, which intentionally narrows what the Auto router may pick. Judged as the desired behaviour, since it is what makes the hold's coverage proof structural rather than nominal. On reconciliation: litellmconfig.mergeParams preserves extra_body and the generator owns only model, api_base and api_key, so the YAML change survives the sync and reaches the box.

7. SDK contract and disclosure. Status, type and code unchanged in both writers. The new message discloses no provider, route, currency amount or balance figure.

Verdict: LGTM.

My responses

On finding 5, rebutted, with a code change so it is not re-raised. Those two tests guard the under-reservation direction, and the old hold over-reserved, so over-reserving has to satisfy them. A guard against holding too little that went red when the hold was too large would be the wrong guard. The revert is caught by TestASolventAccountIsNotRefusedForAnEnvelopeItIsNotUsing, which I confirmed goes red by mutating variablePriceRequestHold to always fall back (hold 2000000000 against a 455000000 balance). The opposite direction is caught by these two: halving the computed hold fails the coverage table on three of its four cases, the fourth sitting under the endpoint floor. I have added a comment on the test file stating exactly this, so the next reader does not spend the same time on it.

On finding 6, accepted and recorded as a deliberate consequence. Without a rate ceiling on the route hive-auto actually serves, the hold has no coverage proof at all: the flat 2.00 USD figure was itself derived from these same two numbers, in 20260822_30_openrouter_auto_variable_pricing.sql, and the live route simply never enforced them. Adding it makes the running system match the design that was already written down. The narrowing is real and worth stating plainly: an Auto request that resolves above 3.00 USD per million prompt or 15.00 per million completion now fails rather than being served at an unbounded price. The alias is preview visibility, and its beta twin has run with the identical ceiling since 2026-08-22.

Findings 1 through 4 and 7: agreed, no change.

Antigravity, finding 5: two of the three sizing tests pass with the fix
reverted. Rebutted rather than changed, with a comment on the test file so
the next reader does not spend the same time on it. They guard the
under-reservation direction and the old hold over-reserved, so
over-reserving has to satisfy them; the revert is caught by
TestASolventAccountIsNotRefusedForAnEnvelopeItIsNotUsing, and shrinking too
far is caught by the coverage table.

CodeRabbit, major: the Hive fallback rung in pickQuickstartAlias could
still return an embeddings-only alias if the catalog listed one ahead of
every chat alias. Accepted; the chat check now applies to both rungs, with
a regression test for that ordering.

CodeRabbit, major: a float64 conversion in a test log line. Accepted;
big.Rat.FloatString instead. Nothing in this repo converts credits to money
through a float, log line included.

CodeRabbit, critical (duplicate test declaration) is a false positive:
there is exactly one declaration of TestAFixedPriceAliasKeepsTheFlatEndpointHold
and the package compiles. CodeRabbit, major (move the rate ceilings and the
model slug to environment variables) is rebutted in the PR thread; the
slug being a literal is a recorded decision from issue #689, and an
environment variable for the ceilings would add a silent drift path where
a cross-file test currently forbids one.

The LiteLLM comment on usage.include is corrected to match what the live
run actually showed. A hive-auto turn settled with terminal_usage_confirmed
true without it, so setting it is belt and braces rather than a repair.
max_price remains the load-bearing half.
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Adversarial review, stream 2 of 2: CodeRabbit CLI

coderabbit review --base main --committed, 16 files reviewed, 5 findings: 1 critical, 4 major. Each below with my response. Two accepted and fixed in 323e457d, three rebutted.


Critical, Functional Correctness, reservation_sizing_test.go:201-209: "Remove the duplicate test declaration. TestAFixedPriceAliasKeepsTheFlatEndpointHold is declared again at lines 210-218. Go rejects duplicate package-level function declarations, so this package cannot compile."

Rebutted, false positive. There is exactly one declaration:

$ grep -c 'func TestAFixedPriceAliasKeepsTheFlatEndpointHold' apps/edge-api/internal/inference/reservation_sizing_test.go
1

at line 210, in a 218-line file. The package compiles and go test ./apps/edge-api/... ./apps/control-plane/... -count=1 -short is green. This reads like the added block being counted once from the diff and once from the file.


Major, Functional Correctness, quickstart-model.ts:37: "Keep the Hive fallback chat-capable. If the catalog lists hive-embedding-default before hive-auto, this selects the embeddings-only alias."

Accepted, fixed. Correct, and it is the same class of bug as the one this function exists to stop: rung 2 could hand a /chat/completions snippet an embeddings id. The chat check now applies to both rungs, and there is a regression test for exactly that ordering (keeps that fallback chat-capable when an embeddings alias sorts first). Five tests in that file now, all passing.


Major, Maintainability, reservation_sizing_test.go:78-79: "Keep diagnostic money conversion exact. Line 79 converts credits to float64 before formatting USD."

Accepted, fixed. new(big.Rat).SetFrac64(hold, CreditsPerUSD).FloatString(4) instead. The rule is that floats never touch money paths, and a log line that prints a credit figure as dollars is close enough to that path that the exception is not worth having.


Major, Maintainability, pricing.go:319-320: "Move the pricing ceilings into validated environment configuration. A pricing change can then deploy with a different LiteLLM ceiling and reservation ceiling."

Rebutted. The failure mode named is real, and an environment variable is the one shape that cannot prevent it. The number these constants have to agree with is provider.max_price in deploy/litellm/config.yaml, which is a tracked deployment artifact, not an environment value: it is itself a literal, deliberately (see the block comment above route-openrouter-auto-beta). Moving the Go side to an env var would put the two halves in places nothing compares, and drift would surface as a mispriced hold in production. Keeping it a constant is what lets TestVariablePriceCeilingsMatchTheLiteLLMConfig read the YAML and fail the build when they disagree, which is the same cross-file guard TestTheHoldProvablyCoversTheWorstBoundedRequest has held over these numbers since 2026-08-22. A silent runtime drift traded for a loud build failure is the wrong direction.


Major, Maintainability, config.yaml:424: "Move the model slug to a validated environment variable. This literal model slug violates the environment-only configuration rule."

Rebutted, and this one is a recorded decision rather than a preference. From the comment on the twin route sixteen lines above the flagged one:

The model is a LITERAL, not os.environ/..._MODEL like every other OpenRouter route here. Issue #689 is what an env var costs on a priced route: the catalog priced one model while the deployment quietly called another. For an alias billed at actual cost the slug IS the product decision, and it must not be repointable without a code review.

route-openrouter-auto-live is billed at actual cost for the same reason and takes the same treatment. The live box already runs this slug as a literal. Making it repointable by environment is the specific mistake #689 records.


Both streams have now run. No stream was skipped.

@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Visual proof

Live demo box, chat-hive.scubed.co, signed in as qa-tester whose composer footer reads $0.455 remaining. Before: Hive Auto refuses every turn with "You exceeded your current quota, please check your plan and billing details". After, same account and same model with this branch's edge-api running on the box: "Hello. Hive auto works." The reservation for that turn held 344,853,600 credits instead of 2,000,000,000, consumed 43,526 and released 344,810,074, which balances to the credit. Full capture log and ledger row in docs/proof/hive-auto-reservation-1372-2026-08-29/.

pr1378-20260829073025-17318-1372-before-hive-auto-quota-refusal.png

pr1378-20260829073036-9531-1372-after-hive-auto-answers.png

@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Review stream availability

Recorded so an absent stream is not read as a clean pass.

  • Antigravity (gemini-3.1-pro-high, effort high): RAN. Findings and my responses above.
  • CodeRabbit: RAN, through the CLI. coderabbit review --base main --committed, 5 findings, all answered above. The GitHub CodeRabbit app posted "Review limit reached" on this PR, so the bot's own inline pass did not happen; the CLI run is what covers this stream, not the app comment.
  • Codex code review: SKIPPED, not clean. chatgpt-codex-connector replied "You have reached your Codex usage limits for code reviews" and produced no review at all. This diff touches the money path, so that is a stream I would have wanted; treat its absence as missing coverage rather than as agreement.

No inline review threads are open on this PR (repos/.../pulls/1378/comments returns zero).

@sakibsadmanshajib
sakibsadmanshajib merged commit 027b375 into main Aug 29, 2026
30 checks passed
@sakibsadmanshajib
sakibsadmanshajib deleted the fix/1372-hive-auto-reservation branch August 29, 2026 07:41
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Visual proof

Post-deploy proof, pipeline stage 10. deploy-demo-box run 33241444493 for merge commit 027b375 completed successfully; edge-api is on the deploy's own image cafdfda15245, and the deployed /etc/litellm/config.yaml now carries max_price 3/15 and usage.include on route-openrouter-auto-live, merged in beside the sync-owned api_base exactly as intended. Same account, same model, $0.45 remaining: Hive Auto answers "hive auto post deploy ok". Reservation for that turn held 346,844,400 credits, consumed 40,110, released 346,804,290, balancing to the credit, terminal_usage_confirmed true.

pr1378-20260829075414-19501-1372-postdeploy-hive-auto-on-merged-main.png

sakibsadmanshajib added a commit that referenced this pull request Aug 29, 2026
…1384)

## What this is

A decision ledger correction, not code. Touches `.wolf/decisions.md`
only, one appended line.

## The stale fact

D-047 (2026-08-23) records that both `hive-auto` and `hive-default` were
converted to fixed price at half their effective customer price. That
stopped being true of `hive-auto` the very next day.

`supabase/migrations/20260824_02_free_pool_router.sql:307` sets
`pricing_mode = 'upstream_actual'` on `hive-auto`, with
`reservation_estimate_credits = 2000000000`. The live database confirms
`upstream_actual` today.

The schema is authoritative on three grounds: it is the later change,
the running system agrees with it, and the reservation mechanism PR
#1378 repairs only exists in `upstream_actual` mode at all.
`hive-default` is unaffected by this correction and remains fixed price
exactly as D-047 records.

This was not a paperwork problem. An agent reasoning from D-047 alone
tonight concluded `hive-auto` was fixed price and missed the flat 2.00
USD hold that made the alias unusable below a two dollar balance, filed
as issue #1372 and repaired in PR #1378.

## What changed

Appended `D-059` to `.wolf/decisions.md`, per the ledger's own
supersede-in-place protocol: D-047 stays exactly as written (it was a
correct record of what was decided on 2026-08-23), and D-059 records
that the `hive-auto` half of it was reversed the following day, with the
schema line, the live confirmation, and the issue and PR that make the
correction matter in practice.

Sources cited in the entry: issue #1372, PR #1378,
`supabase/migrations/20260824_02_free_pool_router.sql`.

## CI expectations

`.wolf/decisions.md` matches the `*.md` arm of the changed-files
allow-list in `.github/workflows/ci.yml` (the `changes` job), so the
required checks should report green without running the full test
matrix. This is a docs-only ledger correction, not a code or schema
change, so no functional verification is applicable.

## Review

No adversarial review streams requested for this PR. It is a
documentation correction against verified facts (migration file, live
database, linked issue and PR), not a code change; a review pipeline
pass is not being skipped by omission, it is being explicitly declared
not applicable here.

Fixes documentation staleness surfaced while working issue #1372 / PR
#1378.
sakibsadmanshajib added a commit that referenced this pull request Aug 29, 2026
## Summary

This is the batched buglog follow-up for the pull requests merged to
`main` on 2026-08-29. Its diff is `.wolf/buglog.jsonl` and nothing else.

Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or
failed build must be logged, but the line may never be appended on a fix
branch. `merge=union` in `.gitattributes` resolves concurrent appends
locally and is ignored by GitHub's server side merge, so two branches
that both appended land in hard conflict there. An unmergeable pull
request gets no `refs/pull/N/merge`, no `pull_request` run and therefore
zero checks, and the required status gate then blocks the merge for a
reason the page never states (issue #873). Each fix accordingly carried
its entry in its own pull request body, and this pull request copies
them onto `main` in one batch, which the protocol explicitly prefers
over one pull request per entry.

## Scope examined

Fifty nine pull requests merged to `main` on 2026-08-29. Forty eight of
them carried at least one entry, for eighty two entries in total. Thirty
two of those were already on `main` and are skipped, leaving fifty
appended here from thirty four pull requests.

The largest block of skips comes from #1342, the equivalent batch for
the 2026-08-28 merges, which merged earlier the same day and already
landed thirty six entries covering #1257, #1268, #1276, #1277, #1287,
#1292, #1293, #1294, #1296, #1301, #1303, #1305, #1313, #1335 and #1337.

## What landed

Fifty entries appended, one JSON object per line, append only. The 232
pre-existing lines are byte identical to `origin/main` (verified by
hashing the first 232 lines of the result against the base file). Every
line in the resulting file parses as JSON and carries `error_message`,
`root_cause`, `fix` and `tags`.

| Source | Entries |
|---|---|
| #1083 | 2 |
| #1277 | 1 |
| #1278 | 1 |
| #1298 | 1 |
| #1334 | 1 |
| #1336 | 3 |
| #1343 | 1 |
| #1346 | 1 |
| #1351 | 1 |
| #1365 | 2 |
| #1368 | 1 |
| #1369 | 1 |
| #1371 | 3 |
| #1375 | 3 |
| #1376 | 1 |
| #1378 | 1 |
| #1379 | 2 |
| #1388 | 5 |
| #1389 | 3 |
| #1390 | 2 |
| #1393 | 1 |
| #1394 | 1 |
| #1410 | 1 |
| #1417 | 1 |
| #1421 | 1 |
| #1423 | 1 |
| #1424 | 1 |
| #1426 | 1 |
| #1429 | 1 |
| #1431 | 1 |
| #1433 | 1 |
| #1434 | 1 |
| #1436 | 1 |
| #1439 | 1 |

Entries are copied verbatim from their source pull request bodies.
Nothing was rewritten, no field was invented, and no field was added. No
JSON needed repair: all eighty two extracted entries parsed on the first
attempt and all four required fields were present on every one.

## Merged pull requests that carried no entry

Eleven of the fifty nine. Recorded here because the gap is itself the
useful signal.

| Pull request | Title | Assessment |
|---|---|---|
| #1013 | chore(deps): bump the go-minor-patch group across 1 directory
with 4 updates | Dependabot bump, no defect fixed, no entry expected |
| #1015 | chore(deps): bump the go-minor-patch group across 1 directory
with 6 updates | Dependabot bump, no entry expected |
| #1016 | chore(deps): bump golang from 1.26-alpine to 1.27-alpine in
/deploy/docker | Dependabot bump, no entry expected |
| #1218 | chore(deps): bump postcss from 8.5.19 to 8.5.26 in
/apps/desktop | Dependabot bump, no entry expected |
| #1219 | chore(deps): bump golang.org/x/crypto from 0.41.0 to 0.52.0 in
/apps/control-plane | Dependabot bump, no entry expected |
| #1342 | chore: batch buglog entries for the 2026-08-28 merges | The
previous batch pull request itself, correctly carries no entry of its
own |
| #1364 | chore: remove four dead skills and record the patterns that
cost time | Protocol gap. The body records patterns that cost time,
which is the shape of a buglog entry, but none was written as one |
| #1383 | test: retire stale expected-failure markers, restore the ones
that are true (#1381, #1382, #1324) | Protocol gap. Stale `it.fails`
markers reading as red is a real defect that was fixed here and should
have carried an entry |
| #1384 | docs: correct D-047, hive-auto reverted to variable pricing
(D-059) | Decision ledger correction, arguably a documentation defect,
no entry written |
| #1387 | chore(deps): bump next from 15.5.23 to 16.3.3 in
/apps/agent-console | Dependabot bump, no entry expected |
| #1398 | docs: rescue the 2026-08-25 parity captures and add the
2026-08-29 QA matrix evidence | Documentation and evidence rescue, no
entry written |

Six of the eleven are Dependabot bumps and one is the previous batch, so
the genuine protocol gaps are #1364, #1383, #1384 and #1398. Of those,
#1383 is the one worth a follow-up: it fixed a real defect class (a
stale expected-failure marker reads as a red "Expect test to fail" and
gets dismissed as pre-existing) and left no record.

## Entries skipped as already present

Thirty two. Thirty of them matched an entry already on `main` on
`error_message`, `id` or `fix`. Two more from #1278 are semantic
duplicates that an exact match would have missed, and were skipped after
reading the landed entries they duplicate:

- #1278's `streaming content_block_start omits text field` entry is
covered by the consolidated
`bug-2026-08-28-anthropic-sdk-wire-conformance` entry landed from #1296,
whose root cause names the same `omitempty` on
`StreamContentBlock.Text`.
- #1278's `GET /v1/models leaked an upstream provider name` entry is
covered by `BUG-1284`, landed from #1300, which names the same
`public.model_aliases.summary` publication path.

#1278's third entry, on `top_k` forwarding producing a 400, is not
covered anywhere on `main` and is appended here. #1342 recorded #1278 as
fully "merged into #1296", which was accurate for two of its three
entries.

## Note on entry quality

One appended entry is thin: #1277's parity re-score record carries
`error_message` of `n/a` and a root cause of "console had no
privacy/data-policy surface at all". It is a parity gap record rather
than a defect record. It is included exactly as written rather than
embellished, per the protocol's preference for the author's own words.

## Test plan

- [x] Branch cut fresh from `origin/main`, diff is `.wolf/buglog.jsonl`
and nothing else
- [x] First 232 lines byte identical to the base file (md5 match)
- [x] All 282 resulting lines parse as JSON and carry `error_message`,
`root_cause`, `fix` and `tags`
- [x] No `.wolf/` telemetry (`anatomy.md`, `memory.md`,
`token-ledger.json`, `hooks/_session.json`, `buglog.json`) in the commit
- [ ] The six required checks report green via the inert path allowlist
in `.github/workflows/ci.yml`

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Hive Auto is unusable below a 2.00 dollar balance and the docs quickstart recommends it

1 participant