Repository navigation
fix(bedrock_mantle): price gpt-5.6-sol at its promotional rates - #41586
kusumakarb wants to merge 1 commit into
Conversation
|
| "model, input_cost, output_cost", | ||
| [ | ||
| ("openai.gpt-5.6-sol", 5.5e-06, 3.3e-05), | ||
| ("openai.gpt-5.6-sol", 4.4e-06, 2.2e-05), |
There was a problem hiding this comment.
Vendor pricing lacks provenance
These vendor-dependent assertions omit the required source and date. Add both before merging so stale pricing is distinguishable from regressions
Context Used: CLAUDE.md (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Good catch, fixed in e7e941a.
Added the source and date directly above the parametrized rates: the three AWS model cards, which column they come from (Commercial In-Region, short context <=272K), the date read, and a note that sol's rates are promotional (announced 2026-08-21, held at least through 2026-11-21) so the staleness has a known horizon.
Worth noting the other half of that CLAUDE.md rule is already covered: this PR also adds test_mantle_matches_in_region_converse_pricing, which pins no vendor literal at all. It asserts the invariant instead -- every cost field on bedrock_mantle/openai.gpt-5.6-* equals its us.openai.gpt-5.6-* counterpart, since mantle serves these In-Region only and the model cards price In-Region and Geo CRIS identically. That is the test that would have caught this bug: the converse keys were updated for the promo and the mantle key was missed, and nothing asserted the two namespaces agree.
I left the literal table in place rather than converting it, since it predates this PR and rewriting it would widen the scope past the one-line price fix.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
The audit in #41597 lists this PR as "not verifiable". I think that rests on reading a different page than the one this PR cites, and the registry already contains the answer. 1. The placeholder problem is on a page this PR does not use#41597 says Bedrock pricing "rendered unresolved This PR cites the model card, which is static server-rendered HTML with no placeholders at all: Its published table, straight out of that HTML: Those are exactly the eight values in this PR. 2. The
|
| key | provider | input | cache read | cache write | output |
|---|---|---|---|---|---|
us.openai.gpt-5.6-sol |
bedrock_converse | 4.4e-06 | 4.4e-07 | 5.5e-06 | 2.2e-05 |
global.openai.gpt-5.6-sol |
bedrock_converse | 4e-06 | 4e-07 | 5e-06 | 2e-05 |
gpt-5.6-sol |
openai | 4e-06 | 4e-07 | 5e-06 | 2e-05 |
bedrock_mantle/openai.gpt-5.6-sol |
bedrock_mantle | 5.5e-06 | 5.5e-07 | 6.875e-06 | 3.3e-05 |
litellm already prices the same model on Bedrock at $4.40 / $22.00 under us., and already encodes the In-Region vs Global CRIS split at exactly the documented 1.1x. The bedrock_mantle key is the only one of the four still on pre-promotion rates.
So accepting this PR does not require verifying anything against AWS. It only requires agreeing that mantle and the us. profile are the same model at the same published price — which is what the parity test added here asserts, and what would have caught the divergence when the converse keys were updated and mantle was missed.
Corroboration
Separately, we reconcile our own Bedrock bill against per-call token counts. Dividing AWS Cost Explorer's per-usage-type dollars by our token counts returns $4.40 / $0.44 / $5.50 / $22.00 to the cent on two independent days. Happy to attach that if useful, though point 3 stands on its own.
|
@devin-ai-integration — correction request on the #41597 audit finding for this PR. The finding recorded was:
Three specific errors in that finding: 1. Wrong source URL. The audit body records reading
The placeholder problem is specific to the marketing pricing page, which resolves prices client-side. The model card is static server-rendered HTML. Both are reproducible: 2. The prices are present in that HTML and are the PR's values. Extracted from the model card, no JS required: All eight values in this PR come from the In-Region rows. 3. The The finding is also contradicted by the registry itself. On
The registry already prices this model on Bedrock at the values this PR proposes, under So the change is verifiable without consulting AWS at all: it makes the mantle key agree with the converse key for the same model at the same published In-Region price. The test added in this PR asserts exactly that parity across sol, terra and luna, which is the check that would have flagged the divergence when the converse keys were updated and mantle was missed. Please re-evaluate against the model card URL rather than the marketing pricing page. Full detail in my previous comment. Edit: the first command above originally read |
AWS reduced GPT-5.6 Sol pricing on 2026-08-21 (20% off input, 33.3% off output), available at least through 2026-11-21. The reduction was applied to the bedrock_converse keys (us.openai.gpt-5.6-sol, global.openai.gpt-5.6-sol) and to the openai key, but the bedrock_mantle/openai.gpt-5.6-sol key still carried the pre-promotion rates, so mantle callers saw reported cost run ~24-43% above what AWS billed, depending on their input/output mix. bedrock-mantle serves Sol In-Region only, and the model card prices In-Region and Geo CRIS identically, so the mantle key now matches the existing us. converse entry exactly: input 5.5e-06 -> 4.4e-06 ($4.40 / 1M) cache write (30m) 6.875e-06 -> 5.5e-06 ($5.50 / 1M) cache read 5.5e-07 -> 4.4e-07 ($0.44 / 1M) output 3.3e-05 -> 2.2e-05 ($22.00 / 1M) plus the matching >272k long-context tier. Terra and Luna were already correct and are unchanged. Ref: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-openai-gpt-56-sol.html Ref: https://aws.amazon.com/about-aws/whats-new/2026/08/bedrock-openai-gpt-56-sol-reduced-pricing/ Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
e7e941a to
e60dbc9
Compare
|
Rebased onto The conflict was in That also retires the second commit on this branch. It existed only to add a source and date above that parametrized table after Greptile flagged the missing provenance — with the literals gone, the annotation has nothing to annotate, so I dropped it rather than carry an empty change. Deleting the literals is the stronger version of the same fix. What remains is the registry change plus Verified after the rebase:
|
|
Absorbed into the rolling registry PR #41754 with the same eight values, your test, and a co-author credit. Thanks! |
TLDR
Problem this solves:
bedrock_mantle/openai.gpt-5.6-solstill has pre-promotion pricingus./global.converse keys were updated; mantle was missedHow it solves it:
us.converse prices stay in syncUser Flow
Before: a developer running GPT-5.6 Sol through bedrock-mantle sees spend far higher than their AWS bill
POST https://litellm-domain/v1/responseswithmodel: bedrock_mantle/openai.gpt-5.6-solhttps://litellm-domain/ui/?page=logsand see the request costed at $0.88 per 100k input + 10k outputAfter: the same request is costed at what AWS charges
POST https://litellm-domain/v1/responseshttps://litellm-domain/ui/?page=logsnow shows $0.66, matching the billRelevant issues
Affected release
Linear ticket
Pre-Submission checklist
Screenshots / Proof of Fix
AWS reduced Sol's price on 2026-08-21, at least through 2026-11-21. Per the model card, Commercial In-Region (which is what mantle serves -- mantle is In-Region only) is $4.40 input / $5.50 30m cache write / $0.44 cache read / $22.00 output per 1M tokens.
Two days of real Sol traffic on our AWS account. Dividing the Cost Explorer per-usage-type dollars by the token counts recovers exactly those rates:
Day totals over that same traffic:
Repro script used below:
Before (4b368bf)
LITELLM_LOCAL_MODEL_COST_MAP=True python proof.pyAfter (336db8b)
LITELLM_LOCAL_MODEL_COST_MAP=True python proof.pyType
🐛 Bug Fix
Caveats (if any)
Low
Final Attestation