Repository navigation
Conversation
…ompletions API bedrock-runtime serves the OpenAI models (gpt-5.4/5.5/5.6/6) on an OpenAI-compatible surface at /openai/v1/chat/completions, alongside Converse. Previously the bedrock provider translated every chat request into Converse, which is lossy for these models: it 400s on unsupported sampling params instead of dropping them, loses reasoning tokens, and hoists mid-conversation system messages (breaking implicit prompt caching). Add BedrockOpenAIChatConfig, the Chat Completions sibling of BedrockOpenAIResponsesConfig. It inherits OpenAIGPT5Config (the same reasoning param mapper the openai provider uses) and overrides only the endpoint URL, Bearer/SigV4 auth, and AWS-param stripping. Routing is data-driven from the model's supported_endpoints, via a shared bedrock_uses_native_openai_chat predicate used by the config selector, the completion dispatcher, and get_optional_params. An explicit bedrock/converse/<model> prefix stays an escape hatch to the Converse translation. Enables the surface for the 16 gpt-5.4/5.5/5.6/6 cross-Region inference-profile entries that advertise /v1/chat/completions.
…sponses API
bedrock-runtime's native chat surface rejects function tools while reasoning is
active on gpt-5.6/6 ("use /v1/responses or set reasoning_effort to 'none'"),
which Converse previously served. Rather than regress that path, extend the
existing responses_api_bridge_check (already used for openai/azure with the same
constraint) to fire for bedrock: a chat request with function tools + active
reasoning on a model that supports the native Responses surface is bridged to
/v1/responses, which handles tools + reasoning and returns proper tool_calls.
The version boundary is measured against bedrock-runtime (gpt-5.6+ reject; 5.4/5.5
serve tools with reasoning natively) via
bedrock_chat_rejects_function_tools_while_reasoning. Custom tools, reasoning_effort
"none", and non-tool requests stay on native chat completions.
|
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 2 · PR risk: 0/10 |
…; trim comment Address review: move the gpt-5.6/6 "native chat rejects function tools while reasoning" boundary out of a hardcoded version check and into a bedrock_chat_rejects_function_tools_while_reasoning flag on the price-map entries, read like the other data-driven Bedrock capabilities, so onboarding a model stays a JSON change. Add a wire-level test asserting the SigV4-signed streaming body carries stream:true, and drop a narrating dispatch comment.
|
@greptileai Pushed 25d0ab3 addressing the review: moved the tools+reasoning boundary to a price-map flag, verified the streaming-body concern is a false positive (added a wire-level test), and trimmed the narrating comment. Please re-review. |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…ta on native chat validate_environment passed caller headers through untouched, so a caller could set X-Amzn-Bedrock-Request-Metadata and have it signed and recorded in AWS billing / CloudTrail, forging identity attribution. Reuse the existing ownership merge (bedrock_request_metadata_headers + merge_bedrock_invoke_headers, as the Invoke OpenAI path does): the proxy-owned header is dropped from caller input (case-insensitively) and re-supplied only from trusted litellm_params.
…ster price-map key - Add coverage for the get_optional_params branch, the _complete_bedrock native dispatch, and predicate edge cases (None model, non-converse route). - Drop untyped **kwargs from BedrockOpenAIChatConfig.__init__ (BaseAWSLLM takes no kwargs; every call site is arg-less) to satisfy the typed-parameter guideline. - Register bedrock_chat_rejects_function_tools_while_reasoning in the schema generator's boolean key table and regenerate the price-map schema.
|
@greptileai Pushed through 55f819c: fixed the request-attribution forgery (proxy now owns the X-Amzn-Bedrock-Request-Metadata header via the existing ownership merge), typed |
… native chat
The model rides in the request body on the native /openai/v1/chat/completions
surface (not the URL like Converse/Invoke), and the generic handler merges
extra_body over the body before SigV4 signing. bedrock isn't in the extra_body
model-stripping path, so a key limited to one model could send
extra_body={"model": "<other>"} and invoke another model on the proxy's AWS
credentials. Pin the authorized model back in sign_request (post-merge, pre-sign),
covering both the SigV4 and Bearer paths.
|
@greptileai Pushed 582fe71 fixing the extra_body model-authorization bypass on the native chat surface (sign_request now pins the authorized model post-merge/pre-sign). Please re-review. |
|
Want your agent to iterate on Greptile's feedback? Try greploops. |
…tation - Add tests for custom_llm_provider, get_error_class, the non-reasoning param fallbacks, and the SigV4 sign_request path (brings the adapter to 100% patch coverage). - Move the model-pin suppression inline on the mutated line as `# rebind-ok` (LIT011) and drop the misplaced block comment.
|
@greptileai Pushed 4466a96. Also merged the latest main (cost-map guard was failing on the merge because the schema was stale against the merged cost map — regenerated), brought the adapter to 100% patch coverage, and added the |
|
@greptileai Pushed c7d2829: dropped the redundant model-pin comment and tightened the signed-body assertion to check the exact |
The cost-map-guard runs under pull_request_target with the BASE branch's schema generator (PR code is never executed, only the cost map is read as data), so a price-map key a PR introduces can never be classified by that generator -- the guard fails every run with "Unclassified keys". Move the gpt-5.6+ boundary back into bedrock_chat_rejects_function_tools_while_reasoning as a version check on the normalized OpenAI model name, matching the openai/azure arms already in the same bridge, and drop the flag from the 12 cost-map entries, the schema generator, and the schema. No new cost-map key, so the guard passes.
|
seems like a dup of #40775 |
|
@krrish-berri-2 mostly, but #40775 doesn't include GPT 6 and it routes function-tools + reasoning to converse rather than responses. I can close this one and create a followup after the #40775 lands |
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer sends a chat request with a temperature to a Bedrock GPT-6 model and gets a hard 400 instead of an answer.
model: bedrock/global.openai.gpt-6-luna,temperature: 0.3,reasoning_effort: low.../model/global.openai.gpt-6-luna/converseThis model doesn't support the temperature field. Remove temperature and try again.After: the same request succeeds; the unsupported sampling param is dropped and the native surface answers.
.../openai/v1/chat/completions, having droppedtemperature(reasoning model) and keptreasoning_effort.../openai/v1/responsesand comes back with normaltool_calls, unchanged from the caller's perspectiveRelevant issues
Type
🆕 New Feature
Screenshots / Proof of Fix
Shared setup: AWS profile with Bedrock access in us-east-1;
LITELLM_LOCAL_MODEL_COST_MAP=Trueto load the price-map entries in this PR.model="bedrock/global.openai.gpt-6-luna".drop_params=True. Same request bodies on both sides.Before (a11a93f)
chat non-stream (temperature + reasoning)
litellm.completion(..., temperature=0.3, reasoning_effort="low", max_completion_tokens=400).../model/global.openai.gpt-6-luna/converselitellm.BadRequestError: BedrockException - {"message":"This model doesn't support the temperature field. Remove temperature and try again."}function tools + reasoning
litellm.completion(..., tools=[get_weather], reasoning_effort="low", max_completion_tokens=600).../model/global.openai.gpt-6-luna/converseget_weathertool call (Converse served this)After (5678141)
chat non-stream (temperature + reasoning)
.../openai/v1/chat/completions(temperaturedropped,reasoning_effortkept)content: "ok"chat stream
stream=True.../openai/v1/chat/completions"stream ok"function tools + reasoning
.../openai/v1/responsesfinish_reason: tool_calls,object: chat.completion,get_weather({"city":"Paris"})— unchanged shape for the caller, no ConverseCaveats (if any)
Low
QA runbook
supported_endpoints: ["/v1/chat/completions", "/v1/responses"]temperature=0.5+reasoning_effort=low; expecttemperaturedropped, request to.../openai/v1/chat/completionstools+reasoning_effort=low; expect the request to.../openai/v1/responsesand normaltool_callsbackbedrock/converse/<model>; expect it stays on Converse