Skip to content

fix(bedrock): avoid invalid auxiliary request parameters - #124937

Closed
JoaoMarcos44 wants to merge 1 commit into
NousResearch:mainfrom
JoaoMarcos44:fix/124923-bedrock-aux-probe
Closed

JoaoMarcos44 wants to merge 1 commit into
NousResearch:mainfrom
JoaoMarcos44:fix/124923-bedrock-aux-probe

Conversation

@JoaoMarcos44

@JoaoMarcos44 JoaoMarcos44 commented Sep 27, 2026 •

Copy link
Copy Markdown

Summary

Fixes #124923 with a narrow Bedrock-only request-shaping change plus regression coverage.

Two failures are addressed:

  1. AnthropicBedrock structured-output 400s. The auxiliary Anthropic adapter currently translates OpenAI-style response_format into Anthropic output_config.format. That is valid on native/compatible Anthropic Messages routes, but AWS Bedrock's Anthropic endpoint rejects the field with output_config.format: Extra inputs are not permitted.

    This PR marks the AnthropicBedrock route explicitly when the Bedrock client is constructed and suppresses the translation only for that route. Native Anthropic / compatible Messages routes continue to receive structured output, and the Bedrock Mantle/OpenAI route is untouched.

  2. Bedrock context probe uses an invalid output minimum. probe_bedrock_context_length() sends maxTokens=8; OpenAI-on-Bedrock models reject output caps below 16 before the prompt-length validation runs, so Hermes never reaches the parseable length error and falls back to the generic 128K context value. The probe now sends maxTokens=16.

Root cause

The structured-output capability decision was being made at the generic Anthropic adapter layer without carrying the fact that this specific Messages client was created for AnthropicBedrock. A provider-wide Bedrock capability flag would be too broad because the same provider slug also owns Mantle/OpenAI and Converse paths.

The context probe independently used an output cap below the minimum accepted by OpenAI-on-Bedrock models, masking the validation error it intentionally provokes.

Changes

  • agent/auxiliary_client.py

    • carry an explicit is_bedrock marker from the Bedrock client-construction boundary;
    • skip response_format -> output_config.format only for AnthropicBedrock;
    • preserve existing translation for native Anthropic / partner Messages routes.
  • agent/bedrock_adapter.py

    • raise the context-probe output cap from 8 to 16.
  • tests/agent/test_bedrock_issue_124923.py

    • Bedrock route omits output_config.format;
    • native Anthropic still translates structured output;
    • context probe sends maxTokens=16 and still parses the provider-reported maximum.

Competing work / provenance

There is overlapping open work in #124931 for the same issue. I am linking it explicitly rather than presenting this as undiscovered territory.

This implementation was developed independently and does not cherry-pick, copy, or reuse commits from #124931 or any other PR. The main implementation difference is ownership: this branch carries an explicit Bedrock-route marker from _build_bedrock_client() instead of inferring Bedrock from the runtime class name inside the generic Anthropic adapter.

No claim is made that the competing PR is broken; maintainers can choose whichever boundary they prefer.

Exact-head validation

Base: 20f01a0cf6d2dbca6141207abb800520c84fbc86

Branch range is exactly one author-owned commit on that base:

21de4de005655607593773efa4133ca6a3d4ee56 — fix(bedrock): avoid invalid aux request parameters

Changed files are limited to the two production owners and one focused regression module.

Three blocker passes

  1. Issue reality / duplication: validated both failures against current source, traced fix(aux): translate response_format for Anthropic wires and retry when providers reject structured output #89589 / fix(aux): structured-output 400s no longer kill fallback candidates or cost a doomed first request on DeepSeek (#83390, #105191, salvage #92908) #113966 / Anthropic adapter: claude-opus-5-5 rejects thinking.type=disabled, so title generation and /reasoning none fail with HTTP 400 #120069 / Bedrock context probe escalates on the wrong branch: tier 2 uploads 12.2 MB that can never return new information #98520, and re-searched issue-linked and symbol-linked PRs before implementation.
  2. Sibling invariants: rejected the tempting provider-wide unsupported_response_formats fix because Bedrock spans AnthropicBedrock, Converse, and Mantle/OpenAI; the final change scopes only the rejecting AnthropicBedrock route and preserves native Anthropic structured output.
  3. Post-implementation / drift: re-checked the final diff, rebased the single commit onto current upstream main, re-searched competing work, and confirmed the branch is exactly one commit ahead with no unrelated files.

The regression tests are committed; I am not claiming a local pytest run from this connector-only environment. Hosted CI on the PR head is the execution gate.

Infographic

file_00000000ff60820ebaba106c913cb3ed.png

@JoaoMarcos44
JoaoMarcos44 force-pushed the fix/124923-bedrock-aux-probe branch from 390bbde to 21de4de Compare September 27, 2026 07:55
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint provider/anthropic Anthropic native Messages API provider/bedrock AWS Bedrock (boto3, IAM) provider/openai OpenAI / Codex Responses API labels Sep 27, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated follow-up for reference; not a maintainer.

The is_bedrock gate works for its target: driving the real _AnthropicCompletionsAdapter.create() with is_bedrock=True and a json_schema format, output_config.format is gone while native routes keep it. Two P2 gaps remain, both in agent/auxiliary_client.py.

1. output_config.effort still ships to the endpoint you are shielding (k<N). The gate wraps only _translate_anthropic_response_format at L1781/L1785. build_anthropic_kwargs runs earlier at L1750, and for an adaptive-thinking model its _thinking_kwargs writes output_config: {"effort": ...} directly (agent/anthropic_adapter.py:601) — never gated. Captured from the real create() path, head 21de4de:

opus-4-7   is_bedrock=False -> {'effort': 'medium', 'format': {...}}
opus-4-7   is_bedrock=True  -> {'effort': 'medium'}
sonnet-4-6 is_bedrock=True  -> {'effort': 'medium'}
sonnet-4-5 is_bedrock=True  -> None

If Bedrock rejects the output_config object, Claude 4.6+ on Bedrock with reasoning still 400s — the same failure, one key narrower. Gate the object, not the translate call.

2. The new test cannot catch #1. test_anthropic_bedrock_omits_structured_output_format asserts "format" not in (kwargs.get("output_config") or {}), which passes while effort remains, and it never sets reasoning, so the adaptive branch never runs. Its other two assertions already hold on base d6864ba — the passthrough filter k not in {"reasoning", "response_format"} is pre-existing. Please confirm by asserting the key set is empty with reasoning on.

3. Detection is a caller flag, not a capability check (please confirm). is_bedrock=True is passed at one of six AnthropicAuxiliaryClient(...) sites (L4833); L1963, L2914, L3084, L5220 default to False. A named custom provider with api_mode=anthropic_messages aimed at a Bedrock-runtime URL takes the L5220 path and gets is_bedrock=False. If reachable, a class-name or base_url check (as #124931 does) avoids a new call-site obligation.

Relation to #124931: same root cause and files, so this covers only what neither closes.

Unverified / please confirm: 1-2 come from running the adapter on the head blob (anthropic_adapter.py verified byte-identical to base); no live Bedrock call was made, so whether the endpoint rejects output_config whole or only format is inferred from #124931. 3 is read from call sites, not executed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists provider/anthropic Anthropic native Messages API provider/bedrock AWS Bedrock (boto3, IAM) provider/openai OpenAI / Codex Responses API type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bedrock: title_generation sends output_config.format (rejected); context probe maxTokens=8 fails on OpenAI-on-Bedrock models

3 participants