Skip to content

fix(azure): match the generation, not one release, for max_completion_tokens - #13007

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.51from
ntdatt812:fix/12981-azure-gpt6
Sep 10, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.51from
ntdatt812:fix/12981-azure-gpt6

Conversation

@ntdatt812

Copy link
Copy Markdown
Contributor

Closes #12981.

azure/gpt-6-astra fails validation and every request:

'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.

Why

AZURE_COMPLETION_TOKEN_DEPLOYMENT matched a literal gpt-5:

/(?:^|[/_-])(?:gpt-5|o(?:1|3|4))(?:[._-]|$)|^gpt-chat-latest$/i

So the max_tokens → max_completion_tokens swap covered GPT-5 and the o-series and nothing after them. The requirement is a property of the generation, not of one release — gpt-6-astra rejects max_tokens for exactly the reason gpt-5 does — so each new family arrives broken and needs another patch. Matching the range fixes GPT-6 and whatever follows in the same edit.

Why a range and not gpt-\d+

Azure's own deployment name for GPT-3.5 is gpt-35-turbo — no dot. A digit-run matches it, and the swap would then strip the max_tokens that deployment actually requires. Breaking a working deployment is worse than the bug being fixed, so the range stops at 19: gpt-(?:[5-9]|1\d). That still leaves room for a future gpt-10.

There is a test for exactly that, and it is the test that fails if someone later widens the pattern.

Verification

tests/unit/azure-param-rules.test.ts — 10 passed. Each claim checked by reverting it:

reverted fails
back to the gpt-5 literal only generations after GPT-5 convert max_tokens too (#12981)
widened to gpt-\d+ only gpt-35-turbo is not a GPT-3.5 deployment caught by the generation range

The pattern lives in one place and has one consumer (applyAzureParamRules, shared by azure-openai and azure-ai), so both wire paths pick this up.

…_tokens

`azure/gpt-6-astra` fails validation and every request with:

  'max_tokens' is not supported with this model. Use
  'max_completion_tokens' instead.

`AZURE_COMPLETION_TOKEN_DEPLOYMENT` matched a literal `gpt-5`, so the
swap applied to GPT-5 and the o-series and to nothing after them. The
requirement belongs to the generation rather than to one release, so
every new family arrives broken and needs another patch. Matching the
range fixes GPT-6 and whatever follows it in the same edit.

The range stops at 19 rather than being a digit-run. Azure's own name for
GPT-3.5 is `gpt-35-turbo`, which requires `max_tokens` and would be
caught by `gpt-\d+` -- the swap would then break a working deployment,
which is worse than the bug. A test pins that, and it is the test that
fails if the range is widened.

Verified: 10 passed in tests/unit/azure-param-rules.test.ts, and each
claim checked by reverting it -- back to the `gpt-5` literal fails only
the new-generation test, widening to `gpt-\d+` fails only the
gpt-35-turbo test.

Closes diegosouzapw#12981
@diegosouzapw
diegosouzapw merged commit 4edc3d5 into diegosouzapw:release/v3.8.51 Sep 10, 2026
8 of 16 checks passed
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…_tokens (diegosouzapw#13007)

Boarded with 13 sibling PRs into one worktree off release/v3.8.51 and validated as a set: 132 focused tests pass across all 15 test files in the batch, typecheck:core is clean, check-changelog-integrity reports no lost base bullets, and check-file-size is green. Your PR merged without conflict against its siblings.

Thank you — the write-up made this reviewable: measuring the behaviour on the release tip and showing the before/after table meant the defect could be confirmed rather than taken on faith.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] Cannot use azure/gpt-6-astra model due to max_tokens parameter

2 participants