[Part 2/3] feat(qwen): add Qwen3.8 Max Preview catalogs - #7874
Conversation
There was a problem hiding this comment.
Code Review
This pull request removes the deprecated Qwen OAuth provider and its associated CLI configuration, migrating Qwen Code to a V4 OpenAI-compatible provider model. It updates documentation, test suites, and internal registries to reflect this change, while also updating the Qwen Code CLI setup to use a dedicated environment file for API keys. The review comments correctly identified inconsistencies in documentation and model references that needed to be updated to align with the new provider configuration.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
3c2f169 to
e3dc312
Compare
3668561 to
d2252b4
Compare
Rebuilt clean on release/v3.8.49 after Part 1 (diegosouzapw#7866) squash-merged — applies only the Part-2 delta (Qwen Web / Qoder qwen3.8-max-preview registration + required-thinking allowlist + Qoder client rework) onto the current tip. No migration in this part (that was Part 1). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
d2252b4 to
aacd7e6
Compare
Rebuilt clean on top of Part 2 (diegosouzapw#7874) over the current release tip — applies only the Part-3 delta (alibaba Model Studio, Alibaba Token Plan, qwen-cloud, qwen-cloud-token-plan with region selector). No migration in this part. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
ccdbc89
into
diegosouzapw:release/v3.8.49
…7882) * feat(qwen): add Qwen3.8 Max Preview catalogs [Part 2/3] Rebuilt clean on release/v3.8.49 after Part 1 (#7866) squash-merged — applies only the Part-2 delta (Qwen Web / Qoder qwen3.8-max-preview registration + required-thinking allowlist + Qoder client rework) onto the current tip. No migration in this part (that was Part 1). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * feat(qwen): add regional Alibaba and Qwen Cloud providers [Part 3/3] Rebuilt clean on top of Part 2 (#7874) over the current release tip — applies only the Part-3 delta (alibaba Model Studio, Alibaba Token Plan, qwen-cloud, qwen-cloud-token-plan with region selector). No migration in this part. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Merged — thanks @backryun! Part 2/3 in. We rebuilt the branch clean on top of the merged Part 1 (the squash-merge broke the original commit stack), applying only your Part-2 delta — your code and authorship are intact. Validated in local merge-train (combined tree, full unit + vitest green) on tomni-proxmox-113 @ 250181493b then squash-merged. |
…er has no native function calling (diegosouzapw#8437) qwen-web is a web-cookie provider that emulates tools via synthetic system prompt text and <tool> XML parsing, never sending a native tools[] field upstream. The filterTargetsByRequestCompatibility gate filters out non-tool- calling targets when the request carries tools, but qwen-web's registry entry had toolCalling=true, so the filter let it through and a tool-using session failing over to qwen-web would silently degrade to text-only chat with 'Tool X does not exists' errors. Sibling web-cookie providers (chatgpt-web, yuanbao-web, claude-web, etc.) all correctly set toolCalling: false — qwen-web was an outlier introduced in PR diegosouzapw#7874. Verification: - LSP diagnostics: clean - Pattern matches chatgpt-web, yuanbao-web, and other web-cookie providers
…ling, empty-response exhaustion (#8476) * test(tail): retire stale i18n __MISSING__ repro + fix qianfan website URL Base-red slice 6, rebased onto the advanced release/v3.8.49 (91fd5f9). The oauth grok-cli #7610 guard was already fixed on the base by #8027 (it reads the warning from grokCliAuthJson.ts) — dropped from this slice to avoid a conflicting duplicate. Remaining two, still red on the current base: - i18n #7258: the "focused repro" asserted zh-TW.json STILL carries raw __MISSING__: placeholders. That backlog was filled (the "no locale has a raw __MISSING__: leaf" invariant is the durable guard); retired the now-inverted repro. - qianfan: Baidu renamed the product page (product/wenxinworkshop -> product-s/ qianfan_home); updated the expected website URL. Validated (clean env): i18n 4/0, qianfan 5/0; oauth-modal-grok 2/0 already green on base. * fix(resilience): short-circuit combo on input-bound failures (context_length_exceeded) (#8375) isInputBoundRequestFailure() predicate detects deterministic input-bound errors (context_length_exceeded/context_window_exceeded). The combo loop propagates the original 400 immediately instead of burning MAX_GLOBAL_ATTEMPTS retrying identical oversized inputs against every account. Test: combo-input-bound-failure-8375.test.ts (1 test, 2 assertions) * fix(resilience): add early-exit in combo dispatcher for input-bound failures (#8375) When isInputBoundRequestFailure detects context_length_exceeded, the combo loop returns {ok:false, response} immediately instead of re-dispatching the oversized request. Test: node --import tsx/esm --test tests/unit/combo-input-bound-failure-8375.test.ts - 1 test, 2 assertions, 0 fail * fix(translator): strip input_image from tool outputs in Responses->Chat downgrade (#8459) toolOutputContentToString() extracts input_text/output_text parts and replaces input_image with a placeholder instead of JSON.stringify'ing the content-part array (which embedded raw ~52KB base64 as inert text). Applied to both function_call_output and custom_tool_call_output branches. Existing translator tests: 88/88 pass. New tests: 4/4 pass. * fix(providers): set qwen-web toolCalling to false — web-cookie provider has no native function calling (#8437) qwen-web is a web-cookie provider that emulates tools via synthetic system prompt text and <tool> XML parsing, never sending a native tools[] field upstream. The filterTargetsByRequestCompatibility gate filters out non-tool- calling targets when the request carries tools, but qwen-web's registry entry had toolCalling=true, so the filter let it through and a tool-using session failing over to qwen-web would silently degrade to text-only chat with 'Tool X does not exists' errors. Sibling web-cookie providers (chatgpt-web, yuanbao-web, claude-web, etc.) all correctly set toolCalling: false — qwen-web was an outlier introduced in PR #7874. Verification: - LSP diagnostics: clean - Pattern matches chatgpt-web, yuanbao-web, and other web-cookie providers * fix(backend): empty upstream response mislabeled as exhausted_connection (#8397) isEmptyContentFailure guard only matched '/empty content/i' but the actual error text from detectMalformedNonStream is 'returned an empty response (no usable choices/output)' — which lacks the word 'content'. Expanded regex to also match '/empty response/i' so these transient upstream glitches don't get classified as connection-level exhaustion in combo diagnostics. Test: 28 existing combo-target-exhaustion tests pass (no new test needed) * test(#8397): add regression test for empty-response 502 not marking provider/connection exhausted * test(qwen-web): align registry snapshot with toolCalling:false Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(resilience): scope #8375 input-bound short-circuit to homogeneous remainders The isInputBoundFailure short-circuit (context_length_exceeded / context_window_exceeded) fired unconditionally on the first target, aborting the whole combo even when later targets are a different model with a larger context window — regressing the intentional heterogeneous-combo fallback that isContextOverflow400 (#6637) protects. Reproduced with a 2-target combo (small-context model fails, larger-context model would have succeeded): the combo never reached target 2. Scope the short-circuit to remainders where every remaining target shares the same modelStr as the one that just failed — the "retrying will fail identically" premise for context_length_exceeded only holds within a homogeneous same-model pool. Rebaselines open-sse/services/combo.ts's frozen file-size cap (3642->3679) for this PR's own combo.ts growth (config/quality/file-size-baseline.json). Refs #8375 Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> --------- Co-authored-by: Probe Test <probe@example.com> Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: ikelvingo <im.kelvinwong@gmail.com> Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
…#7874) Rebuilt clean on release/v3.8.49 after Part 1 (diegosouzapw#7866) squash-merged — applies only the Part-2 delta (Qwen Web / Qoder qwen3.8-max-preview registration + required-thinking allowlist + Qoder client rework) onto the current tip. No migration in this part (that was Part 1). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…iegosouzapw#7882) * feat(qwen): add Qwen3.8 Max Preview catalogs [Part 2/3] Rebuilt clean on release/v3.8.49 after Part 1 (diegosouzapw#7866) squash-merged — applies only the Part-2 delta (Qwen Web / Qoder qwen3.8-max-preview registration + required-thinking allowlist + Qoder client rework) onto the current tip. No migration in this part (that was Part 1). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * feat(qwen): add regional Alibaba and Qwen Cloud providers [Part 3/3] Rebuilt clean on top of Part 2 (diegosouzapw#7874) over the current release tip — applies only the Part-3 delta (alibaba Model Studio, Alibaba Token Plan, qwen-cloud, qwen-cloud-token-plan with region selector). No migration in this part. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…ling, empty-response exhaustion (diegosouzapw#8476) * test(tail): retire stale i18n __MISSING__ repro + fix qianfan website URL Base-red slice 6, rebased onto the advanced release/v3.8.49 (c58f339). The oauth grok-cli diegosouzapw#7610 guard was already fixed on the base by diegosouzapw#8027 (it reads the warning from grokCliAuthJson.ts) — dropped from this slice to avoid a conflicting duplicate. Remaining two, still red on the current base: - i18n diegosouzapw#7258: the "focused repro" asserted zh-TW.json STILL carries raw __MISSING__: placeholders. That backlog was filled (the "no locale has a raw __MISSING__: leaf" invariant is the durable guard); retired the now-inverted repro. - qianfan: Baidu renamed the product page (product/wenxinworkshop -> product-s/ qianfan_home); updated the expected website URL. Validated (clean env): i18n 4/0, qianfan 5/0; oauth-modal-grok 2/0 already green on base. * fix(resilience): short-circuit combo on input-bound failures (context_length_exceeded) (diegosouzapw#8375) isInputBoundRequestFailure() predicate detects deterministic input-bound errors (context_length_exceeded/context_window_exceeded). The combo loop propagates the original 400 immediately instead of burning MAX_GLOBAL_ATTEMPTS retrying identical oversized inputs against every account. Test: combo-input-bound-failure-8375.test.ts (1 test, 2 assertions) * fix(resilience): add early-exit in combo dispatcher for input-bound failures (diegosouzapw#8375) When isInputBoundRequestFailure detects context_length_exceeded, the combo loop returns {ok:false, response} immediately instead of re-dispatching the oversized request. Test: node --import tsx/esm --test tests/unit/combo-input-bound-failure-8375.test.ts - 1 test, 2 assertions, 0 fail * fix(translator): strip input_image from tool outputs in Responses->Chat downgrade (diegosouzapw#8459) toolOutputContentToString() extracts input_text/output_text parts and replaces input_image with a placeholder instead of JSON.stringify'ing the content-part array (which embedded raw ~52KB base64 as inert text). Applied to both function_call_output and custom_tool_call_output branches. Existing translator tests: 88/88 pass. New tests: 4/4 pass. * fix(providers): set qwen-web toolCalling to false — web-cookie provider has no native function calling (diegosouzapw#8437) qwen-web is a web-cookie provider that emulates tools via synthetic system prompt text and <tool> XML parsing, never sending a native tools[] field upstream. The filterTargetsByRequestCompatibility gate filters out non-tool- calling targets when the request carries tools, but qwen-web's registry entry had toolCalling=true, so the filter let it through and a tool-using session failing over to qwen-web would silently degrade to text-only chat with 'Tool X does not exists' errors. Sibling web-cookie providers (chatgpt-web, yuanbao-web, claude-web, etc.) all correctly set toolCalling: false — qwen-web was an outlier introduced in PR diegosouzapw#7874. Verification: - LSP diagnostics: clean - Pattern matches chatgpt-web, yuanbao-web, and other web-cookie providers * fix(backend): empty upstream response mislabeled as exhausted_connection (diegosouzapw#8397) isEmptyContentFailure guard only matched '/empty content/i' but the actual error text from detectMalformedNonStream is 'returned an empty response (no usable choices/output)' — which lacks the word 'content'. Expanded regex to also match '/empty response/i' so these transient upstream glitches don't get classified as connection-level exhaustion in combo diagnostics. Test: 28 existing combo-target-exhaustion tests pass (no new test needed) * test(diegosouzapw#8397): add regression test for empty-response 502 not marking provider/connection exhausted * test(qwen-web): align registry snapshot with toolCalling:false Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(resilience): scope diegosouzapw#8375 input-bound short-circuit to homogeneous remainders The isInputBoundFailure short-circuit (context_length_exceeded / context_window_exceeded) fired unconditionally on the first target, aborting the whole combo even when later targets are a different model with a larger context window — regressing the intentional heterogeneous-combo fallback that isContextOverflow400 (diegosouzapw#6637) protects. Reproduced with a 2-target combo (small-context model fails, larger-context model would have succeeded): the combo never reached target 2. Scope the short-circuit to remainders where every remaining target shares the same modelStr as the one that just failed — the "retrying will fail identically" premise for context_length_exceeded only holds within a homogeneous same-model pool. Rebaselines open-sse/services/combo.ts's frozen file-size cap (3642->3679) for this PR's own combo.ts growth (config/quality/file-size-baseline.json). Refs diegosouzapw#8375 Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> --------- Co-authored-by: Probe Test <probe@example.com> Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: ikelvingo <im.kelvinwong@gmail.com> Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
…#7874) Rebuilt clean on release/v3.8.49 after Part 1 (diegosouzapw#7866) squash-merged — applies only the Part-2 delta (Qwen Web / Qoder qwen3.8-max-preview registration + required-thinking allowlist + Qoder client rework) onto the current tip. No migration in this part (that was Part 1). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…iegosouzapw#7882) * feat(qwen): add Qwen3.8 Max Preview catalogs [Part 2/3] Rebuilt clean on release/v3.8.49 after Part 1 (diegosouzapw#7866) squash-merged — applies only the Part-2 delta (Qwen Web / Qoder qwen3.8-max-preview registration + required-thinking allowlist + Qoder client rework) onto the current tip. No migration in this part (that was Part 1). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * feat(qwen): add regional Alibaba and Qwen Cloud providers [Part 3/3] Rebuilt clean on top of Part 2 (diegosouzapw#7874) over the current release tip — applies only the Part-3 delta (alibaba Model Studio, Alibaba Token Plan, qwen-cloud, qwen-cloud-token-plan with region selector). No migration in this part. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…ling, empty-response exhaustion (diegosouzapw#8476) * test(tail): retire stale i18n __MISSING__ repro + fix qianfan website URL Base-red slice 6, rebased onto the advanced release/v3.8.49 (f662f70). The oauth grok-cli diegosouzapw#7610 guard was already fixed on the base by diegosouzapw#8027 (it reads the warning from grokCliAuthJson.ts) — dropped from this slice to avoid a conflicting duplicate. Remaining two, still red on the current base: - i18n diegosouzapw#7258: the "focused repro" asserted zh-TW.json STILL carries raw __MISSING__: placeholders. That backlog was filled (the "no locale has a raw __MISSING__: leaf" invariant is the durable guard); retired the now-inverted repro. - qianfan: Baidu renamed the product page (product/wenxinworkshop -> product-s/ qianfan_home); updated the expected website URL. Validated (clean env): i18n 4/0, qianfan 5/0; oauth-modal-grok 2/0 already green on base. * fix(resilience): short-circuit combo on input-bound failures (context_length_exceeded) (diegosouzapw#8375) isInputBoundRequestFailure() predicate detects deterministic input-bound errors (context_length_exceeded/context_window_exceeded). The combo loop propagates the original 400 immediately instead of burning MAX_GLOBAL_ATTEMPTS retrying identical oversized inputs against every account. Test: combo-input-bound-failure-8375.test.ts (1 test, 2 assertions) * fix(resilience): add early-exit in combo dispatcher for input-bound failures (diegosouzapw#8375) When isInputBoundRequestFailure detects context_length_exceeded, the combo loop returns {ok:false, response} immediately instead of re-dispatching the oversized request. Test: node --import tsx/esm --test tests/unit/combo-input-bound-failure-8375.test.ts - 1 test, 2 assertions, 0 fail * fix(translator): strip input_image from tool outputs in Responses->Chat downgrade (diegosouzapw#8459) toolOutputContentToString() extracts input_text/output_text parts and replaces input_image with a placeholder instead of JSON.stringify'ing the content-part array (which embedded raw ~52KB base64 as inert text). Applied to both function_call_output and custom_tool_call_output branches. Existing translator tests: 88/88 pass. New tests: 4/4 pass. * fix(providers): set qwen-web toolCalling to false — web-cookie provider has no native function calling (diegosouzapw#8437) qwen-web is a web-cookie provider that emulates tools via synthetic system prompt text and <tool> XML parsing, never sending a native tools[] field upstream. The filterTargetsByRequestCompatibility gate filters out non-tool- calling targets when the request carries tools, but qwen-web's registry entry had toolCalling=true, so the filter let it through and a tool-using session failing over to qwen-web would silently degrade to text-only chat with 'Tool X does not exists' errors. Sibling web-cookie providers (chatgpt-web, yuanbao-web, claude-web, etc.) all correctly set toolCalling: false — qwen-web was an outlier introduced in PR diegosouzapw#7874. Verification: - LSP diagnostics: clean - Pattern matches chatgpt-web, yuanbao-web, and other web-cookie providers * fix(backend): empty upstream response mislabeled as exhausted_connection (diegosouzapw#8397) isEmptyContentFailure guard only matched '/empty content/i' but the actual error text from detectMalformedNonStream is 'returned an empty response (no usable choices/output)' — which lacks the word 'content'. Expanded regex to also match '/empty response/i' so these transient upstream glitches don't get classified as connection-level exhaustion in combo diagnostics. Test: 28 existing combo-target-exhaustion tests pass (no new test needed) * test(diegosouzapw#8397): add regression test for empty-response 502 not marking provider/connection exhausted * test(qwen-web): align registry snapshot with toolCalling:false Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(resilience): scope diegosouzapw#8375 input-bound short-circuit to homogeneous remainders The isInputBoundFailure short-circuit (context_length_exceeded / context_window_exceeded) fired unconditionally on the first target, aborting the whole combo even when later targets are a different model with a larger context window — regressing the intentional heterogeneous-combo fallback that isContextOverflow400 (diegosouzapw#6637) protects. Reproduced with a 2-target combo (small-context model fails, larger-context model would have succeeded): the combo never reached target 2. Scope the short-circuit to remainders where every remaining target shares the same modelStr as the one that just failed — the "retrying will fail identically" premise for context_length_exceeded only holds within a homogeneous same-model pool. Rebaselines open-sse/services/combo.ts's frozen file-size cap (3642->3679) for this PR's own combo.ts growth (config/quality/file-size-baseline.json). Refs diegosouzapw#8375 Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> --------- Co-authored-by: Probe Test <probe@example.com> Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: ikelvingo <im.kelvinwong@gmail.com> Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
Stack
This PR depends on #7866. Until Part 1 merges, GitHub will also show the shared Part 1 commits in this diff; those commits will disappear from the review diff after #7866 lands.
What changed
qwen3.8-max-previewin Qwen Web, Qoder, the free-model catalog, and centralized model specsreasoning_contentwhile keeping answer content separateWhy
Qwen3.8 Max Preview is available on the live Qwen/Qoder surfaces, but OmniRoute did not expose it consistently. Qwen Web also requires thinking mode for this model, which caused a valid upstream response to be interpreted incorrectly by the dashboard model test path.
Validation
node --import tsx/esm --test tests/unit/executor-qwen-web.test.ts tests/unit/qoder-cli.test.ts tests/unit/qoder-executor.test.ts tests/unit/t31-t33-t34-t38-model-specs.test.ts— 62 passednpx prettier --checkon all changed filesnpm run lintCloses the Qwen3.8 catalog portion of the Qwen provider stack; #7854 remains a separate Part 3 provider-family change.