Skip to content

feat(models): align scx-ai-gp Qwen3.8 Max caps - #3439

Merged
smakosh merged 2 commits into
theopenco:mainfrom
SouthernCrossAI:feat/scx-gp-qwen38-max-caps
Aug 6, 2026
Merged

smakosh merged 2 commits into
theopenco:mainfrom
SouthernCrossAI:feat/scx-gp-qwen38-max-caps

Conversation

@bhuvan2134686

@bhuvan2134686 bhuvan2134686 commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

Problem

The scx-ai-gp mapping for qwen3.8-max (added in #3437) declares a narrower capability set than the alibaba mapping for the same upstream model, and asserts a quantization that was never verified.

Two consequences:

  • quantization: "fp8" is a guess. SCX does not publish a quantization for this deployment, and the field is surfaced on the models page.
  • Because the mapping under-declares capabilities, getProviderFilterReasons filters scx-ai-gp out of any request using vision or reasoning_max_tokens — even though the underlying model supports both — so those requests never reach the cheaper SCX deployment.

Approach

Drop quantization entirely rather than guessing a value.

Align the model-intrinsic flags with the alibaba mapping, since both serve the same upstream model:

  • reasoningOutput: "omit"
  • reasoningMaxTokens: true
  • vision: true
  • the supportedParameters allowlist that works around Qwen thinking models rejecting tool_choice: "required" or an object

Deliberately not mirrored

Two fields on the alibaba mapping are platform features, not model capabilities, and would be wrong on SCX's plain OpenAI-compatible endpoint:

  • webSearch / webSearchPrice — Alibaba Cloud's own hosted search (enable_search), not something the model carries with it. Declaring it would route web-search requests to a provider that cannot serve them.
  • cacheReadInputPrice / cacheWriteInputPrice — Alibaba's published off-ratio explicit-cache rates. SCX's cache pricing is already covered by cachedInputPrice.

Notes

The mapping stays test: "skip": the deployment is still absent from SCX's /v1/models listing and chat completions return "Unsupported model". Capability flags are therefore declared from the upstream model's documented behavior, not from live verification — they should be re-checked once SCX brings the deployment online.

Verified with pnpm format and a @llmgateway/models build.

Summary by CodeRabbit

  • New Features

    • Added streaming and web search support for the Qwen 3.8 Max model.
    • Added web search pricing details.
    • Enabled vision inputs, reasoning metadata, and reasoning-effort controls.
  • Bug Fixes

    • Removed unsupported FP8 quantization and deployment-status restrictions.

The scx-ai-gp mapping for qwen3.8-max was added with a narrower
capability set than the Alibaba Cloud mapping for the same model, plus a
quantization claim that was never verified.

Drop `quantization: "fp8"` — SCX does not publish a quantization for this
deployment, so asserting one is a guess that also skews the models page.

Bring the model-intrinsic flags in line with the alibaba mapping, since
both serve the same upstream model: `reasoningOutput: "omit"`,
`reasoningMaxTokens`, `vision`, and the `supportedParameters` list that
works around Qwen thinking models rejecting `tool_choice: "required"`.

Deliberately not mirrored from alibaba: `webSearch`/`webSearchPrice` is
Alibaba Cloud's own hosted search rather than a model capability, and
`cacheReadInputPrice`/`cacheWriteInputPrice` are Alibaba's published
off-ratio explicit-cache rates. Neither applies to SCX's plain
OpenAI-compatible endpoint.

The mapping stays `test: "skip"` — the deployment is still absent from
SCX's /v1/models listing.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 778efe06-e37a-4f2a-9c9a-0763c1261d45

📥 Commits

Reviewing files that changed from the base of the PR and between c83a26a and 2b8d053.

📒 Files selected for processing (1)
  • packages/models/src/models/alibaba.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • packages/models/src/models/alibaba.ts

Walkthrough

The SCX-AI-GP qwen3.8-max configuration now declares cache pricing, reasoning metadata, streaming, vision, web search, and supported generation parameters. It removes FP8 quantization and deployment and vision restrictions.

Changes

Qwen3.8-Max provider configuration

Layer / File(s) Summary
Update provider capabilities and parameters
packages/models/src/models/alibaba.ts
The configuration adds cache read/write pricing, reasoning token metadata, streaming, vision, web search, and supported generation, tool, response-format, and reasoning-effort parameters. It removes FP8 quantization and prior deployment and vision restrictions.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely identifies the main change: aligning SCX-AI-GP Qwen3.8 Max capabilities.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/models/src/models/alibaba.ts`:
- Around line 1698-1703: Update the SCX provider mapping near the capabilities
fields to remove the test skip setting and add deactivatedAt with the review
date (or 2026-08-06), so scx-ai-gp is excluded from production routing.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4b07098d-206b-4b7a-aafb-a301a60209b5

📥 Commits

Reviewing files that changed from the base of the PR and between d5c9489 and c83a26a.

📒 Files selected for processing (1)
  • packages/models/src/models/alibaba.ts

Comment on lines +1698 to 1703
reasoningMaxTokens: true,
reasoningOutput: "omit",
streaming: true,
vision: true,
tools: true,
jsonOutput: true,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

# Expect: determine whether test: "skip" affects production routing or only tests.
rg -n -C 8 'test:\s*"skip"|deactivatedAt|providerId:\s*"scx-ai-gp"|qwen3\.8-max' packages --glob '*.ts'

Repository: theopenco/llmgateway

Length of output: 50377


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '--- Candidate files mentioning test/selection/deactivatedAt ---\n'
rg -n --glob '*.ts' 'type .*Test|deactivatedAt|test:\s*"skip"|skip.*provider|active provider|providers.*deactivated|isDeleted|deletedAt' packages --glob '*.ts' --max-count 200

printf '\n--- Model definition typed definitions ---\n'
fd -e ts 'types\.ts|types.*\.ts|model.*\.ts' packages --max-depth 4 | head -80 | xargs rg -n "export .*ProviderDefinition|deactivatedAt:|interface .*Provider|type .*Provider"

Repository: theopenco/llmgateway

Length of output: 249


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- Candidate files mentioning test/selection/deactivatedAt ---'
rg -n --glob '*.ts' 'type .*Test|deactivatedAt|test:\s*"skip"|skip.*provider|active provider|providers.*deactivated|isDeleted|deletedAt' packages --glob '*.ts' --max-count 200

printf '%s\n' ''
printf '%s\n' '--- Model definition typed definitions ---'
fd -e ts 'types\.ts|types.*\.ts|model.*\.ts' packages --max-depth 4 | head -80 | xargs rg -n "export .*ProviderDefinition|deactivatedAt:|interface .*Provider|type .*Provider"

Repository: theopenco/llmgateway

Length of output: 30600


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- packages/models/src/models.ts ProviderModelMapping lines 170-190 ---'
sed -n '170,190p' packages/models/src/models.ts

printf '%s\n' ''
printf '%s\n' '--- packages/actions/src/get-cheapest-model-for-provider.ts ---'
sed -n '1,80p' packages/actions/src/get-cheapest-model-for-provider.ts

printf '%s\n' ''
printf '%s\n' '--- script around active provider selection ---'
sed -n '170,195p' packages/scripts/src/export-models-dev.ts

printf '%s\n' ''
printf '%s\n' '--- model metadata provider filtering ---'
rg -n --glob '*.ts' 'ProviderModelMapping|deactivatedAt|providers\.find|providers\.filter|isMappingDeactivated|getCheapest' packages/shared packages/actions packages/scripts --max-count 200

Repository: theopenco/llmgateway

Length of output: 21447


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- select-provider-mapping.ts ---'
sed -n '1,100p' packages/actions/src/select-provider-mapping.ts

printf '%s\n' ''
printf '%s\n' '--- routes/tests mentioning test flag ---'
rg -n --glob '*.spec.ts' --glob '*.test.ts' -C 4 'test:\s*"skip"|onlyModel|skipTest|skip.*model|should not|should.*skip' packages --max-count 250

Repository: theopenco/llmgateway

Length of output: 12139


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- exact test flag consumers ---'
rg -n --glob '*.ts' --glob '!packages/**/*.spec.ts' -C 5 'providerMapping|\btest\b|\.find\(\(\s*p\s*\)\s*=>\s*.*providerId' packages --max-count 300

printf '%s\n' ''
printf '%s\n' '--- model availability/export definitions around providers fields ---'
sed -n '70,130p' packages/shared/src/components/models-directory/api-types.ts
sed -n '415,480p' packages/models/src/models.ts

Repository: theopenco/llmgateway

Length of output: 50379


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail
python3 - <<'PY'
import ast, pathlib, re

models_ts = pathlib.Path("packages/models/src/models.ts").read_text()
tree = ast.parse(models_ts)

fields = []
for node in ast.walk(tree):
    if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef, ast.ClassDef)):
        fields.append(node.name)

def parse_decl(text):
    try:
        return ast.parse(text).body[0]
    except Exception:
        return None

for const_name in ["ProviderMapping", "ProviderModelMapping"]:
    decls = []
    for name in fields:
        source = getattr(tree, name, None) if isinstance(tree, ast.AST) else None
    # use direct search around names
PY

printf '%s\n' '--- Locate type alias declarations by line ---'
rg -n 'const ProviderMapping|interface ProviderMapping|type ProviderMapping|interface ProviderModelMapping|type ProviderModelMapping' packages/models src --glob '*.ts'

printf '%s\n' ''
printf '%s\n' '--- Test-flag consumers excluding spec files narrowed ---'
rg -n --glob '*.ts' -C 5 '\bproviderMapping|findProviderMapping|selectProviderMapping|getCheapestModelForProvider|getCheapestFromAvailableProviders' packages/actions packages/shared packages/scripts --max-count 250

Repository: theopenco/llmgateway

Length of output: 462


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- all test flag references in non-spec TypeScript files ---'
rg -n --glob '*.ts' '!packages/**/*.spec.ts' -C 3 'test\b|skip|test:\s*|"skip"' packages/actions packages/shared packages/scripts packages/models/src --max-count 250

Repository: theopenco/llmgateway

Length of output: 292


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- all test flag references in non-spec TypeScript files ---'
rg -n --glob '*.ts' '!**/.spec.ts' -C 3 'test\b|skip|test:\s*|"skip"' packages/actions packages/shared packages/scripts packages/models/src --max-count 250 || true

Repository: theopenco/llmgateway

Length of output: 292


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- all test flag references in non-spec TypeScript files ---'
rg -n --glob 'packages/**/*.ts' --glob '!packages/**/*.spec.ts' -C 3 'test:"skip"|test:\s*"skip"|test: "skip"|skip' packages/actions packages/shared packages/scripts packages/models --max-count 250 || true

printf '%s\n' ''
printf '%s\n' '--- provider mapping selection code ---'
sed -n '1,60p' packages/actions/src/select-provider-mapping.ts
rg -n --glob 'packages/**/*.ts' -C 3 'selectProviderMapping\(|ProviderModelMapping\[\]|\.find\(\(p\)|deactivatedAt' packages/actions packages/shared packages/scripts --max-count 200

Repository: theopenco/llmgateway

Length of output: 50377


Remove test: "skip" on the SCX provider mapping.

test does not exclude a provider from production routing, only from model tests. Since scx-ai-gp returns Unsupported model, add deactivatedAt: new Date("2026-08-06") (or the review date) rather than relying on the test skip flag.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/models/src/models/alibaba.ts` around lines 1698 - 1703, Update the
SCX provider mapping near the capabilities fields to remove the test skip
setting and add deactivatedAt with the review date (or 2026-08-06), so scx-ai-gp
is excluded from production routing.

Source: Coding guidelines

Adds explicit-cache pricing and web search to the scx-ai-gp mapping for
qwen3.8-max, matching the capability surface of the alibaba mapping for
the same upstream model.

Cache: cacheReadInputPrice and cacheWriteInputPrice were previously
unset, so explicit cache_control hits and cache writes both fell back to
cachedInputPrice. Rates mirror the alibaba mapping. SCX's own implicit
rate (0.21e-6) is kept rather than overwritten.

Search: declares webSearch with the standard 0.01 per-search rate, so
routing stops filtering scx-ai-gp out of web-search requests.
@smakosh
smakosh added this pull request to the merge queue Aug 6, 2026
Merged via the queue into theopenco:main with commit acfdb46 Aug 6, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants