Repository navigation
fix(sse): re-run strict system hoist after format translation - #10803
Merged
diegosouzapw merged 1 commit intoAug 21, 2026
Merged
diegosouzapw merged 1 commit into
diegosouzapw merged 1 commit into
Conversation
hoistLeadingSystemMessage() runs on the source message array, before the target translator executes. claudeToOpenAI() then pushes body.system as a fresh leading system message and appends the converted messages after it, so an already-hoisted system message lands back at index 1 and a strict upstream rejects the request with HTTP 400. A Responses-source request is worse off: result.messages does not exist yet at that point, so the call is a no-op and developer->system normalization reintroduces the same problem later. Re-run the same helper on the final outbound array, at the single return of translateRequest(), which is the only shape the upstream receives. The helper is unchanged and idempotent: it returns the same array reference for non-strict providers and for already-compliant requests, so prompt-cache prefixes are unaffected. Reproduced against a vLLM endpoint serving Qwen3.8-27B, whose chat template enforces a single system message at index 0.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…ouzapw#10803) Re-runs hoistLeadingSystemMessage on the final outbound array at translateRequest's single return, instead of only pre-translation. claudeToOpenAI (and the Responses source path, which never ran the pre-translation hoist at all since `messages` doesn't exist yet there) re-introduces/normalizes a leading system message after the hoist already ran, so a strict provider (e.g. vLLM/Qwen3, xiaomi-mimo) could still receive a non-compliant array and 400 with "System message must be at the beginning." Validated live against a vLLM/Qwen3 endpoint (documented in the PR) plus in an isolated worktree boarded onto origin/release/v3.8.50 (0 conflicts, 2 files): - 39/39 focused tests pass (probe-7293-strict-system-hoist including the new Claude-source regression case, memory-system-first-6135, claude-system-role-cache-boundary, memory-cache-safe-injection). - check-file-size, check-changelog-integrity: OK. - typecheck:core: clean. - check-complexity / check-cognitive-complexity: OK, both under baseline. Co-authored-by: Kizuno18 <Kizuno18@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Setting
OMNIROUTE_STRICT_SYSTEM_PROVIDERSwasn't enough to get Claude Code talking to a strict upstream, and the reason turned out to be ordering rather than configuration — the hoist runs before the target translator, andclaudeToOpenAIputs a system message back in front of the one it just hoisted.My setup: a vLLM endpoint serving Qwen3.8-27B, registered as a
vllmconnection, withOMNIROUTE_STRICT_SYSTEM_PROVIDERS=vllm. Qwen3's chat template rejects anything but a single system message at index 0 — same constraintxiaomi-mimoalready has inBUILTIN_PROVIDERS_SYSTEM_MUST_BE_FIRST:Claude Code sends a top-level
systemfield and arole: "system"message insidemessages(its deferred-tools/agents block, ~6.5KB, present on every request — it isn't a hook or a plugin, so there's no way to turn it off client-side). That's the shape that breaks:hoistLeadingSystemMessageatopen-sse/translator/index.ts:321runs on the source array and correctly folds the offender onto index 0claudeToOpenAI(open-sse/translator/request/claude-to-openai.ts:161) pushesbody.systemas a fresh leading system message and appends the converted messages after itA Responses-source request never gets that far —
result.messagesdoesn't exist yet at line 321 (onlyinputdoes), so the call is a plain no-op and Codex CLI hits the same wall throughdeveloper→systemnormalization.The fix re-runs the same helper on the final outbound array, at the single
returnoftranslateRequest, which is the only shape the upstream actually sees. It's the existing helper with no behaviour change: still merge-never-drop, still returns the same array reference for non-strict providers and already-compliant requests, so prompt-cache prefixes are unaffected.Confirmed against the live endpoint — the array
claudeToOpenAIemits today is rejected, and the array the helper produces from it is accepted with both texts preserved:[system, user, system, user](current)System message must be at the beginning.[system, user, user](after the fix)"You are a coding assistant.\ndeferred tools list"Added a case to
tests/unit/probe-7293-strict-system-hoist.test.tscovering the Claude → OpenAI path with both a top-levelsystemand a mid-array one, asserting a single system at index 0 with both texts merged. Stashing theindex.tschange and re-running gives 5 pass / 1 fail with only the new case failing, so it does pin the regression rather than pass vacuously.memory-system-first-6135,claude-system-role-cache-boundaryandmemory-cache-safe-injectionare green too (33 assertions), and ESLint is clean on both files.Worth a look at whether the pre-translation call at line 321 is still needed once this one exists — I left it alone to keep the diff small, but the later call may well subsume it.