Skip to content

fix(sse): re-run strict system hoist after format translation - #10803

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
Kizuno18:fix/strict-system-hoist-after-translation
Aug 21, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
Kizuno18:fix/strict-system-hoist-after-translation

Conversation

@Kizuno18

Copy link
Copy Markdown
Contributor

Setting OMNIROUTE_STRICT_SYSTEM_PROVIDERS wasn't enough to get Claude Code talking to a strict upstream, and the reason turned out to be ordering rather than configuration — the hoist runs before the target translator, and claudeToOpenAI puts a system message back in front of the one it just hoisted.

My setup: a vLLM endpoint serving Qwen3.8-27B, registered as a vllm connection, with OMNIROUTE_STRICT_SYSTEM_PROVIDERS=vllm. Qwen3's chat template rejects anything but a single system message at index 0 — same constraint xiaomi-mimo already has in BUILTIN_PROVIDERS_SYSTEM_MUST_BE_FIRST:

{"error":{"message":"System message must be at the beginning.","type":"BadRequestError","code":400}}

Claude Code sends a top-level system field and a role: "system" message inside messages (its deferred-tools/agents block, ~6.5KB, present on every request — it isn't a hook or a plugin, so there's no way to turn it off client-side). That's the shape that breaks:

  • hoistLeadingSystemMessage at open-sse/translator/index.ts:321 runs on the source array and correctly folds the offender onto index 0
  • then claudeToOpenAI (open-sse/translator/request/claude-to-openai.ts:161) pushes body.system as a fresh leading system message and appends the converted messages after it
  • the hoisted system is back at index 1, and the upstream 400s

A Responses-source request never gets that far — result.messages doesn't exist yet at line 321 (only input does), so the call is a plain no-op and Codex CLI hits the same wall through developer → system normalization.

The fix re-runs the same helper on the final outbound array, at the single return of translateRequest, which is the only shape the upstream actually sees. It's the existing helper with no behaviour change: still merge-never-drop, still returns the same array reference for non-strict providers and already-compliant requests, so prompt-cache prefixes are unaffected.

Confirmed against the live endpoint — the array claudeToOpenAI emits today is rejected, and the array the helper produces from it is accepted with both texts preserved:

messages sent upstream result
[system, user, system, user] (current) 400 System message must be at the beginning.
[system, user, user] (after the fix) 200, system content "You are a coding assistant.\ndeferred tools list"

Added a case to tests/unit/probe-7293-strict-system-hoist.test.ts covering the Claude → OpenAI path with both a top-level system and a mid-array one, asserting a single system at index 0 with both texts merged. Stashing the index.ts change and re-running gives 5 pass / 1 fail with only the new case failing, so it does pin the regression rather than pass vacuously. memory-system-first-6135, claude-system-role-cache-boundary and memory-cache-safe-injection are green too (33 assertions), and ESLint is clean on both files.

Worth a look at whether the pre-translation call at line 321 is still needed once this one exists — I left it alone to keep the diff small, but the later call may well subsume it.

hoistLeadingSystemMessage() runs on the source message array, before the
target translator executes. claudeToOpenAI() then pushes body.system as a
fresh leading system message and appends the converted messages after it,
so an already-hoisted system message lands back at index 1 and a strict
upstream rejects the request with HTTP 400. A Responses-source request is
worse off: result.messages does not exist yet at that point, so the call
is a no-op and developer->system normalization reintroduces the same
problem later.

Re-run the same helper on the final outbound array, at the single return
of translateRequest(), which is the only shape the upstream receives. The
helper is unchanged and idempotent: it returns the same array reference
for non-strict providers and for already-compliant requests, so
prompt-cache prefixes are unaffected.

Reproduced against a vLLM endpoint serving Qwen3.8-27B, whose chat
template enforces a single system message at index 0.
@Kizuno18
Kizuno18 requested a review from diegosouzapw as a code owner August 20, 2026 03:51
@diegosouzapw
diegosouzapw merged commit 1d4c4cd into diegosouzapw:release/v3.8.50 Aug 21, 2026
3 checks passed
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ouzapw#10803)

Re-runs hoistLeadingSystemMessage on the final outbound array at translateRequest's single return, instead of only pre-translation. claudeToOpenAI (and the Responses source path, which never ran the pre-translation hoist at all since `messages` doesn't exist yet there) re-introduces/normalizes a leading system message after the hoist already ran, so a strict provider (e.g. vLLM/Qwen3, xiaomi-mimo) could still receive a non-compliant array and 400 with "System message must be at the beginning."

Validated live against a vLLM/Qwen3 endpoint (documented in the PR) plus in an isolated worktree boarded onto origin/release/v3.8.50 (0 conflicts, 2 files):
- 39/39 focused tests pass (probe-7293-strict-system-hoist including the new Claude-source regression case, memory-system-first-6135, claude-system-role-cache-boundary, memory-cache-safe-injection).
- check-file-size, check-changelog-integrity: OK.
- typecheck:core: clean.
- check-complexity / check-cognitive-complexity: OK, both under baseline.

Co-authored-by: Kizuno18 <Kizuno18@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants