Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 19 additions & 11 deletions experimental/sgl-router/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -198,14 +198,20 @@ prefix queries match the blocks the engine caches. Models the engine encodes in
code but dynamo-render cannot tokenize here (Inkling) route via raw prompt
text, as does any model whose template fails to load or render.

Plain text chat requests (string `content`, no tools, no template kwargs or
reasoning controls or historical `reasoning_content`, no assistant continuation,
no consecutive users or non-leading system turns) additionally forward the
rendered tokens to the engine as `input_ids`, retaining the original messages,
so the engine skips re-tokenizing. Every other request shape is rendered for
routing only: the router renders with dynamo-render and does not replicate
SGLang's request normalization, so forwarding is enabled shape by shape as
parity is verified. Use matching model files on the router and workers, and set
Some chats additionally forward the rendered tokens to the engine as
`input_ids`, retaining the original messages, so the engine skips
re-tokenizing. How many depends on the model's renderer. DeepSeek-V4's native
encoder is fixture-verified against SGLang's request normalization, so it
forwards every chat except multimodal ones and those with caller-provided
`input_ids`. Renderers without that verification (HF Jinja templates, Kimi-K3)
forward only plain text chat requests (string `content`, no tools, no template
kwargs or reasoning controls or historical `reasoning_content`, no assistant
continuation, no consecutive users or non-leading system turns) and warn
`UNVERIFIED` at startup; every other request shape is rendered for routing
only. DeepSeek-V4.1 forwards nothing — its renderer is not verified against
current SGLang — while routing tokenization keeps working.

Use matching model files on the router and workers, and set
the same `--default-chat-template-kwargs`, `SGLANG_DEFAULT_THINKING`,
`SGLANG_DSV4_REASONING_EFFORT`, and `SGLANG_DSV41_REASONING_EFFORT` on both.
The router reads these render defaults from its own configuration and environment;
Expand Down Expand Up @@ -246,9 +252,11 @@ official/preview effort profile is detected from the checkpoint's
and are regenerated by `tests/scripts/generate_deepseek_parity.py`.

V4.1 Flash uses Dynamo's separate V4.1 encoder with SGLang's numeric reasoning
budgets, tool payloads, and `<|System|>` markers. Developer messages and media
are left to the worker (the pinned encoder renders them differently), and a
non-default `SGLANG_DSV41_REASONING_EFFORT` needs the forwarding precautions above.
budgets, tool payloads, and `<|System|>` markers — for routing tokenization
only, since V4.1 never forwards `input_ids`. Developer messages and media are
left to the worker (the pinned encoder renders them differently), and a
non-default `SGLANG_DSV41_REASONING_EFFORT` still matters for cache-aware
routing-hash parity.

## Kimi-K3

Expand Down
Loading