Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 19 additions & 11 deletions experimental/sgl-router/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -198,14 +198,20 @@ prefix queries match the blocks the engine caches. Models the engine encodes in
code but dynamo-render cannot tokenize here (Inkling) route via raw prompt
text, as does any model whose template fails to load or render.

Plain text chat requests (string `content`, no tools, no template kwargs or
reasoning controls or historical `reasoning_content`, no assistant continuation,
no consecutive users or non-leading system turns) additionally forward the
rendered tokens to the engine as `input_ids`, retaining the original messages,
so the engine skips re-tokenizing. Every other request shape is rendered for
routing only: the router renders with dynamo-render and does not replicate
SGLang's request normalization, so forwarding is enabled shape by shape as
parity is verified. Use matching model files on the router and workers, and set
Some chats additionally forward the rendered tokens to the engine as
`input_ids`, retaining the original messages, so the engine skips
re-tokenizing. How many depends on the model's renderer. DeepSeek-V4's native
encoder is fixture-verified against SGLang's request normalization, so it
forwards every chat except multimodal ones and those with caller-provided
`input_ids`. Renderers without that verification (HF Jinja templates, Kimi-K3)
forward only plain text chat requests (string `content`, no tools, no template
kwargs or reasoning controls or historical `reasoning_content`, no assistant
continuation, no consecutive users or non-leading system turns) and warn
`UNVERIFIED` at startup; every other request shape is rendered for routing
only. DeepSeek-V4.1 forwards nothing — its renderer is not verified against
current SGLang — while routing tokenization keeps working.

Use matching model files on the router and workers, and set
the same `--default-chat-template-kwargs`, `SGLANG_DEFAULT_THINKING`,
`SGLANG_DSV4_REASONING_EFFORT`, and `SGLANG_DSV41_REASONING_EFFORT` on both.
The router reads these render defaults from its own configuration and environment;
Expand Down Expand Up @@ -246,9 +252,11 @@ official/preview effort profile is detected from the checkpoint's
and are regenerated by `tests/scripts/generate_deepseek_parity.py`.

V4.1 Flash uses Dynamo's separate V4.1 encoder with SGLang's numeric reasoning
budgets, tool payloads, and `<|System|>` markers. Developer messages and media
are left to the worker (the pinned encoder renders them differently), and a
non-default `SGLANG_DSV41_REASONING_EFFORT` needs the forwarding precautions above.
budgets, tool payloads, and `<|System|>` markers — for routing tokenization
only, since V4.1 never forwards `input_ids`. Developer messages and media are
left to the worker (the pinned encoder renders them differently), and a
non-default `SGLANG_DSV41_REASONING_EFFORT` still matters for cache-aware
routing-hash parity.

## Kimi-K3

Expand Down
6 changes: 4 additions & 2 deletions experimental/sgl-router/src/server/metrics.rs
Original file line number Diff line number Diff line change
Expand Up @@ -90,11 +90,13 @@
//!
//! - `forwarded` — router-rendered `input_ids` replaced engine tokenization.
//! - `disabled` — forwarding is off for the model (`--disable-input-ids-forwarding`,
//! or no chat formatter).
//! no chat formatter, or a renderer not verified against SGLang, such as
//! DeepSeek-V4.1).
//! - `ineligible_multimodal` — the chat carries image, video, or audio content
//! parts, which only the engine's multimodal processor can tokenize.
//! - `ineligible` — the forwarding guard excluded some other request shape
//! (tools, non-string content, caller `input_ids`, template controls, ...).
//! (tools, non-string content, caller `input_ids`, template controls, ...;
//! only caller `input_ids` on models that forward all text chats).
//! - `tokenize_failed` — eligible, but ingress rendering failed (the same
//! requests `sgl_router_ingress_tokenize_errors_total` counts).
//!
Expand Down
68 changes: 44 additions & 24 deletions experimental/sgl-router/src/server/routes/chat/preparation.rs
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ use crate::policies::{has_caller_input_ids, request_tokens_for, RequestTokens};
use crate::server::app_context::AppContext;
use crate::server::error::ApiError;
use crate::server::metrics::{InputIdsForwarding, MetricsRegistry};
use crate::tokenizer::ForwardingScope;
use bytes::Bytes;
use serde::de::IgnoredAny;
use serde::Deserialize;
Expand All @@ -28,7 +29,7 @@ pub(super) struct PreparedChatRequest {
pub(super) input_token_count: usize,
caller_set_rid: bool,
fans_out: bool,
can_forward_input_ids: bool,
forwarding_scope: ForwardingScope,
parsed_body: Option<Value>,
sampling_defaults: Vec<(SamplingField, Number)>,
}
Expand All @@ -44,8 +45,12 @@ impl PreparedChatRequest {
// Validate configured sampling rules and collect missing defaults for forwarding.
let sampling_defaults =
resolve_sampling_defaults(&ctx.config.model.sampling_overrides, &fields, &ctx.metrics)?;
let can_forward_input_ids = !ctx.config.model.disable_input_ids_forwarding
&& ctx.tokenizers.has_chat_formatter(&model.0);
let forwarding_scope = if ctx.config.model.disable_input_ids_forwarding {
ForwardingScope::Never
} else {
ctx.tokenizers.forwarding_scope(&model.0)
};
let can_forward_input_ids = forwarding_scope != ForwardingScope::Never;
let needs_tokens = should_tokenize_request(
can_forward_input_ids,
policy_needs_request_tokens,
Expand Down Expand Up @@ -73,7 +78,7 @@ impl PreparedChatRequest {
input_token_count,
caller_set_rid: fields.caller_set_rid,
fans_out: requests_multiple_samples(&fields, &sampling_defaults),
can_forward_input_ids,
forwarding_scope,
parsed_body,
sampling_defaults,
})
Expand All @@ -95,7 +100,7 @@ impl PreparedChatRequest {
) -> Result<Bytes, ApiError> {
// Routing tokens can replace engine tokenization only for supported chat templates.
let forwarding = input_ids_forwarding(
self.can_forward_input_ids,
self.forwarding_scope,
self.parsed_body.as_ref(),
self.tokens.as_ref(),
);
Expand Down Expand Up @@ -567,20 +572,25 @@ fn can_forward_chat_tokens(value: &Value) -> bool {
///
/// Only chats with forwarding enabled that pass the forwarding guard are
/// eligible; an eligible chat without chat-rendered tokens is a failed offload.
/// Multimodal chats are reported apart from other guard exclusions.
/// Multimodal chats are reported apart from other guard exclusions;
/// `AllText` models skip the guard.
fn input_ids_forwarding(
can_forward_input_ids: bool,
scope: ForwardingScope,
request_value: Option<&Value>,
request_tokens: Option<&RequestTokens>,
) -> InputIdsForwarding {
if !can_forward_input_ids {
if scope == ForwardingScope::Never {
return InputIdsForwarding::Disabled;
}
if request_value.is_some_and(request_has_multimodal_content) {
return InputIdsForwarding::IneligibleMultimodal;
}
let eligible = request_value.is_some_and(|v| {
v.get("messages").is_some_and(|m| m.is_array()) && can_forward_chat_tokens(v)
v.get("messages").is_some_and(|m| m.is_array())
&& match scope {
ForwardingScope::AllText => !has_caller_input_ids(v),
_ => can_forward_chat_tokens(v),
}
});
if !eligible {
InputIdsForwarding::Ineligible
Expand Down Expand Up @@ -918,7 +928,7 @@ mod tests {
]});
assert!(!can_forward_chat_tokens(&value));
assert_eq!(
input_ids_forwarding(true, Some(&value), None),
input_ids_forwarding(ForwardingScope::Guarded, Some(&value), None),
InputIdsForwarding::Ineligible
);
value["messages"][1]["reasoning_content"] = Value::Null;
Expand All @@ -944,7 +954,7 @@ mod tests {
let value = json!({"messages": messages});
assert!(!can_forward_chat_tokens(&value), "{roles:?}");
assert_eq!(
input_ids_forwarding(true, Some(&value), None),
input_ids_forwarding(ForwardingScope::Guarded, Some(&value), None),
InputIdsForwarding::Ineligible
);
}
Expand Down Expand Up @@ -1010,26 +1020,36 @@ mod tests {
]}]});
let text_parts =
json!({"messages":[{"role":"user","content":[{"type":"text","text":"hi"}]}]});
for (enabled, value, rendered, expected) in [
(true, Some(&chat), Some(true), Forwarded),
(true, Some(&chat), Some(false), TokenizeFailed),
(true, Some(&chat), None, TokenizeFailed),
(false, Some(&chat), Some(true), Disabled),
(true, Some(&tools), Some(true), Ineligible),
(true, Some(&prompt), None, Ineligible),
(true, None, None, Ineligible),
(true, Some(&image), None, IneligibleMultimodal),
(false, Some(&image), None, Disabled),
(true, Some(&text_parts), None, Ineligible),
let caller_ids = json!({"messages":[{"role":"user","content":"hi"}], "input_ids":[7]});
let (never, guarded, all) = (
ForwardingScope::Never,
ForwardingScope::Guarded,
ForwardingScope::AllText,
);
for (scope, value, rendered, expected) in [
(guarded, Some(&chat), Some(true), Forwarded),
(guarded, Some(&chat), Some(false), TokenizeFailed),
(guarded, Some(&chat), None, TokenizeFailed),
(never, Some(&chat), Some(true), Disabled),
(guarded, Some(&tools), Some(true), Ineligible),
(guarded, Some(&prompt), None, Ineligible),
(guarded, None, None, Ineligible),
(guarded, Some(&image), None, IneligibleMultimodal),
(never, Some(&image), None, Disabled),
(guarded, Some(&text_parts), None, Ineligible),
(all, Some(&tools), Some(true), Forwarded),
(all, Some(&text_parts), Some(true), Forwarded),
(all, Some(&image), None, IneligibleMultimodal),
(all, Some(&caller_ids), Some(false), Ineligible),
] {
let tokens = rendered.map(|rendered_from_chat| RequestTokens {
ids: vec![1, 2, 3],
rendered_from_chat,
});
assert_eq!(
input_ids_forwarding(enabled, value, tokens.as_ref()),
input_ids_forwarding(scope, value, tokens.as_ref()),
expected,
"enabled={enabled}, request={value:?}, rendered={rendered:?}"
"scope={scope:?}, request={value:?}, rendered={rendered:?}"
);
}
}
Expand Down
9 changes: 9 additions & 0 deletions experimental/sgl-router/src/tokenizer/chat_formatter.rs
Original file line number Diff line number Diff line change
Expand Up @@ -231,6 +231,15 @@ impl ChatFormatter {
})
}

/// DeepSeek-V4 is fixture-verified against SGLang for every text chat; V4.1 stays disabled.
pub fn forwarding_scope(&self) -> super::ForwardingScope {
match self.deepseek {
Some(super::deepseek::Encoder::V4(_)) => super::ForwardingScope::AllText,
Some(super::deepseek::Encoder::V41) => super::ForwardingScope::Never,
None => super::ForwardingScope::Guarded,
}
}

/// Apply the workers' `--default-chat-template-kwargs`; they fill keys the
/// request leaves unset, and a default `reasoning_effort` acts as the request's.
pub fn with_defaults(mut self, defaults: &ChatTemplateKwargs) -> Self {
Expand Down
45 changes: 37 additions & 8 deletions experimental/sgl-router/src/tokenizer/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,17 @@ use dynamo_tokenizers::Tokenizer;
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Arc;

/// Which chats may carry router-rendered `input_ids`, by how well the model's
/// renderer is verified against SGLang.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum ForwardingScope {
Never,
/// Only request shapes the per-shape guard allows.
Guarded,
/// Every non-multimodal chat.
AllText,
}

/// A model's chat formatter plus its fallback-logging state.
struct ChatFormatterEntry {
formatter: ChatFormatter,
Expand Down Expand Up @@ -82,14 +93,26 @@ impl TokenizerRegistry {
tracing::info!(model = %m.id,
"router-generated input_ids forwarding disabled; workers tokenize messages; \
routing tokenization remains available");
} else if me.has_chat_formatter(&m.id) {
tracing::warn!(model = %m.id,
"router-generated input_ids forwarding enabled: requires the workers' model files, \
--default-chat-template-kwargs, SGLANG_DEFAULT_THINKING, and \
SGLANG_DSV4_REASONING_EFFORT / SGLANG_DSV41_REASONING_EFFORT; worker parser overrides \
(including --tool-call-parser deepseekv32), content-format detection, and \
conversation-template stop strings are not replicated. Use \
--disable-input-ids-forwarding for array-only templates or when these assumptions do not hold");
} else {
match me.forwarding_scope(&m.id) {
ForwardingScope::AllText => tracing::info!(model = %m.id,
"router-generated input_ids forwarding enabled for all text chats; requires the \
workers' model files, --default-chat-template-kwargs, SGLANG_DEFAULT_THINKING, \
and SGLANG_DSV4_REASONING_EFFORT"),
ForwardingScope::Guarded => tracing::warn!(model = %m.id,
"UNVERIFIED input_ids forwarding: router rendering is verified against SGLang only \
for DeepSeek-V4, so this model forwards only guarded request shapes (plain text \
chat). Requires the workers' model files and --default-chat-template-kwargs; \
worker parser overrides, content-format detection, and conversation-template stop \
strings are not replicated. Pass --disable-input-ids-forwarding unless you have \
verified parity for this model"),
ForwardingScope::Never if me.has_chat_formatter(&m.id) => {
tracing::warn!(model = %m.id,
"input_ids forwarding disabled: the DeepSeek-V4.1 renderer is not verified against \
current SGLang; workers tokenize messages")
}
ForwardingScope::Never => {}
}
}
Ok(me)
}
Expand All @@ -104,6 +127,12 @@ impl TokenizerRegistry {
self.formatters.contains_key(model_id)
}

pub fn forwarding_scope(&self, model_id: &str) -> ForwardingScope {
self.formatters
.get(model_id)
.map_or(ForwardingScope::Never, |e| e.formatter.forwarding_scope())
}

/// Render with dynamo-render and tokenize; return `None` when unavailable or unsuccessful.
pub fn encode_chat(&self, model_id: &str, request: &serde_json::Value) -> Option<Vec<u32>> {
// Clone the Arc and drop the DashMap guard before the CPU-bound
Expand Down
Loading
Loading