Repository navigation
[Model] Add BerryLM - #39972
Draft
DanilSmorchkov wants to merge 5 commits into
Draft
[Model] Add BerryLM#39972DanilSmorchkov wants to merge 5 commits into
DanilSmorchkov wants to merge 5 commits into
Conversation
DanilSmorchkov
requested review from
Fridge003,
HaiShaw,
JustinTong0323,
Qiaolin-Yu,
hebiao064,
ispobock,
merrymercy,
sogalin,
wisclmy0611,
yizhang2077,
yuan-luo and
zijiexia
as code owners
September 17, 2026 13:27
3 of 5 tasks
DanilSmorchkov
marked this pull request as draft
September 18, 2026 12:11
DanilSmorchkov
force-pushed
the
add-berrylm
branch
from
October 6, 2026 07:56
126bc91 to
5eca54f
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Add BerryLM (
BerryLMForCausalLM,model_type: berrylm), a hybrid linear/full-attention MoE decoder with athinking mode, released as
rwb-ai/BerryLM-OS(~3.2B active / ~18.3B total, 40 layers, 180k text-only vocabulary,256k context; transformers: huggingface/transformers#48901, vLLM: vllm-project/vllm#57384).
Architecture: three linear-attention layers per full-attention layer, the linear layers being a gated delta rule with
a per-channel forget gate (rank-128 projection of the layer input → one log-decay per key channel per head, the
Kimi Delta Attention recurrence, GVA 16 key / 32 value heads); GQA full attention with a sigmoid output gate and
partial rotary; Gated Block AttnRes (every layer reads a softmax mixture of the residual streams committed at
block boundaries, gated toward the identity by a per-layer scalar; token-local, transparent to the caches); 128-expert
top-8 MoE plus a shared expert.
Modifications
python/sglang/srt/models/berrylm.py— self-contained model on the layer-boundary stages (make_stages; the AttnResmixer sits between layers as
residual_batch.fold→ mix →set_written): attention block, KDA block with a mergedin_proj_qkvzbprojection and the low-rank gate head, MoE + shared expert, the mixer as a fused Triton kernel (torchfallback off-GPU). Linear attention goes through the existing
HybridLinearAttnBackend/KDAAttnBackend(Kimi convlayout,
KimiLinearStateShape).python/sglang/srt/configs/berrylm.py+ registration inconfigs/__init__.py,hf_transformers/common.py(the registry plumbing this model needs — the matched config in
attn_backend_wrapper, thesupport_mamba_cache_extra_bufferflag — is on main since Let predicate-registered linear-attention models carry the mamba radix-cache leaves #41165).num_k_heads != num_v_heads, q/k repeated to the value heads in the backend rather thanin the model), and a per-layer opt-out of the fused intra-chunk prefill path (
layer.kda_fused_intra = False→chunk_kda(..., fused_intra=False)): the fused-diagonal kernel clamps theexp2(g_i - g_n)factors to ±126, whichcollapses decays beyond 2^-126 inside a 16-token block; BerryLM's per-channel gates reach that range (issue
[KDA] Fused intra-chunk prefill path (
chunk_kda_fwd_intra(fuse_diagonal=True)) collapses for strong per-channel decays because of the ±126 clamp in the exp2 factorization #39971). Default behaviour for every other model is unchanged.python/sglang/srt/function_call/berrylm_detector.py:--reasoning-parser berrylm(<think>blocks) and--tool-call-parser berrylm(XML<tool_call>with typed JSON arguments), streaming and non-streaming; both namesadded to the CLI name lists (
parser/reasoning_parser_names.py,function_call/parser_names.py).test/registered/unit/function_call/test_berrylm_detector.py: tool-call detector (typed arguments, parallel calls, multilinevalues, streaming) and reasoning detector (think/content, template-opened think,
<tool_call>without</think>).test/registered/e2e/models/test_berrylm.py: GSM8K through the server with the reasoning parser, on the kit's chat backend(
gsm8k_backend = "sgl_eval", thinking, 16k tokens) like the other reasoning-model tests (stage="extra-a",runner_config="1-gpu-large"; registereddisableduntil the release checkpoint is public).supported-models/generative_models.mdx,advanced_features/separate_reasoning.mdx,advanced_features/tool_parser.mdx.Accuracy Tests
Serving-path parity (our harness, H200, release checkpoint): a real
launch_serverscores fixed token ids over/generate(input logprobs + greedy decode) and transformers teacher-forces the same ids; default CUDA graphs, radix + mamba prefix
cache on, short prompts and 6k–8k-token documents, cache-hit and multi-turn continuations.
0.08 is the bf16 floor of this model: vLLM's port sits at the same distance from transformers, the two engines differ
from each other by 0.078, and two transformers builds (5.18.0 and main of 17.09) differ by 0.080 on the same ids. The worst
single prompt position is 2.7 at TP1 and 3.4 at TP2; the same positions deviate in vLLM and between the two transformers
builds (up to 2.8). The model card's serving commands were also replayed from a clean environment on H200 and H100; parser
check over the chat API: 17/17.
Measured on the 17.09 head of this branch (before the port to the layer-boundary stages, earlier checkpoint) and not
repeated on the current head:
state, top-1 95.3 %), the unfused path 8.9 % / top-1 97.3 % — the same profile as vLLM's port of the kernels (6.0 % /
98.8 %) and as transformers' fla vs its own torch reference (5.6 % / 97.3 %);
passes identical. On the current head the longest prompt in the parity run is 8k tokens.
test_berrylm.pyrun as is on the release checkpoint (our cluster, model path and port swapped, offline data): GSM8K0.95 over 200 problems (3 of 200 hit the 16k-token cap), so the threshold is 0.90; about 10 minutes including the
server start. It stays registered
disableduntil the checkpoint is public.Speed Tests and Profiling
vllm bench serveclient againstsglang.launch_serveron H200 (random prompts,ignore_eos, default settings). TP1measured on this branch's tree with the release checkpoint; TP2 is from the 17.09 head of the branch (before the port to
the layer-boundary stages), not re-measured:
No change to any other model's path (the fused-intra switch defaults to the upstream behaviour).
Checklist
supported-models/generative_models.mdx,separate_reasoning.mdx,tool_parser.mdx).CI States
Latest PR Test (Base): ❌ Run #37432786991
Latest PR Test (Extra): ❌ Run #37432786905
Latest PR Test (AMD ROCm 10): ❌ Run #37432787080