Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1590 commits
Select commit Hold shift + click to select a range
b170799
fix(gemma4): quantize MTP bridge projections (#32440)
ayush1399 Aug 19, 2026
1c82955
[HiCache] Split the host-memory budget across co-located ranks (#35540)
cctry Aug 19, 2026
746418a
[DSA] Trim top-k v2 output modes and tighten its PDL waits (#35041)
DarkSharpness Aug 19, 2026
defb2a3
feat(openai): Accept the input_audio content part in chat completions…
jason136 Aug 19, 2026
ed12d68
fix(disagg): allow fake transfer with decode DCP (#35409)
milesial Aug 19, 2026
a6bc053
[Fix] Fix Nemotron-H Mamba illegal memory access under DP attention w…
elvischenv Aug 19, 2026
1270204
Revert "[Feature] Add DeepEPv2 (ElasticBuffer) MoE A2A backend" (#35568)
hnyls2002 Aug 19, 2026
38b74d2
Add docs for TP LMHead optimizaiton (#35283)
SYChen123 Aug 19, 2026
01814e1
[HiCache] Simple style change for buffer mode (#35574)
xiezhq-hermann Aug 19, 2026
082aac8
[Bugfix] Fix min-new-token EOS handling (#31378)
milesial Aug 19, 2026
c7e2c08
fix(constrained): reject NUL bytes in grammar specs to stop an xgramm…
ECMGit Aug 19, 2026
5375bab
[Quant] Load compressed-tensors quantized lm_head instead of value-ca…
Jiminator Aug 19, 2026
1df78c2
chore: bump tilelang to 0.1.12 (#30874)
elvischenv Aug 20, 2026
d216737
[Kernel] Support wider rows in mega_moe_pre_dispatch (#35372)
842974287 Aug 20, 2026
9234e40
[sampling] Fix int32 offset overflow in top-k renorm Triton kernels (…
merrymercy Aug 20, 2026
99c1221
Support custom draft worker classes in DSpark (#35397)
merrymercy Aug 20, 2026
f736895
Make PR babysitter launcher fork-safe (#35575)
merrymercy Aug 20, 2026
e805a8f
[AMD] Keep the PTX-inline-asm diffusion norm fusions off on ROCm (fix…
michaelzhang-ai Aug 20, 2026
1f87d8f
[diffusion] fix: stop reserving nccl device buffers for single-rank g…
mickqian Aug 20, 2026
1cf2b8c
[Spec] Support quantized target lm_head in the DFlash2 selector (#35496)
Jiminator Aug 20, 2026
ab20366
[diffusion] fix: reject unsupported modelopt checkpoint algorithms (#…
mickqian Aug 20, 2026
c747822
[AMD] [Docker] Upgrade Python 3.12 + torch 2.11 + triton 3.7 in ROCm …
chuyeh Aug 20, 2026
dc175b3
[CI][AMD] Run the profiling suite without CUDA graphs on ROCm (#34452)
michaelzhang-ai Aug 20, 2026
f659618
docker: fix CUDA-13 build — rename NCCL_VERSION ARG to avoid base ima…
datdo-msft Aug 20, 2026
a49560c
[misc] Add a comment style rule to .claude/rules (#35597)
hnyls2002 Aug 20, 2026
9db4ba8
[DeepSeek-V4] Add Q8KV8 sparse MLA prefill runtime backend (#32327)
shiyang814-cpu Aug 20, 2026
238ba40
[XPU][CI] key persistent JIT kernel cache by image content ID (#35337)
arathi-hlab Aug 20, 2026
360d10d
[Feature] Add process-local in-memory KV indexer and Router integrati…
wuyl1 Aug 20, 2026
3b22f4f
[diffusion] UX: reduce per-request log noise (#35614)
mickqian Aug 20, 2026
a3e4592
[diffusion] CI: use canonical residency selector in nightly (#35615)
mickqian Aug 20, 2026
db2eb47
[diffusion] fix: keep cosmos3 T=1 fusion on blackwell only (#35612)
mickqian Aug 20, 2026
a5a9d66
[XPU] Fix/kimi linear xpu (#34546)
SKRohit Aug 20, 2026
f744607
[diffusion] UX: report where a component's weights are (#35618)
mickqian Aug 20, 2026
b6dcd39
[Fix] Support 128-aligned hidden sizes in the W4AFP8 DeepEP low-laten…
alexnails Aug 20, 2026
09b7af1
[CI] Surface AMD ROCm 7.2 state in the PR CI-states block (#34813)
michaelzhang-ai Aug 20, 2026
ba433bb
[Docs] Update contribution guide (#35419)
mmangkad Aug 20, 2026
58e3274
update codeowner (#34802)
Aug 20, 2026
02b93e7
[AMD] Add GLM-5.2 MI35x nightly accuracy and perf benchmark (#32570)
michaelzhang-ai Aug 20, 2026
50dae2d
Amd/dsv4 shared experts fusion top6 (#32340)
karverma-amd Aug 20, 2026
32d98aa
[HiCache] Allow a retraction host pool smaller than the device pool (…
cctry Aug 20, 2026
628674a
Remove unused MOONCAKE_COMPILE_ARG argument from Dockerfile (#35649)
ShangmingCai Aug 20, 2026
0bda0b1
[Fix]: exclude SM120 from attn-res TMA dispatch (#35361)
beyondHJM Aug 20, 2026
d287880
Update deepep for SBO feature (#35450)
Fridge003 Aug 20, 2026
7ba3430
Fix Grok-2 nightly: derive image-understanding capability from is_mul…
yctseng0211 Aug 20, 2026
b8996a5
[diffusion] fix: keep large vocab tables in host memory under layerwi…
mickqian Aug 20, 2026
ae23423
Split TRTLLM MHA decode batches by KV sequence length (#34888)
YAMY1234 Aug 20, 2026
21c88f8
[diffusion] quant: support gguf (#35370)
zijiexia Aug 20, 2026
f386e2a
[AMD][CI] Default the ROCm 7.2 PR gate to ROCm 7.2.4 Image (#35602)
bingxche Aug 20, 2026
06ad7b2
[AMD][CI] Run Both ROCm 7.2.4 and ROCm 7.2.0 Images on Nightly Test A…
bingxche Aug 20, 2026
f1b9a1f
[diffusion] feat: support unverified short edge instead of rejecting …
mickqian Aug 20, 2026
17313cf
[diffusion] CI: add minimax-h3 ref2va audio consistency coverage and …
mickqian Aug 20, 2026
cf3813f
[diffusion] feat: add weight source reader (#35668)
mickqian Aug 20, 2026
710267d
[Quant] Load compressed-tensors kv_cache_scheme scales (#35455)
Jiminator Aug 20, 2026
97efc05
[diffusion] feat: plan pinned host memory against the cgroup cap not …
mickqian Aug 20, 2026
82c6fc2
[diffusion] quant: support pruned safetensors checkpoints for minimax…
mickqian Aug 20, 2026
a4ffb99
[Fix] Keep deterministic GDN prefill on Triton (#35632)
mmangkad Aug 20, 2026
c98f1cc
[NPU]Ensure tensors allocated by empty_like are contiguous (#34935)
Estrella-xx Aug 20, 2026
b03ac35
[NPU] [FIX] Fix non-contiguous parameter issue in FIA operator (#34936)
silencejade Aug 20, 2026
9b249a2
test: switch the Inkling-Small NVFP4 deterministic suite to DSPARK (#…
alphabetc1 Aug 20, 2026
04444ee
[diffusion] Refresh eager optimization skills and benchmark safeguard…
BBuf Aug 20, 2026
7f8f030
[diffusion] feat: let every layerwise component be configurable (#35688)
mickqian Aug 20, 2026
be37339
[diffusion] feat: support out-of-tree models and pipelines (#35713)
mickqian Aug 20, 2026
81df6f2
[MUSA] Harden CI dependencies and diffusion warmup (#35610)
yeahdongcn Aug 20, 2026
ba97cc6
Skip empty linear-attention state buffers in PD transfer (#35689)
ispobock Aug 20, 2026
23cb040
fix(multimodal): keep LLaVA image fetch off the CPU-preprocess timeou…
ShangmingCai Aug 20, 2026
2ef0fe4
TP/PP Consensus checker (#34406)
stepinto Aug 20, 2026
61fa64a
feat(grpc): expose KV event discovery metadata (#35714)
Aug 20, 2026
308bc12
📝 [NPU] Clean up quantization comments (#34829)
TallMessiWu Aug 20, 2026
eac91ac
[Fix] Land the decode mamba checkpoint depth on the tree page under D…
kpham-sgl Aug 20, 2026
d9f6861
[docs] Add DFlash2 speculative cells to the Qwen3.8-27B cookbook (#35…
Jiminator Aug 20, 2026
ad367d7
[Kimi K3] Select FlashInfer MXFP4 for SM107 auto MoE (#35554)
leejnau Aug 20, 2026
5a100d9
[misc] Trim restating comments and docstrings in srt/managers (#35622)
hnyls2002 Aug 20, 2026
0f744b6
feat: make mm_inputs msgpack-native (#29656)
ishandhanani Aug 20, 2026
92eeed4
[Docker] Defer CUDA 13 NCCL override until after dependency resolutio…
Fridge003 Aug 20, 2026
0149f56
[CI] Gate `/rerun-test` on commenter trust and remove `/rerun-stage` …
hnyls2002 Aug 20, 2026
94907f0
Add CI permissions for four contributors (#35600)
gongy Aug 20, 2026
a4ef828
fix(openai): avoid duplicate routed expert in response when `return_m…
guapisolo Aug 20, 2026
5a7b26c
[AMD] [sgl-kernel] Bypass caches for peer traffic in ROCm custom all-…
wenkaidu Aug 20, 2026
779e593
Fix _GenerationStreamAccumulator logprob_end off-by-one under retract…
shenxiul Aug 20, 2026
1a138e1
[docs] Tell Qwen3.8-27B DFLASH2 users to build from main (#35753)
Jiminator Aug 20, 2026
ba8e601
[docs] fix note formatting in sglang-d documentation (#35761)
AgainstEntropy Aug 20, 2026
67f6ad6
fix(kernel) Fix Helion small-token prefill bug (#35197)
ethche Aug 20, 2026
14795dc
[docs] Point the Qwen3.8-27B DFLASH2 note back at the rolling dev ima…
Jiminator Aug 20, 2026
f825d72
[Sampling] Restore finite top-k requirement for sampling masks (#35205)
nanjiangwill Aug 20, 2026
5bb981d
[diffusion] chore: read the cgroup this process is actually in (#35707)
mickqian Aug 21, 2026
aa3f766
[CI] Temporarily disable B300 jobs (#35627)
mmangkad Aug 21, 2026
978244d
[P/D disagg] Decode-side radix cache for SWA hybrid models (unified r…
ishandhanani Aug 21, 2026
7e80e88
[diffusion] Fuse LTX-2.5 decoder 3D RoPE (#35698)
BBuf Aug 21, 2026
e0cf75d
[doc] standardize diffusion cookbook model pages (#34247)
mickqian Aug 21, 2026
44806dc
Using unified radix tree by default for all case (#35081)
hzh0425 Aug 21, 2026
6127d1d
[diffusion] feat: allow offloaded weights stay on the checkpoint mapp…
mickqian Aug 21, 2026
bda9952
[AMD] DeepSeek-V4 MI355X: eliminate bpreshuffle fp8-scale copies at p…
karverma-amd Aug 21, 2026
34180a0
[AMD] Improve K3 dspark draft attn kernel perf (#35499)
1am9trash Aug 21, 2026
78c964d
[AMD] Retry transient network failures in ROCm Dockerfile curl fetche…
bingxche Aug 21, 2026
3efa057
[docs] Retune the Qwen3.8-27B RTX 5090 DFLASH2 cells against 1cf2b8c …
Jiminator Aug 21, 2026
f64080f
[AMD] CI: cut two setup cycles from the AMD multimodal-gen lanes (#34…
michaelzhang-ai Aug 21, 2026
73a2c11
Support mxfp8 KV cache in PD transfer (#35718)
ispobock Aug 21, 2026
a688682
[AMD][CI] Fix ROCm 7.0's dead apt index fail the MORI dependency inst…
michaelzhang-ai Aug 21, 2026
44c90c6
[AMD] DSv4: fuse the qk-norm-rope pair on the MTP target-verify path …
karverma-amd Aug 21, 2026
f3fe815
add py env activate in xpu kernel release workflow (#35726)
ZailiWang Aug 21, 2026
6a12583
[AMD] Update ROCm AITER pin to c16d44b (#35810)
bingxche Aug 21, 2026
8a123cb
[Refactor] New EPD (#30398)
Aug 21, 2026
8ff9c2b
[mem_cache][9/N] refactor: move DSAIndexerPoolHost to pool_host.dsa (…
alphabetc1 Aug 21, 2026
4c98759
[AMD] fix(rocm): support flydsl 0.3.0 in the FlyDSL fused norm kernel…
bingxche Aug 21, 2026
896acc8
[Fix] Clear full-to-SWA mapping with `index_fill_` to avoid a blockin…
hnyls2002 Aug 21, 2026
0db2c53
Fix overlap prebuilt row reuse race (#35748)
jasonjk-park Aug 21, 2026
e7a37c8
refactor(disagg): remove dead build_and_send_encode_request (#35843)
ShangmingCai Aug 21, 2026
a5c52a9
[diffusion] Enable LongCat breakable CUDA graphs (#35724)
BBuf Aug 21, 2026
39d4d65
[diffusion] Accelerate SANA-Video linear attention in quality=high (#…
BBuf Aug 21, 2026
dad6fd0
refactor(disagg): remove dead get_embedding_port (#35844)
ShangmingCai Aug 21, 2026
5206f11
[diffusion] fix: stop the mapped-weight store from holding the parame…
mickqian Aug 21, 2026
5a46d65
[diffusion] refactor: resolve lora weight sources deterministically (…
mickqian Aug 21, 2026
4f343ab
refactor(disagg): remove unreferenced dead code (#35838)
ShangmingCai Aug 21, 2026
5ecd6d7
[diffusion] fix: fix quantized qkv scales and missing-param policy fo…
whn09 Aug 21, 2026
a41da99
refactor(disagg): collapse duplicated branches in get_kv_class (#35847)
ShangmingCai Aug 21, 2026
932f632
[diffusion] fix: do not warn that the recommended short edge is unver…
whn09 Aug 21, 2026
0447ade
[diffusion] fix: fall back to a component's default attention backend…
lgy1027 Aug 21, 2026
8658d00
[diffusion] feat: support loading peft lora (#35868)
mickqian Aug 21, 2026
a7ec6b9
Restructure mem_cache auto-labels by layer (#25122)
alphabetc1 Aug 21, 2026
0bdd28d
[mem_cache] docs: add a layer map and placement rules (#35643)
alphabetc1 Aug 21, 2026
70983bd
Add SGLang Granite SWA support via existing Granite models (#35794)
daviswer Aug 21, 2026
61c2da4
[Fix] Pass Anthropic thinking history as reasoning_content for custom…
mmangkad Aug 21, 2026
c373562
fix(grpc): derive choice count before normalization (#35778)
Aug 21, 2026
729a050
refactor(disagg): extract _all_reduce_polls helper (#35886)
ShangmingCai Aug 21, 2026
05c584c
docs: add DSPARK speculative decoding option to Ling-3.0-flash cookbo…
JustinTong0323 Aug 21, 2026
4d42def
chore: bump docs install version to 0.5.18 (#35911)
sglang-bot Aug 21, 2026
590b11a
[Runtime] Don't override CUDA_MODULE_LOADING (#35711)
HanHan009527 Aug 21, 2026
7d89325
[Spec][LoRA] Support multi-adapter LoRA with EAGLE/NEXTN/DFLASH/DSPAR…
jybsuper Aug 21, 2026
8344007
perf: overlap Qwen shared expert with DeepEP routed experts (#34938)
YAMY1234 Aug 21, 2026
2440820
docs: add website link to README header (#35210)
alisonshao Aug 21, 2026
fe8f9d7
[Docs] Add --prerelease=allow to cookbook uv install commands (#35920)
zijiexia Aug 21, 2026
7d7ab4b
Rainj me/rust server refactor2 (#35239)
rainj-me Aug 21, 2026
60ff1e3
[DeepSeek V4] Default FP4 checkpoints to FlashInfer MXFP4 MoE (#35919)
Fridge003 Aug 21, 2026
7fd5454
[DSA] Route the ragged prefill top-k to the v2 kernel (#35175)
DarkSharpness Aug 21, 2026
0ee3749
[Fix] Read the granite sinks dtype from the exec bag, not the legacy …
kpham-sgl Aug 22, 2026
22dafbc
[diffusion] feat: keep a cpu-started vae weights on the checkpoint ma…
mickqian Aug 22, 2026
b26695a
[diffusion] feat: reject unsupported quantized component checkpoints …
mickqian Aug 22, 2026
0be2a20
[diffusion] refactor: hand out pinned host memory per layer (#35867)
mickqian Aug 22, 2026
3b5909d
[DeepSeek V4] Add W4A4 MegaMoE server flag (#35918)
Fridge003 Aug 22, 2026
d90318b
[MLX] Upgrade to Torch 2.13/MLX 0.32+ and redesign the Torch-MLX tens…
yeahdongcn Aug 22, 2026
5662c03
Support CPU offload for mxfp8 KV cache (#35888)
ispobock Aug 22, 2026
fbafd1b
Add sampling observer auxiliary output hooks (#35747)
alecsolder Aug 22, 2026
0db2bdf
Fix buffer-mode HiCache load-back ownership races; add optional prefe…
xiezhq-hermann Aug 22, 2026
ac179ee
fix(disagg): PD transfer-failure injection was silently inert (#35890)
ShangmingCai Aug 22, 2026
5290327
[diffusion] feat: resolve hub component subfolders (#35939)
mickqian Aug 22, 2026
3e09662
[CI] Re-enable B300 jobs (#35607)
mmangkad Aug 22, 2026
83e9ece
[diffusion] Fuse SANA-Video interleaved RoPE (#35695)
BBuf Aug 22, 2026
96bfd24
[diffusion] Enable SANA-Video breakable CUDA graphs (#35729)
BBuf Aug 22, 2026
4cb5aeb
[docs] Re-measure the Qwen3.8-27B RTX 5090, RTX PRO 6000 and DGX Spar…
Jiminator Aug 22, 2026
c35683f
[HiCache] Clamp tombstoned SWA locs in UnifiedSWAKVPool translation (…
xiezhq-hermann Aug 22, 2026
af39ad9
[Model] Complete dots.note.omni support with native encoders, video p…
jianfei-wangg Aug 22, 2026
6fd0384
Make draft attention backends extensible (#35932)
merrymercy Aug 22, 2026
9035432
[VLM] Split Pixtral multi-image features before the CUDA IPC wrap (#3…
mmangkad Aug 22, 2026
b391ef1
[diffusion] feat: load serialized bnb4 components with transformers (…
mickqian Aug 22, 2026
15a4398
refactor(disagg): hoist duplicated _handle_staging_req into a mixin (…
ShangmingCai Aug 22, 2026
5c03069
refactor(disagg): drop dead placeholder overrides in Common KV sender…
ShangmingCai Aug 22, 2026
cb10ca1
[FEAT] Weight Daemon abstraction (#33279)
Aug 22, 2026
382343f
[diffusion] optimization: transfer mapped layers through a courier th…
mickqian Aug 22, 2026
489e605
[diffusion] feat: admit compatible quantized native encoders (#35962)
mickqian Aug 22, 2026
a8c16b2
[diffusion] feature: use the directory for the vae mapping gate (#35946)
mickqian Aug 22, 2026
3c69a4c
fix(test): unbreak test_kv_transfer_replica_metric after #35950 (#35974)
ShangmingCai Aug 22, 2026
61981e1
[diffusion] optimization: keep vae decoder weights in their decode dt…
mickqian Aug 22, 2026
db570fe
[diffusion] chore: let the auto policy select h3's dit for layerwise …
mickqian Aug 22, 2026
d315eb7
[AMD] DeepSeek-V4: add aiter fused mHC post+pre with cross-layer boun…
karverma-amd Aug 22, 2026
46cb12a
[diffusion] fix: fix hunyuan3d stale extension lock hangs (#35989)
mickqian Aug 22, 2026
453b98c
[diffusion] feat: support single-file component weight overrides (#35…
mickqian Aug 22, 2026
0b064e3
[diffusion] comfyui: add a minimax-h3 node and a generic extra-fields…
mickqian Aug 22, 2026
b98d472
[NPU] [Diffusion] Fix critical Ascend NPU Diffusion regression/bugs &…
OrangeRedeng Aug 22, 2026
7d22b7a
[diffusion] docs: add tuning guide for h3 on consumer-level gpu (#35816)
mickqian Aug 22, 2026
cce0a12
refactor(disagg): hoist staging helper imports out of the bootstrap l…
ShangmingCai Aug 22, 2026
eec794b
[AMD][Spec] Fix aiter GQA packing + split-KV routing in NEXTN spec at…
hsthe29 Aug 22, 2026
7f30d66
[Kimi-K3] Fix "wrong grids" crash in DP-sharded vision preprocessing …
fullyz Aug 23, 2026
a36c074
[diffusion] feat: support serialized comfy convrot int8 dits (#35994)
mickqian Aug 23, 2026
59cdc9d
[diffusion] chore: re-home decode-dtype vae weights to a file-backed …
mickqian Aug 23, 2026
bbbcbf9
[diffusion] fix: stabilize ltx-2.3 two-stage cold requests (#35997)
mickqian Aug 23, 2026
70319a0
[diffusion] feat: support loading serialized comfy convrot int8 nativ…
mickqian Aug 23, 2026
97ae278
[diffusion] feat: release a layerwise component's non-layer weights b…
mickqian Aug 23, 2026
bd72190
refactor(disagg): register SGLANG_ENCODER_MM_LOAD_WORKERS in Envs (#3…
ShangmingCai Aug 23, 2026
c9f6b9b
[NPU] [DOC] Refresh supported features and models on Ascend NPU (#35836)
amote-i Aug 23, 2026
849ce71
refactor(disagg): dedupe mooncake failure_exception into a mixin (#36…
ShangmingCai Aug 23, 2026
8270ac4
refactor(disagg): move _is_watermark_ready into StagingManagerMixin (…
ShangmingCai Aug 23, 2026
edd675c
[AMD] Add Radix-4 MoE top-k router kernel for Kimi-K3 routing (#34490)
RolaoDenthu Aug 23, 2026
155aa26
[AMD][DSV4] perf: use full 1024-thread block for indexer top-k on ROC…
karverma-amd Aug 23, 2026
bd3cc97
[diffusion] CI: let the 5090 consumer case runs two warm requests on …
mickqian Aug 23, 2026
6218d6c
config: a defensive publish must not re-project over a live process (…
ch-wan Aug 23, 2026
0e22777
config: record resolution writes in a declaration stash (#35905)
ch-wan Aug 23, 2026
4bc79a1
config: project the config bags from the resolution result (#35906)
ch-wan Aug 23, 2026
64aa859
config: constructing a config no longer resolves it (#35907)
ch-wan Aug 23, 2026
362c2ee
config: borrowed-record reads follow the config bags (#35908)
ch-wan Aug 23, 2026
a43592d
config: pin two orderings resolution relies on (#35909)
ch-wan Aug 23, 2026
340391a
config: publish before the launcher reads effective configuration (#3…
ch-wan Aug 23, 2026
e3a008a
[diffusion] UX: clean up startup and offload logs (#36034)
mickqian Aug 23, 2026
27aa48b
[Fix] lfm2 detector: recover tool calls dropped by common model-outpu…
fatday Aug 23, 2026
44db041
[NVIDIA] Fix SM107 MXFP8 activation prep (#35405)
csahithi Aug 23, 2026
dd15fb5
[diffusion] feat: automatically infer comfy fp8 activation scaling (#…
mickqian Aug 23, 2026
886e37a
[diffusion] CI: guard the anonymous-host budget alongside peak VRAM (…
mickqian Aug 23, 2026
9b1b06b
[NPU] [DOC] Add Ascend NPU (A3) recipe to the Kimi-K3 cookbook (#35508)
amote-i Aug 23, 2026
939c00a
[diffusion] feat: support compact qwen3-vl conditioning for minimax h…
mickqian Aug 23, 2026
8014d9d
[npu] Kill evalscope session by process group and fix report score pa…
pllimax Aug 23, 2026
de6a1db
[diffusion] feat: support hybrid conditioning for minimax h3 (#36080)
mickqian Aug 23, 2026
d1af3c8
[diffusion] feat: support loading native diffusers miniMax h3 compone…
mickqian Aug 23, 2026
95f5ecd
[AMD] Update amd deepseek v4 cookbook 0822 (#35854)
1am9trash Aug 23, 2026
fb6e387
[Mamba] fix mamba index h unexpected assertion for dcp (#36005)
billishyahao Aug 24, 2026
2006462
[AMD][CI] Add the Qwen3.8 MXFP4 MI35x nightly (#35383)
michaelzhang-ai Aug 24, 2026
f4448e6
[diffusion] Reuse SANA fast paths in SANA-Video BCG (#35961)
BBuf Aug 24, 2026
447048d
[diffusion] Reject unsafe quality=high BCG replay (#36008)
BBuf Aug 24, 2026
b2eb0fa
[diffusion] Keep Cosmos3 Nano resident on high-memory GPUs (#36000)
BBuf Aug 24, 2026
e129fe2
[diffusion] Flatten Wan VAE RMSNorm row addressing (#35981)
BBuf Aug 24, 2026
167c339
[CPU] Add check for fused_input_proj in TP=4 (#35669)
yanbing-j Aug 24, 2026
1c1c9d9
[diffusion] refactor: reuse srt quantization contracts and mxfp8 kern…
mickqian Aug 24, 2026
fee00a4
[diffusion] feat: add composable component weight path cli (#36078)
mickqian Aug 24, 2026
4c02584
Add intel_xpu to DETERMINISTIC_ATTENTION_BACKEND_CHOICES (#29143)
kalyank007 Aug 24, 2026
fd73d4b
[CPU] Add graph register for fused_sigmoid_mul_cpu, fused_qk_gemma_rm…
yanbing-j Aug 24, 2026
2d84de5
[diffusion] feat: support loading serialized comfy w4a8 checkpoints (…
mickqian Aug 24, 2026
f6fff25
[diffusion] feat: support vae weight-file overrides (#36085)
mickqian Aug 24, 2026
fbdec28
[XPU] Support INT4 dense linear (AWQ/GPTQ) for XPU (#30236)
YangKai0616 Aug 24, 2026
11b1b4c
[NPU] [DOC] Polish English wording in NPU docs (#36123)
amote-i Aug 24, 2026
514b997
Register CPU CI for 17 e2e tests and partition xeon base-c suite (#35…
1pikachu Aug 24, 2026
a90d770
[Weight Cache] Support static DP/EP layouts (#33684)
UNIDY2002 Aug 24, 2026
1a368ec
[diffusion] optimization: reuse minimax h3 prompt refinement across o…
mickqian Aug 24, 2026
6ca872a
[diffusion] chore: fetch metadata beside nested lora weights (#36057)
mickqian Aug 24, 2026
0c1e9bd
[OpenAI] Drop empty assistant turns for mistral_common tokenizers (#3…
alisonshao Aug 24, 2026
f294d51
[diffusion] fix: fix a refit key error on mapped weights, and stop cl…
mickqian Aug 24, 2026
230c052
[diffusion] chore: reuse srt AutoRound for quantized DiTs (#36068)
mickqian Aug 24, 2026
5ce700a
[diffusion] feat: infer LoRA alpha from safetensors metadata (#36082)
mickqian Aug 24, 2026
7a7b655
[quantization] share bounded post-load device staging (#35180)
mickqian Aug 24, 2026
acba892
[XPU] Support softmax_lse in sgl_kernel::fwd API (#33840)
Valentine233 Aug 24, 2026
09592f5
[diffusion] Keep LongLive2 components resident on large GPUs (#35993)
BBuf Aug 24, 2026
8dcfb3b
[diffusion] Fuse LongCat-Image QKNorm and interleaved RoPE (#35995)
BBuf Aug 24, 2026
5b5b29d
[XPU] Use a fused GDN kernel from sgl-kernel for Qwen3.5 (#33354)
Xia-Weiwen Aug 24, 2026
3e30649
[Fix] Harden FlashAttention CUDA graph metadata bounds (#35454)
aurickq Aug 24, 2026
1daa94a
[CPU] Fix NUMA/core binding for DP ranks (#32856)
chunyuan-w Aug 24, 2026
5683442
[Intel XPU] Add xpu pass for biased_topk and hash_topk (#33323)
gaopengff Aug 24, 2026
b498efc
chore: move cuda_vmm_utils.py under srt/utils/ (#36053)
merrymercy Aug 24, 2026
f98b60d
fix(xpu): enable compressed-tensors FP8 W8A8 on XPU (RedHatAI FP8-dyn…
vshekhawat-hlab Aug 24, 2026
344613c
[diffusion] Default Hunyuan VAE to tiled decode (#36012)
BBuf Aug 24, 2026
8df3b9e
[diffusion] feat: support loading mixed w4a8 text encoders (#36037)
mickqian Aug 24, 2026
3fe18f1
[diffusion] feat: add plain component weight overrides (#36086)
mickqian Aug 24, 2026
77940de
[MoE] Gather the cutlass MoE activation and its scales in one launch …
yuan-luo Aug 24, 2026
b43931e
[diffusion] Refresh quality and BCG benchmark skills (#36016)
BBuf Aug 24, 2026
6d40b8a
[diffusion] Fix Hunyuan QKV pack indexing at production video shapes …
BBuf Aug 24, 2026
ec334b1
xeon ci fail fast strategy change (#36146)
MingxuZh Aug 24, 2026
852b04b
fix(xpu): read enable_deterministic_inference from the config bag (#3…
ShangmingCai Aug 24, 2026
97b176e
Support streaming session on NPU (#32597)
sigama-w Aug 24, 2026
cc74aba
[diffusion] Honor XDG cache for model overlays (#36019)
BBuf Aug 24, 2026
9866fe9
[diffusion] Speed up LingBot high-quality VAE decode (#36024)
BBuf Aug 24, 2026
c8e1ddc
[diffusion] feat: cache LoRA-merged weights in files the page cache c…
mickqian Aug 24, 2026
e28b9cf
Npu single node test timeout config (#36092)
pllimax Aug 24, 2026
f331db3
Merge remote-tracking branch 'upstream/main' into cursor/sync-upstrea…
cursoragent Aug 24, 2026
2aeed3a
docs: refresh DSV4 framework PR checklist as of 2026-08-24
cursoragent Aug 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
198 changes: 198 additions & 0 deletions .claude/rules/comment-style.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,198 @@
---
paths:
- "**/*.py"
- "**/*.cu"
- "**/*.cuh"
- "**/*.cpp"
- "**/*.h"
- "**/*.rs"
---

# Comment Style

Applies to `#` and `//` comments, Python docstrings, and C/C++ Doxygen blocks
(`///`, `/** ... */`) -- Python, CUDA/C++, and Rust alike. Comments are reviewed
like code, with the burden of proof reversed: the author justifies a comment's
existence, not the reviewer its removal.

## Two modes

Writing a comment and editing someone else's are different decisions.

**Writing.** You still have the context that made the comment necessary, so the
judgment is reliable. Everything below `## Cleanup` is written for this case.

**Editing.** The judgment is unreliable and the error is one-way. A one-line
comment carries too little text to tell a restatement apart from the last anchor
for a cross-file fact -- deciding takes the whole function and its callers, more
context than the line costs. And the payoff is asymmetric: keeping a useless
one-liner costs a line of scroll, while deleting a load-bearing one deletes a
fact silently, inside a diff of two hundred deletions where no reviewer will
catch it. So: a bright line, not a judgment call.

> **A comment-only diff does not touch one-line comments.** It removes or
> condenses multi-line prose blocks. A one-liner is rewritten only for a reason
> of its own -- it is wrong, it is stale, or it is commented-out code -- never
> because the line below it says the same thing.

## Cleanup: what a comment-only diff may remove

- **Multi-line prose restating the code** -- the function name, the next line,
the branch condition, or the loop body, written out as a paragraph.
- **Python `Args:` / `Returns:` blocks** on anything that is not a documented API
surface -- see `## Documentation blocks` for the list of surfaces that keep
them.
- **Multi-line history and rationale** -- what the code used to do, why the
change was made, which approach was abandoned. Each of these has a home
elsewhere (see `## Where other explanations live`).
- **Commented-out code**, at any length.

Condense rather than delete when a block carries one fact that is not
recoverable: keep that sentence, drop the enumeration around it.

## What a comment is for

State the fact a reader cannot see from here. Delete the comment and ask what it
costs to recover: a fact that lives in another file costs everything, a grouping
the names only half-encode costs a tedious reconstruction, a restated line costs
nothing.

- **Cross-boundary constraints.** Layout, field order, or call order shared with
a CUDA kernel, the Rust router, an IPC schema, or another process. Nothing in
the Python file shows the other side, so nothing else can warn the next editor.

```python
# The PD wire schema must match on P and D even when only D runs spec decoding;
# a seedless prefill writes the invalid sentinel.
```

- **Units and layout the name cannot carry.** Tokens vs reqs vs pages vs bytes
vs slots, and tensor shape/dtype/layout. Encode it in the name first; comment
only when the name is fixed by an existing interface.

```python
# [num_tokens, num_kv_heads, head_dim], fp8 e4m3, page-major.
```

- **Where a magic number came from** -- not what it means. A hardware
constraint, a measurement, or an admission that it was picked arbitrarily.
The last one is the most valuable: it tells the next person the value is safe
to change.

```python
# Hardware constraint: TMA descriptors require 16B alignment.
# Measured on fp8 GEMM; re-tune when the kernel changes.
# Arbitrary; no evidence this is the right threshold.
```

- **Workarounds, anchored to a verifiable reference and a retirement
condition.** An unanchored workaround is immortal -- nobody can prove it is
safe to delete.

```python
# Workaround for pytorch/pytorch#12345; drop once we require torch >= 2.9.
# Temporary workaround: Event.wait() regresses TPOT on AMD MI355.
```
(second line: `python/sglang/srt/managers/overlap_utils.py:471`)

- **Contracts that are a decision, not a mechanism** -- a sentinel's meaning, a
deliberate omission from a list, an ordering that looks incidental.

```python
# None means the request is excluded from the radix cache, not that lookup failed.
```

- **Structure the names only partly encode** -- a group boundary in a long flat
block, a section split in a long body. `# ===== Helpers =====` above two
functions costs nothing to see past; the same banner splitting a genuinely long
flat module (see `python/sglang/srt/environ.py`) is the only statement of where
one group ends.

Not worth writing, at any length: prose that restates the line below it,
`# Step 1:` / `# Step 2:` numbering over straight-line code (extract named
helpers if the flow needs numbers), the names of the callers (that is what grep
is for), and hedging -- "this should probably be revisited" is either a fact to
establish and state, or a `TODO(<owner>)`.

## Form and tags

- **One or two lines.** A genuinely intricate invariant may run longer; that is a
rare exception, not a licence. Documentation blocks are the separate case
below.
- **ASCII and English only.** No Unicode arrows, math symbols, or CJK.
- **Break at a clause boundary.** If a comment needs a second line, wrap after a
semicolon or a comma -- never mid-phrase. A sentence that will not split
cleanly is a sentence that should be shortened instead.
- **Attach the comment to the line it constrains**, not to the top of the
function. Comments collected into a preamble are the ones that go stale.
- **You change the line, you own its comment.** Update it or delete it -- never
leave it orphaned. A stale comment is worse than no comment.

Two tag spellings only. New code does not use `FIXME`, `XXX`, or `HACK`; existing
occurrences are grandfathered and get folded into these when the line is touched.

- `# NOTE:` -- a constraint or trap. Most of the time the prefix adds nothing;
drop it and just state the fact.
- `# TODO(<gh-handle>):` -- planned work, **always** with an owner or an issue
link. A bare `# TODO: fix this` is rejected in review; an unowned TODO is
never retired. `TODO(perf)` and similar topic tags are acceptable when the work
is a standing category rather than one person's task.

## Where other explanations live

Every explanation has exactly one home. A second copy in a comment drifts from
the first.

- **What the code used to do -> `git log`.** Comments describe the current
state: not what was tried first, not which approach was abandoned. The
exception is a past failure that is still a live constraint -- write it as the
constraint, not as the story.

```python
# Bad: We used to call this before init_new but it broke CUDA graph capture.
# Good: Must run after graph capture; capturing this allocation deadlocks.
```

- **Why the change was made -> the PR body.** Why now, what else was tried,
which benchmark moved. "Now we also handle the case where ..." argues for a
diff, and is written for a reviewer who is long gone.
- **Design rationale -> `docs/` or a module docstring. CLI semantics -> the
`ServerArgs` help text. Test intent -> the test.**
- **Superseded code -> `git show`.**

Also out: review attribution ("as suggested in review") and our own PR numbers
used as a changelog. An upstream issue URL is different -- it is not a changelog
entry but a workaround's retirement condition, and the section above requires it.

What stays next to the line is why *this line* exists, written for a reader who
has neither the PR nor the discussion.

## Documentation blocks

Python docstrings and Doxygen blocks are part of the product on the surfaces
users, integrators, and tooling read, and noise everywhere else. Where warranted
they may run past two lines; every other rule in this file still applies.

- **Yes:** `Engine` and entrypoint public methods, `ServerArgs` help text, base
classes third parties subclass (attention backends, quant methods, model
extension points), and `sgl-kernel` op signatures -- where the shape/dtype
contract is the documentation.
- **Yes:** exported C++ / CUDA entities, as Doxygen with `\brief` / `\param` /
`\tparam` / `\return`. clangd renders these on hover, driven by
`CommentFormat: Doxygen`, so the per-parameter enumeration is the
caller-facing contract rather than filler.
- **Yes:** a bug-regression test, stating in one or two lines which black-box
behavior must not come back -- the live constraint, not the incident. Root
cause and repro belong in the issue and the PR
(see `.claude/rules/unit-test-admission.md`).
- **No:** internal helpers, overrides, private methods. Never hand-write or
generate Python `Args:` / `Returns:` blocks -- the signature and its type
hints already carry the names and types.

The one exception is the orchestration step. In the frozen classes
(`Scheduler`, `TokenizerManager`, `ModelRunner` -- see
`.claude/skills/large-class-style`), each `init_*` / `handle_*` / event-loop
method is an override point a subclass picks from a list; its one-line
docstring is the catalog entry that says what the step does, so it stays even
when it reads close to the method name. A genuine private helper
(`_validate_*`, `_normalize_*`) that no subclass overrides still gets nothing.
2 changes: 1 addition & 1 deletion .claude/rules/general-code-style.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ Default conventions for new and modified Python code. Prefer these unless there

- **Prefer stateless.** Favor pure functions over methods that mutate instance state; pass inputs in, return outputs out.
- **Prefer immutable.** Default to immutable data (frozen structs, tuples, read-only values); mutate only when there is a clear need.
- **Extract init-static values at construction.** When a derived value's inputs are frozen for the object's lifetime (typically configuration: constructor args, env vars, server args), compute it once in `__init__` and store it as a well-named attribute (`self.mtp_enabled`, `self.needs_cpu_seq_lens`); later code reads the attribute instead of re-deriving it. Input immutability is the hard precondition — if inputs can change, recompute in place or funnel mutation through a single override point (the frozen `ServerArgs.override()` pattern). If you can't give the value a meaningful name, the boundary is wrong — don't cache unnameable subexpressions.
- **Extract init-static values at construction.** When a derived value's inputs are frozen for the object's lifetime (typically configuration: constructor args, env vars, server args), compute it once in `__init__` and store it as a well-named attribute (`self.mtp_enabled`, `self.needs_cpu_seq_lens`); later code reads the attribute instead of re-deriving it. Input immutability is the hard precondition — if inputs can change, recompute in place or funnel mutation through a single override point (`get_context().override()` for resolved config). If you can't give the value a meaningful name, the boundary is wrong — don't cache unnameable subexpressions.
- **Functions stay small.** Keep each function under ~100 LOC; split larger ones into named helpers.
- **Files stay small.** Keep each file under ~2k LOC; split larger modules along cohesive boundaries.
- **Core functions read like pseudocode.** The main / orchestration function of a unit should be short and read like algorithm pseudocode — push detail into well-named helpers so the top-level flow is obvious.
Expand Down
Loading
Loading