Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
172 commits
Select commit Hold shift + click to select a range
35faf95
[Bugfix][KV Offload] Ignore stale async lookup results (#54872)
Alex-ai-future Sep 2, 2026
ba6c60e
[Bugfix][KV Offload] Scale UniformTypeKVCacheSpecs groups by DCP (#50…
drakosha Sep 2, 2026
872084f
[CI] Add Kimi-K3-pruned75-DSpark-TP4 gsm8k eval (#54817)
mgoin Sep 2, 2026
c6bca6e
[Bugfix][Multimodal] Scope cache hash kwargs by modality (#54918)
waizuichougou Sep 2, 2026
41848ca
[NIXL] Use int32 array for indices to avoid intermediate conversion (…
iyastreb Sep 2, 2026
3b45d05
[Bugfix][Model] Fix CohereASR streaming audio-token estimate (unit + …
hungnnvidia Sep 2, 2026
d539de1
[Docs] Add missing return annotations flagged by griffe (#54980)
hmellor Sep 2, 2026
605c3dd
[BUILD] Bump cutlass to v4.7.1 (#54190)
Harry-Chen Sep 2, 2026
488e6fd
[CI] Revert flaky `test_quark_int8_w8a8_moe` (#54991)
fxmarty-amd Sep 2, 2026
1356635
[New model][Multimodal] Add DeepSeek-V4-Flash-Vision-Exp support (#54…
Isotr0py Sep 2, 2026
bf7a14d
[CI][ROCm] Add DSpark evals (#54852)
AndreasKaratzas Sep 2, 2026
2691c6c
[Bugfix][CI] Set cudagraph_mode=FULL for the Ernie4.5-VL ViT cudagrap…
stefankoncarevic Sep 2, 2026
3e9d364
[CI/Build][ROCm] Guard the two CUDA-only tests in test_bf16_skinny_ge…
stefankoncarevic Sep 2, 2026
9b38e3a
[CI][MoE] Moe kernels test cleanup (#54954)
stefankoncarevic Sep 2, 2026
1945a94
[Bugfix][Tests] Stabilize B12X linear kernel checks (#54996)
lukealonso Sep 2, 2026
0e3ac49
[ROCm][CI] Fix false multi-node detection on native CI (#54989)
sheralskumar Sep 2, 2026
a56654d
[K3 Perf] Enable DSV3 GEMM for inner-contiguous and row-strided tenso…
yewentao256 Sep 2, 2026
60857ba
[Bugfix][Rust Frontend][Renderer] Align DeepSeek V4 historical develo…
reidliu41 Sep 2, 2026
e3e1241
[ROCm][CI] Extend Multimodal Processor Shard timeout on AMD CI (#55011)
micah-wil Sep 2, 2026
963054e
[CI] Exclude nightly-dev tags from nightly DockerHub cleanup (#55023)
khluu Sep 2, 2026
a0d3e5c
[XPU] Add fused GemmaRMSNorm path for eager execution (#53678)
ccrhx4 Sep 2, 2026
6c6376a
[CI] Remove deleted nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1…
khluu Sep 2, 2026
cf3263d
[Agents] Add Triton kernel-writing skill (#55019)
WoosukKwon Sep 2, 2026
443febe
[Agents] Expose Triton kernel-writing skill to Claude (#55028)
WoosukKwon Sep 2, 2026
ad127d9
[Perf][Rust Frontend] Coalesce decoded chunks per engine update (#55012)
BugenZhao Sep 2, 2026
0e14198
[Skills] Add kernel benchmark sanity references (#54995)
WoosukKwon Sep 3, 2026
5e4e927
[CI] Force HTTP/1.1 for runtime Git installs (#55044)
khluu Sep 3, 2026
e47356c
[ROCm][Installation] Add mooncake package to image using public wheel…
giuseppegrossi Sep 3, 2026
5d09eb2
[Bugfix][MoE] Preserve unquantized weight storage on ROCm (#46009)
aaab8b Sep 3, 2026
27a94d1
[CI] Fix DeepSeek-V4 registry platform guard (#55042)
taneem-ibrahim Sep 3, 2026
3f1af35
[Bugfix][Multimodal] Handle prefix-covered items in SHM worker cache …
waizuichougou Sep 3, 2026
ee17d0d
Use server-generated keys for late-interaction query caches (#51445)
KernelClint Sep 3, 2026
859dd39
[Bugfix][Core] Wait for the previous PP tensor sends before the next …
djw8605 Sep 3, 2026
092334c
[Bugfix] Wait for offload keys before storing chunks (#52923)
982945902 Sep 3, 2026
1f76efa
[Model] Add K2-Horizon model support (#55063)
tanyuqian Sep 3, 2026
096d8e8
[ROCm][CI] Add MiniMax reduce RMS kernel coverage (#55057)
AndreasKaratzas Sep 3, 2026
e6eb907
[CI] Bump Transformers version to 5.16.1 (#53905)
hmellor Sep 3, 2026
4cc0cb6
[CI][AMD] Avoid expandable segments in LoRA TP tests (#55094)
AndreasKaratzas Sep 3, 2026
c21751c
[Kernel] Warm up Qwen GDN gated RMSNorm (#54251)
zupengwang Sep 3, 2026
0d3ede3
[Bugfix][Model] Enable torch.compile for StableLM (#54969)
djramic Sep 3, 2026
facd9a7
[Model Runner V2][Spec Decode] Skip DP sync for all speculator unifor…
TheEpicDolphin Sep 3, 2026
ee0a4c4
[Bugfix] Account for PCP in multi-node world size validation (#55111)
DebugSy Sep 3, 2026
758c79e
[Bugfix] Retain vocab embeddings during replacement (#55083)
taneem-ibrahim Sep 3, 2026
848ab13
[Perf] Accumulate Conformer attention scores with baddbmm (#55062)
Levius-Fubuki Sep 3, 2026
bb363db
feat: Add support for reasoning_token_count to reasoning parser (#54982)
jasonozuzu-cohere Sep 3, 2026
bf95f58
[Core] Triton kernel for small-batch top-p only masking (#54651)
njhill Sep 3, 2026
edc0fb7
Optimize PLE MTP metadata transfers (#55054)
byshiue Sep 3, 2026
31e9c13
[Bugfix][KV Connector] Safely fill circular buffers in DecodeBench (#…
majunze2001 Sep 3, 2026
e55b93f
[Core] Deprecate "all" mamba cache mode (#55041)
njhill Sep 3, 2026
4ae6228
[Bugfix][KV Connector] Populate SimpleCPUOffload BlockStored metadata…
mevince Sep 3, 2026
da8ec28
[Bugfix][KV Offload] Do not let a recurrent group's unhashed block tr…
yifjiang Sep 3, 2026
98ed085
[Model] add GLM-5.3-Flash support (#53906)
ZJY0516 Sep 3, 2026
d410fc1
[Kernel] Enable Kimi-K3 SiTU on the CuteDSL MoE backend and the SM107…
BolinSNLHM Sep 3, 2026
d4d703c
[Bugfix][Model] Fix FP8 PLE loading in mixed ModelOpt checkpoints (#5…
sychen52 Sep 3, 2026
b762406
[Fusion] Manual `ActivationQuantFusionPass` initial application (#51415)
mgoin Sep 3, 2026
8bf3963
[ROCm][Perf] Add low-M FP32 router GEMM for gfx950 (#54845)
Fangzhou-Ai Sep 3, 2026
21a2211
[Rust Frontend] Use token-attributed text in reasoning and unified pa…
BugenZhao Sep 3, 2026
2db1c4d
[Rust Frontend] Support `--lora-modules` for static adapter loading (…
wseaton Sep 3, 2026
fc8f107
Fix DeepSeek V4 FlashMLA auto KV cache dtype (#45091)
Yuzu23 Sep 3, 2026
2a336d8
[warmup] overlap renderer warmup and engine core initialization (#54557)
andyxning Sep 3, 2026
cee0f92
[ROCm][Perf] Optimize MiniMax-M3 decode indexer and top-k (#54682)
Fangzhou-Ai Sep 3, 2026
e410111
[CI] Avoid logging test server environment values (#54379)
taneem-ibrahim Sep 3, 2026
6fdee17
[CI] Zen5 image build (#50314)
andy-neuma Sep 3, 2026
c7e6e36
[Rust Frontend] Add support for TLS in render server (#54999)
zdtsw Sep 3, 2026
9509fc8
[Perf][Kimi-K3] Cut MLA decode concat/cache epilogue latency (#54896)
zyongye Sep 3, 2026
bc2ee48
[Perf] Prefetch the weight before the PDL wait in fused_q_kv_rmsnorm …
zyongye Sep 3, 2026
579aef4
[ROCm] Bump AITER to 0.1.21.post1 (#52826)
Rohan138 Sep 3, 2026
d6bce42
[Rust Frontend] Report reasoning tokens in chat completion usage (#54…
BugenZhao Sep 3, 2026
560ef78
[Perf][Model Runner V2] Compact sampling masks on GPU instead of unpa…
aoshen02 Sep 4, 2026
19c018e
[Bugfix][DCP] Materialize prefill keys on non-owner ranks (#54908)
foraxe Sep 4, 2026
d9e2b52
[Core][KV Events] Echo session_id on GPU BlockStored events (#51381)
xuhuan51 Sep 4, 2026
7dc30f5
[ROCm][CI] Build and publish TheRock nightly docker images (#55014)
Rohan138 Sep 4, 2026
25268f0
[KVConnector] Add retention interval to OffloadingConnector (#51886)
bnellnm Sep 4, 2026
a26b71d
[Bugfix][PD] Pad resumed speculative decode requests (#55126)
ZeldaHuang Sep 4, 2026
69cf055
[Performance][DSv4] Size dequant gather launch grid by rows (#55061)
aoshen02 Sep 4, 2026
8a0a7ee
[EC Connector] P2P NIXL + CPU EC Connector (#47941)
omerpaz95 Sep 4, 2026
1560505
[ROCm][CI] Bump ROCk base to ROCm 10.0 (#55246)
Rohan138 Sep 4, 2026
e352986
[Bugfix][KV Connector] Fix DecodeBenchConnector prefix block selectio…
majunze2001 Sep 4, 2026
78300cd
[Bugfix][Docs] Package glm5next nvidia subtree and fix its docstrings…
hmellor Sep 4, 2026
da55494
docs: note that enforce_eager also disables torch.compile (#55271)
AnshulDesai Sep 4, 2026
8a72866
[Bugfix] Reject tokenizer-less Qwen VL processor init (#54886)
luyixiao95 Sep 4, 2026
ae2d1ca
[Rust Frontend] Record Mooncake/NIXL KV-connector metrics (#52755)
ilmarkov Sep 4, 2026
e862c2f
[Docs] Add example for Renderer.render_cmpl() usage (#54265)
null-Exception1 Sep 4, 2026
29af8bd
[XPU][UT] skip GLM-5.3-Flash test on XPU (#55266)
mayuyuace Sep 4, 2026
8f816a3
[MRV2][Metrics] Support `CUDAGraphStat` in MRV2 (#52358)
yiz-liu Sep 4, 2026
a5c9179
[Qwen3.8-Flash-Next] Compact indexer logits workspace to improve pref…
gau-nernst Sep 4, 2026
a8693df
[SpecDecode]Fix spec decode warmup device selection (#55245)
jikunshang Sep 4, 2026
605ca45
[XPU] Fix device assignment for DP external LB (#53037)
hlin99 Sep 4, 2026
9cd956c
[CI/Test] Add expert parallelism coverage to external LB tests (#53497)
hlin99 Sep 4, 2026
fd4a151
[Kernel] Reuse Qwen4Exp HC combine-norm for MTP input (#54687)
gcanlin Sep 4, 2026
3ff4f02
[AMD][kimik3][ROCm][Perf] Fuse MLA q/kv RMSNorm in AMD Kimi-K3 MLA wr…
rbrugaro-amd Sep 4, 2026
c615b1f
[Mypy] Fix mypy typing for L models (#54177)
taneem-ibrahim Sep 4, 2026
8b6de0e
[Security] Cap GLMGA video sampling to prevent request-driven resourc…
jperezdealgaba Sep 4, 2026
8340fe1
[CI/Build] Fix flaky failures in CPU CI image building (#55317)
bigPYJ1151 Sep 4, 2026
31a8a26
[Qwen3.8-Flash-Next] Improve QSA sparse GQA for prefill and short-ctx…
gau-nernst Sep 4, 2026
07ee97e
[Test] Split test_sampling_mask_preserves_top_k_boundary_ties to remo…
mayuyuace Sep 4, 2026
a69e75b
Fast Start (#54921)
liusy58 Sep 4, 2026
a85d073
[Bugfix] Fix double BOS in LLM.chat() for multimodal models (#55288)
zhang-keliang Sep 4, 2026
f19431e
[Bugfix][MoE] Allow TRTLLM FP8 block-scale MoE with SwiGLU clamp (#55…
aoshen02 Sep 4, 2026
5e729d8
[CI] Only run GitHub Actions on the main repo (#55349)
hmellor Sep 4, 2026
131e028
[Frontend] Warn when removed guided-decoding fields are present in a …
erdholion Sep 4, 2026
6a039f4
[feat] add torchcodec as audio loader and implement selective audio b…
JaredforReal Sep 4, 2026
99a1ab8
[ROCm][CI] Bump ROCk release image build timeout to 3h (#55354)
Rohan138 Sep 4, 2026
8ad2076
[Transformers backend] Find attention with a fuser and attach vLLM's …
bohnstingl Sep 4, 2026
761c586
[Rust Frontend][gRPC] Preserve multimodal metadata for remote-prefill…
Sep 4, 2026
1ff5edb
[Bug-fix] Fix MoE fused sum row offsets (#50220)
happyyzy Sep 4, 2026
8cd95f7
[Bugfix][MLA] Restore DSpark cache-group capability under optimized P…
GirasoleY Sep 4, 2026
701a744
[CI] Raise AMD Spec Decode Eagle 1 job timeout to 35min (#55136)
JaredforReal Sep 4, 2026
a11dfcf
[CI] Surface why the CRCR nightly report cannot read its Buildkite se…
atalman Sep 4, 2026
5690b02
[Quantization] Support online quantization with partially pre-quantiz…
fxmarty-amd Sep 4, 2026
2524051
[Bugfix][Spec Decode] Honour the draft's attention_backend on Model R…
stecasta Sep 4, 2026
3284af6
[Bugfix] Reject 0 or non-positive max concurrency (#54887)
taneem-ibrahim Sep 4, 2026
c81ace1
[Perf] Read VidCom2 frame budgets once (#55331)
Levius-Fubuki Sep 4, 2026
5093e48
[Bugfix][Spec Decode] Drop FlashAttention's AOT schedule for a slidin…
SubSir Sep 4, 2026
784cac7
[Mypy] Fix mypy typing for P models (#54169)
taneem-ibrahim Sep 4, 2026
3f41d10
[1/2][Model Runner V2] DBO support, eager mode (#50945)
specture724 Sep 4, 2026
8f269a9
Revert "[CI] Remove deleted nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reaso…
mgoin Sep 4, 2026
685074c
[Bugfix][V2] Warm up kernels before capturing CUDA graphs (#55341)
aoshen02 Sep 4, 2026
eb74fbb
[Bugfix][NIXL] Don't assert when a failed transfer is cleaned up twic…
jyizheng Sep 4, 2026
874df93
[Bugfix] Preserve Mamba state for padded prompt tails (#55178)
natsala13 Sep 4, 2026
8277c42
[Perf] Ensure async h2d copies are pinned in more places (#55202)
njhill Sep 5, 2026
6cbb3c1
[Perf][GDN] Build cudagraph-capture metadata without a device sync (#…
njhill Sep 5, 2026
385ba6b
[Bugfix][Quantization] Register Quark per-block FP8 scales as weight_…
jimmy-adams Sep 5, 2026
16328c7
[Perf] Kimi K3 nvfp4 Align in_proj weights by 128 to avoid elementwis…
wzhao18 Sep 5, 2026
a7f9be9
[CI][ROCm] Raise AMD job and server readiness timeouts (#55308)
AndreasKaratzas Sep 5, 2026
ea5f41e
[CI][ROCm] Restore Wikitext coverage for Qwen OCP-MX (#55410)
AndreasKaratzas Sep 5, 2026
d4ef5e5
[CI][ROCm] Disable Transformers nightly groups (#55409)
AndreasKaratzas Sep 5, 2026
d87a440
[Core][MRV2] Support eagle3 spec decode with pipeline parallel (#50514)
yongqinwang-cmd Sep 5, 2026
4ee2595
[Attention] Sync FA with upstream (#54819)
StevenWang-CY Sep 5, 2026
e962733
[Security] Validate cache salts before they reach LMCache (#51444)
KernelClint Sep 5, 2026
8369aff
[Feat] Add EPLB support for GLM-5.3-Flash (#55119)
chaunceyjiang Sep 5, 2026
32601ef
[Perf][Multimodal] Avoid duplicate text embedding in Qwen2.5-Omni (#5…
waizuichougou Sep 5, 2026
e473e90
[Docs][Models] Use the official FunASR Nano vLLM checkpoint (#54944)
LauraGPT Sep 5, 2026
2902ca1
[Bugfix] complete VLLMValidationError migration in chat_utils.py (#50…
AdaAibaby Sep 5, 2026
7fbd44c
[Bugfix][DSv4] Seed the -1 sentinel in the prefill sparse index works…
aoshen02 Sep 5, 2026
bc96d76
[Kernel] Fall back from persistent top-k on low-shared-memory GPUs (#…
lucamotz Sep 5, 2026
28e605f
[Bugfix][Qwen4Exp] fix state index strides in fused PLE conv (#55375)
peakcrosser7 Sep 5, 2026
52bc900
[Bugfix][Kernel] Build fused GDN MTP decode for SM110 (#53835)
wei-core Sep 5, 2026
7985444
[Kernel][HY V4] Add Triton iHC pre/post fallback (#55059)
linitra24 Sep 5, 2026
f4eccda
[Bugfix][Multimodal] Bound renderer warmup to the prefill token budge…
lucamotz Sep 5, 2026
039ea82
[Core] Sync DP state on the first step of a wave (#52957)
aoshen02 Sep 6, 2026
1c344ed
[CPU] [Feat] Add native AMX-FP8 attention impl for Diamond Rapids (#…
zhejiangxiaomai Sep 6, 2026
f2e2936
[Kernel] Add fused MoE tuned config for E=256,N=512 on NVIDIA A100 80…
bakiburakogun Sep 6, 2026
1970f3e
Validate scale-out multimodal data before engine handoff (#51898)
KernelClint Sep 6, 2026
144e79c
[Kimi K3] Support internal prefix checkpoints with partial prefix cac…
ZeldaHuang Sep 6, 2026
a1541f5
[Perf] Use SDPA for BLIP-2 Q-Former attention (#55285)
Levius-Fubuki Sep 6, 2026
9afb878
[Bugfix][KV Offload] Skip cleaned-up async lookup batches (#55075)
Alex-ai-future Sep 6, 2026
dd07601
[Bugfix] Fall back to T1 when ARC cannot reclaim enough entries from …
zupengwang Sep 6, 2026
569adb5
[Bugfix][KV Offload] Fix SWA store reachability during chunked prefil…
Whamp Sep 6, 2026
dc02934
[Bugfix][KV Offload] Stop offloading the final sampled token's KV slo…
almogtavor Sep 6, 2026
808f8cd
[Governance] Add aoshen02 as code owner for RL components (#55529)
aoshen02 Sep 6, 2026
52358e6
[Frontend] Expose multimodal metadata for disaggregated prefill (#54659)
zhouyou9505 Sep 6, 2026
722d169
[CI] Align extraction test with canonical auxiliary layer order (#55457)
khluu Sep 6, 2026
6865e67
[Bugfix] Defer adaptive verification until after kernel warmup (#55455)
khluu Sep 6, 2026
4df8018
[Kernel] SM 12.x blockwise FP8: swizzle the CTA raster when the weigh…
jschmied Sep 7, 2026
199cb9b
[Kernel] Remove unused fake implementation (#55535)
jeejeelee Sep 7, 2026
294fbb4
add 2/3/5/6/7 CUDA support in AutoRound format (#52890)
wenhuach21 Sep 7, 2026
de69e82
[ROCm][Perf][DeepSeek V4] Fuse native FP8 shared expert with MXFP4 ro…
Fangzhou-Ai Sep 7, 2026
1f77848
[Bugfix][Offloader] Preserve prefetch static-buffer slot ownership (#…
Big2Wheel Sep 7, 2026
c3ec0d2
[Bugfix] Fix cuda profiler missing bug (#55237)
wzhao18 Sep 7, 2026
6748217
[Bugfix] Fix Kimi K3 NVFP4 MoE weight conversion OOM (#55407)
wzhao18 Sep 7, 2026
3dc7a68
[Perf] Extend Qwen Triton warmup to avoid first-request latency spike…
vhagor Sep 7, 2026
f43ef15
[CI] fix pre-commit (#55630)
ZJY0516 Sep 7, 2026
9cc7793
[Bugfix] Gracefully handle unsupported reasoning_effort in chat templ…
frankie-ys Sep 7, 2026
d9105ea
[Qwen3.8-Flash-Next] Remove torch.compile for NVIDIA implementation (…
gau-nernst Sep 7, 2026
ed29dfa
[XPU] Use fused_input_norm kernel in FusedInputNorm (#52945)
zufangzhu Sep 7, 2026
f7f060d
[Bugfix][Audio] Restore soundfile-first automatic decoding (#55642)
AndreasKaratzas Sep 7, 2026
6fbb00b
[EPD] Add ECMooncakeConnector for encoder cache over Mooncake Transfe…
stmatengss Sep 7, 2026
195bc9c
[CI][ROCm] Temporarily skip unsupported HY-V4 initialization (#55653)
AndreasKaratzas Sep 7, 2026
34b1e9f
[Bugfix][MooncakeStore] Fix finish-time save crash on hybrid models (…
zhewenl Sep 7, 2026
5893426
[Bugfix] DSv4 MXFP4 selector: stop narrowing explicit aliases to thei…
lucifer1004 Sep 7, 2026
fc5805b
re-quantization API, MXFP8 -> FP8 PTPC
fxmarty-amd Sep 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
19 changes: 19 additions & 0 deletions .agents/skills/kernel-microbenchmark/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,25 @@ description: Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUD
- Keep metadata setup, plan construction, allocation, random input generation,
and logging outside the timed region unless that overhead is the experiment.

## Sanity-Check Reference Numbers

Use these as rough reference points for large, well-shaped workloads, not as
gold standards, guaranteed peaks, or hard limits. Hardware SKU, clocks, shape,
precision conventions, and the FLOP/byte accounting can move the result. A
large gap is a prompt to investigate, not proof that a kernel is poor.

| Kernel regime | Hardware | Rough reference |
| --- | --- | ---: |
| Memory-bound, large batch | Blackwell | 6 TB/s |
| BF16 GEMM | Blackwell | 2 PFLOP/s |
| BF16 attention | B200 | 1.6 PFLOP/s |
| FP8 GEMM | Blackwell | 4 PFLOP/s |
| FP8 attention | B300 | 2.8 PFLOP/s |

The BF16 attention reference is approximately the 1613 TFLOP/s result reported
by the FlashAttention-4 paper. Compare kernels only with matching workload and
throughput conventions.

## Multi-GPU Benchmarks

- State whether the run is local or multi-node and report the GPU topology,
Expand Down
62 changes: 62 additions & 0 deletions .agents/skills/triton-kernel-writing/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
---
name: triton-kernel-writing
description: Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.
---

# Triton Kernel Writing

## Implementation

- Follow the official
[Triton semantics](https://triton-lang.org/main/python-api/triton-semantics.html).
Check it when behavior may differ from Python or NumPy, especially type
promotion, integer division and modulo, casts, broadcasting, and variable
scoping.
- Use the Triton kernel generated by `torch.compile` as a possible
implementation to inspect. Print Inductor's generated code with
`TORCH_LOGS="output_code" .venv/bin/python <script>` or enable
`torch._logging.set_logs(output_code=True)` before the compiled function
runs. Treat generated code as a reference, not as proof of correctness or
optimality.
- Find reasonable defaults for compile-time knobs such as `BLOCK_SIZE`, or use
a small, legible heuristic when workloads need different choices. Use
`triton.autotune` only when tuning is critical to performance, such as for a
matrix multiplication. Otherwise prioritize simple code and fast startup.
- Be careful to avoid unintended runtime JIT compilation. For example, put
unimportant runtime integer scalars in `do_not_specialize`, especially those
that may alternate between values such as 0 and 1, which can produce
different specialization keys.
- The Triton compiler does not guarantee safe ordering when a kernel writes to
a pointer and subsequently reads from the same pointer. This pattern must
have a `tl.debug_barrier()` between the write and read. The barrier
synchronizes threads in the block; it does not synchronize separate program
instances.

## Launch and Indexing

- `grid[1]` and `grid[2]` must be at most 65,535. Choose or flatten the grid
order so those dimensions cannot exceed the limit for supported shapes.
For example, `num_tokens` is commonly 8K or 16K, but users may configure 32K
or more. If `num_tokens` is a grid dimension, it is safe to put it in
`grid[0]` (or tile it).
- Use `int64` for offset arithmetic when an index can exceed 32-bit range,
especially for KV-cache addressing. Cast operands before multiplication or
addition so an intermediate does not overflow in 32-bit arithmetic.
- A `[num_tokens, num_heads]` grid can be a good low-latency mapping for decode,
but it can be very slow for prefill. If the kernel serves prefill, consider
tiling tokens or otherwise increasing the work and locality per program.

## Validation

- Check correctness at boundary shapes and at sizes that exercise masks and
large offsets.
- Choose accumulation and intermediate dtypes explicitly. Test numerically
difficult inputs, not only random, well-scaled tensors.
- Use `$kernel-microbenchmark` for benchmark construction, measurement, and
interpretation.
- Benchmark a sweep of `num_tokens` covering decode and representative prefill
workloads. Include relevant head counts and dimensions when they affect the
launch shape, and do not select an implementation or tuning heuristic from a
single setup.
- Include compilation or autotuning overhead when evaluating startup behavior;
report steady-state kernel performance separately.
4 changes: 4 additions & 0 deletions .agents/skills/triton-kernel-writing/agents/openai.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
interface:
display_name: "Triton Kernel Writing"
short_description: "Write robust, practical Triton kernels for vLLM."
default_prompt: "Use $triton-kernel-writing to implement or review this Triton kernel."
138 changes: 138 additions & 0 deletions .buildkite/release-pipeline.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -959,6 +959,144 @@ steps:
DOCKER_BUILDKIT: "1"
DOCKERHUB_USERNAME: "vllmbot"

# =============================================================================
# ROCk Release Pipeline (x86_64 only)
# =============================================================================
#
# Same shape as the ROCm pipeline above (Jobs 1 and 6 only), but built from
# docker/Dockerfile.rock_base and docker/Dockerfile.rock, which install the
# ROCm SDK from TheRock wheels instead of the ROCm apt repos.
#
# Note: ROCm SDK version is determined by ROCM_SDK_VERSION in
# docker/Dockerfile.rock_base
#
# =============================================================================

- group: "Build ROCk Image"
key: "build-rock-image"
depends_on: ~
steps:
# ROCk Job 1: Build ROCk Base Image (with ECR caching)
- label: ":rocm: Build ROCk Base Image"
id: build-rock-base-image
depends_on: ~
agents:
queue: cpu_queue_release
timeout_in_minutes: 360
commands:
- |
set -euo pipefail

# The SDK version, torch wheels and component branches all live in
# the Dockerfile, so its hash is the cache key.
CACHE_KEY=$$(sha256sum docker/Dockerfile.rock_base | cut -c1-16)
ECR_CACHE_TAG="public.ecr.aws/q9t5s3a7/vllm-release-repo:$${CACHE_KEY}-rock-base"

echo "========================================"
echo "ROCk Base Build Configuration"
echo "========================================"
echo " CACHE_KEY: $${CACHE_KEY}"
echo " ECR_CACHE_TAG: $${ECR_CACHE_TAG}"
echo "========================================"

# Login to ECR
aws ecr-public get-login-password --region us-east-1 | \
docker login --username AWS --password-stdin public.ecr.aws/q9t5s3a7

if docker manifest inspect "$${ECR_CACHE_TAG}" > /dev/null 2>&1; then
echo "ECR image cache HIT"
else
echo "CACHE MISS - Building from scratch..."

DOCKER_BUILDKIT=1 docker buildx build \
--file docker/Dockerfile.rock_base \
--tag "$${ECR_CACHE_TAG}" \
--build-arg USE_SCCACHE=1 \
--build-arg SCCACHE_BUCKET_NAME=vllm-build-sccache \
--build-arg SCCACHE_REGION_NAME=us-west-2 \
--build-arg SCCACHE_S3_NO_CREDENTIALS=0 \
--push \
.
fi

# Save ECR tag for downstream jobs
buildkite-agent meta-data set "rock-base-image-tag" "$${ECR_CACHE_TAG}"
env:
DOCKER_BUILDKIT: "1"

# ROCk Job 2: Build ROCk Release Docker Image
- label: ":docker: Build release image - x86_64 - ROCk"
id: build-rock-release-image
depends_on:
- step: block-build-release-images
allow_failure: true
- step: build-rock-base-image
allow_failure: false
agents:
queue: cpu_queue_release
timeout_in_minutes: 180
commands:
- |
set -euo pipefail

# Login to ECR
aws ecr-public get-login-password --region us-east-1 | \
docker login --username AWS --password-stdin public.ecr.aws/q9t5s3a7

ECR_IMAGE_TAG="$$(buildkite-agent meta-data get rock-base-image-tag 2>/dev/null || echo '')"
if [ -z "$${ECR_IMAGE_TAG}" ]; then
echo "ERROR: rock-base-image-tag metadata not found"
echo "This should have been set by the build-rock-base-image job"
exit 1
fi

docker pull "$${ECR_IMAGE_TAG}"

# Pass the base image ECR tag to downstream steps (nightly publish)
buildkite-agent meta-data set "rock-base-ecr-tag" "$${ECR_IMAGE_TAG}"

echo "========================================"
echo "Building vLLM ROCk release image with:"
echo " BASE_IMAGE: $${ECR_IMAGE_TAG}"
echo " BUILDKITE_COMMIT: $${BUILDKITE_COMMIT}"
echo "========================================"

DOCKER_BUILDKIT=1 docker build \
--build-arg max_jobs=16 \
--build-arg BASE_IMAGE="$${ECR_IMAGE_TAG}" \
--build-arg USE_SCCACHE=1 \
--build-arg SCCACHE_BUCKET_NAME=vllm-build-sccache \
--build-arg SCCACHE_REGION_NAME=us-west-2 \
--build-arg SCCACHE_S3_NO_CREDENTIALS=0 \
--build-arg INSTALL_LMCACHE=true \
--tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$${BUILDKITE_COMMIT}-rock \
--target vllm-openai \
--progress plain \
-f docker/Dockerfile.rock .

docker push public.ecr.aws/q9t5s3a7/vllm-release-repo:$${BUILDKITE_COMMIT}-rock
env:
DOCKER_BUILDKIT: "1"

- label: "Publish nightly ROCk image to DockerHub"
depends_on:
- build-rock-release-image
if: build.env("NIGHTLY") == "1"
agents:
queue: small_cpu_queue_release
commands:
- "bash .buildkite/scripts/push-nightly-builds-rocm.sh rock rocm100"
# Clean up old nightly builds (keep only last 14)
- "bash .buildkite/scripts/cleanup-nightly-builds.sh nightly-rocm100- vllm/vllm-openai-rocm"
- "bash .buildkite/scripts/cleanup-nightly-builds.sh base-nightly-rocm100- vllm/vllm-openai-rocm"
plugins:
- docker-login#v3.0.0:
username: vllmbot
password-env: DOCKERHUB_TOKEN
env:
DOCKER_BUILDKIT: "1"
DOCKERHUB_USERNAME: "vllmbot"

# =============================================================================
# Publish to DockerHub and PyPI (at the end so all builds complete first)
# =============================================================================
Expand Down
43 changes: 3 additions & 40 deletions .buildkite/scripts/ci-bake-rocm.sh
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ DEFAULT_REPO_SLUG="vllm-project/vllm"
DEFAULT_CI_HCL_SOURCE="docker/ci-rocm.hcl"
DEFAULT_CI_BASE_CONTENT_FILES=".dockerignore requirements/common.txt requirements/rocm.txt requirements/test/rocm.txt tools/install_torchcodec_rocm.sh rust-toolchain.toml tests/vllm_test_utils"
DEFAULT_CI_BASE_DOCKERFILE="docker/Dockerfile.rocm"
DEFAULT_CI_BASE_DOCKERFILE_STAGES="base rust_toolchain_input_0 rust-toolchain-input rust-toolchain build_nixl lmcache_source build_lmcache build_rocshmem build_deepep mooncake_source build_mooncake package_mooncake export_mooncake mori_base ci_base"
DEFAULT_CI_BASE_DOCKERFILE_STAGES="base rust_toolchain_input_0 rust-toolchain-input rust-toolchain build_nixl lmcache_source build_lmcache build_rocshmem build_deepep mori_base ci_base"
DEFAULT_CI_BASE_METADATA_VERSION="3"
# ROCm CI forces REMOTE_VLLM=0, so content identity covers only the selected
# local-source stages rather than unreachable remote-fetch alternatives.
Expand Down Expand Up @@ -1684,12 +1684,9 @@ ci_base_metadata_pairs() {
metadata_pair "vllm.rocm.deepep_commit" "${DEEPEP_BRANCH:-$(resolve_dockerfile_arg_value "${dockerfile}" "DEEPEP_BRANCH")}"
metadata_pair "vllm.rocm.deepep_nic" "$(resolve_dockerfile_arg_value "${dockerfile}" "DEEPEP_NIC")"
metadata_pair "vllm.rocm.deepep_rocm_arch" "$(resolve_dockerfile_arg_value "${dockerfile}" "DEEPEP_ROCM_ARCH")"
metadata_pair "vllm.rocm.mooncake_repo" "$(resolve_dockerfile_arg_value "${dockerfile}" "MOONCAKE_REPO")"
metadata_pair "vllm.rocm.mooncake_commit" "${MOONCAKE_BRANCH:-$(resolve_dockerfile_arg_value "${dockerfile}" "MOONCAKE_BRANCH")}"
metadata_pair "vllm.rocm.nixl_cache_key" "${NIXL_CACHE_KEY:-}"
metadata_pair "vllm.rocm.rocshmem_cache_key" "${ROCSHMEM_CACHE_KEY:-}"
metadata_pair "vllm.rocm.deepep_cache_key" "${DEEPEP_CACHE_KEY:-}"
metadata_pair "vllm.rocm.mooncake_cache_key" "${MOONCAKE_CACHE_KEY:-}"
}

write_ci_base_metadata_annotations() {
Expand Down Expand Up @@ -2271,7 +2268,7 @@ extract_dependency_pins() {
return 0
fi

for var in NIXL_BRANCH UCX_BRANCH ROCSHMEM_BRANCH DEEPEP_BRANCH MOONCAKE_BRANCH; do
for var in NIXL_BRANCH UCX_BRANCH ROCSHMEM_BRANCH DEEPEP_BRANCH; do
if [[ -n "${!var:-}" ]]; then
echo "Using provided ${var}: ${!var}"
continue
Expand All @@ -2295,19 +2292,16 @@ compute_dependency_cache_keys() {
local ucx_branch=""
local rocshmem_branch=""
local deepep_branch=""
local mooncake_branch=""
local nixl_material=""
local rocshmem_material=""
local deepep_material=""
local mooncake_material=""

bake_dir=$(dirname "${VLLM_BAKE_FILE}")
dockerfile_rocm="${bake_dir}/Dockerfile.rocm"
nixl_branch=$(resolve_dockerfile_arg_value "${dockerfile_rocm}" "NIXL_BRANCH")
ucx_branch=$(resolve_dockerfile_arg_value "${dockerfile_rocm}" "UCX_BRANCH")
rocshmem_branch=$(resolve_dockerfile_arg_value "${dockerfile_rocm}" "ROCSHMEM_BRANCH")
deepep_branch=$(resolve_dockerfile_arg_value "${dockerfile_rocm}" "DEEPEP_BRANCH")
mooncake_branch=$(resolve_dockerfile_arg_value "${dockerfile_rocm}" "MOONCAKE_BRANCH")

if [[ -n "${nixl_branch}" && -n "${ucx_branch}" ]]; then
nixl_material=$(compose_stage_cache_material "${dockerfile_rocm}" "base build_nixl")
Expand Down Expand Up @@ -2341,19 +2335,6 @@ compute_dependency_cache_keys() {
export DEEPEP_CACHE_KEY
echo "DeepEP dependency cache key: ${DEEPEP_CACHE_KEY}"
fi

if [[ -n "${mooncake_branch}" ]]; then
mooncake_material=$(compose_stage_cache_material \
"${dockerfile_rocm}" \
"base mooncake_source build_mooncake package_mooncake export_mooncake")
MOONCAKE_CACHE_KEY=$(
compose_dependency_cache_key \
"${mooncake_branch}" \
"${mooncake_material}"
)
export MOONCAKE_CACHE_KEY
echo "Mooncake dependency cache key: ${MOONCAKE_CACHE_KEY}"
fi
}

compose_stage_cache_material() {
Expand Down Expand Up @@ -2402,13 +2383,6 @@ dependency_cache_ref_for_target() {
printf '%s\n' "${cache_repo}:deepep-rocm-${DEEPEP_BRANCH}-rocshmem-${ROCSHMEM_BRANCH:-}"
fi
;;
mooncake-rocm-ci)
if [[ -n "${MOONCAKE_CACHE_KEY:-}" ]]; then
printf '%s\n' "${cache_repo}:mooncake-rocm-${MOONCAKE_CACHE_KEY}"
elif [[ -n "${MOONCAKE_BRANCH:-}" ]]; then
printf '%s\n' "${cache_repo}:mooncake-rocm-${MOONCAKE_BRANCH}"
fi
;;
esac
}

Expand All @@ -2426,14 +2400,13 @@ resolve_ci_base_dependency_targets() {
local nixl_ref=""
local rocshmem_ref=""
local deepep_ref=""
local mooncake_ref=""

[[ "${TARGET}" == "ci-base-rocm-ci-with-deps" ]] || return 0

case "${mode}" in
always)
echo "ROCM_DEP_CACHE_EXPORT_MODE=always; exporting all dependency caches serially"
for target in nixl-rocm-ci rocshmem-rocm-ci deepep-rocm-ci mooncake-rocm-ci; do
for target in nixl-rocm-ci rocshmem-rocm-ci deepep-rocm-ci; do
if [[ -n "$(dependency_cache_ref_for_target "${target}")" ]]; then
add_dependency_cache_target "${target}"
fi
Expand Down Expand Up @@ -2483,16 +2456,6 @@ resolve_ci_base_dependency_targets() {
fi
fi

if [[ "${mode}" != "always" && -n "${MOONCAKE_CACHE_KEY:-}" ]]; then
mooncake_ref=$(dependency_cache_ref_for_target "mooncake-rocm-ci")
if dependency_cache_ref_exists "${mooncake_ref}"; then
echo "Mooncake dependency cache exists: ${mooncake_ref}"
else
echo "Mooncake dependency cache missing; will seed: ${mooncake_ref}"
add_dependency_cache_target "mooncake-rocm-ci"
fi
fi

# DeepEP inherits from ROCShmem. If ROCShmem is being seeded, seed DeepEP too
# so the pair stays consistent for future ci_base rebuilds.
if printf '%s\n' "${DEPENDENCY_CACHE_TARGETS[@]}" | grep -qx "rocshmem-rocm-ci" \
Expand Down
4 changes: 3 additions & 1 deletion .buildkite/scripts/cleanup-nightly-builds.sh
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ set -ex

# Clean up old nightly builds from DockerHub, keeping only the last 14 builds
# This script uses DockerHub API to list and delete old tags with specified prefix
# Tags starting with "nightly-dev" are always excluded from deletion
# Usage: cleanup-nightly-builds.sh [TAG_PREFIX] [REPO]
# Example: cleanup-nightly-builds.sh "nightly-"
# Example: cleanup-nightly-builds.sh "cu130-nightly-"
Expand Down Expand Up @@ -55,7 +56,8 @@ get_all_tags() {
set -x

# Get both last_updated timestamp and tag name, separated by |
local tags=$(echo "$response" | jq -r --arg prefix "$TAG_PREFIX" '.results[] | select(.name | startswith($prefix)) | "\(.last_updated)|\(.name)"')
# Exclude nightly-dev tags from cleanup
local tags=$(echo "$response" | jq -r --arg prefix "$TAG_PREFIX" '.results[] | select(.name | startswith($prefix)) | select(.name | startswith("nightly-dev") | not) | "\(.last_updated)|\(.name)"')

if [ -z "$tags" ]; then
break
Expand Down
Loading