Skip to content

Attention daily report — 2026-05-10 - #2

Merged
ly865623 merged 1 commit into
mainfrom
claude/determined-euler-8UV1P
May 11, 2026
Merged

Attention daily report — 2026-05-10#2
ly865623 merged 1 commit into
mainfrom
claude/determined-euler-8UV1P

Conversation

@ly865623

Copy link
Copy Markdown
Owner

Summary

新增 2026-05-10 的 attention 技术日报,覆盖过去 24 小时(2026-05-09T01:05Z — 2026-05-10T01:05Z)vllm-project/vllm、sgl-project/sglang、NVIDIA/TensorRT-LLM 三个仓库的 attention / KV cache 相关进展。

Highlights

  • TensorRT-LLM 合并了 DeepSeekV4 attention kernel(#13652)和 Cute-DSL FP8 Paged MQA decode kernel(#13219)。
  • SGLang 合并了 #24097:将 fa_skip_kv_cache 限制为非 MLA 后端,修复 MLA 模型 KV 被错误跳过的问题。
  • vLLM 有两份关键 RFC 在讨论:KV cache 布局标准化(#42082)和 V1 调度器缓存亲和性排序(#42185)。
  • 横向趋势:三家框架今日均聚焦 DeepSeek V4 MLA attention 支持;TensorRT-LLM 在 Blackwell 上 FP8/FP4 attention kernel 持续推进。

File

  • reports/2026-05-10-attention-daily.md

Test plan

  • Markdown 链接全部指向真实 GitHub PR/issue/blog URL
  • 报告控制在 800-1500 字范围内
  • 三仓库各自有独立 section,无相关动态处明确标注

https://claude.ai/code/session_01J2bouHEwcLEicoUh4VjQiW


Generated by Claude Code

Covers vLLM, SGLang, TensorRT-LLM attention-related activity in the
past 24 hours, including TensorRT-LLM's merged DeepSeekV4 attention
kernels and FP8 MQA decode kernel, SGLang's MLA KV-skip fix, and
notable open PRs/issues across all three frameworks.

https://claude.ai/code/session_01J2bouHEwcLEicoUh4VjQiW
@ly865623
ly865623 merged commit 2d30574 into main May 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants