[Klaud Cold] MI325X Qwen3.8-27B BF16 MTP, synthetic AL 2.51 / MI325X Qwen3.8-27B BF16 MTP,合成 AL 2.51 - #3308
[Klaud Cold] MI325X Qwen3.8-27B BF16 MTP, synthetic AL 2.51 / MI325X Qwen3.8-27B BF16 MTP,合成 AL 2.51#3308Oseltamivir wants to merge 3 commits into
Conversation
新增 MI325X Qwen3.8-27B BF16 原生 MTP 配方,以相同镜像和 1k1k 并发范围重跑 #3295 的工作负载。
补充 BF16 原生 MTP 配方对应的 PR 链接。
|
Thanks for the contribution!
中文感谢你的贡献!
|
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=35495280070 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=35495280070 |
Qwen3.8-27B BF16 吞吐测试启用 thinking,使用 3 个原生 MTP 草稿 token 和合成 AL 2.51;准确率评测保留真实验证,并同步中英文说明。
|
InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest |
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
Goal: Rerun MI325X Qwen3.8-27B BF16 with thinking on, synthetic AL 2.51 and 3 native MTP draft tokens.
Baseline recipe: #3295 ·
vllm/vllm-openai-rocm:nightly-a8d1aa9c99b8698a2a78b611b7a10c30e6b3995b1k/1k · TP1 · concurrency 1, 2, 4, 8, 16, 32, 64, 128.
qwen3.827b-bf16-mi325x-vllm-mtpusesQwen/Qwen3.8-27B, BF16 target/native MTP weights and BF16 KV cache. The image, runner, graph mode, prefix-cache setting and concurrency-sized scheduler follow the source recipe.Throughput sets
rejection_sample_method: syntheticandsynthetic_acceptance_length: 2.51, using the BF16thinking_on[3]measurement in #3304. RequireTHINKING_MODE=thinking_on; explicitly enable thinking in server chat-template defaults. The fixed-sequence client retains--use-chat-template, whose checkpoint default renders a thinking-on prompt. Runs requesting accuracy throughEVAL_ONLYorRUN_EVALretain real verification.Validation: Bash syntax, whitespace, append-only changelog and full matrix checks pass (8 throughput points; the current 1k/1k policy selects no default evals). All four actual recipes reject thinking-off before GPU startup. Rendering the checkpoint template confirms its default equals explicit thinking-on. The previous real-verification sweep passed. Synthetic sweep dispatched: run 35495280070 on
452108aed28ef506babacb15ff92f893fa38a178.AI model disclosure
GPT-6 prepared the BF16 and synthetic-acceptance changes and validation. The exact runtime model/version identifier was not exposed and could not be verified. The source FP8 recipe was credited to Claude Code; its underlying model/version was not disclosed and could not be verified.
中文
**目标:**重跑 MI325X Qwen3.8-27B BF16,使用 thinking 开启、合成 AL 2.51 和 3 个原生 MTP 草稿 token。
基线配方:#3295 ·
vllm/vllm-openai-rocm:nightly-a8d1aa9c99b8698a2a78b611b7a10c30e6b3995b1k/1k · TP1 · 并发 1、2、4、8、16、32、64、128。
qwen3.827b-bf16-mi325x-vllm-mtp使用Qwen/Qwen3.8-27B,目标权重、原生 MTP 权重及 KV cache 均为 BF16。沿用原配方的镜像、runner、graph 模式、prefix-cache 设置和按并发配置的 scheduler。吞吐测试设置
rejection_sample_method: synthetic和synthetic_acceptance_length: 2.51,采用 #3304 中 BF16 的thinking_on[3]测量值。要求THINKING_MODE=thinking_on,服务端默认 chat-template 参数显式启用 thinking。固定序列长度客户端保留--use-chat-template,checkpoint 默认模板会生成 thinking 开启的提示词。通过EVAL_ONLY或RUN_EVAL请求准确率评测时保留真实验证。验证:Bash 语法、空白、changelog 只追加约束及完整矩阵检查均通过(8 个吞吐测试点;当前 1k/1k 策略默认不选取 eval)。实际执行四个配方,均在 GPU 启动前拒绝 thinking 关闭。实际渲染 checkpoint 模板,确认默认行为与显式开启 thinking 相同。此前的真实验证 sweep 已通过。合成 AL sweep 已启动:run 35495280070,提交为
452108aed28ef506babacb15ff92f893fa38a178。**AI 模型披露:**GPT-6 完成 BF16、合成接受长度修改与验证;运行时未提供精确模型/版本标识,无法核实。原 FP8 配方署名 Claude Code,未披露底层模型/版本,无法核实。