Pin flip to vLLM 0.29.0: the series re-exported at fuzz 0, KVarN on 0.29, the #114 moves, a cold-boot profiling fix, and 0.28 vs 0.29 on every run setting (#106 part two) - #148
Conversation
… qwen38/0.29 at 337efb79f (35 patches, 0 fuzz, replay identical to the fork tip); KVarN 0.29 patches and modules; retirements listed in PATCHES.md
…uild, docs (carried from port-0.29-bae2023 onto the split README)
…ds them Matches syv-ai#119's shape on the 0.28 line. VLLM_SPEC_DECODE_ATTN, VLLM_SPEC_DECODE_ATTN_QMAX and VLLM_SPEC_ATTN_BLOCK_M move out of the speed-knobs-envs catch-all into spec-decode-attn; VLLM_SPEC_ATTN_DEBUG stays with triton-spec-attn-fp8-kv and is re-anchored beside BLOCK_M's new location, which is what the rebase conflicted on. Done in the fork, not by editing patch files: new branch qwen38/0.29-hq off 337efb79f with the registrations moved between topic commits, then spec-decode-attn, speed-knobs-envs and triton-spec-attn-fp8-kv re-exported. A new branch rather than a rewrite of qwen38/0.29, so no force-push. The envs.py hunk in spec-decode-attn anchors on VLLM_DP_MASTER_PORT - present in pristine v0.29.0 and untouched by any other patch - the same region basecamp anchored the 0.28 moves on, so pass 2's standalone-apply contract is unaffected. Verified, this being a relocation and not a change: fork 38 commits before and after, 0 subjects lost, HEAD attached, branch ref == HEAD envs registry 301 entries before and after, 0 differing lines (same set, same defaults) replay onto pristine v0.29.0: 35 series + 2 kvarn at --fuzz 0, 0 failed, 2 at an offset with exact context, 0 stray .orig/.rej replayed tree IDENTICAL to the new fork tip 03b5c4259 check_vllm_series.sh both passes green on Linux, exit 0, patch integrity: OK
…nflate the profiled transient peak and refuse the KV cache the warm boot grants (cold first boots failed on B and E on a WSL2 4090 and on a native 3090; series 36 + kvarn 2, replay identical to fork tip bf29fa109)
…L2 4090), and the cold-boot falsifier on both boxes (default profile fixed; CTX=huge cold boot still refused on the WSL2 4090, open)
… had no row Two defects found while vetting the rebased 0.29, both mine. The header named one export point (`qwen38/0.29`), but three rows were re-cut on `qwen38/0.29-hq` for the syv-ai#114 registration moves -- so a reader regenerating spec-decode-attn, speed-knobs-envs or triton-spec-attn-fp8-kv from the named branch would get the pre-syv-ai#114 hunks and a --fuzz 0 failure they could not explain from this file. Both branches are now named with their commit ids, and the three re-cut rows are called out by name. Four entries in patches/series had zero mentions in the table: engine-completion-log, engine-stall-sentinel, topk-honour-flashinfer-sampler-switch, triton-spec-attn-fp8-kv. The table is the only place a patch's "retires when" is written down, so an unlisted patch is one nobody knows the exit condition for. Accounting, stated so the next drift is visible: 36 entries in patches/series, 36 series rows, plus 2 kvarn rows that install.sh applies outside patches/series -- 38 rows total.
…oved to bf29fa109 Adds the missing row for the cold-boot fix and repoints the header: qwen38/0.29-hq is now bf29fa109, and it carries four rows, not three -- the three re-cut for the syv-ai#114 registration moves plus memory-profile-after-warmup, which was cut there after that branch had already diverged. Accounting: 37 entries in patches/series, 37 series rows, plus 2 kvarn rows that install.sh applies outside patches/series -- 39 rows total. The commit that added the patch describes the series as 36; it is 37.
The 0.29 port carried four claims written against 0.28.0 without re-checking
them. Three are still true and now say so against the pin actually in use;
the fourth is not re-verified and now says that instead of implying currency.
Verified against v0.29.0 in a checkout, not from memory:
- gotcha 18: DFlashModelTypes is still inside EagleModelTypes, so dflash keeps
async scheduling on (speculative.py:69 on 0.29.0, :67 on 0.28.0 -- the cited
line had drifted by two).
- gotcha 19: async_scheduling still resolves to True in the else branch; one
"= True" and five "= False" assignments in config/vllm.py at both tags.
- the offload residency note: still no gauge; kv_offload is absent from
v1/metrics/loggers.py at both tags.
Not verified: the FlashInfer k=4 illegal-memory-access paragraph in
single-user/README.md was measured on 0.28.0 and has not been re-run on 0.29.0.
Marked unproven either way rather than restated, because the pin moved under it.
Also repointed two capacity-planning TODOs that still said "re-run after the
v0.28.0 upgrade" and were two pins behind.
…sses; the open item narrows to the WSL2 4090's first boot
…e cold default boot on the native 3090; the WSL2 4090 CTX=huge residual is the graph estimate), same hunks, fork tip 512a9699c; docs: warm rows for both boxes, the cache cut that locates the residual, the profiling breakdown, and why an uncounted estimate pass is not the fix
…: the 0.29 row of the field-tests table now names this branch's pin
…CTX=huge path on the WSL2 4090 (cold 4.47 GiB, warm 4.68, actual graph pool 0.0), the knob in the compile-cache key, the equivalent-utilization line scoped by profile
…ctual graph pool reads 0.0 GiB; the native 3090 reads 0.11 to 1.87 GiB by profile, so it under-reserves there
…he allocator's reserved bytes (WSL2's driver reading is pinned at zero after KV allocation and collapses by 5.44 GiB during a cold KVarN compile, which the 0.29 estimate subtracted from the KV budget); memory-profile-after-warmup header corrected (0.28 refuses the native default boot in both cache states); docs: the allocator probe, both fixes' rows, the 0.28 frame
…th both arms, three 0.28 refusals with numbers), the measured driver-to-allocator residue (0.02 and 0.03 GiB), the KV the allocator reading returns, the estimator not made accurate, the CTX=huge cold-to-warm gap open on both boxes; patch header carries the residue
…syv-ai#142 and the four new patches hybrid-sw-block-promote re-exported with syv-ai#142's divisor condition (a whole multiple of the layer's kernel block); bench-probe-errors, serve-404-served-names, serve-model-path-match and tokenize-v1-route applied as-is to the 0.29 fork and re-exported (line offsets only). All five from cpuchip/vllm qwen38/0.29-hq2 at fea76cc82, a new branch so qwen38/0.29-hq stays unrewritten. PATCHES.md: the four rows cut against 0.29.0, the header names both fork points and every row cut from them (memory-profile-after-warmup at c06f8ef11, the file's own export point, and cudagraph-memory-from-allocator, which the header had left out). Series plus KVarN applied to v0.29.0 at --fuzz 0 reproduce fea76cc82's vllm/ tree (0 differing files).
…hat its fork commit message carried The export kept the message's prose and dropped only the header lines of the pasted diff, so 223 hunk-body lines sat in the file's preamble (GNU patch skipped them; a reader and the PR gate did not). The fork commit is reworded on qwen38/0.29-hq2 (now at da6a87935, trees unchanged), and the four new topics are re-exported from their moved commits. Hunks unchanged; series plus KVarN on v0.29.0 at --fuzz 0 reproduce da6a87935's vllm/ tree (0 differing files).
…rted to 0.29) into main main last took syv-ai main at 8d09ec5 and the first port-0.29 branch; the port was re-derived for the PR on top of syv-ai#119's shape, so both lines carried the same files ported twice. The PR line is a superset for every shared file (checked file by file: the same detection with --fuzz 0, the registry check with the seat name reworded out, the retirements listed in PATCHES.md), so it wins there. docker-compose.yml and the image workflow keep main's identity (ghcr.io/cpuchip/qwen38-27b-rtx3090 built from Dockerfile.fork) on top of syv-ai's GPU_COUNT and pull-request build changes; cache-to takes the repository through format().
…CHES.md: dflash2-backport is the one file with no export point, and qwen38/0.29-hq2 is a rewrite of -hq from hybrid-sw-block-promote on
…efix-cache-retention-interval vllm serve on 0.29 reads the deprecated VLLM_PREFIX_CACHE_RETENTION_INTERVAL in the field's default factory, but the flag's unset default is passed explicitly by from_cli_args and wins, so the retention-interval-unset branch forces hybrid + EAGLE to dense (vllm/engine/arg_utils.py at v0.29.0). Measured on a native 3090, CTX=huge SPEC=dflash2 PREFIX_CACHE=1: the export logs the deprecation and boots dense, and an exported 13057 boots; the flag at 13056 boots with no dense line, and the flag at 13057 is refused (not a multiple of 2176). The block sizes on 0.29 match 0.28's (2176 at 7 drafts, 2432 at 15). An exported value is carried over as the flag; the flag in EXTRA_ARGS wins over both.
…apture pinned to the V1 runner gptq_lm_head.py assumed the base model's layout: seven model-0000x shards, a .bak from quant_lm_head.py, a quantization_config.json to copy, a model_extra_tensors.safetensors to link. A checkpoint that went through prepare/quant_heads_stream.py has model.safetensors + model-mtp.safetensors with .bak-orig backups and no quantization_config.json, so the script stopped at its first line (no .bak), and past that the model-0000* glob would have left model-mtp.safetensors, which the index points to, out of the variant. It now reads the bf16 lm_head from .bak or .bak-orig, hardlinks every weight shard the source has, copies only the files that exist, and copies model_extra_tensors.safetensors instead of linking it: build_draft_vocab.py rewrites that file in the variant with save_file, and safetensors 0.4.5 through 0.7.0 write it in place, through the hardlink into the source dir (measured; 0.8.0 replaces the file and leaves the source alone). capture.py hooks the V1 runner (vllm.v1.worker.gpu_model_runner). On the pinned 0.28.0 that is the runner this model gets anyway -- a hybrid architecture outside DEFAULT_V2_MODEL_RUNNER_ARCHITECTURES stays on V1 -- so VLLM_USE_V2_MODEL_RUNNER=0 changes nothing today. vLLM 0.29.0 defaults every model to V2, and there the hooks never fire: the capture finishes with rows=0 and the GPTQ that follows calibrates on zeros (KL 0.00000, round-trip error 1.0000, an lm_head of zeros), measured on the 0.29 port (syv-ai#148). No speculation happens in a capture, so V1 is the right runner on both. It also sets FLASHINFER_DISABLE_VERSION_CHECK=1, as both launchers do, since it runs standalone. Measured on Ar4ikov/Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM after quant_heads_stream.py, 600k teacher-forced UltraChat tokens captured in 10 min on a 3090 (vLLM 0.29.0): RTN int4 lm_head KL 0.00701, GPTQ int4 0.00239 (the base model's published figures: 0.0068 and 0.0029). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y checkpoints The llm-compressor AWQ exports of the uncensored finetune and of the base model, int4 asymmetric g128 with zero points and the vision tower kept, after quant_heads_stream.py and build_draft_vocab.py; the -fast siblings with the int4-GPTQ lm_head (building one from a single-shard export takes the drafter/ fixes in syv-ai#181). The measured rows (vLLM 0.29.0, the syv-ai#148 port with this patch) for MTP, DFlash2, CTX=long, the production line, the fast variant and batch mode, whose shipped 0.972 / 150k did not boot with the tower on (syv-ai#182). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y checkpoints The llm-compressor AWQ exports of the uncensored finetune and of the base model, int4 asymmetric g128 with zero points and the vision tower kept, after quant_heads_stream.py and build_draft_vocab.py; the -fast siblings with the int4-GPTQ lm_head (building one from a single-shard export takes the drafter/ fixes in syv-ai#181). The measured rows (vLLM 0.29.0, the syv-ai#148 port with this patch) for MTP, DFlash2, CTX=long, the production line, the fast variant and batch mode, whose shipped 0.972 / 150k did not boot with the tower on (syv-ai#182). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… on the first reuse after a cold turn, then 99.3-99.4%, against 0% dense) Native 3090, one image, CTX=huge SPEC=dflash2 PREFIX_CACHE=1, syv-ai#174's alternating pair at 32,600 tokens a side, T=0, max_tokens 24, n=1 per arm. The first reuse lands on the last retained snapshot (26,112 = 2 x 13056), so a reader checking that turn against the 0.28 range would read a regression.
|
Ran this on the reference 3090 tonight — a third box, native Linux, 250 W — with the branch head exactly as pushed: pristine Build and apply: clean. 37 patches at exact context, 0 offset, 0 fuzz, on the box's pristine 0.29.0; KVarN installs; Most of your claims reproduce:
One regression you did not report, and it is on the production line. Greedy single-stream acceptance drops on DFlash2, deterministically (greedy tok/step reproduces to two decimals across runs): It is C1-specific — at C4/C8 greedy the two versions accept the same (3.51-3.55 both) — and DFlash2-specific, since MTP goes the other way. The sharper half, 0.29 at 7 drafts on the same production line: On 0.28, 15 drafts is the +9% setting production runs (LABD, the lookup lane filling the verify tail from context). On 0.29 it is now worse than 7 at greedy. That is also the explanation for your own table's oddity — B and C identical to the decimal on both your boxes: the long verify block has stopped paying on 0.29, and at greedy it costs. My guess is the lookup-drafting / adaptive-length patches as re-exported onto 0.29's native DFlash2 (vllm#52816 changed that path), but I have not bisected it and it is only a guess. The logged per-position acceptance cannot show it (both versions log zeros past position 7), so the harness's So this does not merge tonight, for two reasons, neither of them the port's core:
One upgrade hazard to document with the flip. 0.29 ships FlashInfer 0.6.18, which invalidates every cached FlashInfer kernel. On a native install whose system The port itself is in good shape — clean apply, the memory fixes doing exactly what the writeup says, and #64 gone on two machines. Rebase plus an answer on the 15-draft greedy drop, and this is ready. |
|
A second blocker, found through #182: batch mode's shipped defaults do not boot on this branch on a 3090. Reference box, reference checkpoint,
Same checkpoint, same settings, 0.29 gives the KV pool 7.63 GiB where 0.28 gave 6.09. That is your fourth and fifth commits doing what the writeup says — "what the correction returns is KV on every boot" — and at The fix is yours to choose — a lower batch |
PATCHES.md is the one conflict: the 0.29 table kept, syv-ai#172's marlin-int8-asym-zp row added. The 0.28-cut patch applies to 0.29.0 plus this series at exact context (check_vllm_series.sh: 42 at exact context, 0 offset, 0 fuzz).
|
Correction: the greedy DFlash2 "regression" I reported is not one. Please don't spend time on it. It came from my measurement, and I owe you the retraction with the evidence. I went to bisect it this morning and found nothing to bisect. On the rebuilt 0.29 tree (your head plus today's The lookup lane is identical on both versions. A verbatim-copy prompt, the lane's best case, 15 drafts, greedy: 15.45 tok/step on 0.28 and 15.45 on 0.29, to the step. Greedy chat prompts with thinking off: 2.96 vs 2.98. The drop was trajectory divergence on 8 prompts. Harness-shaped requests (chat template, thinking on, 1024 tokens, greedy), sent one at a time, with tok/step and a hash of the output recorded per prompt: 0.28 reproduces itself exactly, text and tok/step. 0.29 writes different text on every prompt, because the numerics moved. Per prompt, that swings acceptance ±25% in both directions, and in aggregate it nets to 3.43 vs 3.41, a 0.6% difference. With the text diverging, the per-prompt deltas have a standard deviation of ~0.55 tok/step, so an 8-prompt greedy cohort's standard error is ~0.19 tok/step, about 6%. My −5.6% was inside one standard error, and so is anything in your campaign table's greedy rows below ~±11%. The same caveat applies to my "15 drafts now below 7" line: that compared different texts too. The batch blocker stands. That one is a hard OOM at the shipped Branch housekeeping, already done: I merged today's |
|
Thanks for the retraction; the per-prompt output hash is the right instrument, and it covers our side too. Our native 3090 read greedy DFlash2 at 15 drafts up 9% (3.21 to 3.50 tok/step, 8 prompts, seed-0 greedy), which sits inside the same ±11% band, so we're not claiming it either. I'll put that caveat on the greedy line in the PR body. On the batch blocker, the ladder is already measured on two native 3090s, which may save you the run. Shipped batch settings (
On the rented box the image's own 0.28 tree at 0.972 boots with 192,525 tokens, your number exactly, and at 0.95 it gets 203,350. Warm throughput at 0.95 is level: 207.5 tok/s on 0.28, 210.4 on 0.29. My proposal is 0.95 for fp8 batch on the 0.29 line: 0.96 boots but has only 468 MiB spare at the peak, and a card that also drives a display loses more than that before vLLM starts. At 0.95 the pool (225,000) is still 17% above what 0.28 gets at 0.972 (192,525). Not measured: The other two batch modes ship at 0.93, and both boot and serve on 0.29 at that value: On the branch: the #172 merge is already staged on our side ( One more thing from your first comment: |
…ody); edit_body.sh chains the edit on the gate, DRY=1 for tests Found 2026-09-23: `pr_gate.py --pr N --body f` read the body from GitHub and ignored the file, so every run meant to gate a new body checked the old one. edit_body.sh gates the PR's live head with the new body and edits only on exit 0; a falsification run of the first version (before this fix) was a live edit and overwrote syv-ai#148's body for about a minute, so the script now says to test with DRY=1 only. Falsified offline: an em-dash body fails, the real body passes.
…i#182) 0.29 with this port's two memory fixes stops over-reserving ~1.5 GiB, and at 0.972 that was the headroom batch mode's unprofiled warmup transients used: 0.972 OOMs in warmup on a 3090 with the tower on and off. Reference-3090 ladder, boot plus 128 requests at C64: 0.93-0.96 all pass with VISION=0/1, 0.972 fails. 0.95 keeps a step below the highest pass, boots cold to the warm pool, and holds at least the pool 0.28 had at 0.972. kvarn and int4pth are unchanged at 0.93. docs/vllm-0.29.md gets the ladder, and a section on why greedy tok/step cannot be compared across the two versions below ~±11%.
|
Both blockers are resolved, so I'm merging this. What I pushed to the branch, and why: 1. Merge of today's 2. Batch
0.95 keeps a full step below the highest pass. It lands on the same pool cold as warm, and it holds at least the pool 0.28 had at 0.972 (220,360 / 192,525), so nobody loses capacity in the upgrade. 3. The greedy "regression" is retracted (details in my previous comment). It was cross-version text divergence on 8 prompts; the per-prompt aggregate is 3.43 vs 3.41, and the lookup lane measures 15.45 tok/step on both versions. CI on Thank you for this port. The memory work in particular — measuring the second profiling pass, and reading the graph pool from the allocator instead of the driver — is what made the batch fix a one-line default rather than a workaround. Closing on merge: #64 (clean on 0.29 on two machines), #66 (the enum flag is in both launchers), #106, #114 (both knobs read through |
…yv-ai#136) into main Conflicts: docs/vllm-0.29.md takes syv's (a superset: the maintainer's greedy-divergence section); the image workflow keeps this fork's identity (Dockerfile.fork into ghcr.io/cpuchip/...); PATCHES.md keeps this fork's header and marlin-int8-asym-zp row, which describe the fork-exported file this main carries (hunks identical to syv's hand-cut file; syv's series replays to fork tip 291980422 with 0 differing files). Batch GPU_UTIL 0.95 comes in from syv-ai#148.
Takes upstream's 0.29.0 series wholesale (patches/series, PATCHES.md, the re-cut patches, KVarN 0.29.0), including the four patches syv-ai#148 retired (int4-mq3d-envs, sse-keep-alive, vllm-pr54282-draft-gumbel-salt, xgrammar-spec-terminated). Drops the fork's auth-deny-default.patch from the series and the tree: it is cut against 0.28.0, both of its target files moved in 0.29.0 (serve/utils/server_utils.py -> serve/middleware/authenticate.py, openai/cli_args.py -> launchers/cli_args.py), so it cannot apply at fuzz 0. It returns with the 0.29 port in syv-ai#169. Keeps from the fork: manual-only image builds and the fork's own buildcache ref (docker-image.yml), the guarded resolver source in the bench scripts, the F12 digest-pinned base, and F04 copy-only variant writes in drafter/gptq_lm_head.py (upstream's syv-ai#181 file handling, copy semantics). verify.sh equals upstream: every fork change to it landed via syv-ai#158/syv-ai#171/syv-ai#172.
…varn, Dockerfile, recipes, docs); adopted auth-deny-default - vllm==0.29.0 + huggingface_hub==1.28.0; Dockerfile/requirements claims re-based; Python 3.14 and our layout kept (upstream runs 3.12) - patches/: the 0.29 re-export verbatim (PR syv-ai#148), 43 files, all --fuzz 0. Retired by 0.29 itself: vllm-pr54282-draft-gumbel-salt and xgrammar-spec-terminated (in 0.29.0), int4-mq3d-envs (absorbed into speed-knobs-envs; VLLM_INT4_MQ_3D stays, the int4 recipes keep working), the graph-memory-reserve hunk of hybrid-kv-groups-v2- cudagraph and the padded-page-view hunk of int4-kv-per-token-head (0.29 covers both natively). VLLM_V2_CUDAGRAPH_MEM_MIB is dead (0.29 profiles graph memory itself), the six exporting recipes drop it - new patches: memory-profile-after-warmup + cudagraph-memory-from- allocator (0.29's cold-cache profiling refused boots; the pair repairs it), topk-honour-flashinfer-sampler-switch (0.29 ships FlashInfer 0.6.18, whose top-k JIT needs nvcc >= 13; the recipes' VLLM_USE_FLASHINFER_SAMPLER=0 now covers the drafter's top-k), compile-key-runtime-knobs, engine-completion-log, engine-stall- sentinel (off by default), serve-404-served-names, serve-model-path- match, tokenize-v1-route, marlin-int8-asym-zp (syv-ai#172), auth-deny- default (syv-ai#169, adopted by user decision) - kvarn trio re-ported to 0.29 (kvarn-0.29.0, kvarn-files-0.29.0 generated from upstream kvarn/files/, kvarn-v2-runner-0.29.0), last in the series in upstream's install.sh order - recipes: PYTORCH_CUDA_ALLOC_CONF defaults expandable_segments:False at TP=2 (custom all-reduce cannot export a VMM graph buffer over CUDA IPC, syv-ai#163/syv-ai#176); w4a16-k4v2-dflash2 gains --prefix-cache- retention-interval 13056 (6 x 2176, the attention block at 7 drafts; upstream syv-ai#174) - README/AGENTS re-based to 0.29.0; upstream's 0.29 measurements (their boxes) carried in the Notes, labeled as upstream measurements; bare-metal note that 0.29's FlashInfer JIT needs nvcc >= 13 (CUDA_HOME at the venv's nvidia/cu13; the Docker image carries CUDA 13) - not adopted (harness / dropped lanes): dflash2-backport, offload- mtp-serve, offload-wsl2-devptr, bench-probe-errors, triton-spec- attn-fp8-kv, upstream's _check_applied.py / check_vllm_series.sh tooling Validated: docker build on Linux (aarch64) applies all 43 in series order at --fuzz 0 and compileall-gates; in-image patch_vllm.py audit reports 43 of 43 in place on vllm 0.29.0.
…rlay without vLLM in the FA-scratch fixture bench-policy.yml asserts the F10 guard in patches/check_vllm_series.sh (refuse this repo, and any tree without .qwen-disposable-series-target unless ALLOW_TREE_RESET=1) and patch-integrity.yml still stamps the sentinel, but the guard never reached main: it lived on validity/build-integrity (0ebd561), and main carried upstream's unguarded script. Ported onto the current script. bench/test_fa_scratch_capacity.py loaded kvarn/.../config.py as a stdlib-only module. Since the 0.29 pin flip (858c3b7, syv-ai#148/syv-ai#114) the overlay reads its knobs through vllm.envs, so the fixture died with ModuleNotFoundError: vllm under the CI python, and with AttributeError KVARN_FA_SCRATCH_CAP in a venv without KVarN installed. It now loads the overlay against a stand-in vllm.envs carrying that one knob with the patch's semantics (raw string, None when unset). bench/test_fa_scratch_capacity.py python3 / venv: RESULT PASS bench-policy mocks policy mocks: OK check_vllm_series.sh . REFUSING (this repo), exit 1 check_vllm_series.sh <no sentinel> REFUSING, exit 1
The pin flip to vLLM 0.29.0, the second half of #106, on top of #119's shape. On top of main: the pin and series, the carried launcher and docs, the #114 moves, the profiling fix, the allocator-measured graph pool, the campaign and falsifier doc, the
PATCHES.mdrows for the four newest entries with the header naming both fork export points, the docs that followed, a merge of main's field-tests commit (#146) with the README's 0.29 row rewritten for this branch, and a merge of main at 13b30ea: #142's divisor condition and the four patches from #165-#168, ported to the 0.29 fork and re-exported.vllm==0.29.0indocker/requirements.txt, and the series re-exported from the fork, one commit per row on v0.29.0 (cpuchip/vllm: 31 rows fromqwen38/0.29, four fromqwen38/0.29-hq, six fromqwen38/0.29-hq2, whose tip is the tree the series reproduces;PATCHES.mdnames each row's export point): 42 series entries (41 applied,dflash2-backportkept and skipped by name since DFlash2 went native in 0.28.0) plus the two KVarN files, every one applied and checked at--fuzz 0, replay onto pristine v0.29.0 identical to the fork tip (0 differing files) on a WSL2 4090 and on a native 3090 independently, both re-run against the fork tip after the last re-cut (the WSL2 replay against the current tip da6a87935, after the merge of main at 13b30ea; the native replay against fea76cc82, whose tree da6a87935 shares, likewise 0 differing files;patches/check_vllm_series.shpasses 1 and 2 green on the native box at 62fd3d4 and 652f99a) (the native replay had first been verified at an earlier tip;patches/check_vllm_series.shresets the tree when it finishes, so the identity diff has to follow a replay of one's own, not the check). Three patches retire because 0.29.0 carries them (vllm-pr54282-draft-gumbel-salt,xgrammar-spec-terminated,sse-keep-alive); two hunks retire because 0.29 does the work itself (the graph-memory reserve inhybrid-kv-groups-v2-cudagraph, the int4 padded-page view).PATCHES.mdhas the row for each and names the retirements.kvarn/kvarn-0.29.0.patch,kvarn/kvarn-v2-runner-0.29.0.patch, the copied modules): 0.29 stopped asking a backend for its KV cache shape and stride, so the backend declares its layout and folds the runner's view into tiles with aviewthat fails on a wrong layout. Same block geometry and the same pool as 0.28 on the reference profile; theKVARN_*knobs are registered under their names, theVLLM_KVARN_*rename stays a follow-up (a decision to reject by name).VLLM_SPEC_DECODE_ATTN,VLLM_SPEC_DECODE_ATTN_QMAXandVLLM_SPEC_ATTN_BLOCK_Mregistered inspec-decode-attn.patch, the patch that reads them, anchored on theVLLM_DP_MASTER_PORTregion no other patch touches; the registry has the same 301 entries before and after.and 0.28 did not(memory-profile-after-warmup.patch): 0.29'sdetermine_available_memoryruns the profiling pass inside the memory-profiling window, so on a cold compile cache the compiler's scratch is read as the model's transient peak. On a fresh cache volume that refused the KV cache on the default profile (4.76 GiB needed for max_model_len 65,536, 4.73 GiB available on the 4090; 4.42 GiB available on the 3090) and onCTX=huge(available read as -1.86 GiB), while the same boot on a warm cache passed. 0.28.0 is a different case on the native 3090: atGPU_UTIL=0.90its default profile does not boot there at all, cold or warm (4.63 GiB against 4.76 needed in both cache states, main's CI image), while it passes on the WSL2 4090; 0.29.0's profiling change introduced a cache-state dependence 0.28.0 never had, and the fix removes it and clears a requirement 0.28.0 does not clear on that card in any state. The patch runs the pass once uncounted, releases the allocator's cached blocks, and measures the second pass. The pinned-pool path (kv_cache_memory_bytes) already did this and is unchanged.cudagraph-memory-from-allocator.patch): 0.29's CUDA-graph memory estimate and its logged "actual" pool are both the driver's free-memory delta across a capture. Read at the same two points by the allocator (torch.cuda.memory_reserved), the real pool on the WSL2 4090 is 0.23 GiB onCTX=hugeand 0.09 on the default profile, while the driver's figure there is pinned at 0.000 once the KV cache fills the budget (so the log said "actual 0.0") and falls from 5.44 GiB to zero during the first compile of the KVarN kernels on a cold cache, which the profiler then subtracted from the KV budget and refused the boot at -0.97 GiB available. The patch hascapture_modelreturn the allocator's reserved delta, logs the driver delta beside it, and reads the FULL-graph samples the estimate extrapolates from the same way. Knowingly unreserved by that reading: the driver-side residue of a capture (graph executables, module loads, local memory, non-torch workspaces), measured on the native 3090 as 0.02 GiB on the default profile and 0.03 onCTX=huge(driver delta minus allocator delta on the same capture). Not claimed: the estimate is not made accurate by this, only the actual is measured (1.03 GiB estimated against 0.09 actual on the default profile after it, on both boxes); what the correction returns is KV on every boot (native default profile 5.37 to 5.48 GiB, pool 73,631 to 75,173;CTX=hugepool 283,185 to 290,265 cold, 292,035 to 299,115 warm). With it, the coldCTX=hugeboot on a fresh volume on the WSL2 4090 passes with the estimate on (4.19 GiB available, pool 270,796, health at 288 s, a request served; estimate 0.28 GiB for a 0.23 GiB pool, the driver's delta across that capture still 5.44), warm reads 4.41 GiB (pool 284,955), and the cold default boot 5.77 GiB (pool 79,414, 0.09 above the profiling fix alone); the logged "actual" pool reads 0.23 and 0.09 instead of 0.0. Still open on both boxes with both patches:CTX=hugecold does not equal warm (0.22 GiB apart on the 4090, 0.14 on the 3090, where the default profile is exactly flat at 5.48 GiB and 75,173 tokens in both states); neither box has an explanation.REQ_METRICS=1also passes--per-request-spec-decode-metrics summary(a three-value enum in 0.29.0, absent from 0.28.0), andREQ_METRICS_DETAILED=1selectsdetailed, the ordered per-step arrays that upstream says are not free to collect, so it stays off in any benchmarked profile. Both are in all three launchers (single-user/start_qwen.sh,single-user/alternative.sh,batch/start_qwen.sh) and theREQ_METRICSrow ofsingle-user/README.md, with then == 1limit and the v0.29.0 shape noted.vllm servestill readsVLLM_PREFIX_CACHE_RETENTION_INTERVALand logs its deprecation, but the--prefix-cache-retention-intervalflag's unset default is passed explicitly and wins, so hybrid + EAGLE falls back to dense (vllm/engine/arg_utils.pyat v0.29.0; vLLM main has no reference to the variable, and [Deprecation] Deprecate items scheduled for 0.29 vllm-project/vllm#55353 removed its registration after the 0.29 branch cut).single-user/start_qwen.shnow passes the flag, and an exported value is carried over as it. Native 3090, one image,CTX=huge SPEC=dflash2 PREFIX_CACHE=1, attention block 2176: the engine takes 13056 with no dense line, and an exported 13057 is refused at boot (not a multiple of 2176); exported through the variable by the previous launcher (8cf70d5, same vLLM tree), the same value had booted. Prefix reuse drops to **0** on every turn when two long conversations alternate (CTX=huge / KVarN k4v2, attention block 2176) #174's alternating pair at 32,600 tokens a side (T=0, 24 tokens out, n=1 per arm): with the interval, the first reuse after the cold turn is 80.0% at 7.0 s (the last retained snapshot, 2 x 13056) and the next two are 99.4% and 99.3% at about 0.5 s; the dense control on the same image reuses 0% and re-prefills in about 29 s every turn. The ~50-55K knee and larger pools were not re-measured on 0.29.Every run setting, 0.28 against 0.29, same box, same harness. Main's own CI image (
ghcr.io/syv-ai/hyperqwen:sha-684e927, vLLM 0.28.0) against this branch's image, both atGPU_UTIL=0.90,bench/run_benchmarks.shin each mode's own mode run twice with the second kept,bench/quality_battery.pyat n=50, one exported made-upVLLM_name as the positive control (exactly one unknown-variable line on every one of the twelve boots, naming it). WSL2 RTX 4090, single-stream decode on the real-prompt cohort (C1, model-default sampling), pool in tokens:DFLASH_TOKENS=15)SPEC=dflash2 CTX=fast PREFIX_CACHE=1 DFLASH_TOKENS=15 INT8_ACT=int8 PREFILL_ATTN=int8)SPEC=mtp CTX=longCTX=huge(KVarN)Greedy rows and C2 to C8 move the same way; the full table with TTFT, tokens per step and the batch cohorts is in
docs/vllm-0.29.md. Batch trades a little steady-state decode (median TPOT 32.4 to 33.6 ms) for faster admission (mean TTFT 5.8 s to 3.8 s at 64 concurrent). Quality: perplexity on the fixed English and Danish corpora agrees to three decimals on every pair; GSM8K at n=50 is within two questions everywhere (the widest gap is the production line, 0.960 against 0.920). The battery's code corpus is the engine under test (it globs the installedvllm/v1/core), so that row is not an A/B number and is not quoted.Native 3090, the same twelve-boot protocol (
GPU_UTIL=0.90, per-arm cache volume, the quality corpus pinned to the same fifteen files, one GPU with production stopped, 13:37 to 16:44Z; the four earlier boots at the launcher default 0.93 are superseded). 0.28 does not boot three of the six settings on this card at 0.90: B and C refuse at 4.63 GiB available against 4.76 needed, cold and warm alike, and A refuses at 4.74 needed for max_model_len 150,000, so those three have a 0.29 row and no A/B, and the table says so rather than leaving a blank that reads as a zero. Single-stream decode on the C1 cohort at the model-default temperature (C over mean TPOT, the same estimator as the card-1 table), tokens per step, pool in tokens:DFLASH_TOKENS=15)SPEC=mtp CTX=longCTX=huge(KVarN)Three readings. The gains come from different places: E gains 11.7% with tokens per step flat (2.54 to 2.53), so that is raw decode plus a 19.4% larger pool, not better speculation; the production line gains 6.0% with tokens per step up 7.2% on an identical pool, which is speculation. The native gains are smaller than the WSL2 4090's and the production line flips sign (4090 -2%, within noise; 3090 +6.0%), with D and E at roughly half the 4090's figures (+4.4% and +11.7% against +8% and +17%); the two boxes are printed as two rows rather than averaged. B and C are identical on 0.29 on both boxes (115.1 and 115.1 here, 138.1 and 138.3 on the 4090), so
DFLASH_TOKENS=15changes nothing at the default temperature. Batch on the 3090 (0.29 only, 64 concurrent 128 in / 512 out): 1213 tok/s decode by C over median TPOT (a different estimator from the C1 rows), 991 e2e, median TPOT 52.76 ms, mean TTFT 4.24 s. Under greedy (T=0) the three settings with both arms read: the production line 133.3 to 140.3 (+5.3%), D 95.5 to 93.5 (-2.1%), E 88.3 to 99.4 (+12.6%); D flips sign under greedy where the production line flipped between boxes at the default temperature, so the small deltas move under sampling and the headline deltas above are the default-temperature ones. None of the greedy deltas is claimed: on an 8-prompt cohort they are noise within about ±11%, because 0.29's numerics change the text on every prompt and per-prompt acceptance swings ±25% either way while netting to 0.6% (the maintainer's measurement in review of this PR, which retracted the greedy DFlash2 regression on the same grounds).The cold-boot falsifier for the profiling fix, both boxes. Same image with and without the patch,
GPU_UTIL=0.90, cold = a fresh cache volume. Native 3090, default profile: cold without the fix refused (torch.compile 44.00 s, 4.42 GiB available); warm without 5.37 GiB, pool 73,631; cold with the fix 5.37 GiB, pool 73,631, health at 221 s; warm with the fix 5.37 GiB, pool 73,631. WSL2 4090, default profile: cold without refused (4.76 GiB, 4.73 after alignment against 4.76 needed); cold with the fix 5.68 GiB, pool 77,872, health at 242 s; warm with the fix 5.68 GiB, same pool. The fix makes the cold boot measure exactly what the warm boot measures, and the equivalent-utilization line (0.8514 in every native row, 0.8522 in every 4090 row) never distinguished a failure from a pass on the default profile. 0.28.0's default boot on the native 3090 refuses cold and warm alike (4.63 GiB against 4.76 needed in both cache states, main's CI image) and passes on the WSL2 4090, so 0.28.0's shortfall there is flat while 0.29.0 without the fix depends on the cache state (4.42 cold, 5.37 warm); with the fix it is 5.37 in both states, 0.74 GiB above 0.28.0's figure, a bar 0.28.0 never cleared on that card. Not fixed by the profiling pass alone on one box:CTX=hugecold on the WSL2 4090, which goes from -1.86 GiB available to -0.97 and still refuses (fixed by the fifth commit, above), while the native 3090 with the profiling pass passes cold (4.38 GiB, pool 283,185, health at 333 s) and warm (4.52 GiB, pool 292,035); the 4090 warm with it reads 4.31 GiB, pool 279,646. Where the 4090's 5.28 GiB cold residual sits: cutting the cache both ways (only the 37 KVarN kernel directories kept and everything else deleted passes at the warm figure; everything kept except those directories refuses at the cold figure) puts it entirely on the KVarN kernels' first compile, and vLLM's own profiling breakdown puts 0.21 GiB of it in the profile (the native box pays 0.14 for the same term) and the other 5.1 GiB in 0.29's CUDA-graph memory estimate, which reads 5.44 GiB on that box cold against 0.37 warm (equivalent utilization 0.673 on every refusing boot, 0.8845 on every passing one) and 0.37 on the native cold boot. Warming that estimate pass once uncounted does not cure it (still refuses, and costs 0.24 GiB on a warm boot), so that is not on the branch; why the estimate inflates on the WSL2 box and not the native one is not identified.VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0boots it (cold 4.47 GiB, pool 290,265; warm 4.68, pool 302,654) but under-reserves a real 0.23 GiB pool, and the knob is part of the compile-cache key; the allocator reading above is what turned that workaround into the fifth commit's fix. On the default profile the equivalent-utilization line is identical across failing and passing boots and is not the explanation; onCTX=hugeit separates them, because the term it reports is the estimate, which is the thing that moves. Rows, the cut, and the log lines:docs/vllm-0.29.md.Checked, and how.
patches/check_vllm_series.shon pristine v0.29.0: pass 1, 41 patches applied with exact context, 2 at an offset, 0 with fuzz (the check prints nothing for a clean patch, only offsets); ten of the rows come from two later fork branches (four fromqwen38/0.29-hqat 4879f94f3, six fromqwen38/0.29-hq2at da6a87935, each cut so the branch before it stays unrewritten), so regenerating from an older export point would give an older tree, andPATCHES.mdnames all three points; pass 2, the five contractual DFlash patches apply standalone withgit apply --check --whitespace=error, exit 0, on Linux (the Windows path cannot run pass 2). The image builds with the--fuzz 0apply loop andverify.sh --installon both boxes. CI on this PR is pending maintainer approval (a fork PR; the workflows are gated on "Approve and run"), so no check has run;patch-integrity's command,bash patches/check_vllm_series.shagainst a pristine v0.29.0 tree, was run locally on both boxes with the result above. The launcher change on main since #119 (resolve_config.sh,resolve_api_key.sh) is untouched; the port's retirement ofVLLM_V2_CUDAGRAPH_MEM_MIBrides on top of it (0 exports on this line, 0 reads in the patched 0.29 tree).Not measured, stated plainly. No int4 or offload boot on this branch (both were booted on the earlier port at f2acb1c with an identical
vllmtree; the rows are indocs/vllm-0.29.md). The campaign's cache volume was keyed per arm, so only the first setting of each arm booted cold; the cold-boot failure can hit whichever setting goes first on a fresh volume, which is why it showed on B and E here. On the WSL2 box the 0.29 huge-context boot delivers its first token early and streams slower than 0.28, finishing sooner at equal quality; the native 3090 shows no such difference. Not understood, WSL2-only, not a regression. The #64 corruption underCTX=huge SPEC=mtp PREFIX_CACHE=1reads clean on the 0.29 image twice and corrupt on the 0.28 image twice (en PPL 10.77 against 14.06 and 12.88), a finding and not a fix claim: the mechanism is not identified and the pool-size hypothesis is excluded in both directions.Decisions a reviewer can reject by name are the twelve in
docs/vllm-0.29.md("Decisions made in this port"), plus the fourth commit's choice to measure the second profiling pass rather than reserve a fixed margin (a margin spends context permanently for a transient that happens once per cache volume).