Skip to content

Pin flip to vLLM 0.29.0: the series re-exported at fuzz 0, KVarN on 0.29, the #114 moves, a cold-boot profiling fix, and 0.28 vs 0.29 on every run setting (#106 part two) - #148

Merged
mhenrichsen merged 23 commits into
syv-ai:mainfrom
cpuchip:port-0.29-hq
Sep 23, 2026

Conversation

@cpuchip

@cpuchip cpuchip commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

The pin flip to vLLM 0.29.0, the second half of #106, on top of #119's shape. On top of main: the pin and series, the carried launcher and docs, the #114 moves, the profiling fix, the allocator-measured graph pool, the campaign and falsifier doc, the PATCHES.md rows for the four newest entries with the header naming both fork export points, the docs that followed, a merge of main's field-tests commit (#146) with the README's 0.29 row rewritten for this branch, and a merge of main at 13b30ea: #142's divisor condition and the four patches from #165-#168, ported to the 0.29 fork and re-exported.

  1. vllm==0.29.0 in docker/requirements.txt, and the series re-exported from the fork, one commit per row on v0.29.0 (cpuchip/vllm: 31 rows from qwen38/0.29, four from qwen38/0.29-hq, six from qwen38/0.29-hq2, whose tip is the tree the series reproduces; PATCHES.md names each row's export point): 42 series entries (41 applied, dflash2-backport kept and skipped by name since DFlash2 went native in 0.28.0) plus the two KVarN files, every one applied and checked at --fuzz 0, replay onto pristine v0.29.0 identical to the fork tip (0 differing files) on a WSL2 4090 and on a native 3090 independently, both re-run against the fork tip after the last re-cut (the WSL2 replay against the current tip da6a87935, after the merge of main at 13b30ea; the native replay against fea76cc82, whose tree da6a87935 shares, likewise 0 differing files; patches/check_vllm_series.sh passes 1 and 2 green on the native box at 62fd3d4 and 652f99a) (the native replay had first been verified at an earlier tip; patches/check_vllm_series.sh resets the tree when it finishes, so the identity diff has to follow a replay of one's own, not the check). Three patches retire because 0.29.0 carries them (vllm-pr54282-draft-gumbel-salt, xgrammar-spec-terminated, sse-keep-alive); two hunks retire because 0.29 does the work itself (the graph-memory reserve in hybrid-kv-groups-v2-cudagraph, the int4 padded-page view). PATCHES.md has the row for each and names the retirements.
  2. KVarN on 0.29 (kvarn/kvarn-0.29.0.patch, kvarn/kvarn-v2-runner-0.29.0.patch, the copied modules): 0.29 stopped asking a backend for its KV cache shape and stride, so the backend declares its layout and folds the runner's view into tiles with a view that fails on a wrong layout. Same block geometry and the same pool as 0.28 on the reference profile; the KVARN_* knobs are registered under their names, the VLLM_KVARN_* rename stays a follow-up (a decision to reject by name).
  3. The spec-decode-attn.patch: VLLM_SPEC_DECODE_ATTN_QMAX and VLLM_SPEC_ATTN_BLOCK_M are read with bare os.environ.get, outside the compile cache key #114 knob moves on this line: VLLM_SPEC_DECODE_ATTN, VLLM_SPEC_DECODE_ATTN_QMAX and VLLM_SPEC_ATTN_BLOCK_M registered in spec-decode-attn.patch, the patch that reads them, anchored on the VLLM_DP_MASTER_PORT region no other patch touches; the registry has the same 301 entries before and after.
  4. One fix the pin needed and 0.28 did not (memory-profile-after-warmup.patch): 0.29's determine_available_memory runs the profiling pass inside the memory-profiling window, so on a cold compile cache the compiler's scratch is read as the model's transient peak. On a fresh cache volume that refused the KV cache on the default profile (4.76 GiB needed for max_model_len 65,536, 4.73 GiB available on the 4090; 4.42 GiB available on the 3090) and on CTX=huge (available read as -1.86 GiB), while the same boot on a warm cache passed. 0.28.0 is a different case on the native 3090: at GPU_UTIL=0.90 its default profile does not boot there at all, cold or warm (4.63 GiB against 4.76 needed in both cache states, main's CI image), while it passes on the WSL2 4090; 0.29.0's profiling change introduced a cache-state dependence 0.28.0 never had, and the fix removes it and clears a requirement 0.28.0 does not clear on that card in any state. The patch runs the pass once uncounted, releases the allocator's cached blocks, and measures the second pass. The pinned-pool path (kv_cache_memory_bytes) already did this and is unchanged.
  5. A second fix the same falsifier led to (cudagraph-memory-from-allocator.patch): 0.29's CUDA-graph memory estimate and its logged "actual" pool are both the driver's free-memory delta across a capture. Read at the same two points by the allocator (torch.cuda.memory_reserved), the real pool on the WSL2 4090 is 0.23 GiB on CTX=huge and 0.09 on the default profile, while the driver's figure there is pinned at 0.000 once the KV cache fills the budget (so the log said "actual 0.0") and falls from 5.44 GiB to zero during the first compile of the KVarN kernels on a cold cache, which the profiler then subtracted from the KV budget and refused the boot at -0.97 GiB available. The patch has capture_model return the allocator's reserved delta, logs the driver delta beside it, and reads the FULL-graph samples the estimate extrapolates from the same way. Knowingly unreserved by that reading: the driver-side residue of a capture (graph executables, module loads, local memory, non-torch workspaces), measured on the native 3090 as 0.02 GiB on the default profile and 0.03 on CTX=huge (driver delta minus allocator delta on the same capture). Not claimed: the estimate is not made accurate by this, only the actual is measured (1.03 GiB estimated against 0.09 actual on the default profile after it, on both boxes); what the correction returns is KV on every boot (native default profile 5.37 to 5.48 GiB, pool 73,631 to 75,173; CTX=huge pool 283,185 to 290,265 cold, 292,035 to 299,115 warm). With it, the cold CTX=huge boot on a fresh volume on the WSL2 4090 passes with the estimate on (4.19 GiB available, pool 270,796, health at 288 s, a request served; estimate 0.28 GiB for a 0.23 GiB pool, the driver's delta across that capture still 5.44), warm reads 4.41 GiB (pool 284,955), and the cold default boot 5.77 GiB (pool 79,414, 0.09 above the profiling fix alone); the logged "actual" pool reads 0.23 and 0.09 instead of 0.0. Still open on both boxes with both patches: CTX=huge cold does not equal warm (0.22 GiB apart on the 4090, 0.14 on the 3090, where the default profile is exactly flat at 5.48 GiB and 75,173 tokens in both states); neither box has an explanation.
  6. REQ_METRICS: add --per-request-spec-decode-metrics once a vLLM release contains upstream #48915 (not in 0.28.0) #66's per-request spec-decode metrics, since the pin now has them: REQ_METRICS=1 also passes --per-request-spec-decode-metrics summary (a three-value enum in 0.29.0, absent from 0.28.0), and REQ_METRICS_DETAILED=1 selects detailed, the ordered per-step arrays that upstream says are not free to collect, so it stays off in any benchmarked profile. Both are in all three launchers (single-user/start_qwen.sh, single-user/alternative.sh, batch/start_qwen.sh) and the REQ_METRICS row of single-user/README.md, with the n == 1 limit and the v0.29.0 shape noted.
  7. launcher: CTX=huge + DFlash2 retains one Mamba snapshot in six (#174) #179's retention interval, which 0.29 would otherwise drop silently: on 0.29, vllm serve still reads VLLM_PREFIX_CACHE_RETENTION_INTERVAL and logs its deprecation, but the --prefix-cache-retention-interval flag's unset default is passed explicitly and wins, so hybrid + EAGLE falls back to dense (vllm/engine/arg_utils.py at v0.29.0; vLLM main has no reference to the variable, and [Deprecation] Deprecate items scheduled for 0.29 vllm-project/vllm#55353 removed its registration after the 0.29 branch cut). single-user/start_qwen.sh now passes the flag, and an exported value is carried over as it. Native 3090, one image, CTX=huge SPEC=dflash2 PREFIX_CACHE=1, attention block 2176: the engine takes 13056 with no dense line, and an exported 13057 is refused at boot (not a multiple of 2176); exported through the variable by the previous launcher (8cf70d5, same vLLM tree), the same value had booted. Prefix reuse drops to **0** on every turn when two long conversations alternate (CTX=huge / KVarN k4v2, attention block 2176) #174's alternating pair at 32,600 tokens a side (T=0, 24 tokens out, n=1 per arm): with the interval, the first reuse after the cold turn is 80.0% at 7.0 s (the last retained snapshot, 2 x 13056) and the next two are 99.4% and 99.3% at about 0.5 s; the dense control on the same image reuses 0% and re-prefills in about 29 s every turn. The ~50-55K knee and larger pools were not re-measured on 0.29.

Every run setting, 0.28 against 0.29, same box, same harness. Main's own CI image (ghcr.io/syv-ai/hyperqwen:sha-684e927, vLLM 0.28.0) against this branch's image, both at GPU_UTIL=0.90, bench/run_benchmarks.sh in each mode's own mode run twice with the second kept, bench/quality_battery.py at n=50, one exported made-up VLLM_ name as the positive control (exactly one unknown-variable line on every one of the twelve boots, naming it). WSL2 RTX 4090, single-stream decode on the real-prompt cohort (C1, model-default sampling), pool in tokens:

setting 0.28 tok/s 0.29 tok/s 0.28 pool 0.29 pool
B, single default 121.8 138.1 66,692 77,872
C, reproduction (DFLASH_TOKENS=15) 127.4 138.3 66,692 77,872
the production line (SPEC=dflash2 CTX=fast PREFIX_CACHE=1 DFLASH_TOKENS=15 INT8_ACT=int8 PREFILL_ATTN=int8) 144.1 141.6 57,669 57,669
D, SPEC=mtp CTX=long 94.4 101.8 163,010 161,479
E, CTX=huge (KVarN) 91.5 107.2 221,238 281,415
A, batch, 64 concurrent 128 in / 512 out: decode / e2e 1976 / 1336 1904 / 1495 152,319 168,556

Greedy rows and C2 to C8 move the same way; the full table with TTFT, tokens per step and the batch cohorts is in docs/vllm-0.29.md. Batch trades a little steady-state decode (median TPOT 32.4 to 33.6 ms) for faster admission (mean TTFT 5.8 s to 3.8 s at 64 concurrent). Quality: perplexity on the fixed English and Danish corpora agrees to three decimals on every pair; GSM8K at n=50 is within two questions everywhere (the widest gap is the production line, 0.960 against 0.920). The battery's code corpus is the engine under test (it globs the installed vllm/v1/core), so that row is not an A/B number and is not quoted.

Native 3090, the same twelve-boot protocol (GPU_UTIL=0.90, per-arm cache volume, the quality corpus pinned to the same fifteen files, one GPU with production stopped, 13:37 to 16:44Z; the four earlier boots at the launcher default 0.93 are superseded). 0.28 does not boot three of the six settings on this card at 0.90: B and C refuse at 4.63 GiB available against 4.76 needed, cold and warm alike, and A refuses at 4.74 needed for max_model_len 150,000, so those three have a 0.29 row and no A/B, and the table says so rather than leaving a blank that reads as a zero. Single-stream decode on the C1 cohort at the model-default temperature (C over mean TPOT, the same estimator as the card-1 table), tokens per step, pool in tokens:

setting 0.28 tok/s 0.29 tok/s 0.28 tok/step 0.29 tok/step 0.28 pool 0.29 pool
B, single default refused (4.63 GiB available, 4.76 needed) 115.1 2.67 73,631
C, reproduction (DFLASH_TOKENS=15) refused (same 4.63, warm cache) 115.1 2.67 73,631
the production line 126.4 134.0 3.19 3.42 57,669 57,669
D, SPEC=mtp CTX=long 88.7 92.6 2.60 2.66 156,122 179,846
E, CTX=huge (KVarN) 83.8 93.6 2.54 2.53 237,168 283,185
A, batch mode, C1 single stream refused (4.74 GiB needed for max_model_len 150,000) 48.1 no tok/step reported 183,247

Three readings. The gains come from different places: E gains 11.7% with tokens per step flat (2.54 to 2.53), so that is raw decode plus a 19.4% larger pool, not better speculation; the production line gains 6.0% with tokens per step up 7.2% on an identical pool, which is speculation. The native gains are smaller than the WSL2 4090's and the production line flips sign (4090 -2%, within noise; 3090 +6.0%), with D and E at roughly half the 4090's figures (+4.4% and +11.7% against +8% and +17%); the two boxes are printed as two rows rather than averaged. B and C are identical on 0.29 on both boxes (115.1 and 115.1 here, 138.1 and 138.3 on the 4090), so DFLASH_TOKENS=15 changes nothing at the default temperature. Batch on the 3090 (0.29 only, 64 concurrent 128 in / 512 out): 1213 tok/s decode by C over median TPOT (a different estimator from the C1 rows), 991 e2e, median TPOT 52.76 ms, mean TTFT 4.24 s. Under greedy (T=0) the three settings with both arms read: the production line 133.3 to 140.3 (+5.3%), D 95.5 to 93.5 (-2.1%), E 88.3 to 99.4 (+12.6%); D flips sign under greedy where the production line flipped between boxes at the default temperature, so the small deltas move under sampling and the headline deltas above are the default-temperature ones. None of the greedy deltas is claimed: on an 8-prompt cohort they are noise within about ±11%, because 0.29's numerics change the text on every prompt and per-prompt acceptance swings ±25% either way while netting to 0.6% (the maintainer's measurement in review of this PR, which retracted the greedy DFlash2 regression on the same grounds).

The cold-boot falsifier for the profiling fix, both boxes. Same image with and without the patch, GPU_UTIL=0.90, cold = a fresh cache volume. Native 3090, default profile: cold without the fix refused (torch.compile 44.00 s, 4.42 GiB available); warm without 5.37 GiB, pool 73,631; cold with the fix 5.37 GiB, pool 73,631, health at 221 s; warm with the fix 5.37 GiB, pool 73,631. WSL2 4090, default profile: cold without refused (4.76 GiB, 4.73 after alignment against 4.76 needed); cold with the fix 5.68 GiB, pool 77,872, health at 242 s; warm with the fix 5.68 GiB, same pool. The fix makes the cold boot measure exactly what the warm boot measures, and the equivalent-utilization line (0.8514 in every native row, 0.8522 in every 4090 row) never distinguished a failure from a pass on the default profile. 0.28.0's default boot on the native 3090 refuses cold and warm alike (4.63 GiB against 4.76 needed in both cache states, main's CI image) and passes on the WSL2 4090, so 0.28.0's shortfall there is flat while 0.29.0 without the fix depends on the cache state (4.42 cold, 5.37 warm); with the fix it is 5.37 in both states, 0.74 GiB above 0.28.0's figure, a bar 0.28.0 never cleared on that card. Not fixed by the profiling pass alone on one box: CTX=huge cold on the WSL2 4090, which goes from -1.86 GiB available to -0.97 and still refuses (fixed by the fifth commit, above), while the native 3090 with the profiling pass passes cold (4.38 GiB, pool 283,185, health at 333 s) and warm (4.52 GiB, pool 292,035); the 4090 warm with it reads 4.31 GiB, pool 279,646. Where the 4090's 5.28 GiB cold residual sits: cutting the cache both ways (only the 37 KVarN kernel directories kept and everything else deleted passes at the warm figure; everything kept except those directories refuses at the cold figure) puts it entirely on the KVarN kernels' first compile, and vLLM's own profiling breakdown puts 0.21 GiB of it in the profile (the native box pays 0.14 for the same term) and the other 5.1 GiB in 0.29's CUDA-graph memory estimate, which reads 5.44 GiB on that box cold against 0.37 warm (equivalent utilization 0.673 on every refusing boot, 0.8845 on every passing one) and 0.37 on the native cold boot. Warming that estimate pass once uncounted does not cure it (still refuses, and costs 0.24 GiB on a warm boot), so that is not on the branch; why the estimate inflates on the WSL2 box and not the native one is not identified. VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 boots it (cold 4.47 GiB, pool 290,265; warm 4.68, pool 302,654) but under-reserves a real 0.23 GiB pool, and the knob is part of the compile-cache key; the allocator reading above is what turned that workaround into the fifth commit's fix. On the default profile the equivalent-utilization line is identical across failing and passing boots and is not the explanation; on CTX=huge it separates them, because the term it reports is the estimate, which is the thing that moves. Rows, the cut, and the log lines: docs/vllm-0.29.md.

Checked, and how. patches/check_vllm_series.sh on pristine v0.29.0: pass 1, 41 patches applied with exact context, 2 at an offset, 0 with fuzz (the check prints nothing for a clean patch, only offsets); ten of the rows come from two later fork branches (four from qwen38/0.29-hq at 4879f94f3, six from qwen38/0.29-hq2 at da6a87935, each cut so the branch before it stays unrewritten), so regenerating from an older export point would give an older tree, and PATCHES.md names all three points; pass 2, the five contractual DFlash patches apply standalone with git apply --check --whitespace=error, exit 0, on Linux (the Windows path cannot run pass 2). The image builds with the --fuzz 0 apply loop and verify.sh --install on both boxes. CI on this PR is pending maintainer approval (a fork PR; the workflows are gated on "Approve and run"), so no check has run; patch-integrity's command, bash patches/check_vllm_series.sh against a pristine v0.29.0 tree, was run locally on both boxes with the result above. The launcher change on main since #119 (resolve_config.sh, resolve_api_key.sh) is untouched; the port's retirement of VLLM_V2_CUDAGRAPH_MEM_MIB rides on top of it (0 exports on this line, 0 reads in the patched 0.29 tree).

Not measured, stated plainly. No int4 or offload boot on this branch (both were booted on the earlier port at f2acb1c with an identical vllm tree; the rows are in docs/vllm-0.29.md). The campaign's cache volume was keyed per arm, so only the first setting of each arm booted cold; the cold-boot failure can hit whichever setting goes first on a fresh volume, which is why it showed on B and E here. On the WSL2 box the 0.29 huge-context boot delivers its first token early and streams slower than 0.28, finishing sooner at equal quality; the native 3090 shows no such difference. Not understood, WSL2-only, not a regression. The #64 corruption under CTX=huge SPEC=mtp PREFIX_CACHE=1 reads clean on the 0.29 image twice and corrupt on the 0.28 image twice (en PPL 10.77 against 14.06 and 12.88), a finding and not a fix claim: the mechanism is not identified and the pool-size hypothesis is excluded in both directions.

Decisions a reviewer can reject by name are the twelve in docs/vllm-0.29.md ("Decisions made in this port"), plus the fourth commit's choice to measure the second profiling pass rather than reserve a fixed margin (a margin spends context permanently for a transient that happens once per cache volume).

… qwen38/0.29 at 337efb79f (35 patches, 0 fuzz, replay identical to the fork tip); KVarN 0.29 patches and modules; retirements listed in PATCHES.md
…uild, docs (carried from port-0.29-bae2023 onto the split README)
…ds them

Matches syv-ai#119's shape on the 0.28 line. VLLM_SPEC_DECODE_ATTN, VLLM_SPEC_DECODE_ATTN_QMAX
and VLLM_SPEC_ATTN_BLOCK_M move out of the speed-knobs-envs catch-all into
spec-decode-attn; VLLM_SPEC_ATTN_DEBUG stays with triton-spec-attn-fp8-kv and is
re-anchored beside BLOCK_M's new location, which is what the rebase conflicted on.

Done in the fork, not by editing patch files: new branch qwen38/0.29-hq off 337efb79f with
the registrations moved between topic commits, then spec-decode-attn, speed-knobs-envs and
triton-spec-attn-fp8-kv re-exported. A new branch rather than a rewrite of qwen38/0.29, so
no force-push.

The envs.py hunk in spec-decode-attn anchors on VLLM_DP_MASTER_PORT - present in pristine
v0.29.0 and untouched by any other patch - the same region basecamp anchored the 0.28 moves
on, so pass 2's standalone-apply contract is unaffected.

Verified, this being a relocation and not a change:
  fork 38 commits before and after, 0 subjects lost, HEAD attached, branch ref == HEAD
  envs registry 301 entries before and after, 0 differing lines (same set, same defaults)
  replay onto pristine v0.29.0: 35 series + 2 kvarn at --fuzz 0, 0 failed, 2 at an offset
    with exact context, 0 stray .orig/.rej
  replayed tree IDENTICAL to the new fork tip 03b5c4259
  check_vllm_series.sh both passes green on Linux, exit 0, patch integrity: OK
…nflate the profiled transient peak and refuse the KV cache the warm boot grants (cold first boots failed on B and E on a WSL2 4090 and on a native 3090; series 36 + kvarn 2, replay identical to fork tip bf29fa109)
…L2 4090), and the cold-boot falsifier on both boxes (default profile fixed; CTX=huge cold boot still refused on the WSL2 4090, open)
… had no row

Two defects found while vetting the rebased 0.29, both mine.

The header named one export point (`qwen38/0.29`), but three rows were
re-cut on `qwen38/0.29-hq` for the syv-ai#114 registration moves -- so a reader
regenerating spec-decode-attn, speed-knobs-envs or triton-spec-attn-fp8-kv
from the named branch would get the pre-syv-ai#114 hunks and a --fuzz 0 failure
they could not explain from this file. Both branches are now named with
their commit ids, and the three re-cut rows are called out by name.

Four entries in patches/series had zero mentions in the table:
engine-completion-log, engine-stall-sentinel,
topk-honour-flashinfer-sampler-switch, triton-spec-attn-fp8-kv. The table
is the only place a patch's "retires when" is written down, so an
unlisted patch is one nobody knows the exit condition for.

Accounting, stated so the next drift is visible: 36 entries in
patches/series, 36 series rows, plus 2 kvarn rows that install.sh applies
outside patches/series -- 38 rows total.
…oved to bf29fa109

Adds the missing row for the cold-boot fix and repoints the header: qwen38/0.29-hq
is now bf29fa109, and it carries four rows, not three -- the three re-cut for the
syv-ai#114 registration moves plus memory-profile-after-warmup, which was cut there after
that branch had already diverged.

Accounting: 37 entries in patches/series, 37 series rows, plus 2 kvarn rows that
install.sh applies outside patches/series -- 39 rows total. The commit that added
the patch describes the series as 36; it is 37.
The 0.29 port carried four claims written against 0.28.0 without re-checking
them. Three are still true and now say so against the pin actually in use;
the fourth is not re-verified and now says that instead of implying currency.

Verified against v0.29.0 in a checkout, not from memory:
  - gotcha 18: DFlashModelTypes is still inside EagleModelTypes, so dflash keeps
    async scheduling on (speculative.py:69 on 0.29.0, :67 on 0.28.0 -- the cited
    line had drifted by two).
  - gotcha 19: async_scheduling still resolves to True in the else branch; one
    "= True" and five "= False" assignments in config/vllm.py at both tags.
  - the offload residency note: still no gauge; kv_offload is absent from
    v1/metrics/loggers.py at both tags.

Not verified: the FlashInfer k=4 illegal-memory-access paragraph in
single-user/README.md was measured on 0.28.0 and has not been re-run on 0.29.0.
Marked unproven either way rather than restated, because the pin moved under it.

Also repointed two capacity-planning TODOs that still said "re-run after the
v0.28.0 upgrade" and were two pins behind.
…sses; the open item narrows to the WSL2 4090's first boot
…e cold default boot on the native 3090; the WSL2 4090 CTX=huge residual is the graph estimate), same hunks, fork tip 512a9699c; docs: warm rows for both boxes, the cache cut that locates the residual, the profiling breakdown, and why an uncounted estimate pass is not the fix
…: the 0.29 row of the field-tests table now names this branch's pin
…CTX=huge path on the WSL2 4090 (cold 4.47 GiB, warm 4.68, actual graph pool 0.0), the knob in the compile-cache key, the equivalent-utilization line scoped by profile
…ctual graph pool reads 0.0 GiB; the native 3090 reads 0.11 to 1.87 GiB by profile, so it under-reserves there
…he allocator's reserved bytes (WSL2's driver reading is pinned at zero after KV allocation and collapses by 5.44 GiB during a cold KVarN compile, which the 0.29 estimate subtracted from the KV budget); memory-profile-after-warmup header corrected (0.28 refuses the native default boot in both cache states); docs: the allocator probe, both fixes' rows, the 0.28 frame
…th both arms, three 0.28 refusals with numbers), the measured driver-to-allocator residue (0.02 and 0.03 GiB), the KV the allocator reading returns, the estimator not made accurate, the CTX=huge cold-to-warm gap open on both boxes; patch header carries the residue
…syv-ai#142 and the four new patches

hybrid-sw-block-promote re-exported with syv-ai#142's divisor condition (a whole multiple of the layer's kernel
block); bench-probe-errors, serve-404-served-names, serve-model-path-match and tokenize-v1-route applied
as-is to the 0.29 fork and re-exported (line offsets only). All five from cpuchip/vllm qwen38/0.29-hq2 at
fea76cc82, a new branch so qwen38/0.29-hq stays unrewritten. PATCHES.md: the four rows cut against 0.29.0,
the header names both fork points and every row cut from them (memory-profile-after-warmup at c06f8ef11,
the file's own export point, and cudagraph-memory-from-allocator, which the header had left out).
Series plus KVarN applied to v0.29.0 at --fuzz 0 reproduce fea76cc82's vllm/ tree (0 differing files).
…hat its fork commit message carried

The export kept the message's prose and dropped only the header lines of the pasted diff, so 223 hunk-body
lines sat in the file's preamble (GNU patch skipped them; a reader and the PR gate did not). The fork
commit is reworded on qwen38/0.29-hq2 (now at da6a87935, trees unchanged), and the four new topics are
re-exported from their moved commits. Hunks unchanged; series plus KVarN on v0.29.0 at --fuzz 0 reproduce
da6a87935's vllm/ tree (0 differing files).
cpuchip added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 22, 2026
…rted to 0.29) into main

main last took syv-ai main at 8d09ec5 and the first port-0.29 branch; the port was re-derived for the PR on
top of syv-ai#119's shape, so both lines carried the same files ported twice. The PR line is a superset for every
shared file (checked file by file: the same detection with --fuzz 0, the registry check with the seat name
reworded out, the retirements listed in PATCHES.md), so it wins there. docker-compose.yml and the image
workflow keep main's identity (ghcr.io/cpuchip/qwen38-27b-rtx3090 built from Dockerfile.fork) on top of
syv-ai's GPU_COUNT and pull-request build changes; cache-to takes the repository through format().
…CHES.md: dflash2-backport is the one file with no export point, and qwen38/0.29-hq2 is a rewrite of -hq from hybrid-sw-block-promote on
…efix-cache-retention-interval

vllm serve on 0.29 reads the deprecated VLLM_PREFIX_CACHE_RETENTION_INTERVAL in the field's default factory,
but the flag's unset default is passed explicitly by from_cli_args and wins, so the retention-interval-unset
branch forces hybrid + EAGLE to dense (vllm/engine/arg_utils.py at v0.29.0). Measured on a native 3090,
CTX=huge SPEC=dflash2 PREFIX_CACHE=1: the export logs the deprecation and boots dense, and an exported 13057
boots; the flag at 13056 boots with no dense line, and the flag at 13057 is refused (not a multiple of 2176).
The block sizes on 0.29 match 0.28's (2176 at 7 drafts, 2432 at 15). An exported value is carried over as
the flag; the flag in EXTRA_ARGS wins over both.
Ar4ikov added a commit to Ar4ikov/HyperQwen that referenced this pull request Sep 22, 2026
…apture pinned to the V1 runner

gptq_lm_head.py assumed the base model's layout: seven model-0000x shards, a .bak from
quant_lm_head.py, a quantization_config.json to copy, a model_extra_tensors.safetensors to
link. A checkpoint that went through prepare/quant_heads_stream.py has model.safetensors +
model-mtp.safetensors with .bak-orig backups and no quantization_config.json, so the script
stopped at its first line (no .bak), and past that the model-0000* glob would have left
model-mtp.safetensors, which the index points to, out of the variant. It now reads the bf16
lm_head from .bak or .bak-orig, hardlinks every weight shard the source has, copies only
the files that exist, and copies model_extra_tensors.safetensors instead of linking it:
build_draft_vocab.py rewrites that file in the variant with save_file, and safetensors
0.4.5 through 0.7.0 write it in place, through the hardlink into the source dir (measured;
0.8.0 replaces the file and leaves the source alone).

capture.py hooks the V1 runner (vllm.v1.worker.gpu_model_runner). On the pinned 0.28.0
that is the runner this model gets anyway -- a hybrid architecture outside
DEFAULT_V2_MODEL_RUNNER_ARCHITECTURES stays on V1 -- so VLLM_USE_V2_MODEL_RUNNER=0 changes
nothing today. vLLM 0.29.0 defaults every model to V2, and there the hooks never fire: the
capture finishes with rows=0 and the GPTQ that follows calibrates on zeros (KL 0.00000,
round-trip error 1.0000, an lm_head of zeros), measured on the 0.29 port (syv-ai#148). No
speculation happens in a capture, so V1 is the right runner on both. It also sets
FLASHINFER_DISABLE_VERSION_CHECK=1, as both launchers do, since it runs standalone.

Measured on Ar4ikov/Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM after quant_heads_stream.py,
600k teacher-forced UltraChat tokens captured in 10 min on a 3090 (vLLM 0.29.0): RTN int4
lm_head KL 0.00701, GPTQ int4 0.00239 (the base model's published figures: 0.0068 and
0.0029).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ar4ikov added a commit to Ar4ikov/HyperQwen that referenced this pull request Sep 22, 2026
…y checkpoints

The llm-compressor AWQ exports of the uncensored finetune and of the base model, int4
asymmetric g128 with zero points and the vision tower kept, after quant_heads_stream.py
and build_draft_vocab.py; the -fast siblings with the int4-GPTQ lm_head (building one from
a single-shard export takes the drafter/ fixes in syv-ai#181). The measured rows (vLLM 0.29.0,
the syv-ai#148 port with this patch) for MTP, DFlash2, CTX=long, the production line, the fast
variant and batch mode, whose shipped 0.972 / 150k did not boot with the tower on (syv-ai#182).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ar4ikov added a commit to Ar4ikov/HyperQwen that referenced this pull request Sep 22, 2026
…y checkpoints

The llm-compressor AWQ exports of the uncensored finetune and of the base model, int4
asymmetric g128 with zero points and the vision tower kept, after quant_heads_stream.py
and build_draft_vocab.py; the -fast siblings with the int4-GPTQ lm_head (building one from
a single-shard export takes the drafter/ fixes in syv-ai#181). The measured rows (vLLM 0.29.0,
the syv-ai#148 port with this patch) for MTP, DFlash2, CTX=long, the production line, the fast
variant and batch mode, whose shipped 0.972 / 150k did not boot with the tower on (syv-ai#182).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… on the first reuse after a cold turn, then 99.3-99.4%, against 0% dense)

Native 3090, one image, CTX=huge SPEC=dflash2 PREFIX_CACHE=1, syv-ai#174's alternating pair at 32,600 tokens a side,
T=0, max_tokens 24, n=1 per arm. The first reuse lands on the last retained snapshot (26,112 = 2 x 13056), so
a reader checking that turn against the 0.28 range would read a regression.
@mhenrichsen

Copy link
Copy Markdown
Contributor

Ran this on the reference 3090 tonight — a third box, native Linux, 250 W — with the branch head exactly as pushed: pristine vllm==0.29.0 reinstalled into a separate venv, this branch's series applied the Dockerfile's way (--fuzz 0, patches/series order), kvarn/install.sh, verify.sh --install. Same-box 0.28 arms from the production tree, and the same harness and the same vllm bench serve client for both versions, so the only variable is the server. Launcher defaults (GPU_UTIL=0.93, the profile pins), bench/run_benchmarks.sh single twice, second run kept.

Build and apply: clean. 37 patches at exact context, 0 offset, 0 fuzz, on the box's pristine 0.29.0; KVarN installs; verify.sh --install 0 failures.

Most of your claims reproduce:

profile (C1 decode, tok/s) 0.28 default / greedy 0.29 default / greedy pool 0.28 → 0.29
production line (CTX=fast, 15 drafts, int8 prefill) 121.4 / 135.9 128.5 / 128.2 57,669 → 57,669 (pinned)
SPEC=dflash2 CTX=huge 114.7 / 128.7 116.3 / 121.2 268,169 → 268,169 (pinned)
D, SPEC=mtp CTX=long 96.3 / 98.0 93.5 / 101.3 176,020 → 202,806
SPEC=mtp CTX=huge MAX_SEQS=4 — — 295,575 → 350,442
  • Default-temperature production line +5.8% (your native 3090: +6.0%), from acceptance, 3.22 → 3.38 tok/step. Same direction and size as yours.
  • The unpinned profiles get the memory: +15% pool on D, +18.6% on MTP at CTX=huge. The pinned DFlash2 profiles are byte-identical, as they should be.
  • Quality: 0.29 production line GSM8K 0.980 over 50, English/Danish PPL 11.04 / 11.39.
  • CTX=huge at MAX_LEN=240000: batched prompt_logprobs returns NaN (quality_battery 400s), plus PPL drift on the same profile #64 reproduces as clean on 0.29 on a second machine — your finding, now with two boxes behind it. Same box, back to back, SPEC=mtp CTX=huge PREFIX_CACHE=1 MAX_SEQS=4, --ppl-only: 0.28 English PPL 13.32 (corrupt), 0.29 10.767 — your 4090 read 10.766 and 10.765. Danish 10.909 against your 10.910 / 10.909. Three decimals across different hardware is not a coincidence.

One regression you did not report, and it is on the production line. Greedy single-stream acceptance drops on DFlash2, deterministically (greedy tok/step reproduces to two decimals across runs):

C1 greedy tok/step          0.28     0.29
production line (15 drafts) 3.58     3.38    (-5.6%, 135.9 -> 128.2 tok/s)
CTX=huge dflash2 (7 drafts) 3.47     3.28    (-5.5%, 128.7 -> 121.2 tok/s)
D, MTP (for contrast)       2.57     2.67    (up)

It is C1-specific — at C4/C8 greedy the two versions accept the same (3.51-3.55 both) — and DFlash2-specific, since MTP goes the other way. The sharper half, 0.29 at 7 drafts on the same production line:

0.29 C1 greedy     7 drafts: 137.6 tok/s, 3.61 tok/step
                  15 drafts: 128.2 tok/s, 3.38 tok/step

On 0.28, 15 drafts is the +9% setting production runs (LABD, the lookup lane filling the verify tail from context). On 0.29 it is now worse than 7 at greedy. That is also the explanation for your own table's oddity — B and C identical to the decimal on both your boxes: the long verify block has stopped paying on 0.29, and at greedy it costs. My guess is the lookup-drafting / adaptive-length patches as re-exported onto 0.29's native DFlash2 (vllm#52816 changed that path), but I have not bisected it and it is only a guess. The logged per-position acceptance cannot show it (both versions log zeros past position 7), so the harness's tok/step is the instrument.

So this does not merge tonight, for two reasons, neither of them the port's core:

  1. The greedy DFlash2 regression above wants a look before production follows the pin — either a fix, or a documented default change (e.g. 7 drafts on 0.29 if that is where it lands).
  2. It needs a rebase onto today's main, which moved a lot. Only three files conflict (PATCHES.md, patches/series, patches/hybrid-sw-block-promote.patch); the launcher changes merge cleanly. What the rebase needs:

One upgrade hazard to document with the flip. 0.29 ships FlashInfer 0.6.18, which invalidates every cached FlashInfer kernel. On a native install whose system nvcc is older than CUDA 13, the first setup-D boot JIT-compiles the fp8-KV prefill kernel with that nvcc and dies (nvcc fatal: Unknown option '--compress-mode=size' — this box's /usr/bin/nvcc is 12.0). Pointing CUDA_HOME at the venv's pip toolchain (…/site-packages/nvidia/cu13) fixes it. The Docker image is unaffected; native users upgrading from 0.28 are the ones who will hit it, and a launcher preflight or a README line would save them the stack trace.

The port itself is in good shape — clean apply, the memory fixes doing exactly what the writeup says, and #64 gone on two machines. Rebase plus an answer on the 15-draft greedy drop, and this is ready.

@mhenrichsen

Copy link
Copy Markdown
Contributor

A second blocker, found through #182: batch mode's shipped defaults do not boot on this branch on a 3090.

Reference box, reference checkpoint, batch/start_qwen.sh untouched (KV=fp8, GPU_UTIL=0.972, MAX_LEN=150000, int8 MLP activations, 64 seats):

0.28.0 0.29.0 (this head)
VISION=0 boots, KV 6.09 GiB, 192,525 tokens, 128 requests at C64 with 0 failed OOM in warmup (Tried to allocate 96.00 MiB, 69 MiB free), KV 7.63 GiB
VISION=1 boots, 220,360 tokens, 0 failed OOM in warmup

Same checkpoint, same settings, 0.29 gives the KV pool 7.63 GiB where 0.28 gave 6.09. That is your fourth and fifth commits doing what the writeup says — "what the correction returns is KV on every boot" — and at GPU_UTIL=0.972 the ~1.5 GiB returned was the headroom batch mode's unprofiled warmup transients (64 seats, the int8 activation workspace) were living in. Your campaign ran every arm at GPU_UTIL=0.90, so the shipped default was never booted on this line. The single-user profiles are unaffected because they pin KV_MEM; batch sizes its pool from GPU_UTIL, which is exactly the path your fixes changed.

The fix is yours to choose — a lower batch GPU_UTIL default on the 0.29 line (0.90 is known to boot; the largest value that does is not measured), or a fixed reserve on top of the corrected accounting — but it has to land with the pin, or setup A breaks for every 24 GB user the day this merges. Together with the greedy DFlash2 drop above, those are the two things between this and a merge; the rest of my verification is in the comment above.

PATCHES.md is the one conflict: the 0.29 table kept, syv-ai#172's
marlin-int8-asym-zp row added. The 0.28-cut patch applies to 0.29.0 plus
this series at exact context (check_vllm_series.sh: 42 at exact context,
0 offset, 0 fuzz).
@mhenrichsen

Copy link
Copy Markdown
Contributor

Correction: the greedy DFlash2 "regression" I reported is not one. Please don't spend time on it. It came from my measurement, and I owe you the retraction with the evidence.

I went to bisect it this morning and found nothing to bisect. On the rebuilt 0.29 tree (your head plus today's main, see below), reference 3090:

The lookup lane is identical on both versions. A verbatim-copy prompt, the lane's best case, 15 drafts, greedy: 15.45 tok/step on 0.28 and 15.45 on 0.29, to the step. Greedy chat prompts with thinking off: 2.96 vs 2.98.

The drop was trajectory divergence on 8 prompts. Harness-shaped requests (chat template, thinking on, 1024 tokens, greedy), sent one at a time, with tok/step and a hash of the output recorded per prompt:

 i   0.28  (rerun)   0.29    text 0.28 vs 0.29
 0   3.23   3.23     4.03    differs
 1   2.53   2.53     2.54    differs
 2   3.56   3.56     2.87    differs
 3   4.10   4.10     3.34    differs
 4   3.89   3.89     3.80    differs
 5   3.96   3.96     3.76    differs
 6   3.05   3.05     3.47    differs
 7   3.75   3.75     4.12    differs
 aggregate  3.43   3.43     3.41

0.28 reproduces itself exactly, text and tok/step. 0.29 writes different text on every prompt, because the numerics moved. Per prompt, that swings acceptance ±25% in both directions, and in aggregate it nets to 3.43 vs 3.41, a 0.6% difference. With the text diverging, the per-prompt deltas have a standard deviation of ~0.55 tok/step, so an 8-prompt greedy cohort's standard error is ~0.19 tok/step, about 6%. My −5.6% was inside one standard error, and so is anything in your campaign table's greedy rows below ~±11%. The same caveat applies to my "15 drafts now below 7" line: that compared different texts too.

The batch blocker stands. That one is a hard OOM at the shipped GPU_UTIL=0.972, not a statistic. I'm measuring the largest GPU_UTIL that boots and serves 64-way concurrency on this line right now, with and without the vision tower, plus the kvarn and int4pth batch modes. The fix will go on your branch with the numbers.

Branch housekeeping, already done: I merged today's main (#172, #181) into port-0.29-hq. The only conflict was PATCHES.md; I kept your 0.29 table and added #172's row. The 0.28-cut marlin-int8-asym-zp.patch applies to 0.29 plus the series at exact context: check_vllm_series.sh gives 42 at exact context, 0 offset, 0 fuzz. It's local for now and goes up with the batch fix.

@cpuchip

cpuchip commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for the retraction; the per-prompt output hash is the right instrument, and it covers our side too. Our native 3090 read greedy DFlash2 at 15 drafts up 9% (3.21 to 3.50 tok/step, 8 prompts, seed-0 greedy), which sits inside the same ±11% band, so we're not claiming it either. I'll put that caveat on the greedy line in the PR body.

On the batch blocker, the ladder is already measured on two native 3090s, which may save you the run. Shipped batch settings (KV=fp8, MAX_LEN=150000, INT8_ACT=int8 INT8_LAYERS=mlp, 64 seats, the reference AutoRound checkpoint, VISION at its default), a fresh cache volume per boot, then 128 requests at 64-way concurrency, about 1,035 tokens in (a unique salt per request, so no prefix hits) and 256 out with ignore_eos. The headless box ran that burst once, straight after boot; the rented box ran it twice, cold then warm:

GPU_UTIL headless native 3090, image from this branch rented native 3090, docs/install.md venv from this branch
0.972 OOM in warmup, 96 MiB asked, 57 MiB free, KV 7.64 GiB OOM in warmup, 96 MiB asked, 63 MiB free, KV 7.65 GiB
0.97 OOM in warmup, 192 MiB asked, 17 MiB free, KV 7.60 GiB not run
0.96 boots, 232,731 tokens, 128/128, 468 MiB free at the burst peak not run
0.95 boots, 225,000 tokens, 128/128, 668 MiB free at the burst peak boots, 225,773 tokens, 128/128 both bursts

On the rented box the image's own 0.28 tree at 0.972 boots with 192,525 tokens, your number exactly, and at 0.95 it gets 203,350. Warm throughput at 0.95 is level: 207.5 tok/s on 0.28, 210.4 on 0.29.

My proposal is 0.95 for fp8 batch on the 0.29 line: 0.96 boots but has only 468 MiB spare at the peak, and a card that also drives a display loses more than that before vLLM starts. At 0.95 the pool (225,000) is still 17% above what 0.28 gets at 0.972 (192,525). Not measured: VISION=1, other MAX_LEN values, and any card other than a 24 GB 3090.

The other two batch modes ship at 0.93, and both boot and serve on 0.29 at that value: kvarn on the rented 3090, 281,706 to 297,357 tokens, warm burst 188.7 to 189.8 tok/s; int4pth on a WSL2 4090, 356,515 to 426,920 tokens, 571 to 618 tok/s.

On the branch: the #172 merge is already staged on our side (port-0.29-hq-sync at ca4a706), with marlin-int8-asym-zp imported as a commit on the vLLM fork and re-exported from it, so every patch file still comes from a fork commit and the series replays onto pristine 0.29.0 identical to the fork tip. To avoid two different merges of the same main landing on one branch: if it suits you, I'll push ca4a706 plus a commit setting the fp8 batch default to 0.95 on the 0.29 line, and you run your boots against that head. If your ladder lands somewhere else, or you'd rather push yours, say so and I'll rebase onto it instead.

One more thing from your first comment: VLLM_PREFIX_CACHE_RETENTION_INTERVAL is not honoured on 0.29 under vllm serve. The flag's unset default wins over the env var, and hybrid + EAGLE falls back to dense. Item 7 of the PR body has the boots; the branch now passes the flag.

cpuchip added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
…ody); edit_body.sh chains the edit on the gate, DRY=1 for tests

Found 2026-09-23: `pr_gate.py --pr N --body f` read the body from GitHub and ignored the file, so every run meant to
gate a new body checked the old one. edit_body.sh gates the PR's live head with the new body and edits only on exit 0;
a falsification run of the first version (before this fix) was a live edit and overwrote syv-ai#148's body for about a
minute, so the script now says to test with DRY=1 only. Falsified offline: an em-dash body fails, the real body passes.
…i#182)

0.29 with this port's two memory fixes stops over-reserving ~1.5 GiB, and at
0.972 that was the headroom batch mode's unprofiled warmup transients used:
0.972 OOMs in warmup on a 3090 with the tower on and off. Reference-3090
ladder, boot plus 128 requests at C64: 0.93-0.96 all pass with VISION=0/1,
0.972 fails. 0.95 keeps a step below the highest pass, boots cold to the warm
pool, and holds at least the pool 0.28 had at 0.972. kvarn and int4pth are
unchanged at 0.93. docs/vllm-0.29.md gets the ladder, and a section on why
greedy tok/step cannot be compared across the two versions below ~±11%.
@mhenrichsen

Copy link
Copy Markdown
Contributor

Both blockers are resolved, so I'm merging this. What I pushed to the branch, and why:

1. Merge of today's main (7496231: #172, #181). The only conflict was PATCHES.md; I kept your 0.29 table and added #172's row. marlin-int8-asym-zp.patch (cut against 0.28) applies to 0.29 plus the series at exact context. check_vllm_series.sh on pristine v0.29.0 gives 42 at exact context, 0 offset, 0 fuzz.

2. Batch GPU_UTIL defaults to 0.95 on this line (20e4d51, fixes #182). Your two memory fixes stop 0.29 over-reserving ~1.5 GiB, and at 0.972 that ~1.5 GiB was batch mode's warmup headroom. Ladder on the reference 3090, each cell a boot plus 128 requests at C64:

GPU_UTIL VISION=1 VISION=0
0.93 / 0.94 / 0.95 pass: 204,896 / 212,628 / 219,587 tokens pass: 210,309 / 217,268 / 225,000
0.95, cold compile cache pass, 219,587 (= warm)
0.96 pass, 227,319
0.972 OOM in warmup OOM in warmup

0.95 keeps a full step below the highest pass. It lands on the same pool cold as warm, and it holds at least the pool 0.28 had at 0.972 (220,360 / 192,525), so nobody loses capacity in the upgrade. KV=kvarn and KV=int4pth pass at their existing 0.93 and are unchanged. The exact pushed launcher, with no GPU_UTIL set: picks 0.95, 225,000-token pool. bench/run_benchmarks.sh batch twice, second run kept: 1,039 tok/s decode / 952 e2e at 64 concurrent, which is the published 0.28 figure (~1,035 / 948), with 0 OOM lines.

3. The greedy "regression" is retracted (details in my previous comment). It was cross-version text divergence on 8 prompts; the per-prompt aggregate is 3.43 vs 3.41, and the lookup lane measures 15.45 tok/step on both versions. docs/vllm-0.29.md now has a section on why cross-version greedy tok/step can't be read below ~±11%, beside the ladder.

CI on 20e4d51: patch integrity and docker image both green.

Thank you for this port. The memory work in particular — measuring the second profiling pass, and reading the graph pool from the allocator instead of the driver — is what made the batch fix a one-line default rather than a workaround.

Closing on merge: #64 (clean on 0.29 on two machines), #66 (the enum flag is in both launchers), #106, #114 (both knobs read through envs on this line), #182.

@mhenrichsen
mhenrichsen merged commit 858c3b7 into syv-ai:main Sep 23, 2026
2 checks passed
cpuchip added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
…yv-ai#136) into main

Conflicts: docs/vllm-0.29.md takes syv's (a superset: the maintainer's greedy-divergence section); the image workflow
keeps this fork's identity (Dockerfile.fork into ghcr.io/cpuchip/...); PATCHES.md keeps this fork's header and
marlin-int8-asym-zp row, which describe the fork-exported file this main carries (hunks identical to syv's hand-cut
file; syv's series replays to fork tip 291980422 with 0 differing files). Batch GPU_UTIL 0.95 comes in from syv-ai#148.
TyroneNel added a commit to TyroneNel/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
Takes upstream's 0.29.0 series wholesale (patches/series, PATCHES.md, the
re-cut patches, KVarN 0.29.0), including the four patches syv-ai#148 retired
(int4-mq3d-envs, sse-keep-alive, vllm-pr54282-draft-gumbel-salt,
xgrammar-spec-terminated).

Drops the fork's auth-deny-default.patch from the series and the tree: it is
cut against 0.28.0, both of its target files moved in 0.29.0
(serve/utils/server_utils.py -> serve/middleware/authenticate.py,
openai/cli_args.py -> launchers/cli_args.py), so it cannot apply at fuzz 0.
It returns with the 0.29 port in syv-ai#169.

Keeps from the fork: manual-only image builds and the fork's own buildcache
ref (docker-image.yml), the guarded resolver source in the bench scripts,
the F12 digest-pinned base, and F04 copy-only variant writes in
drafter/gptq_lm_head.py (upstream's syv-ai#181 file handling, copy semantics).
verify.sh equals upstream: every fork change to it landed via syv-ai#158/syv-ai#171/syv-ai#172.
aaronlockhartdev added a commit to aaronlockhartdev/syv-recipes that referenced this pull request Sep 25, 2026
…varn, Dockerfile, recipes, docs); adopted auth-deny-default

- vllm==0.29.0 + huggingface_hub==1.28.0; Dockerfile/requirements claims
  re-based; Python 3.14 and our layout kept (upstream runs 3.12)
- patches/: the 0.29 re-export verbatim (PR syv-ai#148), 43 files, all
  --fuzz 0. Retired by 0.29 itself: vllm-pr54282-draft-gumbel-salt and
  xgrammar-spec-terminated (in 0.29.0), int4-mq3d-envs (absorbed into
  speed-knobs-envs; VLLM_INT4_MQ_3D stays, the int4 recipes keep
  working), the graph-memory-reserve hunk of hybrid-kv-groups-v2-
  cudagraph and the padded-page-view hunk of int4-kv-per-token-head
  (0.29 covers both natively). VLLM_V2_CUDAGRAPH_MEM_MIB is dead (0.29
  profiles graph memory itself), the six exporting recipes drop it
- new patches: memory-profile-after-warmup + cudagraph-memory-from-
  allocator (0.29's cold-cache profiling refused boots; the pair
  repairs it), topk-honour-flashinfer-sampler-switch (0.29 ships
  FlashInfer 0.6.18, whose top-k JIT needs nvcc >= 13; the recipes'
  VLLM_USE_FLASHINFER_SAMPLER=0 now covers the drafter's top-k),
  compile-key-runtime-knobs, engine-completion-log, engine-stall-
  sentinel (off by default), serve-404-served-names, serve-model-path-
  match, tokenize-v1-route, marlin-int8-asym-zp (syv-ai#172), auth-deny-
  default (syv-ai#169, adopted by user decision)
- kvarn trio re-ported to 0.29 (kvarn-0.29.0, kvarn-files-0.29.0
  generated from upstream kvarn/files/, kvarn-v2-runner-0.29.0),
  last in the series in upstream's install.sh order
- recipes: PYTORCH_CUDA_ALLOC_CONF defaults expandable_segments:False
  at TP=2 (custom all-reduce cannot export a VMM graph buffer over
  CUDA IPC, syv-ai#163/syv-ai#176); w4a16-k4v2-dflash2 gains --prefix-cache-
  retention-interval 13056 (6 x 2176, the attention block at 7 drafts;
  upstream syv-ai#174)
- README/AGENTS re-based to 0.29.0; upstream's 0.29 measurements
  (their boxes) carried in the Notes, labeled as upstream
  measurements; bare-metal note that 0.29's FlashInfer JIT needs
  nvcc >= 13 (CUDA_HOME at the venv's nvidia/cu13; the Docker image
  carries CUDA 13)
- not adopted (harness / dropped lanes): dflash2-backport, offload-
  mtp-serve, offload-wsl2-devptr, bench-probe-errors, triton-spec-
  attn-fp8-kv, upstream's _check_applied.py / check_vllm_series.sh
  tooling

Validated: docker build on Linux (aarch64) applies all 43 in series
order at --fuzz 0 and compileall-gates; in-image patch_vllm.py audit
reports 43 of 43 in place on vllm 0.29.0.
TyroneNel added a commit to TyroneNel/qwen38-27b-rtx3090 that referenced this pull request Sep 25, 2026
…rlay without vLLM in the FA-scratch fixture

bench-policy.yml asserts the F10 guard in patches/check_vllm_series.sh
(refuse this repo, and any tree without .qwen-disposable-series-target
unless ALLOW_TREE_RESET=1) and patch-integrity.yml still stamps the
sentinel, but the guard never reached main: it lived on
validity/build-integrity (0ebd561), and main carried upstream's
unguarded script. Ported onto the current script.

bench/test_fa_scratch_capacity.py loaded kvarn/.../config.py as a
stdlib-only module. Since the 0.29 pin flip (858c3b7, syv-ai#148/syv-ai#114) the
overlay reads its knobs through vllm.envs, so the fixture died with
ModuleNotFoundError: vllm under the CI python, and with AttributeError
KVARN_FA_SCRATCH_CAP in a venv without KVarN installed. It now loads the
overlay against a stand-in vllm.envs carrying that one knob with the
patch's semantics (raw string, None when unset).

  bench/test_fa_scratch_capacity.py   python3 / venv: RESULT PASS
  bench-policy mocks                  policy mocks: OK
  check_vllm_series.sh .              REFUSING (this repo), exit 1
  check_vllm_series.sh <no sentinel>  REFUSING, exit 1
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants