patches: keep a promoted draft block a multiple of the layer's kernel granularity - #142
Conversation
… granularity The promotion picks the smallest divisor of the primary block that still covers the draft layer's page. A divisor can satisfy the scheduler's LCM invariant and still break the layer's own kernel granularity: vLLM requires a framework-level block size to be a multiple of the backend's kernel requirement (Backend.supports_block_size, the MultipleOf branch), and upstream's unify_kv_cache_spec_page_size cannot violate that because it scales by whole ratios only. Derived from the code, no hardware needed: primary block 1840 with a covering block of 96 runs the promotion, the smallest divisor of 1840 at or above 96 is 115, and 115 % 16 == 3 -- not a multiple of the SW drafter's 16-token kernel granularity. 460 and 920 of 1840 fail the same way. Adds the constraint to the divisor search and regenerates the two affected hunk headers. The no-divisor fallback is unchanged: log and leave the config alone rather than risk the scheduler. Verified: bash patches/check_vllm_series.sh against a pristine v0.28.0 checkout -- 38 patches applied with exact context, 0 offsets, 0 fuzz; the five contractual DFlash patches pass git apply --check; patch integrity OK.
|
Reproduced your verification independently, on a fresh Both regenerated hunk headers check out by hand as well: hunk 1 is The invariant is quoted correctly too — if isinstance(supported_size, MultipleOf):
supported_size = supported_size.base
# With hybrid_blocks feature, the framework-level block size
# only needs to be a multiple of the kernel's requirement,
# even if the kernel requires a fixed block_size.
if block_size % supported_size == 0:
return Trueand your worked case holds: Two notes rather than objections:
Related history, if you want the context: #63 is where the divide-the-primary-block rule came from (int4 |
The added condition reads the layer's current framework block, which is the backend's kernel granularity only because nothing has promoted this layer yet -- it is not get_supported_kernel_block_sizes(). Write that down in the patch, since the two diverge for a layer whose block was already promoted upstream of this call. Hunk headers regenerated for the five comment lines (1033,143 -> 1033,148; second hunk start 1060 + (148-6) = 1202).
|
Added the sentence — and thank you for reproducing the series on a pristine 0.28.0 tree and checking the hunk arithmetic by hand; that is more than I had any right to expect. The patch comment now says that Hunk headers regenerated for the five comment lines: Your point 2 is still yours: nothing here changes a shipped geometry, and the boot that confirms Thanks for the #63 pointer — that is the other half of the invariant and I had not connected them. |
|
Booted on the reference box. Every shipped geometry comes out identical to the token, which is what this was waiting on. Same box, same tree, same checkpoint; the only change between arms was swapping the installed
The last row is the one I cared about most. All three Merging. Thank you for arguing this from the invariant rather than waiting for a crash — it is exactly the class of bug that would have surfaced on someone else's card, at a geometry nobody here runs, as an illegal memory access with no pointer back to this function. |
…, the four serving and bench patches from syv-ai#165-syv-ai#168, the triton message fix)
What
hybrid-sw-block-promotepromotes a draft sliding-window layer's block to the smallest divisor of the primary block that still covers that layer's page. This change adds one condition to the search: the divisor must also be a whole multiple of the layer's ownspec.block_size— 16 for the SW drafter. The no-divisor fallback is unchanged.Why
A divisor can satisfy the scheduler's LCM invariant and still break the layer's own kernel granularity. vLLM states the rule itself: a framework-level block size must be a multiple of the backend's kernel requirement —
Backend.supports_block_size(v1/attention/backend.py, theMultipleOfbranch: "the framework-level block size only needs to be a multiple of the kernel's requirement"). Upstream'sunify_kv_cache_spec_page_sizecannot violate it, because it only ever scales by whole ratios (new_block_size = layer_spec.block_size * ratio). Replacing that ratio with an arbitrary divisor of the primary block can.Concrete case, derived from the code, no hardware needed — primary block 1840, covering block 96:
1840 % 96 != 0, so the promotion runs;115 % 16 == 3, so the promoted block is not a multiple of the drafter's 16-token kernel granularity.The same search can also land on 460 or 920 of 1840 (
460 % 16 == 12,920 % 16 == 8).Verification
bash patches/check_vllm_series.sh <pristine vLLM v0.28.0 checkout>:Both hunk headers in the touched file are regenerated to match the added lines (
@@ -1033,6 +1033,135 @@→+1033,143,@@ -1060,6 +1189,12 @@→+1197,12), so the file parses and applies cleanly.Limits
No hardware failure has been observed for this geometry; the argument is invariant-based — the promotion could emit a block size that
supports_block_sizewould reject. The fallback is untouched: when no divisor qualifies, the config is left alone and the layer's page is padded at block 16 (expensive but boots).Review follow-up: what
spec.block_sizestands forThe condition reads the layer's current framework block, which is the backend's kernel granularity only because nothing has promoted this layer yet — it is not
get_supported_kernel_block_sizes(). The patch now says so, since the two diverge for a layer whose block was already promoted upstream of this call, and this condition would then be the weaker of the two.Hunk headers regenerated for the five comment lines:
-1033,6 +1033,143becomes+1033,148, and the second hunk's start becomes1060 + (148-6) = 1202. Counted against the emitted body, and thegit-applyjob is green on the rebuilt branch.The boot on the reference box (that the shipped
CTX=fast/CTX=long/CTX=hugegeometries come out unchanged) is still yours to run — nothing here changes that.