Skip to content

patches: keep the runtime-only env knobs out of the torch.compile cache key (#183) - #186

Merged
mhenrichsen merged 1 commit into
mainfrom
fix/compile-key-runtime-knobs
Sep 23, 2026
Merged

mhenrichsen merged 1 commit into
mainfrom
fix/compile-key-runtime-knobs

Conversation

@mhenrichsen

Copy link
Copy Markdown
Contributor

From #183, point 3: @cch-zuzuche saw a cold compile (~227 s against ~150 s) every time they changed VLLM_MAMBA_ALIGN_KEEP_CHECKPOINTS or VLLM_ENGINE_STALL_SENTINEL_S.

Why

0.29's compile_factors() (envs.py) hashes every registered VLLM_* variable that isn't in ignored_factors. This repo's patches register their knobs so they can be read through envs, and #114 required that for good reason: the spec-decode ones really do change compiled code. But registering a knob that no compiled graph reads puts it in the key anyway. Measured on the installed 0.29 tree:

  • VLLM_ENGINE_STALL_SENTINEL_S: read by the engine loop's watchdog thread (v1/engine/core.py). Registered, not ignored.
  • VLLM_MAMBA_ALIGN_KEEP_CHECKPOINTS: read by the KV block manager (v1/core/single_type_kv_cache_manager.py). Registered, not ignored.
  • VLLM_DFLASH2_CHAIN_LOG_SEC: a log interval.
  • VLLM_MARLIN_TUNE_DIR: a directory path that nothing reads.
  • VLLM_PREFIX_CACHE_RETENTION_INTERVAL: the deprecated env spelling of --prefix-cache-retention-interval, which CacheConfig.compute_hash already leaves out of its own hash for exactly this reason. The launchers use the flag, which is fine, but the env var still cost a cold compile.

The change

A new last entry in patches/series, compile-key-runtime-knobs.patch, adds those five names to ignored_factors beside upstream's own runtime-only entries. It's a separate patch so the fork-exported files stay untouched; it's hand-cut against 0.29.0 with the series applied, and PATCHES.md says so.

Verification

  • check_vllm_series.sh on pristine v0.29.0: 43 at exact context, 0 offset, 0 fuzz.
  • compile_factors() on the reference box's installed 0.29 envs.py, before and after the patch, hashing the key under each setting against the baseline:
setting before after
VLLM_ENGINE_STALL_SENTINEL_S=300 key changes same key
VLLM_MAMBA_ALIGN_KEEP_CHECKPOINTS=1 key changes same key
VLLM_DFLASH2_CHAIN_LOG_SEC=5 key changes same key
VLLM_PREFIX_CACHE_RETENTION_INTERVAL=13056 key changes same key
control VLLM_SPEC_DECODE_ATTN=1 key changes key changes
control KVARN_SPLIT_K=8 key changes key changes
control VLLM_DFLASH2_LOOKUP=1 key changes key changes

The controls are knobs that do reach compiled code, and they still invalidate the cache. So the patch removes exactly the five names and nothing else.

…he key (#183)

compile_factors() hashes every registered VLLM_* variable not in
ignored_factors, so changing the engine-stall watchdog timeout, the Mamba
checkpoint-retention toggle, the DFlash2 chain log interval, an unread
directory path, or the deprecated retention env var forced a cold compile
for a setting no compiled graph reads.
@mhenrichsen
mhenrichsen merged commit 73fd65d into main Sep 23, 2026
3 checks passed
@mhenrichsen
mhenrichsen deleted the fix/compile-key-runtime-knobs branch September 23, 2026 07:39
cpuchip added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
…ai#184 prefix-reuse bench); the new patch imported as a fork commit

compile-key-runtime-knobs applied as-is to cpuchip/vllm qwen38/0.29-hq2 (e9c839b27, author Mads Henrichsen) and
re-exported; PATCHES.md header names it with -hq2's other rows; Dockerfile.fork pins e9c839b27.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant