bench: prefix-reuse reproducer, interleaved-traffic dose-response, and a reuse-correctness needle - #184
Conversation
…d a reuse-correctness needle Three controls for prefix-reuse behaviour, which this repo has no reproducer for. They came out of syv-ai#174 and each one exists because its absence produced a wrong answer first. bench/prefix_alternation.py — two independent long conversations advanced in strict alternation, reading prompt_tokens_details.cached_tokens per request. The pass criterion is "this turn reused the whole prefix the previous turn established" (cached >= 0.9 * prev_prompt), not "a high hit rate": the failure mode is a clean 0%, so an average hides it. It reads the engine's own environment via /proc so an arm label cannot be whatever the caller exported. bench/interleave_dose.py — the two controls that make the above legible. --mode single advances one conversation with nothing else on the box; if that is clean the cache works and the defect is about interleaving. --mode dose injects one independent request of varying size between two turns of a long conversation. The knee this finds is the quantity to compare across arms: it moves with the KV pool, which is how capacity is separated from a fixed structural trigger. Both size their documents after calibrating tokens/char against the running tokenizer — a hardcoded ratio is a silent arm change, and the first version of this asked for 55,000 tokens, produced 48,000, and landed below the knee it was looking for. bench/needle_reuse.py — needle_test.py sends cold prompts and so never exercises the reuse path. Anything that changes which KV gets restored (retention intervals, checkpoint ordering, dtype, block promotion) is invisible to it, and the failure it misses is the worst kind: a high hit rate with a stale restored state. This one puts the passcode inside the reused prefix and requires both a high cache percentage and a correct answer. An optional interleaved second conversation can be advanced between the turns, and a needle at depth 0.95 sits at the edge of the reused region. All three take --tokens/--rounds/--doses and follow bench/ conventions (VLLM_API, VLLM_API_KEY or ../api_key.txt, optional VLLM_MODEL). Verified end to end against a running server.
…, judge turns against syv-ai#102's floor, salt each run - The engine is found by its API port in /proc (or HQ_UNIT's MainPID, user unit first), and the retention interval and max_model_len are read from its command line: this repo's launcher passes both as flags, so the environment alone labelled a 13056 arm '<unset>'. - A healthy turn reuses max(0, prev//B - 1)*B (syv-ai#102: the block holding the previous request's end is never served), B from vllm:cache_config_info, or --block when the engine was launched with an explicit --block-size, where the metric reports that value instead. The flat 90% rule marked every healthy turn under ~20 blocks PREFIX-LOST. - Documents carry a per-run salt, so round 1 is cold even when a previous run's prefix is still cached.
|
Merged, with one commit of mine on top that fixes three things in 1. The arm label read the wrong place for this repo. 2. The 90% pass rule misreads every short conversation. A hybrid model never serves back the block holding the previous request's end (#102), so a healthy turn reuses 3. Round 1 wasn't always cold. The documents were deterministic, so a second run within the cache's lifetime started with a hit. On my first run its round 1 read After the fixes, against production: Worth knowing from the 0.29 Thank you for the harness. The single-conversation and dose-response controls are what made #174 legible, and now anyone can rerun them. |
…ai#184 prefix-reuse bench); the new patch imported as a fork commit compile-key-runtime-knobs applied as-is to cpuchip/vllm qwen38/0.29-hq2 (e9c839b27, author Mads Henrichsen) and re-exported; PATCHES.md header names it with -hq2's other rows; Dockerfile.fork pins e9c839b27.
Three controls for prefix-reuse behaviour. This came out of #174, and each one exists because leaving it out produced a wrong answer first.
You referred to the main one as
bench/test-align-cache.py; renamed to snake_case to match the rest ofbench/. Contents are as run, plus argument handling.bench/prefix_alternation.pyTwo independent long conversations advanced in strict alternation, reading
prompt_tokens_details.cached_tokensper request. For A, each of B's turns is the interleaved traffic, and vice versa; every turn is its own request, which is what an agentic loop does.The pass criterion is not "a high hit rate" but "this turn reused the whole prefix the previous turn established" (
cached >= 0.9 * prev_prompt). The failure is a clean 0%, so any average over turns hides it.It reads the engine's own environment through
/proc/<MainPID>/environto label the arm, rather than trusting the caller's shell — that is what caught an arm label being wrong.bench/interleave_dose.pyThe two controls that make the above interpretable.
--mode singleadvances one conversation with nothing else on the box. If that is clean, the cache works and the defect is about interleaving. Without this arm, "conversation X lost its prefix" cannot be read.--mode doseinjects one independent request of varying size between two turns of a long conversation. The knee this finds is the quantity to compare across arms, because it moves with the KV pool — that is what separates capacity from a fixed structural trigger.Both size their documents after calibrating tokens/char against the running tokenizer. A hardcoded ratio is a silent arm change: the first version of this asked for 55,000 tokens, produced 48,000, and landed below the knee it was trying to find, which cost an afternoon.
bench/needle_reuse.pyneedle_test.pysends cold prompts, so it never exercises the reuse path. Anything that changes which KV is restored — retention intervals, checkpoint ordering, KV dtype, block promotion — is invisible to it, and the failure it misses is the worst kind: a healthy hit rate with a stale restored state.This one puts the passcode inside the reused prefix and requires both conditions: cache percentage above the threshold and the right answer. An optional interleaved second conversation can be advanced between the turns, and a needle at depth 0.95 sits right at the edge of the reused region.
Notes
--tokens/--rounds/--dosesand followbench/conventions:VLLM_API,VLLM_API_KEYor../api_key.txt, plus an optionalVLLM_MODELfor stacks whose served name is not the launcher's default.run_benchmarks.sh, as you asked.1/6in the interpolation is the retention interval from Prefix reuse drops to **0** on every turn when two long conversations alternate (CTX=huge / KVarN k4v2, attention block 2176) #174; the tools do not depend on any particular value.Refs #174.