Skip to content

bench: prefix-reuse reproducer, interleaved-traffic dose-response, and a reuse-correctness needle - #184

Merged
mhenrichsen merged 2 commits into
syv-ai:mainfrom
cch-zuzuche:bench/prefix-reuse-alternation
Sep 23, 2026
Merged

mhenrichsen merged 2 commits into
syv-ai:mainfrom
cch-zuzuche:bench/prefix-reuse-alternation

Conversation

@cch-zuzuche

Copy link
Copy Markdown
Contributor

Three controls for prefix-reuse behaviour. This came out of #174, and each one exists because leaving it out produced a wrong answer first.

You referred to the main one as bench/test-align-cache.py; renamed to snake_case to match the rest of bench/. Contents are as run, plus argument handling.

bench/prefix_alternation.py

Two independent long conversations advanced in strict alternation, reading prompt_tokens_details.cached_tokens per request. For A, each of B's turns is the interleaved traffic, and vice versa; every turn is its own request, which is what an agentic loop does.

The pass criterion is not "a high hit rate" but "this turn reused the whole prefix the previous turn established" (cached >= 0.9 * prev_prompt). The failure is a clean 0%, so any average over turns hides it.

It reads the engine's own environment through /proc/<MainPID>/environ to label the arm, rather than trusting the caller's shell — that is what caught an arm label being wrong.

bench/interleave_dose.py

The two controls that make the above interpretable.

  • --mode single advances one conversation with nothing else on the box. If that is clean, the cache works and the defect is about interleaving. Without this arm, "conversation X lost its prefix" cannot be read.
  • --mode dose injects one independent request of varying size between two turns of a long conversation. The knee this finds is the quantity to compare across arms, because it moves with the KV pool — that is what separates capacity from a fixed structural trigger.

Both size their documents after calibrating tokens/char against the running tokenizer. A hardcoded ratio is a silent arm change: the first version of this asked for 55,000 tokens, produced 48,000, and landed below the knee it was trying to find, which cost an afternoon.

bench/needle_reuse.py

needle_test.py sends cold prompts, so it never exercises the reuse path. Anything that changes which KV is restored — retention intervals, checkpoint ordering, KV dtype, block promotion — is invisible to it, and the failure it misses is the worst kind: a healthy hit rate with a stale restored state.

This one puts the passcode inside the reused prefix and requires both conditions: cache percentage above the threshold and the right answer. An optional interleaved second conversation can be advanced between the turns, and a needle at depth 0.95 sits right at the edge of the reused region.

Notes

  • All three take --tokens / --rounds / --doses and follow bench/ conventions: VLLM_API, VLLM_API_KEY or ../api_key.txt, plus an optional VLLM_MODEL for stacks whose served name is not the launcher's default.
  • Not wired into run_benchmarks.sh, as you asked.
  • Verified end to end against a live 0.29 server: the reuse needle returns 3/3 correct at ~91% cached, and the alternation arm reports the arm label it read from the engine rather than the one exported.
  • The 1/6 in the interpolation is the retention interval from Prefix reuse drops to **0** on every turn when two long conversations alternate (CTX=huge / KVarN k4v2, attention block 2176) #174; the tools do not depend on any particular value.

Refs #174.

…d a reuse-correctness needle

Three controls for prefix-reuse behaviour, which this repo has no reproducer
for. They came out of syv-ai#174 and each one exists because its absence produced a
wrong answer first.

bench/prefix_alternation.py — two independent long conversations advanced in
strict alternation, reading prompt_tokens_details.cached_tokens per request.
The pass criterion is "this turn reused the whole prefix the previous turn
established" (cached >= 0.9 * prev_prompt), not "a high hit rate": the failure
mode is a clean 0%, so an average hides it. It reads the engine's own
environment via /proc so an arm label cannot be whatever the caller exported.

bench/interleave_dose.py — the two controls that make the above legible.
--mode single advances one conversation with nothing else on the box; if that is
clean the cache works and the defect is about interleaving. --mode dose injects
one independent request of varying size between two turns of a long
conversation. The knee this finds is the quantity to compare across arms: it
moves with the KV pool, which is how capacity is separated from a fixed
structural trigger. Both size their documents after calibrating tokens/char
against the running tokenizer — a hardcoded ratio is a silent arm change, and
the first version of this asked for 55,000 tokens, produced 48,000, and landed
below the knee it was looking for.

bench/needle_reuse.py — needle_test.py sends cold prompts and so never exercises
the reuse path. Anything that changes which KV gets restored (retention
intervals, checkpoint ordering, dtype, block promotion) is invisible to it, and
the failure it misses is the worst kind: a high hit rate with a stale restored
state. This one puts the passcode inside the reused prefix and requires both a
high cache percentage and a correct answer. An optional interleaved second
conversation can be advanced between the turns, and a needle at depth 0.95 sits
at the edge of the reused region.

All three take --tokens/--rounds/--doses and follow bench/ conventions
(VLLM_API, VLLM_API_KEY or ../api_key.txt, optional VLLM_MODEL). Verified
end to end against a running server.
…, judge turns against syv-ai#102's floor, salt each run

- The engine is found by its API port in /proc (or HQ_UNIT's MainPID, user
  unit first), and the retention interval and max_model_len are read from its
  command line: this repo's launcher passes both as flags, so the environment
  alone labelled a 13056 arm '<unset>'.
- A healthy turn reuses max(0, prev//B - 1)*B (syv-ai#102: the block holding the
  previous request's end is never served), B from vllm:cache_config_info, or
  --block when the engine was launched with an explicit --block-size, where
  the metric reports that value instead. The flat 90% rule marked every
  healthy turn under ~20 blocks PREFIX-LOST.
- Documents carry a per-run salt, so round 1 is cold even when a previous
  run's prefix is still cached.
@mhenrichsen

Copy link
Copy Markdown
Contributor

Merged, with one commit of mine on top that fixes three things in prefix_alternation.py. All three turned up when I ran it on this repo's own stack. interleave_dose.py and needle_reuse.py went in unchanged: both ran clean against the reference box, and the needle came back OK at 78.5% cached.

1. The arm label read the wrong place for this repo. engine_env() asked systemctl show qwen38-hq-vllm, your unit, at system scope. Here the unit is a user unit (qwen-single), and Docker has none. And even with the right PID, it read EXTRA_ARGS and MAX_LEN from the engine's environment. This repo's launcher computes both and passes them as flags (--prefix-cache-retention-interval 13056 since #179/#148, and --max-model-len), so on a CTX=huge server it would have printed retention=<unset>: exactly the mislabel the function exists to prevent. It now finds the engine by the API port in /proc (or HQ_UNIT's MainPID if set, user unit first) and reads the command line, with the environment as the fallback for the deprecated env spelling. Verified live: retention=13056 MAX_LEN=245760 on a 0.29 CTX=huge server, retention=dense MAX_LEN=57344 on production.

2. The 90% pass rule misreads every short conversation. A hybrid model never serves back the block holding the previous request's end (#102), so a healthy turn reuses max(0, prev // B − 1) × B. At 3K tokens with B = 480 that is 2,400 (79%), which the old rule reported as PREFIX-LOST. With CTX=huge's 2176-token block, every healthy turn under ~44K tokens would have — right at the knee you were mapping. The verdict now uses that floor, with B read from vllm:cache_config_info on /metrics. One wrinkle, found live: when the engine was launched with an explicit --block-size (CTX=huge passes 128, KVarN's tile), the metric reports that value, not the resolved hybrid block. So there's a --block flag for that case, and the metric is ignored whenever --block-size is on the engine's command line.

3. Round 1 wasn't always cold. The documents were deterministic, so a second run within the cache's lifetime started with a hit. On my first run its round 1 read cached=2,400. They now carry a per-run salt, as interleave_dose.py already did.

After the fixes, against production:

engine: retention=dense MAX_LEN=57344   engine attention block: 480
r1 A COLD  cached 0 | r1 B COLD  cached 0 | r2/r3 A,B ok  cached 2,400 (72.6-77.5%)

Worth knowing from the 0.29 CTX=huge run: at ~20K per conversation, with 13056 retention, round 2 reused exactly 13,056 on both sides. That's the first-reuse-after-cold rounding cpuchip documented on #148 (80% first, then 99.3–99.4%), so two rounds show only the first case. The default of 8 rounds is right.

Thank you for the harness. The single-conversation and dose-response controls are what made #174 legible, and now anyone can rerun them.

@mhenrichsen
mhenrichsen merged commit e972eaf into syv-ai:main Sep 23, 2026
3 checks passed
cpuchip added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
…ai#184 prefix-reuse bench); the new patch imported as a fork commit

compile-key-runtime-knobs applied as-is to cpuchip/vllm qwen38/0.29-hq2 (e9c839b27, author Mads Henrichsen) and
re-exported; PATCHES.md header names it with -hq2's other rows; Dockerfile.fork pins e9c839b27.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants