Skip to content

fix(kv-eval): flush the prefix cache between cases, and make --seed actually vary the haystack - #16

Merged
MiaAI-Lab merged 2 commits into
MiaAI-Lab:mainfrom
sethforprivacy:pr-kv-eval-flush-seed
Aug 28, 2026
Merged

MiaAI-Lab merged 2 commits into
MiaAI-Lab:mainfrom
sethforprivacy:pr-kv-eval-flush-seed

Conversation

@sethforprivacy

Copy link
Copy Markdown
Contributor

Two small independent changes to evals/nvfp4_kv_eval.py. Both are about making
the verdict mean what it says; neither changes what a case asks the model.


1. --seed did not change the prompt

filler_line(i) is a pure function of the line index:

def filler_line(i: int) -> str:
    adj = ADJECTIVES[i % len(ADJECTIVES)]
    noun = NOUNS[i % len(NOUNS)]
    serial = hashlib.md5(str(i).encode()).hexdigest()[:8]

and haystack() builds from filler_line(start + i). The only thing the rng
touches is the passkey. So --seed moves the needle and leaves the haystack
byte-identical
— and the haystack is the part attention has to traverse.

Measured on main, three seeds, same depth and position, hashing the prompt with
the needle line removed:

seed 1: needle A6J9-ZWYS-F8Z3   whole-prompt sha c8e589878cb88aaa   filler-only sha 88c7ccc87256c01d
seed 2: needle 577R-CMJF-4CVT   whole-prompt sha e003e87657a35eb2   filler-only sha 88c7ccc87256c01d
seed 7: needle NBT5-68R5-F47V   whole-prompt sha 369b0bd504e88504   filler-only sha 88c7ccc87256c01d

Fix: salt the filler from the run's own rng. Same seed still reproduces the
same prompt exactly — that property is worth keeping — but two seeds now share no
byte of haystack. After:

seed 1: salt 91b7584a2265b1f5  filler-only sha 4c1cc4152f7bcd63
seed 2: salt dcf4bb99f4bea973  filler-only sha 02852fb4c87987d7
seed 7: salt f2a74de452e6b438  filler-only sha f2ceb0ef556505e6
seed 1: salt 91b7584a2265b1f5  filler-only sha 4c1cc4152f7bcd63   <- reproducible

The line shape is unchanged (Record %06d: crate <adj>-<noun> serial <8 hex> …),
so the token-per-line calibration behaves the same. The salt is printed in the
header and written to the JSON report, so a run stays auditable.


2. Nothing flushed the prefix cache between cases

POST /flush_cache is never called, and no case records whether the cache was
cold. With the radix cache on, a case whose prompt shares a prefix with an
earlier one — or a re-run of the same suite — can be answered without
prefilling the haystack at all
, which is the thing the eval exists to measure.

How large the effect is, measured on a live 2× GB10 TP2 pair at native context,
NVFP4 KV, one 16,448-token prompt sent three times:

flush -> Cache flushed.
1 COLD (just flushed)                prompt=16448 tok    6.00s    2,743 tok/s  cache_hit_rate=0.0
2 WARM (identical prompt, no flush)  prompt=16448 tok    0.23s   72,008 tok/s  cache_hit_rate=0.996
flush -> Cache flushed.
3 COLD AGAIN (after flush)           prompt=16448 tok    5.40s    3,047 tok/s  cache_hit_rate=0.0

26× faster and a 99.6% hit rate — request 2 did not traverse the haystack in
any meaningful sense. Request 3 shows the flush restoring cold behaviour, so this
is causal, not drift.

That is a large enough effect that I think it is worth being explicit about what
it means for the result already in the CHANGELOG:

Quick suite on this cluster (2026-08-27): RELIABLE through 128k as a single
huge prompt
(0/50/100%). 64k and 128k NIAH were 3/3 at every position.

I am not suggesting that result is wrong — it reproduced independently on my
own cluster, which is why I trust the eval enough to send patches for it. The
narrower point is that the script does not flush and does not log cache state, so
for any individual case in that run there is no longer a way to tell from the
record whether it was served cold or served from a prefix. That is unknowable
after the fact, and it need not be: one POST per case makes the same verdict
verifiable instead of merely reproducible.

Fix: POST /flush_cache before every case; --no-flush keeps the old
behaviour; per-case cache_flushed and an aggregate cache_flush block go into
the JSON report and the summary line.

One deliberate exception, commented in place: the radix follow-up case flushes
before turn 1 and never between its turns.
Retrieval out of a cached prefix is
exactly what that case is for, so flushing mid-case would delete the thing under
test.


Validation

python3 -c "import ast; ast.parse(...)" and py_compile clean; --help
renders. The seed hashes above are from importing both versions of the module
side by side and comparing.

Live, against a keyed 2× GB10 TP2 pair — native 262,144 context,
kv_cache_dtype=nvfp4, pool 1,907,264 tokens — four quick runs back to back:

run args prompts flushes verdict
A --seed 1 seed-1 haystack 11/11 RELIABLE-RETRIEVAL (niah 7/7)
B --seed 1 --no-flush identical to A 0 RELIABLE-RETRIEVAL (niah 7/7)
C --seed 1 identical to A 11/11 DEGRADED (niah 4/7)
D --seed 2 different haystack 11/11 UNRELIABLE (niah 1/7)

Seed 2 lands on different prompt sizes, which is the change in §1 working
end-to-end rather than only in a unit test: 1104 / 4187 / 4187 / 4191 / 16433 / 16436 / 16436 prompt tokens against seed 1's 1105 / 4156 / 4155 / 4157 / 16315 / 16315 / 16316.

Header and summary lines, from run A:

  seed      1  (filler salt 91b7584a2265b1f5)
  cache     POST /flush_cache before every case (--no-flush to keep it warm)
  …
  cases    9/11 passed
  cache    flushed before 11 case(s)

and with --no-flush:

  cache     NOT flushed between cases — a case may be answered from the radix cache rather than from attention
  …
  cache    NOT flushed — a case may have been served from the radix cache

One thing I should flag, because it is visible in that table

A and C sent byte-identical prompts with identical flushing and disagreed
(7/7 vs 4/7). Every one of the 16 failing cases across all four runs was the
token-0 ! loop
— !!!!!!… instead of an answer, 16/16, no other failure mode
at all. That is the fault #6 is aimed at, it is intermittent on this hardware,
and it is orthogonal to this PR: it is not introduced by these changes, they do
not fix it, and it fires the same way with and without flushing.

I mention it only because it is the second reason a single run's verdict is not a
measurement on this stack, and it is the same remedy — record the conditions with
the result. If it would help I am happy to file that intermittency separately
with the per-case JSON attached; I did not want to bundle it into a patch about
something else.

No serving process was restarted or reconfigured for any of this; the runs are
ordinary inference load plus one POST /flush_cache per case, and every prompt
was ≤ 16.5k tokens.

…ary the haystack

Two independent gaps in the same file.

1. Nothing flushed the cache between cases, so a later case could be
   answered out of the radix cache rather than out of attention — which is
   the thing the eval exists to measure. POST /flush_cache before every
   case; --no-flush keeps the old behaviour. The radix follow-up case
   flushes before turn 1 only, never between its turns: retrieval from a
   cached prefix is what that case is for.

2. filler_line(i) was a pure function of the line index, so every run at a
   given depth sent a byte-identical haystack and --seed moved only the
   passkey. Salt the filler from the run's rng: same seed still reproduces
   the same prompt, different seeds share no byte of haystack.

The report now records seed, filler salt, and whether the cache was
flushed (per case and in aggregate), so a verdict can be audited after the
fact instead of being unknowable.
@sethforprivacy

Copy link
Copy Markdown
Contributor Author

Independent corroboration for §1 from a different engine, plus one narrow point
about how the two halves of this PR interact.

I hit the same defect class today in an unrelated stdlib NIAH script on a
DeepSeek/vLLM stack — fixed random.seed(1234) feeding both the filler and the
needle, so every run rebuilt a byte-identical prompt. Different engine, different
prompt shape, same failure: the anti-cache rationale in the script's own
docstring was false in practice.

The part worth flagging here: a per-run salt is not sufficient on its own.

In this PR the salt is drawn once in run():

filler_salt = f"{rng.getrandbits(64):016x}"
set_filler_salt(filler_salt)

so filler_line(i) is constant for the whole run. Two seeds share no haystack —
which is the fix, and it works — but two cases within one run still share
leading filler wherever their ranges overlap. §2's flush masks that completely,
which is why the combination is correct. It also means the two items are not
really independent: under --no-flush, or if §2 is dropped as
environment-specific, §1 alone still leaves later cases prefix-matching earlier
ones.

I measured that residual before fixing it, at ~252k tokens, 3 depths (10/50/90%),
on two independent 2×GB10 TP2 pairs:

                queries Δ     hits Δ     depth 10%   depth 50%   depth 90%
cluster A       +757,662     +143,360      199.2s      181.3s      109.0s
cluster B       +757,339     +143,360      192.9s      184.0s      111.0s

Different seeds on each, and >99.9% of the in-window queries were mine. The
hits delta is identical on both — 143,360 — which is what makes it structural
rather than incidental traffic.
Mechanism checks out arithmetically: depth 50%
shares the leading ~10% of filler with depth 10%, depth 90% shares ~50% with
depth 50%, predicting ~151k hits against 143,360 observed. The timings agree
independently — 10% cached → 9% faster, 50% cached → 45% faster.

Retrieval correctness was never affected: the needle and everything after it is
unique per depth, so attention still had to traverse. It was the prefill
timings for the later depths that were flattered, which is the same class of
thing §2's table is about.

Deriving the filler per case rather than per run fixes it without depending on
the flush. After that change, three cold rungs on the same cluster:

753,573 tok   PASS    823.9s   915 tok/s   +753,882 queries / +0 hits
879,032 tok   PASS   1020.0s   862 tok/s   +879,341 queries / +0 hits
1,004,785 tok PASS   1248.6s   805 tok/s +1,005,300 queries / +0 hits

Zero hits at every rung, so coldness is evidence rather than assumption.

None of this is a request to change what you have — with §2 in, the run-level
salt is fine, and I would take this PR as-is. It is an argument for not letting
§1 land alone, or for moving set_filler_salt inside the case loop so §1 is
robust by itself. Happy to send that as a follow-up patch if it is useful; it is
a couple of lines and I did not want to bolt it onto a PR that is already
carrying two changes.

@MiaAI-Lab

Copy link
Copy Markdown
Owner

Reviewed and merging.

Both halves belong together and should not be split:

  • --seed previously moved only the needle; the haystack was a pure function of line index, so every run at a given depth sent a byte-identical prompt.
  • Nothing called POST /flush_cache, so a later case (or a re-run) can be answered from the radix cache. The 26× / 99.6% hit-rate measurement on an identical 16,448-token prompt is the reason this eval exists.

Default flush-before-every-case is correct. --no-flush keeps the old behavior. The radix follow-up case flushing before turn 1 only, never between turns, is the right exception.

On the later comment: a per-run salt still shares leading filler across cases inside one run. With the flush on (the default), that is masked. A per-case salt would make --no-flush honest too — optional follow-up, not required to land this.

Eval-only; serving is untouched. Expect colder, slower, possibly worse suite numbers than the CHANGELOG “RELIABLE through 128k” snapshot. The intermittent ! loop is #6, not this PR.

@MiaAI-Lab
MiaAI-Lab merged commit 344f9d0 into MiaAI-Lab:main Aug 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants