Repository navigation
fix(kv-eval): flush the prefix cache between cases, and make --seed actually vary the haystack - #16
Conversation
…ary the haystack Two independent gaps in the same file. 1. Nothing flushed the cache between cases, so a later case could be answered out of the radix cache rather than out of attention — which is the thing the eval exists to measure. POST /flush_cache before every case; --no-flush keeps the old behaviour. The radix follow-up case flushes before turn 1 only, never between its turns: retrieval from a cached prefix is what that case is for. 2. filler_line(i) was a pure function of the line index, so every run at a given depth sent a byte-identical haystack and --seed moved only the passkey. Salt the filler from the run's rng: same seed still reproduces the same prompt, different seeds share no byte of haystack. The report now records seed, filler salt, and whether the cache was flushed (per case and in aggregate), so a verdict can be audited after the fact instead of being unknowable.
|
Independent corroboration for §1 from a different engine, plus one narrow point I hit the same defect class today in an unrelated stdlib NIAH script on a The part worth flagging here: a per-run salt is not sufficient on its own. In this PR the salt is drawn once in filler_salt = f"{rng.getrandbits(64):016x}"
set_filler_salt(filler_salt)so I measured that residual before fixing it, at ~252k tokens, 3 depths (10/50/90%), Different seeds on each, and >99.9% of the in-window queries were mine. The Retrieval correctness was never affected: the needle and everything after it is Deriving the filler per case rather than per run fixes it without depending on Zero hits at every rung, so coldness is evidence rather than assumption. None of this is a request to change what you have — with §2 in, the run-level |
|
Reviewed and merging. Both halves belong together and should not be split:
Default flush-before-every-case is correct. On the later comment: a per-run salt still shares leading filler across cases inside one run. With the flush on (the default), that is masked. A per-case salt would make Eval-only; serving is untouched. Expect colder, slower, possibly worse suite numbers than the CHANGELOG “RELIABLE through 128k” snapshot. The intermittent |
Two small independent changes to
evals/nvfp4_kv_eval.py. Both are about makingthe verdict mean what it says; neither changes what a case asks the model.
1.
--seeddid not change the promptfiller_line(i)is a pure function of the line index:and
haystack()builds fromfiller_line(start + i). The only thing the rngtouches is the passkey. So
--seedmoves the needle and leaves the haystackbyte-identical — and the haystack is the part attention has to traverse.
Measured on
main, three seeds, same depth and position, hashing the prompt withthe needle line removed:
Fix: salt the filler from the run's own rng. Same seed still reproduces the
same prompt exactly — that property is worth keeping — but two seeds now share no
byte of haystack. After:
The line shape is unchanged (
Record %06d: crate <adj>-<noun> serial <8 hex> …),so the token-per-line calibration behaves the same. The salt is printed in the
header and written to the JSON report, so a run stays auditable.
2. Nothing flushed the prefix cache between cases
POST /flush_cacheis never called, and no case records whether the cache wascold. With the radix cache on, a case whose prompt shares a prefix with an
earlier one — or a re-run of the same suite — can be answered without
prefilling the haystack at all, which is the thing the eval exists to measure.
How large the effect is, measured on a live 2× GB10 TP2 pair at native context,
NVFP4 KV, one 16,448-token prompt sent three times:
26× faster and a 99.6% hit rate — request 2 did not traverse the haystack in
any meaningful sense. Request 3 shows the flush restoring cold behaviour, so this
is causal, not drift.
That is a large enough effect that I think it is worth being explicit about what
it means for the result already in the CHANGELOG:
I am not suggesting that result is wrong — it reproduced independently on my
own cluster, which is why I trust the eval enough to send patches for it. The
narrower point is that the script does not flush and does not log cache state, so
for any individual case in that run there is no longer a way to tell from the
record whether it was served cold or served from a prefix. That is unknowable
after the fact, and it need not be: one POST per case makes the same verdict
verifiable instead of merely reproducible.
Fix:
POST /flush_cachebefore every case;--no-flushkeeps the oldbehaviour; per-case
cache_flushedand an aggregatecache_flushblock go intothe JSON report and the summary line.
One deliberate exception, commented in place: the radix follow-up case flushes
before turn 1 and never between its turns. Retrieval out of a cached prefix is
exactly what that case is for, so flushing mid-case would delete the thing under
test.
Validation
python3 -c "import ast; ast.parse(...)"andpy_compileclean;--helprenders. The seed hashes above are from importing both versions of the module
side by side and comparing.
Live, against a keyed 2× GB10 TP2 pair — native 262,144 context,
kv_cache_dtype=nvfp4, pool 1,907,264 tokens — fourquickruns back to back:--seed 1--seed 1 --no-flush--seed 1--seed 2Seed 2 lands on different prompt sizes, which is the change in §1 working
end-to-end rather than only in a unit test:
1104 / 4187 / 4187 / 4191 / 16433 / 16436 / 16436prompt tokens against seed 1's1105 / 4156 / 4155 / 4157 / 16315 / 16315 / 16316.Header and summary lines, from run A:
and with
--no-flush:One thing I should flag, because it is visible in that table
A and C sent byte-identical prompts with identical flushing and disagreed
(7/7 vs 4/7). Every one of the 16 failing cases across all four runs was the
token-0
!loop —!!!!!!…instead of an answer, 16/16, no other failure modeat all. That is the fault #6 is aimed at, it is intermittent on this hardware,
and it is orthogonal to this PR: it is not introduced by these changes, they do
not fix it, and it fires the same way with and without flushing.
I mention it only because it is the second reason a single run's verdict is not a
measurement on this stack, and it is the same remedy — record the conditions with
the result. If it would help I am happy to file that intermittency separately
with the per-case JSON attached; I did not want to bundle it into a patch about
something else.
No serving process was restarted or reconfigured for any of this; the runs are
ordinary inference load plus one
POST /flush_cacheper case, and every promptwas ≤ 16.5k tokens.