server: spec-boundary cache salvage + reasoning-budget round-trip fixes - #69
Conversation
With speculative decoding active the prompt cache was only accepted on an exact token-prefix match (lcp == cached_tokens). Any client-side prompt normalization (e.g. re-sending a conversation without reasoning_content, or trailing-whitespace trimmed by a chat template) diverges the prefix at the tail and forced a full cold reprocessing plus checkpoint invalidation. Both the target and draft MTP contexts report bounded partial sequence removal (RS), the same primitive already used by the context-shift path. On a diverging boundary, roll back the target and draft memory to the common prefix, drop the stale speculative state and reprocess only the remaining tokens. If either memory cannot roll back, keep the previous cold-fallback behavior. server_prompt_cache::load() now also accepts entries with a shorter common prefix when both contexts support trailing removal and the delta fits the RS snapshot bound. Co-Authored-By: Claude <noreply@anthropic.com>
v2: direct llama_memory_seq_rm on the draft context corrupts the MTP pairing state (pending_h rows live outside the context memory), which crashes common_speculative_process with 'missing MTP boundary'. Capture the speculative-impl state together with each context checkpoint (create_checkpoint runs before the pending batch is decoded, so the MTP boundary rows sit exactly at pos_max and stay coherent with the restored target/draft state). On a diverging cache boundary, restore the newest checkpoint that stays within the common prefix and carries a coherent speculative state, then reprocess only the tokens after it. load() accepts entries with a shorter common prefix when such a checkpoint exists (slot and RAM states; disk states keep the exact rule because they restore without live checkpoints). Co-Authored-By: Claude <noreply@anthropic.com>
v3: the checkpoint-based salvage (v2) restored PARTIAL_ONLY state that does not carry the attention KV of hybrid models, so resuming mid-sequence read stale memory and produced nondeterministic output (reproduced: two identical rollback sequences in one container gave different completions). Instead extend the MTP speculative state with a ring of the last RING_N boundary h-rows (serialized in the state blob, format v3): after a bounded trailing llama_memory_seq_rm on target and draft (the same primitive used by verification rollback), rewind pending_h to the last kept position via the ring and reprocess only the remaining tokens. Entries whose divergence exceeds the RS snapshot bound or the ring keep falling back to cold reprocessing as before. Co-Authored-By: Claude <noreply@anthropic.com>
Boundary-only pushes left gaps at the positions skipped by multi-token acceptance rounds, so rollback_state could not find the row at the common prefix even within the RS bound. Push every row at its contiguous position instead. Co-Authored-By: Claude <noreply@anthropic.com>
Chat templates (e.g. Qwen3.8) re-render an assistant turn passed back as '<think>\n' + reasoning_content + '\n</think>\n\n'. When the reasoning budget is exhausted, the forced sequence was just message + end_tag, missing the newline the template inserts before </think>: a truncated turn therefore diverged on client resend at exactly the reasoning boundary, cascading into a cold fallback (spec-boundary-mismatch). Include the leading '\n' in the forced sequence in both the server and CLI call sites so the emitted text matches the re-rendered template. Co-Authored-By: Claude <noreply@anthropic.com>
…t alias pi-ai sends the per-request reasoning budget as 'thinking_token_budget' (vLLM field name) via its openai-completions streamFn; the fork only read 'thinking_budget_tokens', so the client-side thinking level had no matching hard cap and the server default always applied. Fall back to the vLLM name before the server default. Exact name and default behavior unchanged. Co-Authored-By: Claude <noreply@anthropic.com>
When a resent prompt diverges from the cached prefix further back than the speculative boundary ring can rewind, the trailing-rollback path used to fall back to a full cold reprocess of the whole prompt — on a 44k-token agent conversation that is minutes of prefill for a few hundred tokens of divergence. The target and draft memories were already truncated successfully at that point; what cannot be rewound on recurrent architectures is the memory state itself, which only a saved snapshot can restore. Restore the newest context checkpoint at or below the divergence point (same primitive the SWA path uses) and replay from there, so only the tokens after the checkpoint are reprocessed. Boundary bookkeeping is reset and recaptured from the replayed batches. Cold reprocessing remains the fallback when no usable checkpoint exists or the restore fails. Co-Authored-By: Claude <noreply@anthropic.com>
|
Update: added commit 3 — checkpoint-based cache salvage for older divergences The trailing rollback was bounded by the speculative boundary ring, so a resend that diverges further back than the retained boundary window (e.g. a reasoning drop or client-side compaction on a long agent conversation) still paid a full cold reprocess of the whole prompt — minutes of prefill on a 45k-token context, reproducing the "forcing full prompt re-processing" symptom from upstream #21831. New behavior: when the ring cannot rewind, restore the newest context checkpoint at or below the divergence point (target + draft + speculative-impl boundary state, now snapshotted together in Verified on Qwen3.8-27B (hybrid recurrent) + MTP n6, two full scripted runs:
The boundary-state snapshot is the key detail: restoring target+draft KV without the MTP boundary row at the restore position makes every replay batch fail its draft sync ("missing MTP boundary") and eventually aborts the server on empty batches — the state must travel with the checkpoint. |
Summary
Running Qwen3.8-27B (Q4_0_ROCMFP4_STRIX_LEAN) as a coding-agent backend with MTP
speculative decoding (n-max 6) and thinking enabled, every multi-turn conversation
with a client that resends
reasoning_content(OpenAI-style agents such as pi 0.84+)degrades the prompt cache to a cold fallback at each turn. This branch makes the
trailing-rollback of the spec boundary reliable and makes
--reasoning-budgetusable as a hard thinking cap without breaking the cache.
What's in the series
prompt diverges from the cached prefix only in trailing whitespace/merged-token
deltas at the spec boundary, roll back to the last checkpoint instead of a full
cold fallback.
rollback above is O(delta) instead of O(context).
MTP-verified tokens are handled uniformly.
forced sequence was
message + end_tag; chat templates (Qwen3.8 and friends)re-render a resent assistant turn as
<think>\n + reasoning_content + \n</think>\n\n, so the missing leading\nmade every budget-truncated turndiverge on resend at exactly the reasoning boundary. The forced sequence now
includes it (server + CLI call sites).
thinking_token_budgetalias — accepts the vLLM field name for theper-request reasoning budget, used by pi-ai's openai-completions streamFn.
Exact name
thinking_budget_tokensand server default unchanged.Test results (Qwen3.8-27B STRIX_LEAN, container, port 1235, MTP n6)
turn (content + reasoning_content) → 0 cold fallback, 0 trailing rollback
(exact cache hit at token level).
--reasoning-budget 2048): 3 budgetexhaustions + forced closures, 6 natural — 0 cold fallback.
Known limitations (disclosed)
interleaving (documented in our test plan); the common agent flows are stable.
(the truncated turn is not round-trippable) — the checkpoint rollback does not
currently cover client-truncated turns.
Repro/scripts
Test scripts available in the PR discussion on request (deterministic
round-trip, warm-vs-cold, per-request alias checks).