Skip to content

server: spec-boundary cache salvage + reasoning-budget round-trip fixes - #69

Merged
charlie12345 merged 7 commits into
charlie12345:mainfrom
pugant:spec-cache-trailing-rollback
Aug 16, 2026
Merged

charlie12345 merged 7 commits into
charlie12345:mainfrom
pugant:spec-cache-trailing-rollback

Conversation

@pugant

@pugant pugant commented Aug 15, 2026

Copy link
Copy Markdown

Summary

Running Qwen3.8-27B (Q4_0_ROCMFP4_STRIX_LEAN) as a coding-agent backend with MTP
speculative decoding (n-max 6) and thinking enabled, every multi-turn conversation
with a client that resends reasoning_content (OpenAI-style agents such as pi 0.84+)
degrades the prompt cache to a cold fallback at each turn. This branch makes the
trailing-rollback of the spec boundary reliable and makes --reasoning-budget
usable as a hard thinking cap without breaking the cache.

What's in the series

  1. Bounded trailing rollback for spec-boundary cache salvage — when the resent
    prompt diverges from the cached prefix only in trailing whitespace/merged-token
    deltas at the spec boundary, roll back to the last checkpoint instead of a full
    cold fallback.
  2. Checkpoint-based rollback — keeps KV checkpoints at spec boundaries so the
    rollback above is O(delta) instead of O(context).
  3. MTP boundary ring — pushes all batch/verify rows into the rollback ring so
    MTP-verified tokens are handled uniformly.
  4. Forced-end round-trip fix — when the reasoning budget is exhausted, the
    forced sequence was message + end_tag; chat templates (Qwen3.8 and friends)
    re-render a resent assistant turn as <think>\n + reasoning_content + \n</think>\n\n, so the missing leading \n made every budget-truncated turn
    diverge on resend at exactly the reasoning boundary. The forced sequence now
    includes it (server + CLI call sites).
  5. thinking_token_budget alias — accepts the vLLM field name for the
    per-request reasoning budget, used by pi-ai's openai-completions streamFn.
    Exact name thinking_budget_tokens and server default unchanged.

Test results (Qwen3.8-27B STRIX_LEAN, container, port 1235, MTP n6)

  • Budget round-trip: 3 independent runs — resend verbatim of a budget-truncated
    turn (content + reasoning_content) → 0 cold fallback, 0 trailing rollback
    (exact cache hit at token level).
  • No-budget regression: natural thinking closure + resend → 0 cold fallback.
  • Agent workload (pi, 12 requests, --reasoning-budget 2048): 3 budget
    exhaustions + forced closures, 6 natural — 0 cold fallback.

Known limitations (disclosed)

  • The trailing-rollback path has a known flaky edge case under a specific
    interleaving (documented in our test plan); the common agent flows are stable.
  • A manual stop/resume of the client mid-generation still produces a cache skip
    (the truncated turn is not round-trippable) — the checkpoint rollback does not
    currently cover client-truncated turns.

Repro/scripts

Test scripts available in the PR discussion on request (deterministic
round-trip, warm-vs-cold, per-request alias checks).

pugant and others added 6 commits August 15, 2026 16:24
With speculative decoding active the prompt cache was only accepted on an
exact token-prefix match (lcp == cached_tokens). Any client-side prompt
normalization (e.g. re-sending a conversation without reasoning_content,
or trailing-whitespace trimmed by a chat template) diverges the prefix at
the tail and forced a full cold reprocessing plus checkpoint invalidation.

Both the target and draft MTP contexts report bounded partial sequence
removal (RS), the same primitive already used by the context-shift path.
On a diverging boundary, roll back the target and draft memory to the
common prefix, drop the stale speculative state and reprocess only the
remaining tokens. If either memory cannot roll back, keep the previous
cold-fallback behavior.

server_prompt_cache::load() now also accepts entries with a shorter
common prefix when both contexts support trailing removal and the delta
fits the RS snapshot bound.

Co-Authored-By: Claude <noreply@anthropic.com>
v2: direct llama_memory_seq_rm on the draft context corrupts the MTP
pairing state (pending_h rows live outside the context memory), which
crashes common_speculative_process with 'missing MTP boundary'.

Capture the speculative-impl state together with each context checkpoint
(create_checkpoint runs before the pending batch is decoded, so the MTP
boundary rows sit exactly at pos_max and stay coherent with the restored
target/draft state). On a diverging cache boundary, restore the newest
checkpoint that stays within the common prefix and carries a coherent
speculative state, then reprocess only the tokens after it.

load() accepts entries with a shorter common prefix when such a
checkpoint exists (slot and RAM states; disk states keep the exact rule
because they restore without live checkpoints).

Co-Authored-By: Claude <noreply@anthropic.com>
v3: the checkpoint-based salvage (v2) restored PARTIAL_ONLY state that
does not carry the attention KV of hybrid models, so resuming
mid-sequence read stale memory and produced nondeterministic output
(reproduced: two identical rollback sequences in one container gave
different completions).

Instead extend the MTP speculative state with a ring of the last RING_N
boundary h-rows (serialized in the state blob, format v3): after a
bounded trailing llama_memory_seq_rm on target and draft (the same
primitive used by verification rollback), rewind pending_h to the last
kept position via the ring and reprocess only the remaining tokens.
Entries whose divergence exceeds the RS snapshot bound or the ring keep
falling back to cold reprocessing as before.

Co-Authored-By: Claude <noreply@anthropic.com>
Boundary-only pushes left gaps at the positions skipped by multi-token
acceptance rounds, so rollback_state could not find the row at the
common prefix even within the RS bound. Push every row at its contiguous
position instead.

Co-Authored-By: Claude <noreply@anthropic.com>
Chat templates (e.g. Qwen3.8) re-render an assistant turn passed back as
'<think>\n' + reasoning_content + '\n</think>\n\n'. When the reasoning
budget is exhausted, the forced sequence was just message + end_tag,
missing the newline the template inserts before </think>: a truncated
turn therefore diverged on client resend at exactly the reasoning
boundary, cascading into a cold fallback (spec-boundary-mismatch).

Include the leading '\n' in the forced sequence in both the server and
CLI call sites so the emitted text matches the re-rendered template.

Co-Authored-By: Claude <noreply@anthropic.com>
…t alias

pi-ai sends the per-request reasoning budget as 'thinking_token_budget'
(vLLM field name) via its openai-completions streamFn; the fork only
read 'thinking_budget_tokens', so the client-side thinking level had no
matching hard cap and the server default always applied.

Fall back to the vLLM name before the server default. Exact name and
default behavior unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>
When a resent prompt diverges from the cached prefix further back than
the speculative boundary ring can rewind, the trailing-rollback path
used to fall back to a full cold reprocess of the whole prompt — on a
44k-token agent conversation that is minutes of prefill for a few
hundred tokens of divergence.

The target and draft memories were already truncated successfully at
that point; what cannot be rewound on recurrent architectures is the
memory state itself, which only a saved snapshot can restore. Restore
the newest context checkpoint at or below the divergence point (same
primitive the SWA path uses) and replay from there, so only the tokens
after the checkpoint are reprocessed. Boundary bookkeeping is reset and
recaptured from the replayed batches. Cold reprocessing remains the
fallback when no usable checkpoint exists or the restore fails.

Co-Authored-By: Claude <noreply@anthropic.com>
@pugant

pugant commented Aug 15, 2026

Copy link
Copy Markdown
Author

Update: added commit 3 — checkpoint-based cache salvage for older divergences

The trailing rollback was bounded by the speculative boundary ring, so a resend that diverges further back than the retained boundary window (e.g. a reasoning drop or client-side compaction on a long agent conversation) still paid a full cold reprocess of the whole prompt — minutes of prefill on a 45k-token context, reproducing the "forcing full prompt re-processing" symptom from upstream #21831.

New behavior: when the ring cannot rewind, restore the newest context checkpoint at or below the divergence point (target + draft + speculative-impl boundary state, now snapshotted together in common_prompt_checkpoint::data_spec) and replay only the tokens after it.

Verified on Qwen3.8-27B (hybrid recurrent) + MTP n6, two full scripted runs:

  • divergent resend at lcp=41195/cached=45684 → checkpoint rollback at [41191,41191], 4,044 tokens reprocessed instead of 45,236 (~91% saved), generation completes normally
  • altered resends (reasoning dropped / truncated) now go through the checkpoint path instead of cold-reprocessing
  • clean verbatim baseline unchanged (exact hits / small trailing rollbacks), zero aborts

The boundary-state snapshot is the key detail: restoring target+draft KV without the MTP boundary row at the restore position makes every replay batch fail its draft sync ("missing MTP boundary") and eventually aborts the server on empty batches — the state must travel with the checkpoint.

@charlie12345
charlie12345 merged commit ee6479d into charlie12345:main Aug 16, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants