Skip to content

server: per-request reasoning_budget_message and exact cache round-trip with a wrap-up message - #80

Closed
pugant wants to merge 2 commits into
charlie12345:mainfrom
pugant:spec-cache-soft-wrap
Closed

pugant wants to merge 2 commits into
charlie12345:mainfrom
pugant:spec-cache-soft-wrap

Conversation

@pugant

@pugant pugant commented Aug 17, 2026

Copy link
Copy Markdown

Follow-up to #69 — completes the reasoning-budget round-trip when a wrap-up message is configured. Two commits, both running in production since Aug 16 on Qwen3.8-27B (hybrid recurrent) + MTP n-6.

1. reasoning_budget_message per-request on the OAI chat path

The generic copy-through of remaining body keys in oaicompat_chat_params_parse (for (const auto & item : body.items())) skips keys already present in llama_params, so the CLI value pre-populating reasoning_budget_message silently shadowed any per-request one — the body field was ignored on /v1/chat/completions (it worked on /completion, which reads it directly). Read the body field explicitly with the CLI value as default, mirroring what the thinking_token_budget alias handling does for the budget count.

2. Separate newline between wrap-up message and end tag

Chat templates re-render a resent assistant turn as '<think>\n' + reasoning_content|trim + '\n</think>\n\n' (verified against the Qwen3.8 chat template extracted from the GGUF). The |trim strips the trailing newline from the extracted reasoning, so the forced tail '\n' + message + end_tag from #69 diverges at the message/end-tag boundary: the re-render adds a '\n' the generation never emitted, and the resend pays a checkpoint rollback instead of an exact hit.

Fix: emit '\n' + message + '\n' + end_tag when a wrap-up message is set; keep the bare '\n' + end_tag when empty — that is the exact-hit case already verified in #69 and it stays untouched.

Verification (two scripted runs on the production image)

  • Run 1, wiring only (commit 1): the per-request message reaches the forced tail and appears in the extracted reasoning; but a resent budget-truncated turn paid a checkpoint rollback (lcp=280 on cached=534; lcp=1338 on 1398) — divergence exactly at the message→end-tag boundary.
  • Run 2, both commits: regression check with an empty message and wrap-up round-trip with the message set: zero cold fallbacks, zero rollbacks — exact cache hit with the message in the forced tail; multi-turn session with two truncations also clean.
  • Production since Aug 16: 9/10 dense agent tasks hit budget exhausted → forcing → forced complete with the per-request message active, all clean (27/27 marker-matched runs).

Note

Default remains no message (--reasoning-budget-message, default: none). An explicitly configured message should not carry trailing whitespace: the template |trim would strip it and the resend would pay a bounded rollback (covered by the checkpoint salvage from #69).

🤖 Generated with Claude Code

pugant added 2 commits August 17, 2026 09:09
The generic copy-through of remaining body keys skips keys already present
in llama_params, so the message set from the CLI option silently shadowed
any per-request value. Read the body field explicitly with the CLI value
as default, mirroring what the budget alias patch does for the token count.
Chat templates re-render a resent assistant turn as '<think>\n' +
reasoning_content|trim + '\n</think>\n\n'. The trim strips trailing
whitespace from the extracted reasoning, so a forced tail of
'\n' + message + end_tag makes the cache diverge at the message/end-tag
boundary (the resend re-adds a '\n' the generation never emitted) and the
resend pays a checkpoint rollback instead of an exact cache hit.

Emit '\n' + message + '\n' + end_tag when a wrap-up message is set; keep
the bare '\n' + end_tag when empty, which is the exact-hit case already
verified in production.
@pugant pugant closed this by deleting the head repository Aug 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant