Conversation
added 2 commits
August 17, 2026 09:09
The generic copy-through of remaining body keys skips keys already present in llama_params, so the message set from the CLI option silently shadowed any per-request value. Read the body field explicitly with the CLI value as default, mirroring what the budget alias patch does for the token count.
Chat templates re-render a resent assistant turn as '<think>\n' + reasoning_content|trim + '\n</think>\n\n'. The trim strips trailing whitespace from the extracted reasoning, so a forced tail of '\n' + message + end_tag makes the cache diverge at the message/end-tag boundary (the resend re-adds a '\n' the generation never emitted) and the resend pays a checkpoint rollback instead of an exact cache hit. Emit '\n' + message + '\n' + end_tag when a wrap-up message is set; keep the bare '\n' + end_tag when empty, which is the exact-hit case already verified in production.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #69 — completes the reasoning-budget round-trip when a wrap-up message is configured. Two commits, both running in production since Aug 16 on Qwen3.8-27B (hybrid recurrent) + MTP n-6.
1.
reasoning_budget_messageper-request on the OAI chat pathThe generic copy-through of remaining body keys in
oaicompat_chat_params_parse(for (const auto & item : body.items())) skips keys already present inllama_params, so the CLI value pre-populatingreasoning_budget_messagesilently shadowed any per-request one — the body field was ignored on/v1/chat/completions(it worked on/completion, which reads it directly). Read the body field explicitly with the CLI value as default, mirroring what thethinking_token_budgetalias handling does for the budget count.2. Separate newline between wrap-up message and end tag
Chat templates re-render a resent assistant turn as
'<think>\n' + reasoning_content|trim + '\n</think>\n\n'(verified against the Qwen3.8 chat template extracted from the GGUF). The|trimstrips the trailing newline from the extracted reasoning, so the forced tail'\n' + message + end_tagfrom #69 diverges at the message/end-tag boundary: the re-render adds a'\n'the generation never emitted, and the resend pays a checkpoint rollback instead of an exact hit.Fix: emit
'\n' + message + '\n' + end_tagwhen a wrap-up message is set; keep the bare'\n' + end_tagwhen empty — that is the exact-hit case already verified in #69 and it stays untouched.Verification (two scripted runs on the production image)
budget exhausted → forcing → forced completewith the per-request message active, all clean (27/27 marker-matched runs).Note
Default remains no message (
--reasoning-budget-message, default: none). An explicitly configured message should not carry trailing whitespace: the template|trimwould strip it and the resend would pay a bounded rollback (covered by the checkpoint salvage from #69).🤖 Generated with Claude Code