Repository navigation
feat(llm): support root-level thinking_token_budget for chat and responses - #12624
Conversation
|
👋 Hi flpanbin! Thank you for contributing to ai-dynamo/dynamo. Just a reminder: The 🚀 |
WalkthroughThe change adds a root-level ChangesThinking token budget passthrough
Estimated code review effort: 3 (Moderate) | ~20 minutes 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
Comment |
This comment has been minimized.
This comment has been minimized.
a3bbb1b to
762ed10
Compare
762ed10 to
8569eda
Compare
8569eda to
8dad128
Compare
8dad128 to
4bbb455
Compare
|
@flpanbin, please resolve conflicts and rebase with main. |
4bbb455 to
6a4c7f3
Compare
biswapanda
left a comment
There was a problem hiding this comment.
Review feedback on the current head.
…onses Signed-off-by: bin <bin.pan@daocloud.io>
49b42a7 to
b925d6b
Compare
…n_budget Signed-off-by: bin <bin.pan@daocloud.io>
…ken_budget Signed-off-by: bin <bin.pan@daocloud.io>
Signed-off-by: bin <bin.pan@daocloud.io>
|
@flpanbin please take a look - some CI tests are failing |
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
|
/ok to test 152bba5 |
Overview:
Implement root-level
thinking_token_budgetsupport for OpenAI-compatible chat completion and Responses requests, forwarding it to the backend'sthinking_token_budget/max_thinking_tokenssampling parameter while preserving the legacynvext.max_thinking_tokenspassthrough.Details:
thinking_token_budgettoNvCreateChatCompletionRequestandNvCreateResponse.get_thinking_token_budget()and implemented it for both request types.Where should the reviewer start?
lib/llm/src/protocols/openai/chat_completions.rs— new field andOpenAIStopConditionsProviderimpl.lib/llm/src/protocols/openai/responses/mod.rs— Responses API request support and conversion.lib/llm/src/protocols/openai.rs— stop-conditions mapping precedence logic.lib/llm/src/protocols/openai/validate.rs— validation hook.components/src/dynamo/frontend/vllm_processor.py— frontend passthrough.tests/frontend/test_vllm_prepost_integration.py— Python integration tests.Validation
Automated checks
cargo fmt --package dynamo-llmcargo check -p dynamo-llmcargo test -p dynamo-llm thinking_token_budget(5 passed)cargo clippy -p dynamo-llm --all-targetspython3 -m pytest tests/frontend/test_vllm_prepost_integration.pyManual E2E (Frontend + vLLM worker,
Qwen/Qwen3-0.6B)thinking_token_budgetis enforced — request withthinking_token_budget: 16returnsreasoning_tokens: 17):Response:
thinking_token_budgetcompletes normally with no budget override applied.Related Issues
🔗 This PR is linked to an issue:
🚫 This PR is NOT linked to an issue:
Summary by CodeRabbit
New Features
Tests