Skip to content

feat(llm): support root-level thinking_token_budget for chat and responses - #12624

Merged
biswapanda merged 7 commits into
ai-dynamo:mainfrom
flpanbin:support_thinking_token_budget
Sep 19, 2026
Merged

biswapanda merged 7 commits into
ai-dynamo:mainfrom
flpanbin:support_thinking_token_budget

Conversation

@flpanbin

@flpanbin flpanbin commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Overview:

Implement root-level thinking_token_budget support for OpenAI-compatible chat completion and Responses requests, forwarding it to the backend's thinking_token_budget / max_thinking_tokens sampling parameter while preserving the legacy nvext.max_thinking_tokens passthrough.

Details:

  • Added thinking_token_budget to NvCreateChatCompletionRequest and NvCreateResponse.
  • Added get_thinking_token_budget() and implemented it for both request types.
  • Updated the Python frontend vLLM pre/post processor to read the root-level field.
  • Added unit tests for chat completion and Responses conversion.

Where should the reviewer start?

  • lib/llm/src/protocols/openai/chat_completions.rs — new field and OpenAIStopConditionsProvider impl.
  • lib/llm/src/protocols/openai/responses/mod.rs — Responses API request support and conversion.
  • lib/llm/src/protocols/openai.rs — stop-conditions mapping precedence logic.
  • lib/llm/src/protocols/openai/validate.rs — validation hook.
  • components/src/dynamo/frontend/vllm_processor.py — frontend passthrough.
  • tests/frontend/test_vllm_prepost_integration.py — Python integration tests.

Validation

Automated checks

  • cargo fmt --package dynamo-llm
  • cargo check -p dynamo-llm
  • cargo test -p dynamo-llm thinking_token_budget (5 passed)
  • cargo clippy -p dynamo-llm --all-targets
  • python3 -m pytest tests/frontend/test_vllm_prepost_integration.py

Manual E2E (Frontend + vLLM worker, Qwen/Qwen3-0.6B)

  1. Root-level thinking_token_budget is enforced — request with thinking_token_budget: 16 returns reasoning_tokens: 17):
curl -s http://<frontend>:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model "Qwen3-0.6B",
    "messages": [{"role": "user", "content": "What is 37 times 41? Please reason step by step."}],
    "chat_template_kwargs": {"enable_thinking": true},
    "max_tokens": 1000,
    "temperature": 0,
    "thinking_token_budget": 16
  }'

Response:

......
"usage": {
        "prompt_tokens": 24,
        "completion_tokens": 1000,
        "total_tokens": 1024,
        "completion_tokens_details": {
            "reasoning_tokens": 17
        }
    }
  1. Omitted field keeps existing defaults — the same request without thinking_token_budget completes normally with no budget override applied.

Related Issues

⚠️ This section is required. Choose one path below and delete the other.

🔗 This PR is linked to an issue:

🚫 This PR is NOT linked to an issue:

  • Confirmed — no related issue

Open in Devin Review

Summary by CodeRabbit

  • New Features

    • Added support for configuring a thinking token budget directly in chat and response requests.
    • The direct budget setting takes precedence over the legacy configuration when both are provided.
    • The setting is preserved across supported request formats and forwarded to generation controls.
  • Tests

    • Added coverage for propagation, precedence, fallback behavior, and omission of the thinking token budget.

@flpanbin
flpanbin requested a review from a team as a code owner August 4, 2026 05:52
@copy-pr-bot

copy-pr-bot Bot commented Aug 4, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@flpanbin
flpanbin temporarily deployed to external_collaborator August 4, 2026 05:52 — with GitHub Actions Inactive
@flpanbin
flpanbin temporarily deployed to external_collaborator August 4, 2026 05:52 — with GitHub Actions Inactive
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

👋 Hi flpanbin! Thank you for contributing to ai-dynamo/dynamo.

Just a reminder: The NVIDIA Test Github Validation CI runs an essential subset of the testing framework to quickly catch errors.Your PR reviewers may elect to test the changes comprehensively before approving your changes.

🚀

@github-actions github-actions Bot added external-contribution Pull request is from an external contributor frontend `python -m dynamo.frontend` and `dynamo-run in=http|text|grpc` labels Aug 4, 2026
@coderabbitai

coderabbitai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

The change adds a root-level thinking_token_budget to OpenAI-compatible chat and Responses requests. The value is validated, preserved during conversion, resolved against the legacy extension field, and forwarded to vLLM preprocessing.

Changes

Thinking token budget passthrough

Layer / File(s) Summary
Request contract and validation
lib/llm/src/protocols/openai/chat_completions.rs, lib/llm/src/protocols/openai/responses/mod.rs, lib/llm/src/protocols/openai/validate.rs, lib/llm/src/http/service/openai.rs, lib/llm/src/entrypoint/input/text.rs, lib/llm/src/protocols/anthropic/types.rs, lib/llm/src/protocols/unified.rs, lib/llm/src/protocols/openai/chat_completions/delta.rs
Chat and Responses requests define the optional thinking_token_budget field. Validation and default-based test construction support the new field.
Stop-condition budget precedence
lib/llm/src/protocols/openai.rs, lib/llm/src/protocols/openai/chat_completions.rs
Stop-condition extraction prefers the root-level budget and falls back to nvext.max_thinking_tokens. Tests cover both paths and omission.
Responses conversion and preservation
lib/llm/src/protocols/openai/responses/mod.rs
Responses requests expose the budget and forward it to generated chat-completion requests. Tests cover preservation, precedence, and fallback.
vLLM preprocessing forwarding
components/src/dynamo/frontend/vllm_processor.py, tests/frontend/test_vllm_prepost_integration.py
The selected budget is forwarded as stop_conditions.max_thinking_tokens. Integration tests cover forwarding and precedence.

Estimated code review effort: 3 (Moderate) | ~20 minutes

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The PR satisfies most requirements in [#12302], but validation does not enforce backend-supported upper bounds for thinking_token_budget. Add validation for backend-supported upper bounds, or return a clear validation error when the requested budget exceeds them.
✅ Passed checks (4 passed)
Check name Status Explanation
Out of Scope Changes check ✅ Passed The changes remain within the linked issue scope and include only related request handling, conversion, validation, frontend forwarding, and tests.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Title check ✅ Passed The title clearly and concisely describes the main change: adding root-level thinking_token_budget support for chat and Responses requests.
Description check ✅ Passed The description is complete, relevant, and covers the overview, implementation details, review starting points, validation steps, and linked issue. The Related Issues section includes the linked issue…
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch

Comment @coderabbitai help to get the list of available commands.

@flpanbin flpanbin changed the title Support root-level thinking_token_budget for chat and responses feat(llm): support root-level thinking_token_budget for chat and responses Aug 4, 2026

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 3 potential issues.

Open in Devin Review

Comment thread components/src/dynamo/frontend/vllm_processor.py
Comment thread components/src/dynamo/frontend/vllm_processor.py
Comment thread components/src/dynamo/frontend/vllm_processor.py Outdated
@github-actions github-actions Bot added the feat label Aug 4, 2026
@datadog-official

This comment has been minimized.

@flpanbin
flpanbin marked this pull request as draft August 4, 2026 06:01
@flpanbin
flpanbin force-pushed the support_thinking_token_budget branch from a3bbb1b to 762ed10 Compare August 4, 2026 08:14
@flpanbin
flpanbin temporarily deployed to external_collaborator August 4, 2026 08:14 — with GitHub Actions Inactive
@flpanbin
flpanbin force-pushed the support_thinking_token_budget branch from 762ed10 to 8569eda Compare August 4, 2026 08:55
@flpanbin
flpanbin temporarily deployed to external_collaborator August 4, 2026 08:56 — with GitHub Actions Inactive
@flpanbin
flpanbin force-pushed the support_thinking_token_budget branch from 8569eda to 8dad128 Compare August 4, 2026 10:16
@flpanbin
flpanbin temporarily deployed to external_collaborator August 4, 2026 10:16 — with GitHub Actions Inactive
@github-actions github-actions Bot added the backend::vllm Relates to the vllm backend label Aug 4, 2026
@flpanbin
flpanbin marked this pull request as ready for review August 4, 2026 14:20
@flpanbin
flpanbin requested review from a team as code owners August 4, 2026 14:20
@flpanbin
flpanbin force-pushed the support_thinking_token_budget branch from 8dad128 to 4bbb455 Compare August 5, 2026 14:06
@flpanbin
flpanbin temporarily deployed to external_collaborator August 5, 2026 14:07 — with GitHub Actions Inactive
@pskiran1

Copy link
Copy Markdown
Contributor

@flpanbin, please resolve conflicts and rebase with main.

@flpanbin
flpanbin force-pushed the support_thinking_token_budget branch from 4bbb455 to 6a4c7f3 Compare August 27, 2026 06:31
@flpanbin
flpanbin temporarily deployed to external_collaborator September 9, 2026 14:14 — with GitHub Actions Inactive

@biswapanda biswapanda left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review feedback on the current head.

Comment thread lib/llm/src/protocols/openai/responses/mod.rs
Comment thread lib/llm/src/protocols/openai/validate.rs Outdated
Comment thread lib/llm/src/protocols/anthropic/types.rs Outdated
…onses

Signed-off-by: bin <bin.pan@daocloud.io>
@flpanbin
flpanbin force-pushed the support_thinking_token_budget branch from 49b42a7 to b925d6b Compare September 15, 2026 09:16
@flpanbin
flpanbin deployed to external_collaborator September 15, 2026 09:17 — with GitHub Actions Active
…n_budget

Signed-off-by: bin <bin.pan@daocloud.io>
@flpanbin
flpanbin deployed to external_collaborator September 15, 2026 09:41 — with GitHub Actions Active
…ken_budget

Signed-off-by: bin <bin.pan@daocloud.io>
Signed-off-by: bin <bin.pan@daocloud.io>
@flpanbin
flpanbin deployed to external_collaborator September 15, 2026 10:09 — with GitHub Actions Active
@rmccorm4
rmccorm4 requested a review from GuanLuo September 17, 2026 03:51
@biswapanda
biswapanda deployed to external_collaborator September 18, 2026 18:59 — with GitHub Actions Active
@biswapanda

Copy link
Copy Markdown
Contributor

@flpanbin please take a look - some CI tests are failing

Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
@biswapanda
biswapanda deployed to external_collaborator September 18, 2026 19:20 — with GitHub Actions Active
@biswapanda
biswapanda enabled auto-merge (squash) September 18, 2026 19:34
Comment thread lib/llm/src/protocols/unified.rs
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
@biswapanda
biswapanda deployed to external_collaborator September 18, 2026 19:57 — with GitHub Actions Active
@biswapanda
biswapanda disabled auto-merge September 18, 2026 20:10
@biswapanda

Copy link
Copy Markdown
Contributor

/ok to test 152bba5

@biswapanda
biswapanda enabled auto-merge (squash) September 18, 2026 20:12
@biswapanda
biswapanda removed the request for review from GuanLuo September 19, 2026 21:51

This branch was successfully deployed

1 active deployment
external_collaborator — 152bba5f Deployed Sep 18, 2026 by biswapanda via ok-to-test #20268
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend::vllm Relates to the vllm backend external-contribution Pull request is from an external contributor feat frontend `python -m dynamo.frontend` and `dynamo-run in=http|text|grpc` size/L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Pass through thinking_token_budget from OpenAI-compatible requests

4 participants