feat(grpc): add sglang reasoning token usage - #1747
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughAdds ChangesReasoning Token Support End-to-End
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request adds support for SGLang reasoning tokens across the gRPC client, servicer, and model gateway. It introduces a require_reasoning option for generation requests and propagates reasoning_tokens through streaming chunks and completion responses. Feedback on the changes suggests using getattr when accessing reasoning_tokens on BatchTokenIDOutput in Python to maintain backward compatibility with older SGLang versions and avoid potential runtime crashes.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f7216b6fdc
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
model_gateway/src/routers/grpc/regular/streaming.rs (1)
2362-2365:⚠️ Potential issue | 🟠 Major | ⚖️ Poor tradeoffMissing reasoning_tokens tracking in completion streaming endpoint.
The completion streaming endpoint (
process_completion_streaming_chunks) does not track or exposereasoning_tokens, while the chat streaming (lines 213, 494, 567, 573-574) and generate streaming (lines 791-792, 821, 981-982, 1025) endpoints both include reasoning token accounting. This creates an inconsistency where users calling/v1/completionswith thinking-enabled models will not receivereasoning_tokensin the usage statistics.Add reasoning token tracking to the completion streaming endpoint to maintain feature parity:
- At line 2365: Add
let mut total_reasoning = 0u32;- At line 2484: Add
total_reasoning = total_reasoning.max(complete.reasoning_tokens());- At line 2622: Chain
.with_reasoning_tokens(total_reasoning)after.with_cached_tokens(total_cached)Also applies to: 2480-2484, 2613-2623
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@model_gateway/src/routers/grpc/regular/streaming.rs` around lines 2362 - 2365, The process_completion_streaming_chunks function is missing reasoning_tokens tracking that exists in other streaming endpoints (chat streaming and generate streaming). Add reasoning token tracking by: (1) initializing a new mutable variable `total_reasoning` as 0u32 alongside the existing `total_prompt` and `total_cached` variable declarations, (2) accumulating reasoning tokens from each completion chunk using `total_reasoning.max(complete.reasoning_tokens())` in the same section where `total_cached` is being accumulated, and (3) chaining `.with_reasoning_tokens(total_reasoning)` onto the response builder after the existing `.with_cached_tokens(total_cached)` call to expose reasoning tokens in the usage statistics.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@model_gateway/src/routers/grpc/regular/streaming.rs`:
- Around line 2362-2365: The process_completion_streaming_chunks function is
missing reasoning_tokens tracking that exists in other streaming endpoints (chat
streaming and generate streaming). Add reasoning token tracking by: (1)
initializing a new mutable variable `total_reasoning` as 0u32 alongside the
existing `total_prompt` and `total_cached` variable declarations, (2)
accumulating reasoning tokens from each completion chunk using
`total_reasoning.max(complete.reasoning_tokens())` in the same section where
`total_cached` is being accumulated, and (3) chaining
`.with_reasoning_tokens(total_reasoning)` onto the response builder after the
existing `.with_cached_tokens(total_cached)` call to expose reasoning tokens in
the usage statistics.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: cb663fb4-28d5-4017-980a-5e3a21fc6718
📒 Files selected for processing (19)
bindings/golang/src/client.rsbindings/golang/src/policy.rsbindings/golang/src/proto_parse.rscrates/grpc_client/proto/sglang_scheduler.protocrates/grpc_client/python/pyproject.tomlcrates/grpc_client/src/lib.rscrates/grpc_client/src/sglang_scheduler.rscrates/protocols/src/generate.rsgrpc_servicer/pyproject.tomlgrpc_servicer/smg_grpc_servicer/sglang/request_manager.pygrpc_servicer/smg_grpc_servicer/sglang/servicer.pymodel_gateway/src/routers/grpc/client.rsmodel_gateway/src/routers/grpc/common/response_formatting.rsmodel_gateway/src/routers/grpc/harmony/stages/request_building.rsmodel_gateway/src/routers/grpc/proto_wrapper.rsmodel_gateway/src/routers/grpc/regular/processor.rsmodel_gateway/src/routers/grpc/regular/stages/chat/request_building.rsmodel_gateway/src/routers/grpc/regular/stages/messages/request_building.rsmodel_gateway/src/routers/grpc/regular/streaming.rs
|
Hi @Moersity, the DCO sign-off check has failed. All commits must include a To fix existing commits: # Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-leaseTo sign off future commits automatically:
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ea87d900a0
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
ea87d90 to
7caf6fd
Compare
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@bindings/golang/src/client.rs`:
- Around line 220-224: The require_reasoning field is hardcoded to false in two
locations, preventing proper reasoning-token accounting for thinking-enabled
requests. In bindings/golang/src/client.rs at lines 220-224, compute the
require_reasoning value from the incoming chat request and tokenizer using the
same decision logic applied in gateway request stages, then pass the computed
value to SglangGenerateRequestOptions instead of hardcoding false. Apply the
identical computed require_reasoning behavior in bindings/golang/src/policy.rs
at lines 634-638 for the multi-worker request path, replacing the hardcoded
false value with the same computed logic.
In `@crates/grpc_client/src/sglang_scheduler.rs`:
- Around line 226-230: The require_reasoning extraction and propagation logic in
the sglang_scheduler.rs module lacks direct test coverage. Add focused unit
tests that verify: (1) parsing of the require_reasoning boolean from request
JSON body.other, including edge cases where the value is true, false, or a
non-boolean type (should default to false), and (2) the forwarding of
require_reasoning from chat/messages options into the generated proto message.
These tests should exercise the and_then().unwrap_or(false) pattern shown in the
require_reasoning extraction to ensure correct behavior across all input
scenarios.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: f964cea5-4bbc-45dc-b55e-e067d9580f51
📒 Files selected for processing (20)
bindings/golang/src/client.rsbindings/golang/src/grpc_converter.rsbindings/golang/src/policy.rsbindings/golang/src/proto_parse.rscrates/grpc_client/proto/sglang_scheduler.protocrates/grpc_client/python/pyproject.tomlcrates/grpc_client/src/lib.rscrates/grpc_client/src/sglang_scheduler.rscrates/protocols/src/generate.rsgrpc_servicer/pyproject.tomlgrpc_servicer/smg_grpc_servicer/sglang/request_manager.pygrpc_servicer/smg_grpc_servicer/sglang/servicer.pymodel_gateway/src/routers/grpc/client.rsmodel_gateway/src/routers/grpc/common/response_formatting.rsmodel_gateway/src/routers/grpc/harmony/stages/request_building.rsmodel_gateway/src/routers/grpc/proto_wrapper.rsmodel_gateway/src/routers/grpc/regular/processor.rsmodel_gateway/src/routers/grpc/regular/stages/chat/request_building.rsmodel_gateway/src/routers/grpc/regular/stages/messages/request_building.rsmodel_gateway/src/routers/grpc/regular/streaming.rs
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7caf6fdfa6
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
7caf6fd to
412f79f
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 412f79f411
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
|
||
| // Ask SGLang scheduler to count reasoning tokens. | ||
| // Keep false unless the request explicitly enables reasoning/thinking. | ||
| bool require_reasoning = 18; |
There was a problem hiding this comment.
Regenerate Go SDK protos for reasoning fields
The schema adds require_reasoning/reasoning_tokens, but the checked-in pure Go SDK path still uses the stale generated proto under bindings/golang/internal/proto (I checked with rg and there are no RequireReasoning or ReasoningTokens accessors), and bindings/golang/internal/grpc/client_grpc.go builds GenerateRequest/serializes GenerateResponse directly through those types. The fresh path I checked is separate from the fixed Rust FFI call sites, so Go SDK streaming calls still cannot request reasoning and will drop the returned counter before postprocessing; please regenerate/update the Go proto and conversion code along with this schema change.
Useful? React with 👍 / 👎.
412f79f to
404b653
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 404b6535ac
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| Usage::from_counts(prompt_tokens, completion_tokens) | ||
| .with_cached_tokens(complete.cached_tokens), | ||
| .with_cached_tokens(complete.cached_tokens) | ||
| .with_reasoning_tokens(complete.reasoning_tokens), |
There was a problem hiding this comment.
Expose reasoning usage in Go SDK structs
When SGLang returns non-zero complete.reasoning_tokens, this now serializes completion_tokens_details.reasoning_tokens into the JSON chunk, but the Go SDK typed path still unmarshals usage into bindings/golang/client.go's Usage struct, which only has prompt/completion/total fields; CreateChatCompletion and the multi-client path then copy that struct and silently drop the new detail. This affects Go callers using the typed APIs rather than raw RecvJSON; please add matching completion-token detail fields and preserve them while accumulating usage.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Actionable comments posted: 3
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
model_gateway/src/routers/grpc/regular/streaming.rs (1)
2612-2627:⚠️ Potential issue | 🟠 Major | ⚡ Quick winMissing reasoning_tokens in completions streaming usage chunk.
The completions streaming endpoint tracks
total_cachedand includes it in the final usage chunk (line 2622), but does not track or includereasoning_tokens. Chat streaming (lines 567-574) correctly tracks reasoning tokens per index and includes them via.with_reasoning_tokens(total_reasoning). Since both endpoints use the sameUsagestructure fromopenai_protocol::common::Usage, completions should also track and report reasoning tokens.📊 Proposed fix to add reasoning_tokens tracking
Add reasoning token tracking at the top of the function, similar to
total_cached:let mut total_prompt = 0u32; let mut total_cached = 0u32; + let mut total_reasoning = 0u32; let mut total_completion = CompletionTokenTracker::new();Track reasoning tokens when processing Complete messages:
ProtoResponseVariant::Complete(complete) => { let index = complete.index(); total_prompt = total_prompt.max(complete.prompt_tokens()); total_cached = total_cached.max(complete.cached_tokens()); + total_reasoning = total_reasoning.max(complete.reasoning_tokens()); total_completion.record_complete(&complete);Include reasoning tokens in the final usage chunk:
if include_usage { let usage_chunk = CompletionStreamResponse { id: request_id.clone(), object: "text_completion".to_string(), created, choices: vec![], model: model.clone(), system_fingerprint: system_fingerprint.map(String::from), usage: Some( Usage::from_counts(total_prompt, total_completion.total()) .with_cached_tokens(total_cached) + .with_reasoning_tokens(total_reasoning), ), };🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@model_gateway/src/routers/grpc/regular/streaming.rs` around lines 2612 - 2627, The completions streaming endpoint is not tracking reasoning_tokens in the final usage chunk, while the chat streaming endpoint correctly implements this. Add reasoning token tracking to match the pattern used for total_cached: initialize a variable to accumulate reasoning tokens at the top of the function containing the CompletionStreamResponse creation, update this variable when processing Complete messages (similar to how total_cached is updated), and then chain a call to .with_reasoning_tokens(total_reasoning) after the .with_cached_tokens(total_cached) call on the Usage::from_counts() chain to include reasoning tokens in the usage chunk sent via Self::format_completion_sse_into.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@crates/grpc_client/proto/sglang_scheduler.proto`:
- Line 249: Add an inline comment to the reasoning_tokens field in the proto
file to document its purpose. The comment should clarify that this field
represents the count of reasoning or thinking tokens used in the completed
response, similar to how other token count fields are documented in the same
message. Place the comment directly above or inline with the uint32
reasoning_tokens = 12 field definition to improve code maintainability and
readability.
- Line 223: Add an inline comment to the reasoning_tokens field in the proto
file to document that it represents cumulative reasoning/thinking token usage.
Place the comment directly above or next to the field definition to clarify its
semantics and maintain consistency with the documentation style of adjacent
token fields like prompt_tokens, completion_tokens, and cached_tokens.
In `@model_gateway/src/routers/grpc/client.rs`:
- Around line 45-50: The GenerateRequestBuildOptions struct lacks documentation
explaining its purpose and the purpose of its fields. Add a doc comment above
the struct definition explaining that it consolidates request-building
parameters, then add individual doc comments for each field (multimodal_inputs,
tool_constraints, and require_reasoning) describing what each field represents
and how it's used in request building.
---
Outside diff comments:
In `@model_gateway/src/routers/grpc/regular/streaming.rs`:
- Around line 2612-2627: The completions streaming endpoint is not tracking
reasoning_tokens in the final usage chunk, while the chat streaming endpoint
correctly implements this. Add reasoning token tracking to match the pattern
used for total_cached: initialize a variable to accumulate reasoning tokens at
the top of the function containing the CompletionStreamResponse creation, update
this variable when processing Complete messages (similar to how total_cached is
updated), and then chain a call to .with_reasoning_tokens(total_reasoning) after
the .with_cached_tokens(total_cached) call on the Usage::from_counts() chain to
include reasoning tokens in the usage chunk sent via
Self::format_completion_sse_into.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 26d62cbf-1dfc-428c-afdf-a5f7d16aae2d
⛔ Files ignored due to path filters (3)
bindings/golang/internal/proto/common.pb.gois excluded by!**/*.pb.gobindings/golang/internal/proto/sglang_scheduler.pb.gois excluded by!**/*.pb.gobindings/golang/internal/proto/sglang_scheduler_grpc.pb.gois excluded by!**/*.pb.go
📒 Files selected for processing (24)
bindings/golang/internal/ffi/preprocessor.gobindings/golang/internal/grpc/client_grpc.gobindings/golang/src/client.rsbindings/golang/src/grpc_converter.rsbindings/golang/src/policy.rsbindings/golang/src/preprocessor.rsbindings/golang/src/proto_parse.rsbindings/golang/src/utils.rscrates/grpc_client/proto/sglang_scheduler.protocrates/grpc_client/python/pyproject.tomlcrates/grpc_client/src/lib.rscrates/grpc_client/src/sglang_scheduler.rscrates/protocols/src/generate.rsgrpc_servicer/pyproject.tomlgrpc_servicer/smg_grpc_servicer/sglang/request_manager.pygrpc_servicer/smg_grpc_servicer/sglang/servicer.pymodel_gateway/src/routers/grpc/client.rsmodel_gateway/src/routers/grpc/common/response_formatting.rsmodel_gateway/src/routers/grpc/harmony/stages/request_building.rsmodel_gateway/src/routers/grpc/proto_wrapper.rsmodel_gateway/src/routers/grpc/regular/processor.rsmodel_gateway/src/routers/grpc/regular/stages/chat/request_building.rsmodel_gateway/src/routers/grpc/regular/stages/messages/request_building.rsmodel_gateway/src/routers/grpc/regular/streaming.rs
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f2017a3ea5
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| let require_reasoning = ctx.tokenizer_arc().is_some_and(|tokenizer| { | ||
| utils::should_mark_reasoning_started( | ||
| utils::extract_thinking_from_kwargs( | ||
| chat_request.chat_template_kwargs.as_ref(), | ||
| tokenizer.as_ref(), |
There was a problem hiding this comment.
Set require_reasoning for always-reasoning models
For chat requests whose tokenizer reports ThinkingToggle::None because the model always reasons (the enum documents DeepSeek R1 as that case), this predicate returns false, so the SGLang request is sent with require_reasoning=false. Those models can still emit reasoning that the parser separates, but SGLang's scheduler only fills reasoning_tokens when the request bit is true, so usage stays zero for DeepSeek-R1-style chat requests; please cover the analogous Messages/Go call sites that use the same predicate too.
Useful? React with 👍 / 👎.
|
Sorry for the earlier CI noise. I missed running the full pre-commit checks locally before pushing. I’ve fixed the issue, rerun the checks locally, and pushed the updated commit. This should be ready for review again. Thanks for taking another look. |
|
@coderabbitai review |
|
|
||
| // Ask SGLang scheduler to count reasoning tokens. | ||
| // Keep false unless the request explicitly enables reasoning/thinking. | ||
| bool require_reasoning = 18; |
There was a problem hiding this comment.
Is this a requirement for sglang to start count for reasoning tokens? Why do they need such a field instead of turning it on by default? Is it due to performance concerns?
There was a problem hiding this comment.
Yes, if not set, reasoning tokens will always zero, see sglang, This field is describe: Note: For native API, as a work-around, you need to set require_reasoningargument toTrue to ensure the model will think before generating the structured output. It's not required for chat-completion API.
|
Hi @Moersity, this PR has merge conflicts that must be resolved before it can be merged. Please rebase your branch: git fetch origin main
git rebase origin/main
# resolve any conflicts, then:
git push --force-with-lease |
847fe8c to
e1f430d
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e1f430dbb1
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| top_logprobs_num: body.logprobs.unwrap_or(0) as i32, | ||
| return_hidden_states: body.return_hidden_states, | ||
| stream: body.stream, | ||
| require_reasoning: false, |
There was a problem hiding this comment.
Honor require_reasoning on completion requests
For /v1/completions callers that pass the engine-specific require_reasoning: true extra field (the request type preserves unknown fields in CompletionRequest.other, and this path already forwards SGLang-specific constraints such as json_schema), the new proto bit is always sent as false. SGLang only fills reasoning_tokens when this request flag is true, so structured/reasoning completion calls will still report zero reasoning usage even though the generate/chat paths now propagate the flag; please read the boolean from body.other here as the plain generate builder does.
Useful? React with 👍 / 👎.
Signed-off-by: Moersity <lixiang0417.cq@gmail.com>
e1f430d to
02bca87
Compare
|
Hi everyone, this PR is ready for review. Please let me know if there are any additional requirements or improvements I should address to get this approved and merged. Looking forward to your feedback! |
Description
Problem
SGLang gRPC responses did not expose
reasoning_tokensusage through SMG, even when SGLang itself was able to return reasoning token counts.The gRPC path also did not forward the
require_reasoningflag to the SGLang engine, so SGLang would not count reasoning tokens for thinking-enabled requests.Solution
This PR propagates reasoning token usage through the SGLang gRPC stack:
reasoning_tokensto SGLang gRPC stream and completion responses.require_reasoningto SGLang gRPC generate requests.require_reasoningfrom the gateway only when thinking/reasoning is effectively enabled.reasoning_tokensin regular generate/chat responses and streaming usage.0reasoning tokens for unsupported engines.Changes
reasoning_tokensfields to SGLang gRPC proto response messages.require_reasoningto SGLang gRPC generate request messages.require_reasoningin the Python SGLang gRPC servicer.reasoning_tokensfrom SGLang batch outputs and returned them through gRPC responses.Test Plan
Validated against an actual SGLang gRPC-backed SMG deployment with thinking-capable models.
Tested cases:
Validated against an actual SGLang gRPC-backed SMG deployment.
Tested cases:
reasoning_content.usage.completion_tokens_details.reasoning_tokensis present and greater than0.reasoning_tokens.data: {"id":"chatcmpl-019ecf88-94a3-76f2-a7a8-e4551c9fb35d","object":"chat.completion.chunk","created":1781598295,"model":"/tmp/GLM-5.1-FP8","system_fingerprint":"default","choices":[],"usage":{"prompt_tokens":131,"completion_tokens":453,"total_tokens":584,"prompt_tokens_details":{"cached_tokens":128},"completion_tokens_details":{"reasoning_tokens":397}}}require_reasoningis not forwarded astrue.reasoning_tokensis absent or0.reasoning_tokens=0.Checklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit
require_reasoningflag derived from chat content and tokenizer settings, and propagated it through gRPC request building for chat and messages.0.