Flush dropped reasoning at stream end when stream_reasoning=False - #32225
Merged
JustinTong0323 merged 3 commits intoJul 31, 2026
Merged
JustinTong0323 merged 3 commits into
JustinTong0323 merged 3 commits into
Conversation
BaseReasoningFormatDetector.finish() only flushed buffered reasoning for force_nonempty_content. With stream_reasoning=False the whole thinking block stays in _buffer until </think>; a stream cut short (e.g. max_tokens) before the end token dropped it entirely. Flush it as reasoning on finish, matching the non-streaming detect_and_parse path and the InklingDetector fix in sgl-project#31787. CohereCommand4Detector keeps _in_reasoning pinned True and tracks phase via _reasoning_done, so it gets its own finish() that flushes the buffered tail to the bucket its phase implies (reasoning while thinking, normal text after) instead of letting the base heuristic misfile a truncated answer as reasoning.
Contributor
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
Collaborator
Author
|
/tag-and-rerun-ci |
Collaborator
Author
|
/rerun-failed-ci |
1 similar comment
Collaborator
Author
|
/rerun-failed-ci |
Collaborator
Author
|
/rerun-failed-ci |
b8zhong
approved these changes
Jul 31, 2026
saturn-acc
pushed a commit
to saturn-acc/sglang
that referenced
this pull request
Aug 16, 2026
Atituiset
pushed a commit
to Atituiset/sglang
that referenced
this pull request
Sep 10, 2026
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Sep 26, 2026
…cks, chronologisch) Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum (Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity; 479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 -> S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name (Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)). Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859 sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579 sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438 (SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild), #DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release; P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6 + F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49 Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens). Auflösungen (Gabel -> Form, Grund): - L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität" nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker- Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags. Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity. - RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser CalibrationIdentity.d_max_running_requests; nextflash unverändert. - URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie). - scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP komponiert (beide Instrumente getrennt schaltbar). - schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF computed_input_len. - gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab. - flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt). - checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation. - xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen. - vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read. - FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9). - argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only). - census_checkpoint_decision (W161) jetzt für beide Profile aktiv. Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu. Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Follow-up to #31787, which fixed dropped reasoning at stream end for
InklingDetector. As noted in the review of that PR (#31787 (comment)), the base classBaseReasoningFormatDetectorhas the same bug.With
stream_reasoning=False,parse_streaming_incrementaccumulates the whole thinking block inself._bufferand only emits it when the end token (</think>) arrives. If generation is cut short before the end token (e.g.max_tokensmid-thinking), the stream ends with the trace still stuck in_buffer.finish()only flushed it whenforce_nonempty_content=True; otherwise the buffered reasoning was dropped entirely.This affects every detector that uses the base streaming path —
DeepSeekR1Detector,Qwen3Detector(and its subclasses),Glm45Detector,KimiK2Detector,MistralDetector,Nemotron3Detector,MiniMaxM3Detector, etc. — for streaming requests withstream_reasoning=False(the value the scheduler and responses API use).Modifications
BaseReasoningFormatDetector.finish(): flush the buffered trace when the stream ends mid-reasoning.force_nonempty_contentstill reclassifies it as normal text (unchanged); otherwise it is now emitted asreasoning_textinstead of being dropped, matching the non-streamingdetect_and_parsepath. The leading think-start token is stripped in both cases (extracted into a small_strip_leading_think_starthelper).CohereCommand4Detector.finish()(new override): this detector keeps_in_reasoningpinnedTruefor its whole run and tracks phase via_reasoning_done, so the base heuristic (keyed on_in_reasoning) would misfile a truncated answer tail as reasoning. It now flushes the buffered tail to the bucket its phase implies — reasoning while still thinking, normal text once the answer block has started — which also recovers the answer tail when a stream stops without<|END_TEXT|>.I audited the other detectors that override the streaming state machine:
InklingDetectoralready overridesfinish()(Fix dropped Inkling reasoning at stream end #31787) — unaffected.GptOssDetector/MiniMaxAppendThinkDetectornever writeself._buffer, so the base flush is a no-op for them.Apertus2509Detectormaintains_in_reasoninghonestly, so the base flush correctly recovers its buffered reasoning.Accuracy Tests
N/A — pure Python parser change, no model forward / kernel code touched.
Speed Tests and Profiling
N/A.
Tests
Added regression cases to
test/registered/unit/parser/test_reasoning_parser.py:finish()withstream_reasoning=False(and strips the opening<think>);stream_reasoning=Trueonly drops a lingering partial end-tag fragment, never re-emits streamed reasoning;DeepSeekR1Detectorforced reasoning (no<think>start token) flushes on truncation;CohereCommand4Detectorflushes a truncated thinking block as reasoning, and a truncated answer tail as normal text.Verified each new case fails on the pre-fix code and passes after the fix; the pre-existing
force_nonempty_contentcases are unchanged.Checklist
CI States
Latest PR Test (Base): ✅ Run #30155713666
Latest PR Test (Extra): ❌ Run #30155713585