Skip to content

Flush dropped reasoning at stream end when stream_reasoning=False - #32225

Merged
JustinTong0323 merged 3 commits into
sgl-project:mainfrom
JustinTong0323:xinyuan/reasoning-finish-flush
Jul 31, 2026
Merged

JustinTong0323 merged 3 commits into
sgl-project:mainfrom
JustinTong0323:xinyuan/reasoning-finish-flush

Conversation

@JustinTong0323

@JustinTong0323 JustinTong0323 commented Jul 23, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Follow-up to #31787, which fixed dropped reasoning at stream end for InklingDetector. As noted in the review of that PR (#31787 (comment)), the base class BaseReasoningFormatDetector has the same bug.

With stream_reasoning=False, parse_streaming_increment accumulates the whole thinking block in self._buffer and only emits it when the end token (</think>) arrives. If generation is cut short before the end token (e.g. max_tokens mid-thinking), the stream ends with the trace still stuck in _buffer. finish() only flushed it when force_nonempty_content=True; otherwise the buffered reasoning was dropped entirely.

This affects every detector that uses the base streaming path — DeepSeekR1Detector, Qwen3Detector (and its subclasses), Glm45Detector, KimiK2Detector, MistralDetector, Nemotron3Detector, MiniMaxM3Detector, etc. — for streaming requests with stream_reasoning=False (the value the scheduler and responses API use).

Modifications

  • BaseReasoningFormatDetector.finish(): flush the buffered trace when the stream ends mid-reasoning. force_nonempty_content still reclassifies it as normal text (unchanged); otherwise it is now emitted as reasoning_text instead of being dropped, matching the non-streaming detect_and_parse path. The leading think-start token is stripped in both cases (extracted into a small _strip_leading_think_start helper).
  • CohereCommand4Detector.finish() (new override): this detector keeps _in_reasoning pinned True for its whole run and tracks phase via _reasoning_done, so the base heuristic (keyed on _in_reasoning) would misfile a truncated answer tail as reasoning. It now flushes the buffered tail to the bucket its phase implies — reasoning while still thinking, normal text once the answer block has started — which also recovers the answer tail when a stream stops without <|END_TEXT|>.

I audited the other detectors that override the streaming state machine:

  • InklingDetector already overrides finish() (Fix dropped Inkling reasoning at stream end #31787) — unaffected.
  • GptOssDetector / MiniMaxAppendThinkDetector never write self._buffer, so the base flush is a no-op for them.
  • Apertus2509Detector maintains _in_reasoning honestly, so the base flush correctly recovers its buffered reasoning.

Accuracy Tests

N/A — pure Python parser change, no model forward / kernel code touched.

Speed Tests and Profiling

N/A.

Tests

Added regression cases to test/registered/unit/parser/test_reasoning_parser.py:

  • base detector flushes truncated reasoning at finish() with stream_reasoning=False (and strips the opening <think>);
  • stream_reasoning=True only drops a lingering partial end-tag fragment, never re-emits streamed reasoning;
  • DeepSeekR1Detector forced reasoning (no <think> start token) flushes on truncation;
  • CohereCommand4Detector flushes a truncated thinking block as reasoning, and a truncated answer tail as normal text.

Verified each new case fails on the pre-fix code and passes after the fix; the pre-existing force_nonempty_content cases are unchanged.

Checklist

  • Format your code according to pre-commit.
  • Add unit tests.

CI States

Latest PR Test (Base): ✅ Run #30155713666
Latest PR Test (Extra): ❌ Run #30155713585

BaseReasoningFormatDetector.finish() only flushed buffered reasoning for
force_nonempty_content. With stream_reasoning=False the whole thinking block
stays in _buffer until </think>; a stream cut short (e.g. max_tokens) before
the end token dropped it entirely. Flush it as reasoning on finish, matching
the non-streaming detect_and_parse path and the InklingDetector fix in sgl-project#31787.

CohereCommand4Detector keeps _in_reasoning pinned True and tracks phase via
_reasoning_done, so it gets its own finish() that flushes the buffered tail to
the bucket its phase implies (reasoning while thinking, normal text after)
instead of letting the base heuristic misfile a truncated answer as reasoning.
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@JustinTong0323

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Jul 24, 2026
@JustinTong0323

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

1 similar comment
@JustinTong0323

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@JustinTong0323

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@JustinTong0323
JustinTong0323 merged commit 68d4429 into sgl-project:main Jul 31, 2026
216 of 254 checks passed
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 26, 2026
…cks, chronologisch)

Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum
(Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst
anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity;
479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 ->
S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name
(Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)).
Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc

Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859
sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579
sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438
(SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool
sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild),
#DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release;
P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6
+ F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49
Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens).

Auflösungen (Gabel -> Form, Grund):
- L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität"
  nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker-
  Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags.
  Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity.
- RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser
  CalibrationIdentity.d_max_running_requests; nextflash unverändert.
- URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und
  Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie).
- scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP
  komponiert (beide Instrumente getrennt schaltbar).
- schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF
  computed_input_len.
- gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab.
- flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt).
- checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation.
- xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen.
- vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read.
- FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9).
- argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only).
- census_checkpoint_decision (W161) jetzt für beide Profile aktiv.

Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber
dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu.
Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants