Fix KeyError on batch requests whose state is freed before it is read - #36638
Merged
Merged
Conversation
mmangkad
requested review from
Ying1123,
hnyls2002,
merrymercy and
xiezhq-hermann
as code owners
August 27, 2026 06:54
Collaborator
Author
|
/rerun-test test/registered/perf/test_vlms_perf.py |
Contributor
|
Results for 🚀 |
2 tasks done
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Sep 24, 2026
… abort dispatched requests on handler failure Upstream f478b2b (sgl-project#35255, tokenizer hunks) and 70088aa (sgl-project#36638), both v0.5.20. sgl-project#35255: generate_request's failure cleanup only dropped the rid states. For a request that had already reached the scheduler the scheduler request ran on with its KV to the end, and a later abort for the rid was swallowed as an unknown rid. On this line the handler caught only Exception, so a client disconnect (CancelledError) skipped the cleanup altogether. Now: - ReqState.dispatched, set by _mark_state_dispatched after the single and the batch dispatch; - except BaseException (as upstream since sgl-project#32588) -> _release_req_states_on_ failure(request_rids): undelivered states are dropped, dispatched requests are aborted and their state kept until the scheduler's answer removes it; - _handle_batch_request adds the rids it mints for n>1 to request_rids. Adapted / not carried: - upstream's ReqState.abort_sent dedupe: it ships with the scheduler-side retry of a deferred chunked abort that moved queues (process_pending_ chunked_abort). That retry is not taken here -- this line's deferred abort runs the weg2 PP countdown (xsn324) and a rank-local re-abort is a rank consistency question; on P the moved case ends with the request finishing its prefill, on D prompts are not chunk-prefilled. Without the retry a second AbortReq (e.g. the front's /abort_request) can be the one that lands, and a repeated AbortReq is a no-op, so no dedupe. - the disconnect task's `if rid in rid_to_state` filter: it would re-open weg2xsn276 (an abort for a rid this manager no longer knows must still reach every weg2 scheduler rank); abort_request keeps that path verbatim. sgl-project#36638: _wait_one_response captures the ReqState when the waiter is built (plain function returning _stream_one_response) instead of when the async generator is first advanced -- batch dispatch builds every waiter before advancing any, and a sub-request finishing in between had its state deleted (KeyError). Tests (CPU, incg cgroup): test/registered/unit/managers/test_tokenizer_manager_rid_cleanup.py, upstream tests ported (release helper incl. dispatched/undelivered split and a failing abort, n>1 minted-rid cleanup, cancel-after-dispatch sends AbortReq and keeps the state, waiter built before finish still delivers): 21 passed; on the base file the 9 new tests fail.
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Sep 24, 2026
…r part, sgl-project#36638) into the B1 staging line Agent B: dispatched requests are aborted on handler failure/disconnect (except BaseException as upstream), the waiter holds the ReqState from construction (no KeyError for batch requests). Not taken with reason: abort_sent dedup, rid filter in the disconnect task (would reopen weg2xsn276). sgl-project#30986 already covered by the fork, sgl-project#35957 unreachable, sgl-project#37143 not applicable (timeouts -1). Tests: test_tokenizer_manager_rid_cleanup 21 passed (hermetic, cgroup 3G).
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Sep 26, 2026
…cks, chronologisch) Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum (Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity; 479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 -> S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name (Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)). Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859 sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579 sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438 (SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild), #DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release; P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6 + F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49 Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens). Auflösungen (Gabel -> Form, Grund): - L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität" nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker- Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags. Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity. - RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser CalibrationIdentity.d_max_running_requests; nextflash unverändert. - URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie). - scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP komponiert (beide Instrumente getrennt schaltbar). - schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF computed_input_len. - gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab. - flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt). - checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation. - xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen. - vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read. - FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9). - argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only). - census_checkpoint_decision (W161) jetzt für beide Profile aktiv. Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu. Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
_wait_one_responsewas anasync defgenerator, so itsrid_to_state[obj.rid]lookupran only when the generator was first advanced. Batch dispatch builds a generator per
request and advances them all after the loop, so a request that finishes meanwhile has
its entry deleted by
_handle_batch_output, and advancing its generator then raisedKeyErrorand failed the whole batch. Multimodal batches hit this every time, which iswhy
test_vlms_perf.pyfails in the nightly.The fix resolves the state eagerly and hands it to the generator; holding the
ReqStatekeeps the output deliverable after the dict entry is gone (
.get()would have droppedcompleted responses silently). All five call sites are unchanged.
TestWaitOneResponseAfterStateFreedfails pre-fix and passes here. On 2xH200 the nightlytest goes 3/3 fail to 3/3 pass.
CI States
Latest PR Test (Base): ✅ Run #33055585814
Latest PR Test (Extra): ❌ Run #33055585564
Latest PR Test (AMD ROCm 7.2): ❌ Run #33055585721