Skip to content

Fix KeyError on batch requests whose state is freed before it is read - #36638

Merged
hnyls2002 merged 4 commits into
mainfrom
mmangkad/fix-batch-request-state-keyerror
Aug 28, 2026
Merged

hnyls2002 merged 4 commits into
mainfrom
mmangkad/fix-batch-request-state-keyerror

Conversation

@mmangkad

@mmangkad mmangkad commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

_wait_one_response was an async def generator, so its rid_to_state[obj.rid] lookup
ran only when the generator was first advanced. Batch dispatch builds a generator per
request and advances them all after the loop, so a request that finishes meanwhile has
its entry deleted by _handle_batch_output, and advancing its generator then raised
KeyError and failed the whole batch. Multimodal batches hit this every time, which is
why test_vlms_perf.py fails in the nightly.

The fix resolves the state eagerly and hands it to the generator; holding the ReqState
keeps the output deliverable after the dict entry is gone (.get() would have dropped
completed responses silently). All five call sites are unchanged.

TestWaitOneResponseAfterStateFreed fails pre-fix and passes here. On 2xH200 the nightly
test goes 3/3 fail to 3/3 pass.


CI States

Latest PR Test (Base): ✅ Run #33055585814
Latest PR Test (Extra): ❌ Run #33055585564
Latest PR Test (AMD ROCm 7.2): ❌ Run #33055585721

@mmangkad

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/perf/test_vlms_perf.py

@github-actions

github-actions Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/perf/test_vlms_perf.py:

🚀 2-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/perf/test_vlms_perf.py

@mmangkad mmangkad added the run-ci CI: run the baseline test suite on this PR label Aug 27, 2026
@hnyls2002
hnyls2002 merged commit 70088aa into main Aug 28, 2026
158 of 185 checks passed
@hnyls2002
hnyls2002 deleted the mmangkad/fix-batch-request-state-keyerror branch August 28, 2026 19:57
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 24, 2026
… abort dispatched requests on handler failure

Upstream f478b2b (sgl-project#35255, tokenizer hunks) and 70088aa (sgl-project#36638), both
v0.5.20.

sgl-project#35255: generate_request's failure cleanup only dropped the rid states. For a
request that had already reached the scheduler the scheduler request ran on
with its KV to the end, and a later abort for the rid was swallowed as an
unknown rid. On this line the handler caught only Exception, so a client
disconnect (CancelledError) skipped the cleanup altogether. Now:
- ReqState.dispatched, set by _mark_state_dispatched after the single and
  the batch dispatch;
- except BaseException (as upstream since sgl-project#32588) -> _release_req_states_on_
  failure(request_rids): undelivered states are dropped, dispatched requests
  are aborted and their state kept until the scheduler's answer removes it;
- _handle_batch_request adds the rids it mints for n>1 to request_rids.

Adapted / not carried:
- upstream's ReqState.abort_sent dedupe: it ships with the scheduler-side
  retry of a deferred chunked abort that moved queues (process_pending_
  chunked_abort). That retry is not taken here -- this line's deferred abort
  runs the weg2 PP countdown (xsn324) and a rank-local re-abort is a rank
  consistency question; on P the moved case ends with the request finishing
  its prefill, on D prompts are not chunk-prefilled. Without the retry a
  second AbortReq (e.g. the front's /abort_request) can be the one that
  lands, and a repeated AbortReq is a no-op, so no dedupe.
- the disconnect task's `if rid in rid_to_state` filter: it would re-open
  weg2xsn276 (an abort for a rid this manager no longer knows must still
  reach every weg2 scheduler rank); abort_request keeps that path verbatim.

sgl-project#36638: _wait_one_response captures the ReqState when the waiter is built
(plain function returning _stream_one_response) instead of when the async
generator is first advanced -- batch dispatch builds every waiter before
advancing any, and a sub-request finishing in between had its state deleted
(KeyError).

Tests (CPU, incg cgroup):
test/registered/unit/managers/test_tokenizer_manager_rid_cleanup.py, upstream
tests ported (release helper incl. dispatched/undelivered split and a failing
abort, n>1 minted-rid cleanup, cancel-after-dispatch sends AbortReq and keeps
the state, waiter built before finish still delivers): 21 passed; on the base
file the 9 new tests fail.
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 24, 2026
…r part, sgl-project#36638) into the B1 staging line

Agent B: dispatched requests are aborted on handler failure/disconnect
(except BaseException as upstream), the waiter holds the ReqState from
construction (no KeyError for batch requests). Not taken with reason:
abort_sent dedup, rid filter in the disconnect task (would reopen
weg2xsn276). sgl-project#30986 already covered by the fork, sgl-project#35957 unreachable, sgl-project#37143
not applicable (timeouts -1).
Tests: test_tokenizer_manager_rid_cleanup 21 passed (hermetic, cgroup 3G).
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 26, 2026
…cks, chronologisch)

Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum
(Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst
anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity;
479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 ->
S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name
(Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)).
Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc

Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859
sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579
sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438
(SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool
sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild),
#DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release;
P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6
+ F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49
Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens).

Auflösungen (Gabel -> Form, Grund):
- L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität"
  nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker-
  Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags.
  Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity.
- RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser
  CalibrationIdentity.d_max_running_requests; nextflash unverändert.
- URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und
  Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie).
- scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP
  komponiert (beide Instrumente getrennt schaltbar).
- schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF
  computed_input_len.
- gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab.
- flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt).
- checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation.
- xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen.
- vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read.
- FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9).
- argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only).
- census_checkpoint_decision (W161) jetzt für beide Profile aktiv.

Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber
dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu.
Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants