[DFLASH] Support grammar-constrained decoding in speculative verify - #30096
Conversation
DFLASH previously rejected any grammar-constrained request (json_schema / regex / ebnf / structural_tag -- i.e. tool_choice=required, named tool_choice, response_format=json_*) with a 400 error, because the DFLASH verify path never applied the grammar vocab mask to the target logits. DFLASH verify is a linear draft chain, which is a degenerate case of EAGLE's draft tree, so we reuse the existing EAGLE helper generate_token_bitmask() by building a chain topology (retrieve_next_token=[1,2,...,-1], no siblings) and apply the resulting vocab mask to the target logits before the accept/argmax step. Grammar FSM advancement over committed tokens is already handled generically for spec-v2 in SchedulerOutputProcessorMixin._resolve_spec_v2_tokens, so no change is needed there. Verified on Qwen3.6-35B-A3B + DFLASH (block-size 8): forced/named tool calls, response_format json_object/json_schema, and regex all return valid, schema-correct output, and DFLASH stays active on these requests (spec_accept_length ~6.5/8) -- i.e. not a silent non-spec fallback.
Unit (CPU): cover the linear-chain topology DFlashWorkerV2 feeds to generate_token_bitmask -- one case walks the full chain in order, one stops traversal when a chain token is disallowed by the grammar. Integration: TestDFlashServerBase.test_grammar_constrained_decoding sends a response_format=json_schema request (asserts schema-valid JSON) and a regex request (asserts the output matches), across all DFLASH server configs.
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
Hi @Ying1123 @merrymercy @hnyls2002 @Qiaolin-Yu, could you please trigger CI for this PR? Thank you! |
|
Heads up from deploying this in production. We rolled DSpark onto DeepSeek-V4-Flash (TP2) and DeepSeek-V4-Pro (TP8) on B300 and this PR is the exact fix we needed, but as-is it is incomplete for 1. The guard removal covers DSPARK, the masking does not. 2. Likely off-by-one in the DFlash masking itself. Our DSPARK fix (validated in production, 0/20 -> 20/20 schema-valid, GSM8K unchanged, speculation still active): three changes to
Running live on real client traffic ( |
Thanks for the detailed investigation and for validating this in production! Feel free to cherry-pick my commit and build the DSPARK changes on top of it. That should preserve the original commit history/authorship while letting you iterate on the DSPARK-specific fixes. I'd be happy to review the follow-up PR once it's ready. |
|
Opened the DSPARK counterpart as #31753. Following on from my note above: this PR lifts the grammar guard out of the Happy to reconcile whichever of the two lands first; the only overlap is the one |
|
/rerun-test test_dflash.py test_spec_utils_traverse_tree.py |
|
Results for 🚀 🚀 |
|
/tag-and-rerun-ci |
…rks (#1230) ## Problem Every grammar-constrained request to the local DeepSeek-V4-Flash H200 replicas is rejected by sglang `v0.5.16`: ``` DFLASH speculative decoding does not support grammar-constrained decoding yet. ``` The guard is `validate_dflash_request()`. `is_dflash_family()` is `is_dflash() or is_dspark()`, so our `"speculative_algorithm": "DSPARK"` is in scope despite the DFLASH wording. Verified against `localhost:8005`: | request | v0.5.16 | |---|---| | `tool_choice: "auto"` | ✅ | | `tool_choice: "required"` | ❌ 400 | | `response_format: json_object` | ❌ 400 | | `response_format: json_schema` | ❌ 400 | | `regex` / `ebnf` / `structural_tag` | ❌ 400 | On the **streaming** path sglang emits that 400 as an in-band SSE `error` frame under an **HTTP 200**, then a placeholder `usage` of `prompt_tokens: 1, completion_tokens: 1`, then `[DONE]`. The gateway logs it as a success (`status_code=200`, `error` NULL, tokens 1/1, billed) and never fails over — even though the DeepSeek API and Ollama Cloud routes for this same model both handle `json_object` fine. Measured on prod over six hours: **81,880 of 81,945** `response_format` requests to the sglang routes came back this way, ~100%, across 2 users. ## Fix [sgl-project/sglang#30096](sgl-project/sglang#30096) makes the rejection conditional on `spec_algorithm.supports_grammar_overlap()` (true for the DFlash family) and adds the verify-time grammar bitmask. It merged to `main` on **2026-07-25 11:36 UTC** — eleven hours after `v0.5.16` was cut at **00:13 UTC** the same day. `git compare v0.5.16...d021990` → diverged, not an ancestor. v0.5.16 is still the latest release, so the only way to get the fix today is a nightly. This pins `lmsysorg/sglang:nightly-dev-20260806-ae5f8c94` — 456 commits past the fix, 0 behind. ## Rollback plan Revert the `sglang_image` line and restart the units. **Move back to a tag when v0.5.17 ships** (~2026-08-08 on the fortnightly cadence: 0.5.13 Jun 13 → 0.5.14 Jun 26 → 0.5.15 Jul 10 → 0.5.16 Jul 25). A nightly is 456 unreviewed commits of drift on a 1M-context MoE — a bridge, not a destination. ## Rollout Replica B derives its config from this file, so both h200a replicas pick it up on their next unit restart. Rolling them one at a time: restart `h200_idle_proxy` (8003), soak, then `h200_idle_proxy_b` (8005). h200b (8004) stays on v0.5.16 as a control until the soak passes. ## Not in this PR The gateway-side bug — an in-band SSE `error` frame is treated as a successful stream, so no failover and bogus 1/1 usage gets logged and billed — is independent and still open. It will bite again for any upstream that reports errors this way. ## Test `tests/test_local_deployment_proxy.py` + `tests/test_h200_hicache_config.py`: 111 passed. Co-authored-by: Juncheng Yang <Juncheng Yang> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Supersedes the nightly pinned in #1230, before it was ever deployed. ## Why now `v0.5.17` shipped **2026-08-08 00:19 UTC** — the fortnightly cadence #1230 predicted. It is the first release carrying the DFlash grammar fix: ``` $ gh api repos/sgl-project/sglang/compare/d021990...v0.5.17 {"status":"ahead","ahead_by":398,"behind_by":0} ``` `behind_by: 0` → [sgl-project/sglang#30096](sgl-project/sglang#30096) is an ancestor of the tag. By that release the guard is not merely conditional, it is **gone**: ```python # v0.5.17 python/sglang/srt/speculative/dflash_utils.py def validate_dflash_request(req: Req, enable_overlap: bool) -> Optional[str]: if req.return_logprob: ... if enable_overlap and req.return_hidden_states: ... return None # grammar rejection block removed ``` `is_dflash_family()` is still `is_dflash() or is_dspark()`, so our `DSPARK` config is in scope. ## Why this is strictly better than the nightly | | `nightly-dev-20260806-ae5f8c94` | `v0.5.17` | |---|---|---| | commits past the fix | 456 | **398** | | release testing | none | full | | follow-up needed | yes — "move back to a tag" | none | #1230's own rollback note said to do exactly this when v0.5.17 shipped. ## Cost of the swap: zero Nothing was deployed on the nightly. All three replicas still return the 400 as of this PR, so this replaces one pending rollout with another rather than adding a second ~14-minute cold start per replica. ## Rollout (unchanged from #1230) Replica B derives its config from this file, so both h200a replicas follow on their next unit restart. Roll one at a time: `h200_idle_proxy` (8003) first — it is currently taking ~0.2% of traffic against 8005's ~99.7% — soak, then `h200_idle_proxy_b` (8005). h200b (8004) stays on v0.5.16 as a control. `docker pull lmsysorg/sglang:v0.5.17` before the restart; the proxy has no pull step, so otherwise `docker run` fetches it inline and stretches the cold start. ## Test `tests/test_local_deployment_proxy.py` + `tests/test_h200_hicache_config.py`: 111 passed. Co-authored-by: Juncheng Yang <Juncheng Yang> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…gl-project#30096) Co-authored-by: hnyls2002 <lsyincs@gmail.com>
…mmar with DFlash (group D); sgl-project#31488/sgl-project#32409 checked Applicability (27B weg2 serving): the front forwards the whole client body to D in leg 2, and D validated every request with validate_dflash_request, which ABORTED return_logprob and every grammar (json_schema / regex / ebnf / structural_tag). An Anthropic tool_choice any|tool becomes OpenAI required/named and a structure constraint (serving_chat), so such agent requests and every logprob request died on D. Full-feature serving -> port. sgl-project#33459 (36853b8): validate_dflash_request admits return_logprob for DFLASH (DSpark keeps the refusal, as upstream); the ValueError at the top of DFlashWorkerV2.forward_batch_generation is gone; the verify computes the committed run's logprobs with the fork's compute_spec_v2_logprobs (output_indices = arange(bs*block).view(bs, block), steps = block-1: row j predicts out_tokens[:, j]) -> the spec-v2 layout batch_result_processor. _apply_decode_logprobs already reads (next_token_logprobs[i][:commit_len]). sgl-project#30096 (d021990), ADAPTED: upstream builds a GrammarTree and overlaps the FSM advance via the sgl-project#31488 grammar barrier; the fork has neither and runs EAGLE v2 grammar on the synchronous generate_token_bitmask path (overlap is already disabled for spec+grammar decode batches by is_disable_overlap_for_batch, and the result processor already advances the grammar for any spec-v2 result). So: DFlashVerifyInput.grammar field; the rank-synced draft tokens go to the host before the verify launch (grammar batches only); after the forward, _dflash_grammar_vocab_mask walks the LINEAR block (next = i+1, no siblings) with generate_token_bitmask, clears a stale sampling_info.vocab_mask (as EAGLE does), and the mask is applied to the verify logits after apply_dflash_verify_logits_adjustments and before any accept path (greedy/Triton, sampling, selector). Masks are identical on every rank (synced tokens, identical grammar states); accepts stay synced. validate_dflash_request admits grammar for DFLASH, DSpark keeps refusing. The scheduler passes its spec_algorithm (new optional 3rd arg). sgl-project#31488 (e7e8aaa, overlap grammar with spec verify: grammar barrier, GrammarTree, eagle/frozen-kv/multi-layer workers, 214 lines) and sgl-project#32409 (3da1071, GrammarMask type across 16 files) are checked and NOT ported: throughput/refactor over infrastructure the fork does not have; the synchronous form above is correct without them. Risk: none for requests without logprobs/grammar (both paths are gated on batch.return_logprob / batch.has_grammar). Grammar batches pay one D2H of the draft block and a CPU chain walk per verify step. Tests (CPU, incg 3 GiB cgroup): - new test_dflash_logprobs_grammar_33459_30096.py (7): base 1e49837 6 failed / 1 passed (the unchanged hidden-states refusal), with the port 7 passed. - base == port: test_spec_utils_traverse_tree 1/1, test_batch_result_processor_spec_grammar 2/2; test_dflash_*.py group 66 passed (base 61 + the 5 red sgl-project#37818 tests of this branch).
…ec ring 27B ReplaySSM package, slice S3 of REPLAYSSM_PLAN.md. With the flag off every path is unchanged (no rows planned -> recurrent route, intermediate commit). Ported: upstream GDNAttnBackend._replayssm_target_verify (same kernel call, launch_mode="verify", null block -1, request rows as replay indices) and the compact commit sequence of upstream's spec_utils (commit_gdn_replayssm_spec + commit_gdn_replayssm_circular + conv-window rollback). Adapted (this line's own wiring; upstream refuses GDN + DFLASH): - ForwardMetadata.replayssm_spec_rows: the verify metadata plans the request rows (req_pool_indices) in all three builders (eager, graph capture, graph replay -- there the static buffer with padded rows zeroed). The commit reads the rows the verify used from the metadata, like mamba_cache_indices, so no caller changes: DFLASH (dflash_worker_v2 untouched), the EAGLE/MTP commit and the lane all go through update_mamba_state_after_mtp_verify. - The commit sits in that shared hook (ring branch when intermediate_ssm is None). Linear chain -> accepted = last step + 1 (bonus included); -1 folds nothing, like the masked scatter. - FOLD EVERY COMMIT for any SSM dtype (upstream defers the fold for fp32): radix insert, HiCache backup, tail adopt and the flip all read `temporal` and assume it is the committed state after every verify, as the recurrent route makes it. The ring is then per-step scratch (write_pos 0 at every verify). - Heal: outside a capture the verify metadata re-states write_pos = is_flush = 0 for the step's rows (true by construction after every commit), so a verify never trusts cursor bytes from an earlier step (TMS restore, flip, reused row). cache_base is circular and harmless at write_pos 0. - forward_extend routes the verify to the ring when rows are planned; a draft tree on the ring, or a pool with neither intermediate state nor planned rows, is refused. The ring kernel honours token strides, so the verify split stays torch.split views (sgl-project#33778). - Advance cursor null block 0 (request row 0 is the ReqToTokenPool padding row); the widest verify window is kept on the pool as the advance's constant (one compiled variant across the adaptive ladder). - Ring length must be >= 16 (the compact commit's tl.dot) -- pool and static check; S2 test moved to L=16. - _replayssm_spec_for: a draft-KV-only producer (no verify workspace, sgl-project#1233 FIX 4) treats the flag as a no-op instead of hitting the pool's missing-window refusal. - getattr on the new metadata field in the two readers (sgl-project#624 stub drift: test_gdn_verify_strided_qkv_33778 doubles predate it). Tests (CPU, incg cgroup, one file at a time): test/registered/unit/layers/attention/test_replayssm_spec_route_s3.py: 10 passed. Hermetic: rows planned + healed (eager, replay; not in capture; padded row -> row 0), flag off / decode plan nothing, the three refusals, producer no-op. Interpreter (production route functions on real CPU pools, ring vs intermediate, 2 layers, head ratio 3, 2 requests + padded row, track crossing, second step after garbage cursors + heal): fp32 verify rel 3.7e-7 / 2.4e-7, commit (incl. track slot) 2.9e-7 / 3.1e-7 vs the recurrent route; fp16 (bf16 stand-in) verify 5.0e-4 vs recurrent 3.8e-4 of the fp32 truth, commit 3.6e-4 == recurrent 3.6e-4; conv rollback equal; untouched slots bit-exact; cursors 0 after every commit. On the S2 sources: 7 failed + 3 errors. Regression: test_replayssm_spec_ring_pool_s2 10, test_mamba_checkpoint_interval 58, test_dflash_mamba_track_post_verify_37818 5, test_conv_verify_private_ window_444 13, test_prefill_graph_stale_track_rows_34184 2, test_weg2_prefill_only_capture_1233 15, test_forward_metadata_plan_record 9, test_fused_replay_state_indices_32219 3, test_gdn_verify_strided_qkv_33778 4, test_verify_intermediate_row_ownership_450 12, test_mamba2_conv_verify_ private_window_450 10, test_dual_group_concurrency 159 -- all passed. ruff F: no new findings. 27B line (desk/27b-up-replayssm-line-0924): cherry-picked from 7d69bd8 (desk/27b-up-replayssm-0924, on the NF-based staging line fa757eb), clean. gdn_backend.py and mamba2_metadata.py are identical on both lines; dflash_worker_v2 differs here (sgl-project#31468, sgl-project#33459/sgl-project#30096 ports) but not in _update_target_mamba_state_after_verify, so the DFLASH commit reaches the shared hook exactly as on NF. One adaptation, text only: the commit docstring listed "tail adopt" among the readers of `temporal` -- that is the NF line's H21 install (weg2.tail_adopt), absent here; the 27B readers are the radix insert, the HiCache backup (the weg2 L2 arena write, xsn351) and the P/D flip. Tests on this line (CPU, incg cgroup, S3 tree, one file at a time): test_replayssm_spec_route_s3.py 10 passed -- interpreter, production route vs the recurrent route: fp32 verify rel 3.74e-7 / 2.40e-7 (step 1 / step 2 after garbage cursors + heal), commit incl. the track slot 2.94e-7 / 3.10e-7 (all < 4e-7); fp16 (bf16 stand-in) verify 5.03e-4 / 2.72e-4 vs recurrent 3.84e-4 / 2.55e-4 of the fp32 truth, commit 3.61e-4 / 3.27e-4 == recurrent; conv rollback equal, padded outputs zero, untouched slots bit-exact, cursors 0 after every commit; test_replayssm_spec_ring_pool_s2.py 10 passed (8 subtests). Regression (package tip = S6 tree, same cgroup, one file at a time): test_dflash_mamba_track_post_verify_37818 5, test_conv_verify_private_window_444 13, test_prefill_graph_stale_track_rows_34184 2, test_weg2_prefill_only_capture_1233 15, test_forward_metadata_plan_record 9, test_fused_replay_state_indices_32219 3, test_gdn_verify_strided_qkv_33778 4, test_verify_intermediate_row_ownership_450 12, test_mamba2_conv_verify_private_window_450 10, test_dual_group_concurrency 159, test_mamba_checkpoint_interval 58 -- all passed.
…ec ring 27B ReplaySSM package, slice S3 of REPLAYSSM_PLAN.md. With the flag off every path is unchanged (no rows planned -> recurrent route, intermediate commit). Ported: upstream GDNAttnBackend._replayssm_target_verify (same kernel call, launch_mode="verify", null block -1, request rows as replay indices) and the compact commit sequence of upstream's spec_utils (commit_gdn_replayssm_spec + commit_gdn_replayssm_circular + conv-window rollback). Adapted (this line's own wiring; upstream refuses GDN + DFLASH): - ForwardMetadata.replayssm_spec_rows: the verify metadata plans the request rows (req_pool_indices) in all three builders (eager, graph capture, graph replay -- there the static buffer with padded rows zeroed). The commit reads the rows the verify used from the metadata, like mamba_cache_indices, so no caller changes: DFLASH (dflash_worker_v2 untouched), the EAGLE/MTP commit and the lane all go through update_mamba_state_after_mtp_verify. - The commit sits in that shared hook (ring branch when intermediate_ssm is None). Linear chain -> accepted = last step + 1 (bonus included); -1 folds nothing, like the masked scatter. - FOLD EVERY COMMIT for any SSM dtype (upstream defers the fold for fp32): radix insert, HiCache backup, tail adopt and the flip all read `temporal` and assume it is the committed state after every verify, as the recurrent route makes it. The ring is then per-step scratch (write_pos 0 at every verify). - Heal: outside a capture the verify metadata re-states write_pos = is_flush = 0 for the step's rows (true by construction after every commit), so a verify never trusts cursor bytes from an earlier step (TMS restore, flip, reused row). cache_base is circular and harmless at write_pos 0. - forward_extend routes the verify to the ring when rows are planned; a draft tree on the ring, or a pool with neither intermediate state nor planned rows, is refused. The ring kernel honours token strides, so the verify split stays torch.split views (sgl-project#33778). - Advance cursor null block 0 (request row 0 is the ReqToTokenPool padding row); the widest verify window is kept on the pool as the advance's constant (one compiled variant across the adaptive ladder). - Ring length must be >= 16 (the compact commit's tl.dot) -- pool and static check; S2 test moved to L=16. - _replayssm_spec_for: a draft-KV-only producer (no verify workspace, sgl-project#1233 FIX 4) treats the flag as a no-op instead of hitting the pool's missing-window refusal. - getattr on the new metadata field in the two readers (sgl-project#624 stub drift: test_gdn_verify_strided_qkv_33778 doubles predate it). Tests (CPU, incg cgroup, one file at a time): test/registered/unit/layers/attention/test_replayssm_spec_route_s3.py: 10 passed. Hermetic: rows planned + healed (eager, replay; not in capture; padded row -> row 0), flag off / decode plan nothing, the three refusals, producer no-op. Interpreter (production route functions on real CPU pools, ring vs intermediate, 2 layers, head ratio 3, 2 requests + padded row, track crossing, second step after garbage cursors + heal): fp32 verify rel 3.7e-7 / 2.4e-7, commit (incl. track slot) 2.9e-7 / 3.1e-7 vs the recurrent route; fp16 (bf16 stand-in) verify 5.0e-4 vs recurrent 3.8e-4 of the fp32 truth, commit 3.6e-4 == recurrent 3.6e-4; conv rollback equal; untouched slots bit-exact; cursors 0 after every commit. On the S2 sources: 7 failed + 3 errors. Regression: test_replayssm_spec_ring_pool_s2 10, test_mamba_checkpoint_interval 58, test_dflash_mamba_track_post_verify_37818 5, test_conv_verify_private_ window_444 13, test_prefill_graph_stale_track_rows_34184 2, test_weg2_prefill_only_capture_1233 15, test_forward_metadata_plan_record 9, test_fused_replay_state_indices_32219 3, test_gdn_verify_strided_qkv_33778 4, test_verify_intermediate_row_ownership_450 12, test_mamba2_conv_verify_ private_window_450 10, test_dual_group_concurrency 159 -- all passed. ruff F: no new findings. 27B line (desk/27b-up-replayssm-line-0924): cherry-picked from 7d69bd8 (desk/27b-up-replayssm-0924, on the NF-based staging line fa757eb), clean. gdn_backend.py and mamba2_metadata.py are identical on both lines; dflash_worker_v2 differs here (sgl-project#31468, sgl-project#33459/sgl-project#30096 ports) but not in _update_target_mamba_state_after_verify, so the DFLASH commit reaches the shared hook exactly as on NF. One adaptation, text only: the commit docstring listed "tail adopt" among the readers of `temporal` -- that is the NF line's H21 install (weg2.tail_adopt), absent here; the 27B readers are the radix insert, the HiCache backup (the weg2 L2 arena write, xsn351) and the P/D flip. Tests on this line (CPU, incg cgroup, S3 tree, one file at a time): test_replayssm_spec_route_s3.py 10 passed -- interpreter, production route vs the recurrent route: fp32 verify rel 3.74e-7 / 2.40e-7 (step 1 / step 2 after garbage cursors + heal), commit incl. the track slot 2.94e-7 / 3.10e-7 (all < 4e-7); fp16 (bf16 stand-in) verify 5.03e-4 / 2.72e-4 vs recurrent 3.84e-4 / 2.55e-4 of the fp32 truth, commit 3.61e-4 / 3.27e-4 == recurrent; conv rollback equal, padded outputs zero, untouched slots bit-exact, cursors 0 after every commit; test_replayssm_spec_ring_pool_s2.py 10 passed (8 subtests). Regression (package tip = S6 tree, same cgroup, one file at a time): test_dflash_mamba_track_post_verify_37818 5, test_conv_verify_private_window_444 13, test_prefill_graph_stale_track_rows_34184 2, test_weg2_prefill_only_capture_1233 15, test_forward_metadata_plan_record 9, test_fused_replay_state_indices_32219 3, test_gdn_verify_strided_qkv_33778 4, test_verify_intermediate_row_ownership_450 12, test_mamba2_conv_verify_private_window_450 10, test_dual_group_concurrency 159, test_mamba_checkpoint_interval 58 -- all passed. NF line (H64, onto 88dec66): one conflict, resolved line by line in gdn_backend.forward_extend. This line carries no sgl-project#33778 port (no kernel_dispatcher.target_verify_supports_strided_qkv), so the 27B hunk that widens use_strided_verify_qkv for the ring has nothing to widen. Kept the NF condition `(is_cuda() or is_hip()) and qkv_dim <= MAX_FUSED_QKV_SPLIT_DIM`: the ring verify gets the same q/k/v preparation as the recurrent verify on this line (dense [1, T, H, D] from fused_qkv_split_gdn_prefill on CUDA, torch.split views elsewhere); _replayssm_target_verify flattens both layouts by numel. Text only: the commit docstring names the NF H21 tail adopt (weg2/tail_adopt.py reads and installs cache.temporal) among the readers of temporal again -- the 27B line had dropped it because it lacks H21. Tests (CPU, incg cgroup, one file at a time): test_replayssm_spec_route_s3.py 10 passed; test_replayssm_spec_ring_pool_s2.py 10 passed (8 subtests). (cherry picked from commit 974254d)
…cks, chronologisch) Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum (Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity; 479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 -> S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name (Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)). Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859 sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579 sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438 (SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild), #DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release; P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6 + F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49 Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens). Auflösungen (Gabel -> Form, Grund): - L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität" nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker- Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags. Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity. - RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser CalibrationIdentity.d_max_running_requests; nextflash unverändert. - URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie). - scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP komponiert (beide Instrumente getrennt schaltbar). - schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF computed_input_len. - gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab. - flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt). - checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation. - xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen. - vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read. - FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9). - argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only). - census_checkpoint_decision (W161) jetzt für beide Profile aktiv. Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu. Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
Motivation
Under
--speculative-algorithm DFLASH, any grammar-constrained request is rejected with HTTP 400:This makes DFLASH unusable for structured / forced tool calling and JSON output — a very common serving path (agents, function calling, document extraction). Verified trigger surface:
tool_choice: "auto"tool_choice: "required"tool_choice: {named function}response_format: {type: json_object}response_format: {type: json_schema}regex/ebnf/structural_tagvalidate_dflash_request()rejects any request whosesampling_paramscarriesjson_schema/regex/ebnf/structural_tag, because the DFLASH verify path never applied the grammar vocab mask (the DFLASH worker had no grammar handling at all). Simply removing the guard would be unsafe — grammar requests would then run through an unconstrained verify and emit malformed output silently.Modifications
Two observations keep the change small and let it reuse the existing, proven EAGLE machinery:
draft_tokens[bs, block_size], where column 0 is the current (already-committed) token and columns1:are the draft proposals. A chain is a degenerate case of EAGLE's draft tree — each node's only child is the next position, no siblings — so the existing helpersglang.srt.speculative.spec_utils.generate_token_bitmask()/traverse_tree()is reused as-is by constructingretrieve_next_token = [1, 2, …, -1]andretrieve_next_sibling = [-1, …].SchedulerOutputProcessorMixin._resolve_spec_v2_tokens→_accept_grammar_tokens. DFLASH is spec-v2, so no change is needed there; the only missing piece was the mask application during verify.Changes:
python/sglang/srt/speculative/dflash_utils.py— remove the grammar-constraint rejection invalidate_dflash_request(thereturn_logprob/return_hidden_stateschecks are kept).python/sglang/srt/speculative/dflash_worker_v2.py— in the verify block, right afterapply_dflash_verify_logits_adjustments, whenbatch.has_grammar: build the linear-chain topology, callgenerate_token_bitmask(...), andgrammar.apply_vocab_mask(next_token_logits, vocab_mask)before the accept/argmax step. This masks both the greedy-argmax and the sampling-softmax accept paths, since both readnext_token_logits.The masking mirrors EAGLE's grammar-aware verify (
eagle_worker_v2.py+eagle_utils.py), so accepted tokens along the chain are grammar-valid by construction, and the downstream spec-v2 grammar FSM advancement stays correct.Accuracy Tests
Model: Qwen3.6-35B-A3B,
--speculative-algorithm DFLASH --speculative-dflash-block-size 8,--grammar-backend xgrammar,--tool-call-parser qwen3_coder(Tested on AMD ROCm / triton attention backend).tool_choice: required+ nested object/array schema), 10 trials @temperature=0, streaming: 10/10 schema-valid, deterministic (1 distinct output), and 7/7 ground-truth field checks (invoice number, grand total, currency, line-item count, seller, SWIFT, Σ line_totals == grand_total).autoselection,tool_choice: requiredwith mixed types (int / bool / enum / array), parallel tool calls (3 correct calls for a 3-city prompt), multi-turn with a returned tool result, Vietnamese unicode arguments preserved, streaming reassembly, and the no-tool-needed case (model does not emit a spurious call).accept_token failed, no traceback / abort.Tests added:
test/registered/unit/spec/test_spec_utils_traverse_tree.py: two cases for the linear-chain topology thatDFlashWorkerV2feeds togenerate_token_bitmask(retrieve_next_token = [1, 2, …, -1], no siblings) — one asserting the whole chain is walked in order, one asserting traversal stops when a chain token is disallowed by the grammar.test/registered/spec/dflash/test_dflash.py::test_grammar_constrained_decoding: launches the DFLASH server and sends aresponse_format: json_schemarequest (asserts schema-valid JSON) and aregexrequest (asserts the output matches the pattern). Runs across all DFLASH server configs in the file (page size, chunked prefill, no-cuda-graph, spec-v2, plan-stream).Benchmarking and Profiling
DFLASH remains genuinely active on grammar-constrained requests — it is not a silent fallback to non-speculative decoding:
spec_accept_rate ≈ 0.79,spec_accept_length ≈ 6.5 / 8, single-request decode throughput ~620–660 tok/s.So grammar requests keep the speculative-decoding speedup while now producing correct, constrained output.
Reproduce
Launch the server with DFLASH + a grammar backend (
--log-requestsso per-request spec metrics are printed):Send a grammar-constrained request (before this PR: HTTP 400; after: schema-valid output):
Read the speculative acceptance for that request from the server log (confirms DFLASH is still speculating under grammar, not falling back to 1 token/step):
Aggregate metrics are also available via
--enable-metrics(sglang:spec_accept_lengthin Prometheus). Swappingresponse_formatfortool_choice: "required", a namedtool_choice,response_format: {type: json_object}, or a top-level"regex": "..."exercises the other previously-rejected grammar paths.Checklist
CI States
Latest PR Test (Base): ⏳ Run #30155825631
Latest PR Test (Extra): ❌ Run #30155825531