fix(h200): pin sglang to a post-0.5.16 nightly so grammar decoding works - #1230
Merged
Merged
Conversation
v0.5.16 rejects every grammar-constrained request on the DeepSeek-V4-Flash H200 replicas with HTTP 400 "DFLASH speculative decoding does not support grammar-constrained decoding yet". The guard is validate_dflash_request(), and is_dflash_family() is is_dflash() or is_dspark() -- so our DSPARK config is in scope despite the DFLASH wording. It covers response_format (json_object and json_schema), regex, ebnf, structural_tag, and tool_choice: "required" / named function. On the streaming path sglang returns that 400 as an in-band SSE error frame under an HTTP 200, with a placeholder usage of prompt_tokens 1 / completion_tokens 1. The gateway logged those as successful 200s with 1/1 tokens and never failed over to the DeepSeek API or Ollama Cloud routes, both of which handle json_object fine. Measured on prod: 82k requests in six hours, ~100% of all response_format traffic to the sglang routes. sgl-project/sglang#30096 makes the rejection conditional on spec_algorithm.supports_grammar_overlap() and adds the verify-time bitmask. It merged to main on 2026-07-25 11:36 UTC, eleven hours after v0.5.16 was cut at 00:13 UTC the same day, and no release carries it yet. Pin the 2026-08-06 nightly (ae5f8c94, 456 commits past the fix) and move back to a tag when v0.5.17 ships (~2026-08-08 on the fortnightly cadence). Replica B derives its config from this file, so both h200a replicas follow on their next unit restart -- roll them one at a time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Copilot stopped reviewing on behalf of
1a1a11a due to an error
August 7, 2026 06:00
Code Review SummaryStatus: No Issues Found | Recommendation: Merge Files Reviewed (2 files)
Reviewed by nemotron-3-ultra-550b-a55b:free · Input: 72.3K · Output: 2.3K · Cached: 67.6K |
1a1a11a
added a commit
that referenced
this pull request
Aug 8, 2026
Supersedes the nightly pinned in #1230, before it was ever deployed. ## Why now `v0.5.17` shipped **2026-08-08 00:19 UTC** — the fortnightly cadence #1230 predicted. It is the first release carrying the DFlash grammar fix: ``` $ gh api repos/sgl-project/sglang/compare/d021990...v0.5.17 {"status":"ahead","ahead_by":398,"behind_by":0} ``` `behind_by: 0` → [sgl-project/sglang#30096](sgl-project/sglang#30096) is an ancestor of the tag. By that release the guard is not merely conditional, it is **gone**: ```python # v0.5.17 python/sglang/srt/speculative/dflash_utils.py def validate_dflash_request(req: Req, enable_overlap: bool) -> Optional[str]: if req.return_logprob: ... if enable_overlap and req.return_hidden_states: ... return None # grammar rejection block removed ``` `is_dflash_family()` is still `is_dflash() or is_dspark()`, so our `DSPARK` config is in scope. ## Why this is strictly better than the nightly | | `nightly-dev-20260806-ae5f8c94` | `v0.5.17` | |---|---|---| | commits past the fix | 456 | **398** | | release testing | none | full | | follow-up needed | yes — "move back to a tag" | none | #1230's own rollback note said to do exactly this when v0.5.17 shipped. ## Cost of the swap: zero Nothing was deployed on the nightly. All three replicas still return the 400 as of this PR, so this replaces one pending rollout with another rather than adding a second ~14-minute cold start per replica. ## Rollout (unchanged from #1230) Replica B derives its config from this file, so both h200a replicas follow on their next unit restart. Roll one at a time: `h200_idle_proxy` (8003) first — it is currently taking ~0.2% of traffic against 8005's ~99.7% — soak, then `h200_idle_proxy_b` (8005). h200b (8004) stays on v0.5.16 as a control. `docker pull lmsysorg/sglang:v0.5.17` before the restart; the proxy has no pull step, so otherwise `docker run` fetches it inline and stretches the cold start. ## Test `tests/test_local_deployment_proxy.py` + `tests/test_h200_hicache_config.py`: 111 passed. Co-authored-by: Juncheng Yang <Juncheng Yang> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Every grammar-constrained request to the local DeepSeek-V4-Flash H200 replicas is rejected by sglang
v0.5.16:The guard is
validate_dflash_request().is_dflash_family()isis_dflash() or is_dspark(), so our"speculative_algorithm": "DSPARK"is in scope despite the DFLASH wording. Verified againstlocalhost:8005:tool_choice: "auto"tool_choice: "required"response_format: json_objectresponse_format: json_schemaregex/ebnf/structural_tagOn the streaming path sglang emits that 400 as an in-band SSE
errorframe under an HTTP 200, then a placeholderusageofprompt_tokens: 1, completion_tokens: 1, then[DONE]. The gateway logs it as a success (status_code=200,errorNULL, tokens 1/1, billed) and never fails over — even though the DeepSeek API and Ollama Cloud routes for this same model both handlejson_objectfine.Measured on prod over six hours: 81,880 of 81,945
response_formatrequests to the sglang routes came back this way, ~100%, across 2 users.Fix
sgl-project/sglang#30096 makes the rejection conditional on
spec_algorithm.supports_grammar_overlap()(true for the DFlash family) and adds the verify-time grammar bitmask.It merged to
mainon 2026-07-25 11:36 UTC — eleven hours afterv0.5.16was cut at 00:13 UTC the same day.git compare v0.5.16...d021990→ diverged, not an ancestor. v0.5.16 is still the latest release, so the only way to get the fix today is a nightly.This pins
lmsysorg/sglang:nightly-dev-20260806-ae5f8c94— 456 commits past the fix, 0 behind.Rollback plan
Revert the
sglang_imageline and restart the units. Move back to a tag when v0.5.17 ships (~2026-08-08 on the fortnightly cadence: 0.5.13 Jun 13 → 0.5.14 Jun 26 → 0.5.15 Jul 10 → 0.5.16 Jul 25). A nightly is 456 unreviewed commits of drift on a 1M-context MoE — a bridge, not a destination.Rollout
Replica B derives its config from this file, so both h200a replicas pick it up on their next unit restart. Rolling them one at a time: restart
h200_idle_proxy(8003), soak, thenh200_idle_proxy_b(8005). h200b (8004) stays on v0.5.16 as a control until the soak passes.Not in this PR
The gateway-side bug — an in-band SSE
errorframe is treated as a successful stream, so no failover and bogus 1/1 usage gets logged and billed — is independent and still open. It will bite again for any upstream that reports errors this way.Test
tests/test_local_deployment_proxy.py+tests/test_h200_hicache_config.py: 111 passed.