Skip to content

fix(h200): pin sglang to a post-0.5.16 nightly so grammar decoding works - #1230

Merged
1a1a11a merged 1 commit into
devfrom
jason/claude/sglang-nightly-grammar
Aug 7, 2026
Merged

1a1a11a merged 1 commit into
devfrom
jason/claude/sglang-nightly-grammar

Conversation

@1a1a11a

@1a1a11a 1a1a11a commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Problem

Every grammar-constrained request to the local DeepSeek-V4-Flash H200 replicas is rejected by sglang v0.5.16:

DFLASH speculative decoding does not support grammar-constrained decoding yet.

The guard is validate_dflash_request(). is_dflash_family() is is_dflash() or is_dspark(), so our "speculative_algorithm": "DSPARK" is in scope despite the DFLASH wording. Verified against localhost:8005:

request v0.5.16
tool_choice: "auto" ✅
tool_choice: "required" ❌ 400
response_format: json_object ❌ 400
response_format: json_schema ❌ 400
regex / ebnf / structural_tag ❌ 400

On the streaming path sglang emits that 400 as an in-band SSE error frame under an HTTP 200, then a placeholder usage of prompt_tokens: 1, completion_tokens: 1, then [DONE]. The gateway logs it as a success (status_code=200, error NULL, tokens 1/1, billed) and never fails over — even though the DeepSeek API and Ollama Cloud routes for this same model both handle json_object fine.

Measured on prod over six hours: 81,880 of 81,945 response_format requests to the sglang routes came back this way, ~100%, across 2 users.

Fix

sgl-project/sglang#30096 makes the rejection conditional on spec_algorithm.supports_grammar_overlap() (true for the DFlash family) and adds the verify-time grammar bitmask.

It merged to main on 2026-07-25 11:36 UTC — eleven hours after v0.5.16 was cut at 00:13 UTC the same day. git compare v0.5.16...d021990 → diverged, not an ancestor. v0.5.16 is still the latest release, so the only way to get the fix today is a nightly.

This pins lmsysorg/sglang:nightly-dev-20260806-ae5f8c94 — 456 commits past the fix, 0 behind.

Rollback plan

Revert the sglang_image line and restart the units. Move back to a tag when v0.5.17 ships (~2026-08-08 on the fortnightly cadence: 0.5.13 Jun 13 → 0.5.14 Jun 26 → 0.5.15 Jul 10 → 0.5.16 Jul 25). A nightly is 456 unreviewed commits of drift on a 1M-context MoE — a bridge, not a destination.

Rollout

Replica B derives its config from this file, so both h200a replicas pick it up on their next unit restart. Rolling them one at a time: restart h200_idle_proxy (8003), soak, then h200_idle_proxy_b (8005). h200b (8004) stays on v0.5.16 as a control until the soak passes.

Not in this PR

The gateway-side bug — an in-band SSE error frame is treated as a successful stream, so no failover and bogus 1/1 usage gets logged and billed — is independent and still open. It will bite again for any upstream that reports errors this way.

Test

tests/test_local_deployment_proxy.py + tests/test_h200_hicache_config.py: 111 passed.

v0.5.16 rejects every grammar-constrained request on the DeepSeek-V4-Flash
H200 replicas with HTTP 400 "DFLASH speculative decoding does not support
grammar-constrained decoding yet". The guard is validate_dflash_request(),
and is_dflash_family() is is_dflash() or is_dspark() -- so our DSPARK config
is in scope despite the DFLASH wording. It covers response_format
(json_object and json_schema), regex, ebnf, structural_tag, and
tool_choice: "required" / named function.

On the streaming path sglang returns that 400 as an in-band SSE error frame
under an HTTP 200, with a placeholder usage of prompt_tokens 1 /
completion_tokens 1. The gateway logged those as successful 200s with 1/1
tokens and never failed over to the DeepSeek API or Ollama Cloud routes,
both of which handle json_object fine. Measured on prod: 82k requests in six
hours, ~100% of all response_format traffic to the sglang routes.

sgl-project/sglang#30096 makes the rejection conditional on
spec_algorithm.supports_grammar_overlap() and adds the verify-time bitmask.
It merged to main on 2026-07-25 11:36 UTC, eleven hours after v0.5.16 was
cut at 00:13 UTC the same day, and no release carries it yet. Pin the
2026-08-06 nightly (ae5f8c94, 456 commits past the fix) and move back to a
tag when v0.5.17 ships (~2026-08-08 on the fortnightly cadence).

Replica B derives its config from this file, so both h200a replicas follow
on their next unit restart -- roll them one at a time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings August 7, 2026 06:00

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The job was not started because recent GitHub Actions payments have failed or your spending limit needs to be increased.

@kilo-code-bot

kilo-code-bot Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Code Review Summary

Status: No Issues Found | Recommendation: Merge

Files Reviewed (2 files)
  • ops/h200_idle_proxy/models.json
  • ops/h200_idle_proxy/README.md

Reviewed by nemotron-3-ultra-550b-a55b:free · Input: 72.3K · Output: 2.3K · Cached: 67.6K

@1a1a11a
1a1a11a merged commit c6e2ed6 into dev Aug 7, 2026
15 of 16 checks passed
1a1a11a added a commit that referenced this pull request Aug 8, 2026
Supersedes the nightly pinned in #1230, before it was ever deployed.

## Why now

`v0.5.17` shipped **2026-08-08 00:19 UTC** — the fortnightly cadence
#1230 predicted. It is the first release carrying the DFlash grammar
fix:

```
$ gh api repos/sgl-project/sglang/compare/d021990...v0.5.17
{"status":"ahead","ahead_by":398,"behind_by":0}
```

`behind_by: 0` →
[sgl-project/sglang#30096](sgl-project/sglang#30096)
is an ancestor of the tag.

By that release the guard is not merely conditional, it is **gone**:

```python
# v0.5.17 python/sglang/srt/speculative/dflash_utils.py
def validate_dflash_request(req: Req, enable_overlap: bool) -> Optional[str]:
    if req.return_logprob: ...
    if enable_overlap and req.return_hidden_states: ...
    return None          # grammar rejection block removed
```

`is_dflash_family()` is still `is_dflash() or is_dspark()`, so our
`DSPARK` config is in scope.

## Why this is strictly better than the nightly

|  | `nightly-dev-20260806-ae5f8c94` | `v0.5.17` |
|---|---|---|
| commits past the fix | 456 | **398** |
| release testing | none | full |
| follow-up needed | yes — "move back to a tag" | none |

#1230's own rollback note said to do exactly this when v0.5.17 shipped.

## Cost of the swap: zero

Nothing was deployed on the nightly. All three replicas still return the
400 as of this PR, so this replaces one pending rollout with another
rather than adding a second ~14-minute cold start per replica.

## Rollout (unchanged from #1230)

Replica B derives its config from this file, so both h200a replicas
follow on their next unit restart. Roll one at a time: `h200_idle_proxy`
(8003) first — it is currently taking ~0.2% of traffic against 8005's
~99.7% — soak, then `h200_idle_proxy_b` (8005). h200b (8004) stays on
v0.5.16 as a control.

`docker pull lmsysorg/sglang:v0.5.17` before the restart; the proxy has
no pull step, so otherwise `docker run` fetches it inline and stretches
the cold start.

## Test

`tests/test_local_deployment_proxy.py` +
`tests/test_h200_hicache_config.py`: 111 passed.

Co-authored-by: Juncheng Yang <Juncheng Yang>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
@realtmxi
realtmxi deleted the jason/claude/sglang-nightly-grammar branch August 28, 2026 08:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants