Skip to content

[DFLASH] Support grammar-constrained decoding in speculative verify - #30096

Merged
hnyls2002 merged 10 commits into
sgl-project:mainfrom
hsthe29:dflash-grammar-constrained-decoding
Jul 25, 2026
Merged

hnyls2002 merged 10 commits into
sgl-project:mainfrom
hsthe29:dflash-grammar-constrained-decoding

Conversation

@hsthe29

@hsthe29 hsthe29 commented Jul 4, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Under --speculative-algorithm DFLASH, any grammar-constrained request is rejected with HTTP 400:

DFLASH speculative decoding does not support grammar-constrained decoding yet.

This makes DFLASH unusable for structured / forced tool calling and JSON output — a very common serving path (agents, function calling, document extraction). Verified trigger surface:

Request before after
tool_choice: "auto" ✅ OK ✅ OK
tool_choice: "required" ❌ 400 ✅ OK
tool_choice: {named function} ❌ 400 ✅ OK
response_format: {type: json_object} ❌ 400 ✅ OK
response_format: {type: json_schema} ❌ 400 ✅ OK
regex / ebnf / structural_tag ❌ 400 ✅ OK
plain chat (no constraint) ✅ OK ✅ OK

validate_dflash_request() rejects any request whose sampling_params carries json_schema / regex / ebnf / structural_tag, because the DFLASH verify path never applied the grammar vocab mask (the DFLASH worker had no grammar handling at all). Simply removing the guard would be unsafe — grammar requests would then run through an unconstrained verify and emit malformed output silently.

Modifications

Two observations keep the change small and let it reuse the existing, proven EAGLE machinery:

  1. DFLASH verify is a linear draft chain: draft_tokens[bs, block_size], where column 0 is the current (already-committed) token and columns 1: are the draft proposals. A chain is a degenerate case of EAGLE's draft tree — each node's only child is the next position, no siblings — so the existing helper sglang.srt.speculative.spec_utils.generate_token_bitmask() / traverse_tree() is reused as-is by constructing retrieve_next_token = [1, 2, …, -1] and retrieve_next_sibling = [-1, …].
  2. Grammar FSM advancement over committed tokens is already generic for spec-v2 in SchedulerOutputProcessorMixin._resolve_spec_v2_tokens → _accept_grammar_tokens. DFLASH is spec-v2, so no change is needed there; the only missing piece was the mask application during verify.

Changes:

  • python/sglang/srt/speculative/dflash_utils.py — remove the grammar-constraint rejection in validate_dflash_request (the return_logprob / return_hidden_states checks are kept).
  • python/sglang/srt/speculative/dflash_worker_v2.py — in the verify block, right after apply_dflash_verify_logits_adjustments, when batch.has_grammar: build the linear-chain topology, call generate_token_bitmask(...), and grammar.apply_vocab_mask(next_token_logits, vocab_mask) before the accept/argmax step. This masks both the greedy-argmax and the sampling-softmax accept paths, since both read next_token_logits.

The masking mirrors EAGLE's grammar-aware verify (eagle_worker_v2.py + eagle_utils.py), so accepted tokens along the chain are grammar-valid by construction, and the downstream spec-v2 grammar FSM advancement stays correct.

Accuracy Tests

Model: Qwen3.6-35B-A3B, --speculative-algorithm DFLASH --speculative-dflash-block-size 8, --grammar-backend xgrammar, --tool-call-parser qwen3_coder (Tested on AMD ROCm / triton attention backend).

  • All previously-400 request shapes now return 200 / well-formed output.
  • Complex nested extraction (document-KIE shape, tool_choice: required + nested object/array schema), 10 trials @ temperature=0, streaming: 10/10 schema-valid, deterministic (1 distinct output), and 7/7 ground-truth field checks (invoice number, grand total, currency, line-item count, seller, SWIFT, Σ line_totals == grand_total).
  • Diverse tool-call suite (all pass): multi-tool auto selection, tool_choice: required with mixed types (int / bool / enum / array), parallel tool calls (3 correct calls for a 3-city prompt), multi-turn with a returned tool result, Vietnamese unicode arguments preserved, streaming reassembly, and the no-tool-needed case (model does not emit a spurious call).
  • Server log is clean: no accept_token failed, no traceback / abort.

Tests added:

  • Unit (CPU CI) — test/registered/unit/spec/test_spec_utils_traverse_tree.py: two cases for the linear-chain topology that DFlashWorkerV2 feeds to generate_token_bitmask (retrieve_next_token = [1, 2, …, -1], no siblings) — one asserting the whole chain is walked in order, one asserting traversal stops when a chain token is disallowed by the grammar.
  • Integration — test/registered/spec/dflash/test_dflash.py::test_grammar_constrained_decoding: launches the DFLASH server and sends a response_format: json_schema request (asserts schema-valid JSON) and a regex request (asserts the output matches the pattern). Runs across all DFLASH server configs in the file (page size, chunked prefill, no-cuda-graph, spec-v2, plan-stream).

Benchmarking and Profiling

DFLASH remains genuinely active on grammar-constrained requests — it is not a silent fallback to non-speculative decoding:

  • spec_accept_rate ≈ 0.79, spec_accept_length ≈ 6.5 / 8, single-request decode throughput ~620–660 tok/s.

So grammar requests keep the speculative-decoding speedup while now producing correct, constrained output.

Reproduce

Launch the server with DFLASH + a grammar backend (--log-requests so per-request spec metrics are printed):

python3 -m sglang.launch_server \
  --model-path Qwen/Qwen3.6-35B-A3B \
  --speculative-algorithm DFLASH \
  --speculative-draft-model-path modal-labs/Qwen3.6-35B-A3B-DFlash \
  --speculative-dflash-block-size 8 \
  --grammar-backend xgrammar \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 \
  --attention-backend triton \
  --log-requests --log-requests-level 1 \
  --host 0.0.0.0 --port 30000

Send a grammar-constrained request (before this PR: HTTP 400; after: schema-valid output):

curl -s http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3.6-35B-A3B",
    "messages": [{"role": "user", "content": "What is the capital of France? Reply as JSON."}],
    "max_tokens": 128, "temperature": 0,
    "chat_template_kwargs": {"enable_thinking": false},
    "response_format": {"type": "json_schema", "json_schema": {"name": "ans",
      "schema": {"type": "object", "properties": {"capital": {"type": "string"}},
                 "required": ["capital"], "additionalProperties": false}}}
  }'
# -> {"capital": "Paris"}

Read the speculative acceptance for that request from the server log (confirms DFLASH is still speculating under grammar, not falling back to 1 token/step):

# a `Finish:` line reports e.g. spec_accept_length ~6.5, spec_accept_rate ~0.79
grep "spec_accept_length" server.log | tail -1

Aggregate metrics are also available via --enable-metrics (sglang:spec_accept_length in Prometheus). Swapping response_format for tool_choice: "required", a named tool_choice, response_format: {type: json_object}, or a top-level "regex": "..." exercises the other previously-rejected grammar paths.

Checklist


CI States

Latest PR Test (Base): ⏳ Run #30155825631
Latest PR Test (Extra): ❌ Run #30155825531

hsthe29 added 2 commits July 4, 2026 04:31
DFLASH previously rejected any grammar-constrained request (json_schema /
regex / ebnf / structural_tag -- i.e. tool_choice=required, named tool_choice,
response_format=json_*) with a 400 error, because the DFLASH verify path never
applied the grammar vocab mask to the target logits.

DFLASH verify is a linear draft chain, which is a degenerate case of EAGLE's
draft tree, so we reuse the existing EAGLE helper generate_token_bitmask() by
building a chain topology (retrieve_next_token=[1,2,...,-1], no siblings) and
apply the resulting vocab mask to the target logits before the accept/argmax
step. Grammar FSM advancement over committed tokens is already handled
generically for spec-v2 in
SchedulerOutputProcessorMixin._resolve_spec_v2_tokens, so no change is needed
there.

Verified on Qwen3.6-35B-A3B + DFLASH (block-size 8): forced/named tool calls,
response_format json_object/json_schema, and regex all return valid,
schema-correct output, and DFLASH stays active on these requests
(spec_accept_length ~6.5/8) -- i.e. not a silent non-spec fallback.
Unit (CPU): cover the linear-chain topology DFlashWorkerV2 feeds to
generate_token_bitmask -- one case walks the full chain in order, one stops
traversal when a chain token is disallowed by the grammar.

Integration: TestDFlashServerBase.test_grammar_constrained_decoding sends a
response_format=json_schema request (asserts schema-valid JSON) and a regex
request (asserts the output matches), across all DFLASH server configs.
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@hsthe29

hsthe29 commented Jul 4, 2026

Copy link
Copy Markdown
Contributor Author

Hi @Ying1123 @merrymercy @hnyls2002 @Qiaolin-Yu, could you please trigger CI for this PR? Thank you!

@shanemort1982

Copy link
Copy Markdown
Contributor

Heads up from deploying this in production. We rolled DSpark onto DeepSeek-V4-Flash (TP2) and DeepSeek-V4-Pro (TP8) on B300 and this PR is the exact fix we needed, but as-is it is incomplete for DSPARK and silently corrupts output rather than fixing it. Two issues:

1. The guard removal covers DSPARK, the masking does not.
The grammar check is deleted from the shared validate_dflash_request(), which the scheduler gates on is_dflash_family() (is_dflash() or is_dspark()), so it stops rejecting DSPARK requests too. But the masking is only added to DFlashWorkerV2.forward_batch_generation. DSPARK runs DSparkWorkerV2, a sibling under BaseSpecWorker with its own _forward_decode, so no mask is ever applied. Net result on DSPARK: grammar requests are accepted and return unconstrained output. We measured 20/20 response_format: json_schema requests coming back as truncated invalid JSON. A silent correctness failure is worse than the original 400.

2. Likely off-by-one in the DFlash masking itself.
generate_token_bitmask is built over draft_tokens (block_size columns), but the TARGET_VERIFY forward returns logits over the full chain verify_ids_2d = cat([anchor, drafts]) = block_size + 1 columns. The bitmask rows and the logit rows do not line up, so the FSM advances over positions that were never committed. The symptom is exactly the mid-value truncation we saw. Suggest building the mask over verify_ids_2d.

Our DSPARK fix (validated in production, 0/20 -> 20/20 schema-valid, GSM8K unchanged, speculation still active): three changes to DSparkWorkerV2._forward_decode -

  • build the bitmask over verify_ids_2d so mask rows align 1:1 with the target logit rows;
  • force the eager path when batch.has_grammar (the folded DsparkVerifyEpilogue runs accept inside the CUDA graph and silently ignores a mask applied to next_token_logits in Python);
  • apply the mask before accept_and_finalize, mirroring EAGLE.

Running live on real client traffic (response_format json_schema / json_object, tool_choice: required / named) with zero grammar rejections. Happy to open a PR with the DSPARK side on top of this, or fold it in here, whichever you prefer.

cc @hsthe29 @merrymercy @Ying1123 @Qiaolin-Yu

@hsthe29

hsthe29 commented Jul 20, 2026

Copy link
Copy Markdown
Contributor Author

Heads up from deploying this in production. We rolled DSpark onto DeepSeek-V4-Flash (TP2) and DeepSeek-V4-Pro (TP8) on B300 and this PR is the exact fix we needed, but as-is it is incomplete for DSPARK and silently corrupts output rather than fixing it. Two issues:

1. The guard removal covers DSPARK, the masking does not.
The grammar check is deleted from the shared validate_dflash_request(), which the scheduler gates on is_dflash_family() (is_dflash() or is_dspark()), so it stops rejecting DSPARK requests too. But the masking is only added to DFlashWorkerV2.forward_batch_generation. DSPARK runs DSparkWorkerV2, a sibling under BaseSpecWorker with its own _forward_decode, so no mask is ever applied. Net result on DSPARK: grammar requests are accepted and return unconstrained output. We measured 20/20 response_format: json_schema requests coming back as truncated invalid JSON. A silent correctness failure is worse than the original 400.

2. Likely off-by-one in the DFlash masking itself.
generate_token_bitmask is built over draft_tokens (block_size columns), but the TARGET_VERIFY forward returns logits over the full chain verify_ids_2d = cat([anchor, drafts]) = block_size + 1 columns. The bitmask rows and the logit rows do not line up, so the FSM advances over positions that were never committed. The symptom is exactly the mid-value truncation we saw. Suggest building the mask over verify_ids_2d.

Our DSPARK fix (validated in production, 0/20 -> 20/20 schema-valid, GSM8K unchanged, speculation still active): three changes to DSparkWorkerV2._forward_decode -

  • build the bitmask over verify_ids_2d so mask rows align 1:1 with the target logit rows;
  • force the eager path when batch.has_grammar (the folded DsparkVerifyEpilogue runs accept inside the CUDA graph and silently ignores a mask applied to next_token_logits in Python);
  • apply the mask before accept_and_finalize, mirroring EAGLE.

Running live on real client traffic (response_format json_schema / json_object, tool_choice: required / named) with zero grammar rejections. Happy to open a PR with the DSPARK side on top of this, or fold it in here, whichever you prefer.

cc @hsthe29 @merrymercy @Ying1123 @Qiaolin-Yu

Thanks for the detailed investigation and for validating this in production!

Feel free to cherry-pick my commit and build the DSPARK changes on top of it. That should preserve the original commit history/authorship while letting you iterate on the DSPARK-specific fixes. I'd be happy to review the follow-up PR once it's ready.

@hsthe29

hsthe29 commented Jul 20, 2026

Copy link
Copy Markdown
Contributor Author

@shanemort1982

@shanemort1982

Copy link
Copy Markdown
Contributor

Opened the DSPARK counterpart as #31753.

Following on from my note above: this PR lifts the grammar guard out of the
shared validate_dflash_request (gated on is_dflash_family(), so it covers
both DFLASH and DSPARK) but only masks in DFlashWorkerV2. A DSPARK deployment
routes to the sibling DSparkWorkerV2, so with the guard gone it would accept
grammar requests and then return unconstrained output. #31753 keeps the guard
for DFLASH and adds the masking in DSparkWorkerV2 (linear verify chain, mask
built over verify_ids_2d so rows line up, plus a has_grammar exclusion from
the folded epilogue). Validated on a DeepSeek-V4 DSPARK deployment: 0/20 to 20/20
schema-conforming.

Happy to reconcile whichever of the two lands first; the only overlap is the one
if in validate_dflash_request.

@hnyls2002

Copy link
Copy Markdown
Collaborator

/rerun-test test_dflash.py test_spec_utils_traverse_tree.py

@github-actions

github-actions Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_dflash.py test_spec_utils_traverse_tree.py:

🚀 1-gpu-5090 (1 test): ✅ View workflow run

cd test/ && python3 registered/spec/dflash/test_dflash.py

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/spec/test_spec_utils_traverse_tree.py

@hnyls2002

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Jul 25, 2026
@hnyls2002
hnyls2002 merged commit d021990 into sgl-project:main Jul 25, 2026
159 of 194 checks passed
1a1a11a added a commit to HarvardMadSys/hybridInference that referenced this pull request Aug 7, 2026
…rks (#1230)

## Problem

Every grammar-constrained request to the local DeepSeek-V4-Flash H200
replicas is rejected by sglang `v0.5.16`:

```
DFLASH speculative decoding does not support grammar-constrained decoding yet.
```

The guard is `validate_dflash_request()`. `is_dflash_family()` is
`is_dflash() or is_dspark()`, so our `"speculative_algorithm": "DSPARK"`
is in scope despite the DFLASH wording. Verified against
`localhost:8005`:

| request | v0.5.16 |
|---|---|
| `tool_choice: "auto"` | ✅ |
| `tool_choice: "required"` | ❌ 400 |
| `response_format: json_object` | ❌ 400 |
| `response_format: json_schema` | ❌ 400 |
| `regex` / `ebnf` / `structural_tag` | ❌ 400 |

On the **streaming** path sglang emits that 400 as an in-band SSE
`error` frame under an **HTTP 200**, then a placeholder `usage` of
`prompt_tokens: 1, completion_tokens: 1`, then `[DONE]`. The gateway
logs it as a success (`status_code=200`, `error` NULL, tokens 1/1,
billed) and never fails over — even though the DeepSeek API and Ollama
Cloud routes for this same model both handle `json_object` fine.

Measured on prod over six hours: **81,880 of 81,945** `response_format`
requests to the sglang routes came back this way, ~100%, across 2 users.

## Fix


[sgl-project/sglang#30096](sgl-project/sglang#30096)
makes the rejection conditional on
`spec_algorithm.supports_grammar_overlap()` (true for the DFlash family)
and adds the verify-time grammar bitmask.

It merged to `main` on **2026-07-25 11:36 UTC** — eleven hours after
`v0.5.16` was cut at **00:13 UTC** the same day. `git compare
v0.5.16...d021990` → diverged, not an ancestor. v0.5.16 is still the
latest release, so the only way to get the fix today is a nightly.

This pins `lmsysorg/sglang:nightly-dev-20260806-ae5f8c94` — 456 commits
past the fix, 0 behind.

## Rollback plan

Revert the `sglang_image` line and restart the units. **Move back to a
tag when v0.5.17 ships** (~2026-08-08 on the fortnightly cadence: 0.5.13
Jun 13 → 0.5.14 Jun 26 → 0.5.15 Jul 10 → 0.5.16 Jul 25). A nightly is
456 unreviewed commits of drift on a 1M-context MoE — a bridge, not a
destination.

## Rollout

Replica B derives its config from this file, so both h200a replicas pick
it up on their next unit restart. Rolling them one at a time: restart
`h200_idle_proxy` (8003), soak, then `h200_idle_proxy_b` (8005). h200b
(8004) stays on v0.5.16 as a control until the soak passes.

## Not in this PR

The gateway-side bug — an in-band SSE `error` frame is treated as a
successful stream, so no failover and bogus 1/1 usage gets logged and
billed — is independent and still open. It will bite again for any
upstream that reports errors this way.

## Test

`tests/test_local_deployment_proxy.py` +
`tests/test_h200_hicache_config.py`: 111 passed.

Co-authored-by: Juncheng Yang <Juncheng Yang>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
1a1a11a added a commit to HarvardMadSys/hybridInference that referenced this pull request Aug 8, 2026
Supersedes the nightly pinned in #1230, before it was ever deployed.

## Why now

`v0.5.17` shipped **2026-08-08 00:19 UTC** — the fortnightly cadence
#1230 predicted. It is the first release carrying the DFlash grammar
fix:

```
$ gh api repos/sgl-project/sglang/compare/d021990...v0.5.17
{"status":"ahead","ahead_by":398,"behind_by":0}
```

`behind_by: 0` →
[sgl-project/sglang#30096](sgl-project/sglang#30096)
is an ancestor of the tag.

By that release the guard is not merely conditional, it is **gone**:

```python
# v0.5.17 python/sglang/srt/speculative/dflash_utils.py
def validate_dflash_request(req: Req, enable_overlap: bool) -> Optional[str]:
    if req.return_logprob: ...
    if enable_overlap and req.return_hidden_states: ...
    return None          # grammar rejection block removed
```

`is_dflash_family()` is still `is_dflash() or is_dspark()`, so our
`DSPARK` config is in scope.

## Why this is strictly better than the nightly

|  | `nightly-dev-20260806-ae5f8c94` | `v0.5.17` |
|---|---|---|
| commits past the fix | 456 | **398** |
| release testing | none | full |
| follow-up needed | yes — "move back to a tag" | none |

#1230's own rollback note said to do exactly this when v0.5.17 shipped.

## Cost of the swap: zero

Nothing was deployed on the nightly. All three replicas still return the
400 as of this PR, so this replaces one pending rollout with another
rather than adding a second ~14-minute cold start per replica.

## Rollout (unchanged from #1230)

Replica B derives its config from this file, so both h200a replicas
follow on their next unit restart. Roll one at a time: `h200_idle_proxy`
(8003) first — it is currently taking ~0.2% of traffic against 8005's
~99.7% — soak, then `h200_idle_proxy_b` (8005). h200b (8004) stays on
v0.5.16 as a control.

`docker pull lmsysorg/sglang:v0.5.17` before the restart; the proxy has
no pull step, so otherwise `docker run` fetches it inline and stretches
the cold start.

## Test

`tests/test_local_deployment_proxy.py` +
`tests/test_h200_hicache_config.py`: 111 passed.

Co-authored-by: Juncheng Yang <Juncheng Yang>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Atituiset pushed a commit to Atituiset/sglang that referenced this pull request Sep 10, 2026
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 24, 2026
…mmar with DFlash (group D); sgl-project#31488/sgl-project#32409 checked

Applicability (27B weg2 serving): the front forwards the whole client body to
D in leg 2, and D validated every request with validate_dflash_request,
which ABORTED return_logprob and every grammar (json_schema / regex / ebnf /
structural_tag). An Anthropic tool_choice any|tool becomes OpenAI
required/named and a structure constraint (serving_chat), so such agent
requests and every logprob request died on D. Full-feature serving -> port.

sgl-project#33459 (36853b8): validate_dflash_request admits return_logprob for DFLASH
(DSpark keeps the refusal, as upstream); the ValueError at the top of
DFlashWorkerV2.forward_batch_generation is gone; the verify computes the
committed run's logprobs with the fork's compute_spec_v2_logprobs
(output_indices = arange(bs*block).view(bs, block), steps = block-1: row j
predicts out_tokens[:, j]) -> the spec-v2 layout batch_result_processor.
_apply_decode_logprobs already reads (next_token_logprobs[i][:commit_len]).

sgl-project#30096 (d021990), ADAPTED: upstream builds a GrammarTree and overlaps the
FSM advance via the sgl-project#31488 grammar barrier; the fork has neither and runs
EAGLE v2 grammar on the synchronous generate_token_bitmask path (overlap is
already disabled for spec+grammar decode batches by
is_disable_overlap_for_batch, and the result processor already advances the
grammar for any spec-v2 result). So: DFlashVerifyInput.grammar field; the
rank-synced draft tokens go to the host before the verify launch (grammar
batches only); after the forward, _dflash_grammar_vocab_mask walks the
LINEAR block (next = i+1, no siblings) with generate_token_bitmask, clears a
stale sampling_info.vocab_mask (as EAGLE does), and the mask is applied to
the verify logits after apply_dflash_verify_logits_adjustments and before
any accept path (greedy/Triton, sampling, selector). Masks are identical on
every rank (synced tokens, identical grammar states); accepts stay synced.
validate_dflash_request admits grammar for DFLASH, DSpark keeps refusing.
The scheduler passes its spec_algorithm (new optional 3rd arg).

sgl-project#31488 (e7e8aaa, overlap grammar with spec verify: grammar barrier,
GrammarTree, eagle/frozen-kv/multi-layer workers, 214 lines) and sgl-project#32409
(3da1071, GrammarMask type across 16 files) are checked and NOT ported:
throughput/refactor over infrastructure the fork does not have; the
synchronous form above is correct without them.

Risk: none for requests without logprobs/grammar (both paths are gated on
batch.return_logprob / batch.has_grammar). Grammar batches pay one D2H of
the draft block and a CPU chain walk per verify step.

Tests (CPU, incg 3 GiB cgroup):
- new test_dflash_logprobs_grammar_33459_30096.py (7): base 1e49837
  6 failed / 1 passed (the unchanged hidden-states refusal), with the port
  7 passed.
- base == port: test_spec_utils_traverse_tree 1/1,
  test_batch_result_processor_spec_grammar 2/2; test_dflash_*.py group
  66 passed (base 61 + the 5 red sgl-project#37818 tests of this branch).
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 24, 2026
…ec ring

27B ReplaySSM package, slice S3 of REPLAYSSM_PLAN.md. With the flag off every
path is unchanged (no rows planned -> recurrent route, intermediate commit).

Ported: upstream GDNAttnBackend._replayssm_target_verify (same kernel call,
launch_mode="verify", null block -1, request rows as replay indices) and the
compact commit sequence of upstream's spec_utils (commit_gdn_replayssm_spec +
commit_gdn_replayssm_circular + conv-window rollback).

Adapted (this line's own wiring; upstream refuses GDN + DFLASH):
- ForwardMetadata.replayssm_spec_rows: the verify metadata plans the request
  rows (req_pool_indices) in all three builders (eager, graph capture, graph
  replay -- there the static buffer with padded rows zeroed). The commit reads
  the rows the verify used from the metadata, like mamba_cache_indices, so no
  caller changes: DFLASH (dflash_worker_v2 untouched), the EAGLE/MTP commit
  and the lane all go through update_mamba_state_after_mtp_verify.
- The commit sits in that shared hook (ring branch when intermediate_ssm is
  None). Linear chain -> accepted = last step + 1 (bonus included); -1 folds
  nothing, like the masked scatter.
- FOLD EVERY COMMIT for any SSM dtype (upstream defers the fold for fp32):
  radix insert, HiCache backup, tail adopt and the flip all read `temporal`
  and assume it is the committed state after every verify, as the recurrent
  route makes it. The ring is then per-step scratch (write_pos 0 at every
  verify).
- Heal: outside a capture the verify metadata re-states write_pos = is_flush
  = 0 for the step's rows (true by construction after every commit), so a
  verify never trusts cursor bytes from an earlier step (TMS restore, flip,
  reused row). cache_base is circular and harmless at write_pos 0.
- forward_extend routes the verify to the ring when rows are planned; a draft
  tree on the ring, or a pool with neither intermediate state nor planned
  rows, is refused. The ring kernel honours token strides, so the verify split
  stays torch.split views (sgl-project#33778).
- Advance cursor null block 0 (request row 0 is the ReqToTokenPool padding
  row); the widest verify window is kept on the pool as the advance's
  constant (one compiled variant across the adaptive ladder).
- Ring length must be >= 16 (the compact commit's tl.dot) -- pool and static
  check; S2 test moved to L=16.
- _replayssm_spec_for: a draft-KV-only producer (no verify workspace, sgl-project#1233
  FIX 4) treats the flag as a no-op instead of hitting the pool's
  missing-window refusal.
- getattr on the new metadata field in the two readers (sgl-project#624 stub drift:
  test_gdn_verify_strided_qkv_33778 doubles predate it).

Tests (CPU, incg cgroup, one file at a time):
test/registered/unit/layers/attention/test_replayssm_spec_route_s3.py: 10
passed. Hermetic: rows planned + healed (eager, replay; not in capture;
padded row -> row 0), flag off / decode plan nothing, the three refusals,
producer no-op. Interpreter (production route functions on real CPU pools,
ring vs intermediate, 2 layers, head ratio 3, 2 requests + padded row, track
crossing, second step after garbage cursors + heal): fp32 verify rel 3.7e-7 /
2.4e-7, commit (incl. track slot) 2.9e-7 / 3.1e-7 vs the recurrent route;
fp16 (bf16 stand-in) verify 5.0e-4 vs recurrent 3.8e-4 of the fp32 truth,
commit 3.6e-4 == recurrent 3.6e-4; conv rollback equal; untouched slots
bit-exact; cursors 0 after every commit. On the S2 sources: 7 failed + 3
errors.
Regression: test_replayssm_spec_ring_pool_s2 10, test_mamba_checkpoint_interval
58, test_dflash_mamba_track_post_verify_37818 5, test_conv_verify_private_
window_444 13, test_prefill_graph_stale_track_rows_34184 2,
test_weg2_prefill_only_capture_1233 15, test_forward_metadata_plan_record 9,
test_fused_replay_state_indices_32219 3, test_gdn_verify_strided_qkv_33778 4,
test_verify_intermediate_row_ownership_450 12, test_mamba2_conv_verify_
private_window_450 10, test_dual_group_concurrency 159 -- all passed.
ruff F: no new findings.

27B line (desk/27b-up-replayssm-line-0924): cherry-picked from 7d69bd8
(desk/27b-up-replayssm-0924, on the NF-based staging line fa757eb), clean.
gdn_backend.py and mamba2_metadata.py are identical on both lines;
dflash_worker_v2 differs here (sgl-project#31468, sgl-project#33459/sgl-project#30096 ports) but not in
_update_target_mamba_state_after_verify, so the DFLASH commit reaches the
shared hook exactly as on NF. One adaptation, text only: the commit
docstring listed "tail adopt" among the readers of `temporal` -- that is the
NF line's H21 install (weg2.tail_adopt), absent here; the 27B readers are the
radix insert, the HiCache backup (the weg2 L2 arena write, xsn351) and the
P/D flip.

Tests on this line (CPU, incg cgroup, S3 tree, one file at a time):
test_replayssm_spec_route_s3.py 10 passed -- interpreter, production route vs
the recurrent route: fp32 verify rel 3.74e-7 / 2.40e-7 (step 1 / step 2 after
garbage cursors + heal), commit incl. the track slot 2.94e-7 / 3.10e-7 (all <
4e-7); fp16 (bf16 stand-in) verify 5.03e-4 / 2.72e-4 vs recurrent 3.84e-4 /
2.55e-4 of the fp32 truth, commit 3.61e-4 / 3.27e-4 == recurrent; conv
rollback equal, padded outputs zero, untouched slots bit-exact, cursors 0
after every commit; test_replayssm_spec_ring_pool_s2.py 10 passed (8
subtests).

Regression (package tip = S6 tree, same cgroup, one file at a time):
test_dflash_mamba_track_post_verify_37818 5,
test_conv_verify_private_window_444 13,
test_prefill_graph_stale_track_rows_34184 2,
test_weg2_prefill_only_capture_1233 15, test_forward_metadata_plan_record 9,
test_fused_replay_state_indices_32219 3, test_gdn_verify_strided_qkv_33778 4,
test_verify_intermediate_row_ownership_450 12,
test_mamba2_conv_verify_private_window_450 10, test_dual_group_concurrency
159, test_mamba_checkpoint_interval 58 -- all passed.
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 25, 2026
…ec ring

27B ReplaySSM package, slice S3 of REPLAYSSM_PLAN.md. With the flag off every
path is unchanged (no rows planned -> recurrent route, intermediate commit).

Ported: upstream GDNAttnBackend._replayssm_target_verify (same kernel call,
launch_mode="verify", null block -1, request rows as replay indices) and the
compact commit sequence of upstream's spec_utils (commit_gdn_replayssm_spec +
commit_gdn_replayssm_circular + conv-window rollback).

Adapted (this line's own wiring; upstream refuses GDN + DFLASH):
- ForwardMetadata.replayssm_spec_rows: the verify metadata plans the request
  rows (req_pool_indices) in all three builders (eager, graph capture, graph
  replay -- there the static buffer with padded rows zeroed). The commit reads
  the rows the verify used from the metadata, like mamba_cache_indices, so no
  caller changes: DFLASH (dflash_worker_v2 untouched), the EAGLE/MTP commit
  and the lane all go through update_mamba_state_after_mtp_verify.
- The commit sits in that shared hook (ring branch when intermediate_ssm is
  None). Linear chain -> accepted = last step + 1 (bonus included); -1 folds
  nothing, like the masked scatter.
- FOLD EVERY COMMIT for any SSM dtype (upstream defers the fold for fp32):
  radix insert, HiCache backup, tail adopt and the flip all read `temporal`
  and assume it is the committed state after every verify, as the recurrent
  route makes it. The ring is then per-step scratch (write_pos 0 at every
  verify).
- Heal: outside a capture the verify metadata re-states write_pos = is_flush
  = 0 for the step's rows (true by construction after every commit), so a
  verify never trusts cursor bytes from an earlier step (TMS restore, flip,
  reused row). cache_base is circular and harmless at write_pos 0.
- forward_extend routes the verify to the ring when rows are planned; a draft
  tree on the ring, or a pool with neither intermediate state nor planned
  rows, is refused. The ring kernel honours token strides, so the verify split
  stays torch.split views (sgl-project#33778).
- Advance cursor null block 0 (request row 0 is the ReqToTokenPool padding
  row); the widest verify window is kept on the pool as the advance's
  constant (one compiled variant across the adaptive ladder).
- Ring length must be >= 16 (the compact commit's tl.dot) -- pool and static
  check; S2 test moved to L=16.
- _replayssm_spec_for: a draft-KV-only producer (no verify workspace, sgl-project#1233
  FIX 4) treats the flag as a no-op instead of hitting the pool's
  missing-window refusal.
- getattr on the new metadata field in the two readers (sgl-project#624 stub drift:
  test_gdn_verify_strided_qkv_33778 doubles predate it).

Tests (CPU, incg cgroup, one file at a time):
test/registered/unit/layers/attention/test_replayssm_spec_route_s3.py: 10
passed. Hermetic: rows planned + healed (eager, replay; not in capture;
padded row -> row 0), flag off / decode plan nothing, the three refusals,
producer no-op. Interpreter (production route functions on real CPU pools,
ring vs intermediate, 2 layers, head ratio 3, 2 requests + padded row, track
crossing, second step after garbage cursors + heal): fp32 verify rel 3.7e-7 /
2.4e-7, commit (incl. track slot) 2.9e-7 / 3.1e-7 vs the recurrent route;
fp16 (bf16 stand-in) verify 5.0e-4 vs recurrent 3.8e-4 of the fp32 truth,
commit 3.6e-4 == recurrent 3.6e-4; conv rollback equal; untouched slots
bit-exact; cursors 0 after every commit. On the S2 sources: 7 failed + 3
errors.
Regression: test_replayssm_spec_ring_pool_s2 10, test_mamba_checkpoint_interval
58, test_dflash_mamba_track_post_verify_37818 5, test_conv_verify_private_
window_444 13, test_prefill_graph_stale_track_rows_34184 2,
test_weg2_prefill_only_capture_1233 15, test_forward_metadata_plan_record 9,
test_fused_replay_state_indices_32219 3, test_gdn_verify_strided_qkv_33778 4,
test_verify_intermediate_row_ownership_450 12, test_mamba2_conv_verify_
private_window_450 10, test_dual_group_concurrency 159 -- all passed.
ruff F: no new findings.

27B line (desk/27b-up-replayssm-line-0924): cherry-picked from 7d69bd8
(desk/27b-up-replayssm-0924, on the NF-based staging line fa757eb), clean.
gdn_backend.py and mamba2_metadata.py are identical on both lines;
dflash_worker_v2 differs here (sgl-project#31468, sgl-project#33459/sgl-project#30096 ports) but not in
_update_target_mamba_state_after_verify, so the DFLASH commit reaches the
shared hook exactly as on NF. One adaptation, text only: the commit
docstring listed "tail adopt" among the readers of `temporal` -- that is the
NF line's H21 install (weg2.tail_adopt), absent here; the 27B readers are the
radix insert, the HiCache backup (the weg2 L2 arena write, xsn351) and the
P/D flip.

Tests on this line (CPU, incg cgroup, S3 tree, one file at a time):
test_replayssm_spec_route_s3.py 10 passed -- interpreter, production route vs
the recurrent route: fp32 verify rel 3.74e-7 / 2.40e-7 (step 1 / step 2 after
garbage cursors + heal), commit incl. the track slot 2.94e-7 / 3.10e-7 (all <
4e-7); fp16 (bf16 stand-in) verify 5.03e-4 / 2.72e-4 vs recurrent 3.84e-4 /
2.55e-4 of the fp32 truth, commit 3.61e-4 / 3.27e-4 == recurrent; conv
rollback equal, padded outputs zero, untouched slots bit-exact, cursors 0
after every commit; test_replayssm_spec_ring_pool_s2.py 10 passed (8
subtests).

Regression (package tip = S6 tree, same cgroup, one file at a time):
test_dflash_mamba_track_post_verify_37818 5,
test_conv_verify_private_window_444 13,
test_prefill_graph_stale_track_rows_34184 2,
test_weg2_prefill_only_capture_1233 15, test_forward_metadata_plan_record 9,
test_fused_replay_state_indices_32219 3, test_gdn_verify_strided_qkv_33778 4,
test_verify_intermediate_row_ownership_450 12,
test_mamba2_conv_verify_private_window_450 10, test_dual_group_concurrency
159, test_mamba_checkpoint_interval 58 -- all passed.

NF line (H64, onto 88dec66): one conflict, resolved line by line in
gdn_backend.forward_extend. This line carries no sgl-project#33778 port (no
kernel_dispatcher.target_verify_supports_strided_qkv), so the 27B hunk that
widens use_strided_verify_qkv for the ring has nothing to widen. Kept the NF
condition `(is_cuda() or is_hip()) and qkv_dim <= MAX_FUSED_QKV_SPLIT_DIM`:
the ring verify gets the same q/k/v preparation as the recurrent verify on
this line (dense [1, T, H, D] from fused_qkv_split_gdn_prefill on CUDA,
torch.split views elsewhere); _replayssm_target_verify flattens both layouts
by numel. Text only: the commit docstring names the NF H21 tail adopt
(weg2/tail_adopt.py reads and installs cache.temporal) among the readers of
temporal again -- the 27B line had dropped it because it lacks H21.
Tests (CPU, incg cgroup, one file at a time): test_replayssm_spec_route_s3.py
10 passed; test_replayssm_spec_ring_pool_s2.py 10 passed (8 subtests).

(cherry picked from commit 974254d)
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 26, 2026
…cks, chronologisch)

Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum
(Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst
anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity;
479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 ->
S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name
(Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)).
Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc

Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859
sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579
sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438
(SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool
sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild),
#DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release;
P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6
+ F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49
Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens).

Auflösungen (Gabel -> Form, Grund):
- L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität"
  nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker-
  Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags.
  Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity.
- RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser
  CalibrationIdentity.d_max_running_requests; nextflash unverändert.
- URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und
  Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie).
- scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP
  komponiert (beide Instrumente getrennt schaltbar).
- schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF
  computed_input_len.
- gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab.
- flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt).
- checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation.
- xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen.
- vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read.
- FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9).
- argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only).
- census_checkpoint_decision (W161) jetzt für beide Profile aktiv.

Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber
dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu.
Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants