perf(dflash): capture the draft's per-step context ingest and use fused fp8 kernels - #584
Open
myshytf wants to merge 63 commits into
Conversation
Record scheduler-side speculative widths in GrammarOutput so worker-side draft trimming cannot shift flattened grammar masks onto later requests. Destination logits continue to use the worker-visible width, while source offsets use the serialized scheduler width. Validated with focused unit coverage and a 160-request concurrent DeepSeek V4 structured-output workload.
KimiK3ToolParser.extract_tool_calls_streaming matched calls with _call_re, which requires the closing <|close|>call<|sep|> marker. Until that marker arrived nothing was emitted for the call, so a long tool call produced no SSE deltas for the whole generation and then dumped the entire arguments JSON in one delta. Track the call from its <|open|>call ...<|sep|> marker instead. The name goes out immediately, and _partial_arguments serializes the arguments seen so far as a prefix of the final JSON, so each step can stream the difference against what it already sent. String argument bodies are raw text, so they are forwarded as they arrive with a trailing partial close marker held back; other types still need the whole literal to decode and are held until their block closes. The concatenated deltas are byte-identical to the non-streaming extract_tool_calls output. Signed-off-by: guptaishaan <guptaishaan@users.noreply.github.com>
Withhold whitespace-tolerant argument-close fragments until they form a complete XTML marker. This keeps streamed JSON argument deltas prefix-stable for every marker form accepted by the parser. Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Codex <codex@openai.com>
Document the target model input and optional NeoX layout result using the repository's Google-style docstring contract. This is documentation-only and does not change runtime behavior. Co-authored-by: OpenAI Codex <codex@openai.com>
Initialize fresh assistant generations in the reasoning channel when Kimi thinking is enabled, while preserving rendered marker state for continued assistant messages. Filter complete and split XTML control markers at the composed parser boundary so malformed model transitions cannot expose protocol syntax as API content. The thinking-disabled path and continuation semantics remain unchanged. Validation: 72 Kimi K3 reasoning and tool-parser tests; Ruff format and lint; git diff whitespace validation.
Signed-off-by: jungjiyu <libraryofjiyu@gmail.com> Assisted-by: ChatGPT
Model a 17-group hybrid KV layout and report a load failure from the final group. The test requires failure_policy=fail to finish only the affected request, emit an error result, and schedule a subsequent healthy request.\n\nValidation: 20 KV load-failure tests and 7 hybrid/Mamba scheduler tests pass in the CUDA 13.3 PyTorch 2.13 runtime.
Stop accepting speculative token batches when the grammar matcher reaches its terminal state. Preserve terminal-state tracking across validation and acceptance calls so tokens after a complete structured value cannot be committed. This is the Infernal Invocation backport of vllm-project#52805 commits d8cde608cf1f3de406c75f081a76a0e6eb55a9cb, 1cf6f25351357354cf8c520c0b2976b029429668, and 1856abd22452c3da67364986ece7245fce52c950. Signed-off-by: Martin Vit <martin@voipmonitor.org>
Structured-output masks are prepared before speculative verification. An accepted block can cross reasoning activation or grammar termination, so its suffix may have been sampled under a grammar state that no longer applies at commit time. Validate the accepted block without advancing the matcher, commit only its valid prefix, and roll scheduler accounting back for resampling. Preserve the unstructured and single-token fast paths, and report only committed draft tokens in speculative metrics. Co-authored-by: Adam Moisa <adammoisa@gmail.com> Assisted-by: OpenAI Codex Signed-off-by: Martin Vit <martin@voipmonitor.org> (cherry picked from commit fa0777f) Signed-off-by: Martin Vit <martin@voipmonitor.org>
Infernal Invocation exposes prompt inspection through is_reasoning_end_for_prompt. Make the upstream structured-output regression fixture implement the branch contract so it exercises the production method instead of a stale mock interface. Signed-off-by: Martin Vit <martin@voipmonitor.org>
Type the conditional Kimi compact-RoPE protection scope through the shared context-manager interface. Both the Kimi protection context and the no-op context retain their existing runtime behavior. Signed-off-by: Martin Vit <martin@voipmonitor.org>
The debug branch initializes the event list before every sweep point. Assert that invariant after detaching the list from the model runner so static analysis can verify indexed event access. Profiling and warmup behavior are unchanged. Signed-off-by: Martin Vit <martin@voipmonitor.org>
…DFlash aux state (vllm-project#50487) Signed-off-by: Rahul Chalamala <22563365+rchalamala@users.noreply.github.com> Co-authored-by: Janelle Cai <janelle.cai@modal.com> (cherry picked from commit 03a8d0b)
Verify that disabled AttnRes capture returns before reading unavailable weights and that enabled capture selects both normalization and projection weights from the correct consumer. Document the capture interface parameters and return value.
Compute MoonViT rotary frequencies only for the image grid sizes present in each request instead of materializing the configured 512x512 ceiling. This reduces the measured first-image CUDA allocation peak from 340,018,176 bytes to 1,990,656 bytes for a 36x36 grid while preserving bit-identical CPU and CUDA output. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
Project independent Kimi vision features separately so MXFP8/Marlin workspace scales with the largest image instead of the sum of all scheduled images. Preserve output order, shape, activation dtype, and numerical results while reducing the measured TP16 three-image transient peak by 32.52 MiB. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
Define token-position DCP shard count on each cache specification and use max_num_blocks_per_req as the worker block-table width contract. Attention caches retain full, partial, or replicated DCP layouts; recurrent caches report one token-position shard and preserve their mode-specific table width. This removes the model runner's cache-type special case while retaining the 1,310-column Mamba align table required by a 1,000,000-token model length with 768-token blocks and seven speculative blocks. Assisted-by: OpenAI Codex <noreply@openai.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
Signed-off-by: Martin Vit <martin@voipmonitor.org>
Gather each tensor-parallel vision shard at its produced row count instead of padding every rank to the largest shard. This preserves embedding order and the uniform-size fast path while preventing the transient allocation from scaling with TP size when a request contains fewer images than ranks. Validate zero-length PyNccl inputs, single-image output parity, empty inputs, uneven four-GPU assignments, and multi-image assignments. A TP16 Kimi-K3-shaped harness reduces the collective output from 224 MiB to 14 MiB per GPU with bit-exact gathered content. Signed-off-by: Martin Vit <martin@voipmonitor.org>
Signed-off-by: Martin Vit <martin@voipmonitor.org>
Cache each head's prefix and suffix log-sum-exp values before any output write when the thread group fits inside a CUDA block. This preserves chunked-attention accumulators that pass the running LSE tensor as both prefix input and output destination, while retaining the direct-load path for head groups that cross block boundaries. Index all cached values through the declared tensor strides.\n\nAdd exact in-place versus disjoint-output coverage for the six-head, 128-element MLA geometry at 256 and 4096 tokens.\n\nThe shared-memory loading structure adapts vLLM PR vllm-project#45778 (commit c71576f) to the strided-LSE kernel contract.\n\nCo-authored-by: nicole-lihui <nicole.li@daocloud.io> Signed-off-by: Martin Vit <martin@voipmonitor.org>
… GPU Adds a verifier-side proxy (RemoteK3DSparkSpeculator) and a standalone draft server (vllm.entrypoints.k3_dspark_standalone + k3_dspark_rpc) so the DSpark draft model executes on its own single GPU while the target runs TP/DCP on separate GPUs. Draft weights, KV, Markov head, and CUDA graphs live entirely on the draft process; the target exchanges context and proposals over a versioned ZMQ/TCP protocol (PROTOCOL_VERSION=2). Behavior and invariants: - VLLM_K3_DRAFT_REMOTE_ADDRESS selects the remote path at speculator construction; unset preserves the existing local DSpark/DFlash path. - propose() matches BaseSpeculator's signature; rank 0 performs RPC and all ranks consume the broadcast result. - Fail closed: any RPC failure fills draft tokens with -1 (no speculation for the step) and disables affected requests until they leave the batch; FREE remains safe for never-created remote state. - Retained-prefix reconnection validates a target prefix-cache hit against retained draft state via a host-visible view of the request token table (InputBatch.all_token_ids_cpu, backed by StagedWriteTensor.cpu). - CUDA-graph capture interface preserved: init_cudagraph_manager and capture(capture_phase=...) conform to BaseSpeculator. Compatibility: no change when the remote address is unset; draft side supports DSpark and DFlash checkpoints on a single GPU including Ampere-class cards. Validation: 19 new CPU unit tests pass (test_k3_dspark_remote_speculator.py, test_k3_dspark_standalone.py); production-qualified serving lukealonso/Kimi-K3-QSRT-K2 TP8/DCP8 with an Inferact BF16 DSpark draft on a dedicated RTX 3090. Limitations: one remote draft process (draft TP1); TCP transport; greedy draft sampling with block rejection sampling on the verifier. AI assistance was used in the preparation of this change; every line was reviewed and the listed tests were run by the submitter. Signed-off-by: myshytf <9619163+myshytf@users.noreply.github.com>
Signed-off-by: myshytf <9619163+myshytf@users.noreply.github.com>
Keep DSpark and DFlash scheduling lookahead semantics while applying EAGLE's last-hash target-cache drop only when an actual target KV group is marked as EAGLE. This preserves fine target APC tails for remote/disaggregated drafts and retains the legacy fallback for classic EAGLE. Signed-off-by: myshytf <9619163+myshytf@users.noreply.github.com>
Signed-off-by: myshytf <9619163+myshytf@users.noreply.github.com>
… <=1 output token When a batch contains only new requests (no running ones) and every one has max_tokens <= 1, set num_spec_tokens_to_schedule = 0. Speculative decoding cannot help a 1-token output, so the draft pass and verification are pure overhead. This is the shape of every max_tokens=1 API call, every prefill-throughput benchmark, and every embedding/classification-style request. Measured on RTX 5090 (31.4 GiB), Qwen3.8-27B EXL3, MTP=6: 1-token request latency 141 ms -> 127 ms 2051-token prefill bench 7445 -> 7635 tok/s (+2.5%) TG on normal requests 189.8 tok/s (unchanged) The guard is conservative: it requires scheduled_running_reqs to be empty, so an in-flight multi-token generation can never lose its draft tokens. Signed-off-by: Michel Belleau <michel.belleau@malaiwah.com>
…rmless for single-token requests
Call the finalized FlashInfer workspace prepare API during vLLM graph warmup so autotune and cache lookup complete before CUDA graph capture. Signed-off-by: myshytf <9619163+myshytf@users.noreply.github.com>
The six-layer DFlash draft for Kimi-K3 (hidden 7,168, intermediate 14,336, BF16, 4.9 GB including the 2.35 GB shared lm_head) is weight-bandwidth bound on its dedicated GPU: the target's per-proposal timing showed server_gpu_query 5.5 ms and server_gpu_context 1.4 ms of a 7.3 ms server_total per decode step (K=3). The standalone draft server gains --draft-quantization (vLLM online quantization shorthand for the draft's linear layers; fp8_per_channel is one float8_e4m3 scale per output channel with dynamic per-token activation scaling through torch._scaled_mm) and --draft-fp8-head, which scores proposals with the rowwise-fp8 lm_head copy from vllm/model_executor/layers/fp8_draft_head.py. DFlashQwen3ForCausalLM implements maybe_init_fp8_draft_head / the fp8 compute_logits branch (the same contract as the DSpark draft), load_dflash_model calls the hook after lm_head aliasing and before CUDA-graph capture, and the fused context K/V projection dequantizes the online-fp8 qkv weight back to the model dtype (the transposed [in, out] float8 tensor with [out, 1] scales) instead of slicing the raw parameter. Draft-time only: the target's verification pass never sees these weights, so accepted tokens keep the target distribution (rejection sampling is exact for any proposal distribution, and the proposal logits shipped for probabilistic acceptance are the fp8 draft's own); only the acceptance length can move. Production (TP8/DCP8 target, draft on a ninth RTX PRO 6000, K=3): server_gpu_query 5.5 -> 3.45 ms, server_total 7.3 -> 5.1 ms, rpc_roundtrip 8.1 -> 5.8 ms; mean acceptance length 2.51-2.68 versus 2.45-2.61 before; decode c1 43.6/47.1 -> 46.0/49.2 tok/s (ITL 22.6/20.9 -> 21.3/19.8 ms); draft device memory 22.1 -> 18.6 GB. Smoke-test token unchanged (198).
vllm/model_executor/models/qwen3_dflash.py.orig was committed by mistake alongside the fp8 draft change; it is a copy of the pre-change module and has no consumer.
…ed fp8 kernels The standalone K3 DFlash draft server spends most of a proposal outside its captured query graph. Two changes: k3_dspark_rpc.py: the per-step context append (the accepted rows of one decode step, 1..max_num_seqs x (1 + K) rows, at most one KV block) replays a CUDA graph per row count: raw target auxiliary rows -> combined draft states -> every layer's K/V -> RoPE -> KV cache write at the staged slots. The eager path issued about 30 launches with host gaps between them. Larger appends (prefill-sized, reconnects) and projected rows keep the eager path. K3_DRAFT_INGEST_GRAPH=0 disables the capture. Draft tokens are identical with the graphs on and off. k3_dspark_standalone.py: the draft engine config selects the CUDA per-token fp8 quantization op, the CUTLASS fp8 GEMM (per-token x per-channel scaling in the epilogue) and the vLLM C RMSNorm kernels instead of their torch-native forms, which issued 8-13 tiny kernels per projection (about 600 per proposal) inside the captured graph. The runtime now installs the configured IR op priorities; vLLM workers do this at init and the standalone runtime has no worker, so every IR op had been running torch-native regardless of the configuration. Draft-side only: the target's verification is untouched; draft logits change at fp8-rounding level. Validation (RTX PRO 6000, fp8 per-channel DFlash draft, K=3, second instance on the production draft GPU): proposal 5.89 -> 3.72 ms at two requests x 3 accepted rows (gpu_context 1.15 -> 0.51, gpu_query 4.13 -> 2.81, host_submit 1.85 -> 0.57 ms); 3.42 ms at one request; 3.95 ms at four requests x 4 rows. Token digests equal between the captured and eager ingest over 24 proposals.
|
Warning Review limit reachedNext included review available in 33 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Team Run ID: 📒 Files selected for processing (64)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…ong line `--draft-fp8-head` now sets VLLM_DSPARK_FP8_DRAFT_HEAD in `main()` before `_load_runtime` builds the draft head; `_build_vllm_config` only builds configuration. The behaviour is unchanged: the draft loader reads the variable while the head is created. The fp8 qkv dequantization helper in qwen3_dflash.py keeps its lines within the 88-column limit. Validation: tests/v1/spec_decode/test_k3_dspark_standalone.py (10 passed in the SM120 production image). Co-Authored-By: Claude Code <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HPWxmKzfikaemyykd3p89D
…kimi-k3-draft-ingest-graph-20260902-pr
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #570 (fp8 DFlash draft); the first two commits are that PR.
The standalone Kimi-K3 DFlash draft server (
vllm/entrypoints/k3_dspark_rpc.py,k3_dspark_standalone.py) spends most of a proposal outside its captured query graph. This change:max_num_seqsx (1 + K) rows, at most one KV block) replay a CUDA graph per row count: raw target auxiliary rows -> combined draft states -> every layer's K/V -> RoPE -> KV cache write at the staged slots. The eager path issued about 30 launches with host gaps between them. Prefill-sized appends, reconnects and projected rows keep the eager path.K3_DRAFT_INGEST_GRAPH=0disables the capture.custom_ops: ["none", "+quant_fp8"]), the CUTLASS fp8 GEMM with per-token x per-channel scaling in its epilogue (linear_backend: cutlass) and the vLLM C RMSNorm kernels. The torch-native forms issued 8-13 tiny kernels per projection (about 600 per proposal) inside the captured graph.kernel_config.ir_op_priority.set_default()). vLLM workers do this at init; the standalone runtime has no worker, so every IR op had been running torch-native regardless of the configuration.Draft-side only: the target's verification is untouched; draft logits change at fp8-rounding level (GEMM kernel).
Validation
RTX PRO 6000 (SM120), fp8 per-channel DFlash draft with fp8 head, K=3, second draft instance on the production draft GPU, proposals of the production shape (
research harness: reset context of 512 rows, then steady-state appends):Draft token digests over 24 proposals are identical with the captured ingest on and off (
K3_DRAFT_INGEST_GRAPH=0). The GPU span of a proposal is now 3.4 ms with 6 % idle (was 5.3-6.0 ms with 30-38 % idle); the remainder is the fp8 GEMMs (25 projections 1.71 ms, lm_head 0.89 ms) at about 75 % of HBM bandwidth.🤖 Generated with Claude Code
https://claude.ai/code/session_01HPWxmKzfikaemyykd3p89D