Repository navigation
[FlyDSL][DSv4] Prototype FP4 MQA streaming TopK - #5282
AMD-yanfeiwang wants to merge 4 commits into
Conversation
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
|
MI355X runtime update (gfx950, ROCm 7.2):
Production-grid microbenchmark (
So the prototype removes 96–99% of score-path storage, but is currently about 2.4–2.5x slower at the core operator. That does not meet the default-path performance gate yet; this PR remains Draft while the 8-GPU end-to-end run finishes. The result supports keeping sgl-project/sglang#38086 as the near-term alternative unless further local-selection optimization closes the gap. |
e369bec to
b400855
Compare
…y looked right Pointing the FlyDSL collectors at $WORK/head was correct and insufficient: that tree was populated only inside `if grep -q "invariant-removed\|api-signature"`. On a PR deriving neither family it is an empty directory, and a collector that reads files finds nothing to read. #5207 has six buffer calls in its diff -- five bounded, one a bare `make_buffer_tensor(W_scale, max_size=True)` -- and flydslbounds returned zero. #5301, which the previous commit was verified against, deletes a guard and so derives invariant-removed. That is the only reason the fix appeared to work. A fixture cannot see this: every test here builds its own root and hands it over, so the tree is always populated. Materialising head is now unconditional; only the `evidence` call stays behind the family test. The cost is a few `git show` calls against objects already fetched at line 125. On #5207 B8 now reports exactly one candidate and it is worth asking about: `W_scale` is bound with `num_records_bytes=n_heads * self.n_per_head * 4` at line 184 of the same file, and with `max_size=True` at line 92, inside a BlockScale whose constructor already receives head and n_scale_cols. Same tensor, two treatments, one of them unbounded. Also verified silent for the right reason on #5146 (its one added buffer call is buffer_ops.py forwarding variables, not a kernel binding a tensor) and on #5282 (no buffer construction in the diff at all).
Motivation
The DeepSeek-V4 FP4 indexer currently materializes an FP32
[query_rows, c4_context]score matrix before TopK. At 768K full-token context this is about 0.75 MiB per query row; a 1024-row local prefill is about 768 MiB per C4 layer invocation. SGLang currently avoids allocator fragmentation with a persistent 2 GiB slab (sgl-project/sglang#37660), but that reserves context-sized storage instead of removing the intermediate.This draft adds an AITER operator that computes exact TopK without materializing the full logits matrix.
Changes
flydsl_pa_mqa_topk_fp4_prefillfor gfx950 FP4 paged MQA.(score descending, logical index ascending);+0 > -0is preserved;(candidate_value, raw_index, valid_count).length <= topk, return every logical position sequentially and pad with-1.For the common SGLang prefill shape
rows=1024, parallel_unit_num=1024, topk=1024, internal candidate + merge scratch is about 20 MiB, independent of context length, versus roughly 768 MiB of logits at 768K context. No device-global context-sized allocation is introduced.Scope
Initial production specialization:
Graph capture is rejected explicitly in this first version; SGLang will retain its existing graph path as fallback. A dependent SGLang draft will capability-probe this API and keep old AITER compatibility. sgl-project/sglang#38086 remains an independent workspace-based alternative.
Validation
Completed:
Queued on MI355X:
This stays draft until those GPU and end-to-end long-context results are attached.