Fix DFlash sampling bias and GDN extreme-decay state corruption - #8
Terrydaktal wants to merge 2 commits into
Conversation
The DFlash2 selector used the same Philox noise stream as the target's replacement draw after rejection. This conditions replacements on the rejected proposal and changes the intended output distribution. Adapt the independent draft stream from vllm-project/vllm#54282 (fe755c88995ad468882517b6c4bdd60138d46a3a) to the selector's direct Gumbel call. Offset only its local RNG index, preserve model/cache positions and greedy behavior, and install the guarded idempotent patch in the release, development and patch-image build paths. The pinned libr4d GDN scan's midpoint factorization clamps growing exponents at 80. Large decay spans can erase valid contributions and final recurrent state without producing nonfinite values. Extend the existing libr4d build patch with a bounded FP32 recurrence for affected sequence/head pairs, using a conservative span threshold of 128. Preserve ordinary fast-scan outputs, the entry-point ABI, caller stream, tensor layouts and decode path. Propagate the original launch error before launching the correction. No host readback, synchronization or additional temporary tensor is introduced. Add public, checkpoint-free GPU regressions for native selector/rejection sampling and an independent FP64 GDN recurrence. Include ordinary and extreme decays, unequal sequences, partial chunks, mixed heads, large values, both 48/16 and 24/8 head layouts, and graph replays with changed inputs. Record compact numeric qualification evidence and document reproduction, provenance, the need to rebuild old computed state, and remaining performance limits. Link the correction documentation from the README. Validation on R9700/gfx1201 with the pinned Radiance 1.0.16 image: - Patch applies to pinned libr4d source; both scan units compile with its production compiler flags; sampler installation and repeat installation pass. - At three positions with 200,000 draws per arm, shared-noise probability error is 0.01781-0.01953 versus 0.000965-0.001180 after correction. Target-only and greedy controls pass. - Original GDN reproduces 46.4% and 82.8% output error and 100% state error on two extreme-decay cases. Both corrected layouts pass all 12 cases and three graph replays, with output error below 0.345% and state error below 0.252% against FP64. This commit is limited to the two numerical defects and their validation. End-to-end throughput, distributed TP2 and repetition-rate improvements are not claimed. The GDN correction adds prefill work that still needs workload performance qualification. No loop guard, history recovery or serving-policy change is included.
Correct the numerical-corrections documentation: upstream #54282 already fixes the same DFlash2 selector through IS_DRAFTING=True. The pinned vLLM v0.28.0 helper predates that API, so our patch applies the same salt directly to its local RNG index. Remove wording suggesting that current upstream still needs a separate selector correction. Validation: inspected the merged upstream selector and Gumbel-helper diffs against the pinned API. This changes documentation only; the qualified arithmetic corrections and recorded results are unchanged.
|
Thank you for submitting this PR! I will be taking a look at this soon :) |
|
@magiccodingman before using yours I had already managed to code my own kernel up to 62 t/s from scratch, but I noticed that your one was 77 t/s for me on just one R9700, and it uses dense attention when I was using Quest96, so I was outplayed, mind you there are probably quite a few contributors on this one, but there are definitely quite a few correctness issues here, I get quite a few loops and some bullshits at times. At the moment I'm running a massive conformance test suite that I've ported over from my own effort and it's picked up a few more mismatches so there will be lot more to come. I have 5 more coming I think. |
|
@Terrydaktal lol your story is the same as many of us haha. I was managing my own custom stuff until I ran into DeadCodes project and was like, "okay this guy smoked me". Then I forked his branch, began to follow Brians fork for a lot of his work, fixes, and mxfp4 work. Then Rob and Deadcodes amazing work unifying RDNA4 library. And that's where I started down the Dflash to build out some of my own stuff. Recently yoinked some ideas from tclaiver as well. Like the newest update has fp8 kv-cache calibrations added, though it's not default on yet because I've not fully tested it yet, even though I've added it. I wanted that for fidelity but also to do some testing before I started experimenting with a calibrated TurboQuant KV cache system. I have a new experiment I'm still working on that in theory if it works. This project could potentially run Qwen3.8 Next Flash with 2 R9700's and only 30 GB of allocated RAM and "hopefully" have pretty crazy speeds. But I assume you found this project from the discord community right? That community is awesome. But those conformance test suites. Honestly if you're willing to make a PR sharing them. I'd absolutely love to have it. A big part of this fork wasn't just to yoink the brilliant work of others while also experimenting down more paths + having more flexibility. But you'll see I have over 100+ tests right now that's critical to the system. It's tests I built to capture issues I experienced in production and want to see blow up early, not while it's in prod. But my whole project right now is also in GitLab, and my GitHub is a cloned mirror of gitlab. I had no idea people would actually follow or use my project lol. So, if I'm getting PR's. I may get this project mirrored with my latest on gitlab in the coming days when I finish my experiment. Then just use GitHub so I don't have mirror + PR bs. So don't let the PR hanging for a few days be discouraging. I need to sync up gitlab to github and handle my weird bs. My experimental branch is just held up right now due to this horrendous bug I've been chasing that showed itself with the new qwen4 architecture. |
…eme decay spans Repair the native scan's extreme-decay path, which could clamp valid recurrent contributions to zero. Recompute affected heads with the bounded recurrence and preserve the verified kernel/source binding. Keep this localized numerical repair separate from the later complete M1/M8 and eager/compiled arithmetic alignment. Provenance Upstream PR: magiccodingman/vllm-radiance#8 (author: Terrydaktal). The Coherence extreme-decay guard is retained as a separately qualified adaptation.
…pans Repair the native scan's extreme-decay path, which could clamp valid recurrent contributions to zero. Recompute affected heads with the bounded recurrence and preserve the verified kernel/source binding. Keep this localized numerical repair separate from the later complete M1/M8 and eager/compiled arithmetic alignment. Provenance Upstream PR: magiccodingman/vllm-radiance#8 (author: Terrydaktal). The Coherence extreme-decay guard is retained as a separately qualified adaptation.
This fixes two independent numerical errors: biased replacement-token sampling in the DFlash2 selector path, and silent corruption of recurrent state during GDN prefill with large decay spans. The PR contains only those corrections, build integration, synthetic regression tests and numerical evidence.
DFlash proposal/replacement noise
Backport vLLM #54282, commit
fe755c88995ad468882517b6c4bdd60138d46a3a, to Radiance's pinned vLLM v0.28.0. Upstream already fixes this exact DFlash2 selector by passingIS_DRAFTING=Trueto an updatedgumbel_noised_argmaxAPI; this adaptation applies the same RNG salt directly because the pinned API predates that change. This is a backport, not an additional fix for current upstream vLLM. Its proposal and the target's replacement currently reuse(seed, position), which biases probabilistic rejection sampling. Salt the proposal's local RNG index by1 << 30; model/cache positions and greedy selection are preserved. The new guarded, idempotent patch is included in all three applicable Docker build paths.GDN extreme-decay correction
The pinned scan factorizes bounded decay products around a midpoint, then clamps growing exponents at 80. Spans above 160 can erase valid diagonal contributions and carried state while all values remain finite.
Extend the existing libr4d patch with a same-stream correction kernel. For sequence/head pairs exceeding a conservative span of 128, recompute from the original initial state using the bounded FP32 defining recurrence. Ordinary heads retain the fast scan result. The scan ABI, layouts and decode path are preserved; there is no host readback, synchronization or temporary tensor allocation.
Validation
All inputs are synthetic probabilities or random tensors; no model checkpoint or conversation is required.
benchmarks/;docs/NUMERICAL_CORRECTIONS.mdexplains reproduction and provenance.These are single-R9700 numerical tests, including the two per-rank layouts; they do not constitute distributed TP2 or complete image-build qualification. The extra correction launch and recurrence add prefill work; end-to-end throughput still needs qualification. Persisted context state computed with the old math must be rebuilt.
This PR makes no claim that these defects fully explain model repetition or that the fixes eliminate it. It introduces no loop guard, sampling penalty, history recovery or scheduling change.
Related later qualification
The later compiled M1/M8 and eager/compiled arithmetic repairs are tracked separately in #10/#11 and the public two-fix report. They are not added to this two-issue PR, and their finite replay results do not establish that either defect here was the cause of a particular model loop.