Skip to content

Fix DFlash sampling bias and GDN extreme-decay state corruption - #8

Open
Terrydaktal wants to merge 2 commits into
magiccodingman:mainfrom
Terrydaktal:main
Open

Terrydaktal wants to merge 2 commits into
magiccodingman:mainfrom
Terrydaktal:main

Conversation

@Terrydaktal

@Terrydaktal Terrydaktal commented Sep 14, 2026 •

Copy link
Copy Markdown

This fixes two independent numerical errors: biased replacement-token sampling in the DFlash2 selector path, and silent corruption of recurrent state during GDN prefill with large decay spans. The PR contains only those corrections, build integration, synthetic regression tests and numerical evidence.

DFlash proposal/replacement noise

Backport vLLM #54282, commit fe755c88995ad468882517b6c4bdd60138d46a3a, to Radiance's pinned vLLM v0.28.0. Upstream already fixes this exact DFlash2 selector by passing IS_DRAFTING=True to an updated gumbel_noised_argmax API; this adaptation applies the same RNG salt directly because the pinned API predates that change. This is a backport, not an additional fix for current upstream vLLM. Its proposal and the target's replacement currently reuse (seed, position), which biases probabilistic rejection sampling. Salt the proposal's local RNG index by 1 << 30; model/cache positions and greedy selection are preserved. The new guarded, idempotent patch is included in all three applicable Docker build paths.

GDN extreme-decay correction

The pinned scan factorizes bounded decay products around a midpoint, then clamps growing exponents at 80. Spans above 160 can erase valid diagonal contributions and carried state while all values remain finite.

Extend the existing libr4d patch with a same-stream correction kernel. For sequence/head pairs exceeding a conservative span of 128, recompute from the original initial state using the bounded FP32 defining recurrence. Ordinary heads retain the fast scan result. The scan ABI, layouts and decode path are preserved; there is no host readback, synchronization or temporary tensor allocation.

Validation

All inputs are synthetic probabilities or random tensors; no model checkpoint or conversation is required.

  • The full libr4d patch applies to the pinned source; the original and corrected scan translation units compile using the pinned image/compiler and production flags. Sampler installation and repeat installation pass.
  • Native sampling: 200,000 draws at each of three positions. Maximum absolute probability error is 0.01781–0.01953 before versus 0.000965–0.001180 after. Target-only and greedy controls pass.
  • GDN: the original reproduces 46.4%/82.8% output error and 100% state error on two extreme-decay cases. Both 48/16-head and 24/8-head layouts pass all 12 corrected cases and three graph replays with changed inputs. Maximum corrected output error is 0.345%, state error 0.252%, against an independent FP64 recurrence.
  • Coverage includes unequal sequences, incomplete chunks, mixed affected heads and large value amplitudes. The scripts, source/binary hashes and compact results are included under benchmarks/; docs/NUMERICAL_CORRECTIONS.md explains reproduction and provenance.

These are single-R9700 numerical tests, including the two per-rank layouts; they do not constitute distributed TP2 or complete image-build qualification. The extra correction launch and recurrence add prefill work; end-to-end throughput still needs qualification. Persisted context state computed with the old math must be rebuilt.

This PR makes no claim that these defects fully explain model repetition or that the fixes eliminate it. It introduces no loop guard, sampling penalty, history recovery or scheduling change.

Related later qualification

The later compiled M1/M8 and eager/compiled arithmetic repairs are tracked separately in #10/#11 and the public two-fix report. They are not added to this two-issue PR, and their finite replay results do not establish that either defect here was the cause of a particular model loop.

The DFlash2 selector used the same Philox noise stream as the target's
replacement draw after rejection. This conditions replacements on the rejected
proposal and changes the intended output distribution. Adapt the independent
draft stream from vllm-project/vllm#54282 (fe755c88995ad468882517b6c4bdd60138d46a3a)
to the selector's direct Gumbel call. Offset only its local RNG index, preserve
model/cache positions and greedy behavior, and install the guarded idempotent
patch in the release, development and patch-image build paths.

The pinned libr4d GDN scan's midpoint factorization clamps growing exponents
at 80. Large decay spans can erase valid contributions and final recurrent
state without producing nonfinite values. Extend the existing libr4d build
patch with a bounded FP32 recurrence for affected sequence/head pairs, using
a conservative span threshold of 128. Preserve ordinary fast-scan outputs,
the entry-point ABI, caller stream, tensor layouts and decode path. Propagate
the original launch error before launching the correction. No host readback,
synchronization or additional temporary tensor is introduced.

Add public, checkpoint-free GPU regressions for native selector/rejection
sampling and an independent FP64 GDN recurrence. Include ordinary and extreme
decays, unequal sequences, partial chunks, mixed heads, large values, both
48/16 and 24/8 head layouts, and graph replays with changed inputs. Record
compact numeric qualification evidence and document reproduction, provenance,
the need to rebuild old computed state, and remaining performance limits.
Link the correction documentation from the README.

Validation on R9700/gfx1201 with the pinned Radiance 1.0.16 image:
- Patch applies to pinned libr4d source; both scan units compile with its
  production compiler flags; sampler installation and repeat installation pass.
- At three positions with 200,000 draws per arm, shared-noise probability error
  is 0.01781-0.01953 versus 0.000965-0.001180 after correction. Target-only and
  greedy controls pass.
- Original GDN reproduces 46.4% and 82.8% output error and 100% state error
  on two extreme-decay cases. Both corrected layouts pass all 12 cases and
  three graph replays, with output error below 0.345% and state error below
  0.252% against FP64.

This commit is limited to the two numerical defects and their validation.
End-to-end throughput, distributed TP2 and repetition-rate improvements are
not claimed. The GDN correction adds prefill work that still needs workload
performance qualification. No loop guard, history recovery or serving-policy
change is included.
Correct the numerical-corrections documentation: upstream #54282 already
fixes the same DFlash2 selector through IS_DRAFTING=True. The pinned
vLLM v0.28.0 helper predates that API, so our patch applies the same salt
directly to its local RNG index. Remove wording suggesting that current
upstream still needs a separate selector correction.

Validation: inspected the merged upstream selector and Gumbel-helper
diffs against the pinned API. This changes documentation only; the
qualified arithmetic corrections and recorded results are unchanged.
@magiccodingman

Copy link
Copy Markdown
Owner

Thank you for submitting this PR! I will be taking a look at this soon :)

@Terrydaktal

Terrydaktal commented Sep 15, 2026 •

Copy link
Copy Markdown
Author

@magiccodingman before using yours I had already managed to code my own kernel up to 62 t/s from scratch, but I noticed that your one was 77 t/s for me on just one R9700, and it uses dense attention when I was using Quest96, so I was outplayed, mind you there are probably quite a few contributors on this one, but there are definitely quite a few correctness issues here, I get quite a few loops and some bullshits at times. At the moment I'm running a massive conformance test suite that I've ported over from my own effort and it's picked up a few more mismatches so there will be lot more to come. I have 5 more coming I think.

@magiccodingman

magiccodingman commented Sep 15, 2026 •

Copy link
Copy Markdown
Owner

@Terrydaktal lol your story is the same as many of us haha. I was managing my own custom stuff until I ran into DeadCodes project and was like, "okay this guy smoked me". Then I forked his branch, began to follow Brians fork for a lot of his work, fixes, and mxfp4 work. Then Rob and Deadcodes amazing work unifying RDNA4 library. And that's where I started down the Dflash to build out some of my own stuff. Recently yoinked some ideas from tclaiver as well. Like the newest update has fp8 kv-cache calibrations added, though it's not default on yet because I've not fully tested it yet, even though I've added it. I wanted that for fidelity but also to do some testing before I started experimenting with a calibrated TurboQuant KV cache system.

I have a new experiment I'm still working on that in theory if it works. This project could potentially run Qwen3.8 Next Flash with 2 R9700's and only 30 GB of allocated RAM and "hopefully" have pretty crazy speeds.

But I assume you found this project from the discord community right? That community is awesome.


But those conformance test suites. Honestly if you're willing to make a PR sharing them. I'd absolutely love to have it. A big part of this fork wasn't just to yoink the brilliant work of others while also experimenting down more paths + having more flexibility. But you'll see I have over 100+ tests right now that's critical to the system. It's tests I built to capture issues I experienced in production and want to see blow up early, not while it's in prod.


But my whole project right now is also in GitLab, and my GitHub is a cloned mirror of gitlab. I had no idea people would actually follow or use my project lol. So, if I'm getting PR's. I may get this project mirrored with my latest on gitlab in the coming days when I finish my experiment. Then just use GitHub so I don't have mirror + PR bs.

So don't let the PR hanging for a few days be discouraging. I need to sync up gitlab to github and handle my weird bs.

My experimental branch is just held up right now due to this horrendous bug I've been chasing that showed itself with the new qwen4 architecture.

Terrydaktal added a commit to Terrydaktal/vllm-coherence that referenced this pull request Sep 18, 2026
…eme decay spans

Repair the native scan's extreme-decay path, which could clamp valid recurrent contributions to zero. Recompute affected heads with the bounded recurrence and preserve the verified kernel/source binding.

Keep this localized numerical repair separate from the later complete M1/M8 and eager/compiled arithmetic alignment.

Provenance
Upstream PR: magiccodingman/vllm-radiance#8 (author: Terrydaktal).
The Coherence extreme-decay guard is retained as a separately qualified adaptation.
Terrydaktal added a commit to Terrydaktal/vllm-coherence that referenced this pull request Sep 18, 2026
…pans

Repair the native scan's extreme-decay path, which could clamp valid recurrent contributions to zero. Recompute affected heads with the bounded recurrence and preserve the verified kernel/source binding.

Keep this localized numerical repair separate from the later complete M1/M8 and eager/compiled arithmetic alignment.

Provenance
Upstream PR: magiccodingman/vllm-radiance#8 (author: Terrydaktal).
The Coherence extreme-decay guard is retained as a separately qualified adaptation.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants