Repository navigation
Conversation
…scale check) (sgl-project#97) * quant: reject non-finite fp8 K/V scales at load (sibling of the zero-scale check) Co-Authored-By: Rahul Chalamala <22563365+rchalamala@users.noreply.github.com> * kv_cache: raise ValueError for non-finite scales instead of assert Co-Authored-By: Rahul Chalamala <22563365+rchalamala@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Relationship to upstream: sgl-project#40243 (draft, closed without merge 2026-09-25) proposed the same check; main still accepts non-finite scales. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…roject#99) Check numel before the finite check so multi-element scales raise a clear ValueError instead of an ambiguous RuntimeError, and evaluate isfinite on CPU copies. Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Relationship to upstream: sgl-project#40243 (draft, closed without merge 2026-09-25) proposed the same check. Main only rejects multi-element scales after .tolist(), after the comparisons that raise an ambiguous RuntimeError. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
rodamani
marked this pull request as ready for review
September 29, 2026 21:13
rodamani
requested review from
Alisehen,
AniZpZ,
BBuf,
Edwardf0t1,
FlamingoPg,
HaiShaw,
OrangeRedeng,
b8zhong,
ch-wan and
mmangkad
as code owners
September 29, 2026 21:13
Contributor
Author
|
/tag-and-rerun-ci |
3 of 5 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
FP8 KV-cache scale loading in
layers/quantization/kv_cache.pyrejects zero scales but accepts non-finite ones, so a checkpoint with a NaN / Infk_scaleorv_scaleloads and silently corrupts attention. Multi-element (non-per-tensor) scales only fail later with an ambiguousRuntimeErrorfrom a tensor comparison.Modifications
numel() == 1first and raise a clearValueErrorfor non-per-tensor scales.ValueError(evaluated on CPU copies).test/registered/unit/layers/quantization/test_kv_cache_scale_validation.py.#40243 proposed the finite check and was closed without merge.
Accuracy Tests
CPU: 7 passed. With the first commit's source change reverted, 3 of its 6 tests fail.
Speed Tests and Profiling
Load-time validation only.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ✅ Run #36767676542
Latest PR Test (Extra): ❌ Run #36767675843
Latest PR Test (AMD ROCm 10): ❌ Run #36767675960