Skip to content

[Fix] Validate FP8 KV scales before updating layers - #25

Open
rchalamala wants to merge 1 commit into
integration/kimi-k3-v0.5.20from
fix/fp8-kv-scale-validation
Open

rchalamala wants to merge 1 commit into
integration/kimi-k3-v0.5.20from
fix/fp8-kv-scale-validation

Conversation

@rchalamala

@rchalamala rchalamala commented Sep 25, 2026 •

Copy link
Copy Markdown

Motivation

Invalid FP8 KV checkpoint scales can reach tensor comparisons or be stored as non-finite values. FNUZ conversion can also double an otherwise finite scale beyond float32 range.

Modifications

Backport the scale-validation behavior from sgl-project/sglang#40243: reject unsupported scale shapes and non-finite values, then verify converted scales fit the destination dtype before updating the layer. Preserve existing default-scale and single-scale handling.

Accuracy Tests

CPU validation:

  • Direct scale-validation runner: 10 tests passed.
  • Scale-validation and neighboring FP4 KV suite: 28 passed, 1 skipped, 9 subtests passed.
  • With only the source fix reverted, the regression suite fails on non-finite values, unsupported shapes, and FNUZ conversion overflow.
  • Registered-test checks, Ruff 0.15.1 lint/format, isort 7.0.0, codespell 2.4.1, and applicable pre-commit checks passed.
PYTHONPATH=python python test/registered/unit/layers/quantization/test_kv_cache_scale_validation.py
PYTHONPATH=python python -m pytest -q \
  test/registered/unit/layers/quantization/test_kv_cache_scale_validation.py \
  test/registered/unit/layers/quantization/test_fp4_kv_cache_quant_method.py

The neighboring compiled CPU tests were run with the system C++ compiler. The default compiler produced a glibc compatibility error on both base and head. No GPU inference tests were run; the change is checkpoint-load validation.

Speed Tests and Profiling

Not applicable: no inference kernel or forward-path changes.

Checklist

  • Code formatting and applicable lint checks passed.
  • CPU unit tests registered and directly runnable.
  • Existing scale defaults and single-scale behavior covered.
  • SGLang code style guidance followed.

CI States

Latest PR Test (Base): ❌ Run #36094995231
Latest PR Test (Extra): ❌ Run #36094995129
Latest PR Test (AMD ROCm 10): ❌ Run #36094995278

Reject non-scalar and non-finite checkpoint scales with explicit errors.
Validate converted scales before storing them so FNUZ scale doubling cannot
write infinities into the layer. Preserve existing default and single-scale
handling.

Backport the final scale-validation behavior from
sgl-project#40243.
Extend regression coverage to both scale tensors, empty tensors, unchanged
parameters after rejection, and existing default-scale behavior.

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@rchalamala

Copy link
Copy Markdown
Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-25T06:40:22.813022Z 542f75a Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. More of your lovely PRs please.

Reviewed commit: 542f75a9f3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@rchalamala
rchalamala marked this pull request as ready for review September 25, 2026 06:37

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no bugs or issues to report.

Devin Review

@rchalamala rchalamala changed the title [Quantization] Validate FP8 KV scales before updating layers [Fix] Validate FP8 KV scales before updating layers Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant