Skip to content

Release 2.8.3.post2: backport weights_only=True torch.load fix (CVE-2026-31253) - #2974

Merged
Johnsonms merged 5 commits into
Dao-AILab:ko3n1g/release/v2.8.3-cu13from
Johnsonms:release/v2.8.3-post2-cve-31253
Oct 6, 2026
Merged

Johnsonms merged 5 commits into
Dao-AILab:ko3n1g/release/v2.8.3-cu13from
Johnsonms:release/v2.8.3-post2-cve-31253

Conversation

@Johnsonms

@Johnsonms Johnsonms commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Why

No published FA2 version has either of the two security fixes that are on main:

Advisory Fix on main In v2.8.3.post1?
CVE-2026-31253 / GHSA-7g5w-pq96-8c5w — torch.load without weights_only=True #2622 (12f0ce1d, 2026-06-05) no
CVE-2026-62239 / GHSA-w9r7-4gwr-j959 — bare tarfile.extractall in hopper/setup.py #2702 (0816ef1, 2026-07-12) no

v2.8.3.post1 (current PyPI) was tagged from this release branch, which forked from v2.8.3 before both; main has __version__ = "2.8.4" but v2.8.4 was never tagged. Neither advisory has a fixed version, so scanners flag every release (#2972).

The torch.load fix touches files that ship in the wheel (flash_attn/utils/pretrained.py, flash_attn/models/llama.py). hopper/setup.py is not in the sdist/wheel, so that one only affects source builds of FA3, but backporting it lets both advisories be closed with the same version.

What

On top of v2.8.3.post1: cherry-pick -x of #2622 and #2702 (both clean), plus __version__ → 2.8.3.post2. Nothing else is backported.

GitHub shows 5 commits because the branch tip (ae13dfa4) is a different commit for #2618 than the one v2.8.3.post1 was tagged on (035c7231 + a8aa52b1); they are content-identical apart from the version line, so the effective diff is just the three commits above.

Release

After merge: git tag v2.8.3.post2 && git push origin v2.8.3.post2 runs publish.yml the same way as post1. Once released, I'll update both GHSA/PYSEC entries with first_patched_version.

Related: #2972, #2583, #2637

ko3n1g and others added 3 commits June 9, 2026 08:40
* ci: use 1 ninja job for cu13.2

Signed-off-by: oliver könig <okoenig@nvidia.com>

* fix(setup): request cu13 prebuilt wheels for CUDA 13 torch

get_wheel_url() binned every CUDA >= 12 to major '12', so under a CUDA 13
torch it requested cu12 wheels and never matched the published cu13
artifacts, falling back to a multi-hour source build. Add a CUDA 13
branch so the guessed wheel name uses cu13, matching WHEEL_CUDA_VERSION
in _build.yml.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>

---------

Signed-off-by: oliver könig <okoenig@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>
Passing weights_only=False (the pre-2.4 default) to torch.load allows
arbitrary Python object deserialization from the checkpoint file.
A malicious .pt/.pth file can execute arbitrary code on the machine
loading it — a well-known PyTorch deserialization vector (CWE-502).

Four call sites updated:
  training/src/utils/checkpoint.py  load_checkpoint()
  training/src/eval.py               eval checkpoint loader
  flash_attn/utils/pretrained.py    partial(torch.load, ...) loader
  flash_attn/models/llama.py        state_dicts_from_checkpoint()

weights_only=True restricts deserialization to tensors, dicts, lists,
tuples, and other primitive types — no arbitrary Python objects.
Requires PyTorch >= 1.13; FA4's CuTeDSL dependency already requires
a modern PyTorch 2.x build, so no compatibility regression.

Fixes Dao-AILab#2583

(cherry picked from commit 12f0ce1)
@Johnsonms
Johnsonms requested review from drisspg and tridao October 5, 2026 19:03
@katwang

katwang commented Oct 5, 2026

Copy link
Copy Markdown

Is it possible to backport the fix for GHSA-w9r7-4gwr-j959 in this release as well? It looks like that was also fixed in main: 0816ef1

aryanputta and others added 2 commits October 5, 2026 22:35
… symlink escape (Dao-AILab#2702)

* hopper/setup.py: harden tarfile extraction against path traversal and symlink escape

download_and_copy() extracted NVIDIA toolchain archives with a bare
tarfile.extractall() into the predictable ~/.flashattn/nvidia/<name> cache,
allowing arbitrary file write at build time via a pre-planted symlink or a
malicious archive member (issue Dao-AILab#2637).

- Add safe_extractall(): use the PEP 706 data filter when available (3.12,
  backported to 3.10.12/3.11.4), else fall back to per-member path containment
  and link rejection (stream-safe, single pass).
- Refuse extraction into a symlinked cache path, closing the primary
  pre-planted-symlink vector on all Python versions.

Signed-off-by: Aryan Putta <aryansputta@gmail.com>

* hopper/setup.py: allow in-destination links in extractall fallback

Address review on Dao-AILab#2702:

1. The no-data-filter fallback rejected every link member, which
   regressed real builds: the cuda_nvcc archives ship intra-package
   symlinks (e.g. libnvvm.so -> libnvvm.so.4) that the data filter
   permits. Allow links whose resolved target stays inside the extract
   dir instead, matching the data-filter behavior, and keep rejecting
   escaping and absolute-target links.

2. Harden the cache-path check: os.path.islink only inspects the leaf,
   so also require the fully resolved tmp_path to stay under the cache
   root, catching a symlinked parent directory.

Signed-off-by: Aryan <aryansputta@gmail.com>

---------

Signed-off-by: Aryan Putta <aryansputta@gmail.com>
Signed-off-by: Aryan <aryansputta@gmail.com>
(cherry picked from commit 0816ef1)
@Johnsonms
Johnsonms force-pushed the release/v2.8.3-post2-cve-31253 branch from b91e585 to 6333044 Compare October 5, 2026 22:36
@Johnsonms
Johnsonms merged commit fa74eb7 into Dao-AILab:ko3n1g/release/v2.8.3-cu13 Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants