fix(security): harden SafeUnpickler with exact-name allowlist for generic modules - #35818
Merged
Merged
Conversation
…eric modules
The prefix-based allowlist for builtins./operator./pickletools./types./... let
an attacker chain reflective primitives (builtins.getattr + builtins.__import__,
or operator.attrgetter + pickletools.sys) into os.system() even though
("os", "system") was deny-listed, causing unauthenticated RCE via
/load_lora_adapter_from_tensors (CVE-2026-15969).
Replace prefix allowlisting of those generic building-block modules with an
exact-name allowlist of the side-effect-free primitives SGLang's own payloads
need (dict/list/tuple/bytes/...). operator. and pickletools. are fully rejected.
Trusted internal prefixes (torch.*, sglang.srt.*, peft.*,
multiprocessing.reduction.*, ...) keep prefix allowlisting.
Add unit tests covering the original chain, the operator/pickletools bypass
chains, benign payload round-trips, and a torch tensor round-trip.
This endpoint deserializes attacker-controlled pickle payloads but was the only admin-style LoRA endpoint missing @auth_level, so with only --admin-api-key configured (no --api-key) the auth middleware (NORMAL level) allowed unauthenticated calls. Align it with /load_lora_adapter and /unload_lora_adapter (AuthLevel.ADMIN_OPTIONAL).
…d of crashing A malicious serialized_named_tensors payload rejected by SafeUnpickler (e.g. the CVE-2026-15969 chain) raised an unhandled RuntimeError in the scheduler event loop, taking down the whole server (DoS). Wrap the deserialization entry points in try/except so the request returns a failure result and the server keeps serving. Covers both serialized-tensor endpoints: - scheduler.load_lora_adapter_from_tensors - scheduler_components.weight_updater.update_weights_from_tensor
…s on the wire The first-round fix removed builtins./operator./pickletools. prefix trust but kept prefix allowlisting for sglang.srt.*, which still allowed two full RCE bypasses (confirmed against a patched server): - sglang.srt.utils.common.dynamic_import (also re-exported into the weight_updater module) lets a pickle chain import any module and fetch any attribute, then call it -> os.system; - sglang.srt.managers.io_struct._maybe_unwrap_pickle(PickleWrapper(evil)) runs a plain pickle.loads on attacker bytes, defeating SafeUnpickler entirely. This round: - Drop ALL sglang.srt.* prefix allowlisting. Only exact (module, symbol) entries for classes real payloads reference stay: FlattenedTensorBucket, FlattenedTensorMetadata, LocalSerializedTensor. - Deny high-risk code-loading entry points inside the kept torch.* prefix (torch.load, torch.hub.load, torch.utils.cpp_extension.load*, torch.jit.load). - Require admin auth (ADMIN_FORCE) for /load_lora_adapter_from_tensors and /update_weights_from_tensor: rejected unconditionally unless --admin-api-key is configured. - Switch the tensor wire format to safetensors (no code-execution semantics) for sglang's own clients; the server accepts both safetensors and legacy pickle (legacy goes through the hardened SafeUnpickler) for a smooth transition. - Regression tests: dynamic_import chains (both module paths), the io_struct nested-pickle chain, torch.load/hub.load denial, safetensors wire round-trip and legacy-pickle compatibility.
Removed comment about NPU storage from allowlist.
JinyanYi
marked this pull request as ready for review
August 21, 2026 07:21
JinyanYi
requested review from
CatherineSue,
JustinTong0323,
Ying1123,
hnyls2002,
ispobock,
merrymercy,
slin1237 and
xiezhq-hermann
as code owners
August 21, 2026 07:21
sglang-npu-bot
approved these changes
Aug 21, 2026
sglang-npu-bot
merged commit Aug 21, 2026
27e5e97
into
sgl-project:qwen_optimize
182 of 207 checks passed
gjsheu
added a commit
to gjsheu/sglang
that referenced
this pull request
Sep 4, 2026
… for generic modules (sgl-project#35818)" This reverts commit 27e5e97.
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
CVE-2026-15969 is an unauthenticated RCE via
/load_lora_adapter_from_tensors:SafeUnpickler.find_class()used prefix allowlists + a deny-list, so reflective chains (builtins.__import__+getattr,operator.attrgetter+pickletools.sys,sglang.srt.utils.common.dynamic_import,io_struct._maybe_unwrap_pickle(PickleWrapper(evil))) all reachos.system()despite the deny-list. A rejected malicious payload also crashed the scheduler loop (DoS).Related work
#30423 (fixes #30165) targets the same root cause by extending the deny-list. As a blocklist it is inherently incomplete — we confirmed on a patched build that the
dynamic_importandio_structnested-pickle chains still execute. This PR takes the allowlist direction and closes those remaining bypasses.Modifications
common.py: exact-name allowlist for generic modules (dropbuiltins./operator./pickletools.prefix trust); drop allsglang.srt.*prefix trust, keep an exact(module, symbol)allowlist forFlattenedTensorBucket/FlattenedTensorMetadata/LocalSerializedTensor; denytorch.load/hub.load/cpp_extension.load*/jit.load; adddeserialize_tensor_payload(safetensors preferred, hardened-pickle fallback).http_server_engine.py: HTTP client sends base64 safetensors.tp_worker.py: deserialize viadeserialize_tensor_payload; normalize safetensorsdictto(name, tensor)pairs.scheduler.py/weight_updater.py: try/except around deserialization (rejected payloads no longer crash the server).http_server.py:@auth_level(ADMIN_OPTIONAL)on/load_lora_adapter_from_tensors, matching the other admin LoRA endpoints.test_safe_unpickler.py: tests for the original chain, operator/pickletools, dynamic_import, io_struct nested-pickle, torch.load denial, safetensors round-trip, benign round-trips.Accuracy / Speed Tests
Not applicable. Security hardening; no inference-path impact.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ci.CI States
Latest PR Test (Base): ❌ Run #32458264768
Latest PR Test (Extra): ❌ Run #32458264312
Latest PR Test (AMD ROCm 7.2): ❌ Run #32458264346