Repository navigation
[Bugfix] Make GPU sync checks safe under torch.compile - #56904
Conversation
|
Exact isolated validation is running in Buildkite #88941 at literal signed head |
b9077f5 to
9d6c9de
Compare
|
Exact validation update:
The PR remains draft and still requires @khluu to review and defend every line before it is made ready. |
9d6c9de to
80290ca
Compare
Co-authored-by: Sherlock <sherlock@raft.ai> Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
80290ca to
c2fbacb
Compare
|
Replacement exact #88962 caught a genuine flaw in the previous revision before tests collected: I corrected that by registering a weak-referenceable Python wrapper instead. The wrapper calls the original descriptor, preserving the no-recursion property; Dynamo treats the wrapper as opaque and AOT/Inductor can trace/lower the underlying call. I also rebased the draft onto current main and reran all applicable pre-commit hooks plus Current signed head: #88969 is pinned to that literal head and only the H200 Async/Inputs/Utils/Worker and Multimodal Extended Generation 1 lanes. It is currently building the image; this draft remains held until both gates pass. |
|
Exact H200 build #88969 at
Commit |
|
Current replacement exact H200 build #88988 uses literal signed head |
bee6beb to
544bf9d
Compare
Co-authored-by: Sherlock <sherlock@raft.ai> Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
544bf9d to
a6b806a
Compare
|
Exact Buildkite #89000 update:
I changed the decorator to the helper's explicit AI assistance was used; human line review remains required before readiness. |
|
Correction: I copied an invalid full suffix after the displayed short SHA into #89010, so bootstrap correctly rejected The verified local, fork, GitHub pull-ref, and PR head are all exactly |
Co-authored-by: Sherlock <sherlock@raft.ai> Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
|
Exact #89011 closed red: Async/Inputs/Utils/Worker passed, while both new Voxtral parameters failed before engine compilation with Current signed/DCO head The draft remains held for terminal exact evidence and human line review. AI assistance was used. |
|
Terminal exact validation at current head
This confirms explicit AI disclosure: This validation report was prepared with AI assistance. |
|
/ci test |
Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com> Co-authored-by: Sherlock <sherlock@raft.ai> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Summary
gpu_sync_debug's patched tensor-copy wrappers currently read aContextVarbefore checking whether Dynamo is tracing. A full-graph compile therefore fails
with an unsupported
ContextVar.get()instead of bypassing sync checks asintended.
This change:
torch.compiler.is_compiling()before accessing eitherContextVar;Tensor.todescriptor through a weak-referenceablePython callable registered with
torch.compiler.allow_in_graph, so Dynamotreats the call as opaque while AOT/Inductor can still lower it;
torch.compileregression with device and dtype deriveddynamically from another tensor;
processes so neither prior pytest CUDA state nor the other graph mode can
contaminate them.
CI incident and test-selection evidence
Daily main build #88920 first exposed this in
test_voxtral_realtime_cudagraph[FULL_DECODE_ONLY]in the H200 MultimodalModels (Extended Generation 1) job. PR #51167 added the compiled Voxtral path,
but that exact job was blocked in both its PR build #88857 and post-merge main
build #88899, so the failure was not exercised before the daily run.
I searched open vLLM issues and PRs for the exact
ContextVar.get/ Voxtralfull-graph failure and did not find an existing fix.
Validation
pre-commit run --files vllm/utils/gpu_sync_debug.py tests/utils_/test_gpu_sync_debug.py tests/models/multimodal/generation/test_voxtral_realtime.py— passed all applicable hooksgit diff --check— passedsaved C++
Tensor.todescriptor as Dynamo-unsupported in the Voxtralfull-graph test.
is invalid at import time because
allow_in_graphstores a weak reference.FULL_DECODE_ONLYcompile/CUDA-graph surface at literal headc2fbacb1.The following
FULL_AND_PIECEWISEparameter then hit the normal startupguard with only 25.73/32.5 GiB free after the prior in-process compiled
runner retained about 6.5 GiB; later failures were cascades from that state.
isolation helper chose
forkafter the parent pytest process had initializedCUDA. Both Voxtral cases therefore failed before engine compilation with
Cannot re-initialize CUDA in forked subprocess.no test evidence.
pre-compilation CUDA-fork failures. Its literal head still contained the
helper's default mode because the intended one-line
spawnedit had notbeen staged into the prior commit.
literal signed head
f7f21d97caecf810db069402ed02c4624311af32, whichexplicitly passes
method="spawn"to the isolation helper.No model evaluation is required because the production change only alters
debug-wrapper control flow while Dynamo is tracing. The exact affected
generation test is included in targeted CI.
AI assistance disclosure
This change was prepared with AI assistance. Human review is required before
this draft is marked ready or merged.