[Kernel] Enable PDL for per_token_group_quant_8bit_kernel - #46508
Merged
Conversation
mgoin
reviewed
Jun 23, 2026
Comment on lines
122
to
+163
| @@ -153,6 +157,10 @@ __global__ void per_token_group_quant_8bit_kernel( | |||
|
|
|||
| QuantizeGroup<T, DST_DTYPE>(smem_group, group_output, group_size, lane_id, | |||
| threads_per_group, y_s, min_8bit, max_8bit); | |||
|
|
|||
| #if (defined(__CUDA_ARCH__) && (__CUDA_ARCH__ >= 900)) | |||
| asm volatile("griddepcontrol.launch_dependents;"); | |||
| #endif | |||
Member
There was a problem hiding this comment.
I think the wait here needs a memory clobber, otherwise the compiler is free to hoist the input loads above the wait.
I don't think we have a reason for using the raw asm, so I'd recommend using cudaGridDependencySynchronize to replace wait and cudaTriggerProgrammaticLaunchCompletion to replace launch_dependents. We do this in fused_deepseek_v4_qnorm_rope_kv_insert_kernel.cu
zyongye
approved these changes
Jun 23, 2026
qli88
pushed a commit
to qli88/vllm
that referenced
this pull request
Jun 26, 2026
…ct#46508) Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai> Signed-off-by: Qiang Li <qiang.li2@amd.com>
wincent8
pushed a commit
to wincent8/vllm
that referenced
this pull request
Jun 29, 2026
…ct#46508) Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
Dao007forever
pushed a commit
to Dao007forever/vllm
that referenced
this pull request
Jul 18, 2026
…ct#46508) Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
philippesic
pushed a commit
to philippesic/vllm-semantic-cache
that referenced
this pull request
Jul 19, 2026
…ct#46508) Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
plasticchris
pushed a commit
to plasticchris/vllm
that referenced
this pull request
Jul 20, 2026
…ct#46508) Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
follow up on #42996
Test Plan
Test Result
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.