Skip to content

[CI]Main2Main 0807 - #13477

Merged
wangxiyuan merged 6 commits into
vllm-project:mainfrom
LQDLove:main2main-20260804
Aug 11, 2026
Merged

wangxiyuan merged 6 commits into
vllm-project:mainfrom
LQDLove:main2main-20260804

Conversation

@LQDLove

@LQDLove LQDLove commented Aug 4, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Upgrade baseline

  • Update the verified vLLM main anchor from 2e09247c2d7b6b97d13af6e71a85bf8d1271deb6 to 58d3918e3ea0a544ffedadad2ba84559e9c51d8f. The full upstream range is available in this comparison.
  • Preserve the vLLM 0.26.0 compatibility lane while adapting the main lane to the new upstream contracts. Version gates use vllm_version_is("0.26.0") and are limited to real contract differences.
  • The changes are organized in the same order as the changed files in this PR. Each item identifies the upstream change, the downstream adaptation, and why the adaptation is required.

Changes by file

1. .github/vllm-main-verified.commit
Change Upstream change Downstream adaptation Why
Update anchor to 58d3918e Upgrade window 0351e9aa...58d3918e. Set anchor. Source of truth for main2main workflow.
2. vllm_ascend/patch/platform/patch_fused_moe.py
Change Upstream change Downstream adaptation Why
Version-gate FusedMoE → FusedMoEFactory rename vllm #44941 renamed FusedMoE to FusedMoEFactory. On main, capture and patch FusedMoEFactory; on v0.26.0, also patch legacy FusedMoE. Both lanes need the Ascend runner patch at the correct binding.
3. vllm_ascend/models/deepseek_v4.py / vllm_ascend/models/minimax_m3/minimax_m3.py / vllm_ascend/ops/fused_moe/fused_moe.py / vllm_ascend/ops/fused_moe/routed_experts.py
Change Upstream change Downstream adaptation Why
Import and use FusedMoEFactory; remove dead FusedMoE re-export vllm #44941. Replace FusedMoE with FusedMoEFactory. Old symbol no longer exists on main. Remove stale FusedMoE re-export from fused_moe.py and dead reference in routed_experts.py comment.
4. tests/ut/models/test_deepseek_v4_moe.py / tests/ut/models/minimax_m3/test_minimax_m3.py
Change Upstream change Downstream adaptation Why
Update monkeypatch target to FusedMoEFactory vllm #44941. "FusedMoE" → "FusedMoEFactory". Must match the symbol imported by models.
5. vllm_ascend/worker/model_runner_v1.py
Change Upstream change Downstream adaptation Why
Version-gate calculate_kv_scales removal vllm #49389 removed runtime KV-scale calculation. Add vllm_version_is("0.26.0") guard. Ascend MRV1 still supports it on v0.26.0; attribute absent on main.
Version-gate clear_buffer() removal vllm #50721 removed clear_buffer() from RoutedExpertsCapturer. Wrap in vllm_version_is("0.26.0") guard. On main, each routed layer overwrites current step's token rows.
6. vllm_ascend/models/layer/attention/layer.py
Change Upstream change Downstream adaptation Why
Remove dead Q/K/V_SCALE_CONSTANT references vllm #49389 removed env var registrations. Module-level constants still exist. Remove unused q_range/k_range/v_range initializations and dead import envs. Dead-code cleanup; DSAAttention.forward() never used these attributes.
7. tests/ut/patch/platform/test_deepseek_v4_thinking.py
Change Upstream change Downstream adaptation Why
Version-gate reasoning effort expectations vllm #50580 maps low/minimal/medium → low. vllm_version_is("0.26.0") guard. v0.26.0 keeps old mapping; main uses new.
8. tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py
Change Upstream change Downstream adaptation Why
Enable chunked_prefill for hybrid model vllm #50991 enabled prefix cache by default for Mamba/hybrid models. False → True. Hybrid model now requires chunked prefill.
9. vllm_ascend/patch/platform/patch_vision.py (new) + vllm_ascend/patch/platform/__init__.py
Change Upstream change Downstream adaptation Why
Patch FusedInputNorm.forward eps=0.0 → eps=1e-5 vllm #50411 added FusedInputNorm with F.batch_norm(eps=0.0). Monkey-patch forward to use eps=1e-5; guarded with contextlib.suppress(ImportError). Upstream PyTorch 2.13.0 allows eps >= 0 for inference; vllm-ascend PyTorch 2.10.0 requires eps > 0 always. Release wheels lack FusedInputNorm. Remove this patch once bundled PyTorch >= 2.13.0.
10. vllm_ascend/ops/triton/mamba/postprocess.py
Change Upstream change Downstream adaptation Why
Add kernel signature parameters vllm #50432 changed signature. Add CONV_STATE_DIM_FIRST, HAS_IDX_MAPPING, PRECOMPUTED_NEW_COMPUTED, state_dim_row_count/stride, idx_mapping_ptr parameters, and num_loops for DS conv copy. Must match upstream kernel contract.
11. tests/e2e/conftest.py / tests/ut/spec_decode/test_speculators_vwn_eagle3.py
Change Upstream change Downstream adaptation Why
HunyuanVL placeholder version gate; remove unused maybe_calc_kv_scales mock vllm #49691, #49389. Version gate and dead-mock removal. Adapt to upstream contract changes.
12. vllm_ascend/patch/__init__.py
Change Upstream change Downstream adaptation Why
Document FusedMoE → FusedMoEFactory rename and new patch_vision entry — Update patch registry documentation. Keep the patch manifest in sync with reality.

Compatibility and review notes

  • Version gates use vllm_version_is("0.26.0") exclusively; no hasattr fallbacks beyond the explicitly justified clear_buffer guard (where the upstream change is a method removal, not a rename).
  • The FusedMoE → FusedMoEFactory rename is applied consistently across all call sites: deepseek_v4.py, minimax_m3.py, fused_moe.py, routed_experts.py, and patch_fused_moe.py.
  • The layer.py Q/K/V_SCALE_CONSTANT removal is a dead-code cleanup: the DSAAttention class initialized these tensors from envs module-level constants (which still exist), but never used them in forward().

Does this PR introduce any user-facing change?

No. This is a compatibility update; no new Ascend-specific public API is introduced.

How was this patch tested?

CI on the branch. See Buildkite workflow run for detailed results.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request performs a routine synchronization of the local vllm dependency tracking file with the upstream vllm main branch. This ensures that the CI environment is aligned with the latest developments and fixes from the upstream repository.

Highlights

  • Dependency Update: Updated the vllm-main-verified.commit file to track the latest upstream vllm main commit (72cd5424).
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the verified commit hash for the vllm-main dependency in .github/vllm-main-verified.commit to 72cd5424da80a4a9caa3f42fd65bc0b94e61cbf0. There are no review comments, so I have no feedback to provide.

Suggested PR Title:

[CI][Misc] Update vllm-main-verified commit hash

Suggested PR Summary:

### What this PR does / why we need it?
This pull request updates the verified commit hash for the `vllm-main` dependency in `.github/vllm-main-verified.commit` to `72cd5424da80a4a9caa3f42fd65bc0b94e61cbf0`.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
No tests were added as this is a simple dependency commit hash update.

@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@LQDLove LQDLove changed the title [CI] vllm main2main 0804 72cd5424 [CI]Main2Main 0804 Aug 4, 2026
@LQDLove
LQDLove force-pushed the main2main-20260804 branch from 8d04145 to c1d02a4 Compare August 4, 2026 08:41
@LQDLove
LQDLove force-pushed the main2main-20260804 branch 4 times, most recently from fd22101 to 3da8ea9 Compare August 4, 2026 11:55
@LQDLove LQDLove closed this Aug 4, 2026
@LQDLove LQDLove reopened this Aug 4, 2026
@LQDLove
LQDLove force-pushed the main2main-20260804 branch from 3da8ea9 to 7ffba89 Compare August 4, 2026 12:46
@LQDLove
LQDLove force-pushed the main2main-20260804 branch from 7ffba89 to b9fe5e6 Compare August 4, 2026 13:07
@github-actions github-actions Bot removed the ci/build label Aug 4, 2026
@LQDLove
LQDLove force-pushed the main2main-20260804 branch from b9fe5e6 to 0e1c88d Compare August 4, 2026 13:26
@LQDLove LQDLove closed this Aug 4, 2026
@LQDLove LQDLove reopened this Aug 4, 2026
@LQDLove
LQDLove force-pushed the main2main-20260804 branch from 0e1c88d to 5ec06ac Compare August 4, 2026 14:15
@LQDLove
LQDLove force-pushed the main2main-20260804 branch from 5ec06ac to f1fd765 Compare August 4, 2026 14:18
@LQDLove LQDLove closed this Aug 4, 2026
@LQDLove LQDLove reopened this Aug 4, 2026
Comment thread vllm_ascend/ops/fused_moe/fused_moe.py Outdated
# main branch renamed FusedMoE -> FusedMoEFactory; v0.26.0 keeps FusedMoE.
# The FusedMoEFactory import is kept for re-export compatibility on v0.26.0 only.
if vllm_version_is("0.26.0"):
from vllm.model_executor.layers.fused_moe.layer import FusedMoE # noqa: F401

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we don't need it. Please delete it.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK

@LQDLove

LQDLove commented Aug 7, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

Comment thread vllm_ascend/worker/model_runner_v1.py Outdated
if self.routed_experts_initialized:
self.routed_experts_capturer.clear_buffer()
if vllm_version_is("0.26.0"):
self.routed_experts_capturer.clear_buffer()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This change is incorrect, it shoud be:

if vllm_version_is("0.26.0"):
    if self.vllm_config.model_config.enable_return_routed_experts:
        if self.routed_experts_initialized:
            self.routed_experts_capturer.clear_buffer()

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok



def _patched_fused_input_norm_forward(self, grid_thw, visual_dtype):
if self.is_identity:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please add patch descriptions to vllm_ascend/patch/__init__.py

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok

state_conv_widths_ptr, # conv width for conv states (0 for temporal)
state_group_indices_ptr, # maps state_idx to group index in block table
# DS conv row metadata. Zero keeps the single-region copy path.
state_dim_row_count_ptr, # int32: per-block dim row count for DS conv

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please have the model owner confirm that the modifications are correct and there is no performance degradation.

@LQDLove

LQDLove commented Aug 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Rerun:

  • E2E

@github-actions

github-actions Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@LQDLove

LQDLove commented Aug 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun

[Bot]: rerun completed.

Rerun:

  • E2E

@LQDLove

LQDLove commented Aug 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Rerun:

  • E2E

@LQDLove

LQDLove commented Aug 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@@ -1 +1 @@
# Adapt from https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/mamba_utils.py

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

补充单算子用例和MD

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

好的,后面补充一下该用例和MD覆盖到涉及的场景

Update vllm-main-verified.commit to track upstream vllm main commit
8d9b52f7c2514490bdadfd5eb0c931e58625df2e.

Adaptations for upstream main changes between 2e09247c and 8d9b52f7c:

- FusedMoE → FusedMoEFactory rename (vllm#44941): version-gated patch in
  patch_fused_moe.py, updated imports/usage in deepseek_v4.py,
  minimax_m3.py, fused_moe.py, routed_experts.py, and test monkeypatch
- Removed calculate_kv_scales (vllm#49389): version-gated in
  model_runner_v1.py
- Removed Q/K/V_SCALE_CONSTANT env var registrations (vllm#49389):
  remove unused dead-code from attention layer.py. Module-level typed
  constants still exist in envs.py.
- RoutedExpertsCapturer.clear_buffer() removed (vllm#50721): version-gated
  with vllm_version_is("0.26.0") and merged inner conditions
- FusedInputNorm eps=0.0 (vllm#50411): patch FusedInputNorm.forward to
  replace eps=0.0 with eps=1e-5 for Ascend torch_npu compatibility.
  Guarded with contextlib.suppress(ImportError) for release wheels.
- Triton postprocess_mamba_fused_kernel signature changed (vllm#50432):
  add CONV_STATE_DIM_FIRST, HAS_IDX_MAPPING, PRECOMPUTED_NEW_COMPUTED
  parameters to Ascend's custom kernel
- Mamba prefix cache enabled by default (vllm#50991): enable
  chunked_prefill in extract_hidden_states E2E test
- DeepSeek V4 reasoning effort mapping (vllm#50580): version-gated
  thinking test expectations
- HunyuanVL placeholder format changed (vllm#49691): version-gated in
  conftest.py
- Remove unused maybe_calc_kv_scales mock from spec_decode test
- Remove stale FusedMoE re-export from fused_moe.py

Signed-off-by: liaoqidan <1107297340@qq.com>
The FusedMoEFactory adaptation in 93bc205 introduced an unused
_DefaultAscendMoERunner/_DefaultAscendRoutedExperts block whose
type[RoutedExperts] annotation references an undefined name, failing
ruff F821 in pre-commit. The block was never referenced; 310P runner
selection already happens via REGISTERED_ASCEND_OPS in vllm_ascend/utils.py.

Signed-off-by: liaoqidan <1107297340@qq.com>
Clarify that upstream vLLM uses PyTorch 2.13.0 (eps>0 training, eps>=0
inference) while vllm-ascend bundles PyTorch 2.10.0 (eps>0 always), so
upstream passing eps=0.0 works there but fails on Ascend. The patch can
be removed once bundled PyTorch >= 2.13.0.

Signed-off-by: liaoqidan <1107297340@qq.com>
Sync .github/vllm-main-verified.commit to the latest vLLM main HEAD.

Signed-off-by: liaoqidan <1107297340@qq.com>

@LoganJane LoganJane left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for five P1 correctness and compatibility issues in vllm_ascend/ops/triton/mamba/postprocess.py. Details are provided inline. This review is intentionally limited to P1 findings.

if not PRECOMPUTED_NEW_COMPUTED:
num_tokens_running_state = num_computed + num_scheduled - num_draft
else:
num_tokens_running_state = (new_num_computed // block_size) * block_size

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Restore the MRv2 running-state formula.

Under PRECOMPUTED_NEW_COMPUTED, this assigns num_tokens_running_state to the same aligned value as aligned_new_computed. Consequently, needs_copy is always true and accept_token_bias is always zero. For example, with block_size=16, new_num_computed=32, and num_accepted=3, the correct values are num_tokens_running_state=30 and accept_token_bias=2, while this code produces 32 and 0. The upstream contract is:

new_num_computed = tl.load(num_computed_tokens_ptr + req_idx)
num_tokens_running_state = new_num_computed - num_accepted + 1

Please use that formula; otherwise MRv2 align can skip required state shifts or copy at the wrong boundary.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok

num_scheduled = tl.load(num_scheduled_tokens_ptr + req_idx)
num_computed = tl.load(num_computed_tokens_ptr + req_idx)
num_draft = tl.load(num_draft_tokens_ptr + req_idx)
if PRECOMPUTED_NEW_COMPUTED:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Do not load num_scheduled_tokens_ptr on the precomputed path.

run_fused_postprocess_align() passes None for both num_scheduled_tokens_ptr and num_draft_tokens_ptr when PRECOMPUTED_NEW_COMPUTED=True, but line 80 performs tl.load(num_scheduled_tokens_ptr + req_idx) unconditionally before this branch. This can fail during Triton specialization/compilation by doing pointer arithmetic and a load through None; relying on dead-code elimination is unsafe. Move the num_scheduled, num_computed, and num_draft loads entirely into the else branch, matching upstream.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok

state_idx = tl.program_id(1)

if HAS_IDX_MAPPING:
req_idx = tl.load(idx_mapping_ptr + batch_idx).to(tl.int32)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Validate the batch index, not the mapped request slot, against num_reqs.

num_reqs is the number of active batch rows, whereas req_idx is a request-state slot and may be sparse/non-contiguous. For example, num_reqs=2 with idx_mapping=[5, 1] is valid, but the current req_idx >= num_reqs check drops slot 5 entirely. It also fails to reject the -1 skip sentinel and can read decision buffers at a negative offset. The control flow should be:

if batch_idx >= num_reqs:
    return
if HAS_IDX_MAPPING:
    req_idx = tl.load(idx_mapping_ptr + batch_idx)
    if req_idx < 0:
        return
else:
    req_idx = batch_idx

A positive request-slot bound would require the actual buffer capacity, not num_reqs.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok

dim_row_count = tl.load(state_dim_row_count_ptr + state_idx)
dim_row_stride = tl.load(state_dim_row_stride_ptr + state_idx)
num_rows_to_copy = (conv_width - accept_token_bias).to(tl.int64)
copy_size = dim_row_count * dim_row_stride

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Fix the DS convolution row copy dimensions and advance the row address.

For DS layout (state[block, dim, state_len]), the kernel must iterate dim_row_count rows and copy (conv_width - accept_token_bias) * state_elem_size bytes per row, advancing source and destination by dim_row_stride for each row. This code swaps those quantities: it uses conv_width - bias as the loop count and dim_row_count * dim_row_stride as the per-loop copy size. The loop at lines 192-196 then never applies a row offset, so it repeats the same oversized copy.

For dim=4, state_len=3, fp16, and bias=1, this copies 24 bytes from source offset 2, reading 2 bytes past a 24-byte block, and repeats it twice. Please use num_loops = dim_row_count, copy_size = (conv_width - bias) * state_elem_size, and add row * dim_row_stride to both pointers while keeping the pointer-type cast hoisted.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok

CONV_STATE_DIM_FIRST: tl.constexpr,
# HAS_IDX_MAPPING: when True, program_id(0) is a batch index resolved to a
# req-state slot via idx_mapping_ptr (V2). When False, it is the req index.
HAS_IDX_MAPPING: tl.constexpr = False,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Preserve the vLLM v0.26.0 output-buffer contract.

This Ascend kernel is installed for both main and v0.26.0, but their MRv2 callers have different output contracts. Main (after vLLM #50432) passes an accepted-token snapshot as input and a non-null output buffer. v0.26.0 passes None as num_accepted_tokens_out_ptr and expects the HAS_IDX_MAPPING path to update num_accepted_tokens_ptr in place. The unconditional store to num_accepted_tokens_out_ptr at line 176 therefore compiles/writes through None on the v0.26.0 lane.

Please either backport the #50432 host-side snapshot/output-buffer flow to the v0.26.0 compatibility patch, or select the write contract explicitly at host/version specialization time. Using one unconditional output-pointer store for both lanes is not compatible.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok

- PRECOMPUTED_NEW_COMPUTED: restore running-state formula
  (new_num_computed - num_accepted + 1) instead of the aligned value.
- Move num_scheduled/num_computed/num_draft loads into the else branch
  to avoid loading through None pointers on the precomputed path.
- Validate batch_idx against num_reqs (not the mapped req slot) and
  handle the -1 skip sentinel under HAS_IDX_MAPPING.
- Fix DS conv row copy: num_loops = dim_row_count, copy_size =
  (conv_width - bias) * elem_size, advancing src/dst by dim_row_stride.
- Select the num_accepted write target by whether an output buffer is
  provided (main snapshot+out-buffer vs v0.26.0 in-place under V2).

Signed-off-by: liaoqidan <1107297340@qq.com>
@LQDLove

LQDLove commented Aug 10, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

1 similar comment
@LQDLove

LQDLove commented Aug 10, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

…ernel

dim_row_count/dim_row_stride were loaded inside the CONV_STATE_DIM_FIRST
branch but referenced from the copy loop's CONV_STATE_DIM_FIRST and
is_conv_state branch. Triton's control-flow scoping cannot prove the
variable is defined across the narrower condition, so compilation failed
with an undefined dim_row_stride. Load the DS row metadata inside the
copy branch itself (same scope as the loop) and drop the num_loops
intermediate, matching upstream's self-contained DS copy.

Signed-off-by: liaoqidan <1107297340@qq.com>
@LQDLove

LQDLove commented Aug 10, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

2 similar comments
@LQDLove

LQDLove commented Aug 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@LQDLove

LQDLove commented Aug 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants