[FIX_FOR_VLLM_CUSTOM=793ca6998adfa5a3ab22d6012dee78044b8ba901] Track MoERunner shared-experts refactor - #1723
Merged
iboiko-habana merged 1 commit intoAug 12, 2026
Conversation
pawel-olejniczak
requested review from
PatrykWo,
adobrzyn,
afierka-intel,
iboiko-habana,
jbyczkow,
mgawarkiewicz-intel and
michalkuligowski
as code owners
August 12, 2026 08:43
pawel-olejniczak
had a problem deploying
to
pre-merge-approval
August 12, 2026 08:43 — with
GitHub Actions
Error
…MoERunner shared-experts refactor Root cause: vllm PR #51838 removed MoERunner._maybe_sync_shared_experts_stream and the is_internal_router property. Upstream: vllm-project/vllm#51838 Fix: launch shared experts via SharedExperts.maybe_forward_async and thread fused_output_is_reduced in the HPU MoE monkeypatch; inline `gate is not None` for the qwen3-next internal-router check. Signed-off-by: Paweł Olejniczak <pawelx.olejniczak@intel.com>
pawel-olejniczak
force-pushed
the
fix/moerunner-shared-experts-stream
branch
from
August 12, 2026 08:46
69f6bd0 to
5a65268
Compare
Contributor
There was a problem hiding this comment.
Pull request overview
Updates the Gaudi MoE integration to match upstream vLLM’s MoERunner shared-experts refactor (vLLM PR #51838), preventing runtime AttributeErrors caused by removed MoERunner attributes while keeping the fast-path behavior aligned with upstream.
Changes:
- Replaces the removed
MoERunner._maybe_sync_shared_experts_streamusage withSharedExperts.maybe_forward_asyncand threads the resulting overlap flag into_apply_quant_method. - Mirrors upstream’s updated routed/shared reduction flow by threading a
fused_output_is_reducedflag through the reduce/transform steps in the dp_size==1 fast path. - Replaces the removed
MoERunner.is_internal_routercheck in Qwen3Next sparse-MoE forward with an equivalentgate is not Nonecheck.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
vllm_gaudi/ops/hpu_fused_moe.py |
Updates the dp_size==1 patched MoE forward to match upstream shared-experts launch and reduction sequencing changes. |
vllm_gaudi/models/qwen3_next.py |
Removes dependency on upstream-removed is_internal_router by inlining the gate-presence check. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
iboiko-habana
approved these changes
Aug 12, 2026
Contributor
✅ CI PassedAll checks passed successfully against the following vllm commit: |
iboiko-habana
approved these changes
Aug 12, 2026
iboiko-habana
pushed a commit
that referenced
this pull request
Aug 31, 2026
…e-detect MoE shared-experts overlap API (#1773) This PR consolidates 1 hourly-CI fix against vllm@`488b6da105222f4f8130b3aae4a0e67cfc61f522`. ## Bug 1: Feature-detect MoE shared-experts overlap API - **State machine id**: moerunner_dual_stream_relanded_sync_attr_missing - **Commit**: 326038e ### Root cause vllm#52033 (2c7d7dd64a) re-landed the dual-stream shared-experts decode work that vllm#52024 had reverted, deleting MoERunner._maybe_sync_shared_experts_stream and re-introducing SharedExperts.maybe_forward_async plus the shared_experts_overlapping kwarg on MoERunner._apply_quant_method. patched_fused_moe_forward in vllm_gaudi/ops/hpu_fused_moe.py still called the deleted method, so every MoE forward raised AttributeError: 'MoERunner' object has no attribute '_maybe_sync_shared_experts_stream' and 19 of 59 jobs went red in CI run 33314657602 (run_unit_tests, run_pd_disaggregate_test and 17 MoE e2e legs). The preceding run 33291778445 was green at the same vllm-gaudi commit, so the break is purely upstream-side. This is the third flip of the same upstream API: vllm#51838 introduced it (adapted in #1723), vllm#52024 reverted it (adapted in #1732), and vllm#52033 has now re-landed it. ### Culprit Regression introduced by [PR #52033](vllm-project/vllm#52033). ### Fix route the aux-stream launch through _launch_shared_experts_overlap, which mirrors upstream MoERunner._forward_impl: it feature-detects _maybe_sync_shared_experts_stream and SharedExperts.maybe_forward_async, and only forwards shared_experts_overlapping when an async launch actually happened. maybe_forward_async gates on current_platform.is_cuda_alike(), so on Gaudi it returns False, the kwarg is omitted, and the path stays correct on both API shapes - a fourth flip in either direction will not re-break the hourly suite. Verified on a fresh Gaudi3 four-card pod at the pinned vLLM SHA: the two MoE unit tests CI reported red fail at baseline with the same AttributeError and pass with the fix, and the full run_unit_tests suite shows no MoE-related failures. Signed-off-by: Paweł Olejniczak <pawelx.olejniczak@intel.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Root cause
Upstream vLLM PR #51838 refactored the MoERunner shared-experts path: it removed
MoERunner._maybe_sync_shared_experts_streamand theis_internal_routerproperty. The HPU
patched_fused_moe_forwardmonkeypatch and the qwen3-nextsparse-MoE forward both called those now-missing attributes, raising
AttributeErroracross the MoE jobs.Upstream PR
vllm-project/vllm#51838
Reworks MoERunner to launch shared experts asynchronously via
SharedExperts.maybe_forward_asyncand to thread afused_output_is_reducedflag through the reduce/transform steps.
Fix
Launch shared experts via
SharedExperts.maybe_forward_async(which returnsFalse on HPU, so they run synchronously — behaviourally identical to the old
no-op stream sync) and thread the resulting overlap flag plus
_fused_output_is_reducedthrough_apply_quant_methodand thereduce/transform steps, mirroring upstream
MoERunner.forward. Inline thegate is not Nonecheck in the qwen3-next internal-router branch.