Skip to content

[FIX_FOR_VLLM_CUSTOM=793ca6998adfa5a3ab22d6012dee78044b8ba901] Track MoERunner shared-experts refactor - #1723

Merged
iboiko-habana merged 1 commit into
vllm-project:mainfrom
pawel-olejniczak:fix/moerunner-shared-experts-stream
Aug 12, 2026
Merged

iboiko-habana merged 1 commit into
vllm-project:mainfrom
pawel-olejniczak:fix/moerunner-shared-experts-stream

Conversation

@pawel-olejniczak

Copy link
Copy Markdown
Collaborator

Root cause

Upstream vLLM PR #51838 refactored the MoERunner shared-experts path: it removed
MoERunner._maybe_sync_shared_experts_stream and the is_internal_router
property. The HPU patched_fused_moe_forward monkeypatch and the qwen3-next
sparse-MoE forward both called those now-missing attributes, raising
AttributeError across the MoE jobs.

Upstream PR

vllm-project/vllm#51838
Reworks MoERunner to launch shared experts asynchronously via
SharedExperts.maybe_forward_async and to thread a fused_output_is_reduced
flag through the reduce/transform steps.

Fix

Launch shared experts via SharedExperts.maybe_forward_async (which returns
False on HPU, so they run synchronously — behaviourally identical to the old
no-op stream sync) and thread the resulting overlap flag plus
_fused_output_is_reduced through _apply_quant_method and the
reduce/transform steps, mirroring upstream MoERunner.forward. Inline the
gate is not None check in the qwen3-next internal-router branch.

…MoERunner shared-experts refactor

Root cause: vllm PR #51838 removed MoERunner._maybe_sync_shared_experts_stream and the is_internal_router property.
Upstream: vllm-project/vllm#51838
Fix: launch shared experts via SharedExperts.maybe_forward_async and thread fused_output_is_reduced in the HPU MoE monkeypatch; inline `gate is not None` for the qwen3-next internal-router check.

Signed-off-by: Paweł Olejniczak <pawelx.olejniczak@intel.com>
@pawel-olejniczak
pawel-olejniczak force-pushed the fix/moerunner-shared-experts-stream branch from 69f6bd0 to 5a65268 Compare August 12, 2026 08:46
@pawel-olejniczak
pawel-olejniczak deployed to pre-merge-approval August 12, 2026 08:46 — with GitHub Actions Active

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates the Gaudi MoE integration to match upstream vLLM’s MoERunner shared-experts refactor (vLLM PR #51838), preventing runtime AttributeErrors caused by removed MoERunner attributes while keeping the fast-path behavior aligned with upstream.

Changes:

  • Replaces the removed MoERunner._maybe_sync_shared_experts_stream usage with SharedExperts.maybe_forward_async and threads the resulting overlap flag into _apply_quant_method.
  • Mirrors upstream’s updated routed/shared reduction flow by threading a fused_output_is_reduced flag through the reduce/transform steps in the dp_size==1 fast path.
  • Replaces the removed MoERunner.is_internal_router check in Qwen3Next sparse-MoE forward with an equivalent gate is not None check.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

File Description
vllm_gaudi/ops/hpu_fused_moe.py Updates the dp_size==1 patched MoE forward to match upstream shared-experts launch and reduction sequencing changes.
vllm_gaudi/models/qwen3_next.py Removes dependency on upstream-removed is_internal_router by inlining the gate-presence check.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@github-actions

Copy link
Copy Markdown
Contributor

✅ CI Passed

All checks passed successfully against the following vllm commit:
793ca6998adfa5a3ab22d6012dee78044b8ba901

@iboiko-habana
iboiko-habana merged commit 587a521 into vllm-project:main Aug 12, 2026
3 checks passed
iboiko-habana pushed a commit that referenced this pull request Aug 31, 2026
…e-detect MoE shared-experts overlap API (#1773)

This PR consolidates 1 hourly-CI fix against
vllm@`488b6da105222f4f8130b3aae4a0e67cfc61f522`.

## Bug 1: Feature-detect MoE shared-experts overlap API

- **State machine id**: moerunner_dual_stream_relanded_sync_attr_missing
- **Commit**: 326038e

### Root cause
vllm#52033 (2c7d7dd64a) re-landed the dual-stream shared-experts decode
work that vllm#52024 had reverted, deleting
MoERunner._maybe_sync_shared_experts_stream and re-introducing
SharedExperts.maybe_forward_async plus the shared_experts_overlapping
kwarg on MoERunner._apply_quant_method. patched_fused_moe_forward in
vllm_gaudi/ops/hpu_fused_moe.py still called the deleted method, so
every MoE forward raised AttributeError: 'MoERunner' object has no
attribute '_maybe_sync_shared_experts_stream' and 19 of 59 jobs went red
in CI run 33314657602 (run_unit_tests, run_pd_disaggregate_test and 17
MoE e2e legs). The preceding run 33291778445 was green at the same
vllm-gaudi commit, so the break is purely upstream-side. This is the
third flip of the same upstream API: vllm#51838 introduced it (adapted
in #1723), vllm#52024 reverted it (adapted in #1732), and vllm#52033 has
now re-landed it.

### Culprit
Regression introduced by [PR
#52033](vllm-project/vllm#52033).

### Fix
route the aux-stream launch through _launch_shared_experts_overlap,
which mirrors upstream MoERunner._forward_impl: it feature-detects
_maybe_sync_shared_experts_stream and SharedExperts.maybe_forward_async,
and only forwards shared_experts_overlapping when an async launch
actually happened. maybe_forward_async gates on
current_platform.is_cuda_alike(), so on Gaudi it returns False, the
kwarg is omitted, and the path stays correct on both API shapes - a
fourth flip in either direction will not re-break the hourly suite.
Verified on a fresh Gaudi3 four-card pod at the pinned vLLM SHA: the two
MoE unit tests CI reported red fail at baseline with the same
AttributeError and pass with the fix, and the full run_unit_tests suite
shows no MoE-related failures.

Signed-off-by: Paweł Olejniczak <pawelx.olejniczak@intel.com>

This branch was successfully deployed

1 active deployment
pre-merge-approval — 5a652683 Deployed Aug 12, 2026 by pawel-olejniczak via gate #1316
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants