Skip to content

[FIX_FOR_VLLM_CUSTOM=488b6da105222f4f8130b3aae4a0e67cfc61f522] Feature-detect MoE shared-experts overlap API - #1773

Merged
iboiko-habana merged 1 commit into
vllm-project:mainfrom
pawel-olejniczak:fix/moe-shared-experts-overlap-version-tolerant
Aug 31, 2026
Merged

iboiko-habana merged 1 commit into
vllm-project:mainfrom
pawel-olejniczak:fix/moe-shared-experts-overlap-version-tolerant

Conversation

@pawel-olejniczak

@pawel-olejniczak pawel-olejniczak commented Aug 31, 2026 •

Copy link
Copy Markdown
Collaborator

This PR consolidates 1 hourly-CI fix against vllm@488b6da105222f4f8130b3aae4a0e67cfc61f522.

Bug 1: Feature-detect MoE shared-experts overlap API

  • State machine id: moerunner_dual_stream_relanded_sync_attr_missing
  • Commit: 326038e

Root cause

vllm#52033 (2c7d7dd) re-landed the dual-stream shared-experts decode work that vllm#52024 had reverted, deleting MoERunner._maybe_sync_shared_experts_stream and re-introducing SharedExperts.maybe_forward_async plus the shared_experts_overlapping kwarg on MoERunner._apply_quant_method. patched_fused_moe_forward in vllm_gaudi/ops/hpu_fused_moe.py still called the deleted method, so every MoE forward raised AttributeError: 'MoERunner' object has no attribute '_maybe_sync_shared_experts_stream' and 19 of 59 jobs went red in CI run 33314657602 (run_unit_tests, run_pd_disaggregate_test and 17 MoE e2e legs). The preceding run 33291778445 was green at the same vllm-gaudi commit, so the break is purely upstream-side. This is the third flip of the same upstream API: vllm#51838 introduced it (adapted in #1723), vllm#52024 reverted it (adapted in #1732), and vllm#52033 has now re-landed it.

Culprit

Regression introduced by PR #52033.

Fix

route the aux-stream launch through _launch_shared_experts_overlap, which mirrors upstream MoERunner._forward_impl: it feature-detects _maybe_sync_shared_experts_stream and SharedExperts.maybe_forward_async, and only forwards shared_experts_overlapping when an async launch actually happened. maybe_forward_async gates on current_platform.is_cuda_alike(), so on Gaudi it returns False, the kwarg is omitted, and the path stays correct on both API shapes - a fourth flip in either direction will not re-break the hourly suite. Verified on a fresh Gaudi3 four-card pod at the pinned vLLM SHA: the two MoE unit tests CI reported red fail at baseline with the same AttributeError and pass with the fix, and the full run_unit_tests suite shows no MoE-related failures.

…e-detect MoE shared-experts overlap API

Root cause: vllm#52033 re-landed the dual-stream shared-experts work that vllm#52024 had reverted, removing MoERunner._maybe_sync_shared_experts_stream that our dp_size==1 MoE fast path calls, so every MoE forward raised AttributeError (19 of 59 hourly jobs red).
Upstream: vllm-project/vllm#52033
Fix: Route the aux-stream launch through _launch_shared_experts_overlap, which feature-detects _maybe_sync_shared_experts_stream vs SharedExperts.maybe_forward_async and only forwards shared_experts_overlapping when an async launch actually happened, so a future flip in either direction no longer breaks HPU.

Signed-off-by: Paweł Olejniczak <pawelx.olejniczak@intel.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Warning

Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.

Adds a version-tolerant launch path for MoE shared-experts overlap to handle upstream vLLM API shape changes without breaking the HPU fused MoE forward.

Changes:

  • Introduces _launch_shared_experts_overlap to feature-detect upstream overlap APIs.
  • Updates patched_fused_moe_forward to use the helper and conditionally forward shared_experts_overlapping.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +726 to 736
# Call it exactly as upstream _forward_impl does. Only pass
# shared_experts_overlapping when an async launch actually happened: that
# is the only case in which the kwarg is guaranteed to exist upstream.
extra_quant_kwargs = {"shared_experts_overlapping": True} if shared_experts_overlapping else {}
shared_output, fused_hidden = self._apply_quant_method(
hidden_states=hidden_states,
router_logits=router_logits,
shared_experts_input=shared_experts_input,
input_ids=input_ids,
**extra_quant_kwargs,
)
Comment on lines +651 to +665
sync_stream = getattr(runner, "_maybe_sync_shared_experts_stream", None)
if sync_stream is not None:
# vllm#52024 shape: runner-side sync, overlap decided internally.
sync_stream(shared_experts_input)
return False

shared_experts = getattr(runner, "_shared_experts", None)
if shared_experts is None or shared_experts_input is None:
return False

# vllm#51838 / vllm#52033 shape: launch on the aux stream, await it later.
forward_async = getattr(shared_experts, "maybe_forward_async", None)
if forward_async is None:
return False
return bool(forward_async(shared_experts_input))
@github-actions

Copy link
Copy Markdown
Contributor

✅ CI Passed

All checks passed successfully against the following vllm commit:
488b6da105222f4f8130b3aae4a0e67cfc61f522

@iboiko-habana
iboiko-habana merged commit 1953e1a into vllm-project:main Aug 31, 2026
3 of 5 checks passed

This branch was successfully deployed

1 active deployment
pre-merge-approval — 326038ed Deployed Aug 31, 2026 by pawel-olejniczak via gate #1506
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants