Skip to content

[Fix] State the configured MoE-DP width in the weight-cache fingerprint - #41806

Merged
ch-wan merged 1 commit into
mainfrom
cheng/hot-fix/weight-cache-moe-dp-fingerprint
Sep 30, 2026
Merged

ch-wan merged 1 commit into
mainfrom
cheng/hot-fix/weight-cache-moe-dp-fingerprint

Conversation

@ch-wan

@ch-wan ch-wan commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

This PR is part of a stack (oldest at bottom):

Motivation

The engine fills CacheConfig.moe_dp_size with get_moe_cp_size(), the width of the MoE communicator. The daemon fills it with the configured moe_dp_size. CacheConfig.matches() compares for equality. Whenever attn_cp_size > moe_dp_size, the MoE-DP group aliases the attention-CP group and is wider than configured, so every fetch from the daemon fails with Daemon config mismatch!.

Modifications

  • The engine now states the configured moe_dp_size, as the daemon does. The ranks already agree: both sides read moe_dp_rank from the published placement, which takes the attention-CP rank under the same aliasing.
  • test_weight_cache_daemon.py gains TestWeightCacheDaemonQwen3MoeAttnCP: Qwen3-30B-A3B served through the daemon with --attn-cp-size 2 --enable-prefill-cp --cp-strategy zigzag. The 2-GPU suite's est_time rises to cover it.

Accuracy Tests

H200:

  • New test class on the parent commit: the daemon logs Config mismatch: {'moe_dp_size': (1, 2)} and the client fails to load. On this PR: 3 passed, about two minutes.
  • Qwen3-30B-A3B, daemon and client with --tp 2 --attn-cp-size 2 --enable-prefill-cp --cp-strategy zigzag --cuda-graph-backend-prefill=disabled: weights load over IPC, GSM8K (200 questions) 0.910.
  • test/registered/unit at this PR's head, compared with main: no new failures.

Speed Tests and Profiling

Not applicable.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): 🚫 Run #36772507415
Latest PR Test (Extra): 🚫 Run #36772507303
Latest PR Test (AMD ROCm 10): ❌ Run #36772507201

@ch-wan
ch-wan force-pushed the cheng/hot-fix/tensorcast-world-placement branch from 96e0a09 to f573c0e Compare September 30, 2026 04:24
@ch-wan
ch-wan force-pushed the cheng/hot-fix/weight-cache-moe-dp-fingerprint branch from ea4c1ed to 29921f3 Compare September 30, 2026 04:24
@ch-wan
ch-wan force-pushed the cheng/hot-fix/tensorcast-world-placement branch from f573c0e to eed1566 Compare September 30, 2026 20:24
Base automatically changed from cheng/hot-fix/tensorcast-world-placement to main September 30, 2026 20:25
The engine filled `CacheConfig.moe_dp_size` with `get_moe_cp_size()`, the
width of the MoE communicator, while the daemon fills it with the
configured `moe_dp_size`. `CacheConfig.matches()` compares for equality,
so whenever `attn_cp_size > moe_dp_size` -- where the MoE-DP group aliases
the attention-CP group and is wider than configured -- every fetch failed
with a config mismatch.

Both sides now state the configured width. The ranks already agreed: both
read `moe_dp_rank` from the published placement, which takes the
attention-CP rank under the same aliasing.

The weight-cache daemon e2e test gains an attention-CP case.
@ch-wan
ch-wan force-pushed the cheng/hot-fix/weight-cache-moe-dp-fingerprint branch from 29921f3 to 49b65a8 Compare September 30, 2026 20:25
@ch-wan
ch-wan merged commit 3f20738 into main Sep 30, 2026
@ch-wan
ch-wan deleted the cheng/hot-fix/weight-cache-moe-dp-fingerprint branch September 30, 2026 20:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant