Repository navigation
Conversation
Signed-off-by: recky-c <ruiqicheng510@gmail.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request addresses a critical data type mismatch in the indexer C8 slot mapping logic when using PCP. By explicitly casting the slot_mapping to int32 before invoking the StoreKvBlockMetadata kernel, the fix ensures that write destinations are calculated correctly regardless of the input data type, preventing potential cache corruption. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [BugFix] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
There was a problem hiding this comment.
Code Review
This pull request ensures that the slot mapping tensor is explicitly cast to int32 before being passed to the native metadata kernel in the Ascend attention backend, as the kernel expects int32 input. The change includes corresponding updates to the unit tests to verify this behavior across different input dtypes. As there were no review comments provided, I have no feedback to offer.
| torch.ops._C_ascend.store_kv_block_metadata( | ||
| slot_mapping, | ||
| # The native metadata kernel reads slot indices as int32. | ||
| slot_mapping.to(torch.int32), |
There was a problem hiding this comment.
Not recommended to make changes in common code paths. Instead, add type conversion at the points where PCP introduces differences.
Signed-off-by: recky-c <ruiqicheng510@gmail.com>
840b73e to
0da6735
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
What this PR does / why we need it?
Fixes two C8 cache-write slot mapping issues exposed by PCP and PD disaggregation:
slot_mapping, while the nativeStoreKvBlockMetadatakernel reads it asint32_t*. Convert only the operator argument to int32, preserving the original metadata.AscendSFAPCPDCPImplso the gathered KV rows use the complete DCP slot mapping. The ordinary DCP path keeps its existing local slice.The tests extend the existing indexer dtype coverage and verify that the combined PCP+DCP implementation keeps the complete SFA slot mapping.
Does this PR introduce any user-facing change?
Fixes incorrect indexer and SFA C8 cache writes that can produce incorrect completions when PCP, DCP, and PD disaggregation are combined. No configuration changes.
How was this patch tested?
git diff --checkpassed.28acc786+ [BugFix][Attention] Build SFA indexer DCP metadata independently #16325 (f56cb5b2) + [BugFix][Mooncake] Fix DCP mapping and support P/D-only DCP and PCP+DCP in PD disaggregation #16492 (037f5143), paired with vLLM84030bbe, GLM-5.2 W4A8C8, MRV2, EP, async scheduling, FULL_DECODE_ONLY, MooncakeConnectorV1, and P TP8 x PCP2 -> D TP8.The hardware results apply to the stated combined checkout and topology.