Add PD test for inkling with mxfp8 KV - #35840
Conversation
|
/rerun-test test/registered/disaggregation/test_disaggregation_inkling_mxfp8.py |
|
Results for 🚀 |
|
/rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py |
|
Results for 🚀 |
|
/rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py |
|
Results for 🚀 |
…ce' into add-inkling-pd-mxfp8-test
|
/rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py |
|
Results for 🚀 |
Motivation
Inkling is the widest single-model exercise of the multi-component state transfer we have. One request carries the main KV list plus four state components, and with MXFP8 KV they cover three of the five index-payload builders:
SWAMAMBABLOCK_SCALEBLOCK_SCALE_SWAA model that ships one or two components leaves most of that dispatch untested. Nothing launched a PD pair for this one, so registration, per-component index payloads, and the page-granular slicing behind them had no end-to-end guard on the path where they interact -- which is where both of the last two defects in this area lived.
MXFP8 KV needs SM100+, so the case is registered on a Blackwell runner.
Only the MXFP8 configuration is covered. It is a superset of bf16 at the orchestration level -- same components, same transfer, same hierarchical prefill cache, plus the two scale components -- and bf16 PD is already exercised by the existing disaggregation tests, so a second case would spend another server pair on largely duplicated coverage.
HiCache rides the prefill role only: the decode role forces chunk cache, and its radix opt-in is refused for sliding-window models.
start_prefill/start_decodeare overridden because the fixture pins--tp 1, the same reasontest_disaggregation_dsv4.pyoverrides them.Verification
4xB200, one prefill and one decode role at TP=2 each, mooncake over IB: 667s, gsm8k 0.855, with 97 retract-and-resume cycles along the way.
est_time=800covers that.The device pool is bounded rather than left to the memory fraction. Write-through wants the host pool above the device pool and the default ratio puts it at 2x, so the device pool is what keeps host memory in range for two roles on one node -- and a bounded pool is also what pushes the host tier into use instead of everything staying resident on device.
That bound is what made this test worth writing twice: retraction only happens once the pool fills, and the first run took the decode scheduler down on a path MXFP8 KV had never reached (fixed in #35888, gsm8k 0.12 before, 0.855 after). A pool sized to avoid retraction would have passed and guarded nothing.
CI States
Latest PR Test (Base): ❌ Run #32687861250
Latest PR Test (Extra): ❌ Run #32687861055
Latest PR Test (AMD ROCm 7.2): ❌ Run #32687861374