Skip to content

Add PD test for inkling with mxfp8 KV - #35840

Merged
ispobock merged 13 commits into
mainfrom
add-inkling-pd-mxfp8-test
Aug 24, 2026
Merged

ispobock merged 13 commits into
mainfrom
add-inkling-pd-mxfp8-test

Conversation

@ispobock

@ispobock ispobock commented Aug 21, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Inkling is the widest single-model exercise of the multi-component state transfer we have. One request carries the main KV list plus four state components, and with MXFP8 KV they cover three of the five index-payload builders:

Component Index payload
SWA window, translated into the SWA index space
MAMBA per-request ShortConv state slot
BLOCK_SCALE whole sequence, full KV index space
BLOCK_SCALE_SWA window, SWA index space

A model that ships one or two components leaves most of that dispatch untested. Nothing launched a PD pair for this one, so registration, per-component index payloads, and the page-granular slicing behind them had no end-to-end guard on the path where they interact -- which is where both of the last two defects in this area lived.

MXFP8 KV needs SM100+, so the case is registered on a Blackwell runner.

Only the MXFP8 configuration is covered. It is a superset of bf16 at the orchestration level -- same components, same transfer, same hierarchical prefill cache, plus the two scale components -- and bf16 PD is already exercised by the existing disaggregation tests, so a second case would spend another server pair on largely duplicated coverage.

HiCache rides the prefill role only: the decode role forces chunk cache, and its radix opt-in is refused for sliding-window models.

start_prefill / start_decode are overridden because the fixture pins --tp 1, the same reason test_disaggregation_dsv4.py overrides them.

Verification

4xB200, one prefill and one decode role at TP=2 each, mooncake over IB: 667s, gsm8k 0.855, with 97 retract-and-resume cycles along the way. est_time=800 covers that.

The device pool is bounded rather than left to the memory fraction. Write-through wants the host pool above the device pool and the default ratio puts it at 2x, so the device pool is what keeps host memory in range for two roles on one node -- and a bounded pool is also what pushes the host tier into use instead of everything staying resident on device.

That bound is what made this test worth writing twice: retraction only happens once the pool fills, and the first run took the decode scheduler down on a path MXFP8 KV had never reached (fixed in #35888, gsm8k 0.12 before, 0.855 after). A pool sized to avoid retraction would have passed and guarded nothing.


CI States

Latest PR Test (Base): ❌ Run #32687861250
Latest PR Test (Extra): ❌ Run #32687861055
Latest PR Test (AMD ROCm 7.2): ❌ Run #32687861374

@ispobock

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/disaggregation/test_disaggregation_inkling_mxfp8.py

@github-actions

github-actions Bot commented Aug 21, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/disaggregation/test_disaggregation_inkling_mxfp8.py:

🚀 4-gpu-b200 (1 test): ❌ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_inkling_mxfp8.py

@ispobock ispobock added the run-ci-extra CI: also run the extra suite (requires run-ci) label Aug 21, 2026
@ispobock

Copy link
Copy Markdown
Collaborator Author

/rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py

@github-actions

github-actions Bot commented Aug 21, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py:

🚀 4-gpu-b200 (1 test): ❌ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_inkling_mxfp8.py

@ispobock

Copy link
Copy Markdown
Collaborator Author

/rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py

@github-actions

github-actions Bot commented Aug 22, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py:

🚀 4-gpu-b200 (1 test): ❌ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_inkling_mxfp8.py

@ispobock
ispobock changed the base branch from main to fix-host-pool-retraction-inference August 24, 2026 03:04
@ispobock

Copy link
Copy Markdown
Collaborator Author

/rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py

@github-actions

github-actions Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test registered/disaggregation/test_disaggregation_inkling_mxfp8.py:

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_inkling_mxfp8.py

Base automatically changed from fix-host-pool-retraction-inference to main August 24, 2026 16:11
@ispobock
ispobock merged commit 586211b into main Aug 24, 2026
86 of 94 checks passed
@ispobock
ispobock deleted the add-inkling-pd-mxfp8-test branch August 24, 2026 16:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant