Skip to content

[XPU] Let Inkling use the Mamba extra_buffer strategy on XPU - #42814

Draft
jmunetong wants to merge 1 commit into
sgl-project:mainfrom
jmunetong:xpu/inkling-extra-buffer
Draft

jmunetong wants to merge 1 commit into
sgl-project:mainfrom
jmunetong:xpu/inkling-extra-buffer

Conversation

@jmunetong

@jmunetong jmunetong commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

On current main, Inkling cannot launch on Intel XPU at all. The server dies during server-arg resolution, before any tp worker starts:

File ".../srt/arg_groups/mamba_hook.py", line 103, in validate_mamba_extra_buffer
    assert supports_mamba_cache_extra_buffer(view, hf_config), (
AssertionError: extra_buffer is not supported for InklingForConditionalGeneration; use no_buffer.

supports_mamba_cache_extra_buffer returns False for every architecture on XPU, and it does so before the _MAMBA_EXTRA_BUFFER_ARCHS check, whose comment says Inkling is an arch the validator "must accept". Inkling has no no_buffer path:

  • its model override pins mamba_radix_cache_strategy=extra_buffer
  • models/inkling.py asserts enable_mamba_extra_buffer

So no launch flag avoids the assert.

The early-return came in with #30345 (XPU LoRA), whose description doesn't mention it. Elsewhere XPU is a supported extra_buffer platform: #32227 added XPU to the platform check, and test_mamba_extra_buffer_platform.py (#36410) asserts the platform layer accepts is_xpu.

Modifications

arg_groups/overrides.py: exempt only InklingForConditionalGeneration and InklingForConditionalGenerationMTP from the XPU early-return. Inkling then takes the existing linear_attn_backend == "triton" check, which is the default.

  • Every other architecture, including the rest of _MAMBA_EXTRA_BUFFER_ARCHS (e.g. KimiK3, NemotronH), is still rejected on XPU exactly as before.
  • The line only matters when is_xpu is true, so CUDA and ROCm are unaffected.

A broader alternative would honour _MAMBA_EXTRA_BUFFER_ARCHS on XPU in general, by moving the XPU check below it. I kept this narrow because Inkling is the only one of those architectures I've run with extra_buffer on XPU. Happy to widen it if reviewers prefer.

Accuracy Tests

Two cases added to test/registered/unit/server_args/test_mamba_extra_buffer_platform.py (CPU CI, base-a-test-cpu), using the file's existing override_platform:

case guards result without this fix result with it
test_inkling_is_admitted_on_xpu the launch failure above, both Inkling archs fails (2 subtests) passes
test_other_archs_stay_rejected_on_xpu the XPU check not degrading to "allow everything": KimiK3 and NemotronH accepted off XPU, rejected on XPU passes passes; fails if the XPU check is removed entirely

End to end: with this change (plus the in-flight #42670/#42677), a 6-layer reduced Inkling checkpoint (--load-format dummy, bf16, tp=4, Intel Arc Pro B60) launched and served cleanly across 6 benchmark sessions and 2 smoke launches. The server log shows mamba_radix_cache_strategy: extra_buffer, linear_attn_backend: triton, and the Mamba cache allocated on all 4 ranks. That run used dummy weights, so it is a launch-and-serve check, not an accuracy check.

Speed Tests and Profiling

N/A: this changes server-arg resolution only, with no kernel or forward-path change.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #37536172990
Latest PR Test (Extra): ❌ Run #37536172695
Latest PR Test (AMD ROCm 10): ❌ Run #37536173357

supports_mamba_cache_extra_buffer returned False for every architecture on
XPU, ahead of the _MAMBA_EXTRA_BUFFER_ARCHS check that lists Inkling as an
arch the validator must accept. Inkling has no no_buffer path: its model
override pins mamba_radix_cache_strategy=extra_buffer and inkling.py asserts
enable_mamba_extra_buffer. So on XPU the server died in arg resolution with
"extra_buffer is not supported for InklingForConditionalGeneration; use
no_buffer." before any worker started, and no flag avoided it.

Exempt only the two Inkling architectures from the XPU early-return; every
other architecture is still rejected on XPU as before. CUDA and ROCm are
unaffected.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant