Repository navigation
Conversation
supports_mamba_cache_extra_buffer returned False for every architecture on XPU, ahead of the _MAMBA_EXTRA_BUFFER_ARCHS check that lists Inkling as an arch the validator must accept. Inkling has no no_buffer path: its model override pins mamba_radix_cache_strategy=extra_buffer and inkling.py asserts enable_mamba_extra_buffer. So on XPU the server died in arg resolution with "extra_buffer is not supported for InklingForConditionalGeneration; use no_buffer." before any worker started, and no flag avoided it. Exempt only the two Inkling architectures from the XPU early-return; every other architecture is still rejected on XPU as before. CUDA and ROCm are unaffected.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
On current
main, Inkling cannot launch on Intel XPU at all. The server dies during server-arg resolution, before any tp worker starts:supports_mamba_cache_extra_bufferreturnsFalsefor every architecture on XPU, and it does so before the_MAMBA_EXTRA_BUFFER_ARCHScheck, whose comment says Inkling is an arch the validator "must accept". Inkling has nono_bufferpath:mamba_radix_cache_strategy=extra_buffermodels/inkling.pyassertsenable_mamba_extra_bufferSo no launch flag avoids the assert.
The early-return came in with #30345 (XPU LoRA), whose description doesn't mention it. Elsewhere XPU is a supported
extra_bufferplatform: #32227 added XPU to the platform check, andtest_mamba_extra_buffer_platform.py(#36410) asserts the platform layer acceptsis_xpu.Modifications
arg_groups/overrides.py: exempt onlyInklingForConditionalGenerationandInklingForConditionalGenerationMTPfrom the XPU early-return. Inkling then takes the existinglinear_attn_backend == "triton"check, which is the default._MAMBA_EXTRA_BUFFER_ARCHS(e.g. KimiK3, NemotronH), is still rejected on XPU exactly as before.is_xpuis true, so CUDA and ROCm are unaffected.A broader alternative would honour
_MAMBA_EXTRA_BUFFER_ARCHSon XPU in general, by moving the XPU check below it. I kept this narrow because Inkling is the only one of those architectures I've run withextra_bufferon XPU. Happy to widen it if reviewers prefer.Accuracy Tests
Two cases added to
test/registered/unit/server_args/test_mamba_extra_buffer_platform.py(CPU CI,base-a-test-cpu), using the file's existingoverride_platform:test_inkling_is_admitted_on_xputest_other_archs_stay_rejected_on_xpuEnd to end: with this change (plus the in-flight #42670/#42677), a 6-layer reduced Inkling checkpoint (
--load-format dummy, bf16, tp=4, Intel Arc Pro B60) launched and served cleanly across 6 benchmark sessions and 2 smoke launches. The server log showsmamba_radix_cache_strategy: extra_buffer,linear_attn_backend: triton, and the Mamba cache allocated on all 4 ranks. That run used dummy weights, so it is a launch-and-serve check, not an accuracy check.Speed Tests and Profiling
N/A: this changes server-arg resolution only, with no kernel or forward-path change.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #37536172990
Latest PR Test (Extra): ❌ Run #37536172695
Latest PR Test (AMD ROCm 10): ❌ Run #37536173357