Skip to content

[Intel][XPU]Enable chunked prefill scnearios for XPU with UT - #33804

Merged
mingfeima merged 2 commits into
sgl-project:mainfrom
AnuSajikumar6264:chunked_prefill
Sep 29, 2026
Merged

mingfeima merged 2 commits into
sgl-project:mainfrom
AnuSajikumar6264:chunked_prefill

Conversation

@AnuSajikumar6264

@AnuSajikumar6264 AnuSajikumar6264 commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

The chunked-prefill scripted-runtime suite under test/manual/chunked_prefill/ could not run at
all on Intel XPU: the harness killed every engine during startup. Three argument-resolution paths
also ended in an unhelpful failure rather than a clear rejection. This makes the suite runnable on
--device xpu and reports honestly where a path cannot work there.

Rebased on f1a512c51c. 23 files, +389/-125.

Modifications

Engine startup and runtime

  • test/scripted_runtime/http_server.py — kv-canary JIT-compiles its plan/verify/write kernels
    from .cuh against the arch reported by torch.cuda, so install_canary raised inside
    ModelRunner.alloc_memory_pool; the server died before the dispatch-loop handshake and the test
    side waited out LISTENER_ACCEPT_TIMEOUT_S. The harness now launches with the canary off wherever
    its kernels cannot build, and passes nothing rather than restating the ServerArgs defaults.
  • managers/scheduler.py — continue_generation read torch.cuda.memory_reserved(), which
    returns 0 when CUDA was never initialized, so the reclaim log asserted a call that never happened.
    It now reads the active device module, and when the device exposes no allocator hook it says so
    instead of reporting a 0 MB reclaim.
  • utils/common.py — new device_memory_reserved, next to the existing empty_device_cache,
    returning 0 where the device module has no memory_reserved (torch.cpu and torch.mps do not).
  • layers/attention/{aiter,wave}_backend.py — both read value buffer 0, but the pool indexes as
    layer_id - start_layer, so a non-first PP rank read a negative index. Both now index from
    start_layer, reusing the wording of the same fix already on main in triton_backend.py.
    wave_backend uses the self.token_to_kv_pool reference the constructor already captured;
    aiter_backend keeps a local because its own attribute is not assigned until later.

Argument resolution

  • arg_groups/kv_cache_hook.py — a host KV pool resolves its page movers from
    sgl_kernel.kvcacheio, imported only under if _is_cuda or _is_hip with no else branch, then
    called unconditionally (pool_host/mha.py:278,331,475,863). So on XPU
    --enable-hierarchical-cache, --disaggregation-decode-enable-offload-kvcache and an explicit
    --disaggregation-decode-retraction-backup=host_pool each died minutes into startup on a bare
    NameError. All three are now rejected in handle_cache_compatibility, where the other host-pool
    combinations are already validated and which runs unconditionally.
  • mem_cache/kv_cache_builder.py — a decode server that leaves the retraction backup unset
    degrades to cpu_tensor on XPU instead of being rejected for a pool it never asked for, matching
    what HIP already gets. Keyed on the resolved device, not a hardware probe.
  • arg_groups/overrides.py — deterministic inference fell through to fa3, which asserts
    SM 8x/9x in its own factory, so --device xpu --enable-deterministic-inference could not start.
    XPU now resolves to triton, which unlike intel_xpu is in
    RADIX_SUPPORTED_DETERMINISTIC_ATTENTION_BACKEND and so does not silently disable the radix cache.
    Two lines, inserted after the SM branches; no other device's fallback changes.
  • layers/quantization/mxfp4.py — mxfp4 is registered on cpu and xpu, where the weight-load
    tail called torch.cuda.empty_cache() and freed nothing.

Test-side fixes the above exposed

  • scripted_runtime/background_http_poster.py, context/http_post.py — a rejected control POST
    surfaced only as an unrelated socket-arrival timeout. submit_coro now returns its future and the
    timeout handler interrogates that one future, so the server's own complaint is reported. This
    removes the process-global failure list, whose entries could be attributed to a later request, and
    restores _log_coro_exception as a staticmethod — which also unbreaks four pre-existing CI tests
    in test/registered/unit/scripted_runtime/test_background_http_poster.py.
  • test_scripted_abort.py — a repeat abort is dropped at the TokenizerManager once
    abort_sent is set, so awaiting a second AbortReq could only time out; and reusing a rid needs
    the TokenizerManager to have released it, which no number of scheduler steps in the same yield
    can achieve.
  • scripted_runtime_chunked_helpers.py — _drain_until_released had drifted into four copies,
    three byte-identical; it is now one drain_until_released helper that also waits on lock_refs,
    which drops an iteration after the KV pages. Call sites in test_scripted_abort.py,
    test_scripted_lifecycle.py, test_scripted_regression.py and test_scripted_pp.py are rewired to
    it (test_scripted_pp.py passes max_steps=16 to keep its longer drain). The rid settle constant
    moved up with the other module constants and states where its value came from and how a slower
    device fails; test_scripted_multi_req.py loses a comment that duplicated it.
  • test_scripted_regression.py — the abort-release regression guards read through
    find_req_by_rid, which stops seeing a req whether it released or leaked, so they passed on a
    leaked row. They now compare pool free counts, with the per-req probes kept only on the branch
    where they can still fail.
  • test_scripted_special_case.py — the flashinfer and HiCache classes skip on devices that
    cannot run them, instead of each burning a 300 s listener timeout.

Accuracy Tests

New CI unit tests, on CPU runners:

python -m pytest -q \
  test/registered/unit/server_args/ \
  test/registered/unit/scripted_runtime/ \
  test/registered/unit/utils/test_device_module_probes.py
# 331 passed, 99 subtests passed

Added as test/registered/unit/utils/test_device_module_probes.py (new file), plus cases appended to
test/registered/unit/scripted_runtime/test_http_server.py and
test/registered/unit/server_args/test_server_args.py. They cover the four runtime behaviours this
PR adds: the device-module probes
(device_memory_reserved / empty_device_cache with and without the hooks), the canary knobs per
device string, the host-pool rejection per opt-in path, and the deterministic attention fallback with
the platform pinned so the result holds on any runner.

The manual suite, on one Intel Arc Pro B60 (23.9 GiB):

source /opt/intel/oneapi/2026.1/oneapi-vars.sh --force
export PYTHONPATH=$PWD/python
ZE_AFFINITY_MASK=7 CUDA_VISIBLE_DEVICES=7 MASTER_PORT=29700 SGLANG_TEST_PORT_BASE=29720 \
  python -m pytest -v -ra --timeout=1200 test/manual/chunked_prefill/

302 collected. Latest full run on this branch: 301 passed, 11 skipped, 19 errors, 0 failures
(baseline before this PR on the same hardware: 252 passed, 23 skipped, 27 errors). No test-body
assertion failures in either run.

End-to-end output accuracy

The scripted tests verify bookkeeping (chunk counts, KV pages, lock refs, pool accounting); none of
them checks that the emitted tokens are correct, so a chunked KV write landing at the wrong position
would pass all 301. GSM8K covers that gap. The number that matters is chunked vs unchunked on the
same model and seed, not the absolute score.

Qwen/Qwen3-0.6B, --device xpu --attention-backend triton, 200 examples, 32 threads:

--chunked-prefill-size GSM8K score eval latency (s) output tok/s
256 0.575 296.76 551.44
-1 (off) 0.605 320.56 536.68

Delta -0.030 against a standard error of 0.049 on the difference (0.61 sigma), i.e. indistinguishable
at this sample size; chunking is also 7 % faster end to end, consistent with the sweep above.

Two honest caveats on this figure. Only 11 of 200 requests actually spanned a chunk boundary, because
most GSM8K prompts are shorter than 256 tokens — so it is a weak probe of the chunked path. And
n=200 cannot resolve a difference smaller than about 0.10, so it rules out a gross corruption, not a
subtle one. The stronger measurement is the suite's own test_mixed_prefix_gsm8k_chunked, which uses
24-shot mixed-prefix prompts specifically so every request spans several chunks; it is currently
blocked on XPU by the Llama-3.2-1B-Instruct failure described below, not by anything in this PR.

Suite errors

The 19 remaining errors are environmental or out of scope, not chunked-prefill defects:

Cause Count
openai/gpt-oss-20b not cached (mxfp4 is registered on XPU upstream, so the SWA classes attempt the load) 10
Multi-card classes on a single-card run (device index is out of range) 5
meta-llama/Llama-3.2-1B-Instruct fails on XPU on current main: RuntimeError: mat1 and mat2 shapes cannot be multiplied (1x14336 and 2048x2048), raised from the plain model forward (eager_runner._execute_extend), with no LoRA in the stack. This PR touches no model or LoRA file, and it reproduces with these commits applied to the parent of #30345, so it is neither this PR's doing nor that commit's. It is the sole cause of the four lora_overlap errors and of every e2e GSM8K error 4

Speed Tests

This PR alters no kernel and no scheduling path, so it cannot move a number by construction. What is
worth showing for an enablement PR is that chunked prefill behaves sensibly on XPU once the suite can
run, so the sweep below is a characterisation of the feature on this hardware rather than a
before/after for this diff.

One Intel Arc Pro B60 (23.9 GiB), Qwen/Qwen3-0.6B, --device xpu --attention-backend triton,
--mem-fraction-static 0.7. 65 536 total input tokens in every row.

python -m sglang.benchmark.offline_throughput \
  --model-path Qwen/Qwen3-0.6B --device xpu --attention-backend triton \
  --dataset-name random --random-input 4096 --random-output 8 --random-range-ratio 1.0 \
  --num-prompts 16 --chunked-prefill-size <N> --mem-fraction-static 0.7

--random-input 4096, 16 prompts:

--chunked-prefill-size duration (s) input tok/s vs chunking off
512 18.90 3467.58 +9.3 %
1024 18.66 3512.29 +10.7 %
2048 19.20 3412.92 +7.6 %
4096 (= input, one chunk) 20.80 3151.15 -0.6 %
-1 (off) 20.66 3171.58 baseline

--random-input 8192, 8 prompts:

--chunked-prefill-size duration (s) input tok/s vs chunking off
1024 33.14 1977.25 +16.7 %
2048 33.94 1931.09 +13.9 %
8192 (= input, one chunk) 38.84 1687.22 -0.4 %
-1 (off) 38.67 1694.78 baseline

Chunked prefill is worth 10-17 % input throughput on XPU at these shapes, best around 1024, and the
benefit grows with prompt length. Setting the chunk size equal to the input length reproduces the
unchunked number to within 0.6 %, which is the expected degenerate case and a useful sanity check that
the sweep is measuring what it claims.

Note for anyone reproducing this: sglang.benchmark.one_batch is not usable here. It constructs a
ModelRunner directly and never enters the scheduler, so every --chunked-prefill-size yields an
identical number. offline_throughput drives a real Engine, and the scheduler log confirms chunking
engages (Prefill batch, #new-token: 512 ... #pending-token: 3583).

Notes

  • test/manual/ is excluded from CI by test/README.md, so the runtime behaviours added here are
    covered by the test/registered/unit/ tests above rather than by the manual suite.
  • A median/MAD latency-outlier filter for ChunkSizePredictor.fit was considered and dropped.
    Replayed against the only real 127-sample XPU profile in the tree, it moved the quadratic
    coefficient from -1.24e-05 to -3.65e-05 — further from the a > 0 acceptance test — by discarding
    the largest and most curvature-informative sample. On that sweep corr(seq_len, latency) is
    -0.108, so the quadratic term is unidentifiable and the problem is the sweep range, not outliers.
    TestPPDynamic therefore still cannot pass on XPU; a wider sweep or a linear fallback is the real
    fix and is not attempted here.

CI States

Latest PR Test (Base): ✅ Run #36384585722
Latest PR Test (Extra): ❌ Run #36384585473
Latest PR Test (AMD ROCm 10): ❌ Run #36384585714

@siju-samuel

Copy link
Copy Markdown
Contributor

/tag-run-ci-label

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Aug 13, 2026
@github-actions github-actions Bot added lora hicache Hierarchical Caching for SGLang labels Sep 7, 2026
@AnuSajikumar6264
AnuSajikumar6264 marked this pull request as ready for review September 8, 2026 03:14
@AnuSajikumar6264

Copy link
Copy Markdown
Contributor Author

Thanks — all twelve are addressed. Rebased onto f1a512c51c (was 1357 commits behind at the
first pass, then 42, then 27), which by itself dropped the hunks you flagged as already landed.
Per-comment below; #10 has a longer reply in its own thread.


1. arg_groups/hicache_hook.py — wrong function, incomplete coverage. Agreed on both. Moved out
of the layout/IO resolver, but into handle_cache_compatibility (kv_cache_hook.py) rather than
validate_hicache_host_memory_mode: it already groups these three host-pool flags, already raises,
and runs unconditionally, so the veto can't be skipped by an unrelated gate. Broadened to all three
opt-ins including an explicit --disaggregation-decode-retraction-backup=host_pool, which your
proposal would still have let through since the selector only runs when the backup is None. The
auto-resolved case degrades to cpu_tensor in kv_cache_builder.py instead, keyed on
get_device().device rather than is_xpu() — the probe is hardware presence, so it would silently
degrade a --device cpu config on any XPU-capable host. hicache_hook.py now drops out of the diff
entirely. Left the try/except-stub idea for a follow-up; it is the better long-term shape but a much
larger change.

2. scheduler.py — false-success reporting one device over. Correct, and thank you: I had traded
an XPU bug for a CPU one. Now branches on empty_device_cache's return, exactly as you suggested,
and device_memory_reserved was added because torch.cpu has no memory_reserved either.

3. background_http_poster.py — return the future. Implemented as described; _failures,
_failures_lock, _record_failure and take_failures() are gone and _log_coro_exception is a
staticmethod again. Settle timeout is 1.0 s: loopback POST into the same process, so anything
unsettled a second past an already-elapsed 60 s deadline is stuck rather than racing.
This turned up something neither of us had flagged: the PR was breaking four pre-existing CI tests in
test/registered/unit/scripted_runtime/test_background_http_poster.py (9 pass on main, 4 failed on
the branch) because the file calls _log_coro_exception unbound and post() had started reading
resp.status. That directory is 39 passing now.
One correction to the diagnosis: neither failing abort test was caused by this. start_req posts
stream=True, so a duplicate-rid ValueError comes back inside a 200 SSE body and the future
carries no exception — resp.status >= 400 can never fire for /generate. The real causes were
abort_sent short-circuiting a repeat abort, and same-step rid reuse being impossible because the
release is cross-process. Both fixed, both passing.

4. http_server.py — return {} and the policy split. Took the return {}. On the policy: rather
than add a production rejection, I made the two harness paths share one predicate —
CANARY_SUPPORTED_DEVICES is now public and chunked_prefill_test_utils.py gates on it too. That
mattered more than it looked: use_kv_canary defaults True, so every e2e GSM8K server was
launching with --kv-canary raise and dying on XPU. Rejecting kv_canary != "none" in
validation_hook.py is the cleaner end state but changes production launch behaviour, which I did not
want to smuggle into this PR. Skipped (c) — a log line is enough now that a test asserts the kwargs.

5. test_case.py — one call site, shared base class. Agreed; moved to
@unittest.skipUnless(is_flashinfer_available(), ...) on the single class that sets that backend.
test_case.py is back to its upstream state and out of the diff. Commit messages rewritten — the
quantization claim was stale, and I removed the mxfp4 guard entirely since upstream now registers
mxfp4 for XPU, so it was dead code.

6. RID_RELEASE_SETTLE_STEPS — magic number, wrong placement. Moved up with the other module
constants, and it now records where 40 came from and that a slower device fails as a duplicate-rid
rejection rather than a timeout. I did not take the retry option: the rejection is invisible to the
harness for the reason in #3 (200 + SSE body), so a retry loop would have nothing to key on.

7. _drain_until_released — four copies. Hoisted to drain_until_released in
scripted_runtime_chunked_helpers.py with DRAIN_RELEASE_STEPS = 12; abort, lifecycle,
regression and pp all call it, pp passing max_steps=16. Comment reworded to state the
consequence rather than the mechanism.

8. Vacuous asserts, and >=. Both correct. The per-req probes are now inside
if req is not None:, and the pool comparisons use == — you were right that >= could only fire on
the impossible over-free side. A verifier reproduced the leak you described by no-op'ing
release_kv_cache, which is what convinced me the pool deltas had to carry the check.

9. wave_backend.py — shadowed reference, weaker comment. Both fixed. Uses the
self.token_to_kv_pool the constructor already captured, and both hunks now carry triton_backend.py's
wording verbatim so the three copies read alike. Kept the local in aiter_backend.py, since as you
noted its own attribute is not assigned until later.

10. Description drift. See the dedicated reply in that thread.

11. overrides.py — drifted mirror, and why triton. Took the substance, not the shared-resolver
call: routing through get_default_attn_backend breaks the arg-resolution unit tests on a non-CUDA
host (measured — it reddens test_cuda_host_keeps_its_fa3_default and the CPU case), so this stays a
two-line elif view.device == "xpu": backend = "triton" after the SM branches, scoped to XPU only.
Your guess about the reason was right and it is now in the comment: intel_xpu is absent from
RADIX_SUPPORTED_DETERMINISTIC_ATTENTION_BACKEND, so choosing it would silently trip the
disable_radix_cache fallback. ROCm/MPS/CPU deliberately unchanged — widening this in an XPU PR is
the drive-by you flagged elsewhere.

12. Nothing here is covered by CI. Fair, and the sharpest of the twelve. Added
test/registered/unit/utils/test_device_module_probes.py plus cases in test_http_server.py and
test_server_args.py, covering the device probes, the canary knobs per device string, the host-pool
rejection per opt-in path, and the deterministic fallback with the platform pinned so it holds on any
runner. 329 passing.
Two things I want to be straight about rather than let the green tick imply more than it should.
test/manual/ is excluded from CI and the CI runners are CUDA/CPU, so a green run on this PR carries
no XPU signal
— the new tests are device-agnostic by design and would pass identically on a CUDA
box. The XPU evidence is a manual A/B on one Arc Pro B60: on pristine main,
TestChunkSizeDefault is 7 errors in 322 s (install_canary -> "HTTP server did not connect"); on
this branch it is 7 passed in 189 s, and the full 302 run is 301 passed / 0 failures / 3 h 28 m.
And register_xpu_ci exists (47 files use it), while this PR registers nothing for it — so the
enablement is not currently CI-enforced. The cheap fix is registering
test/registered/chunked_prefill/test_scripted_core_1gpu.py and
test/registered/scripted_runtime/test_scripted_runtime_core.py for stage-b-test-1-gpu-xpu; happy
to add it here or as a follow-up, but I cannot validate the XPU runners myself, so the first CI run
would be the real test.


Also worth flagging, unrelated to the review: meta-llama/Llama-3.2-1B-Instruct currently fails on
XPU on main with RuntimeError: mat1 and mat2 shapes cannot be multiplied (1x14336 and 2048x2048),
raised from the plain model forward (eager_runner._execute_extend) with no LoRA in the stack. This
PR touches no model or LoRA file, and it reproduces with these commits applied to the parent of
#30345, so it is neither this PR's doing nor that commit's. It is the sole cause of the four
lora_overlap errors and of every e2e GSM8K error in the run above, and I will file it separately.

@siju-samuel

Copy link
Copy Markdown
Contributor

/tag-run-ci-label

@AnuSajikumar6264

Copy link
Copy Markdown
Contributor Author
  File "/usr/lib/python3.12/asyncio/runners.py", line 194, in run
    return runner.run(main)
           ^^^^^^^^^^^^^^^^
  File "/usr/lib/python3.12/asyncio/runners.py", line 118, in run
    return self._loop.run_until_complete(task)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "uvloop/loop.pyx", line 1512, in uvloop.loop.Loop.run_until_complete
  File "uvloop/loop.pyx", line 1505, in uvloop.loop.Loop.run_until_complete
  File "uvloop/loop.pyx", line 1379, in uvloop.loop.Loop.run_forever
  File "uvloop/loop.pyx", line 557, in uvloop.loop.Loop._run
  File "uvloop/loop.pyx", line 476, in uvloop.loop.Loop._on_idle
  File "uvloop/cbhandles.pyx", line 83, in uvloop.loop.Handle._run
  File "uvloop/cbhandles.pyx", line 63, in uvloop.loop.Handle._run
  File "/opt/venv/lib/python3.12/site-packages/sglang/srt/managers/tokenizer_manager.py", line 3701, in print_exception_wrapper
    await func()
  File "/opt/venv/lib/python3.12/site-packages/sglang/srt/managers/tokenizer_manager.py", line 3253, in sigterm_watchdog
    sys.exit(0)
SystemExit: 0

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/opt/venv/lib/python3.12/site-packages/starlette/routing.py", line 655, in lifespan
    await receive()
  File "/opt/venv/lib/python3.12/site-packages/uvicorn/lifespan/on.py", line 137, in receive
    return await self.receive_queue.get()
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/lib/python3.12/asyncio/queues.py", line 158, in get
    await getter
asyncio.exceptions.CancelledError

[2026-09-12 19:59:54] kill_process_tree called: parent_pid=15847, include_parent=False, pid=15847

======================================================================
ERROR: test_run (__main__.TestEncoderAttention.test_run)
----------------------------------------------------------------------
Traceback (most recent call last):
  File "/opt/venv/lib/python3.12/site-packages/sglang/srt/utils/common.py", line 3539, in retry
    return fn()
           ^^^^
  File "/opt/venv/lib/python3.12/site-packages/sglang/test/test_utils.py", line 2195, in <lambda>
    lambda: super(CustomTestCase, self)._callTestMethod(method),
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/lib/python3.12/unittest/case.py", line 589, in _callTestMethod
    if method() is not None:
       ^^^^^^^^
  File "/sglang-checkout/test/registered/xpu/test_encoder_attention_backend.py", line 109, in test_run
    self.run_decode()
  File "/sglang-checkout/test/registered/xpu/test_encoder_attention_backend.py", line 100, in run_decode
    assert_one_item(ret)
  File "/sglang-checkout/test/registered/xpu/test_encoder_attention_backend.py", line 87, in assert_one_item
    if item["meta_info"]["finish_reason"]["type"] == "stop":
       ~~~~^^^^^^^^^^^^^
KeyError: 'meta_info'

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/opt/venv/lib/python3.12/site-packages/sglang/test/test_utils.py", line 2194, in _callTestMethod
    retry(
  File "/opt/venv/lib/python3.12/site-packages/sglang/srt/utils/common.py", line 3554, in retry
    raise Exception(f"retry() exceed maximum number of retries.")
Exception: retry() exceed maximum number of retries.

----------------------------------------------------------------------
Ran 1 test in 45.250s

FAILED (errors=1)
{
  "error": {
    "message": "Qwen-VL image artifacts do not match prompt placeholders"
  }
}
.
.
End (5/29):
filename='/sglang-checkout/test/registered/xpu/test_encoder_attention_backend.py', elapsed=52, estimated_time=360.0
.
.


✗ FAILED: /sglang-checkout/test/registered/xpu/test_encoder_attention_backend.py returned exit code 1

Fail. Time elapsed: 616.07s

The xpu failure is unrelated to my changes.

@mingfeima mingfeima left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

keep change smaller.

Comment thread python/sglang/srt/layers/attention/aiter_backend.py Outdated
Comment thread python/sglang/srt/layers/attention/wave_backend.py Outdated
Comment thread python/sglang/srt/utils/common.py Outdated
@mingfeima mingfeima added intel xpu intel gpu with device `torch.xpu` labels Sep 16, 2026
@mingfeima

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@mingfeima mingfeima left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

keep changes minimal and focused.

Comment thread python/sglang/srt/managers/scheduler.py Outdated
Comment thread python/sglang/test/scripted_runtime/context/http_post.py Outdated
Comment thread python/sglang/test/scripted_runtime/background_http_poster.py Outdated
Comment thread python/sglang/test/scripted_runtime/http_server.py Outdated
@AnuSajikumar6264
AnuSajikumar6264 force-pushed the chunked_prefill branch 4 times, most recently from 42875ba to e70ca36 Compare September 24, 2026 03:28
@mingfeima

Copy link
Copy Markdown
Collaborator

LGTM now. XPU ci fail on main has been fixed, need rebase.

The scripted runtime could not bring up an engine on --device xpu, so none of
test/manual/chunked_prefill/ ran there. Every change here is test-side.

- kv-canary now has an XPU path: its write / verify / plan kernels are CUDA-JIT
  only, so they route to a torch reference. install_canary refuses to capture a
  decode over that reference, so the harness turns the decode graph off where
  the canary falls back, instead of turning the canary off. The decision is made
  after engine_kwargs merge and only when the canary is actually enabled, so a
  class that opts out of the canary keeps its decode graph. That reference folds
  on the host, so the script timeout is scaled where it is in play; a 100-chunk
  prefill needs about 220s of the 120s the CUDA path was budgeted.
- TestSpecialCaseHiCache needs the MHA host-pool movers, which pool_host/mha.py
  builds only for CUDA and ROCm, so it is skipped off those two platforms rather
  than off XPU alone.
- TestSpecialCaseDeterministicFlashInfer needs flashinfer, which ships NVIDIA-only
  kernels.

Test-side fixes the above exposed:

- Reusing a rid needs the TokenizerManager to have released it, which is not
  observable from the scheduler, so the settle helper is named for the wait it
  performs rather than for a condition it never checks.
- _drain_until_released had drifted into four copies; it is now one helper, and
  it waits on lock_refs, which drops an iteration after the KV pages.
- The abort-release regression guards read through find_req_by_rid, which stops
  seeing a req whether it released or leaked, so they passed on a leaked row.
  They now compare the pool free counts.
- Dropped the double-abort test: since sgl-project#35255 the TokenizerManager returns early
  on a repeat abort, so the scheduler never sees the second AbortReq and the
  test could not exercise the idempotency it was named for.

New CPU-CI unit tests cover that the canary stays enabled on every device and
that an opted-out canary leaves the decode graph alone.
@mingfeima
mingfeima merged commit 2ff52e3 into sgl-project:main Sep 29, 2026
162 of 180 checks passed
arbi-dev added a commit to arbicity/sglang-turbo that referenced this pull request Oct 7, 2026
…5.21 (#12)

* [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG (sgl-project#40384)

* [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (sgl-project#40405)

* Fix TBO child batch missing dp_spec_prefill_coordination_applied (sgl-project#41096)

* [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe (sgl-project#39059)

* [Diffusion] Fuse lossless LingBot World FP32 normalization (sgl-project#40425)

* [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens (sgl-project#40878)

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>

* [AMD] Add tuned dsv4 shape (sgl-project#40996)

* [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency (sgl-project#40388)

* [Diffusion] migrate the whole _register_configs from registry.py to the model own config file (sgl-project#40612)

* Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.) (sgl-project#39816)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>

* [HiCache] Make host reclamation independent of transfer order (sgl-project#40512)

* [Refactor] Retire the model-specific Kimi K3 kernel namespace (sgl-project#40922)

* [diffusion] update code owner (sgl-project#41130)

* [NPU] Update CANN version to 9.1.0 (sgl-project#40524)

* [PD] Enable deferred decode-side KV release by default (sgl-project#41023)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [chore] surface the cookbook to users who pip install sglang (sgl-project#40866)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(sampling): validate sampling_seed is an int within int64 range (sgl-project#28960)

Signed-off-by: Ting Sun <suntcrick@gmail.com>

* [HiCache] Demote SWA KV to host on write_back eviction instead of dropping it (sgl-project#40712)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [AMD] ci: move the Miles ROCm 7.2 nightly build to 7.2.4 (sgl-project#40387)

Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>

* [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size (sgl-project#38726)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: Triton 3.8 compatbility to support DSV4.1-Flash in CUDA 13.4 image (Rubin) (sgl-project#40805)

* Support unified memory decode host pools (sgl-project#39478)

Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>

* [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune (sgl-project#41138)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [DP attention] Publish DP buffer sizes from a ForwardBatch (sgl-project#40858)

Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>

* [Test] Run the Qwen3.5 Triton DCP nightly with the radix cache enabled (sgl-project#41155)

* [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260923 daily (sgl-project#41109)

* [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send (sgl-project#41079)

* [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture (sgl-project#41080)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Pass each layer stack's output through a communicator exit (sgl-project#41081)

* [Refactor] Share the MoE output all-reduce between models (sgl-project#41097)

* [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention (sgl-project#41082)

* [Fix] Broadcast requests along attention CP before attention TP (sgl-project#41083)

* [Refactor] Split prepare_attn into a reduction step and per-quant-format residual steps (sgl-project#41084)

* [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt (sgl-project#41120)

* [qwen 3.8 next] Fuse Qwen PLE gate and convolution preparation for target verify (sgl-project#40041)

* MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950 (sgl-project#36574)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>

* MoE: small-batch sorting path with fused mxfp8 quantisation (sgl-project#36559)

Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting (sgl-project#41090)

* [Fix] Patch set_dp_buffer_len_from_batch in DP spec prefill coordination test (sgl-project#41162)

* [Docs] Enable Qwen3.8 Flash Next NVIDIA NVFP4 on B200/B300/GB300 (sgl-project#41046)

* [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator (sgl-project#41049)

* [Bugfix] Align DeepSeek-V4.1 reasoning effort budgets (sgl-project#39929)

Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>

* Fix mixed chunk prefill with DP speculative coordination (sgl-project#41179)

* Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++ (sgl-project#37442)

Co-authored-by: Caio Rocha <caiorocha@microsft.com>

* [HiCache] Batch buffer-only KV backups within each flush (sgl-project#40960)

Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>

* [PD] Add a `none` decode retraction backup and subclass seams in the PD queues (sgl-project#41103)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>

* [Score API] Setwise Scoring Support (sgl-project#38965)

* [DSV4] Account for FlashMLA physical KV page padding in memory budgets (sgl-project#41091)

Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [mem_cache] Drop `is_insert` from `cache_finished_req`; release rows from `release_kv_cache` (sgl-project#40988)

* [AMD] Restore non-DCP Mamba checkpoint donation to fix agent-mode cache hit at high conc with HiCache (sgl-project#40907)

* [diffusion] feat: add opt-in SRT prompt enhancement to image and video APIs (sgl-project#41095)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* Support XQA backend for SpecDec verify (sgl-project#32269)

* [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM (sgl-project#40793)

Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>

* [Perf] Lazy-load built-in model definitions and nixl_ep at startup (sgl-project#41061)

Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>

* [CI] Move GLM-5.2 layer-split test to extra-b-test-8-gpu-b300 (sgl-project#41207)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Fix] Keep the target's DP sync slot in draft scopes (sgl-project#41062)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>

* [feat] add a system one compatible /v1/systemone route (sgl-project#41208)

* fix(openai): reject request-supplied chat_template by default (sgl-project#28135)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels (sgl-project#41159)

* [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound (sgl-project#41201)

Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>

* [diffusion] model: support Ming-Image Design and Design-Layer (sgl-project#41067)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain (sgl-project#40456)

* [HiCache] fix: Drain pending backups before internal Mamba write-back (sgl-project#41092)

* [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4 (sgl-project#35619)

* [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch (sgl-project#41193)

* [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft (sgl-project#41194)

* [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv (sgl-project#41195)

* [Refactor] Carry a deferred FFN all-reduce as UnreducedOutput and complete it in the next layer without the fused kernel (sgl-project#41196)

* [Refactor] Leave the FFN reduction to the next layer under attention DP (sgl-project#41197)

* [Refactor] Move Step-3.5, GLM5-Next, Dots3, MiniMax-M3 and Qwen3.5 onto ffn_exit (sgl-project#41198)

* [Refactor] Build prepare_mlp and the layout moves from named steps (sgl-project#41191)

* [Refactor] Take a layer's last-layer fact from its scatter-mode plan (sgl-project#41199)

* [Refactor] Decide an FFN exit's completion once and declare the group it owes (sgl-project#41200)

* [XPU] Disable test_ngram_corpus on XPU and extend XPU CI path filter (sgl-project#41224)

* [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager (sgl-project#34417)

Co-authored-by: mickqian <mickqian@users.noreply.github.com>

* [diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (sgl-project#40486)

* Revert "[NPU] Fuse FIA KV-cache K/V writes into one npu_scatter_pa_kv_cache call" (sgl-project#41132)

* [ROCm][Bugfix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints (sgl-project#39064)

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>

* fix(multimodal): return 400 for corrupt image inputs (sgl-project#28131)

Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>

* [unified-memory] Honor move gates in float relocation and size auto HiCache from host capacity (sgl-project#41248)

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>

* Fix multimodal feature offload races (sgl-project#40621)

Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>

* [Sampling] Stream sampling masks as per-request arrays (sgl-project#40986)

* [Rust frontend] Decode input_ids without untagged buffering (sgl-project#41246)

Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>

* [Test] Remove dead eval modules and point GSM8K/MMLU docs to sgl-eval (sgl-project#41215)

* [Test] Add in-process sgl-eval adapter and move validated GSM8K tests to it (sgl-project#41216)

* [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard (sgl-project#39379)

* [KVCache] Support lmcache unified radix cache (sgl-project#38652)

Signed-off-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sglang][lora] Support DP attention in LoRA backends (sgl-project#36389)

Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>

* dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers (sgl-project#41018)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [Test] Route all sgl-eval benchmarks through run_sgl_eval and deprecate run_eval (sgl-project#41280)

* [Fix] Fix cpu CI fail introduced by pr sgl-project#38652 (sgl-project#41287)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [PD] Share one head-slice helper across mooncake, mori, and nixl (sgl-project#39660)

Co-authored-by: BBuf <1182563586@qq.com>

* [misc] Call all_gather_single / reduce_scatter_single to drop torch deprecation warnings (sgl-project#41284)

* [Test] Remove unit tests that only mirror implementation or never run in CI (sgl-project#41286)

* [DCP] Use logical token capacity for PD admission and load reporting (sgl-project#39731)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [CI] Harden `/rerun-test` dispatch and partitioning (sgl-project#41285)

* [diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper (sgl-project#40490)

* [DSv4.1] Move the ratio-1/2 index top-k ops into kernels/ops/attention/dsv4 (sgl-project#41291)

Co-authored-by: DarkSharpness <2040703891@qq.com>

* [diffusion] model: support Anima Base v1.0 (sgl-project#41011)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget (sgl-project#40854)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match (sgl-project#41261)

* [Test] Remove more unit tests that mirror implementation or never run in CI (sgl-project#41297)

* [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill (sgl-project#36505)

* [MemCache] Unify component eviction cursors and lock receipts (sgl-project#41276)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [diffusion] fix: fix Qwen-Image 2.1 default RGBA output (sgl-project#41150)

Co-authored-by: jacky.cheng <yichiche@amd.com>

* [diffusion] fix: copy small files into overlay materialized trees instead of linking them (sgl-project#41298)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Refactor] Restore logical kernel groups and test organization (sgl-project#41243)

* [sgl-router] Track input_ids forwarding outcomes per chat request (sgl-project#41185)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [DSv4.1] Move the low-ratio index top-k into dsv4/low_ratio_indexer (sgl-project#41125)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention (sgl-project#41311)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Score API] Setwise scoring: CausalLM support (batched + --enable-mis) (sgl-project#41188)

* [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving) (sgl-project#41223)

Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>

* [DeepSeek V4.1] Add DeepSelect JIT kernel. (sgl-project#40556)

Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Refactor] Choose prepare_attn / prepare_mlp steps and fused kernels at construction (sgl-project#41252)

* [Refactor] Drive an FFN exit's flags and its completion from one selection (sgl-project#41253)

* [Refactor] Declare at construction the layers whose FFN completes its own reduction (sgl-project#41254)

* [Refactor] Run the LayerNorm SP region's boundary steps in the communicator itself (sgl-project#41255)

* [Refactor] Pick the two-batch-overlap split's layout moves once and remove execute (sgl-project#41256)

* [Refactor] Choose a dense layer's boundaries under attention DP from both sides' declarations (sgl-project#41257)

* [CI] Add __main__ entry to test_deep_select.py (sgl-project#41341)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [sgl-router] Book the input_ids forwarding outcome only for built bodies (sgl-project#41342)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [CI] Move DeepSelect into the attention kernel group to fix the namespace test (sgl-project#41349)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090) (sgl-project#41292)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* [CI] Merge the Kimi-Linear PD DCP4 nightly tests and drop exact-token parity (sgl-project#41321)

* [AMD] Fix jit broken on rocm env (sgl-project#41356)

* Make sliding-window caching and speculative batch padding extensible (sgl-project#41325)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>

* [MemCache] Fix LMCache component cursors and per-cache backend selection (sgl-project#41328)

Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>

* [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm (sgl-project#41377)

* dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels (sgl-project#41019)

Co-authored-by: x <x>

* [sgl-router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs (sgl-project#41221)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [mem_cache] Never free the protected prefix on request release (sgl-project#41312)

* [DSV4.1][HiCache] fix: wait for the layer transfer before reading low-ratio index-K (sgl-project#41345)

* [Diffusion] Preserve per-sample rollout trajectories across multi-output merge (sgl-project#34416)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>

* [kv-shard 3/4] Enable Control Plane B (sgl-project#39964)

Co-authored-by: Zhangheng <hzh0425@apache.org>

* [qwen 3.8 next] Fuse small CUDA graph input buffer copies (sgl-project#41166)

* [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200 (sgl-project#41305)

Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>

* [mem_cache] Replace `cache_finished_req` with `insert_req`; `release_kv_cache` frees and unpins (sgl-project#41281)

* [diffusion] feat: support in-place lora merge/unmerge under layerwise offload (sgl-project#36192)

* [diffusion] feat: minimax-h3 spectrum skip-step + fused RMSNorm/AdaLN (sgl-project#35684)

* [diffusion] feat: add MiniMax-H3 to ComfyUI integrated mode (sgl-project#35990)

* [Diffusion] Apply latent-ids and packing to caller-provided initial latents (sgl-project#34418)

Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>

* [diffusion] fix: lora-wrapped linears crash qwen-image and minimax-h3 inference (sgl-project#41272)

Signed-off-by: rockdu <kangrdu@gmail.com>

* [diffusion] fix: run an all-valid attention mask on the backend's unmasked kernel (sgl-project#41309)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* [VLM] Introduce FA4 into ViT for SM100/SM103 (sgl-project#41344)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>

* [bench] Take each request's prompt length from the server (sgl-project#39889)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>

* [Diffusion] Reuse bit-exact packed SwiGLU for Ming-Image (sgl-project#41266)

* Add MiniMax arch fallback to auto parser resolution (sgl-project#40930)

* [HiCache] Add the page-unified KV load-back JIT kernel  (sgl-project#39726)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>

* [AMD] Update v4 cookbook for megamoe, fp8 kv attn, BCG (sgl-project#41458)

* [PD] Fan drain abort ACKs out to every decode peer of the room (sgl-project#41402)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts (sgl-project#37462)

Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Port chat_parsing core (sgl-project#40477)

Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>

* [Refactor] Choose the boundaries of plain-TP dense layers and MoE layers from declarations (sgl-project#41417)

* [Refactor] Give the fused prepare_mlp kernels an explicit contract (sgl-project#41418)

* [Refactor] Choose the LayerNorm SP region's and input-scattered batches' steps from declarations (sgl-project#41419)

* [Refactor] Run every batch of a layer from one BoundarySteps (sgl-project#41420)

* [Refactor] Choose a GQA prefill CP extend's steps from declarations (sgl-project#41421)

* [Fix] Keep one copy of CP-replicated rows in the DP gather (sgl-project#41422)

* [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum (sgl-project#41423)

* [Refactor] Choose fully-DP dense and DSA / MLA prefill CP layers' steps from declarations (sgl-project#41424)

* [Refactor] Run MHC layers on the shared boundary steps with MHC's residual operations (sgl-project#41425)

* [Refactor] Run two-batch-overlap layers on the declared boundaries (sgl-project#41426)

* [Refactor] Let the MoE declare whether its skipped reduction is one TP all-reduce (sgl-project#41427)

* [Refactor] Publish the LoRA token layout from the batch's FFN input rows (sgl-project#41428)

* [Refactor] Build each decoder boundary from the declarations of its two sides (sgl-project#41429)

* [Refactor] Nemotron-H: build each layer's boundaries from its stage and the previous one (sgl-project#41430)

* [Refactor] Move CuTe DSL-fused layers onto the declared boundaries and FFN-exit kernel entries (sgl-project#41431)

* [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather (sgl-project#41432)

* [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism (sgl-project#41433)

* [Refactor] Falcon-H1: complete the FFN's sum through ffn_exit (sgl-project#41434)

* [Refactor] Choose MoE layers' boundaries from declarations when moe_dp_size equals attn_cp_size (sgl-project#41435)

* [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators (sgl-project#41436)

* [Refactor] Step-3.5: complete the dense MLP's sum through ffn_exit (sgl-project#41437)

* [Refactor] Choose every layer's boundaries from declarations and remove the scatter-mode selection (sgl-project#41438)

* [Refactor] Split the layer communicator into a package (move only) (sgl-project#41439)

* [Refactor] Declare each stage's residual read and update, and give each stage its own entry (sgl-project#41440)

* [Refactor] Build the boundary into any stage with one construction (sgl-project#41441)

* [Refactor] Give a layer its CuTe DSL kernels at construction instead of a subclass (sgl-project#41442)

* [Refactor] Replace LayerScatterModes with LayerFacts and remove ScatterMode (sgl-project#41443)

* [XPU] Disable test_ngram_corpus on XPU and detect XPU tests dynamically in CI filter (sgl-project#41075)

* [Fix][NPU] Fix performance degradation caused by serial execution of two-stage ACL op calls on graph (sgl-project#40814)

* [diffusion] feat: support multiple task types for pipelines (sgl-project#38762)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [CI] fix CI regression on xeon (sgl-project#41002)

* [CI] Real-model Kimi-Linear PD parity at page, DCP virtual-page, chunk and cached-prefix boundaries (sgl-project#41378)

* [AMD] Register Triton data movement tests in PR CI (sgl-project#41137)

* [AMD] Add GLM-5.3-Flash MI35x nightly test (sgl-project#36903)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>

* [AMD] [Docker] Remove unused LLVM 18 setup from ROCm TileLang build (sgl-project#41387)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>

* [Intel GPU] Xpu/weekly simple model enablement 2026 09 21 (sgl-project#40664)

Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Bugfix] fix(hicache): wait for decode offload before retraction (sgl-project#30899)

Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Rust server unify datapath for mm and generate requests (sgl-project#39679)

* feat(npu): Support returning indexer top-k results (sgl-project#39060)

* Fix chat template cache key order (sgl-project#41517)

* docs: add prefill context parallelism guide and design draft (sgl-project#39354)

* [XPU] Bump sglang-kernel-xpu wheel to v0.3.0 (sgl-project#41220)

Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>

* [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping (sgl-project#41164)

* [Router] Give the cache-aware tree a snapshot surface (1/13) (sgl-project#40687)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Let predicate-registered linear-attention models carry the mamba radix-cache leaves (sgl-project#41165)

* [diffusion] optimization: populate cpu weight stores before host registration to speedup layerwise-offload initialization (sgl-project#40439)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* perf(multimodal): offload CPU feature hashing with bounded admission (sgl-project#39539)

* [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe) (sgl-project#40943)

* [sgl-router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4 (sgl-project#41226)

Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>

* [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13) (sgl-project#40688)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [AMD] [GLM5] Fuse shared expert into AITER MoE on gfx950 (sgl-project#41161)

Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>

* [Router] Name a replica's siblings with --kv-peer-selector (3/13) (sgl-project#40689)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [diffusion] docs: consolidate Qwen-Image 2.1 guidance in its cookbook (sgl-project#41540)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Unified Memory] fix: preserve FP8 dtype in unified MHA pool (sgl-project#38133)

* [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers (sgl-project#41144)

* [Doc] Add kernel benchmark rule on L2 cache reuse (sgl-project#41545)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [DeepSelect] Add page-table transform to top-k and tighten the layout contract (sgl-project#41364)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* [diffusion] Qwen-Image 2.1: fuse Q/K RMSNorm + RoPE + KV packing into one CUDA kernel and project Q/K/V with one packed GEMM (sgl-project#41339)

* [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU (sgl-project#40446)

* [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism (sgl-project#39643)

* [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts (sgl-project#41404)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* [PD] Keep the sampling mask of a replayed rebootstrap token (sgl-project#41235)

* [NVIDIA] Update deepgemm, deep-ep, sgl-kernel in CUDA 13.4 image, use cuda base image (sgl-project#40987)

* [Docs][AMD] Update GLM-5.2 MI355X daily image (sgl-project#41597)

* dsv4.1-amd: gfx950 sparse decode attention and sorted top-k (sgl-project#41020)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels (sgl-project#41021)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>

* [Model Loader] Stop checkpoint prefetch after iterator completion (sgl-project#41588)

Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>

* [mem cache] refactor: remove the index-K continuous getters orphaned by the CP v1 removal (sgl-project#41469)

* [PD] Keep EAGLE DP graph and token metadata consistent (sgl-project#32196)

Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>

* [Docs] Add GigaChat 3.5 and GigaChat 3.5 Reasoning cookbook pages (sgl-project#41118)

Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>

* [Model] Add IQuest Q1 support and MTP draft (sgl-project#41590)

Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>

* [diffusion] model: support flux 3 action robot policies (sgl-project#41066)

* [Radix Cache] Sync Rust TreeCore and make it the default (sgl-project#39627)

Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>

* [npu]support NPU 910C L2 memcache offload (sgl-project#41527)

* [Test] Fail fast when PD test RDMA devices are not openable by ibverbs (sgl-project#41600)

* Remove GLM-4.1V-9B-Thinking from encoder DP MMMU test (sgl-project#41608)

Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>

* [Refactor] Group communicator fusion and CP adapters (sgl-project#41547)

* [Refactor] Centralize decoder output access (sgl-project#41548)

* [Refactor] Carry residual state across stage boundaries (sgl-project#41549)

* [Refactor] Capture auxiliary states at residual reads (sgl-project#41550)

* [Refactor] Select reduction fusion at the consumer (sgl-project#41551)

* [Refactor] Construct independent decoder stage boundaries (sgl-project#41552)

* [Refactor] Migrate specialized decoder and overlap boundaries (sgl-project#41553)

* [Refactor] Retire the layer facade and simplify boundary internals (sgl-project#41554)

* [Refactor] Rename the module to layer_boundary (sgl-project#41555)

* [Refactor] Group layer boundary unit tests (sgl-project#41556)

* [Refactor] Document layer boundary contracts and integration (sgl-project#41557)

* [Rust] Extract a transport-neutral frontend core (sgl-project#39385)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* [PD] Give FakeKVReceiver ensure_abort_notified (sgl-project#41618)

* [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path (sgl-project#40828)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [Intel][XPU]Enable chunked prefill scnearios for XPU with UT (sgl-project#33804)

Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>

* [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts (sgl-project#34528)

* [diffusion] feat: support bounded exact conditioning cache across native models (sgl-project#40470)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>

* [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B (sgl-project#41059)

* [Cherry-pick to release/v0.5.21] [Fix] Restore deferred layer dumps and pin SentencePiece for InternVL (sgl-project#41777) (sgl-project#41788)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>

* [Cherry-pick to release/v0.5.21] [Test] Demote PD test RDMA openability check to a warning (sgl-project#41681) (sgl-project#41802)

Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>

* feat(plugin-seam): forward-port the turbo-attn seam onto upstream v0.5.21

The whole carry of arbicity/sglang-turbo main @71bcc9df9b over upstream
v0.5.18, squashed and replayed onto upstream v0.5.21 (e00930c, the
commit lmsysorg/sglang:v0.5.21 is built from) as a 3-way merge. Folds in
the post-port commits #9 (hybrid sliding-window through the kv-cache
plugin), #10 (Gemma 4 on a plugin backend) and #11 (draft backend factory
reads attention_backends()).

Upstream's structure wins and the seam is re-applied on it:
- server_args.py is now a thin record over arg_groups/: the
  --kv-cache-dtype choices are hoisted into arg_groups/choices.py as
  KV_CACHE_DTYPE_CHOICES + add_kv_cache_dtype_choices (re-exported from
  server_args), and the gpt-oss / Gemma 4 backend whitelists that accept a
  plugin backend move to arg_groups/model_hook.py.
- ServerArgs._handle_plugin_kv_cache_pairing is dropped: v0.5.21's
  resolution hooks (arg_groups/resolution_hooks.py) are the official slot,
  and turbo-attn's plugin now pairs the flags there.
- Plugin dtype reads go through the resolved model bag (get_model()), not
  the raw ServerArgs record, which v0.5.21 leaves as operator input.
- HybridSWAPoolConfigurator: the plugin prices the full and SWA layers it
  holds and stands in as the per-layer cost, so upstream's new draft-SWA
  and unified-pool formulas apply unchanged; layer ids come from
  kvc.layer_info, as upstream's own SWA pool takes them.
- Qwen3.5 NEXTN embed on meta only on a single pipeline stage: from v0.5.21
  a PP draft may load its own embedding (pp_draft_embedding).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Yanfei Wang <yanfwang@crsuse2-m2m-v2-017.us-east2-a.compute.internal>
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Signed-off-by: chunxiaozheng <1179548172@qq.com>
Signed-off-by: rockdu <kangrdu@gmail.com>
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
Co-authored-by: AMD-yanfeiwang <yanfei.wang@amd.com>
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: paulzhang-tm <paulzhang@thinkingmachines.ai>
Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>
Co-authored-by: 黄孝君 <dingfangsu23@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Ting SUN <suntcrick@gmail.com>
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
Co-authored-by: Jackey Hua <107608053+zhendonghua@users.noreply.github.com>
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: hanminglu <hanminglu@fb.com>
Co-authored-by: oss-sync bot (pranjalssh) <pranjalssh@users.noreply.github.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: ChangLiu0709 <cliu1004@amd.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: mikevin920 <mikevin920@yahoo.com>
Co-authored-by: Kevin Mi <mikevin920@gmail.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Co-authored-by: Apexsf <50563213+Apexsf@users.noreply.github.com>
Co-authored-by: fengtinglei <fengtinglei@bytedance.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Caio Rocha <164253795+caiocbr@users.noreply.github.com>
Co-authored-by: Caio Rocha <caiorocha@microsft.com>
Co-authored-by: metamergebot <metamergebot@gmail.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <zqx@meta.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: akhilg-nv <165961486+akhilg-nv@users.noreply.github.com>
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>
Co-authored-by: Richard Wang <wangrichard08@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyi Song <xinyis10@illinois.edu>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: ashwini rathi <ashwini.rathi@intel.com>
Co-authored-by: Jianghai <72591262+CjhHa1@users.noreply.github.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Justin Tong <justintong0323@users.noreply.github.com>
Co-authored-by: cctry <csycfl@gmail.com>
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Yongji Wu <yongji@meta.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com>
Co-authored-by: Lucia Fang <116399278+luccafong@users.noreply.github.com>
Co-authored-by: Ethan (Yusheng) Su <yushengsu@radixark.ai>
Co-authored-by: chunxiaozheng <1179548172@qq.com>
Co-authored-by: Erik Wijmans <erik@thinkingmachines.ai>
Co-authored-by: James Liu <51351043+chromecast56@users.noreply.github.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
Co-authored-by: WMC <tnwilly@gmail.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: Kan Wu <kan.wu@radixark.ai>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Hexu Zhao <45677459+TarzanZhao@users.noreply.github.com>
Co-authored-by: hexu.zhao <zhaohexu2001@gmail.com>
Co-authored-by: yuyu5333 <77156718+yuyu5333@users.noreply.github.com>
Co-authored-by: Ariel <32339430+arieller@users.noreply.github.com>
Co-authored-by: jasonjk-park <jasonjk@fb.com>
Co-authored-by: Ankith Averineni <saverine@amd.com>
Co-authored-by: aimicahchen <aimicahchen@tencent.com>
Co-authored-by: Shunkangz <182541032+Shunkangz@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: kernel-design-agents <334018530+kernel-design-agents@users.noreply.github.com>
Co-authored-by: Kangrui Du <89372739+Rockdu@users.noreply.github.com>
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Nguyễn Văn Cao Nguyên <nvcnvn1@gmail.com>
Co-authored-by: CuzMi <simon.weijie@gmail.com>
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: mrusanovsky <mrusanovsky@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com>
Co-authored-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>
Co-authored-by: Kurkur <102506892+litmei@users.noreply.github.com>
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Xinguo Zhu <xinguo.zhu@intel.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
Co-authored-by: quitenode <quitenode@users.noreply.github.com>
Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com>
Co-authored-by: Amrutha M <amrutha.m@intel.com>
Co-authored-by: charlesxu91 <charlesxu.mi@gmail.com>
Co-authored-by: KMS07 <meher.sai.kotthagattu@intel.com>
Co-authored-by: Avi Fenesh <aviarchi1994@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Kevin Flansburg <kflansburg@cloudflare.com>
Co-authored-by: agent <agent@local>
Co-authored-by: Kevin Flansburg <6134007+kflansburg@users.noreply.github.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>
Co-authored-by: flb_ <floatlibai@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Pramod Kumar <144990617+pramodkumar-habanalabs@users.noreply.github.com>
Co-authored-by: Peng Wu <peng@thinkingmachines.ai>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Denji-kk <51020465+Denji-kk@users.noreply.github.com>
Co-authored-by: karverma-amd <karan.verma@amd.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: triple-mu <gpu@163.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Zhang, Jiejing <280536569+jiejingzhangamd@users.noreply.github.com>
Co-authored-by: Hrithvik Alex <halex623@gmail.com>
Co-authored-by: Hrithvik Alex <hrithvik@baseten.co>
Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <50638944+GungnirAP@users.noreply.github.com>
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
Co-authored-by: zelong huang <yzhu@ubiquant.com>
Co-authored-by: zelong518 <zelonghuang05@gmail.com>
Co-authored-by: chenwenxiaolive <cocoshirolive@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: alphabetc1 <alphabetc1@users.noreply.github.com>
Co-authored-by: James <445169590@qq.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Rohit Kumar Singh <9626333+SKRohit@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Singh <rohitsi2@iil-gnrap01.iind.intel.com>
Co-authored-by: Singh <rohitsi2@iil-login.iind.intel.com>
Co-authored-by: Anupa Sajikumar <anupa.sajikumar@intel.com>
Co-authored-by: eeecho <82921722+AliceChenyy@users.noreply.github.com>
Co-authored-by: Jiacong Fang <zldrobit@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang intel lora run-ci CI: run the baseline test suite on this PR xpu intel gpu with device `torch.xpu`

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants