Skip to content

[Test] Add get_swa_key_page_size to the Q8KV8 sparse-prefill fake KV pool - #42211

Merged
nvpohanh merged 1 commit into
sgl-project:mainfrom
nvpohanh:fix/q8kv8-sparse-prefill-swa-page-size
Oct 2, 2026
Merged

nvpohanh merged 1 commit into
sgl-project:mainfrom
nvpohanh:fix/q8kv8-sparse-prefill-swa-page-size

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

[by Claude Code]

Motivation

#42128 changed DeepseekV4AttnBackend to read the SWA page size through token_to_kv_pool.get_swa_key_page_size() instead of token_to_kv_pool.swa_kv_pool.page_size. The fake _TokenToKVPool in test/registered/kernels/ops/attention/test_q8kv8_sparse_prefill_backend.py does not define that method. Since #42128 merged, four tests fail on main in the jit-kernel-unit-test-1-gpu-large shard:

AttributeError: '_TokenToKVPool' object has no attribute 'get_swa_key_page_size'. Did you mean: 'get_extra_key_page_size'?
  python/sglang/srt/layers/attention/deepseek_v4_backend.py:2423 (_build_sparse_prefill_chunk_cache)
  • test_q8kv8_sparse_prefill_helper_builds_fp8_workspace_matching_bf16_path[0|4|128]
  • test_q8kv8_sparse_prefill_real_kernel_matches_bf16_sparse_path

This breaks every PR that runs the jit-kernel suite, regardless of what the PR changes. Examples:

Also, the shard stops at the first failing file, so the tests scheduled after this one in that shard never run.

Modifications

Add get_swa_key_page_size() to the fake pool. It returns swa_kv_pool.page_size, matching the real DeepSeekV4TokenToKVPool when no request window is configured. This is a 3-line, test-only change with no production code changes.

This is the same change as commit 94d21e3 in #41251, split out so main is unblocked without waiting for that larger perf PR. The hunk is identical, so #41251 should merge cleanly on top of it.

Test Escape Postmortem

Summary: #42128 broke this test on main, but its own PR Test Base run passed. The suite that contains this test never ran on #42128, and the same gap let #38954 break the same fake pool in September.

Timeline (UTC)

Time Event
2026-10-01 21:28 #42128's PR Test Base run passes (run 36928953138). Its check-changes step logs Filter jit_kernel = false, so call-jit-kernel-tests was skipped.
2026-10-01 23:10 The scheduled all-suites run on main passes, including both jit-kernel-unit-test-1-gpu-large shards (run 36939300008). It ran on f17f7705a5, which predates #42128.
2026-10-02 01:50 #42128 merges as ef867fa40d.
2026-10-02 02:24 First failure, on an unrelated PR that runs the jit-kernel suite (job 110676994618). The same signature then appears on other unrelated PRs (links in Motivation).
2026-10-02 07:47 This fix PR is opened.
2026-10-02 11:10 Next scheduled main run, the first one that would run this suite with #42128 included.

Why CI didn't catch it

  • test_q8kv8_sparse_prefill_backend.py lives under test/registered/kernels/ and is registered in the base-b-kernel-unit stage. That stage runs only inside call-jit-kernel-tests.
  • call-jit-kernel-tests runs only when check-changes sets jit_kernel=true.
    • The filter in .github/workflows/_pr-test-check-changes.yml matches only kernel and CI paths: test/registered/kernels/**, python/sglang/kernels/**, .github/workflows/pr-test.yml, .github/workflows/pr-test-jit-kernel.yml, and python/pyproject.toml.
  • [Fix][DSV4.1] SWA page size with bounded replay #42128 changed none of those paths. It touched only python/sglang/srt/layers/attention/deepseek_v4_backend.py, python/sglang/srt/mem_cache/deepseek_v4_memory_pool.py, and test/registered/unit/mem_cache/test_dsv4_compressed_pools.py.
  • Despite its location, this is not a pure kernel test. It drives DeepseekV4AttnBackend._forward_prefill_sparse in srt through a hand-written fake _TokenToKVPool. Any change to the pool interface that the backend calls can therefore break the test without triggering it.
  • What the gating covers ✅:
    • PRs that touch kernel or kernel-test paths run this test.
    • The scheduled main runs (11:10 and 23:10 UTC) run every suite.
  • What it misses ❌:
    • PRs that change only the DSV4 srt backend or memory pool this test depends on.
    • The scheduled runs catch such breakage only after merge, and by then unrelated PRs are already failing.
  • This is a repeat escape. [Refactor] Generalize DeepSeek V4 compressed pool management #38954 changed only srt files and unit tests, renamed a pool field, and broke the same fake pool with an AttributeError; [Test] Fix Q8KV8 sparse-prefill pool fixture after page-size rename #39019 fixed it. Both times the fake pool drifted from the real DeepSeekV4TokenToKVPool, and the path filter kept the breaking PR from running the test.

Remediation (follow-ups, not in this PR)

  • Close the trigger gap. Two options:
    • Add the srt paths this test depends on to the jit_kernel filter: at least python/sglang/srt/layers/attention/deepseek_v4_backend.py, python/sglang/srt/layers/attention/dsv4/**, and python/sglang/srt/mem_cache/deepseek_v4_memory_pool.py.
    • Or move the backend-level tests in this file into a suite that main-package changes already trigger.
  • Reduce fake-pool drift. Derive the fake from the real DeepSeekV4TokenToKVPool interface, for example by subclassing it or using create_autospec, instead of hand-listing methods. Accessors added to the real pool are then inherited or fail at the fixture, rather than surfacing only when the suite happens to run.

Accuracy Tests

Speed Tests and Profiling

Not applicable: test-only change.

Checklist

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ❌ Run #36980423210
Latest PR Test (Extra): ❌ Run #36980422882
Latest PR Test (AMD ROCm 10): ❌ Run #36980423204

…pool

sgl-project#42128 made DeepseekV4AttnBackend read the SWA page size through
token_to_kv_pool.get_swa_key_page_size(), but the fake _TokenToKVPool in
test_q8kv8_sparse_prefill_backend.py does not define it. Four tests in the
jit-kernel 1-gpu-large shard now fail on main with AttributeError. Add the
accessor to the fake pool, mirroring the real pool's default branch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nvpohanh

nvpohanh commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/kernels/ops/attention/test_q8kv8_sparse_prefill_backend.py

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/kernels/ops/attention/test_q8kv8_sparse_prefill_backend.py:

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/kernels/ops/attention/test_q8kv8_sparse_prefill_backend.py

@nvpohanh

nvpohanh commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

force-merging since this is a very local test change and the test has passed. and this is to fix a wide-spread CI failure.

@nvpohanh
nvpohanh merged commit 41bd213 into sgl-project:main Oct 2, 2026
102 of 114 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant