Skip to content

[Fix] Vacuous marker writes in the cache tests, and an undebited Mamba admission slot - #36415

Merged
ispobock merged 23 commits into
sgl-project:mainfrom
alphabetc1:fix/mamba-gap-reserve-double-read
Sep 4, 2026
Merged

ispobock merged 23 commits into
sgl-project:mainfrom
alphabetc1:fix/mamba-gap-reserve-double-read

Conversation

@alphabetc1

@alphabetc1 alphabetc1 commented Aug 26, 2026 •

Copy link
Copy Markdown
Collaborator

Two independent pre-existing defects, both found while adding Mamba support to
HiCache buffer-only mode (#36345, stacked on this). They are unrelated to each
other; the two commits are separately reviewable.


1. fix: unified radix cache tests write markers into the pool, not a copy

Two helpers in the unified radix cache unit suite write their marker into a
copy of the pool, not the pool:

k_buf[indices].fill_(marker)          # advanced indexing -> new tensor
mamba_cache.temporal[:, mamba_indices].fill_(marker)

buf[indices] with a tensor index is advanced indexing, which returns a fresh
tensor; .fill_() mutates that temporary and the pool keeps its zeros.

Every downstream assertion of the form "the loaded bytes equal the producer's"
is therefore comparing zeros to zeros — test_hicache_load_back_restores_data,
test_hicache_l3_prefetch_roundtrip, test_buffer_only_read_path_roundtrip
and the SWA data checks all pass unconditionally. A load-back that restored the
wrong slot, or restored nothing at all, would not turn any of them red.

Reproduced standalone:

>>> t = torch.zeros(3, 8); t[:, torch.tensor([1, 2])].fill_(5); t.sum()
tensor(0.)

Fixed by assigning through __setitem__, which is an in-place index_put_.


2. fix: charge the Mamba slot a prefill admission reserved but never debited

PrefillAdder.add_one_req reads _mamba_gap_budget_for_req(req) twice, and the
two reads straddle a call that changes its answer.

total_tokens += self._mamba_gap_budget_for_req(req)     # gate
...
if req.needs_host_load_back():
    self.tree_cache.init_load_back(...)                 # binds req.mamba_pool_idx
...
self._update_prefill_budget(..., mamba_gap_reserve=self._mamba_gap_budget_for_req(req))

_mamba_gap_budget_for_req returns 0 once req.mamba_pool_idx is not None, and
MambaComponent.prepare_load_back sets exactly that field during
init_load_back. So on any request that takes a host load-back, the gate
reserves a slot's worth of shared-gap bytes and the debit that follows charges 0
— rem_total_tokens, cur_rem_tokens and rem_mamba_slots all stay
un-decremented for a slot the request is now holding.

The next request in the same batch is then admitted against a budget that
believes that slot is still free. On the shared Mamba pool the residual is what
alloc_req_slots' fail-loud RuntimeError backstops.

Fixed by reading it once, before init_load_back, and reusing the value at both
debit sites. No behaviour change for requests that already own a slot, or off
the shared Mamba pool where _mamba_slot_cost is 0.


Verification

(1) H200, -k 'buffer or hicache' over the whole unified-radix-cache config
matrix: 418 passed, 403 skipped, 0 failed — identical to main, so the
now-live byte comparisons hold as they stand. The point is that they can fail
from here on.

(2) is inert unless the allocator is the shared
UnifiedMambaTokenToKVPoolAllocator composite, which is only built under
--enable-unified-memory (kv_cache_configurator.py:390). Everywhere else
_mamba_slot_cost is 0, so both reads return 0 and the value passed to
_update_prefill_budget is unchanged. The runs above — and the Inkling classes
in test_unified_radix_cache_kl_hybrid_bitexact.py, which do not set that flag
— therefore show the change is harmless, not that it is exercised. I have
not run a unified-memory hybrid-Mamba configuration; the argument for the change
is the read-after-mutation above, which is visible in the source. Happy to add a
targeted PrefillAdder unit case if reviewers want a guard.


CI States

Latest PR Test (Base): ✅ Run #33832677017
Latest PR Test (Extra): ❌ Run #33832676907
Latest PR Test (AMD ROCm 7.2): ⏳ Run #33832677009

alphabetc1 and others added 2 commits August 26, 2026 09:50
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ited

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c12432c8b9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread python/sglang/srt/managers/schedule_policy.py
@alphabetc1 alphabetc1 added the run-ci CI: run the baseline test suite on this PR label Aug 26, 2026
@alphabetc1 alphabetc1 changed the title [Scheduler] Charge the Mamba slot a prefill admission reserved but never debited [Fix] Vacuous marker writes in the cache tests, and an undebited Mamba admission slot Aug 26, 2026
@alphabetc1

Copy link
Copy Markdown
Collaborator Author

/rerun-group radix_cache/unified_radix_tree

@github-actions

github-actions Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-group radix_cache/unified_radix_tree:

🚀 4-gpu-h100 (4 tests): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_hicache_pp_kl.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_cp.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_dsv4.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_mamba.py

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_dcp.py

🚀 8-gpu-h200 (3 tests): ❌ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_dsv4_pp.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_mimo.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_nightly.py

🚀 2-gpu-h100 (2 tests): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_full.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_swa.py

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_hybrid_bitexact.py

@alphabetc1

Copy link
Copy Markdown
Collaborator Author

/rerun-group radix_cache/unified_radix_tree

@github-actions

github-actions Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-group radix_cache/unified_radix_tree:

🚀 4-gpu-h100 (4 tests): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_hicache_pp_kl.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_cp.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_dsv4.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_mamba.py

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_dcp.py

🚀 8-gpu-h200 (3 tests): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_dsv4_pp.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_glm52.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_mimo.py

🚀 2-gpu-h100 (2 tests): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_full.py
cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_swa.py

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_hybrid_bitexact.py

@alphabetc1
alphabetc1 enabled auto-merge (squash) September 1, 2026 17:13
@alphabetc1
alphabetc1 disabled auto-merge September 4, 2026 11:09
@ispobock
ispobock merged commit e4adf63 into sgl-project:main Sep 4, 2026
159 of 187 checks passed
@alphabetc1
alphabetc1 deleted the fix/mamba-gap-reserve-double-read branch September 4, 2026 11:53
StevenChenSE pushed a commit to StevenChenSE/sglang that referenced this pull request Sep 6, 2026
…a admission slot (sgl-project#36415)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 26, 2026
…cks, chronologisch)

Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum
(Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst
anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity;
479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 ->
S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name
(Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)).
Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc

Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859
sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579
sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438
(SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool
sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild),
#DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release;
P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6
+ F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49
Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens).

Auflösungen (Gabel -> Form, Grund):
- L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität"
  nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker-
  Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags.
  Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity.
- RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser
  CalibrationIdentity.d_max_running_requests; nextflash unverändert.
- URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und
  Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie).
- scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP
  komponiert (beide Instrumente getrennt schaltbar).
- schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF
  computed_input_len.
- gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab.
- flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt).
- checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation.
- xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen.
- vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read.
- FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9).
- argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only).
- census_checkpoint_decision (W161) jetzt für beide Profile aktiv.

Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber
dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu.
Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fastfail run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants