Skip to content

[FEAT] Allow local preemption during speculative decoding - #52

Open
0z5a wants to merge 12 commits into
ThinkFlowLab:mainfrom
0z5a:rlt-spec-sched-only
Open

0z5a wants to merge 12 commits into
ThinkFlowLab:mainfrom
0z5a:rlt-spec-sched-only

Conversation

@0z5a

@0z5a 0z5a commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Keep local priority preemption and resume at committed speculative-round boundaries. Reuse the existing KV snapshot/restore interface and move suspension resource cleanup into the engine. Remove intra-round scheduling and request migration, including their APIs and tests.

Resolve the conflicts against current main, retaining its attention backend and KV admission changes. Add regressions for preemption/resume, per-request sampling RNG, aborting suspended requests, and request-ID reuse.

Add a paired traffic benchmark and profiling recipe covering equal/mixed priority, uniform/mixed short-long requests, and low/high load. Preserve per-request TTFT/E2E, outputs, memory, and preemption counts; profile separately from timed measurements. enable_preemption remains False pending evidence of consistent TTFT benefits with a small E2E penalty.

Validation for this revision:

  • Ruff and git diff --check: passed.
  • Benchmark matrix, CLI refusal paths, and aggregation contracts: passed with stdlib only; no Torch/model/GPU execution.
  • CPU regressions: 118 distinct cases passed; 26 GPU cases deselected. The original child/supervisor exited naturally with 0, and an independent stdlib audit checked the XML and all 132 committed source files. CUDA stayed uninitialized (Torch 2.13/Python 3.12; tiny CPU models only). Full runtime mapping remains incomplete.
  • GPU regressions, real-model traffic measurements, and the requested rlt-perf-opt profiling: not run. The referenced skill/Developer Must-Read location has not been found in the available repository or skills.

@0z5a
0z5a force-pushed the rlt-spec-sched-only branch 3 times, most recently from a471ea3 to 8e345e9 Compare September 24, 2026 07:07
@0z5a
0z5a marked this pull request as ready for review September 24, 2026 09:08
@0z5a
0z5a marked this pull request as draft September 24, 2026 09:09
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Merge the cleaned Graph parent while preserving the scheduler changes and existing commit history. Remove benchmark and result files from the final PR diff.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
@bjf-frz bjf-frz changed the title Support speculative scheduling, preemption, and KV migration [FEAT] Support speculative scheduling, preemption, and KV migration Oct 8, 2026
Comment thread vllm_rlt/engine/preemption.py Outdated
Comment on lines +165 to +171
if not self._restore_state(request, packet.state):
return False
if packet.rng_state is not None:
request.generator = torch.Generator(device=e.cache_manager.device)
request.generator.set_state(packet.rng_state)
e.scheduler.requests[request.request_id] = request
e.scheduler.enqueue(request, packet.state["stage"])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please route migration imports through scheduler admission before restoring KV.
_restore_state() only checks whether the snapshot’s current pages fit, and lines 170–171 then register the request directly as active. This bypasses max_num_seqs and the applicable KV growth budget.

I reproduced this with max_num_seqs=1, incremental KV allocation, preemption disabled, and a 32-block pool. Importing a second request returns True; both requests later exhaust the pool and fail with scheduler made no progress. The same workload completes through normal admission.

Please add a scheduler-owned admission operation for migration. If admission is temporarily blocked, return False without consuming the packet or leaving target-side state. Once admitted, restore the saved execution stage.

Comment thread vllm_rlt/engine/preemption.py Outdated
Comment on lines +79 to +89
def _signature(self):
e, cache = self.engine, self.engine.cache_manager
return (
e.model.config,
cache.layout,
cache.block_size,
cache.dtype,
cache.key_cache.shape[1:],
cache.attention_info,
cache.device.type,
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please validate the destination execution mode before allocating or restoring migration state.
_signature() checks model/cache compatibility but does not check whether the destination supports Stage.SPECULATIVE.

With the same model and cache configuration, a destination created without speculative_config accepts the packet and returns True. The request is then queued as SPECULATIVE, which the normal scheduling policy never handles; the next step() fails with scheduler made no progress.

Please reject this incompatible destination with a clear error before _restore_state(). Supporting migration into ordinary decoding would require an explicit state/stage conversion.

Comment thread vllm_rlt/config.py Outdated

@bjf-frz bjf-frz Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replace K with the number of speculative tokens to make the docstring clear without requiring readers to infer the notation.

@bjf-frz

bjf-frz commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

The current benchmarks do not demonstrate a meaningful benefit from intra-round scheduling. Leave it out of this PR to avoid adding scheduling and state-management complexity without a demonstrated payoff. We can revisit it once benchmarks show a clear throughput or latency improvement.

@bjf-frz

bjf-frz commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Please add benchmarks for equal-priority traffic, mixed-priority arrivals, mixed short and long requests, and low versus high load. If they show consistent TTFT improvements with only a small E2E penalty, enable preemption by default (True). Users prioritizing minimum E2E latency can disable it.

Comment thread vllm_rlt/engine/preemption.py Outdated
allocation = cache._get_allocation(victim.request_id)
allocation = cache._get_allocation(request.request_id)
blocks = [b for table in allocation.block_tables for b in table]
# Copy views one page at a time: a pressure recovery must not allocate

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a comment here explaining that copying KV pages individually to CPU avoids allocating a temporary GPU tensor to gather them before transfer.

Comment thread vllm_rlt/engine/preemption.py Outdated
Comment on lines +116 to +119
e = self.engine
e.model_runner.release(request.request_id)
e.cache_manager.poll_prefixes()
e.cache_manager.free(request.request_id)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Move the request cleanup logic in _release() into an engine method. Releasing runner state, freeing KV cache, removing queue entries, and clearing request tensors are engine responsibilities; PreemptionManager should delegate this cleanup to the engine.

@bjf-frz

bjf-frz commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Please remove the request migration feature and its related tests from this PR. A separate KV migration mechanism is unnecessary; the existing KV transfer infrastructure should be reused. Keep local preemption and resume support.

@bjf-frz

bjf-frz commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

resolve conflict please

@bjf-frz

bjf-frz commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

After making these changes, please run the rlt-perf-opt skill referenced in the Developer Must-Read guide to validate performance and provide profiling results.

Remove intra-round scheduling and request migration, reuse the current
KV snapshot/restore interface, and move suspension cleanup into the engine.
Merge current main to preserve the attention backend and admission changes.

Add speculative preemption lifecycle/RNG regressions and paired traffic
benchmarks with separate profiling. Keep preemption opt-in pending the
requested performance evidence.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a 0z5a changed the title [FEAT] Support speculative scheduling, preemption, and KV migration [FEAT] Allow local preemption during speculative decoding Oct 9, 2026
@0z5a
0z5a marked this pull request as ready for review October 9, 2026 12:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants