perf(scheduler): coalesce prefills on TP - #2238
Conversation
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
c084d85 to
41da984
Compare
e768a0e to
d75dad4
Compare
d75dad4 to
4346669
Compare
|
Read this against the tree at head in an isolated worktree. The goal — let a The anchor is a state-cache position, not a prefix-cache boundary
seq.checkpoint_end_pos = 0
if not self.state.enabled:
returnand returns again if Where the state cache is enabled, the anchor is still not the position the waiter I would resolve this before weighing anything below, because it decides whether the The skip is in the wrong place in the Phase-2 loopFour separate consequences follow from where line 1531 sits, so they are worth taking It runs after It is upstream of both The deferred seq goes to And because the skip consumes neither a seq slot nor token budget before The fill signal was capped at four entries
The bound is also spent by entries that contribute nothing: the aborted / Two smaller things in the same function: What it costs per tick
The prefix compare is np.array_equal(
np.frombuffer(seq.token_ids, dtype=np.int32, count=anchor),
np.frombuffer(producer.token_ids, dtype=np.int32, count=anchor),
)which evaluates "Local mode" has three definitions
Relatedly, in that branch That decode-interval HOLD also returns before every must-fire bound the module docstring Docs and tests
There are no tests for any of it. What I would want to seeThe anchor question first, since a corrected anchor may change what the rest should look Happy to look at a revision. |
4346669 to
5c220c6
Compare
Review — schedule optimization for agentic workloadsReviewed The two ideas here are good ones — hold a TP-local prefill batch until it is worth firing, and let a consumer wait for the producer that is about to hand it a checkpoint. But both are implemented by predicting what Phase 2 will do rather than by asking it, and that has a single clean statement:
The wait has the mirror-image problem: its bound is a property of the slot, not of the request. Provenance: [verified] means the deciding code was read in the PR-head worktree or the behavior was measured. [reported] means the shape matches the code but was not traced end to end. Line numbers are PR-head. 1. The estimator and Phase 2 disagree in five places [verified]The estimator's only job is to predict how much Phase 2 will admit. Where it diverges:
Each row is independently sufficient to produce The The offload row is a plain arithmetic overstatement: Nothing asserts the two agree. If the estimator is to stay, the honest shape is one walk that both the signal and the admission read, or at minimum a debug-mode assertion that the predicted admission matches the realized one. 2.
|
Update the local coalescing flag when installing or removing a delayer, and compare shared token prefixes through NumPy buffer views to avoid copying both arrays. Keep the existing cross-DP test stub explicit about its topology.
Keep per-request deadlines through retries and allocation rollbacks while allowing bounded lookahead for independent requests. Wait for useful in-flight checkpoints beyond smaller offload hits, without letting prior queue age disable prefix reuse. Clarify coalescing and checkpoint wait limits.
Three abort tests bypass Scheduler.__init__ and missed its new deadline map. Initialize the map so all six parameterized offload cancellation cases exercise resource cleanup with a complete scheduler fixture.
371071f to
66e6bd0
Compare
A state load's lifecycle held three facts in three owners across two
processes: "this request owes a report on hash H" (StateOffloadIndex,
engine), "its state slot stays off the free list"
(BlockManager._orphan_load_slots, engine) and "have both legs landed"
(_JointPark, worker). No object could state
dispatched == settled + outstanding
so no test could assert it.
The state leg now rides the request as `LMCacheReqMeta.state_load_spec`
and runs inside the KV leg's own task, so one dispatch emits exactly one
completion on every path, including a raise. Dense's `_do_load_req` is
split into `_load_kv_bytes` + `_finish_load` to give it that seam. A
state-only load travels the ordinary load path on a no-op KV spec
(hbm == lmc). `StateOffloadIndex` becomes the sole engine-side owner: it
takes the destination slot, absorbs orphan parking, and audits its own
invariant (surfaced as `state_offload_invariant_violations`).
Deleted: `_JointPark`, `metadata.state_loads`, the state-load disposition
channel, the state tier's load executor and staging-lane semaphore, the
engine-side load transport on BlockManager, `_publish_state_loads` /
`_settle_state_load` / `_abandon_state_load`, and the state-only park
branch. PP > 1 now refuses at startup instead of warning.
Rebased onto main (ba51495). The dense save producer fence this branch
carried is dropped in favour of main's #2339, which orders the pack stream
on the device rather than host-synchronizing; `producer_event` on
LMCacheReqMeta and its tests go with it, and staging.py's fence docs now
point at #2339's path. Conflicts with #2238 and #2305 resolved keeping both
sides.
Squashed from 20 commits; the pre-rebase history is at ea3f48d.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A state load's lifecycle held three facts in three owners across two
processes: "this request owes a report on hash H" (StateOffloadIndex,
engine), "its state slot stays off the free list"
(BlockManager._orphan_load_slots, engine) and "have both legs landed"
(_JointPark, worker). No object could state
dispatched == settled + outstanding
so no test could assert it.
The state leg now rides the request as `LMCacheReqMeta.state_load_spec`
and runs inside the KV leg's own task, so one dispatch emits exactly one
completion on every path, including a raise. Dense's `_do_load_req` is
split into `_load_kv_bytes` + `_finish_load` to give it that seam. A
state-only load travels the ordinary load path on a no-op KV spec
(hbm == lmc). `StateOffloadIndex` becomes the sole engine-side owner: it
takes the destination slot, absorbs orphan parking, and audits its own
invariant (surfaced as `state_offload_invariant_violations`).
Deleted: `_JointPark`, `metadata.state_loads`, the state-load disposition
channel, the state tier's load executor and staging-lane semaphore, the
engine-side load transport on BlockManager, `_publish_state_loads` /
`_settle_state_load` / `_abandon_state_load`, and the state-only park
branch. PP > 1 now refuses at startup instead of warning.
Rebased onto main (ba51495). The dense save producer fence this branch
carried is dropped in favour of main's #2339, which orders the pack stream
on the device rather than host-synchronizing; `producer_event` on
LMCacheReqMeta and its tests go with it, and staging.py's fence docs now
point at #2339's path. Conflicts with #2238 and #2305 resolved keeping both
sides.
Squashed from 20 commits; the pre-rebase history is at ea3f48d.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Small agentic prefills interrupt decode.
Enable existing PrefillDelayer on TP/DCP (DP=1, PP=1) via nonzero
ATOM_PREFILL_DECODE_INTERVAL. Add bounded HBM probes for batching, skip probing during decode protection, and wait for reusable in-flight prompt-end checkpoints. Interval 0 preserves default TP scheduling.Kimi-K3 AgentX C64, TP8/DCP8; identical clients, full warmup + 3600s. Candidates use
ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000and--state-checkpoint-interval-tokens 32768.Percentages are changes from baseline; TTFT increases are regressions. Recommend interval 4 for smaller TTFT cost. C64 combines code/configuration changes; no ablation was run for that comparison.
Kimi-K3 AgentX C16
Percentages are changes from the C16 interval 0 control. Interval 4 lowers ITL while increasing total throughput, at a TTFT P50/P90 cost of 583/312 ms. Both runs have zero measurement errors and two end-of-run cancellations. All 1,462 common requests have identical output lengths; the table uses the full official AIPerf results. Total throughput is AIPerf
total_token_throughput.avg / 8.C16 scheduling diagnostics over the 3600s window: prefill batches decrease from 2,854 to 2,508 (-12.12%) while actual prefill tokens increase by 7.09%; the share of batches containing multiple requests rises from 10.62% to 31.06%.
C16's larger percentage ITL gain partly reflects its lower starting ITL: P90 drops by 7.01 ms (-22.18%), versus 7.22 ms (-6.34%) for C64 interval 4. The recipes also differ: C16 uses DSpark while C64 has no speculation, and the interval counts scheduling passes rather than emitted tokens. These results do not isolate concurrency as the cause of the different gains.
Validation: 355 passed, 1 skipped; GPU shared-prefix checks passed; full GSM8K C64, 5-shot: 1257/1319 and 1256/1319. Black passes; Ruff has existing findings, none added. Full-suite collection is blocked by missing SGLang/vLLM dependencies.