Skip to content

[KVConnector][Tiering] Drive the tiering control plane from a dedicated thread - #58113

Closed
liranschour wants to merge 1 commit into
vllm-project:mainfrom
liranschour:fix/55179-tiering-control-thread
Closed

liranschour wants to merge 1 commit into
vllm-project:mainfrom
liranschour:fix/55179-tiering-control-thread

Conversation

@liranschour

@liranschour liranschour commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Addresses #55179. The P2P secondary tier only advances its control plane from the
per-step scheduler hooks: TieringOffloadingManager.on_schedule_end polls each
tier for finished jobs and then lets each one serve inbound peer requests. The
p2p tier's transports are non-blocking by design (ZmqTransport.poll() returns
immediately with no traffic), so nothing moves between steps and a peer's
control-plane round trip waits for a step boundary.

This moves that sweep onto a dedicated thread owned by the tiering manager.

Honest summary of what this does and does not buy

It does not, on its own, produce a measurable latency win. I benchmarked it
loaded (numbers below) and fetch RTT is parity-within-noise. I am submitting it
because the mechanism it removes is real and because the measurement localises
the actual bottleneck, not because it makes the benchmark faster.

The residual wait appears to be on the producer side: the fetch is not
waiting to be answered, it is waiting for the producer's scheduler thread to
reach complete_store → tier.submit_store and park the blocks. This change
deliberately leaves _flush_pending_promotions and the cascade submit_store
calls on the scheduler thread, so that wait is untouched. Getting the win
probably requires moving the producer-side block parking off the step boundary
as well.

Relationship to the existing PRs on this issue (per AGENTS.md)

This overlaps open work, and I want that on the record rather than buried:

In #55179 @nilig asked for a read on "a thread touching the connector from
inside EngineCore, even phase-exclusive, versus a control thread owned by the
tier". This PR is a working implementation of the second option, offered as
input to that question:

#57165 / #55962 this PR
Driver EngineCore poller thread thread owned by TieringOffloadingManager
API surface new KVConnectorBase_V1.poll_pending_work none
Files touched 17 1 source + 3 test
Exclusion lock hands connector between engine phases lock held for a whole scheduler step

If the maintainers prefer either existing PR, close this one. It is up as a
concrete alternative and as a carrier for the measurement, not a land grab.

Design: why the lock is step-scoped

Per-call locking is not sufficient. The connector learns a chunk is a HIT from
lookup() (under get_num_new_matched_tokens) and only reads it in a later
prepare_load() hook (under update_state_after_alloc). In between the chunk
still has ref_cnt == 0, i.e. it is evictable. A promotion started while serving
a peer lookup in that window can evict it via
prepare_write → CPUOffloadingManager.prepare_store, and prepare_load then
trips assert chunk is not None.

So the lock is taken on the first scheduler-side call of a step and released at
the end of on_schedule_end. That places the control thread's window over model
execution. The read-mostly hooks (has_pending_work, get_stats,
take_events) are the exception: they hold the lock only for their own duration,
because the engine ticks them on every iteration including ones that schedule
nothing and run no model, and folding them into the step-wide hold leaves an idle
engine with no gap at all — the state a producer waiting for a consumer sits in.

Tiers re-enter through _SecondaryTierFacingParent, whose caller already holds
the lock, so the four ParentManager methods gained *_unlocked variants.
exclude_tier_idx moved onto those, which restores the public signatures to
match OffloadingManager. A sweep that raises is logged and re-raised on the
next scheduler-side call, matching the engine crash these failures produced when
they ran on the scheduler thread rather than leaving the control plane silently
dead.

No interface changes, and no changes under tiering/p2p/.

Tests

Unit, on this branch (upstream main base):

venv/bin/python -m pytest tests/v1/kv_offload/tiering/ -q
  -> 387 passed, 12 skipped   (includes the 215 untouched p2p tests)

venv/bin/python -m pytest tests/v1/kv_connector/unit/offloading_connector/ -q
  -> 462 passed

Also run with the change cherry-picked onto the #57139 base for the live tests
below: tests/v1/kv_offload/tiering/ -> 387 passed, 11 skipped.

Ten new tests in tests/v1/kv_offload/tiering/test_tiering_offloading.py cover:
sweeping with no further steps, sweeps running off the scheduler thread, the step
locking out the control thread, on_schedule_end not double-serving, a failed
sweep re-raising on the scheduler thread, shutdown() joining, reset_cache()
with the thread live, idle-engine hooks not starving the thread, and the
disabled/no-tier cases.

Two of those were verified by mutation rather than assertion, since they are the
substantive claims: reverting lookup to per-call locking fails
test_step_locks_out_the_control_thread, and reverting _observation_lock to a
step-wide hold fails
test_idle_engine_hooks_do_not_block_the_control_thread.

Existing step-driven suites pass control_plane_thread=False to stay
deterministic; that flag is the only new knob and it is constructor-only.

Notes

  • Model evaluation on accuracy-sensitive benchmarks was not run; the change does
    not alter what is computed, only which thread services the tier control plane.
    The correctness checks above are the serving-level evidence.
  • AI assistance was used to write this change, its tests, and this
    description. Per AGENTS.md it stays a draft until I have reviewed every changed
    line and can defend it end to end.

🤖 Generated with Claude Code

…ed thread

The secondary tiers' control plane only advanced from the per-step
scheduler hooks: on_schedule_end polled every tier for finished jobs and
then let each one serve inbound peer requests. The p2p tier's transports
are non-blocking by design, so nothing moved between steps, and a rank
busy with long prefills made every peer control-plane round trip wait for
a step boundary — several round trips per transfer.

TieringOffloadingManager now runs both from its own daemon thread, and
serializes it against the scheduler with one lock held for a whole
scheduler step rather than per call. Step granularity is required, not
just tidier: the connector learns a chunk is a HIT in lookup() and only
reads it in a later prepare_load() hook, and the chunk is evictable in
between, so a promotion started while serving a peer lookup in that
window could evict it out from under the pending load. The lock is
released at the end of on_schedule_end, which puts the control plane's
window over model execution — the time a rank mid-prefill has to spare.

The read-mostly hooks (has_pending_work, get_stats, take_events) instead
hold the lock only for their own duration. The engine ticks them every
iteration, including iterations that schedule nothing and run no model,
so folding them into the step-wide hold would leave an idle engine with
no gap between steps at all — the state a producer waiting for a consumer
to connect is in, and exactly when the control plane matters.

Tiers re-enter the manager through _SecondaryTierFacingParent, whose
caller already holds the lock, so the four ParentManager methods gained
*_unlocked variants; exclude_tier_idx moved onto those, restoring the
public signatures to match OffloadingManager. A sweep that raises is
logged and re-raised on the next scheduler-side call, matching the engine
crash these failures produced when they ran on the scheduler thread,
instead of leaving the control plane silently dead.

No interface changes, and no changes under tiering/p2p. The thread is on
whenever a secondary tier is configured; control_plane_thread=False keeps
everything on the calling thread, which is what the step-driven tests need.

Signed-off-by: Liran Schour <lirans@il.ibm.com>
@mergify

mergify Bot commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @liranschour.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant