Repository navigation
feat(grpc): add TokenSpeed KV cache event support (SubscribeKvEvents bridge) - #1771
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughExtracts engine-neutral KV-event ZMQ→proto conversion helpers ( ChangesTokenSpeed KV-event bridge
Sequence Diagram(s)sequenceDiagram
participant Client as gRPC Client
participant Servicer as TokenSpeedSchedulerServicer
participant Config as resolve_kv_events_config
participant ZMQ as ZMQ SUB Socket
participant Shared as stream_kv_events / convert_batch
Client->>Servicer: SubscribeKvEvents(request)
Servicer->>Config: resolve_kv_events_config(server_args)
Config-->>Servicer: ResolvedKvEventsConfig (endpoint, topic) or None
alt config is None
Servicer-->>Client: UNIMPLEMENTED abort
else config present
Servicer->>ZMQ: connect(endpoint_for_rank(endpoint, rank=0))
Servicer->>ZMQ: subscribe(topic)
Servicer->>Shared: stream_kv_events(sub_socket, msgspec.decode, ...)
Shared-->>Client: send_initial_metadata()
loop until cancelled
ZMQ-->>Shared: multipart [seq_bytes, msgpack payload]
Shared->>Shared: convert_batch(decoded_batch)
Shared-->>Client: yield KvEventBatch proto
end
Servicer->>ZMQ: close()
end
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces support for TokenSpeed KV-event cache-aware routing by implementing the SubscribeKvEvents gRPC stream. It refactors the existing engine-neutral ZMQ-to-proto conversion and streaming logic into a shared kv_events.py module, which is now utilized by both vLLM and TokenSpeed bridges. Additionally, it adds config resolution, optional package dependencies, documentation, and unit/integration tests. Feedback on the implementation highlights the need to lazy-import KVEventBatch inside the streaming function rather than at the top level to prevent import errors when TokenSpeed is not installed, and to wrap the socket connection and subscription setup inside the try block to ensure proper resource cleanup in the finally block.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| config = self._kv_events_config | ||
|
|
||
| # DP attention publishes one PUB socket per rank (port + rank) with | ||
| # independent sequence counters; subscribing to several on one socket | ||
| # interleaves them and breaks gap detection. Subscribe to rank 0 only. | ||
| pub_endpoint = endpoint_for_rank(config.endpoint, 0) | ||
|
|
||
| zmq_ctx = zmq.asyncio.Context.instance() | ||
| sub_socket = zmq_ctx.socket(zmq.SUB) | ||
| sub_socket.subscribe(config.topic.encode("utf-8")) | ||
| sub_socket.connect(pub_endpoint) | ||
| logger.info("SubscribeKvEvents: connected to ZMQ endpoint %s", pub_endpoint) | ||
|
|
||
| decoder = msgspec.msgpack.Decoder(KVEventBatch) | ||
|
|
||
| try: | ||
| async for proto_batch in stream_kv_events( | ||
| sub_socket, | ||
| decoder.decode, | ||
| lambda: context.send_initial_metadata(()), | ||
| context.cancelled, | ||
| ): | ||
| yield proto_batch | ||
| except asyncio.CancelledError: | ||
| pass | ||
| except Exception as e: # noqa: BLE001 | ||
| logger.exception("SubscribeKvEvents failed") | ||
| await context.abort(grpc.StatusCode.INTERNAL, str(e)) | ||
| finally: | ||
| sub_socket.close(linger=0) | ||
| logger.info("SubscribeKvEvents: stream closed") |
There was a problem hiding this comment.
Wrap the socket subscription, connection, and decoder initialization inside the try block. If any of these steps fail (e.g., connection error or missing tokenspeed package), the socket will still be properly closed in the finally block, preventing resource leaks. Additionally, lazy-import KVEventBatch here to keep the module importable without tokenspeed installed.
config = self._kv_events_config
# DP attention publishes one PUB socket per rank (port + rank) with
# independent sequence counters; subscribing to several on one socket
# interleaves them and breaks gap detection. Subscribe to rank 0 only.
pub_endpoint = endpoint_for_rank(config.endpoint, 0)
zmq_ctx = zmq.asyncio.Context.instance()
sub_socket = zmq_ctx.socket(zmq.SUB)
try:
sub_socket.subscribe(config.topic.encode("utf-8"))
sub_socket.connect(pub_endpoint)
logger.info("SubscribeKvEvents: connected to ZMQ endpoint %s", pub_endpoint)
from tokenspeed.runtime.pd.kv_events import KVEventBatch
decoder = msgspec.msgpack.Decoder(KVEventBatch)
async for proto_batch in stream_kv_events(
sub_socket,
decoder.decode,
lambda: context.send_initial_metadata(()),
context.cancelled,
):
yield proto_batch
except asyncio.CancelledError:
pass
except Exception as e: # noqa: BLE001
logger.exception("SubscribeKvEvents failed")
await context.abort(grpc.StatusCode.INTERNAL, str(e))
finally:
sub_socket.close(linger=0)
logger.info("SubscribeKvEvents: stream closed")References
- To support optional dependencies, import modules from the extra inside the functions that use them (lazy import) rather than at the top level to prevent ImportError.
- When managing resources that are acquired and need to be released, ensure that all acquired resources are tracked immediately and consistently, especially when exceptions might occur during subsequent operations, to prevent resource leaks.
There was a problem hiding this comment.
Applied the try-wrapping in 23844a9 — subscribe/connect and the Decoder(...) now run inside the try, so the finally always closes the socket (e.g. on a malformed endpoint). Kept KVEventBatch as a top-level import per the project's import convention (the module already requires tokenspeed).
| if self._kv_events_config is None: | ||
| await context.abort( | ||
| grpc.StatusCode.UNIMPLEMENTED, | ||
| "KV cache events not enabled. Start TokenSpeed with " | ||
| "--kv-events-config " | ||
| '\'{"enable_kv_cache_events": true, "publisher": "zmq"}\'', | ||
| ) |
There was a problem hiding this comment.
🟡 Nit: Missing defensive return after context.abort(). Every other early-abort in this file (e.g., lines 146–147, 150–151, 489) follows the abort → return pattern. While context.abort() raises AbortError so the return is technically unreachable, it serves as defense-in-depth and keeps the pattern consistent — without it, if abort ever doesn't raise, config would be None and config.endpoint on line 582 would AttributeError.
| if self._kv_events_config is None: | |
| await context.abort( | |
| grpc.StatusCode.UNIMPLEMENTED, | |
| "KV cache events not enabled. Start TokenSpeed with " | |
| "--kv-events-config " | |
| '\'{"enable_kv_cache_events": true, "publisher": "zmq"}\'', | |
| ) | |
| await context.abort( | |
| grpc.StatusCode.UNIMPLEMENTED, | |
| "KV cache events not enabled. Start TokenSpeed with " | |
| "--kv-events-config " | |
| '\'{"enable_kv_cache_events": true, "publisher": "zmq"}\'', | |
| ) | |
| return |
There was a problem hiding this comment.
Done in 23844a9 — added return after the UNIMPLEMENTED abort to match the abort → return pattern used elsewhere in this file.
| except asyncio.CancelledError: | ||
| pass | ||
| except Exception as e: # noqa: BLE001 | ||
| logger.exception("SubscribeKvEvents failed") | ||
| await context.abort(grpc.StatusCode.INTERNAL, str(e)) |
There was a problem hiding this comment.
🟡 Nit: Missing except grpc.aio.AbortError: raise guard — the Generate method (line 222) has one so that an AbortError propagated from inside the try block (e.g., from send_initial_metadata on an already-aborted RPC) doesn't get caught here, logged as "SubscribeKvEvents failed", and re-wrapped as INTERNAL.
| except asyncio.CancelledError: | |
| pass | |
| except Exception as e: # noqa: BLE001 | |
| logger.exception("SubscribeKvEvents failed") | |
| await context.abort(grpc.StatusCode.INTERNAL, str(e)) | |
| except asyncio.CancelledError: | |
| pass | |
| except grpc.aio.AbortError: | |
| raise | |
| except Exception as e: # noqa: BLE001 | |
| logger.exception("SubscribeKvEvents failed") | |
| await context.abort(grpc.StatusCode.INTERNAL, str(e)) |
There was a problem hiding this comment.
Done in 23844a9 — added except grpc.aio.AbortError: raise so an abort propagated from inside the try isn't caught, logged as "SubscribeKvEvents failed", and re-wrapped as INTERNAL. Matches Generate.
There was a problem hiding this comment.
Clean, well-structured PR. The refactoring to promote engine-neutral ZMQ→proto helpers to a shared module is done carefully — vLLM re-exports preserve backwards compatibility, and the TokenSpeed bridge mirrors the vLLM one closely. Config resolution, conversion, and streaming are all well-tested without requiring an engine install.
Two nits on exception-handling consistency in SubscribeKvEvents (defensive return after abort + AbortError re-raise guard), both matching patterns already established in the same file's Generate method.
0 🔴 Important · 2 🟡 Nit · 0 🟣 Pre-existing
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
grpc_servicer/pyproject.toml (1)
32-41:⚠️ Potential issue | 🟠 MajorUpdate to current versions of
pyzmqandmsgspecin thetokenspeedoptional-dependency group.The floor constraints
pyzmq>=25.0.0andmsgspec>=0.18.0are significantly outdated. As of June 2026, the latest stable versions arepyzmq==27.1.0andmsgspec==0.21.1respectively. The declared versions are from 2023 (approximately 2–3 years old) and, while no critical security vulnerabilities are documented for these specific versions, they lack numerous bug fixes, performance improvements, and refinements released since.Update the floor constraints to more recent versions—ideally
pyzmq>=27.0.0andmsgspec>=0.21.0, or at minimum closer to current releases.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@grpc_servicer/pyproject.toml` around lines 32 - 41, The tokenspeed optional-dependency group in pyproject.toml has outdated floor constraints for its dependencies. Update the pyzmq constraint from pyzmq>=25.0.0 to pyzmq>=27.0.0 and the msgspec constraint from msgspec>=0.18.0 to msgspec>=0.21.0 to align with more recent stable versions that include bug fixes and performance improvements instead of versions from 2–3 years ago.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@grpc_servicer/tests/test_tokenspeed_kv_events_stream.py`:
- Around line 72-106: The test uses fixed timing delays (asyncio.sleep calls at
line 72 and within the batches loop at line 105) which cause nondeterministic
failures due to ZMQ slow-joiner behavior dropping early frames. Replace the
initial 0.2 second sleep before publishing with a handshake mechanism that
confirms the subscriber in the consume() function is actively listening (for
example, by having the consumer signal readiness or by sending and awaiting an
initial acknowledgment message). Similarly, replace the 0.05 second sleep
between batch sends with deterministic synchronization, such as waiting for each
batch to be collected before sending the next one, to ensure the consumer
processes each message before the publisher sends the next batch.
---
Outside diff comments:
In `@grpc_servicer/pyproject.toml`:
- Around line 32-41: The tokenspeed optional-dependency group in pyproject.toml
has outdated floor constraints for its dependencies. Update the pyzmq constraint
from pyzmq>=25.0.0 to pyzmq>=27.0.0 and the msgspec constraint from
msgspec>=0.18.0 to msgspec>=0.21.0 to align with more recent stable versions
that include bug fixes and performance improvements instead of versions from 2–3
years ago.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 007ac31d-5c69-4fdd-b396-79954e8935fb
📒 Files selected for processing (9)
docs/getting-started/kv-events-cache-aware.mddocs/proposals/2026-06-16-tokenspeed-kv-events.mdgrpc_servicer/pyproject.tomlgrpc_servicer/smg_grpc_servicer/kv_events.pygrpc_servicer/smg_grpc_servicer/tokenspeed/kv_events.pygrpc_servicer/smg_grpc_servicer/tokenspeed/servicer.pygrpc_servicer/smg_grpc_servicer/vllm/kv_events.pygrpc_servicer/tests/test_tokenspeed_kv_events.pygrpc_servicer/tests/test_tokenspeed_kv_events_stream.py
f4c0b1b to
64310da
Compare
There was a problem hiding this comment.
Actionable comments posted: 5
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@grpc_servicer/smg_grpc_servicer/kv_events.py`:
- Around line 154-155: In the frame validation block where len(frames) < 3 is
checked, add a debug-level log statement before the continue statement to
capture information about the malformed frame. Include relevant context such as
the frame length and any identifying details that would help diagnose protocol
mismatches during debugging, without impacting production performance.
In `@grpc_servicer/smg_grpc_servicer/tokenspeed/kv_events.py`:
- Around line 68-71: After assigning the default values for endpoint and topic
in the configuration resolution logic, add type validation to ensure both
endpoint and topic are strings. If either endpoint or topic is not a string
type, log a warning message indicating that KV-event streaming is being disabled
due to invalid configuration types, and return None from the function to cleanly
disable this feature instead of allowing invalid types to propagate to
SubscribeKvEvents where they would cause runtime errors later.
In `@grpc_servicer/smg_grpc_servicer/tokenspeed/servicer.py`:
- Around line 584-607: The socket setup code including zmq context
instantiation, socket creation, subscription, and connection to the endpoint are
all executed outside the try/finally block in the SubscribeKvEvents method. This
means if any of those operations fail, the exception handling logic and socket
cleanup will not be executed. Move all socket setup code (zmq_ctx instantiation,
sub_socket socket creation, subscribe call, and connect call) into the try block
before the stream_kv_events call, and add a guard in the finally block to check
if sub_socket was successfully created before attempting to close it.
- Around line 24-29: Move the imports of msgspec and zmq-related modules (zmq,
zmq.asyncio) from module scope into the SubscribeKvEvents method where they are
actually used. Wrap these imports in a try-except block within SubscribeKvEvents
to catch ImportError when optional dependencies are not installed, and use gRPC
abort with UNIMPLEMENTED status when the import fails. Additionally, move the
socket setup code (currently at lines 584-587) from before the try block into
the try-except block so socket creation failures are also properly handled and
reported through gRPC error handling instead of causing an unhandled exception
at import time.
In `@grpc_servicer/tests/test_tokenspeed_kv_events.py`:
- Around line 57-59: The test test_none_when_publisher_null is passing the
string "null" instead of an actual None value for the publisher parameter, which
tests the wrong code path. Change the publisher parameter in the _cfg function
call from the string "null" to None (or None value as appropriate for your
language) to properly validate the JSON null contract behavior where publisher
should default to "zmq" when enabled.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: e330510c-db99-496c-9f5e-876d7941a406
📒 Files selected for processing (8)
docs/getting-started/kv-events-cache-aware.mdgrpc_servicer/pyproject.tomlgrpc_servicer/smg_grpc_servicer/kv_events.pygrpc_servicer/smg_grpc_servicer/tokenspeed/kv_events.pygrpc_servicer/smg_grpc_servicer/tokenspeed/servicer.pygrpc_servicer/smg_grpc_servicer/vllm/kv_events.pygrpc_servicer/tests/test_tokenspeed_kv_events.pygrpc_servicer/tests/test_tokenspeed_kv_events_stream.py
…bridge) Implements TokenSpeedSchedulerServicer.SubscribeKvEvents so SMG's cache_aware router can route against TokenSpeed workers' actual KV-cache state instead of the approximate token tree. The entire Rust/SMG consumer side (proto RPC, gRPC client dispatch, KvEventMonitor, PositionalIndexer, cache_aware) and the TokenSpeed ZMQ publisher (--kv-events-config) already existed; only the Python servicer bridge was missing — mirroring the vLLM bridge in #1652. - Promote the engine-neutral ZMQ->proto conversion (to_int64, endpoint_for_rank, convert_event, convert_batch, stream_kv_events) from vllm/kv_events.py to a shared smg_grpc_servicer/kv_events.py; vllm/kv_events.py re-exports them and keeps its vLLM-specific resolver. convert_batch now reads the DP rank from data_parallel_rank (vLLM) or attn_dp_rank (TokenSpeed). - Add tokenspeed/kv_events.py resolver: parses server_args.kv_events_config (JSON string) and returns the ZMQ endpoint iff enabled with publisher=zmq. - Implement SubscribeKvEvents in the TokenSpeed servicer (rank-0 subscription, no replay), matching the vLLM bridge. - Add the [tokenspeed] pyproject extra (pyzmq, msgspec). - Tests (engine-free, no TokenSpeed install): resolver + shared conversion unit tests and an in-process ZMQ PUB/SUB integration test using msgspec structs that mirror the TokenSpeed wire layout. - Docs: TokenSpeed worker launch section in kv-events-cache-aware.md. No Rust, proto, or tokenspeed-lib changes (the pinned tokenspeed already ships the KV-event publisher). Signed-off-by: key4ng <rukeyang@gmail.com>
…repo root test_vllm_kv_events* load vllm/kv_events.py by file path, which now re-exports from the shared smg_grpc_servicer.kv_events module. When pytest runs from the repo root (CI), grpc_servicer/ is not on sys.path so that package import fails at collection. Add a tests/conftest.py that prepends this repo's grpc_servicer/ to sys.path (taking precedence over any stale installed copy). Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
- defensive return after the UNIMPLEMENTED abort (matches Generate/others) - re-raise grpc.aio.AbortError instead of re-wrapping it as INTERNAL - move socket subscribe/connect/decoder into the try so the finally always closes the socket (e.g. on a malformed endpoint) - resolver: reject non-string endpoint/topic and disable cleanly instead of failing later as an opaque INTERNAL - tests: cover explicit JSON-null publisher (-> zmq) and non-string config Signed-off-by: key4ng <rukeyang@gmail.com>
5acf68c to
23844a9
Compare
|
Rebased onto latest Applied
Skipped (with reason)
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 23844a9642
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Signed-off-by: key4ng <rukeyang@gmail.com>
Gateway-sweep benchmark — 1 / 2 / 4 gatewaysReplicated the PR-1652 methodology for TokenSpeed: 1 gateway
2 gateways
4 gateways
What it shows
Test setup & methodologyHardware: 8×H100 80GB, CUDA 13. Workers (3, TP=2, GPUs 0–5): CUDA_VISIBLE_DEVICES=0,1 python -m smg_grpc_servicer.tokenspeed \
--model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B --tensor-parallel-size 2 \
--host 127.0.0.1 --port 31001 --disable-kvstore \
--kv-events-config '{"enable_kv_cache_events": true, "publisher": "zmq", "endpoint": "tcp://*:6001", "topic": "kv-events"}'
# (workers 2/3 on GPUs 2,3 / 4,5, ports 31002/31003, zmq 6002/6003)
# events-off runs (round_robin, approx) launch the same workers without --kv-events-configGateways (N ∈ {1,2,4}): RUST_LOG=warn ./target/release/smg launch \
--worker-urls grpc://127.0.0.1:31001 grpc://127.0.0.1:31002 grpc://127.0.0.1:31003 \
--model-path deepseek-ai/DeepSeek-R1-Distill-Qwen-32B --policy {round_robin|cache_aware} \
--block-size 64 --host 127.0.0.1 --port 30000 --prometheus-port 29001
# N gateways on ports 30000/30010/30020/30030, fronted by nginx round-robin
# (proxy_buffering off, upstream keepalive, worker_rlimit_nofile raised to avoid FD exhaustion)Client: vllm bench serve --backend openai --model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B \
--base-url http://127.0.0.1:<nginx-or-gateway-port> --endpoint /v1/completions \
--dataset-name prefix_repetition \
--prefix-repetition-num-prefixes 600 --prefix-repetition-prefix-len 1000 \
--prefix-repetition-suffix-len 50 --prefix-repetition-output-len 128 \
--num-prompts 1800 --request-rate 32 --burstiness 1.0 --ignore-eos \
--percentile-metrics ttft,e2el --metric-percentiles 95,99Strategy:
|
Description
Problem
KV-event-driven cache-aware routing — where the gateway routes each request to the worker whose KV cache already holds the longest prefix, using the worker's actual cache state — worked for SGLang and vLLM but not for TokenSpeed. TokenSpeed gRPC workers silently fell back to the approximate token tree because
TokenSpeedSchedulerServicernever implementedSubscribeKvEvents(the RPC returnedUNIMPLEMENTED, so SMG'sKvEventMonitordisabled the subscription for that worker).The entire Rust/SMG side is already engine-agnostic — the proto RPC (
common.protoSubscribeKvEvents/KvEventBatch), the TokenSpeed gRPC client (impl_subscribe_kv_events!),GrpcClient::TokenSpeeddispatch,KvEventMonitor, thePositionalIndexer, and thecache_awarepolicy. The only missing piece was the Python servicer bridge.Solution
Implement
SubscribeKvEventsonTokenSpeedSchedulerServicer, bridging TokenSpeed's in-process ZMQ KV-cache events (msgpackBlockStored/BlockRemoved/AllBlocksCleared) to the engine-neutralcommon.protoKvEventBatchstream SMG already consumes. No Rust, proto, or TokenSpeed-engine changes.The ZMQ→proto conversion is engine-neutral, so it is promoted to a shared module reused by both the vLLM and TokenSpeed bridges; each engine keeps only its own config resolver.
Changes
grpc_servicer/smg_grpc_servicer/kv_events.py(new): engine-neutral ZMQ→proto helpers (to_int64,endpoint_for_rank,convert_event,convert_batch,stream_kv_events), promoted fromvllm/kv_events.py.convert_batchreads the DP rank fromdata_parallel_rank(vLLM) orattn_dp_rank(TokenSpeed).grpc_servicer/smg_grpc_servicer/vllm/kv_events.py: re-exports the shared helpers; keeps the vLLM-specific resolver.grpc_servicer/smg_grpc_servicer/tokenspeed/kv_events.py(new): TokenSpeed config resolver — parsesserver_args.kv_events_configand returns the ZMQ endpoint iff events are enabled withpublisher=zmq.grpc_servicer/smg_grpc_servicer/tokenspeed/servicer.py: implementSubscribeKvEvents(rank-0 subscription, no replay), mirroring the vLLM bridge.grpc_servicer/pyproject.toml: add the[tokenspeed]extra (pyzmq,msgspec).docs/getting-started/kv-events-cache-aware.md.Test Plan
cargo +nightly fmt,ruff check/ruff format, andpytest grpc_servicer/tests/pass for the new/changed files. Existingtest_vllm_kv_events*.pyremain green (the shared-module refactor is import-compatible).Manual E2E (H100)
Built SMG + the servicer from this branch, ran live TokenSpeed gRPC workers (
DeepSeek-R1-Distill-Qwen-32B) under SMG--policy cache_aware:Starting KV event subscription→KV event stream connected start_seq=0per worker (RPC streams instead ofUNIMPLEMENTED)Learned block_size from KV event ... block_size=64(block size learned from realBlockStoredevents through the bridge)--kv-events-config→Backend does not implement SubscribeKvEvents, disabling KV event subscription(graceful fallback)Benchmark (event-driven vs approximate token tree)
vllm bench serve --dataset-name prefix_repetition(600 prefixes × 1000-token shared prefix + 50-token suffix, 128-token output, 3000 requests, open-loop Poisson, saturating rate 32), through 4 SMG gateways (nginx round-robin, no mesh) in front of 3 TokenSpeed workers (DeepSeek-R1-Distill-Qwen-32B, TP=2). TokenSpeed workers launched with--disable-kvstoreand--kv-events-config '{"enable_kv_cache_events": true, "publisher": "zmq", ...}'. 8×H100, 3000/3000 successful per run.Event-driven cache-aware routing roughly doubles the prefix-cache hit rate and halves TTFT versus the approximate token tree for TokenSpeed workers. The win is concentrated in prefill/TTFT (expected — prefix-cache reuse cuts prefill work); end-to-end latency converges at this scale since the 128-token decode dominates.
Checklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspasses (no Rust changes)🤖 Generated with Claude Code
Summary by CodeRabbit
Release Notes
New Features
Documentation
Enhancements
Tests