Skip to content

feat(grpc): add subscribe_kv_events to all backend clients - #558

Merged
slin1237 merged 1 commit into
mainfrom
slin/kv-event-grpc-client
Feb 27, 2026
Merged

slin1237 merged 1 commit into
mainfrom
slin/kv-event-grpc-client

Conversation

@slin1237

@slin1237 slin1237 commented Feb 27, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Add impl_subscribe_kv_events!() shared macro for KV event subscription across all three backend clients
  • Add unified subscribe_kv_events() to GrpcClient enum for backend-agnostic usage
  • Part of KV cache event-driven routing (Task 2 of 9)

What changed

File Change
grpc_client/src/lib.rs New impl_subscribe_kv_events!() macro (follows impl_get_tokenizer!() pattern)
grpc_client/src/sglang_scheduler.rs Invoke macro to add subscribe_kv_events() method
grpc_client/src/vllm_engine.rs Invoke macro to add subscribe_kv_events() method
grpc_client/src/trtllm_service.rs Invoke macro to add subscribe_kv_events() method
model_gateway/src/routers/grpc/client.rs Unified subscribe_kv_events() on GrpcClient enum

Why

The gateway needs to subscribe to real-time KV cache events from backends for event-driven cache-aware routing. All three backends expose the same SubscribeKvEvents RPC (added in #557), so a shared macro eliminates duplication — same pattern as the existing impl_get_tokenizer!().

How

  • Macro in lib.rs generates a subscribe_kv_events(start_sequence_number) -> Result<Streaming<KvEventBatch>> method
  • Each backend client invokes the macro inside its impl block
  • GrpcClient::subscribe_kv_events() dispatches to the correct backend variant
  • Uses $crate::common_proto::* types for the request/response (shared proto types)

Test plan

  • cargo build -p smg-grpc-client compiles
  • cargo build -p smg compiles
  • No runtime tests — requires a live backend with the SubscribeKvEvents RPC implemented

Refs: .claude/kv-event/DESIGN.md Section 7

Summary by CodeRabbit

  • New Features
    • Added KV cache event subscription allowing clients to stream KV events from all supported inference engines.
    • Unified client API to subscribe starting from a specified sequence number, providing consistent access to KV event streams across backends.

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Feb 27, 2026
@coderabbitai

coderabbitai Bot commented Feb 27, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 1545371 and 1fb7cad.

📒 Files selected for processing (5)
  • grpc_client/src/lib.rs
  • grpc_client/src/sglang_scheduler.rs
  • grpc_client/src/trtllm_service.rs
  • grpc_client/src/vllm_engine.rs
  • model_gateway/src/routers/grpc/client.rs

📝 Walkthrough

Walkthrough

Adds a shared macro impl_subscribe_kv_events that generates a subscribe_kv_events streaming method for engine clients; invokes the macro in SglangScheduler, TrtllmService, and VllmEngine clients; and exposes a unified subscribe_kv_events(start_seq: u64) async method on GrpcClient that delegates to backend clients.

Changes

Cohort / File(s) Summary
Macro Definition & Export
grpc_client/src/lib.rs
Adds impl_subscribe_kv_events macro that implements a subscribe_kv_events method returning tonic::Streaming<KvEventBatch> and re-exports it via pub(crate) use impl_subscribe_kv_events;.
Client Macro Invocations
grpc_client/src/sglang_scheduler.rs, grpc_client/src/trtllm_service.rs, grpc_client/src/vllm_engine.rs
Invokes crate::impl_subscribe_kv_events!() inside each client impl block, adding the generated subscription method to those clients (two insertion points in vllm_engine.rs).
Public Gateway Interface
model_gateway/src/routers/grpc/client.rs
Adds pub async fn subscribe_kv_events(start_seq: u64) on GrpcClient that delegates to the backend client variants and returns Result<tonic::Streaming<...::KvEventBatch>, Box<dyn std::error::Error + Send + Sync>>.

Sequence Diagram(s)

sequenceDiagram
  participant App
  participant GrpcClient as "GrpcClient (gateway)"
  participant Backend as "Engine Client\n(Sglang/Vllm/Trtllm)"
  participant Engine as "gRPC Engine Server"

  App->>GrpcClient: subscribe_kv_events(start_seq)
  GrpcClient->>Backend: subscribe_kv_events(start_seq)
  Backend->>Engine: gRPC SubscribeKvEventsRequest(start_sequence_number)
  Engine-->>Backend: stream KvEventBatch...
  Backend-->>GrpcClient: forward stream
  GrpcClient-->>App: tonic::Streaming<KvEventBatch>
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Poem

🐰 I dug a macro hole, tidy and neat,
I stitched event streams with flopsy feet.
From gateway to backend the batches hop by,
Kv events tumble—oh my, oh my! 🥕✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately and concisely describes the main change: adding subscribe_kv_events functionality to all backend clients via a shared macro pattern.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch slin/kv-event-grpc-client

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the grpc_client and model_gateway components by introducing a standardized mechanism for subscribing to real-time KV cache events from various backend services. By leveraging a shared macro and a unified GrpcClient enum method, it streamlines the integration of event-driven cache-aware routing, which is crucial for improving the efficiency and responsiveness of the model gateway.

Highlights

  • Shared Macro for KV Event Subscription: A new shared macro, impl_subscribe_kv_events!(), has been introduced to standardize KV event subscription across all backend clients, following the pattern of the existing impl_get_tokenizer!() macro.
  • Unified subscribe_kv_events(): A unified subscribe_kv_events() method has been added to the GrpcClient enum, allowing for backend-agnostic usage of KV event subscriptions.
  • Event-Driven Routing: This change is a foundational step towards enabling event-driven KV cache-aware routing, specifically addressing Task 2 of 9 in the overall effort.
Changelog
  • grpc_client/src/lib.rs
    • New impl_subscribe_kv_events!() macro added for consistent KV event subscription.
  • grpc_client/src/sglang_scheduler.rs
    • Macro invoked to add subscribe_kv_events() method to the SGLang scheduler client.
  • grpc_client/src/trtllm_service.rs
    • Macro invoked to add subscribe_kv_kv_events() method to the TensorRT-LLM service client.
  • grpc_client/src/vllm_engine.rs
    • Macro invoked to add subscribe_kv_events() method to the vLLM engine client.
  • model_gateway/src/routers/grpc/client.rs
    • Unified subscribe_kv_events() method implemented on the GrpcClient enum for backend dispatch.
Activity
  • Initial build checks passed for smg-grpc-client and smg packages.
  • Referenced a design document (.claude/kv-event/DESIGN.md Section 7) for architectural context.
  • Noted that runtime tests are dependent on a live backend implementing the SubscribeKvEvents RPC.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request effectively adds the subscribe_kv_events functionality to all backend gRPC clients by using a shared macro. This is a great approach to reduce code duplication and maintain consistency. The implementation is clean and follows existing patterns in the codebase. The suggestion to improve naming consistency in the unified GrpcClient wrapper aligns with prioritizing external API specifications.

Comment thread model_gateway/src/routers/grpc/client.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_client/src/lib.rs`:
- Around line 67-74: The subscribe_kv_events wrapper currently builds a
tonic::Request and calls client.subscribe_kv_events without injecting trace
context, breaking distributed traces; update the code that constructs
$crate::common_proto::SubscribeKvEventsRequest into a tonic::Request and inject
the current trace/OTel context into the request's metadata (using the same
metadata propagation helper or approach used by the other RPC wrappers like
generate/embed) before calling client.subscribe_kv_events(request). Ensure you
modify the code around Request::new, the SubscribeKvEventsRequest construction,
and the client.subscribe_kv_events call so the request carries the trace headers
exactly as the other traced RPC entrypoints do.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 63e2baa and 1545371.

📒 Files selected for processing (5)
  • grpc_client/src/lib.rs
  • grpc_client/src/sglang_scheduler.rs
  • grpc_client/src/trtllm_service.rs
  • grpc_client/src/vllm_engine.rs
  • model_gateway/src/routers/grpc/client.rs

Comment thread grpc_client/src/lib.rs Outdated
Add impl_subscribe_kv_events!() macro in grpc_client/src/lib.rs
following the existing impl_get_tokenizer!() pattern. The macro
provides a shared subscribe_kv_events() method that returns
tonic::Streaming<KvEventBatch> for long-lived event streams.

The macro is invoked in all three backend clients:
- grpc_client/src/sglang_scheduler.rs
- grpc_client/src/vllm_engine.rs
- grpc_client/src/trtllm_service.rs

Add unified subscribe_kv_events() to GrpcClient enum in
model_gateway/src/routers/grpc/client.rs for backend-agnostic
event subscription.

Refs: DESIGN.md Section 7
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@slin1237
slin1237 force-pushed the slin/kv-event-grpc-client branch from 1545371 to 1fb7cad Compare February 27, 2026 15:55
@slin1237
slin1237 merged commit fbc93c1 into main Feb 27, 2026
20 of 22 checks passed
@slin1237
slin1237 deleted the slin/kv-event-grpc-client branch February 27, 2026 16:06
slin1237 added a commit that referenced this pull request Feb 27, 2026
…uting

Add a positional indexer that uses DashMap<(position, ContentHash), SeqEntry>
for O(1) random access to any depth position, replacing tree pointer-chasing
with direct positional lookup. Jump search skips positions in strides of
jump_size (default 64), yielding amortized O(D/J + W) matching complexity.

What changed:
- kv_index/Cargo.toml: add rustc-hash (FxHash) and xxhash-rust (XXH3) deps
- kv_index/src/event_tree.rs: new 1250-line module with PositionalIndexer
- kv_index/src/lib.rs: export new types (PositionalIndexer, ContentHash,
  SequenceHash, StoredBlock, OverlapScores, WorkerId, compute_content_hash)

Key design decisions:
- Dual-hash scheme: ContentHash (XXH3-64, position-independent, from token
  IDs) for indexing; SequenceHash (position-aware, from backend proto
  block_hash) for disambiguation
- SeqEntry enum with Single/Multi optimization avoids HashMap allocation in
  the common single-sequence-hash case
- Per-worker LevelIndex reverse lookup enables O(1) block removal
- FxHashMap/FxHashSet throughout for 3-5x faster hashing on non-adversarial
  data
- DashMap entries cleaned up when last worker is removed (no memory leak)
- TOCTOU-safe worker lookup uses graceful fallback instead of unwrap
- tracing::warn on unresolvable parent hash, tracing::debug on untracked
  worker removal
- All methods take &self with internal DashMap + parking_lot::RwLock for
  thread safety

Public API:
- apply_stored(worker, blocks, parent_seq_hash): process store events
- apply_removed(worker, seq_hashes): process remove events
- apply_cleared(worker): process cache-clear events
- remove_worker(worker): full worker removal
- find_matches(content_hashes) -> OverlapScores: jump-search matching
- compute_content_hash(token_ids) -> ContentHash: XXH3 hashing
- current_size() -> usize: total blocks across all workers

Tests: 43 tests covering store/match/remove/clear operations, jump search
with various configurations (jump_size=1, 3, 4, 64), concurrent read+write,
DashMap cleanup verification, hash computation, and edge cases.

Refs: #557, #558
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Comment thread grpc_client/src/lib.rs
tonic::Streaming<$crate::common_proto::KvEventBatch>,
Box<dyn std::error::Error + Send + Sync>,
> {
let request = tonic::Request::new($crate::common_proto::SubscribeKvEventsRequest {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wish rust have a way to check this kind of imports. It technically follows the current setting but still a bit ugly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants