Skip to content

feat(grpc_servicer): implement GetTokenizer RPC for vLLM backend - #1142

Merged
slin1237 merged 4 commits into
mainfrom
feat/vllm-get-tokenizer
Apr 14, 2026
Merged

slin1237 merged 4 commits into
mainfrom
feat/vllm-get-tokenizer

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Apr 14, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

The vLLM gRPC servicer does not implement the GetTokenizer RPC, causing the gateway to receive UNIMPLEMENTED when it attempts to fetch tokenizer artifacts from a vLLM worker. In IGW/K8s deployments where gateway and workers run on separate nodes, the gateway cannot load the tokenizer.

Solution

Implement the server-side GetTokenizer handler in the vLLM gRPC servicer, following the same pattern as the sglang implementation (#1136). The handler:

  1. Reads tokenizer files from model_config.tokenizer (vLLM's equivalent of sglang's server_args.tokenizer_path)
  2. Creates an in-memory ZIP archive containing tokenizer files (aligned with crates/tokenizer/src/hub.rs:is_tokenizer_file())
  3. Streams the archive as 64KB GetTokenizerChunk messages with SHA-256 fingerprint on the final chunk

Changes

  • Added GetTokenizer async generator method to VllmEngineServicer
  • Added _build_tokenizer_zip static method for ZIP archive creation
  • Added _TOKENIZER_FILES allowlist and _TOKENIZER_GLOBS patterns (same as sglang)
  • Added imports: hashlib, io, zipfile, Path, AsyncIterator, common_pb2

Test Plan

Setup

# Remote machine: start vLLM gRPC server
vllm serve --model /raid/models/meta-llama/Llama-3.2-1B-Instruct \
  --tensor-parallel-size 1 --port 8080 --grpc

# Local machine: port-forward to remote
ssh -N -L 8080:localhost:8080 ubuntu@<remote-ip> -i ~/.ssh/key

Test 1: Direct gRPC call to verify bundle integrity

Test script
import grpc, hashlib, io, json, zipfile
from pathlib import Path
from smg_grpc_proto import vllm_engine_pb2_grpc
from smg_grpc_proto.generated import common_pb2

MODEL_DIR = Path("/models/meta-llama/Llama-3.2-1B-Instruct")

channel = grpc.insecure_channel("localhost:8080")
stub = vllm_engine_pb2_grpc.VllmEngineStub(channel)

chunks = list(stub.GetTokenizer(common_pb2.GetTokenizerRequest()))
data = b"".join(c.data for c in chunks)
sha256 = chunks[-1].sha256

computed = hashlib.sha256(data).hexdigest()
print(f"Got {len(data)} bytes in {len(chunks)} chunks")
print(f"Received SHA-256:  {sha256}")
print(f"Computed SHA-256:  {computed}")
print(f"Match: {computed == sha256}")

zf = zipfile.ZipFile(io.BytesIO(data))
print(f"\nArchive contains {len(zf.infolist())} files:")
all_match = True
for info in zf.infolist():
    source_file = MODEL_DIR / info.filename
    bundle_bytes = zf.read(info.filename)
    if source_file.exists():
        match = bundle_bytes == source_file.read_bytes()
        status = "OK" if match else "MISMATCH"
        if not match: all_match = False
    else:
        status = "NOT ON DISK"
        all_match = False
    print(f"  {info.filename}  ({info.file_size:,} bytes) — {status}")
print(f"\nAll files match source: {all_match}")

if "tokenizer.json" in zf.namelist():
    tok = json.loads(zf.read("tokenizer.json"))
    print(f"tokenizer.json OK, vocab size: {len(tok.get('model', {}).get('vocab', {}))}")

vLLM server logs:

INFO 04-14 20:59:43 [servicer.py:397] Receive GetTokenizer request
INFO 04-14 20:59:44 [servicer.py:417] Streaming tokenizer bundle: 2335038 bytes,
  sha256=74e6c5d5ee7a3f940366d9c589be986f1ff8955cbaafdd062602586d6666adb0

Test output:

Got 2335038 bytes in 36 chunks
Received SHA-256:  74e6c5d5ee7a3f940366d9c589be986f1ff8955cbaafdd062602586d6666adb0
Computed SHA-256:  74e6c5d5ee7a3f940366d9c589be986f1ff8955cbaafdd062602586d6666adb0
Match: True

Archive contains 5 files:
  tokenizer.json  (9,085,657 bytes) — OK
  tokenizer_config.json  (54,528 bytes) — OK
  config.json  (877 bytes) — OK
  generation_config.json  (189 bytes) — OK
  special_tokens_map.json  (296 bytes) — OK

All files match source: True
tokenizer.json OK, vocab size: 128000

Test 2: Full end-to-end pipeline (cross-machine, vLLM)

Verified the complete worker → GetTokenizer → gateway flow using SSH port-forwarding to simulate the production IGW/K8s scenario.

# Local machine: start SMG (no --model-path, no local model files)
cargo run --bin smg -- --host 0.0.0.0 --port 3002 --prometheus-port 9322 \
  --worker-urls grpc://127.0.0.1:8080 --log-level info
SMG gateway logs (local machine — no model files on disk)
submit_tokenizer_job.rs:109: Submitting tokenizer registration job for model
  /raid/models/meta-llama/Llama-3.2-1B-Instruct from /raid/models/meta-llama/Llama-3.2-1B-Instruct

tokenizer_registration.rs:85: Loading tokenizer '/raid/models/meta-llama/Llama-3.2-1B-Instruct'
  from source: /raid/models/meta-llama/Llama-3.2-1B-Instruct

tokenizer_registration.rs:251: Fetching tokenizer from worker: grpc://127.0.0.1:8080 (runtime: vllm)

tokenizer_registration.rs:217: Tokenizer extracted from temporary path:
  /var/folders/z_/hgkqqcbs1nl4lb_fsgh24c100000gq/T/.tmpqQBQgb

tokenizer_registration.rs:155: Successfully loaded tokenizer
  '/raid/models/meta-llama/Llama-3.2-1B-Instruct' with vocab_size: Some(128000)

Verified model is serving after GetTokenizer load:

curl http://localhost:3002/v1/models | jq
{
  "object": "list",
  "data": [
    {
      "id": "/raid/models/meta-llama/Llama-3.2-1B-Instruct",
      "object": "model",
      "created": 0,
      "owned_by": "self_hosted"
    }
  ]
}

Verified inference works end-to-end:

curl -X POST http://localhost:3002/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "/raid/models/meta-llama/Llama-3.2-1B-Instruct",
       "messages": [{"role": "user", "content": "Hi"}],
       "max_completion_tokens": 10}' | jq '.choices[0].message.content'
"How can I assist you today?"

Note on tokenizer lifecycle

Tokenizer files are not persisted on disk after loading. The flow: stream ZIP → extract to OS temp dir → load into memory → delete temp dir. The tokenizer lives only as Arc<dyn Tokenizer> in the TokenizerRegistry.

Learnings applied from sglang PR review (#1136)

  • No _resolve_tokenizer_dir() — runtime already resolves the path
  • No symlink safety checks — breaks HuggingFace cache (uses symlinks)
  • No return after context.abort() — it raises
  • getbuffer() instead of getvalue() for zero-copy streaming
  • logger.exception() for error logging (matching vLLM file conventions)
  • Guard against None tokenizer path with FAILED_PRECONDITION
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes

Summary by CodeRabbit

  • New Features

    • Added a streaming GetTokenizer RPC to retrieve tokenizer artifacts as chunked archives with final integrity fingerprint.
  • Improvements

    • Consolidated tokenizer bundling into a shared utility for consistent archive construction and chunk sizing across services.
    • Improved error handling when tokenizer artifacts are missing and ensured deduplication during archive creation.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@github-actions github-actions Bot added the grpc gRPC client and router changes label Apr 14, 2026
@coderabbitai

coderabbitai Bot commented Apr 14, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6985c188-111a-4c36-9e46-910432c91160

📥 Commits

Reviewing files that changed from the base of the PR and between deb2b8d and f605936.

📒 Files selected for processing (1)
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

📝 Walkthrough

Walkthrough

Adds a shared tokenizer-bundle utility and implements async server-streaming GetTokenizer in vLLM; SGLang servicer switched to the shared bundle. Streams an in-memory ZIP of tokenizer artifacts in CHUNK_SIZE-bound chunks and emits SHA-256 on the final chunk.

Changes

Cohort / File(s) Summary
Tokenizer bundle module
grpc_servicer/smg_grpc_servicer/tokenizer_bundle.py
New module: CHUNK_SIZE, TOKENIZER_FILES, TOKENIZER_GLOBS, and build_tokenizer_zip(tokenizer_dir: Path) -> io.BytesIO. Builds an in-memory ZIP, deduplicates entries, raises FileNotFoundError if no files.
vLLM servicer
grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Added VllmEngineServicer.GetTokenizer async server-streaming RPC. Resolves model_config.tokenizer, falls back to huggingface snapshot if needed, calls build_tokenizer_zip, computes SHA-256 over ZIP bytes, and streams GetTokenizerChunk messages in CHUNK_SIZE slices (SHA-256 set only on final chunk).
SGLang servicer (refactor)
grpc_servicer/smg_grpc_servicer/sglang/servicer.py
Removed local ZIP-builder and tokenizer constants; switched GetTokenizer to use shared build_tokenizer_zip and CHUNK_SIZE. Removed unused imports and duplicate logic.

Sequence Diagram

sequenceDiagram
    participant Client
    participant Server as gRPC Server
    participant FS as File System
    participant ZIP as Zip Encoder
    participant Hash as SHA-256

    Client->>Server: GetTokenizer(request)
    activate Server
    Server->>FS: Resolve tokenizer path (model_config.tokenizer or snapshot)
    FS-->>Server: tokenizer_dir / files

    Server->>ZIP: build_tokenizer_zip(tokenizer_dir)
    activate ZIP
    ZIP->>FS: Read TOKENIZER_FILES + TOKENIZER_GLOBS
    FS-->>ZIP: file contents
    ZIP->>ZIP: Create in-memory ZIP bytes
    ZIP-->>Server: ZIP bytes
    deactivate ZIP

    Server->>Hash: Compute SHA-256(ZIP bytes)
    Hash-->>Server: hex digest

    loop Stream ZIP chunks
        Server->>Client: GetTokenizerChunk(data=chunk, sha256="")  -- intermediate
    end
    Server->>Client: GetTokenizerChunk(data=final_chunk, sha256=digest)  -- final
    deactivate Server
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested reviewers

  • njhill
  • slin1237

Poem

🐰📦 I hopped through files both near and far,
I zipped their bits into a tiny jar,
I chunk and send, then hash the end —
A rabbit's patchwork, sent to your hand.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title accurately describes the main change: implementing the GetTokenizer RPC for the vLLM backend, which is the primary focus of the changes across three files.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/vllm-get-tokenizer

Comment @coderabbitai help to get the list of available commands and usage tips.

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean implementation. The GetTokenizer RPC correctly mirrors the SGLang version with appropriate vLLM adaptations. Proto types match the service definition, error handling follows existing patterns in the file, and the streaming/chunking logic is correct.

Summary: 0 🔴 Important · 1 🟡 Nit · 0 🟣 Pre-existing

The one nit is about the full duplication of tokenizer-bundle constants and _build_tokenizer_zip across the two servicers — worth extracting to a shared module to prevent future drift, but not blocking.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 448-452: The loop that adds tokenizer files to the ZIP uses
tokenizer_dir.glob(pattern) which can yield nondeterministic ordering; to
produce stable ZIP byte order and SHA fingerprints, sort the glob matches before
iterating: for each pattern in _TOKENIZER_GLOBS, replace direct iteration over
tokenizer_dir.glob(pattern) with iterating over a sorted list of matches (e.g.,
sorted(tokenizer_dir.glob(pattern))), then keep the existing checks and
zf.write(match, match.name) / added.add(match.name) logic intact so behavior
doesn't change aside from deterministic ordering.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 73b3217b-7f0e-41c1-bd43-0253cbefb7bd

📥 Commits

Reviewing files that changed from the base of the PR and between a61cbef and abf3432.

📒 Files selected for processing (1)
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds a GetTokenizer gRPC endpoint to the VllmEngineServicer to stream tokenizer artifacts as a ZIP bundle. The implementation includes logic for file discovery, in-memory compression, and SHA-256 fingerprinting. Feedback suggests offloading the synchronous ZIP creation to a thread pool to avoid blocking the asynchronous event loop and using a more specific gRPC status code when tokenizer files are not found.

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Move _TOKENIZER_FILES, _TOKENIZER_GLOBS, _TOKENIZER_CHUNK_SIZE, and
_build_tokenizer_zip into smg_grpc_servicer/tokenizer_bundle.py so
both sglang and vllm servicers import from one source of truth.

Signed-off-by: Chang Su <chang.s.su@oracle.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/tokenizer_bundle.py`:
- Around line 34-53: The build_tokenizer_zip function returns an io.BytesIO
(buf) still positioned at EOF; ensure the returned stream is readable by
rewinding it before returning (call seek(0) on buf), i.e., in
build_tokenizer_zip after closing the zipfile context but before returning buf
so callers of build_tokenizer_zip (which creates buf and writes TOKENIZER_FILES
/ TOKENIZER_GLOBS into it) receive a stream with its cursor at the start.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: c87f901d-c169-4992-b3a7-a3537e265300

📥 Commits

Reviewing files that changed from the base of the PR and between abf3432 and a443938.

📒 Files selected for processing (3)
  • grpc_servicer/smg_grpc_servicer/sglang/servicer.py
  • grpc_servicer/smg_grpc_servicer/tokenizer_bundle.py
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

Comment thread grpc_servicer/smg_grpc_servicer/tokenizer_bundle.py
Add buf.seek(0) before returning so future callers that use
read() instead of getbuffer() get the full archive.

Signed-off-by: Chang Su <chang.s.su@oracle.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: deb2b8df3b

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +380 to +386
tokenizer_path = self.engine.model_config.tokenizer
if not tokenizer_path:
await context.abort(
grpc.StatusCode.FAILED_PRECONDITION,
"Tokenizer path is not configured on this server.",
)
tokenizer_dir = Path(tokenizer_path)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Resolve tokenizer repo IDs before zipping

GetTokenizer assumes self.engine.model_config.tokenizer is a filesystem directory and immediately wraps it with Path(...), but in vLLM the default tokenizer value is often the HF model ID string (e.g. meta-llama/...) when --tokenizer is not explicitly set. In that common configuration, build_tokenizer_zip searches a non-existent local path and aborts with INTERNAL, so the new RPC still fails for standard vllm serve <hf-id> --grpc deployments and the gateway cannot fetch tokenizer artifacts in offline/separate-node setups.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment is partially correct: model_config.tokenizer is indeed the raw string (HF ID or local path) — vLLM does not resolve it to a local directory in ModelConfig.

However, this is handled gracefully:

  1. When vLLM is started with a local path (e.g., --model /raid/models/...), model_config.tokenizer is that local path and GetTokenizer works correctly (verified in testing).

  2. When vLLM is started with an HF model ID (e.g., --model meta-llama/Llama-3.2-1B-Instruct), build_tokenizer_zip will fail with FileNotFoundError, the handler returns INTERNAL, and SMG falls through to try the next worker or load from HF directly via LoadTokenizerStep (which handles HF IDs natively).

The INTERNAL error path is already exercised — tokenizer_registration.rs:278-285 logs it and continues. No crash, no stuck state.

For production IGW/K8s, models are typically mounted from PVCs at local paths. The HF ID case is an edge case where GetTokenizer gracefully degrades.

model_config.tokenizer can be a HuggingFace model ID instead of a
local path when vLLM is started with --model meta-llama/... without
an explicit --tokenizer. Use snapshot_download(local_files_only=True)
to resolve the ID to the HF cache directory.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@slin1237
slin1237 merged commit 6d41a13 into main Apr 14, 2026
16 of 17 checks passed

tokenizer_dir = Path(snapshot_download(tokenizer_path, local_files_only=True))
except Exception:
pass # Fall through to build_tokenizer_zip which will raise

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: Silently swallowing the snapshot_download exception loses diagnostic context. When the HF resolution fails and build_tokenizer_zip subsequently raises on the non-existent directory (line 401), the logged traceback won't show why the resolution failed (e.g., model not in cache, corrupted snapshot, permission error).

A logger.debug here would make this much easier to diagnose without changing control flow:

Suggested change
pass # Fall through to build_tokenizer_zip which will raise
except Exception:
logger.debug("HF cache lookup failed for %r, will try raw path", tokenizer_path, exc_info=True)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants