Skip to content

perf(multimodal): reduce TokenSpeed encoder input transport overhead - #1879

Merged
slin1237 merged 7 commits into
smg-project:mainfrom
yechank-nvidia:yechan/tokenspeed-encoder-serialization
Jul 7, 2026
Merged

slin1237 merged 7 commits into
smg-project:mainfrom
yechank-nvidia:yechan/tokenspeed-encoder-serialization

Conversation

@yechank-nvidia

@yechank-nvidia yechank-nvidia commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Problem

Qwen image and video preprocessing spent substantial CPU time on repeated buffer materialization, resize setup, pixel indexing, and serial frame processing. The overhead became more pronounced for videos and concurrent requests.

Solution

Optimize the existing Qwen preprocessing path while preserving its public API, output layout, and preprocessing semantics.

The implementation reuses a shared worker pool, parallelizes independent temporal groups, specializes PIL-compatible RGB resize kernels, and reduces indexing and allocation overhead in resize and patchification.

Changes

  • Reuse a bounded shared preprocessing worker pool.
  • Skip identity bicubic resize passes.
  • Process independent video temporal groups in parallel.
  • Optimize PIL-compatible horizontal and vertical RGB bicubic resize kernels.
  • Process adjacent RGB pixels together in the vertical resize pass.
  • Reduce indexing overhead in raw RGB video patchification.
  • Add deterministic image and video preprocessing regression fixtures.
  • Preserve upstream SMG output shapes and SHA-256 fingerprints.
  • Correct Qwen video resize budgeting for temporally padded frame counts.

Performance

Input Concurrency Upstream throughput SMG throughput Throughput speedup
Image 512x512 1 Baseline 1.89x 1.89x
Image 512x512 8 Baseline 5.23x 5.23x
Image 512x512 32 Baseline 7.15x 7.15x
Image 3840x2160 1 Baseline 1.05x 1.05x
Image 3840x2160 8 Baseline 1.49x 1.49x
Image 3840x2160 32 Baseline 2.01x 2.01x
Video 720p x20 frames 1 Baseline 7.29x 7.29x
Video 720p x20 frames 8 Baseline 3.24x 3.24x
Video 720p x20 frames 32 Baseline 3.24x 3.24x
Video 1080p x20 frames 1 Baseline 5.70x 5.70x
Video 1080p x20 frames 8 Baseline 2.79x 2.79x
Video 1080p x20 frames 32 Baseline 2.91x 2.91x
Video 4K x20 frames 1 Baseline 3.57x 3.57x
Video 4K x20 frames 8 Baseline 2.14x 2.14x
Video 4K x20 frames 32 Baseline 2.21x 2.21x

Test Plan

  • Run Qwen image and video preprocessing tests.
  • Run deterministic preprocessing fingerprint tests.
  • Verify optimized and upstream outputs have identical shapes and SHA-256 fingerprints.
  • Benchmark image and video preprocessing at concurrency 1, 8, and 32.
  • Run:

cargo +nightly fmt --check
cargo test -p llm-multimodal qwen
cargo test -p llm-multimodal --test qwen_preprocess_golden
cargo clippy -p llm-multimodal --all-targets -- -D warnings

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Performance Improvements
    • Improved multimodal request assembly with async TokenSpeed generation, reduced tensor copying, and view-based SHM encoder payload serialization.
    • TokenSpeed shared-memory writes now fill a pre-sized, memory-mapped byte buffer for lower overhead.
  • Bug Fixes
    • Fixed TokenSpeed SHM cleanup semantics when assembly tasks are dropped early.
    • Corrected array slicing/serialization to preserve logical row-major order across different memory layouts and dtype conversions.
  • Tests
    • Added/updated coverage for SHM cleanup, mapped-payload writer correctness, parallel u16 (float16/bfloat16) byte parity, and borrowed-slice behavior (including Fortran-contiguous inputs).

Signed-off-by: yechank-nvidia <161688079+yechank-nvidia@users.noreply.github.com>
Signed-off-by: yechank-nvidia <161688079+yechank-nvidia@users.noreply.github.com>
Signed-off-by: yechank-nvidia <161688079+yechank-nvidia@users.noreply.github.com>
@github-actions github-actions Bot added dependencies Dependency updates grpc gRPC client and router changes model-gateway Model gateway crate changes labels Jul 6, 2026
@coderabbitai

coderabbitai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: cabb51d6-b6a4-45b5-9f2a-6d281f09cdd0

📥 Commits

Reviewing files that changed from the base of the PR and between 6f5d9d8 and c836159.

📒 Files selected for processing (1)
  • model_gateway/src/routers/grpc/multimodal.rs

📝 Walkthrough

Walkthrough

This PR makes multimodal tensor handling view-based, moves TokenSpeed assembly onto async/blocking execution, and replaces SHM writes with fixed-size memory-mapped buffers. It also updates request-building call sites and adds coverage for the new serialization and cleanup paths.

Changes

Zero-copy tensor and SHM serialization

Layer / File(s) Summary
Dependencies for mmap and parallelism
model_gateway/Cargo.toml
Adds memmap2, rayon, and rustix with the fs feature.
Async multimodal assembly wiring
model_gateway/src/routers/grpc/multimodal.rs, model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs, model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs
Changes assemble_multimodal_data to async, adds TokenSpeedAssemblyOptions, runs TokenSpeed assembly in spawn_blocking, and updates chat/message request building to await assembly only when multimodal input exists.
Borrowed tensor slicing and serialization
model_gateway/src/routers/grpc/multimodal.rs
Switches encoder slicing and serialization helpers to ArrayViewD, removes owned-array copies, and refactors dtype serialization into buffer-filling paths with rayon-parallel u16 conversion for large outputs.
Mapped TokenSpeed SHM writes
model_gateway/src/routers/grpc/proto_wrapper.rs
Reworks TokenSpeed SHM file creation to reserve a fixed-size /dev/shm file, memory-map it, and write payload bytes directly through the mapped buffer.
Serialization and SHM tests
model_gateway/src/routers/grpc/multimodal.rs, model_gateway/src/routers/grpc/proto_wrapper.rs
Adds tests for parallel u16 serialization parity, borrowed-slice row-major serialization, and Linux-only mapped SHM payload cleanup.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related PRs

  • lightseekorg/smg#588: Both PRs refactor the multimodal assembly pipeline and assemble_multimodal_data flow in multimodal.rs and request building.
  • lightseekorg/smg#776: Also updates the Messages gRPC request-building flow around assemble_multimodal_data(...) and async integration.
  • lightseekorg/smg#1515: Closely aligned with the TokenSpeed multimodal assembly and SHM payload plumbing changed here.

Suggested labels: multimodal

Suggested reviewers: CatherineSue, key4ng, slin1237

Poem

A rabbit hopped through bytes and views,
No copied tensors left to chew.
In shm the mmap softly glowed,
While parallel carrots neatly strode.
🐇 Zero-copy dreams take flight tonight.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main change: optimizing TokenSpeed multimodal input transport overhead.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the serialization of multimodal encoder inputs and shared memory (SHM) writing to use memory-mapped files (memmap2) and parallelized conversions (rayon), avoiding unnecessary allocations and copying. Specifically, write_tokenspeed_shm_with now maps the SHM file directly, and on Linux, it pre-allocates space using rustix::fs::fallocate. The review feedback recommends adding a fallback to file.set_len when fallocate is unsupported (e.g., in certain containerized or virtualized environments) to prevent hard failures.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +553 to +561
#[cfg(target_os = "linux")]
fn reserve_tokenspeed_shm_file(file: &std::fs::File, nbytes: usize) -> std::io::Result<()> {
Ok(rustix::fs::fallocate(
file,
FallocateFlags::empty(),
0,
nbytes as u64,
)?)
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

On Linux, rustix::fs::fallocate can fail with EOPNOTSUPP (if the underlying filesystem, such as certain container mounts or virtualized environments like WSL, does not support allocation) or ENOSYS (on older kernels). To prevent hard failures in these environments, fall back to file.set_len if fallocate returns either of these errors.

Suggested change
#[cfg(target_os = "linux")]
fn reserve_tokenspeed_shm_file(file: &std::fs::File, nbytes: usize) -> std::io::Result<()> {
Ok(rustix::fs::fallocate(
file,
FallocateFlags::empty(),
0,
nbytes as u64,
)?)
}
#[cfg(target_os = "linux")]
fn reserve_tokenspeed_shm_file(file: &std::fs::File, nbytes: usize) -> std::io::Result<()> {
match rustix::fs::fallocate(file, FallocateFlags::empty(), 0, nbytes as u64) {
Ok(()) => Ok(()),
Err(err) if err == rustix::io::Errno::OPNOTSUPP || err == rustix::io::Errno::NOSYS => {
file.set_len(nbytes as u64)
}
Err(err) => Err(err.into()),
}
}

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@coderabbitat Check if addressed

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@coderabbitai Check if addressed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 1551-1581: Offload the TokenSpeed assembly work from the async
request-building path, since assemble_multimodal_data() is still being invoked
inline from the execute() methods in chat/request_building.rs and
messages/request_building.rs and can block a Tokio worker. Move the branch that
calls fill_f32_slice_as_u16_bytes into spawn_blocking or a dedicated CPU-bound
executor, keeping the async flow responsive while the rayon-backed serialization
runs. Use assemble_multimodal_data() and fill_f32_slice_as_u16_bytes() as the
main touchpoints when wiring the offload.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 681aca0c-f3b1-40dd-b544-cdd2e4103445

📥 Commits

Reviewing files that changed from the base of the PR and between 1cb3236 and 7b20d10.

📒 Files selected for processing (3)
  • model_gateway/Cargo.toml
  • model_gateway/src/routers/grpc/multimodal.rs
  • model_gateway/src/routers/grpc/proto_wrapper.rs

Comment thread model_gateway/src/routers/grpc/multimodal.rs
@yechank-nvidia yechank-nvidia changed the title Yechan/tokenspeed encoder serialization perf(multimodal): reduce TokenSpeed encoder input transport overhead Jul 6, 2026
Signed-off-by: yechank-nvidia <161688079+yechank-nvidia@users.noreply.github.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 971-979: Move the TokenSpeed option resolution into the blocking
work so the `/dev/shm` probe is not done on the async path. Update the
GrpcClient::TokenSpeed branch in multimodal.rs to compute
tokenspeed_assembly_options inside the spawn_blocking closure together with
assemble_tokenspeed_with_options, and keep the
resolve_tokenspeed_shm_enabled/tokenspeed_shm_dev_writable call chain off the
runtime thread.

In `@model_gateway/src/routers/grpc/proto_wrapper.rs`:
- Around line 555-563: The fallocate handling in the proto wrapper should retry
interrupted syscalls instead of immediately falling back or returning an error.
Update the logic around the `rustix::fs::fallocate` call in `proto_wrapper.rs`
to use `rustix::io::retry_on_intr` (or equivalent retry handling) for
`Errno::INTR`, while keeping the existing `OPNOTSUPP` and `NOSYS` fallback to
`file.set_len` and preserving the current error conversion path for other
failures.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: c4f5fc04-0680-4802-922d-80d96c1dd7bf

📥 Commits

Reviewing files that changed from the base of the PR and between 7b20d10 and d094d5b.

📒 Files selected for processing (4)
  • model_gateway/src/routers/grpc/multimodal.rs
  • model_gateway/src/routers/grpc/proto_wrapper.rs
  • model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs
  • model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs

Comment thread model_gateway/src/routers/grpc/multimodal.rs
Comment thread model_gateway/src/routers/grpc/proto_wrapper.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6f5d9d8bbc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread model_gateway/src/routers/grpc/multimodal.rs Outdated
lightseek-bot and others added 2 commits July 6, 2026 13:42
Signed-off-by: yechank-nvidia <161688079+yechank-nvidia@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Dependency updates grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants