Skip to content

feat(completions): add CompletionPreparationStage for gRPC pipeline - #907

Merged
CatherineSue merged 2 commits into
mainfrom
mourya/cmp-2
Mar 26, 2026
Merged

CatherineSue merged 2 commits into
mainfrom
mourya/cmp-2

Conversation

@vschandramourya

@vschandramourya vschandramourya commented Mar 25, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Problem

#840 introduced native /v1/completions typing in the gRPC pipeline by adding RequestType::Completion, FinalResponse::Completion, RequestContext::for_completion(), and execute_completion().

What is still missing is the first endpoint-specific stage for completions, so completion requests still do not have a dedicated Stage 1 preparation flow.

Solution

Add CompletionPreparationStage as the Stage 1 preparation step for native completion requests in the regular gRPC pipeline. This keeps CompletionRequest native in the pipeline and prepares the request directly from completion fields instead of routing through chat-style message templating or laundering through GenerateRequest.

What changed

New: regular/stages/completion/

  • Add model_gateway/src/routers/grpc/regular/stages/completion/mod.rs
  • Add model_gateway/src/routers/grpc/regular/stages/completion/preparation.rs

New: CompletionPreparationStage

Created as the completion-specific Stage 1 preparation step, parallel in role to MessagePreparationStage from the Messages rollout, but much simpler because /v1/completions does not need message-block conversion or chat-template rendering.
It currently:

  • reads CompletionRequest natively from RequestContext
  • resolves the tokenizer from the registry
  • accepts a single prompt string
  • rejects prompt arrays for now with a clear bad-request error
  • tokenizes the prompt text
  • creates the stop decoder from completion request fields
  • stores PreparationOutput and stop_decoder in context for later stages

Modified

  • model_gateway/src/routers/grpc/regular/stages/mod.rs
    • export the new completion stages module
  • model_gateway/src/routers/grpc/regular/stages/preparation.rs
    • delegate RequestType::Completion(_) to CompletionPreparationStage
  • model_gateway/src/routers/grpc/context.rs
    • remove stale #[expect(dead_code)] from completion_request_arc() now that it is used by the new completion preparation stage

Test plan

  • cargo fmt -p smg
  • cargo check -p smg
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • New Features

    • Added support for the /v1/completions endpoint with request validation and preparation.
    • Implemented prompt tokenization and error handling for completion requests.
    • Added stop-sequence decoder setup and configuration for completion responses.
    • Validation added to reject batched (array) prompts for completions.
  • Bug Fixes

    • Minor lint/expectation cleanup (no runtime behavior changes).

Signed-off-by: VS Chandra Mourya <msrinivasa@together.ai>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Mar 25, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the gRPC pipeline by introducing a dedicated preparation stage for native completion requests. This change streamlines the processing of /v1/completions by providing an endpoint-specific flow that directly handles completion requests, tokenizes prompts, and manages stop decoders, thereby improving the efficiency and clarity of the completion API's internal handling.

Highlights

  • New Completion Preparation Stage: Introduced CompletionPreparationStage as the dedicated Stage 1 preparation step for native /v1/completions requests within the gRPC pipeline. This stage handles the CompletionRequest natively, tokenizes the prompt, and creates a stop decoder.
  • Native Request Handling: The new stage directly processes CompletionRequest fields, avoiding conversion through chat-style message templating or GenerateRequest, ensuring a native completion flow.
  • Prompt Processing: The CompletionPreparationStage resolves the tokenizer, accepts a single prompt string, and explicitly rejects prompt arrays with a bad-request error. It then tokenizes the prompt text and stores the PreparationOutput.
  • Pipeline Integration: The main PreparationStage now delegates RequestType::Completion to the new CompletionPreparationStage, integrating it seamlessly into the existing gRPC request processing pipeline.
  • Code Cleanup: A stale #[expect(dead_code)] attribute was removed from the completion_request_arc() method in RequestContext, as the method is now actively used by the new completion preparation stage.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 25, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9eb5f59b-132d-46f5-ae9e-e35e2a608fb1

📥 Commits

Reviewing files that changed from the base of the PR and between 50be037 and 4179d0c.

📒 Files selected for processing (2)
  • model_gateway/src/routers/grpc/regular/stages/completion/preparation.rs
  • model_gateway/src/routers/grpc/regular/stages/preparation.rs

📝 Walkthrough

Walkthrough

Adds Completion API handling to the gRPC pipeline: new CompletionPreparationStage that validates and tokenizes prompts, constructs stop decoders, stores preparation state, and integrates this stage into the regular preparation dispatch. Also removes a dead_code lint expectation from completion_request_arc.

Changes

Cohort / File(s) Summary
Context Configuration
model_gateway/src/routers/grpc/context.rs
Removed #[expect(dead_code, ...)] from RequestContext::completion_request_arc, leaving the clippy::panic expectation.
Completion Stage Module
model_gateway/src/routers/grpc/regular/stages/completion/mod.rs, model_gateway/src/routers/grpc/regular/stages/mod.rs
Added new completion module and declared it in regular stages; re-exports CompletionPreparationStage at crate visibility.
Completion Preparation Stage Implementation
model_gateway/src/routers/grpc/regular/stages/completion/preparation.rs
New CompletionPreparationStage: reads CompletionRequest, rejects array/batched prompts, resolves tokenizer, tokenizes prompt (errors -> bad_request), builds stop decoder, writes PreparationOutput and stop decoder into request state, returns Ok(None).
Preparation Stage Integration
model_gateway/src/routers/grpc/regular/stages/preparation.rs
PreparationStage now contains and delegates to CompletionPreparationStage for RequestType::Completion(_) in its execute routing; constructor initializes the new stage.

Sequence Diagram

sequenceDiagram
    participant Client
    participant PreparationStage
    participant CompletionPreparationStage
    participant Tokenizer
    participant StopDecoder
    participant RequestContext

    Client->>PreparationStage: execute(ctx: CompletionRequest)
    PreparationStage->>CompletionPreparationStage: execute(ctx)
    CompletionPreparationStage->>RequestContext: completion_request_arc()
    RequestContext-->>CompletionPreparationStage: CompletionRequest
    CompletionPreparationStage->>CompletionPreparationStage: validate prompt (no batches)
    CompletionPreparationStage->>Tokenizer: tokenize(prompt_text)
    alt tokenization success
        Tokenizer-->>CompletionPreparationStage: token_ids
        CompletionPreparationStage->>StopDecoder: build(stop, stop_token_ids,...)
        StopDecoder-->>CompletionPreparationStage: stop_decoder
        CompletionPreparationStage->>RequestContext: store PreparationOutput + stop_decoder
        CompletionPreparationStage-->>PreparationStage: Ok(None)
    else tokenization failure
        Tokenizer-->>CompletionPreparationStage: error
        CompletionPreparationStage->>CompletionPreparationStage: emit trace error
        CompletionPreparationStage-->>PreparationStage: bad_request(tokenization_failed)
    end
    PreparationStage-->>Client: final response
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐰 I hopped into the pipeline bright and quick,
Tokenized prompts with a nimble kick,
Built stop decoders with a twitch and a cheer,
Prep stored safe — the flow is clear! ✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding CompletionPreparationStage to the gRPC pipeline for completions requests.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch mourya/cmp-2

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces the initial CompletionPreparationStage for the /v1/completions gRPC endpoint, integrating it into the existing request processing pipeline. This new stage is responsible for resolving prompts, tokenizing input, and creating a stop decoder for completion requests. The review feedback suggests improving code consistency by removing the new() method from CompletionPreparationStage and instantiating it directly as a struct literal, mirroring the approach used for other similar pipeline stages.

Comment thread model_gateway/src/routers/grpc/regular/stages/completion/preparation.rs Outdated
Comment thread model_gateway/src/routers/grpc/regular/stages/preparation.rs Outdated
Signed-off-by: VS Chandra Mourya <msrinivasa@together.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants