Skip to content

feat(messages api): support streaming - #280

Merged
key4ng merged 19 commits into
mainfrom
message-api-stream
Feb 5, 2026
Merged

key4ng merged 19 commits into
mainfrom
message-api-stream

Conversation

@key4ng

@key4ng key4ng commented Feb 2, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

The Anthropic Messages API (/v1/messages) lacked streaming support, limiting real-time response delivery for clients using the stream: true parameter.

Solution

Implement full SSE (Server-Sent Events) streaming for the Anthropic Messages API, following the Anthropic streaming specification with proper event types and delta handling.

Changes

  • Add pipeline and context structure for Message API request processing
  • Implement streaming response parser for Anthropic SSE events (message_start, content_block_start, content_block_delta, message_delta,
    message_stop)
  • Add streaming metrics tracking (TTFT, throughput, token estimation)
  • Create stage-based pipeline architecture:
    • validation - Request validation
    • worker_selection - Route to appropriate backend
    • request_building - Transform request for backend
    • request_execution - Execute request with streaming support
    • response_processing - Handle streaming SSE response
    • dispatch_metadata - Track request metadata
  • Support content block types: text, thinking, tool_use
  • Clean up legacy non-streaming implementation

Test Plan

cargo run --bin smg -- --backend anthropic --worker-urls https://api.anthropic.com

curl http://localhost:30000/v1/messages \
    -H 'Content-Type: application/json' \
    -H 'anthropic-version: 2023-06-01' \
    -H "X-Api-Key: $ANTHROPIC_API_KEY" \
    --max-time 600 \
    -d '{
          "max_tokens": 1024,
          "messages": [
            {
              "content": "Hello, world",
              "role": "user"
            }
          ],
          "model": "claude-3-haiku-20240307",
          "stream": true
        }'

event: message_start
data: {"type":"message_start","message":{"model":"claude-3-haiku-20240307","id":"msg_01H3LPXrHV6SVKE36tZ1GPvd","type":"message","role":"assistant","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":10,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard"}}     }

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""} }

event: ping
data: {"type": "ping"}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}              }

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"!"}               }

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":" It"}      }

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"'s nice to meet"}             }

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":" you."}  }

event: content_block_stop
data: {"type":"content_block_stop","index":0      }

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":10,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"output_tokens":12}              }

event: message_stop
data: {"type":"message_stop"     }
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • New Features

    • Runtime validation for message submissions.
  • Improvements

    • New stage-based message pipeline improving routing reliability and observability.
    • Safer response handling with size limits and incremental reads to prevent oversized responses.
    • Better worker selection and load tracking for improved request stability and clearer streaming vs non-streaming handling.
    • Tighter header propagation controls and clearer gateway timeout/error responses.
  • Chores

    • Removed legacy in-file streaming/non-streaming handlers and associated tests.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @key4ng, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a comprehensive pipeline architecture for the Anthropic router, enhancing its functionality, security, and observability. The changes include a structured approach to processing Messages API requests, improved streaming support with detailed metrics, and security measures to prevent credential leakage and DoS attacks. The new pipeline architecture provides a more modular and maintainable design, allowing for easier addition of new features and improvements in the future.

Highlights

  • Pipeline Architecture: Introduces a 6-stage pipeline for processing Messages API requests in the Anthropic router, including validation, worker selection, request building, dispatch metadata, request execution, and response processing.
  • Request Context: Adds a RequestContext struct to manage request data, shared components, and processing state throughout the pipeline stages.
  • Streaming Support: Implements SSE streaming support with metrics collection, including time-to-first-token (TTFT) and token throughput.
  • Security Enhancements: Improves security by filtering workers based on the provider in multi-provider setups to prevent credential leakage and adds size limits to error responses to prevent DoS attacks.
  • Metrics and Observability: Adds detailed metrics for both streaming and non-streaming requests, including request duration, token usage, and error counts.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a significant refactoring of the Anthropic router, replacing separate streaming and non-streaming handlers with a unified, stage-based pipeline architecture, which improves maintainability and extensibility. The current implementation has critical reliability issues, including flawed worker load management leading to inaccurate routing decisions, and a data corruption bug in MetricsStream that can cause duplicated data in streaming responses. A critical security vulnerability related to unbounded memory consumption in the SSE parser has been identified, but due to its cross-cutting nature and design implications, it should be addressed in a dedicated pull request. While there are excellent security enhancements in models.rs preventing credential leakage and adding response size limits, the remaining critical issues, along with a minor documentation issue, need to be addressed before merging.

Comment thread model_gateway/src/routers/anthropic/stages/response_processing.rs Outdated
Comment thread model_gateway/src/routers/anthropic/stages/request_execution.rs Outdated
Comment thread model_gateway/src/routers/anthropic/stages/response_processing.rs Outdated
@github-actions github-actions Bot added the model-gateway Model gateway crate changes label Feb 2, 2026
@key4ng
key4ng force-pushed the message-api-stream branch 2 times, most recently from 554941f to 79636aa Compare February 4, 2026 03:41
@key4ng key4ng changed the title Message api stream feat(messages api): support streaming Feb 4, 2026
@key4ng
key4ng marked this pull request as ready for review February 4, 2026 03:51
@slin1237

slin1237 commented Feb 4, 2026

Copy link
Copy Markdown
Member

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces streaming support for the Anthropic Messages API by refactoring request handling into a well-structured, multi-stage pipeline, enhancing modularity, maintainability, and extensibility. It also adds detailed streaming metrics. Several minor cleanups and improvements are suggested, such as removing unused code, eliminating leftover helper methods, and ensuring consistent metrics calculation for both streaming and non-streaming responses.

Comment thread model_gateway/src/routers/anthropic/context.rs Outdated
Comment thread model_gateway/src/routers/anthropic/stages/response_processing.rs
Comment thread model_gateway/src/routers/anthropic/streaming/metrics.rs Outdated
Comment thread model_gateway/src/routers/anthropic/streaming/metrics.rs Outdated
Comment thread model_gateway/src/routers/anthropic/router.rs Outdated
Comment thread model_gateway/src/routers/anthropic/streaming/metrics.rs Outdated
Comment thread model_gateway/src/routers/anthropic/context.rs Outdated
Comment thread model_gateway/src/routers/anthropic/context.rs Outdated
Comment thread model_gateway/src/routers/anthropic/context.rs Outdated
Comment thread model_gateway/src/routers/anthropic/context.rs Outdated
@key4ng
key4ng force-pushed the message-api-stream branch from cee7677 to 8156b27 Compare February 4, 2026 18:28
@key4ng
key4ng requested a review from slin1237 February 4, 2026 18:34
@coderabbitai

coderabbitai Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds a staged Messages pipeline for the Anthropic router: introduces RequestContext and SharedComponents, a MessagesPipeline wiring WorkerSelection, RequestBuilding, RequestExecution, and ResponseProcessing stages; removes legacy streaming/non-streaming handlers; adds response size limits, Anthropic worker filtering, header propagation rules, and request validation.

Changes

Cohort / File(s) Summary
Pipeline core
model_gateway/src/routers/anthropic/context.rs, model_gateway/src/routers/anthropic/pipeline.rs
Add per-request RequestContext and SharedComponents; implement MessagesPipeline with ordered stage execution, early-return semantics, and masked Debug representations.
Stages
model_gateway/src/routers/anthropic/stages/mod.rs, model_gateway/src/routers/anthropic/stages/request_building.rs, model_gateway/src/routers/anthropic/stages/request_execution.rs, model_gateway/src/routers/anthropic/stages/response_processing.rs, model_gateway/src/routers/anthropic/stages/worker_selection.rs
Introduce PipelineStage trait and four stage implementations: worker selection, request building (URL + header propagation), request execution (HTTP call, timeout/error mapping, worker load handling), and response processing (streaming + non-streaming flows, size limits, LoadTrackingStream, metrics).
Router integration & refactor
model_gateway/src/routers/anthropic/router.rs, model_gateway/src/routers/anthropic/mod.rs
AnthropicRouter now constructs/uses MessagesPipeline and SharedComponents; remove direct http_client and legacy worker-helper methods; route_messages requires model_id: &str and delegates to pipeline.
Removed handlers & tests
model_gateway/src/routers/anthropic/messages/mod.rs, model_gateway/src/routers/anthropic/messages/non_streaming.rs, model_gateway/src/routers/anthropic/messages/streaming.rs, model_gateway/src/routers/anthropic/messages/tools.rs
Delete previous non-streaming and streaming message handler modules and associated tests; keep tools::extract_tool_calls re-export; tests for extract_tool_calls removed.
Utilities & safety
model_gateway/src/routers/anthropic/utils.rs, model_gateway/src/routers/anthropic/models.rs
Add Anthropic-aware worker filtering (get_healthy_anthropic_workers, find_best_worker_for_model), response body size enforcement (MAX_RESPONSE_SIZE, read_response_body_limited, ReadBodyResult), content-length prechecks, and refined header propagation (should_propagate_header).
Router manager & API signatures
model_gateway/src/routers/mod.rs, model_gateway/src/routers/router_manager.rs, model_gateway/src/server.rs
Change route_generate/route_messages to require model_id: &str; RouterManager resolves model IDs in IGW mode and forwards to selected router; call sites updated accordingly.
Request validation
protocols/src/messages.rs
Derive Validate for CreateMessageRequest with validations: non-empty model, non-empty messages, and max_tokens >= 1; add no-op Normalizable impl.
Errors
model_gateway/src/routers/error.rs
Add gateway_timeout(...) helper returning 504 responses.

Sequence Diagram

sequenceDiagram
    participant Client
    participant Router as AnthropicRouter
    participant Pipeline as MessagesPipeline
    participant WS as WorkerSelectionStage
    participant RB as RequestBuildingStage
    participant RE as RequestExecutionStage
    participant RP as ResponseProcessingStage
    participant Registry as WorkerRegistry
    participant Worker as RemoteWorker

    Client->>Router: route_messages(request, headers, model_id)
    Router->>Pipeline: execute(request, headers, model_id)

    Pipeline->>WS: execute(ctx)
    WS->>Registry: find_best_worker_for_model(model_id)
    Registry-->>WS: Arc<Worker>
    WS-->>Pipeline: Ok(None)

    Pipeline->>RB: execute(ctx)
    RB->>RB: build URL and propagate headers
    RB-->>Pipeline: Ok(None)

    Pipeline->>RE: execute(ctx)
    RE->>Worker: POST /v1/messages (JSON + headers) with timeout
    alt worker responds
        Worker-->>RE: Response
        RE-->>Pipeline: Ok(None)
    else error / timeout
        RE-->>Pipeline: Err(Response)
    end

    Pipeline->>RP: execute(ctx)
    alt streaming
        RP->>RP: wrap stream with LoadTrackingStream
        RP-->>Pipeline: Ok(Some(streaming Response))
    else non-streaming
        RP->>RP: read/parse body, enforce limits, record metrics
        RP-->>Pipeline: Ok(Some(Response))
    end

    Pipeline-->>Router: Response
    Router-->>Client: Response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Poem

🐰 I tunneled through a stub to see,
Stages stitched in sequence, one-two-three,
Headers hopped, the bytes did stream,
Workers hummed inside the scheme,
A rabbit cheers — pipeline's a dream!

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 66.15% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'feat(messages api): support streaming' directly summarizes the main objective: adding streaming support to the Anthropic Messages API, which is the primary change across all modified files.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch message-api-stream

Comment @coderabbitai help to get the list of available commands and usage tips.

@slin1237

slin1237 commented Feb 4, 2026

Copy link
Copy Markdown
Member

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Feb 4, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Fix all issues with AI agents
In `@model_gateway/src/routers/anthropic/models.rs`:
- Around line 73-92: The code currently calls response.bytes().await (via
req_builder.send().await -> response) and only then checks bytes.len(), which
can buffer an unbounded body; change this to first inspect
response.content_length() and immediately return the error if Some(len) >
MAX_RESPONSE_SIZE, and if content_length is unknown use response.bytes_stream()
to consume the body incrementally while summing chunk lengths and aborting as
soon as the accumulated size exceeds MAX_RESPONSE_SIZE; update the code paths
around req_builder.send().await, response.content_length(), and
response.bytes_stream() so the body is never fully buffered beyond the
configured MAX_RESPONSE_SIZE.

In `@model_gateway/src/routers/anthropic/stages/request_execution.rs`:
- Around line 58-95: The circuit breaker is only recorded in
request_execution.rs via worker.record_outcome(status.is_success()), which
misreports successes when later parsing or streaming fails; move or defer
recording success into ResponseProcessingStage so the outcome is recorded only
after successful parsing/consumption, and add explicit failure recordings: call
worker.record_outcome(false) inside handle_success_response when JSON parsing
fails, and inside LoadTrackingStream::drop when the stream is interrupted
(ensure you can detect interruption vs successful completion), and remove or
guard the existing record_outcome call in request_execution.rs to avoid
double-reporting.
- Around line 17-68: The request builder currently calls
.timeout(Duration::from_secs(DEFAULT_WORKER_TIMEOUT_SECS)) which hard-codes 120s
and overrides the gateway/client timeout; remove the hard-coded
DEFAULT_WORKER_TIMEOUT_SECS usage in RequestExecutionStage::execute and instead
obtain the configured timeout from the shared configuration (e.g.
SharedComponents or request_timeout_secs) — either by adding a timeout field to
RequestExecutionStage via new(http_client, request_timeout_secs) or by reading
ctx.shared_components.request_timeout_secs inside execute — then apply that
value (converted to Duration) to the reqwest request builder (or omit
per-request timeout to rely on the client-level timeout) so gateway-wide timeout
configuration is respected.

In `@model_gateway/src/routers/anthropic/stages/response_processing.rs`:
- Around line 280-346: The handler handle_error_response currently calls
response.bytes().await which buffers the entire body before checking
MAX_ERROR_RESPONSE_SIZE; change this to consume response.bytes_stream() and read
chunks up to MAX_ERROR_RESPONSE_SIZE (accumulating into a Vec<u8> and stopping
when the limit is exceeded), returning a truncated/overflow message if exceeded,
and handling stream errors similarly to the Err branch; ensure you still produce
a lossily-decoded body_preview, log the body_size and overflow event, and
preserve the existing metrics, worker.decrement_load() and final Err((status,
body).into_response()) behavior.

In `@model_gateway/src/routers/anthropic/stages/worker_selection.rs`:
- Around line 28-74: The find_best_worker_for_model function currently calls
worker_registry.get_workers_filtered with None provider and thus can pick
wildcard workers across providers; change it to be provider-aware like
models.rs: detect if multiple providers exist in the cluster and, when Anthropic
credentials may be present for the request, pass Some(ProviderType::Anthropic)
(or the appropriate ProviderType) into get_workers_filtered instead of None;
update the call site in find_best_worker_for_model (and keep using
supports_model and min_by_key(load) as before) so only workers from the intended
provider are considered and credentials cannot leak to other providers.
🧹 Nitpick comments (2)
model_gateway/src/routers/anthropic/stages/response_processing.rs (1)

225-278: Return Ok(Some) for successful responses to match StageResult contract.

Err(...) is documented as the error path; consider using Ok(Some(...)) for success (and making the pipeline log message more neutral if you adopt this).

♻️ Suggested adjustment
-                Err((StatusCode::OK, Json(message)).into_response())
+                Ok(Some((StatusCode::OK, Json(message)).into_response()))
model_gateway/src/routers/anthropic/context.rs (1)

120-155: Prefer validated streaming flag when available.
If the validation stage ever normalizes or overrides streaming, is_streaming() could diverge from the pipeline state. Consider consulting state.validation first and falling back to the request flag.

♻️ Suggested adjustment
     pub fn is_streaming(&self) -> bool {
-        self.input.request.stream.unwrap_or(false)
+        self.state
+            .validation
+            .as_ref()
+            .map(|v| v.is_streaming)
+            .unwrap_or_else(|| self.input.request.stream.unwrap_or(false))
     }

Comment thread model_gateway/src/routers/anthropic/models.rs
Comment thread model_gateway/src/routers/anthropic/stages/request_execution.rs Outdated
Comment thread model_gateway/src/routers/anthropic/stages/request_execution.rs
Comment thread model_gateway/src/routers/anthropic/stages/response_processing.rs Outdated
Comment thread model_gateway/src/routers/anthropic/stages/worker_selection.rs Outdated
@github-actions github-actions Bot added documentation Improvements or additions to documentation python-bindings Python bindings changes ci CI/CD configuration changes tests Test changes tool-parser Tool/function call parser changes labels Feb 4, 2026
@key4ng
key4ng force-pushed the message-api-stream branch from ea2628d to 39bc681 Compare February 4, 2026 22:49
@github-actions github-actions Bot removed documentation Improvements or additions to documentation python-bindings Python bindings changes ci CI/CD configuration changes tests Test changes tool-parser Tool/function call parser changes labels Feb 4, 2026
@key4ng
key4ng force-pushed the message-api-stream branch from 346ac2b to acfe956 Compare February 5, 2026 20:42

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Fix all issues with AI agents
In `@model_gateway/src/routers/anthropic/stages/response_processing.rs`:
- Around line 168-176: The Content-Type check for SSE is case-sensitive and can
miss mixed-case headers; update the check in the response handling (where
content_type is derived from response.headers().get(header::CONTENT_TYPE)) to
perform a case-insensitive comparison — e.g., convert content_type to lowercase
with to_ascii_lowercase() (or use a case-insensitive comparison) before calling
contains("text/event-stream") so the if condition that detects SSE streams
correctly matches values like "Text/Event-Stream" or "text/Event-Stream".

In `@model_gateway/src/routers/anthropic/utils.rs`:
- Around line 115-142: The function filter_by_anthropic_provider currently
treats missing provider (default_provider() -> None) as compatible with a single
explicit provider which can leak Anthropic credentials; change the
provider-detection logic to treat None as a distinct provider value so any mix
of Some(...) and None counts as multiple providers. Concretely, in
filter_by_anthropic_provider update the tracking variable (first_provider) to
hold Option<Option<ProviderType>> and compare the full Option<ProviderType>
returned by Worker::default_provider() when scanning workers; if you detect
differing Option<ProviderType> values set has_multiple_providers true as before,
and keep the existing filtering branch that only keeps workers where
default_provider() == Some(ProviderType::Anthropic).
- Around line 85-108: The current read_response_body_limited function decodes
each chunk with from_utf8_lossy which can corrupt multibyte UTF-8 sequences
split across chunks; fix it by accumulating raw bytes into a Vec<u8> (instead of
appending chunk-decoded Strings), enforcing the max_size cap by checking
total_size after adding each chunk and returning ReadBodyResult::TooLarge if
exceeded, and only after the stream completes decode the entire buffer once
using String::from_utf8 (or map the UTF-8 error to ReadBodyResult::Error) and
return ReadBodyResult::Ok(body) on success; keep the existing chunk read error
handling (Err branch) returning ReadBodyResult::Error(e.to_string()) and update
references within read_response_body_limited accordingly.
🧹 Nitpick comments (1)
model_gateway/src/routers/anthropic/stages/request_execution.rs (1)

37-118: Emit router metrics on dispatch failures.

When send() fails, ResponseProcessingStage never runs, so request/duration/error metrics may be missing for timeouts/connect errors. Consider recording them here for observability parity.

🧭 Suggested metrics hook for send failures
@@
-        let url = &http_request.url;
+        let url = &http_request.url;
+        let model_id = &ctx.input.model_id;
+        let start_time = ctx.start_time;
+        let is_streaming = ctx.is_streaming();
@@
                 // Record circuit breaker failure
                 worker.record_outcome(false);
+
+                // Record router metrics for dispatch failures (response_processing won't run)
+                Metrics::record_router_request(
+                    metrics_labels::ROUTER_HTTP,
+                    metrics_labels::BACKEND_EXTERNAL,
+                    metrics_labels::CONNECTION_HTTP,
+                    model_id,
+                    "messages",
+                    bool_to_static_str(is_streaming),
+                );
+                Metrics::record_router_duration(
+                    metrics_labels::ROUTER_HTTP,
+                    metrics_labels::BACKEND_EXTERNAL,
+                    metrics_labels::CONNECTION_HTTP,
+                    model_id,
+                    "messages",
+                    start_time.elapsed(),
+                );
+                Metrics::record_router_error(
+                    metrics_labels::ROUTER_HTTP,
+                    metrics_labels::BACKEND_EXTERNAL,
+                    metrics_labels::CONNECTION_HTTP,
+                    model_id,
+                    "messages",
+                    metrics_labels::ERROR_BACKEND,
+                );
➕ Import metrics utilities
-use crate::routers::{anthropic::context::RequestContext, error};
+use crate::{
+    observability::metrics::{bool_to_static_str, metrics_labels, Metrics},
+    routers::{anthropic::context::RequestContext, error},
+};

Comment thread model_gateway/src/routers/anthropic/stages/response_processing.rs Outdated
Comment thread model_gateway/src/routers/anthropic/utils.rs Outdated
Comment thread model_gateway/src/routers/anthropic/utils.rs
@key4ng
key4ng merged commit 0438842 into main Feb 5, 2026
20 checks passed
@key4ng
key4ng deleted the message-api-stream branch February 5, 2026 21:47
ppraneth pushed a commit that referenced this pull request Feb 18, 2026
Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes protocols Protocols crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants