Skip to content

fix(metrics): engine metrics from PROMETHEUS_MULTIPROC_DIR + graceful bind - #1075

Closed
ConnorLi96 wants to merge 9 commits into
mainfrom
connorli/engine-metrics
Closed

ConnorLi96 wants to merge 9 commits into
mainfrom
connorli/engine-metrics

Conversation

@ConnorLi96

@ConnorLi96 ConnorLi96 commented Apr 9, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  1. In gRPC mode, engine-level metrics (TPS, queue depth, etc.) from the backend worker are not collected because the gateway doesn't read from PROMETHEUS_MULTIPROC_DIR.
  2. Empty prometheus output during startup causes parse error log spam on every scrape interval.
  3. Metrics server panics on port conflict instead of degrading gracefully.

Solution

  • Read and merge engine metrics from PROMETHEUS_MULTIPROC_DIR .db files in gRPC mode
  • Guard against empty prometheus output at both collector and call site
  • On bind failure, log error and return no-op handle instead of panicking

Changes

  • bindings/python/src/smg/serve.py — set PROMETHEUS_MULTIPROC_DIR
  • model_gateway/src/core/worker_manager.rs — metrics collection + empty output guard
  • model_gateway/src/observability/metrics_server.rs — graceful bind failure

Co-authored-by: Scott Lee scott@together.ai

Checklist:

  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes

Made with Cursor

Summary by CodeRabbit

  • New Features

    • Worker metrics & load monitoring with per-worker load collection, aggregated engine metrics, distributed cache flush, and new health endpoints for generation routes.
  • Bug Fixes

    • Metrics endpoint now tolerates bind failures and engine-metrics errors; improved backend timing, selection logging, and request-level token timing (TTFT/inter-token) metrics.
  • Chores

    • Temporary Prometheus multiprocess lifecycle support for gRPC workers (creation & cleanup).

@github-actions github-actions Bot added python-bindings Python bindings changes model-gateway Model gateway crate changes labels Apr 9, 2026
@coderabbitai

coderabbitai Bot commented Apr 9, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds Python ServeOrchestrator Prometheus multiprocess dir lifecycle for gRPC mode; introduces a Rust WorkerManager and LoadMonitor for worker telemetry, load polling, and engine-metrics aggregation (including a Python subprocess path); integrates engine metrics into /metrics; and adds timing/tracing and health endpoints.

Changes

Cohort / File(s) Summary
Python Prometheus dir
bindings/python/src/smg/serve.py
Create temporary PROMETHEUS_MULTIPROC_DIR for gRPC connection mode, store path on the orchestrator, and remove the directory during orchestrator cleanup.
Rust worker lifecycle & telemetry
model_gateway/src/core/worker_manager.rs
New module: WorkerManager (list worker URLs, flush HTTP worker caches concurrently, fetch per-worker loads via HTTP/gRPC, collect engine metrics including spawning python3 to read PROMETHEUS_MULTIPROC_DIR) and LoadMonitor (per-group polling, watch subscription, start/stop group loops).
Metrics server integration
model_gateway/src/observability/metrics_server.rs, model_gateway/src/server.rs
Register engine-metrics dependencies at startup; /metrics handler appends aggregated engine metrics from WorkerManager::get_engine_metrics; bind failures log and return a no-op task instead of panicking.
Worker metrics collection (alternate path)
model_gateway/src/worker/manager.rs
Split metrics collection by ConnectionMode: fan-out /metrics for HTTP workers, derive HTTP metrics URL for gRPC workers (grpc port + 1) and scrape separately; added grpc_worker_metrics_url helper and unit tests.
Metrics aggregation fixes
model_gateway/src/worker/metrics_aggregator.rs, model_gateway/tests/metrics_aggregator_test.rs
Replace global colon rewrite with sentinel encoding/decoding around parse/serialize; update tests/expected output to preserve namespace:metric names and label colons.
Routing instrumentation
model_gateway/src/routers/.../request_execution.rs, .../worker_selection.rs, http/pd_router.rs, http/router.rs
Add timing instrumentation and info/warn tracing for backend/PD/gRPC requests and worker selection events; logging only, no selection logic changes.
Health endpoints & OpenAI health changes
model_gateway/src/routers/grpc/.../pd_router.rs, model_gateway/src/routers/grpc/router.rs, model_gateway/src/routers/openai/health.rs
Add health_generate handlers returning 200/503 based on worker health and listing unhealthy workers; remove external-only filtering in OpenAI health responses and adjust returned fields.
gRPC servicer metrics
grpc_servicer/smg_grpc_servicer/sglang/request_manager.py
Add request-level TTFT tracking and TokenizerMetricsCollector usage; record TTFT, inter-token latency, and finished-request metrics.

Sequence Diagram(s)

sequenceDiagram
    participant MetricsServer as SMG Metrics Server
    participant WM as WorkerManager
    participant HTTP as HTTP Worker
    participant GRPC as gRPC Worker
    participant Py as Python Subprocess
    participant Dir as PROMETHEUS_MULTIPROC_DIR

    MetricsServer->>WM: get_engine_metrics()
    activate WM
    WM->>HTTP: HTTP GET /metrics (fan-out for HTTP workers)
    HTTP-->>WM: metrics text

    Note over WM,GRPC: gRPC workers use multiprocess files
    WM->>GRPC: identify gRPC workers
    WM->>Py: spawn `python3` subprocess to read Dir and emit Prometheus text
    activate Py
    Py->>Dir: read multiprocess files
    Py-->>WM: metrics text
    deactivate Py

    WM->>WM: aggregate_metrics(all texts)
    WM-->>MetricsServer: aggregated metrics text
    deactivate WM
    MetricsServer-->>MetricsServer: render combined /metrics response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested labels

tests

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237
  • gongwei-130

Poem

🐰 I dug a tiny metrics den,
For prometheus files from all my friends,
Workers hum and counters sing,
I twitch my nose at every ping,
Carrots, logs, and telemetry—what fun!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 77.36% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main changes: reading engine metrics from PROMETHEUS_MULTIPROC_DIR and implementing graceful handling of bind failures.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch connorli/engine-metrics

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Comment on lines +339 to +347
Ok(text) if !text.trim().is_empty() => {
metric_packs.push(MetricPack {
labels: vec![],
metrics_text: text,
});
}
Ok(_) => {
// No metrics available yet from gRPC workers — skip silently
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: The Ok(_) arm on line 345 is dead code. collect_prometheus_multiproc_metrics() already returns Err when the output is empty (lines 116-118), so the Ok(text) variant will always contain non-empty text, making the if !text.trim().is_empty() guard always true and the Ok(_) branch unreachable.

Consider simplifying:

Suggested change
Ok(text) if !text.trim().is_empty() => {
metric_packs.push(MetricPack {
labels: vec![],
metrics_text: text,
});
}
Ok(_) => {
// No metrics available yet from gRPC workers — skip silently
}
Ok(text) => {
metric_packs.push(MetricPack {
labels: vec![],
metrics_text: text,
});
}

Comment on lines +94 to +107
let output = tokio::process::Command::new("python3")
.args([
"-c",
"import sys\n\
from prometheus_client import CollectorRegistry, generate_latest\n\
from prometheus_client.multiprocess import MultiProcessCollector\n\
registry = CollectorRegistry()\n\
MultiProcessCollector(registry)\n\
sys.stdout.buffer.write(generate_latest(registry))\n",
])
.env("PROMETHEUS_MULTIPROC_DIR", &dir)
.output()
.await
.map_err(|e| format!("failed to run python3 prometheus collector: {e}"))?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: No timeout on the python3 subprocess. If the process hangs (e.g., corrupt .db file in the multiproc dir, or python3 not found and the OS stalls), the metrics endpoint request will hang indefinitely. Since this runs on every /metrics scrape, a stuck process could accumulate and exhaust resources.

Consider adding a timeout, e.g. via tokio::time::timeout:

let output = tokio::time::timeout(
    std::time::Duration::from_secs(5),
    tokio::process::Command::new("python3")
        .args([...])
        .env("PROMETHEUS_MULTIPROC_DIR", &dir)
        .output(),
)
.await
.map_err(|_| "prometheus collector timed out after 5s".to_string())?
.map_err(|e| format!("failed to run python3 prometheus collector: {e}"))?;

Comment on lines +686 to +699
def _cleanup_prometheus_dir(self) -> None:
"""Remove the temporary prometheus multiprocess directory and its .db files."""
if self._prometheus_dir is None:
return
try:
shutil.rmtree(self._prometheus_dir)
logger.info("Cleaned up PROMETHEUS_MULTIPROC_DIR=%s", self._prometheus_dir)
except OSError as e:
logger.warning(
"Failed to clean up PROMETHEUS_MULTIPROC_DIR=%s: %s",
self._prometheus_dir,
e,
)
self._prometheus_dir = None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: After removing the directory, the PROMETHEUS_MULTIPROC_DIR environment variable still points to the now-deleted path. If any code (e.g., a straggling atexit handler or a library import) reads the env var after cleanup, it will reference a non-existent directory, which can cause confusing errors.

Consider also clearing the env var:

Suggested change
def _cleanup_prometheus_dir(self) -> None:
"""Remove the temporary prometheus multiprocess directory and its .db files."""
if self._prometheus_dir is None:
return
try:
shutil.rmtree(self._prometheus_dir)
logger.info("Cleaned up PROMETHEUS_MULTIPROC_DIR=%s", self._prometheus_dir)
except OSError as e:
logger.warning(
"Failed to clean up PROMETHEUS_MULTIPROC_DIR=%s: %s",
self._prometheus_dir,
e,
)
self._prometheus_dir = None
def _cleanup_prometheus_dir(self) -> None:
"""Remove the temporary prometheus multiprocess directory and its .db files."""
if self._prometheus_dir is None:
return
try:
shutil.rmtree(self._prometheus_dir)
logger.info("Cleaned up PROMETHEUS_MULTIPROC_DIR=%s", self._prometheus_dir)
except OSError as e:
logger.warning(
"Failed to clean up PROMETHEUS_MULTIPROC_DIR=%s: %s",
self._prometheus_dir,
e,
)
os.environ.pop("PROMETHEUS_MULTIPROC_DIR", None)
self._prometheus_dir = None

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Clean implementation for gRPC metrics collection via prometheus multiproc and graceful metrics server bind failure.

3 minor nits flagged — all 🟡:

  • Dead Ok(_) arm in the metrics match (unreachable due to earlier empty check)
  • No timeout on the python3 subprocess for metrics collection
  • PROMETHEUS_MULTIPROC_DIR env var not cleared on cleanup

None are blocking.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7f618a4642

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +116 to +117
if text.trim().is_empty() {
return Err("no metrics available from gRPC workers yet".to_string());

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Treat empty gRPC multiprocess output as a non-error

When the multiprocess scrape has no samples yet (a normal startup state), this branch returns an Err, but get_engine_metrics logs all Err results as warnings. That makes the "skip silently" Ok(_) branch unreachable and produces warning spam on every scrape until metrics appear. Returning an empty Ok (or a distinct non-warning outcome) would preserve the intended quiet behavior.

Useful? React with 👍 / 👎.

Comment on lines +57 to +61
#[expect(
clippy::disallowed_methods,
reason = "no-op task for graceful degradation"
)]
return tokio::spawn(async {});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Run Prometheus upkeep even on metrics bind failure

This early return bypasses the upkeep task started later in start_metrics_server, even though the recorder is already installed with an upkeep timeout. In the degraded path (port conflict), the router keeps serving traffic but never calls run_upkeep(), so histogram/metric maintenance is skipped for the lifetime of the process.

Useful? React with 👍 / 👎.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enhances metrics collection by integrating Prometheus multi-process metrics for gRPC workers. The Python serve module now manages a temporary directory for Prometheus metrics, ensuring proper setup and cleanup. The Rust WorkerManager has been updated to collect these gRPC metrics by executing a Python subprocess. Additionally, the metrics HTTP server startup in Rust is made more resilient by gracefully handling port binding failures. A critical issue was identified in the inline Python script used for gRPC metrics collection, where an indentation error would prevent the script from executing correctly, thus hindering metrics collection.

Comment on lines +97 to +102
"import sys\n\
from prometheus_client import CollectorRegistry, generate_latest\n\
from prometheus_client.multiprocess import MultiProcessCollector\n\
registry = CollectorRegistry()\n\
MultiProcessCollector(registry)\n\
sys.stdout.buffer.write(generate_latest(registry))\n",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The inline Python script has an indentation error. Each line of the script after the first one starts with a space, which will cause a Python IndentationError: unexpected indent when the script is executed. This will prevent metrics from being collected from gRPC workers.

To fix this, please remove the leading spaces from each line of the Python script within the string literal.

            "import sys\nfrom prometheus_client import CollectorRegistry, generate_latest\nfrom prometheus_client.multiprocess import MultiProcessCollector\nregistry = CollectorRegistry()\nMultiProcessCollector(registry)\nsys.stdout.buffer.write(generate_latest(registry))\n",

@mergify

mergify Bot commented Apr 9, 2026

Copy link
Copy Markdown
Contributor

Hi @ConnorLi96, this PR has merge conflicts that must be resolved before it can be merged. Please rebase your branch:

git fetch origin main
git rebase origin/main
# resolve any conflicts, then:
git push --force-with-lease

@mergify mergify Bot added the needs-rebase PR has merge conflicts that need to be resolved label Apr 9, 2026
Scott Lee and others added 5 commits April 14, 2026 14:46
Signed-off-by: Scott Lee <scott@together.ai>
Signed-off-by: Scott Lee <scott@together.ai>
When the gRPC worker hasn't written any .db files yet (startup phase or
no requests processed), generate_latest() returns empty bytes.  Passing
this empty string to parse_prometheus() triggers a parse error WARN log
on every scrape interval.

Add guards at both the collector function (return Err for empty output)
and the call site (skip empty text silently) to prevent log spam.

Signed-off-by: ConnorLi96 <ConnorLi96@users.noreply.github.com>
Made-with: Cursor
When the metrics port is already in use, log an error and return a no-op
handle instead of panicking, so the router can still operate.

Signed-off-by: ConnorLi96 <ConnorLi96@users.noreply.github.com>
Made-with: Cursor
The /metrics endpoint on the metrics port now returns both SMG's own
metrics (smg_*) and engine metrics (vllm_*, sglang_*, nv_trt_*) in a
single Prometheus text response. One scrape target per pod.

Engine metrics deps (WorkerRegistry + reqwest::Client) are registered
via OnceLock after AppContext init. Before registration, the handler
returns SMG metrics only.

No namespace collision — engine and SMG use distinct metric prefixes.
The /engine_metrics endpoint on the main port stays as-is.

Signed-off-by: ConnorLi96 <ConnorLi96@users.noreply.github.com>
Made-with: Cursor
@ConnorLi96
ConnorLi96 force-pushed the connorli/engine-metrics branch from 7f618a4 to 63eff15 Compare April 14, 2026 22:23
@mergify mergify Bot removed the needs-rebase PR has merge conflicts that need to be resolved label Apr 14, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 63eff15f22

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +23 to +25
use crate::worker::{
manager::{EngineMetricsResult, WorkerManager},
registry::WorkerRegistry,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Wire /metrics to the new gRPC multiprocess collector

This import still routes /metrics through crate::worker::manager::WorkerManager, but the new gRPC multiprocess logic was added in model_gateway/src/core/worker_manager.rs and is never wired into the crate (there is no core module export). As a result, prometheus_handler continues to call the old get_engine_metrics implementation (HTTP fan-out to .../metrics), which fails for grpc:// workers and leaves engine metrics absent in gRPC mode.

Useful? React with 👍 / 👎.

Comment on lines +323 to +334
for resp in responses {
if let Ok(r) = resp.result {
if r.status().is_success() {
if let Ok(text) = r.text().await {
metric_packs.push(MetricPack {
labels: vec![("worker_addr".into(), resp.url)],
metrics_text: text,
});
}
}
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: Three levels of silent error swallowing — network errors (resp.result), HTTP errors (r.status()), and body-read errors (r.text()) are all silently dropped. If a worker is consistently failing, the operator will see fewer metrics in the aggregated output but have no signal about which worker is down or why.

Consider adding at least a debug! log for failures so operators can diagnose missing metrics with RUST_LOG=debug:

Suggested change
for resp in responses {
if let Ok(r) = resp.result {
if r.status().is_success() {
if let Ok(text) = r.text().await {
metric_packs.push(MetricPack {
labels: vec![("worker_addr".into(), resp.url)],
metrics_text: text,
});
}
}
}
}
for resp in responses {
match resp.result {
Ok(r) if r.status().is_success() => {
if let Ok(text) = r.text().await {
metric_packs.push(MetricPack {
labels: vec![("worker_addr".into(), resp.url)],
metrics_text: text,
});
}
}
Ok(r) => {
debug!("Worker {} returned HTTP {} for /metrics", resp.url, r.status());
}
Err(e) => {
debug!("Failed to fetch /metrics from {}: {e}", resp.url);
}
}

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
bindings/python/src/smg/serve.py (1)

660-699: ⚠️ Potential issue | 🟡 Minor

Always clean up PROMETHEUS_MULTIPROC_DIR, even when no worker was recorded.

Line 662 returns before _cleanup_prometheus_dir() runs. If Line 607 created the temp directory and the first Popen fails before self.workers.append(...), the finally/atexit path leaks the directory.

Suggested fix
 def _cleanup_workers(self) -> None:
     """SIGTERM all worker process groups, wait, then SIGKILL stragglers."""
     if not self.workers:
+        self._cleanup_prometheus_dir()
         return
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@bindings/python/src/smg/serve.py` around lines 660 - 699, The early return in
_cleanup_workers prevents _cleanup_prometheus_dir from running when self.workers
is empty, leaking a previously created PROMETHEUS_MULTIPROC_DIR; update
_cleanup_workers so it always calls self._cleanup_prometheus_dir() (e.g., remove
the early return or call self._cleanup_prometheus_dir() before returning) so the
temporary directory is removed even if no worker was recorded, keeping the rest
of the SIGTERM/SIGKILL logic unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/core/worker_manager.rs`:
- Around line 114-119: The current code turns a cold-start empty scrape into
Err("no metrics available from gRPC workers yet"), causing repeated warning
logs; change the empty-output branch to return Ok(String::new()) instead of
Err(...) (i.e., when text.trim().is_empty() => Ok(String::new())), or
alternatively introduce a distinct sentinel error type and handle it silently at
the caller; update the callers that currently warn on any Err (the code that
iterates over the scraper results and logs warnings) to treat an empty String
(or the sentinel) as a non-error/no-op and skip logging so cold-start scrapes do
not spam warnings.
- Around line 556-595: The watch snapshot isn't pruned when all fetches fail
because the code returns early on group_loads.is_empty(); change
group_monitor_loop so that you always compute all_group_urls (from
workers.iter().map(|w| w.url().to_string())) and call tx.send_modify to remove
those URLs before the empty-check/continue, then only perform the
successful-path work (inserting group_loads into the map, calling
policy.update_loads and worker_load_manager.update_dp_loads) when group_loads is
non-empty; ensure worker_load_manager.update_dp_loads(&group_dp_loads) is not
called on failed ticks so the DP-rank cache remains unchanged.
- Around line 94-107: Wrap the tokio::process::Command invocation that currently
calls .output().await in a tokio::time::timeout using the existing
REQUEST_TIMEOUT constant: call tokio::time::timeout(REQUEST_TIMEOUT,
the_command.output()).await, then map the Timeout error to a descriptive Err
(e.g., "timed out running python3 prometheus collector") and unwrap the inner
Result from the command to preserve its original I/O error mapping (the existing
map_err("failed to run python3 prometheus collector: {e}") logic). Update the
code around the Command::new("python3") / .output().await block to perform this
two-level error handling so a hung subprocess or FS stall is bounded by
REQUEST_TIMEOUT.

---

Outside diff comments:
In `@bindings/python/src/smg/serve.py`:
- Around line 660-699: The early return in _cleanup_workers prevents
_cleanup_prometheus_dir from running when self.workers is empty, leaking a
previously created PROMETHEUS_MULTIPROC_DIR; update _cleanup_workers so it
always calls self._cleanup_prometheus_dir() (e.g., remove the early return or
call self._cleanup_prometheus_dir() before returning) so the temporary directory
is removed even if no worker was recorded, keeping the rest of the
SIGTERM/SIGKILL logic unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: da945273-c257-4551-b2f1-f6f482b42c26

📥 Commits

Reviewing files that changed from the base of the PR and between d8be56c and 63eff15.

📒 Files selected for processing (4)
  • bindings/python/src/smg/serve.py
  • model_gateway/src/core/worker_manager.rs
  • model_gateway/src/observability/metrics_server.rs
  • model_gateway/src/server.rs

Comment on lines +94 to +107
let output = tokio::process::Command::new("python3")
.args([
"-c",
"import sys\n\
from prometheus_client import CollectorRegistry, generate_latest\n\
from prometheus_client.multiprocess import MultiProcessCollector\n\
registry = CollectorRegistry()\n\
MultiProcessCollector(registry)\n\
sys.stdout.buffer.write(generate_latest(registry))\n",
])
.env("PROMETHEUS_MULTIPROC_DIR", &dir)
.output()
.await
.map_err(|e| format!("failed to run python3 prometheus collector: {e}"))?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

fd -type f -name "worker_manager.rs" | head -20

Repository: lightseekorg/smg

Length of output: 230


🏁 Script executed:

if [ -f "model_gateway/src/core/worker_manager.rs" ]; then
  wc -l "model_gateway/src/core/worker_manager.rs"
fi

Repository: lightseekorg/smg

Length of output: 104


🏁 Script executed:

head -120 "model_gateway/src/core/worker_manager.rs" | tail -40

Repository: lightseekorg/smg

Length of output: 1675


🏁 Script executed:

# Find the function containing the prometheus collector code
rg -A 5 "CollectorRegistry" model_gateway/src/core/worker_manager.rs

Repository: lightseekorg/smg

Length of output: 460


🏁 Script executed:

# Search for REQUEST_TIMEOUT or similar timeout constants
rg "REQUEST_TIMEOUT|timeout.*=" model_gateway/src/ | head -20

Repository: lightseekorg/smg

Length of output: 2022


🏁 Script executed:

# Check what function calls the Python subprocess and if it's on metrics path
rg -B 15 "python3.*prometheus_client" model_gateway/src/core/worker_manager.rs | head -50

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Find where collect_prometheus_multiproc_metrics is called
rg "collect_prometheus_multiproc_metrics" model_gateway/src/core/worker_manager.rs

Repository: lightseekorg/smg

Length of output: 200


🏁 Script executed:

# Get the context around the function definition to see line numbers
cat -n model_gateway/src/core/worker_manager.rs | grep -A 35 "collect_prometheus_multiproc_metrics"

Repository: lightseekorg/smg

Length of output: 3311


🏁 Script executed:

# Check if there are timeout patterns used elsewhere in this file
rg "tokio::time::timeout|Duration::from" model_gateway/src/core/worker_manager.rs

Repository: lightseekorg/smg

Length of output: 285


🏁 Script executed:

# Look for where this function is used and what calls it
rg -B 5 -A 5 "collect_prometheus_multiproc_metrics" model_gateway/src/

Repository: lightseekorg/smg

Length of output: 1739


🏁 Script executed:

# Get the function signature for the function containing line 338
sed -n '280,340p' model_gateway/src/core/worker_manager.rs | head -60

Repository: lightseekorg/smg

Length of output: 2100


🏁 Script executed:

# Get more context to understand the function name
sed -n '250,340p' model_gateway/src/core/worker_manager.rs

Repository: lightseekorg/smg

Length of output: 3148


🏁 Script executed:

# Look for "get_engine_metrics" or similar functions
rg "fn .*engine.*metric|pub.*fn.*metric" model_gateway/src/core/worker_manager.rs -A 3

Repository: lightseekorg/smg

Length of output: 203


Add a timeout to the Python collector subprocess.

The subprocess call at lines 94–107 uses unbounded .output().await, which can block metrics scrapes indefinitely if python3 or the filesystem hangs. The codebase already defines REQUEST_TIMEOUT (5 seconds) and applies it to similar async operations elsewhere in this file (e.g., HTTP requests); use the same pattern here.

Suggested fix
+    let output = tokio::time::timeout(
+        REQUEST_TIMEOUT,
+        tokio::process::Command::new("python3")
+            .args([
+                "-c",
+                "import sys\n\
+                 from prometheus_client import CollectorRegistry, generate_latest\n\
+                 from prometheus_client.multiprocess import MultiProcessCollector\n\
+                 registry = CollectorRegistry()\n\
+                 MultiProcessCollector(registry)\n\
+                 sys.stdout.buffer.write(generate_latest(registry))\n",
+            ])
+            .env("PROMETHEUS_MULTIPROC_DIR", &dir)
+            .output(),
+    )
+    .await
+    .map_err(|_| "python3 prometheus collector timed out".to_string())?
+    .map_err(|e| format!("failed to run python3 prometheus collector: {e}"))?;
-    let output = tokio::process::Command::new("python3")
-        .args([
-            "-c",
-            "import sys\n\
-             from prometheus_client import CollectorRegistry, generate_latest\n\
-             from prometheus_client.multiprocess import MultiProcessCollector\n\
-             registry = CollectorRegistry()\n\
-             MultiProcessCollector(registry)\n\
-             sys.stdout.buffer.write(generate_latest(registry))\n",
-        ])
-        .env("PROMETHEUS_MULTIPROC_DIR", &dir)
-        .output()
-        .await
-        .map_err(|e| format!("failed to run python3 prometheus collector: {e}"))?;
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
let output = tokio::process::Command::new("python3")
.args([
"-c",
"import sys\n\
from prometheus_client import CollectorRegistry, generate_latest\n\
from prometheus_client.multiprocess import MultiProcessCollector\n\
registry = CollectorRegistry()\n\
MultiProcessCollector(registry)\n\
sys.stdout.buffer.write(generate_latest(registry))\n",
])
.env("PROMETHEUS_MULTIPROC_DIR", &dir)
.output()
.await
.map_err(|e| format!("failed to run python3 prometheus collector: {e}"))?;
let output = tokio::time::timeout(
REQUEST_TIMEOUT,
tokio::process::Command::new("python3")
.args([
"-c",
"import sys\n\
from prometheus_client import CollectorRegistry, generate_latest\n\
from prometheus_client.multiprocess import MultiProcessCollector\n\
registry = CollectorRegistry()\n\
MultiProcessCollector(registry)\n\
sys.stdout.buffer.write(generate_latest(registry))\n",
])
.env("PROMETHEUS_MULTIPROC_DIR", &dir)
.output(),
)
.await
.map_err(|_| "python3 prometheus collector timed out".to_string())?
.map_err(|e| format!("failed to run python3 prometheus collector: {e}"))?;
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/core/worker_manager.rs` around lines 94 - 107, Wrap the
tokio::process::Command invocation that currently calls .output().await in a
tokio::time::timeout using the existing REQUEST_TIMEOUT constant: call
tokio::time::timeout(REQUEST_TIMEOUT, the_command.output()).await, then map the
Timeout error to a descriptive Err (e.g., "timed out running python3 prometheus
collector") and unwrap the inner Result from the command to preserve its
original I/O error mapping (the existing map_err("failed to run python3
prometheus collector: {e}") logic). Update the code around the
Command::new("python3") / .output().await block to perform this two-level error
handling so a hung subprocess or FS stall is bounded by REQUEST_TIMEOUT.

Comment on lines +114 to +119
let text = String::from_utf8(output.stdout)
.map_err(|e| format!("prometheus collector output is not valid UTF-8: {e}"))?;
if text.trim().is_empty() {
return Err("no metrics available from gRPC workers yet".to_string());
}
Ok(text)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

The expected empty-output case still logs once per scrape.

Lines 116-118 turn the cold-start “no metrics yet” state into Err(...), and Lines 348-351 warn on every Err. That replaces the old parse spam with warning spam. Either return Ok(String::new()) for the empty case or distinguish that sentinel error and skip it silently at the call site.

Also applies to: 337-352

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/core/worker_manager.rs` around lines 114 - 119, The current
code turns a cold-start empty scrape into Err("no metrics available from gRPC
workers yet"), causing repeated warning logs; change the empty-output branch to
return Ok(String::new()) instead of Err(...) (i.e., when text.trim().is_empty()
=> Ok(String::new())), or alternatively introduce a distinct sentinel error type
and handle it silently at the caller; update the callers that currently warn on
any Err (the code that iterates over the scraper results and logs warnings) to
treat an empty String (or the sentinel) as a non-error/no-op and skip logging so
cold-start scrapes do not spam warnings.

Comment on lines +556 to +595
let results = future::join_all(futures).await;

// Collect successful loads
let mut group_loads: HashMap<String, WorkerLoadResponse> = HashMap::new();
let mut group_dp_loads: HashMap<String, HashMap<isize, isize>> = HashMap::new();
for (url, response) in results {
if let Some(load) = response {
group_loads.insert(url.clone(), load.clone());
let dp_rank_loads = load.dp_rank_loads();
group_dp_loads.insert(url, dp_rank_loads);
}
}

if group_loads.is_empty() {
debug!("No loads fetched for group {group_key}");
continue;
}

debug!(
"Fetched loads from {}/{} workers in group {group_key}",
group_loads.len(),
workers.len()
);

// Update policies with this group's loads
for policy in &power_of_two_policies {
policy.update_loads(&group_loads);
}
worker_load_manager.update_dp_loads(&group_dp_loads);

// Atomically merge into the shared watch channel.
// Remove all group URLs first to clear stale entries from workers
// that failed this tick, then insert successful loads.
let all_group_urls: Vec<String> = workers.iter().map(|w| w.url().to_string()).collect();
tx.send_modify(|map| {
for url in &all_group_urls {
map.remove(url);
}
map.extend(group_loads);
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Prune the watch snapshot even when every fetch in a tick fails.

Line 569 exits before the group's URLs are removed from the shared watch map, so subscribers keep seeing stale per-worker loads after a full miss. The DP-rank cache can stay last-known-good, but the watch snapshot still needs to be cleared on failed ticks.

Suggested fix
-            let results = future::join_all(futures).await;
+            let results = future::join_all(futures).await;
+            let all_group_urls: Vec<String> = workers.iter().map(|w| w.url().to_string()).collect();

             // Collect successful loads
             let mut group_loads: HashMap<String, WorkerLoadResponse> = HashMap::new();
             let mut group_dp_loads: HashMap<String, HashMap<isize, isize>> = HashMap::new();
             for (url, response) in results {
@@
             }

             if group_loads.is_empty() {
+                tx.send_modify(|map| {
+                    for url in &all_group_urls {
+                        map.remove(url);
+                    }
+                });
                 debug!("No loads fetched for group {group_key}");
                 continue;
             }
@@
-            let all_group_urls: Vec<String> = workers.iter().map(|w| w.url().to_string()).collect();
             tx.send_modify(|map| {
                 for url in &all_group_urls {
                     map.remove(url);
                 }
                 map.extend(group_loads);

Based on learnings, group_monitor_loop should prune the watch snapshot for the group’s URLs on both successful and failed ticks, while leaving the DP-rank cache untouched on failed fetches.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
let results = future::join_all(futures).await;
// Collect successful loads
let mut group_loads: HashMap<String, WorkerLoadResponse> = HashMap::new();
let mut group_dp_loads: HashMap<String, HashMap<isize, isize>> = HashMap::new();
for (url, response) in results {
if let Some(load) = response {
group_loads.insert(url.clone(), load.clone());
let dp_rank_loads = load.dp_rank_loads();
group_dp_loads.insert(url, dp_rank_loads);
}
}
if group_loads.is_empty() {
debug!("No loads fetched for group {group_key}");
continue;
}
debug!(
"Fetched loads from {}/{} workers in group {group_key}",
group_loads.len(),
workers.len()
);
// Update policies with this group's loads
for policy in &power_of_two_policies {
policy.update_loads(&group_loads);
}
worker_load_manager.update_dp_loads(&group_dp_loads);
// Atomically merge into the shared watch channel.
// Remove all group URLs first to clear stale entries from workers
// that failed this tick, then insert successful loads.
let all_group_urls: Vec<String> = workers.iter().map(|w| w.url().to_string()).collect();
tx.send_modify(|map| {
for url in &all_group_urls {
map.remove(url);
}
map.extend(group_loads);
});
let results = future::join_all(futures).await;
let all_group_urls: Vec<String> = workers.iter().map(|w| w.url().to_string()).collect();
// Collect successful loads
let mut group_loads: HashMap<String, WorkerLoadResponse> = HashMap::new();
let mut group_dp_loads: HashMap<String, HashMap<isize, isize>> = HashMap::new();
for (url, response) in results {
if let Some(load) = response {
group_loads.insert(url.clone(), load.clone());
let dp_rank_loads = load.dp_rank_loads();
group_dp_loads.insert(url, dp_rank_loads);
}
}
if group_loads.is_empty() {
tx.send_modify(|map| {
for url in &all_group_urls {
map.remove(url);
}
});
debug!("No loads fetched for group {group_key}");
continue;
}
debug!(
"Fetched loads from {}/{} workers in group {group_key}",
group_loads.len(),
workers.len()
);
// Update policies with this group's loads
for policy in &power_of_two_policies {
policy.update_loads(&group_loads);
}
worker_load_manager.update_dp_loads(&group_dp_loads);
// Atomically merge into the shared watch channel.
// Remove all group URLs first to clear stale entries from workers
// that failed this tick, then insert successful loads.
tx.send_modify(|map| {
for url in &all_group_urls {
map.remove(url);
}
map.extend(group_loads);
});
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/core/worker_manager.rs` around lines 556 - 595, The watch
snapshot isn't pruned when all fetches fail because the code returns early on
group_loads.is_empty(); change group_monitor_loop so that you always compute
all_group_urls (from workers.iter().map(|w| w.url().to_string())) and call
tx.send_modify to remove those URLs before the empty-check/continue, then only
perform the successful-path work (inserting group_loads into the map, calling
policy.update_loads and worker_load_manager.update_dp_loads) when group_loads is
non-empty; ensure worker_load_manager.update_dp_loads(&group_dp_loads) is not
called on failed ticks so the DP-rank cache remains unchanged.

@github-actions github-actions Bot added grpc gRPC client and router changes openai OpenAI router changes labels Apr 15, 2026
error = %e,
"PD backend request failed"
);
error!("PD request transport error, both sides aborted: {e}");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: This push removed the comment that was here explaining why record_outcome is not called in this error path:

// Don't record_outcome here — the caller (execute_dual_dispatch)
// records outcomes from the response status after we return.

That comment documented a non-obvious design decision. Without it, a future reader might add a record_outcome call here (mirroring other error paths), which would cause double-counting with the caller. Consider restoring it above the return.

pub(super) fn get_server_info(registry: &WorkerRegistry) -> Response {
let stats = registry.stats();
let workers = external_workers(registry);
let workers = registry.get_all();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: This change removes the RuntimeType::External filter and drops the "external_workers" field from the /info JSON response. If any monitoring dashboards or scripts consume that field, they'll silently get null/missing. The behavioral change (health checks now cover all workers, not just external ones) seems intentional, but the field removal is a breaking API change worth noting in the PR description.

@ConnorLi96
ConnorLi96 force-pushed the connorli/engine-metrics branch from 018f984 to f4d1ecb Compare April 15, 2026 22:05
Add structured INFO logs for routing decisions and backend response
times to both HTTP and gRPC router paths.

Also fix health_generate to check all workers (not just External),
so gRPC-connected backends (TRT-LLM, SGLang, vLLM) are visible.

Signed-off-by: ConnorLi96 <ConnorLi96@users.noreply.github.com>
Made-with: Cursor
Signed-off-by: ConnorLi96 <ConnorLi96@users.noreply.github.com>
Made-with: Cursor
Signed-off-by: ConnorLi96 <ConnorLi96@users.noreply.github.com>
Made-with: Cursor
Comment on lines +391 to +419
async fn health_generate(&self, _req: Request<Body>) -> Response {
let workers = self.worker_registry.get_all();
if workers.is_empty() {
return (StatusCode::SERVICE_UNAVAILABLE, "No workers registered").into_response();
}
let (healthy, unhealthy): (Vec<_>, Vec<_>) = workers.iter().partition(|w| w.is_healthy());
if unhealthy.is_empty() {
(
StatusCode::OK,
format!("OK - {} workers healthy", healthy.len()),
)
.into_response()
} else {
let info: Vec<_> = unhealthy
.iter()
.map(|w| format!("{} ({})", w.model_id(), w.url()))
.collect();
(
StatusCode::SERVICE_UNAVAILABLE,
format!(
"{}/{} workers unhealthy: {}",
unhealthy.len(),
workers.len(),
info.join(", ")
),
)
.into_response()
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Important: This PD router health check doesn't verify that both worker types (prefill and decode) are present and healthy. get_all() returns all workers regardless of type, so if all prefill workers are down (or none are registered) but decode workers are healthy, this will report "OK - N workers healthy" even though the router can't actually serve any requests (it needs at least one of each).

Compare with the HTTP PD router (http/pd_router.rs:1213) which selects a PD pair and tests both sides, and with the Debug impl just above this method (lines 364-377) which already filters by WorkerType::Prefill / WorkerType::Decode.

Consider checking both types:

Suggested change
async fn health_generate(&self, _req: Request<Body>) -> Response {
let workers = self.worker_registry.get_all();
if workers.is_empty() {
return (StatusCode::SERVICE_UNAVAILABLE, "No workers registered").into_response();
}
let (healthy, unhealthy): (Vec<_>, Vec<_>) = workers.iter().partition(|w| w.is_healthy());
if unhealthy.is_empty() {
(
StatusCode::OK,
format!("OK - {} workers healthy", healthy.len()),
)
.into_response()
} else {
let info: Vec<_> = unhealthy
.iter()
.map(|w| format!("{} ({})", w.model_id(), w.url()))
.collect();
(
StatusCode::SERVICE_UNAVAILABLE,
format!(
"{}/{} workers unhealthy: {}",
unhealthy.len(),
workers.len(),
info.join(", ")
),
)
.into_response()
}
}
async fn health_generate(&self, _req: Request<Body>) -> Response {
let prefill_workers = self.worker_registry.get_workers_filtered(
None,
Some(WorkerType::Prefill),
Some(ConnectionMode::Grpc),
None,
false,
);
let decode_workers = self.worker_registry.get_workers_filtered(
None,
Some(WorkerType::Decode),
Some(ConnectionMode::Grpc),
None,
false,
);
if prefill_workers.is_empty() || decode_workers.is_empty() {
return (
StatusCode::SERVICE_UNAVAILABLE,
format!(
"Missing worker type: {} prefill, {} decode",
prefill_workers.len(),
decode_workers.len()
),
)
.into_response();
}
let workers = self.worker_registry.get_all();
let (healthy, unhealthy): (Vec<_>, Vec<_>) = workers.iter().partition(|w| w.is_healthy());
if unhealthy.is_empty() {
(
StatusCode::OK,
format!("OK - {} workers healthy", healthy.len()),
)
.into_response()
} else {
let info: Vec<_> = unhealthy
.iter()
.map(|w| format!("{} ({})", w.model_id(), w.url()))
.collect();
(
StatusCode::SERVICE_UNAVAILABLE,
format!(
"{}/{} workers unhealthy: {}",
unhealthy.len(),
workers.len(),
info.join(", ")
),
)
.into_response()
}
}

Comment on lines +514 to +542
async fn health_generate(&self, _req: Request<Body>) -> Response {
let workers = self.worker_registry.get_all();
if workers.is_empty() {
return (StatusCode::SERVICE_UNAVAILABLE, "No workers registered").into_response();
}
let (healthy, unhealthy): (Vec<_>, Vec<_>) = workers.iter().partition(|w| w.is_healthy());
if unhealthy.is_empty() {
(
StatusCode::OK,
format!("OK - {} workers healthy", healthy.len()),
)
.into_response()
} else {
let info: Vec<_> = unhealthy
.iter()
.map(|w| format!("{} ({})", w.model_id(), w.url()))
.collect();
(
StatusCode::SERVICE_UNAVAILABLE,
format!(
"{}/{} workers unhealthy: {}",
unhealthy.len(),
workers.len(),
info.join(", ")
),
)
.into_response()
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: This is byte-for-byte identical to GrpcPDRouter::health_generate. Consider extracting a shared helper (similar to openai/health.rs:health_generate) that both gRPC router impls can call, to avoid the two copies drifting apart over time.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f4d1ecb185

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

let smg_text = state.handle.render();

let engine_text = if let Some(deps) = ENGINE_METRICS_DEPS.get() {
match WorkerManager::get_engine_metrics(&deps.worker_registry, &deps.client).await {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Restrict engine metrics scrape to local workers

This handler now calls WorkerManager::get_engine_metrics on every /metrics scrape, and that helper fans out GET /metrics to all registered workers (model_gateway/src/worker/manager.rs, fan_out(&workers, ...)). In OpenAI/IGW deployments, those workers can be external provider URLs, so each Prometheus scrape generates outbound requests (with bearer auth when configured) to upstream APIs that do not expose engine metrics, creating avoidable external traffic and possible rate-limit/credential-exposure risk. Filter to worker types/runtime modes that actually export local engine metrics before invoking the fan-out.

Useful? React with 👍 / 👎.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/pd_router.rs`:
- Around line 392-417: The health check is using worker_registry.get_all() so
non-PD workers can affect PD health; modify the logic in this block to first
filter the workers list to only PD gRPC workers (e.g., those with the
PD/Prefill/Decode role or a protocol field indicating gRPC) before partitioning
by is_healthy(); keep the rest of the flow (partition, formatting using
is_healthy(), model_id(), url()) the same so only PD gRPC workers are counted
for OK vs SERVICE_UNAVAILABLE responses.

In `@model_gateway/src/routers/grpc/router.rs`:
- Around line 515-540: The health-check currently calls
self.worker_registry.get_all() and evaluates all workers; change it to only
consider gRPC-routable workers by filtering the collection before partitioning
so non-gRPC workers don't affect the gRPC router status. Update the code that
builds workers (from self.worker_registry.get_all()) to filter using the worker
predicate that indicates gRPC routability (e.g., a method like
is_grpc_routable() or supports_grpc() on the worker), then continue to use the
same partitioning (is_healthy()), model_id(), and url() logic on that filtered
list to produce the OK/503 response.

In `@model_gateway/src/routers/http/pd_router.rs`:
- Around line 647-669: The code currently logs "PD backend request completed" at
info! for any Ok(pd_result) even when prefill_resp or decode_resp is a 5xx;
change the Ok branch to check prefill_resp.status().is_server_error() ||
decode_resp.status().is_server_error() and, if true, emit warn! (include the
same fields: prefill_url = prefill.url(), decode_url = decode.url(),
prefill_status = prefill_resp.status().as_u16(), decode_status =
decode_resp.status().as_u16(), duration_ms = backend_duration_ms, streaming =
context.is_stream) with a message like "PD backend returned 5xx", otherwise keep
the info! "PD backend request completed"; locate this change in the match on
pd_result that assigns (prefill_resp, decode_resp).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: e057761d-8eaa-4979-8526-2cde7de3fd15

📥 Commits

Reviewing files that changed from the base of the PR and between 018f984 and f4d1ecb.

📒 Files selected for processing (7)
  • model_gateway/src/routers/grpc/common/stages/request_execution.rs
  • model_gateway/src/routers/grpc/common/stages/worker_selection.rs
  • model_gateway/src/routers/grpc/pd_router.rs
  • model_gateway/src/routers/grpc/router.rs
  • model_gateway/src/routers/http/pd_router.rs
  • model_gateway/src/routers/http/router.rs
  • model_gateway/src/routers/openai/health.rs

Comment on lines +392 to +417
let workers = self.worker_registry.get_all();
if workers.is_empty() {
return (StatusCode::SERVICE_UNAVAILABLE, "No workers registered").into_response();
}
let (healthy, unhealthy): (Vec<_>, Vec<_>) = workers.iter().partition(|w| w.is_healthy());
if unhealthy.is_empty() {
(
StatusCode::OK,
format!("OK - {} workers healthy", healthy.len()),
)
.into_response()
} else {
let info: Vec<_> = unhealthy
.iter()
.map(|w| format!("{} ({})", w.model_id(), w.url()))
.collect();
(
StatusCode::SERVICE_UNAVAILABLE,
format!(
"{}/{} workers unhealthy: {}",
unhealthy.len(),
workers.len(),
info.join(", ")
),
)
.into_response()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Scope PD health to PD gRPC workers only.

Line 392 uses get_all(), so unrelated workers can incorrectly flip PD health to 503. The PD health endpoint should evaluate only Prefill/Decode gRPC workers.

🔧 Proposed fix
-        let workers = self.worker_registry.get_all();
+        let prefill_workers = self.worker_registry.get_workers_filtered(
+            None,
+            Some(WorkerType::Prefill),
+            Some(ConnectionMode::Grpc),
+            None,
+            false,
+        );
+        let decode_workers = self.worker_registry.get_workers_filtered(
+            None,
+            Some(WorkerType::Decode),
+            Some(ConnectionMode::Grpc),
+            None,
+            false,
+        );
+        let workers: Vec<_> = prefill_workers
+            .into_iter()
+            .chain(decode_workers.into_iter())
+            .collect();
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/grpc/pd_router.rs` around lines 392 - 417, The
health check is using worker_registry.get_all() so non-PD workers can affect PD
health; modify the logic in this block to first filter the workers list to only
PD gRPC workers (e.g., those with the PD/Prefill/Decode role or a protocol field
indicating gRPC) before partitioning by is_healthy(); keep the rest of the flow
(partition, formatting using is_healthy(), model_id(), url()) the same so only
PD gRPC workers are counted for OK vs SERVICE_UNAVAILABLE responses.

Comment on lines +515 to +540
let workers = self.worker_registry.get_all();
if workers.is_empty() {
return (StatusCode::SERVICE_UNAVAILABLE, "No workers registered").into_response();
}
let (healthy, unhealthy): (Vec<_>, Vec<_>) = workers.iter().partition(|w| w.is_healthy());
if unhealthy.is_empty() {
(
StatusCode::OK,
format!("OK - {} workers healthy", healthy.len()),
)
.into_response()
} else {
let info: Vec<_> = unhealthy
.iter()
.map(|w| format!("{} ({})", w.model_id(), w.url()))
.collect();
(
StatusCode::SERVICE_UNAVAILABLE,
format!(
"{}/{} workers unhealthy: {}",
unhealthy.len(),
workers.len(),
info.join(", ")
),
)
.into_response()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Restrict gRPC health checks to gRPC-routable workers.

Line 515 currently inspects all registered workers. Unhealthy non-gRPC workers can incorrectly make this gRPC router report 503.

🔧 Proposed fix
-        let workers = self.worker_registry.get_all();
+        let workers = self.worker_registry.get_workers_filtered(
+            None,
+            None,
+            Some(crate::worker::ConnectionMode::Grpc),
+            None,
+            false,
+        );
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/grpc/router.rs` around lines 515 - 540, The
health-check currently calls self.worker_registry.get_all() and evaluates all
workers; change it to only consider gRPC-routable workers by filtering the
collection before partitioning so non-gRPC workers don't affect the gRPC router
status. Update the code that builds workers (from
self.worker_registry.get_all()) to filter using the worker predicate that
indicates gRPC routability (e.g., a method like is_grpc_routable() or
supports_grpc() on the worker), then continue to use the same partitioning
(is_healthy()), model_id(), and url() logic on that filtered list to produce the
OK/503 response.

Comment on lines 647 to +669
let (prefill_response, decode_response) = match pd_result {
Ok((prefill_resp, decode_resp)) => (prefill_resp, decode_resp),
Ok((prefill_resp, decode_resp)) => {
info!(
target: "smg::upstream",
prefill_url = prefill.url(),
decode_url = decode.url(),
prefill_status = prefill_resp.status().as_u16(),
decode_status = decode_resp.status().as_u16(),
duration_ms = backend_duration_ms,
streaming = context.is_stream,
"PD backend request completed"
);
(prefill_resp, decode_resp)
}
Err(e) => {
warn!(
target: "smg::upstream",
prefill_url = prefill.url(),
decode_url = decode.url(),
duration_ms = backend_duration_ms,
error = %e,
"PD backend request failed"
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Don't log PD 5xx responses as completed.

tokio::try_join! only tells us both HTTP sends returned a Response. If either prefill_resp.status() or decode_resp.status() is 5xx, this block still emits "PD backend request completed" at info!, even though the request is about to fail later in this method. That makes PD failure logs misleading and inconsistent with the regular HTTP router's 5xx warn! path in model_gateway/src/routers/http/router.rs Lines 350-373.

🔧 Suggested fix
-        let (prefill_response, decode_response) = match pd_result {
-            Ok((prefill_resp, decode_resp)) => {
-                info!(
-                    target: "smg::upstream",
-                    prefill_url = prefill.url(),
-                    decode_url = decode.url(),
-                    prefill_status = prefill_resp.status().as_u16(),
-                    decode_status = decode_resp.status().as_u16(),
-                    duration_ms = backend_duration_ms,
-                    streaming = context.is_stream,
-                    "PD backend request completed"
-                );
-                (prefill_resp, decode_resp)
-            }
+        let (prefill_response, decode_response) = match pd_result {
+            Ok((prefill_resp, decode_resp)) => {
+                let prefill_status = prefill_resp.status();
+                let decode_status = decode_resp.status();
+                let server_error =
+                    prefill_status.is_server_error() || decode_status.is_server_error();
+
+                if server_error {
+                    warn!(
+                        target: "smg::upstream",
+                        prefill_url = prefill.url(),
+                        decode_url = decode.url(),
+                        prefill_status = prefill_status.as_u16(),
+                        decode_status = decode_status.as_u16(),
+                        duration_ms = backend_duration_ms,
+                        streaming = context.is_stream,
+                        "PD backend request failed"
+                    );
+                } else {
+                    info!(
+                        target: "smg::upstream",
+                        prefill_url = prefill.url(),
+                        decode_url = decode.url(),
+                        prefill_status = prefill_status.as_u16(),
+                        decode_status = decode_status.as_u16(),
+                        duration_ms = backend_duration_ms,
+                        streaming = context.is_stream,
+                        "PD backend request completed"
+                    );
+                }
+
+                (prefill_resp, decode_resp)
+            }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/http/pd_router.rs` around lines 647 - 669, The code
currently logs "PD backend request completed" at info! for any Ok(pd_result)
even when prefill_resp or decode_resp is a 5xx; change the Ok branch to check
prefill_resp.status().is_server_error() ||
decode_resp.status().is_server_error() and, if true, emit warn! (include the
same fields: prefill_url = prefill.url(), decode_url = decode.url(),
prefill_status = prefill_resp.status().as_u16(), decode_status =
decode_resp.status().as_u16(), duration_ms = backend_duration_ms, streaming =
context.is_stream) with a message like "PD backend returned 5xx", otherwise keep
the info! "PD backend request completed"; locate this change in the match on
pd_result that assigns (prefill_resp, decode_resp).

…kers

Two fixes to make /metrics on :9900 include sglang engine metrics when the
router is run with --connection-mode grpc:

1. model_gateway: in WorkerManager::get_engine_metrics, split workers by
   connection mode. HTTP workers keep the existing fan_out path. gRPC
   workers are scraped at http://{host}:{grpc_port+1}/metrics, which is the
   sglang convention when --enable-metrics is set. Previously an HTTP GET
   was fired at the gRPC port and silently failed, so the entire engine
   metrics block ended up empty. Scrape errors are now logged instead of
   swallowed. Adds grpc_worker_metrics_url helper + unit tests covering
   grpc://, grpcs://, and @dp_rank suffix variants.

2. grpc_servicer.sglang.request_manager: gRPC mode launches only the
   scheduler process (no TokenizerManager), so TokenizerMetricsCollector
   was never initialized and request-level metrics (TTFT, e2e latency,
   inter-token latency, prompt/generation token totals) never appeared.
   GrpcRequestManager now creates its own TokenizerMetricsCollector when
   server_args.enable_metrics is set and observes TTFT / inter-token
   latency / e2e latency in _handle_batch_output, mirroring the logic in
   TokenizerManager.

Verified end-to-end against a sglang worker in gRPC mode: sglang_* now
appears on :9900/metrics alongside smg_*, and counters increment on
inference traffic (5 requests: num_requests_total 5 -> 16; TTFT and
e2e_latency histograms populated).

Made-with: Cursor
@mergify

mergify Bot commented Apr 20, 2026

Copy link
Copy Markdown
Contributor

Hi @ConnorLi96, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

let (host, port_and_suffix) = stripped.rsplit_once(':')?;
let port_str = port_and_suffix.split(['/', '@']).next()?;
let port: u16 = port_str.parse().ok()?;
Some(format!("http://{host}:{}/metrics", port + 1))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Important: port + 1 will panic in debug builds (or silently wrap to 0 in release) when the gRPC worker is on port 65535. While 65535 is uncommon for gRPC, u16 arithmetic overflow is undefined behavior in safe Rust — it's a panic in debug and a silent wrap in release. Use checked_add so the function returns None (consistent with the other parse-failure paths) instead of panicking.

Suggested change
Some(format!("http://{host}:{}/metrics", port + 1))
let port: u16 = port_str.parse().ok()?;
let metrics_port = port.checked_add(1)?;
Some(format!("http://{host}:{metrics_port}/metrics"))

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0f93a4cfa6

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +844 to +845
for worker in &grpc_workers {
if let Some(metrics_url) = grpc_worker_metrics_url(worker.url()) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Scrape gRPC metrics in parallel

This loop processes gRPC workers serially, and each iteration awaits an HTTP scrape with REQUEST_TIMEOUT (5s). If one or more worker metrics endpoints are slow/unreachable (a common startup or partial-failure case), total /metrics latency scales with worker count and can exceed Prometheus scrape timeouts, causing the whole scrape to fail. The gRPC branch should fan out requests concurrently like the HTTP branch.

Useful? React with 👍 / 👎.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
model_gateway/src/worker/manager.rs (1)

842-878: ⚠️ Potential issue | 🟠 Major

Scrape gRPC workers in parallel — sequential await will block /metrics.

The HTTP branch uses fan_out with buffer_unordered(MAX_CONCURRENT), but this loop awaits each gRPC worker serially with a 5s REQUEST_TIMEOUT. With N gRPC workers where some are slow/unreachable, total latency approaches N * 5s, which will trivially exceed a typical Prometheus scrape timeout (10s default) once you have more than ~2 unreachable workers, and will also stall prometheus_handler in observability/metrics_server.rs. Fan these out the same way as the HTTP path.

♻️ Suggested fix: run gRPC scrapes concurrently
-        // gRPC workers expose Prometheus metrics on a separate HTTP port
-        // (gRPC port + 1) when --enable-metrics is passed to sglang.
-        for worker in &grpc_workers {
-            if let Some(metrics_url) = grpc_worker_metrics_url(worker.url()) {
-                match client
-                    .get(&metrics_url)
-                    .timeout(REQUEST_TIMEOUT)
-                    .send()
-                    .await
-                {
-                    Ok(r) if r.status().is_success() => {
-                        if let Ok(text) = r.text().await {
-                            metric_packs.push(MetricPack {
-                                labels: vec![(
-                                    "worker_addr".into(),
-                                    worker.url().to_string(),
-                                )],
-                                metrics_text: text,
-                            });
-                        }
-                    }
-                    Ok(r) => {
-                        warn!(
-                            "gRPC worker metrics endpoint {} returned {}",
-                            metrics_url,
-                            r.status()
-                        );
-                    }
-                    Err(e) => {
-                        warn!(
-                            "Failed to scrape gRPC worker metrics from {}: {e}",
-                            metrics_url
-                        );
-                    }
-                }
-            }
-        }
+        // gRPC workers expose Prometheus metrics on a separate HTTP port
+        // (gRPC port + 1) when --enable-metrics is passed to sglang.
+        let grpc_futures = grpc_workers.iter().filter_map(|worker| {
+            let metrics_url = grpc_worker_metrics_url(worker.url())?;
+            let client = client.clone();
+            let worker_url = worker.url().to_string();
+            Some(async move {
+                let result = client
+                    .get(&metrics_url)
+                    .timeout(REQUEST_TIMEOUT)
+                    .send()
+                    .await;
+                (worker_url, metrics_url, result)
+            })
+        });
+        let grpc_results: Vec<_> = stream::iter(grpc_futures)
+            .buffer_unordered(MAX_CONCURRENT)
+            .collect()
+            .await;
+        for (worker_url, metrics_url, result) in grpc_results {
+            match result {
+                Ok(r) if r.status().is_success() => {
+                    if let Ok(text) = r.text().await {
+                        metric_packs.push(MetricPack {
+                            labels: vec![("worker_addr".into(), worker_url)],
+                            metrics_text: text,
+                        });
+                    }
+                }
+                Ok(r) => warn!(
+                    "gRPC worker metrics endpoint {} returned {}",
+                    metrics_url,
+                    r.status()
+                ),
+                Err(e) => warn!(
+                    "Failed to scrape gRPC worker metrics from {metrics_url}: {e}"
+                ),
+            }
+        }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/worker/manager.rs` around lines 842 - 878, The gRPC
scraping loop currently awaits each request sequentially causing high total
latency; replace the for-loop over grpc_workers with the same fan_out pattern
used by the HTTP branch: build a stream over grpc_workers (using
futures::stream::iter or similar), map each worker into an async closure that
calls grpc_worker_metrics_url(worker.url()), does the
client.get(...).timeout(REQUEST_TIMEOUT).send().await, processes the response
into an Option<MetricPack> (using grpc_worker_metrics_url, client,
REQUEST_TIMEOUT, MetricPack), then use .buffer_unordered(MAX_CONCURRENT) to run
them concurrently and collect/for_each the resulting MetricPack options into
metric_packs; ensure you preserve the existing log paths (warn on non-success
and Err) and avoid shared-mutable push races by returning Option<MetricPack>
from each future and collecting them into metric_packs after the concurrent
execution.
grpc_servicer/smg_grpc_servicer/sglang/request_manager.py (1)

804-863: ⚠️ Potential issue | 🟡 Minor

Aborted/preempted requests are not reflected in finished-request metrics.

observe_one_finished_request is only emitted in _handle_batch_output on the success-finish branch (L718-736). Requests that terminate via _handle_abort_req (explicit client abort, scheduler-side priority preemption, queue-full, KV-cache pressure) set state.finished = True and push an abort response to the queue, but never record an end-to-end finished observation. Depending on how downstream dashboards compare "started" vs. "finished" counts this will look like a leak, and abort-path latencies will be invisible.

If the intent is "only count successful completions", consider at minimum emitting a distinct counter for aborted/preempted terminations (e.g., with a finish_reason label) so that totals reconcile. Otherwise, mirror the finished-request observation in _handle_abort_req with appropriate token counts (0/cumulative from state).

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@grpc_servicer/smg_grpc_servicer/sglang/request_manager.py` around lines 804 -
863, The abort path in _handle_abort_req currently marks state.finished and
pushes an abort_response but does not emit the finished-request metric
(observe_one_finished_request) used in _handle_batch_output, so
aborted/preempted requests are missing from finished-request metrics; update
_handle_abort_req to call the same observation routine (or emit a distinct
counter) immediately after marking the state finished and before/after putting
to state.out_queue: include request_id/state.rid, finish_reason (use
recv_obj.finished_reason or a synthesized {"type":"abort","message":"Abort
before prefill"}), and token counts (use state.cumulative_prompt_tokens and
state.cumulative_completion_tokens if available, else 0) so dashboards reconcile
started vs finished counts; reference symbols: _handle_abort_req,
_handle_batch_output, observe_one_finished_request,
state.cumulative_prompt_tokens, state.cumulative_completion_tokens,
rid_to_state.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/sglang/request_manager.py`:
- Around line 209-235: Initialize self._metrics_labels = {} before the try so
it's always defined, then in the try continue to assign labels and create
TokenizerMetricsCollector (using server_args, labels,
bucket_time_to_first_token, bucket_e2e_request_latency, and optional
bucket_inter_token_latency) and set self.metrics_collector and
self._metrics_labels on success; replace the broad except Exception: with
narrower handling—catch ImportError (log a warning with exc_info) and catch
TypeError (or other specific runtime/API-mismatch errors) separately so you log
the exception type and message via logger.warning, ensuring programming errors
aren't silently swallowed while preserving graceful degradation when
observability imports or APIs are missing.

In `@model_gateway/src/worker/manager.rs`:
- Around line 895-903: The function grpc_worker_metrics_url can overflow when
port == 65535; change its logic in grpc_worker_metrics_url to use a checked
addition (e.g., port.checked_add(1)) or explicitly return None if port == 65535
before formatting, so you never wrap to 0; update the code path that computes
port (port: u16 = port_str.parse().ok()?) to handle this check and return None
on overflow, and add a unit test asserting
grpc_worker_metrics_url("grpc://host:65535") -> None.

---

Outside diff comments:
In `@grpc_servicer/smg_grpc_servicer/sglang/request_manager.py`:
- Around line 804-863: The abort path in _handle_abort_req currently marks
state.finished and pushes an abort_response but does not emit the
finished-request metric (observe_one_finished_request) used in
_handle_batch_output, so aborted/preempted requests are missing from
finished-request metrics; update _handle_abort_req to call the same observation
routine (or emit a distinct counter) immediately after marking the state
finished and before/after putting to state.out_queue: include
request_id/state.rid, finish_reason (use recv_obj.finished_reason or a
synthesized {"type":"abort","message":"Abort before prefill"}), and token counts
(use state.cumulative_prompt_tokens and state.cumulative_completion_tokens if
available, else 0) so dashboards reconcile started vs finished counts; reference
symbols: _handle_abort_req, _handle_batch_output, observe_one_finished_request,
state.cumulative_prompt_tokens, state.cumulative_completion_tokens,
rid_to_state.

In `@model_gateway/src/worker/manager.rs`:
- Around line 842-878: The gRPC scraping loop currently awaits each request
sequentially causing high total latency; replace the for-loop over grpc_workers
with the same fan_out pattern used by the HTTP branch: build a stream over
grpc_workers (using futures::stream::iter or similar), map each worker into an
async closure that calls grpc_worker_metrics_url(worker.url()), does the
client.get(...).timeout(REQUEST_TIMEOUT).send().await, processes the response
into an Option<MetricPack> (using grpc_worker_metrics_url, client,
REQUEST_TIMEOUT, MetricPack), then use .buffer_unordered(MAX_CONCURRENT) to run
them concurrently and collect/for_each the resulting MetricPack options into
metric_packs; ensure you preserve the existing log paths (warn on non-success
and Err) and avoid shared-mutable push races by returning Option<MetricPack>
from each future and collecting them into metric_packs after the concurrent
execution.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 51056492-142c-476f-b97c-0ab020b8708a

📥 Commits

Reviewing files that changed from the base of the PR and between f4d1ecb and 0f93a4c.

📒 Files selected for processing (2)
  • grpc_servicer/smg_grpc_servicer/sglang/request_manager.py
  • model_gateway/src/worker/manager.rs

Comment on lines +209 to +235
self.metrics_collector = None
if server_args.enable_metrics:
try:
from sglang.srt.observability.metrics_collector import (
TokenizerMetricsCollector,
)

labels = {
"model_name": server_args.served_model_name,
}
self.metrics_collector = TokenizerMetricsCollector(
server_args=server_args,
labels=labels,
bucket_time_to_first_token=server_args.bucket_time_to_first_token,
bucket_e2e_request_latency=server_args.bucket_e2e_request_latency,
bucket_inter_token_latency=getattr(
server_args, "bucket_inter_token_latency", None
),
)
self._metrics_labels = labels
logger.info("TokenizerMetricsCollector initialized for gRPC request-level metrics")
except Exception:
logger.warning(
"Failed to initialize TokenizerMetricsCollector, "
"request-level metrics will be unavailable",
exc_info=True,
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Metrics init: minor hardening — _metrics_labels default and narrower except.

Two small points:

  1. self._metrics_labels is only assigned inside the successful try branch. All current read sites are gated by self.metrics_collector is not None, so this is safe today, but any future reader that forgets the guard will hit AttributeError. Defining self._metrics_labels = {} before the try makes the invariant robust to refactors.
  2. except Exception: swallows everything including KeyboardInterrupt subclasses (not in Py3, but) and, more importantly, masks programming errors like TypeError from passing an unexpected kwarg (e.g., bucket_inter_token_latency not being accepted by older sglang versions). Consider logging the exception type at least, or catching ImportError + a narrower runtime error separately so API-mismatch bugs surface during development.
Proposed diff
         self.metrics_collector = None
+        self._metrics_labels: dict[str, str] = {}
         if server_args.enable_metrics:
             try:
                 from sglang.srt.observability.metrics_collector import (
                     TokenizerMetricsCollector,
                 )
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@grpc_servicer/smg_grpc_servicer/sglang/request_manager.py` around lines 209 -
235, Initialize self._metrics_labels = {} before the try so it's always defined,
then in the try continue to assign labels and create TokenizerMetricsCollector
(using server_args, labels, bucket_time_to_first_token,
bucket_e2e_request_latency, and optional bucket_inter_token_latency) and set
self.metrics_collector and self._metrics_labels on success; replace the broad
except Exception: with narrower handling—catch ImportError (log a warning with
exc_info) and catch TypeError (or other specific runtime/API-mismatch errors)
separately so you log the exception type and message via logger.warning,
ensuring programming errors aren't silently swallowed while preserving graceful
degradation when observability imports or APIs are missing.

Comment on lines +895 to +903
fn grpc_worker_metrics_url(worker_url: &str) -> Option<String> {
let stripped = worker_url
.strip_prefix("grpc://")
.or_else(|| worker_url.strip_prefix("grpcs://"))?;
let (host, port_and_suffix) = stripped.rsplit_once(':')?;
let port_str = port_and_suffix.split(['/', '@']).next()?;
let port: u16 = port_str.parse().ok()?;
Some(format!("http://{host}:{}/metrics", port + 1))
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Guard against port + 1 overflow on u16.

If port_str parses to 65535, port + 1 panics in debug and wraps to 0 in release, producing a bogus http://host:0/metrics. Unlikely in practice, but trivial to harden and avoids a hidden footgun if someone ever configures a high ephemeral port.

🛡️ Proposed fix
-    let port: u16 = port_str.parse().ok()?;
-    Some(format!("http://{host}:{}/metrics", port + 1))
+    let port: u16 = port_str.parse().ok()?;
+    let metrics_port = port.checked_add(1)?;
+    Some(format!("http://{host}:{metrics_port}/metrics"))

Optionally add a test case asserting grpc_worker_metrics_url("grpc://host:65535") returns None.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
fn grpc_worker_metrics_url(worker_url: &str) -> Option<String> {
let stripped = worker_url
.strip_prefix("grpc://")
.or_else(|| worker_url.strip_prefix("grpcs://"))?;
let (host, port_and_suffix) = stripped.rsplit_once(':')?;
let port_str = port_and_suffix.split(['/', '@']).next()?;
let port: u16 = port_str.parse().ok()?;
Some(format!("http://{host}:{}/metrics", port + 1))
}
fn grpc_worker_metrics_url(worker_url: &str) -> Option<String> {
let stripped = worker_url
.strip_prefix("grpc://")
.or_else(|| worker_url.strip_prefix("grpcs://"))?;
let (host, port_and_suffix) = stripped.rsplit_once(':')?;
let port_str = port_and_suffix.split(['/', '@']).next()?;
let port: u16 = port_str.parse().ok()?;
let metrics_port = port.checked_add(1)?;
Some(format!("http://{host}:{metrics_port}/metrics"))
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/worker/manager.rs` around lines 895 - 903, The function
grpc_worker_metrics_url can overflow when port == 65535; change its logic in
grpc_worker_metrics_url to use a checked addition (e.g., port.checked_add(1)) or
explicitly return None if port == 65535 before formatting, so you never wrap to
0; update the code path that computes port (port: u16 = port_str.parse().ok()?)
to handle this check and return None on overflow, and add a unit test asserting
grpc_worker_metrics_url("grpc://host:65535") -> None.

@github-actions github-actions Bot added the tests Test changes label Apr 22, 2026
@mergify

mergify Bot commented Apr 22, 2026

Copy link
Copy Markdown
Contributor

Hi @ConnorLi96, this PR has merge conflicts that must be resolved before it can be merged. Please rebase your branch:

git fetch origin main
git rebase origin/main
# resolve any conflicts, then:
git push --force-with-lease

@mergify mergify Bot added the needs-rebase PR has merge conflicts that need to be resolved label Apr 22, 2026
…cs endpoint

Option A (sentinel round-trip): encode ':' to '__smgcolon48f__' before
passing text to openmetrics_parser (which rejects colons), then decode
the sentinel back to ':' after serialization. This preserves colons in
both metric names (sglang:num_running_reqs) and label values
(grpc://host:9001).

Option A chosen over B (passthrough-when-no-labels) because labels are
always injected by get_engine_metrics, making B a no-op. Option C
(raw text injection) was rejected as the largest diff with full
Prometheus syntax responsibility.

The sentinel uses only lowercase to avoid conflicts with
openmetrics_parser's case-sensitive keyword grammar (HELP/TYPE).

Before: sglang_num_running_reqs (colon clobbered to underscore at
metrics_aggregator.rs:19 via .replace(":", "_"))
After:  sglang:num_running_reqs (original Prometheus namespace:name
format preserved)

Tracks: consolidation-doc §3.A-ish — colon preservation follow-up to #873
Signed-off-by: Connor Li <ConnorLi96@users.noreply.github.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/worker/metrics_aggregator.rs`:
- Line 23: The current replace call (let encoded = metrics_text.replace(':',
COLON_SENTINEL);) is not input-safe because an input containing the literal
COLON_SENTINEL will be incorrectly decoded; fix by making the transformation
reversible: either (A) pre-escape any existing COLON_SENTINEL occurrences in
metrics_text (e.g., replace COLON_SENTINEL with an escaped variant) before
replacing ':' so round-trip decode distinguishes originals from encoded colons,
or (B) switch to a truly reversible encoding (e.g., percent-encoding or base64)
for metrics_text; alternatively, if you accept the negligible risk, add a clear
comment next to the COLON_SENTINEL constant and the use in encoded/replace
explaining the trust assumption and why collisions are acceptable. Ensure
references to metrics_text, COLON_SENTINEL, and the encoded variable are updated
accordingly.
- Around line 13-16: The comment above the COLON_SENTINEL constant uses the
`SAFETY:` marker but this is a safe-code invariant; change the comment marker to
`INVARIANT:` and keep the rest of the explanation verbatim so it documents the
round-trip assumption about encoding/decoding colons for `namespace:metric_name`
while reserving `SAFETY:` for true unsafe blocks; update the comment that
precedes `const COLON_SENTINEL: &str = "__smgcolon48f__";` accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: e6539697-1c95-47bc-875f-288083279259

📥 Commits

Reviewing files that changed from the base of the PR and between 0f93a4c and 3f6eec0.

📒 Files selected for processing (2)
  • model_gateway/src/worker/metrics_aggregator.rs
  • model_gateway/tests/metrics_aggregator_test.rs

Comment on lines +13 to +16
// SAFETY: openmetrics_parser rejects colons in metric names. We encode colons to
// this placeholder before parsing and decode back after serialization, preserving
// the original `namespace:metric_name` colon format in the output.
const COLON_SENTINEL: &str = "__smgcolon48f__";

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Use INVARIANT: instead of SAFETY: for this safe-code assumption.

Per repo convention, SAFETY: is reserved for unsafe blocks; safe-code assumptions should be documented with INVARIANT:. There is no unsafe here — this is a round-trip invariant over the input text.

✏️ Proposed change
-// SAFETY: openmetrics_parser rejects colons in metric names. We encode colons to
-// this placeholder before parsing and decode back after serialization, preserving
-// the original `namespace:metric_name` colon format in the output.
+// INVARIANT: openmetrics_parser rejects colons in metric names. We encode colons to
+// this placeholder before parsing and decode back after serialization, preserving
+// the original `namespace:metric_name` colon format in the output. The sentinel is
+// chosen to be extremely unlikely to collide with real metric text.
 const COLON_SENTINEL: &str = "__smgcolon48f__";

As per learnings, use the marker INVARIANT: to document assumptions in safe code; reserve SAFETY: for explaining why unsafe blocks are sound.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/worker/metrics_aggregator.rs` around lines 13 - 16, The
comment above the COLON_SENTINEL constant uses the `SAFETY:` marker but this is
a safe-code invariant; change the comment marker to `INVARIANT:` and keep the
rest of the explanation verbatim so it documents the round-trip assumption about
encoding/decoding colons for `namespace:metric_name` while reserving `SAFETY:`
for true unsafe blocks; update the comment that precedes `const COLON_SENTINEL:
&str = "__smgcolon48f__";` accordingly.

let metrics_text = &metric_pack.metrics_text;
// openmetrics_parser doesn't handle colons in metric names; replace with underscores
let metrics_text = metrics_text.replace(":", "_");
let encoded = metrics_text.replace(':', COLON_SENTINEL);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Sentinel round-trip is robust in practice but not input-safe.

metrics_text.replace(':', COLON_SENTINEL) followed by a reverse replace(COLON_SENTINEL, ":") will misbehave if the literal string __smgcolon48f__ ever appears in input metric text (e.g., inside a label value) — it would be decoded to : in the output. The probability is negligible given the sentinel shape, and upstream sources are trusted Prometheus exposition, so this is a minor caveat rather than a bug. Consider noting this assumption alongside the constant if you plan to take metric text from less-trusted sources later.

Also applies to: 40-40

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/worker/metrics_aggregator.rs` at line 23, The current
replace call (let encoded = metrics_text.replace(':', COLON_SENTINEL);) is not
input-safe because an input containing the literal COLON_SENTINEL will be
incorrectly decoded; fix by making the transformation reversible: either (A)
pre-escape any existing COLON_SENTINEL occurrences in metrics_text (e.g.,
replace COLON_SENTINEL with an escaped variant) before replacing ':' so
round-trip decode distinguishes originals from encoded colons, or (B) switch to
a truly reversible encoding (e.g., percent-encoding or base64) for metrics_text;
alternatively, if you accept the negligible risk, add a clear comment next to
the COLON_SENTINEL constant and the use in encoded/replace explaining the trust
assumption and why collisions are acceptable. Ensure references to metrics_text,
COLON_SENTINEL, and the encoded variable are updated accordingly.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3f6eec0ebd

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

let (host, port_and_suffix) = stripped.rsplit_once(':')?;
let port_str = port_and_suffix.split(['/', '@']).next()?;
let port: u16 = port_str.parse().ok()?;
Some(format!("http://{host}:{}/metrics", port + 1))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Guard gRPC metrics port increment against overflow

When a worker is configured on port 65535, port + 1 overflows in grpc_worker_metrics_url. In debug builds this can panic during /metrics scraping, and in release builds it wraps to 0, producing an invalid metrics URL and silently dropping that worker’s engine metrics. Using a checked increment (and returning None on overflow) avoids both crash and silent mis-scrape behavior.

Useful? React with 👍 / 👎.

Signed-off-by: Connor Li <ConnorLi96@users.noreply.github.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9a9b6e0329

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +43 to +46
let _ = ENGINE_METRICS_DEPS.set(EngineMetricsDeps {
worker_registry,
client,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Allow engine metrics deps to be refreshed per startup

register_engine_metrics_deps writes into a process-global OnceLock and ignores set failures, so only the first AppContext ever becomes visible to /metrics. In embedded/test scenarios that start the router more than once in the same process, subsequent servers will scrape engine metrics from stale registry/client state (and keep old Arcs alive) instead of their own workers. Use an updatable global (e.g., RwLock<Option<_>>) or explicit reset path rather than a one-shot lock here.

Useful? React with 👍 / 👎.

@github-actions

github-actions Bot commented May 6, 2026

Copy link
Copy Markdown

This pull request has been automatically marked as stale because it has not had any activity within 14 days. It will be automatically closed if no further activity occurs within 16 days. Leave a comment if you feel this pull request should remain open. Thank you!

@github-actions github-actions Bot added the stale PR has been inactive for 14+ days label May 6, 2026
@github-actions

Copy link
Copy Markdown

This pull request has been automatically closed due to inactivity. Please feel free to reopen if you intend to continue working on it. Thank you!

@github-actions github-actions Bot closed this May 22, 2026
@lightseek-bot
lightseek-bot deleted the connorli/engine-metrics branch May 23, 2026 00:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes needs-rebase PR has merge conflicts that need to be resolved openai OpenAI router changes python-bindings Python bindings changes stale PR has been inactive for 14+ days tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant