Skip to content

fix(worker): query engine-specific HTTP load endpoint for vLLM/SGLang - #1867

Merged
slin1237 merged 2 commits into
mainfrom
xz/http-load-engine-aware
Jul 3, 2026
Merged

slin1237 merged 2 commits into
mainfrom
xz/http-load-engine-aware

Conversation

@XinyueZhang369

@XinyueZhang369 XinyueZhang369 commented Jul 2, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

/get_loads reported load: -1 for every HTTP worker, and load-aware routing (power_of_two / cache_aware) had no signal to act on.

WorkerMonitor::fetch_http_load unconditionally called GET /v1/loads?include=…. Verified against the dev ORD cross-region-router fleet, stock vLLM v0.7.3 and SGLang images 404 that path — it's an SGLang-custom/mock endpoint. So every fetch failed and load fell back to -1.

Change

Dispatch fetch_http_load on runtime_type and query the endpoint each engine actually serves:

Runtime Load source
vLLM GET /metrics → vllm:gpu_cache_usage_perc → token_usage, num_requests_running/waiting
SGLang GET /v1/loads first (custom builds), fall back to GET /metrics (sglang:token_usage, num_running_reqs, num_queue_reqs, …)
Unspecified / other unchanged GET /v1/loads (mock worker, custom engines)

A small PromScrape parser reads the flat Prometheus gauge lines (sum for counts, mean for ratios). Both metric fetchers require the KV-usage gauge and return None otherwise, so token_usage is never a fake 0.0.

Ratio vs. absolute tokens

vLLM/SGLang /metrics expose the KV-usage ratio, not an absolute used-token count. Metric-derived snapshots therefore carry token_usage but leave num_used_tokens/max_total_num_tokens at 0. New WorkerLoadResponse::has_absolute_token_data() (any rank with max_total_num_tokens > 0) gates the two consumers that assume absolute tokens:

  • /get_loads scalar reports -1 (unavailable) instead of a misleading 0; the ratio is still available in details.loads[].token_usage.
  • DP-rank cache is fed only real per-rank tokens, and workers that degrade from /v1/loads to ratio-only are evicted so stale per-rank loads can't keep driving MinimumTokensPolicy.

power_of_two / cache_aware are unaffected — they read token_usage from the watch-channel group loads.

Testing

  • cargo test -p smg --lib worker::monitor → 17 pass (6 new: parser, engine mappings, absolute-token discriminator).
  • cargo clippy -p smg -p openai-protocol --lib -- -D warnings → clean.
  • Endpoint behavior verified against live vLLM/SGLang workers in dev ORD cross-region-router.

Notes / follow-ups

  • vLLM's absolute used-tokens is derivable from vllm:cache_config_info labels (num_gpu_blocks × block_size × ratio) — deliberately not included here to keep the parser simple and avoid version-fragile label parsing. Could populate the /get_loads absolute scalar for vLLM in a follow-up.
  • Lexicographic (ratio → waiting) routing tie-break for saturated workers is out of scope for this PR.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Improved worker load reporting to better handle different engine/runtime metric formats.
    • Added support for recognizing when absolute token counts are actually available.
  • Bug Fixes

    • Worker load values now avoid showing misleading token counts when only ratio-based metrics are present.
    • Stale worker load data is now cleared more reliably when absolute token data is unavailable.

@coderabbitai

coderabbitai Bot commented Jul 2, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds a has_absolute_token_data predicate to WorkerLoadResponse that checks for positive max_total_num_tokens. This gates load computation in get_all_worker_loads and DP-rank cache updates in WorkerMonitor. Also reworks HTTP load fetching to dispatch by runtime type, scraping Prometheus /metrics for vLLM/SGLang via a new PromScrape parser, with centralized authenticated GET requests and new tests.

Changes

Absolute Token Data Gating and Prometheus Load Scraping

Layer / File(s) Summary
Absolute token data predicate
crates/protocols/src/worker.rs
Adds has_absolute_token_data() to WorkerLoadResponse, checking if any DP-rank snapshot has max_total_num_tokens > 0.
Load calculation gating
model_gateway/src/worker/manager.rs
get_all_worker_loads now only derives load from total_used_tokens() when absolute token data is present, else defaults to -1.
Runtime-dispatched Prometheus load fetching
model_gateway/src/worker/monitor.rs
Adds PromScrape parser and sum/mean/has helpers; fetch_http_load dispatches to vLLM/SGLang Prometheus scraping or native JSON endpoint based on RuntimeType, centralizing requests in authed_get.
DP-rank cache gating and eviction
model_gateway/src/worker/monitor.rs
DP-rank cache insertion is now conditional on has_absolute_token_data(); otherwise the worker URL is queued for eviction and removed after update_dp_loads.
PromScrape and absolute-token tests
model_gateway/src/worker/monitor.rs
Adds prom_scrape_tests validating parsing, aggregation, metric-name compatibility, and absolute-token detection behavior.

Estimated code review effort: 3 (Moderate) | ~30 minutes

Sequence Diagram(s)

sequenceDiagram
  participant WorkerMonitor
  participant fetch_http_load
  participant authed_get
  participant PromScrape
  participant Worker

  WorkerMonitor->>fetch_http_load: poll worker load
  fetch_http_load->>authed_get: GET /metrics or /v1/loads
  authed_get->>Worker: HTTP request
  Worker-->>authed_get: response body
  authed_get-->>fetch_http_load: raw body
  fetch_http_load->>PromScrape: parse(body)
  PromScrape-->>fetch_http_load: metric map
  fetch_http_load-->>WorkerMonitor: WorkerLoadResponse
Loading

Possibly related PRs

  • lightseekorg/smg#554: Introduces the WorkerLoadResponse/SchedulerLoadSnapshot types with max_total_num_tokens that this PR extends via has_absolute_token_data().
  • lightseekorg/smg#1118: Refactors the same DP-rank cache update plumbing in monitor.rs/manager.rs that this PR now gates on absolute token data.
  • lightseekorg/smg#1630: Adds gRPC GetLoads producing WorkerLoadResponse from vLLM scheduler-load fields, consumed by the gating logic added here.

Suggested labels: tests, enhancement

Suggested reviewers: CatherineSue, key4ng, claude

Poem

A rabbit scrapes the metrics fine,
Prometheus gauges, line by line,
When tokens count is truly there,
We trust the load, we do not dare
To guess a cache that's incomplete—
Hop, hop, evict the stale defeat! 🐰📊

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly reflects the main change: worker load fetching now uses engine-specific HTTP endpoints for vLLM and SGLang.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch xz/http-load-engine-aware

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added protocols Protocols crate changes model-gateway Model gateway crate changes labels Jul 2, 2026
@claude

claude Bot commented Jul 2, 2026 •

Copy link
Copy Markdown

👋 The PR description doesn't fully follow
PULL_REQUEST_TEMPLATE.md:

  • Missing header: ## Description
  • Missing header: ### Problem (uses ## Problem instead of ### Problem)
  • Missing header: ### Solution
  • Missing header: ## Changes (uses ## Change instead of ## Changes)
  • Missing header: ## Test Plan (uses ## Testing instead of ## Test Plan)

Please update the PR description so reviewers have the context they need.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a fallback mechanism to fetch worker load metrics from Prometheus /metrics endpoints for vLLM and SGLang runtimes when native /v1/loads endpoints are unavailable. It adds a minimal Prometheus text-format parser, normalizes the scraped metrics into a single-rank load response, and ensures that ratio-only metrics do not poison the absolute token-based DP-rank routing cache. The reviewer feedback is highly constructive, pointing out that a dedicated parsing library like prometheus-parse should be preferred over custom string manipulation, and identifying metric name changes in newer versions of vLLM and SGLang that require backward-compatible handling.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread model_gateway/src/worker/monitor.rs
Comment thread model_gateway/src/worker/monitor.rs
Comment thread model_gateway/src/worker/monitor.rs
Comment thread model_gateway/src/worker/monitor.rs Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean, well-designed PR. The engine-specific dispatch, ratio-vs-absolute-token distinction, and DP-cache eviction logic are all sound. Good test coverage for the new PromScrape parser and the has_absolute_token_data discriminator. One minor nit on the Prometheus parser's timestamp robustness — not a current bug but worth hardening.

@XinyueZhang369
XinyueZhang369 force-pushed the xz/http-load-engine-aware branch 4 times, most recently from e20b125 to 8f89a1f Compare July 2, 2026 17:40
Comment thread model_gateway/src/worker/monitor.rs Outdated
Signed-off-by: XinyueZhang369 <zoeyzhang369@gmail.com>
@XinyueZhang369
XinyueZhang369 force-pushed the xz/http-load-engine-aware branch from 8f89a1f to 33570e2 Compare July 2, 2026 18:16
@XinyueZhang369
XinyueZhang369 marked this pull request as ready for review July 3, 2026 20:43

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
model_gateway/src/worker/manager.rs (1)

952-971: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

successful/failed counts now conflate fetch outcome with token-data availability.

successful/failed are derived purely from l.load >= 0 (line 970-971). Since ratio-only vLLM/SGLang /metrics responses now legitimately report load: -1 (via the new has_absolute_token_data() gate) even though fetch_http_load succeeded and returned real data, these workers will now be counted as failed in WorkerLoadsResult. Any dashboard/alerting or CLI output built on successful/failed will misreport healthy vLLM/SGLang fleets as having fetch failures.

Consider deriving successful/failed from whether details is Some/None (actual fetch outcome) rather than from load >= 0 (token-data availability), since these are now two independent signals.

💡 Proposed fix
         let loads = future::join_all(futures).await;
-        let successful = loads.iter().filter(|l| l.load >= 0).count();
-        let failed = loads.iter().filter(|l| l.load < 0).count();
+        let successful = loads.iter().filter(|l| l.details.is_some()).count();
+        let failed = loads.iter().filter(|l| l.details.is_none()).count();
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/worker/manager.rs` around lines 952 - 971, The `successful`
and `failed` counters in `WorkerLoadsResult` are currently based on `l.load >=
0`, which mixes fetch success with whether absolute token data was present.
Update the aggregation near the `join_all` result in `manager.rs` so it counts
success/failure from the actual fetch outcome (`details` present vs absent in
`WorkerLoadInfo`), while keeping `load` as the separate token-data indicator
returned by `fetch_http_load`/`has_absolute_token_data()`. This will keep
`successful`/`failed` accurate for ratio-only vLLM/SGLang responses that
legitimately set `load` to `-1`.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@model_gateway/src/worker/manager.rs`:
- Around line 952-971: The `successful` and `failed` counters in
`WorkerLoadsResult` are currently based on `l.load >= 0`, which mixes fetch
success with whether absolute token data was present. Update the aggregation
near the `join_all` result in `manager.rs` so it counts success/failure from the
actual fetch outcome (`details` present vs absent in `WorkerLoadInfo`), while
keeping `load` as the separate token-data indicator returned by
`fetch_http_load`/`has_absolute_token_data()`. This will keep
`successful`/`failed` accurate for ratio-only vLLM/SGLang responses that
legitimately set `load` to `-1`.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: d49ce8de-8cf3-4909-98e3-e9362daa77cc

📥 Commits

Reviewing files that changed from the base of the PR and between 5e6952e and cf91872.

📒 Files selected for processing (3)
  • crates/protocols/src/worker.rs
  • model_gateway/src/worker/manager.rs
  • model_gateway/src/worker/monitor.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cf918720ac

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Some(Self::single_rank(SchedulerLoadSnapshot {
num_running_reqs: m.sum("vllm:num_requests_running") as i32,
num_waiting_reqs: m.sum("vllm:num_requests_waiting") as i32,
token_usage: m.mean(kv_usage),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Normalize summed vLLM cache usage gauges

For vLLM 0.x multi-process/multi-GPU servers, vllm:gpu_cache_usage_perc is exported with Prometheus multiprocess_mode="sum" (see the vLLM 0.7.3 metrics docs: https://docs.vllm.ai/en/v0.7.3/serving/metrics.html), so this value can exceed 1.0 when multiple engine processes report the same ratio. Publishing m.mean(kv_usage) directly as the 0–1 token_usage signal makes load-aware policies treat multi-GPU workers as overloaded at normal utilization and skews comparisons against single-GPU workers; please normalize the summed gauge by the engine/GPU count or otherwise derive an average before updating routing state.

Useful? React with 👍 / 👎.

@slin1237
slin1237 merged commit 8155b1f into main Jul 3, 2026
82 of 84 checks passed
@slin1237
slin1237 deleted the xz/http-load-engine-aware branch July 3, 2026 23:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes protocols Protocols crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants