Skip to content

fix(gateway): register unidentified OpenAI-compatible HTTP backends as generic - #2090

Merged
slin1237 merged 2 commits into
mainfrom
fix/2085-generic-runtime-detection
Aug 11, 2026
Merged

slin1237 merged 2 commits into
mainfrom
fix/2085-generic-runtime-detection

Conversation

@pallasathena92

Copy link
Copy Markdown
Collaborator

Description

Problem

smg launch --worker-urls http://<host>:8000 pointed at a server started by tokenspeed serve (kimi-k3 recipe) fails registration with:

Step failed: detect_backend - HTTP backend detection failed for http://...:
Could not detect HTTP backend for http://... (tried /v1/models, /version, /server_info)

while curl http://<host>:8000/v1/models answers fine (#2085).

Root cause: HTTP detection recognizes only owned_by ∈ {sglang, nvidia, vllm} plus the /version (vLLM) and /server_info (SGLang) probes. The endpoint in the issue is the SMG gateway that tokenspeed serve embeds in front of the engine — it reports owned_by: "self_hosted" (SMG's own no-provider model-card value, crates/protocols/src/model_card.rs) and exposes neither fallback endpoint, so a healthy OpenAI-compatible worker was rejected until the startup timeout. The gRPC path already fingerprints tokenspeed; the HTTP path had no way to label "OpenAI-compatible, engine unidentified".

Solution

  • New RuntimeType::Generic ("generic"): a generic OpenAI-compatible HTTP backend whose engine could not be identified. No engine guessed, no sglang default — the label states exactly what was verified.
  • detect_http_backend strategy 3: if /v1/models was live with ≥1 model but neither owned_by nor the /version//server_info probes identify the engine, register as generic and WARN with the per-probe evidence plus an override hint.
  • Unreachable or model-less endpoints still fail (the step keeps retrying), and the failure message now lists per-probe outcomes instead of a bare endpoint list.
  • Explicit runtime_type config still bypasses detection entirely; generic is also accepted as an explicit value (escape hatch).
  • gRPC pipeline treats Generic like External/Unspecified: no PD dispatch, never behind a gRPC/ZMQ client.
  • DP-aware discovery: generic HTTP workers take the existing warn-and-skip path (same as vllm/tokenspeed HTTP).

Changes

  • crates/protocols/src/worker.rs: RuntimeType::Generic + as_str/FromStr/serde round-trip tests
  • model_gateway/src/workflow/steps/local/detect_backend.rs: ModelsProbe outcome enum, generic fallback, evidence-rich warn/error, module docs corrected (they claimed tokenspeed/mlx HTTP detection that never existed)
  • model_gateway/src/routers/grpc/common/stages/{request_execution,encode}.rs: exhaustive-match arms for the new variant
  • No bindings changes needed: Python/Go only construct specific variants; verified by maturin develop + import

Test Plan

Unit tests in detect_backend.rs against a mock axum backend (cargo test -p smg --lib detect_backend):

  • owned_by: "self_hosted", no /version, no /server_info → detects generic (reproduces [Bug]: Step failed: detect_backend - HTTP backend detection failed #2085; failed with the exact issue error before the fix)
  • missing owned_by → generic
  • unrecognized owned_by but live /version → still vllm (probe priority preserved)
  • nvidia owned_by → still sglang
  • unreachable endpoint → still errors (step retries)
  • /v1/models with empty data → still errors (engine not up yet)

cargo test -p openai-protocol: generic round-trips str/serde and counts as specified.

Known coverage boundary: no e2e for tokenspeed-serve-over-HTTP — the e2e infra runs TokenSpeed only in gRPC/ZMQ modes (the same gap that let #2085 slip through).

Gates: cargo +nightly fmt --all clean; cargo clippy --workspace --all-targets -- -D warnings clean (--all-features needs system OpenCV, per CONTRIBUTING fallback); full cargo test green (105 suites); make python-dev equivalent builds and imports.

Closes #2085

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets -- -D warnings passes (workspace fallback; no system OpenCV for --all-features)
  • Documentation updated (module/step docs in detect_backend.rs)

…kends

RuntimeType has no value for a worker that speaks plain OpenAI HTTP but
matches no known engine, so HTTP detection had no honest label to assign
and hard-failed instead (#2085). Add Generic ("generic"), treated by the
gRPC pipeline like External/Unspecified: it never appears behind a
gRPC/ZMQ client and does not support PD dispatch.

Refs: #2085
Signed-off-by: yifeng liu <31553858+pallasathena92@users.noreply.github.com>
…s generic

HTTP backend detection only fingerprints sglang and vllm. Anything else
speaking clean OpenAI — e.g. the SMG gateway that tokenspeed serve embeds
on its main port, which reports owned_by "self_hosted" and exposes
neither /version nor /server_info — failed the detect_backend step until
the startup timeout and the worker never registered, even though a plain
curl of /v1/models against it succeeded.

When /v1/models is live with at least one model but no engine fingerprint
matches after the /version and /server_info probes, register the worker
as generic and log a warning naming the evidence instead of rejecting a
healthy backend. A dead or model-less endpoint still fails (the step
retries), detection errors now spell out per-probe outcomes, and an
explicit runtime_type continues to override detection.

Closes #2085

Signed-off-by: yifeng liu <31553858+pallasathena92@users.noreply.github.com>
@github-actions github-actions Bot added grpc gRPC client and router changes protocols Protocols crate changes model-gateway Model gateway crate changes labels Aug 11, 2026
@coderabbitai

coderabbitai Bot commented Aug 11, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: c5eeecae-6cec-4fe9-a609-a52421ebcfd3

📥 Commits

Reviewing files that changed from the base of the PR and between c2cf59c and c7a9824.

📒 Files selected for processing (4)
  • crates/protocols/src/worker.rs
  • model_gateway/src/routers/grpc/common/stages/encode.rs
  • model_gateway/src/routers/grpc/common/stages/request_execution.rs
  • model_gateway/src/workflow/steps/local/detect_backend.rs

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added support for identifying generic OpenAI-compatible HTTP backends when the provider cannot be determined.
    • Generic backends are now recognized and represented consistently across runtime configuration and requests.
  • Bug Fixes

    • Improved backend detection to distinguish unavailable servers from live servers with unknown engine identifiers.
    • Preserved more specific error details for failed or invalid model-discovery requests.
    • Added handling for generic backends in supported request-routing paths.

Walkthrough

The PR adds RuntimeType::Generic, extends gateway handling, and updates HTTP backend detection. Successful OpenAI-compatible servers with unknown or missing owned_by values register as generic, while probe failures remain errors.

Changes

Generic HTTP runtime detection

Layer / File(s) Summary
Runtime type contract
crates/protocols/src/worker.rs
Adds RuntimeType::Generic with "generic" parsing, display, serialization, deserialization, and specification tests.
HTTP detection outcomes and fallback
model_gateway/src/workflow/steps/local/detect_backend.rs
Separates recognized, unrecognized, and failed /v1/models results. Preserves vLLM and SGLang fallback behavior. Registers unidentified live servers as generic. Adds detection tests and probe helpers.
Gateway runtime handling
model_gateway/src/routers/grpc/common/stages/encode.rs, model_gateway/src/routers/grpc/common/stages/request_execution.rs
Maps generic runtimes to "unknown" backend names and rejects them for disaggregated PD execution.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant DetectBackendStep
  participant HTTPBackend
  participant RuntimeType
  DetectBackendStep->>HTTPBackend: GET /v1/models
  HTTPBackend-->>DetectBackendStep: Recognized or unrecognized owned_by
  DetectBackendStep->>HTTPBackend: Probe /version and /server_info
  HTTPBackend-->>DetectBackendStep: Fallback backend information
  DetectBackendStep->>RuntimeType: Register detected runtime or Generic
Loading

Possibly related PRs

Suggested labels: tests

Suggested reviewers: catherinesue, key4ng

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: registering unidentified OpenAI-compatible HTTP backends as generic.
Description check ✅ Passed The description directly explains the detection failure, proposed generic runtime, implementation changes, and test coverage.
Linked Issues check ✅ Passed The changes address issue #2085 by registering healthy OpenAI-compatible HTTP workers when engine identification is inconclusive.
Out of Scope Changes check ✅ Passed The runtime, routing, detection, documentation, and test changes support the stated objectives and issue #2085.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/2085-generic-runtime-detection

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean, well-scoped change. The ModelsProbe enum is a good design choice — it cleanly separates "server answered but unrecognized" from "couldn't probe" which is exactly the distinction needed for the generic fallback. Strategy 3 fires only when /v1/models was live with ≥1 model but no engine fingerprint matched, so unreachable/model-less endpoints still fail and retry as expected. Exhaustive match coverage verified across the codebase — all wildcard arms in related files handle Generic correctly. Tests are thorough.

@slin1237
slin1237 merged commit f7c304d into main Aug 11, 2026
45 of 46 checks passed
@slin1237
slin1237 deleted the fix/2085-generic-runtime-detection branch August 11, 2026 13:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes protocols Protocols crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Step failed: detect_backend - HTTP backend detection failed

2 participants