fix(gateway): register unidentified OpenAI-compatible HTTP backends as generic - #2090
Conversation
…kends RuntimeType has no value for a worker that speaks plain OpenAI HTTP but matches no known engine, so HTTP detection had no honest label to assign and hard-failed instead (#2085). Add Generic ("generic"), treated by the gRPC pipeline like External/Unspecified: it never appears behind a gRPC/ZMQ client and does not support PD dispatch. Refs: #2085 Signed-off-by: yifeng liu <31553858+pallasathena92@users.noreply.github.com>
…s generic HTTP backend detection only fingerprints sglang and vllm. Anything else speaking clean OpenAI — e.g. the SMG gateway that tokenspeed serve embeds on its main port, which reports owned_by "self_hosted" and exposes neither /version nor /server_info — failed the detect_backend step until the startup timeout and the worker never registered, even though a plain curl of /v1/models against it succeeded. When /v1/models is live with at least one model but no engine fingerprint matches after the /version and /server_info probes, register the worker as generic and log a warning naming the evidence instead of rejecting a healthy backend. A dead or model-less endpoint still fails (the step retries), detection errors now spell out per-probe outcomes, and an explicit runtime_type continues to override detection. Closes #2085 Signed-off-by: yifeng liu <31553858+pallasathena92@users.noreply.github.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (4)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe PR adds ChangesGeneric HTTP runtime detection
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant DetectBackendStep
participant HTTPBackend
participant RuntimeType
DetectBackendStep->>HTTPBackend: GET /v1/models
HTTPBackend-->>DetectBackendStep: Recognized or unrecognized owned_by
DetectBackendStep->>HTTPBackend: Probe /version and /server_info
HTTPBackend-->>DetectBackendStep: Fallback backend information
DetectBackendStep->>RuntimeType: Register detected runtime or Generic
Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Clean, well-scoped change. The ModelsProbe enum is a good design choice — it cleanly separates "server answered but unrecognized" from "couldn't probe" which is exactly the distinction needed for the generic fallback. Strategy 3 fires only when /v1/models was live with ≥1 model but no engine fingerprint matched, so unreachable/model-less endpoints still fail and retry as expected. Exhaustive match coverage verified across the codebase — all wildcard arms in related files handle Generic correctly. Tests are thorough.
Description
Problem
smg launch --worker-urls http://<host>:8000pointed at a server started bytokenspeed serve(kimi-k3 recipe) fails registration with:while
curl http://<host>:8000/v1/modelsanswers fine (#2085).Root cause: HTTP detection recognizes only
owned_by ∈ {sglang, nvidia, vllm}plus the/version(vLLM) and/server_info(SGLang) probes. The endpoint in the issue is the SMG gateway thattokenspeed serveembeds in front of the engine — it reportsowned_by: "self_hosted"(SMG's own no-provider model-card value,crates/protocols/src/model_card.rs) and exposes neither fallback endpoint, so a healthy OpenAI-compatible worker was rejected until the startup timeout. The gRPC path already fingerprints tokenspeed; the HTTP path had no way to label "OpenAI-compatible, engine unidentified".Solution
RuntimeType::Generic("generic"): a generic OpenAI-compatible HTTP backend whose engine could not be identified. No engine guessed, no sglang default — the label states exactly what was verified.detect_http_backendstrategy 3: if/v1/modelswas live with ≥1 model but neitherowned_bynor the/version//server_infoprobes identify the engine, register asgenericand WARN with the per-probe evidence plus an override hint.runtime_typeconfig still bypasses detection entirely;genericis also accepted as an explicit value (escape hatch).Genericlike External/Unspecified: no PD dispatch, never behind a gRPC/ZMQ client.genericHTTP workers take the existing warn-and-skip path (same as vllm/tokenspeed HTTP).Changes
crates/protocols/src/worker.rs:RuntimeType::Generic+ as_str/FromStr/serde round-trip testsmodel_gateway/src/workflow/steps/local/detect_backend.rs:ModelsProbeoutcome enum, generic fallback, evidence-rich warn/error, module docs corrected (they claimed tokenspeed/mlx HTTP detection that never existed)model_gateway/src/routers/grpc/common/stages/{request_execution,encode}.rs: exhaustive-match arms for the new variantmaturin develop+ importTest Plan
Unit tests in
detect_backend.rsagainst a mock axum backend (cargo test -p smg --lib detect_backend):owned_by: "self_hosted", no/version, no/server_info→ detectsgeneric(reproduces [Bug]: Step failed: detect_backend - HTTP backend detection failed #2085; failed with the exact issue error before the fix)owned_by→genericowned_bybut live/version→ stillvllm(probe priority preserved)nvidiaowned_by→ stillsglang/v1/modelswith emptydata→ still errors (engine not up yet)cargo test -p openai-protocol:genericround-trips str/serde and counts as specified.Known coverage boundary: no e2e for tokenspeed-serve-over-HTTP — the e2e infra runs TokenSpeed only in gRPC/ZMQ modes (the same gap that let #2085 slip through).
Gates:
cargo +nightly fmt --allclean;cargo clippy --workspace --all-targets -- -D warningsclean (--all-featuresneeds system OpenCV, per CONTRIBUTING fallback); fullcargo testgreen (105 suites);make python-devequivalent builds and imports.Closes #2085
Checklist
cargo +nightly fmtpassescargo clippy --all-targets -- -D warningspasses (workspace fallback; no system OpenCV for--all-features)detect_backend.rs)