Skip to content

feat(zmq): teach the direct-backend path to speak TokenSpeed - #2032

Closed
slin1237 wants to merge 1 commit into
mainfrom
feat/zmq-tokenspeed
Closed

slin1237 wants to merge 1 commit into
mainfrom
feat/zmq-tokenspeed

Conversation

@slin1237

@slin1237 slin1237 commented Aug 2, 2026

Copy link
Copy Markdown
Member

Description

Problem

The ZMQ direct-backend path (#2000, built on #2015) speaks only vLLM's EngineCore protocol. TokenSpeed (and later sglang) engines need the same same-host ipc:// fast path, but the crate's transport, connector, and gateway adapter all hard-assumed the vLLM wire format — and several workflow steps hard-assumed "ZMQ worker ⇒ vLLM".

Solution

Make the ZMQ stack engine-neutral and add TokenSpeed as the second protocol:

  • engine-zmq-client: a new EngineProtocol trait seams the shared transport/connector (handshake, ROUTER/DEALER identity framing, output loop, abort-on-drop) away from the per-engine wire structs. VllmProtocol keeps the existing behavior; TokenSpeedProtocol adds the sglang-family msgpack tuples (WireTokenizedGenerateReq 5-tuple, WireSamplingParams 13-tuple, WireBatchTokenIDOut 8-tuple with sampled-token logprob columns). Handshake structs move to a neutral protocol/handshake.rs (re-exported for compatibility). A neutral EngineLoad replaces the vLLM-specific stats type in the shared seam.
  • Gateway adapter (zmq_client.rs): ZmqEngineClient selects the protocol from the worker's explicit runtime_type (tokenspeed vs vllm; anything else is rejected before the handshake and at registration). Streams map both protocols to the existing vLLM-proto pipeline — chunks carry incremental tokens/logprobs, the terminal Complete carries the cumulative set, and a finish-tick's tokens are emitted as a chunk first so streaming never loses the last token.
  • No silent narrowing: requests the wire cannot honor fail loudly with invalid_argument instead of degrading — structured-output constraints, n>1, top-k/prompt logprobs, stop strings (TokenSpeed), logit_bias (TokenSpeed), nonzero data_parallel_rank. Sampled-token logprobs are wired end-to-end on both ZMQ protocols.
  • Runtime plumbing: detect_backend/discover_metadata no longer force ZMQ workers to vLLM — an explicitly configured runtime_type survives to the built worker (unspecified still defaults to vLLM with a warning). BackendClient::runtime_type() reports the actual ZMQ backend runtime.
  • wfaas fix: the DAG scheduler could fail a workflow with a spurious "Workflow deadlocked" when a step completed between the completion-drain and the tracker read (instant-completing ZMQ detection steps hit this routinely). The deadlock branch now re-drains the completion channel before failing, and the run_if paths send their completion inside the tracker lock scope. Regression tests included.
  • smg serve: per-user ZMQ socket dir (SMG_ZMQ_SOCKET_DIR override), FNV handshake-port derivation pinned by conformance vectors against the Python mirror.

Changes

  • crates/engine_zmq_client: protocol/mod.rs (EngineProtocol/EngineOutput/EngineBatch/EngineLoad), protocol/tokenspeed/{mod,request,sampling,output}.rs, protocol/handshake.rs, generic connector.rs/transport.rs, EngineCoreReadyResponse.max_num_batched_tokens widened to i64 (TokenSpeed sends -1 = disabled).
  • model_gateway: routers/grpc/zmq_client.rs (ZmqBackend enum, per-protocol translate/map with loud rejection boundary, logprob accumulation), routers/grpc/backend_client.rs (runtime passthrough), worker/worker.rs (runtime param on connect), workflow/steps/local/{detect_backend,discover_metadata,create_worker}.rs (runtime preservation + ZMQ runtime validation).
  • crates/workflow: deadlock-detector re-drain + run_if in-lock completion send + regression tests.
  • bindings/python/src/smg/serve.py: per-user socket dir.

Test Plan

  • cargo clippy --workspace --all-targets --all-features -- -D warnings, cargo +nightly fmt --all -- --check clean; engine-zmq-client (incl. mock-engine e2e for both protocols), wfaas, and smg zmq_client suites green.
  • Cross-language contract: Python msgspec-encoded fixtures decode in the Rust codec and vice versa (field order, arity, float64 logprob columns).
  • Live e2e on GB300: real TokenSpeed engine (Qwen3-0.6B, branch lightseekorg/tokenspeed#feat/zmq-msgpack) registered via POST /workers {"url":"ipc://...","connection_mode":"zmq","runtime_type":"tokenspeed"} — chat (content + reasoning_content), completions, sampled-token logprobs (--enable-output-logprobs), top-k/prompt-logprob rejections, and repeated worker registration with zero workflow-deadlock failures. The existing vLLM ZMQ path re-validated unchanged.

Follow-ups: engine-side handshake retry, DP>1 (coordinator/wave, task tracked), legacy /v1/completions logprobs rendering (pre-existing gap for all backends).

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

@coderabbitai

coderabbitai Bot commented Aug 2, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added TokenSpeed backend support for ZMQ workers, including serving, request handling, streaming, sampling, log probabilities, and multi-request generation.
    • Added backend selection and runtime configuration for vLLM and TokenSpeed.
    • Added configurable ZMQ handshake addresses.
    • Improved runtime and worker metadata reporting across supported backends.
  • Bug Fixes

    • Prevented workflow deadlocks when conditional steps are skipped.
    • Improved validation for unsupported runtime and handshake configurations.
    • Added compatibility for disabling chunked prefill through the handshake settings.
  • Tests

    • Added comprehensive coverage for TokenSpeed integration, worker configuration, streaming, validation, and workflow completion.

Walkthrough

Changes

TokenSpeed CLI and runtime configuration

Layer / File(s) Summary
Backend selection and worker launcher
bindings/python/src/lib.rs, bindings/python/src/smg/router.py, bindings/python/src/smg/router_args.py, bindings/python/src/smg/serve.py, bindings/python/tests/test_serve.py
Adds vllm and tokenspeed backend selection. Adds a ZMQ-only TokenSpeed launcher with argument passthrough and handshake-port derivation.
Startup runtime propagation
model_gateway/src/main.rs, model_gateway/src/config/*
Maps ZMQ backends to explicit vLLM or TokenSpeed runtime types and stores the selection in RouterConfig.

Generic ZMQ protocol and gateway

Layer / File(s) Summary
Protocol contracts and wire formats
crates/engine_zmq_client/src/protocol/*
Adds EngineProtocol, shared batches and load data, the TokenSpeed request/output/sampling formats, and signed handshake token limits.
Generic client and transport
crates/engine_zmq_client/src/connector.rs, crates/engine_zmq_client/src/transport.rs, crates/engine_zmq_client/src/lib.rs
Generalizes the ZMQ client and output loop over engine protocols while preserving vLLM aliases and adding TokenSpeed aliases.
Runtime-aware gateway adapter
model_gateway/src/routers/grpc/zmq_client.rs, model_gateway/src/routers/grpc/backend_client.rs
Selects vLLM or TokenSpeed at connection time, translates requests and streams, supports fan-out, and validates backend-specific features.

Worker and workflow changes

Layer / File(s) Summary
Handshake address propagation
crates/protocols/src/worker.rs, model_gateway/src/worker/*, model_gateway/src/workflow/steps/local/create_worker.rs, model_gateway/src/workflow/job_queue.rs
Adds optional TCP handshake overrides, validates them, and propagates them with runtime types into ZMQ connections.
Workflow completion synchronization
crates/workflow/src/engine.rs, crates/workflow/tests/workflow_test.rs
Synchronizes completion signals with tracker updates and adds regression tests for skipped dependency chains.

Estimated code review effort: 5 (Critical) | ~120 minutes

Possibly related issues

Possibly related PRs

  • smg-project/smg#2015: Extends the vLLM-only ZMQ client, launcher, routing, and runtime handling with TokenSpeed support.
  • smg-project/smg#412: Provides related worker/runtime configuration and RouterConfigBuilder infrastructure.
  • smg-project/smg#438: Modifies local worker runtime detection and propagation in related workflow paths.

Suggested reviewers: key4ng, catherinesue

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the main change: adding TokenSpeed support to the direct-backend ZMQ path.
Description check ✅ Passed The description directly explains the TokenSpeed ZMQ implementation, runtime plumbing, workflow fixes, and validation plan.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/zmq-tokenspeed

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added python-bindings Python bindings changes grpc gRPC client and router changes tests Test changes workflow Workflow crate changes model-gateway Model gateway crate changes labels Aug 2, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thorough review of all 22 changed files (~2,260 lines added). No issues found — this is a high-quality PR.

Key observations:

  • EngineProtocol trait seam is well-designed: the connector and transport are cleanly generic over the protocol, with zero-sized type implementations for each engine family. No runtime dispatch overhead in the hot path.
  • Loud rejection boundary in both translate_request and translate_request_tokenspeed correctly fails requests the wire cannot honor (structured output, n>1, prompt logprobs, stop strings, logit_bias, nonzero DP rank for TokenSpeed) instead of silently narrowing.
  • Finish-tick chunk/complete split correctly handles the case where the terminal output carries new tokens: emits a Chunk first (so streaming frontends decode the last token) and holds back the cumulative Complete for the next poll.
  • Deadlock fix in the workflow engine is correct: the re-drain catches completions that arrive between Phase 0's drain and Phase 1's tracker read, and the run_if paths now atomically send their completion inside the tracker write-lock scope to close the race window. The 50-iteration stress test provides confidence.
  • discover_metadata returning None for ZMQ runtime preserves an explicitly configured runtime (previously it always overwrote with "vllm"), which is the key plumbing change that makes TokenSpeed registration work.
  • Cross-language FNV pinning (same vectors in Rust derive_handshake_port and Python _zmq_handshake_port) is a nice contract test that prevents silent port-agreement breakage.
  • f64→f32 logprob downcast for TokenSpeed is an acceptable precision tradeoff given the proto column type constraint, and is documented.

0 🔴 Important · 0 🟡 Nit · 0 🟣 Pre-existing

@slin1237
slin1237 force-pushed the feat/zmq-tokenspeed branch from d4d092a to 5baa27d Compare August 3, 2026 03:08
@github-actions github-actions Bot added the protocols Protocols crate changes label Aug 3, 2026
sub.request_id = format!("{}-{i}", req.request_id);
if let Some(sp) = sub.sampling_params.as_mut() {
sp.n = 1;
sp.seed = sp.seed.map(|seed| seed.wrapping_add(i as i32));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: wrapping_add can produce a negative i32 when seed is close to i32::MAX. On the vLLM path this is fine (negative i64 seeds are still deterministic), but on the TokenSpeed path translate_sampling_tokenspeed silently drops negative seeds via u64::try_from(seed).ok(), making those sub-requests non-deterministic when the user expected determinism. Practically impossible to hit (requires seed ≈ i32::MAX with n > 1), but saturating_add would be a safer choice — it would clamp to i32::MAX and give identical samples for the overflow subs (undesirable but at least predictable).

@claude

claude Bot commented Aug 3, 2026

Copy link
Copy Markdown

Review Summary

0 🔴 Important · 1 🟡 Nit · 0 🟣 Pre-existing

Thoroughly reviewed all 31 files in this PR. The engine-neutral EngineProtocol trait abstraction, TokenSpeed wire protocol implementation, and the workflow deadlock fix are all well-engineered. The transport generalization (Client<P> / RequestStream<P>) is clean, and the explicit rejection boundaries in translate_request_tokenspeed() are the right call for incremental protocol coverage.

The one nit posted is on seed derivation in fan_out_requests() — minor, non-blocking.

No blocking issues found. Not approving per synchronize-event policy.

@slin1237
slin1237 force-pushed the feat/zmq-tokenspeed branch from 5baa27d to dcedacc Compare August 3, 2026 05:44
Base automatically changed from feat/zmq-direct-backend to main August 3, 2026 15:37
Make the ZMQ direct-backend stack engine-neutral and add TokenSpeed as
the second wire protocol, speaking its msgpack-native tagged msgspec
structs directly — a same-host TokenSpeed scheduler is driven over
ipc:// with no Python servicer hop.

engine-zmq-client:
- EngineProtocol trait seams the shared transport/connector (handshake,
  ROUTER/DEALER identity framing, output loop, abort) away from the
  per-engine wire structs; the vLLM protocol keeps its behavior.
- TokenSpeed protocol speaks the engine's native tagged structs: the
  tokenized generate request is emitted as the tagged positional prefix
  through `stream` (nested native SamplingParams, normalized frontend-
  side), and the per-step output decodes the tagged slim batch struct
  (token ids, finish reasons, token counts, sampled-token logprob
  columns). Tag-validated decode, trailing-field tolerance, and pinned
  cross-language byte vectors from the Python encoder.
- Handshake structs move to a neutral protocol/handshake.rs; a neutral
  EngineLoad replaces engine-specific stats in the shared seam;
  max_num_batched_tokens widened to i64 (-1 = disabled).

Gateway:
- ZmqEngineClient selects the protocol from the worker's explicit
  runtime_type; unsupported runtimes are rejected before the handshake
  and at registration. BackendClient reports the actual runtime.
- n>1 is fanned out frontend-side: n single-sample wire requests with
  per-sub rids and deterministic seed derivation, streams interleaved
  with per-choice proto indexes, cumulative usage counted once, drop
  aborts every sub.
- Streams emit a finish-tick's tokens as a chunk before the cumulative
  Complete, so streaming never drops the last token. Sampled-token
  logprobs are wired end-to-end on both ZMQ protocols.
- No silent narrowing: structured-output constraints, top-k/prompt
  logprobs, stop strings, logit_bias, and nonzero data_parallel_rank
  fail loudly with invalid_argument.
- Workers honor an optional WorkerSpec.zmq_handshake_address bind
  override; the FNV port derivation from the ipc path remains the
  no-config default (doc + conformance vectors pinned against the
  Python launcher mirror).
- detect_backend/discover_metadata preserve an explicitly configured
  ZMQ runtime; unspecified still defaults to vLLM with a warning.

smg serve / config:
- `smg serve --backend tokenspeed --connection-mode zmq` launches the
  engine headless (`python -m tokenspeed.cli serve --headless` with the
  derived --data-parallel-rpc-port), mirroring the vLLM zmq launcher;
  dense data parallelism = N independent workers.
- `--backend` pins the startup ZMQ worker runtime through both config
  conversion paths and the Python bindings; per-user ZMQ socket dir
  (SMG_ZMQ_SOCKET_DIR override).

workflow engine:
- Fix a spurious "Workflow deadlocked" failure: completions landing
  between the drain and the tracker read are now re-drained in the
  deadlock branch, and the run_if paths send their completion inside
  the tracker lock scope. Regression tests included.

Validated live on GB300 against a real TokenSpeed engine (Qwen3-0.6B):
chat (content + reasoning_content), completions, streaming, n=2,
sampled-token logprobs, loud rejections, engine-side invalid-request
aborts, ENGINE_CORE_DEAD death detection, and the handshake-address
override with a bare-default engine.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@slin1237
slin1237 force-pushed the feat/zmq-tokenspeed branch from dcedacc to 4443132 Compare August 3, 2026 15:40
@coderabbitai
coderabbitai Bot requested a review from CatherineSue August 3, 2026 15:41
@claude

claude Bot commented Aug 3, 2026

Copy link
Copy Markdown

👋 The PR description doesn't fully follow
PULL_REQUEST_TEMPLATE.md:

  • Missing header: ## Changes (found ### Changes instead — should be a top-level ## header)
  • Missing header: ## Test Plan (found ### Test Plan instead — should be a top-level ## header)

Please update the PR description so reviewers have the context they need.

@slin1237

slin1237 commented Aug 3, 2026

Copy link
Copy Markdown
Member Author

Superseded by a fresh PR against main. This PR was created stacked on #2015 (feat/zmq-direct-backend); after #2015 merged, GitHub keeps it flagged as a stacked PR and refuses both the sync GraphQL and REST merge endpoints. Reopening the identical rebased commit as a non-stacked PR to unblock the merge.

@slin1237 slin1237 closed this Aug 3, 2026
@slin1237
slin1237 deleted the feat/zmq-tokenspeed branch August 3, 2026 15:45

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 6

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/workflow/src/engine.rs (1)

696-716: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

🔴 Important: Persist the terminal step state before publishing its completion.

The scheduler can process the queued completion and finalize the workflow before state_store.update(...).await finishes. wait_for_completion can then call cleanup_if_terminal, and the ignored update can fail after the workflow reports completion. This can leave a skipped or failed run_if step without its terminal state.

  • crates/workflow/src/engine.rs#L696-L716: Persist StepStatus::Skipped before removing the step from running and sending StepResult::Skip inside the tracker lock.
  • crates/workflow/src/engine.rs#L728-L749: Persist StepStatus::Failed and last_error before removing the step from running and sending StepResult::Failure inside the tracker lock.

Handle state-store and completion-send errors explicitly. As per coding guidelines, do not silently fall back to None or a default when configuration validation should fail loudly.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/workflow/src/engine.rs` around lines 696 - 716, The skip and failure
completion paths in crates/workflow/src/engine.rs#L696-L716 and
crates/workflow/src/engine.rs#L728-L749 must persist terminal state before
publishing completion. In both paths, await and explicitly handle the
state_store.update result for StepStatus::Skipped or StepStatus::Failed with
last_error, then remove the step from running and send the completion while
handling send errors explicitly; do not ignore failures or substitute
None/default values when validation should fail loudly.

Source: Coding guidelines

🧹 Nitpick comments (3)
model_gateway/src/routers/grpc/zmq_client.rs (1)

90-109: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

🟡 Nit: Add tests for the two connect-time rejection branches.

Lines 90-96 reject a TokenSpeed backend with engine_count > 1. Lines 100-109 reject any runtime other than Vllm, TokenSpeed, or Unspecified. The test module covers the request-level rejections but not these two. Both branches return before the handshake, so a unit test needs no mock engine and stays fast.

💚 Proposed tests
#[tokio::test]
async fn tokenspeed_rejects_multi_engine_before_handshake() {
    let err = ZmqEngineClient::connect(
        "tcp://127.0.0.1:1",
        "ipc:///tmp/unused-in",
        "ipc:///tmp/unused-out",
        2,
        "m".to_string(),
        RuntimeType::TokenSpeed,
        Duration::from_millis(50),
    )
    .await
    .expect_err("DP>1 must be rejected");
    assert!(err.to_string().contains("single engine"), "{err}");
}

#[tokio::test]
async fn unsupported_runtime_is_rejected_before_handshake() {
    let err = ZmqEngineClient::connect(
        "tcp://127.0.0.1:1",
        "ipc:///tmp/unused-in",
        "ipc:///tmp/unused-out",
        1,
        "m".to_string(),
        RuntimeType::Sglang,
        Duration::from_millis(50),
    )
    .await
    .expect_err("sglang has no ZMQ engine adapter");
    assert!(err.to_string().contains("no engine implementation"), "{err}");
}

As per coding guidelines: "Run the pr-test-analyzer agent to verify that tests adequately cover new or changed functionality."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/routers/grpc/zmq_client.rs` around lines 90 - 109, Add unit
tests in the existing test module for ZmqEngineClient::connect covering both
connect-time rejection branches: TokenSpeed with engine_count greater than one
and an unsupported runtime such as RuntimeType::Sglang. Use unreachable
endpoints and a short timeout to verify each returns before the handshake, and
assert the error messages identify the single-engine and
missing-engine-implementation conditions; run the pr-test-analyzer agent to
confirm coverage.

Source: Coding guidelines

crates/engine_zmq_client/src/lib.rs (1)

38-42: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

🟡 Nit — Re-export EngineLoad next to the other protocol types.

Client::engine_load returns Option<EngineLoad>, but EngineLoad is not in the crate-root re-export list. Callers must reach it through protocol::EngineLoad while EngineBatch and EngineOutput are available at the root. Add it for a consistent public surface.

♻️ Proposed change
-pub use protocol::{EngineBatch, EngineOutput, EngineProtocol};
+pub use protocol::{EngineBatch, EngineLoad, EngineOutput, EngineProtocol};
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/engine_zmq_client/src/lib.rs` around lines 38 - 42, Update the
crate-root protocol re-exports to include EngineLoad alongside EngineBatch,
EngineOutput, and EngineProtocol, so the type returned by Client::engine_load is
publicly available consistently.
crates/engine_zmq_client/src/protocol/tokenspeed/mod.rs (1)

109-112: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

🟡 Nit — validate does not enforce the documented n == 1 invariant.

crates/engine_zmq_client/src/protocol/tokenspeed/sampling.rs line 90 documents that "the transport validates n == 1". This validate implementation accepts every request unconditionally. A caller that sets sampling_params.n > 1 then reaches the engine, which stores n and never fans out, so the caller silently receives one completion instead of n. Either enforce the check here or correct the doc comment in sampling.rs.

♻️ Proposed fix to enforce the invariant
-    fn validate(_request: &Self::Request) -> Result<()> {
-        // The tokenized text path has no fields this client cannot represent.
-        Ok(())
-    }
+    fn validate(request: &Self::Request) -> Result<()> {
+        // n > 1 is fanned out by the gateway; the engine stores n without
+        // acting on it, so a value above 1 would silently drop completions.
+        if request.sampling_params.n != 1 {
+            return Err(Error::UnsupportedField {
+                context: "TokenizedGenerateReqInput",
+                field: "sampling_params.n",
+            });
+        }
+        Ok(())
+    }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/engine_zmq_client/src/protocol/tokenspeed/mod.rs` around lines 109 -
112, Update tokenspeed’s `validate` method to inspect the request’s sampling
parameters and reject any request whose `n` is not 1, preserving success for
valid single-completion requests. Keep the documented transport invariant in
`sampling.rs` aligned with this enforcement.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/engine_zmq_client/src/connector.rs`:
- Around line 242-253: Update the submit flow around register and P::encode_add
so encoding occurs before the registry entry is created, or otherwise remove the
request_id entry when encoding fails. Preserve the existing rollback for
send_to_engine failures and ensure successful requests still register before
sending.
- Around line 628-634: Replace the std::mem::forget(ns) call in the IpcNamespace
setup with a scoped binding that keeps ns alive for the full test while allowing
it to be dropped normally for cleanup. Preserve the endpoint values and ensure
ns remains in scope through all connection and assertion logic.

In `@crates/workflow/tests/workflow_test.rs`:
- Around line 993-1094: Add a deterministic regression test near
test_run_if_false_mid_chain_completes using a custom StateStore whose
get_context() returns an error; configure a run_if step to exercise that failure
path, then assert the workflow reaches its expected failure status and completes
without a deadlock or timeout.

In `@model_gateway/src/main.rs`:
- Line 1465: Add Go binding support for startup_worker_runtime_type alongside
the existing builder call, matching the Python binding’s ZMQ-only
Vllm/Tokenspeed mapping. Expose the field in the public Go configuration API and
propagate it through the relevant Go-to-Rust conversion or builder path,
preserving existing behavior for other runtime types.
- Around line 1401-1414: Update the ZMQ branch that initializes
startup_worker_runtime_type to reject explicitly selected backends other than
Backend::Vllm and Backend::Tokenspeed before leaving the runtime unpinned;
return the existing startup error type with a clear unsupported-backend message.
Preserve None for non-ZMQ connections and unspecified ZMQ backends, and add
coverage for direct --backend trtllm with --worker-urls ipc://....

In `@model_gateway/src/workflow/steps/local/create_worker.rs`:
- Around line 380-391: Update validate_zmq_handshake_override to validate the
configured zmq_handshake_address scheme for ZMQ workers, accepting only tcp://
values and returning an error for ipc:// or any other non-TCP value. Preserve
the existing error for overrides used with non-ZMQ connection modes and ensure
invalid configurations fail before registration.

---

Outside diff comments:
In `@crates/workflow/src/engine.rs`:
- Around line 696-716: The skip and failure completion paths in
crates/workflow/src/engine.rs#L696-L716 and
crates/workflow/src/engine.rs#L728-L749 must persist terminal state before
publishing completion. In both paths, await and explicitly handle the
state_store.update result for StepStatus::Skipped or StepStatus::Failed with
last_error, then remove the step from running and send the completion while
handling send errors explicitly; do not ignore failures or substitute
None/default values when validation should fail loudly.

---

Nitpick comments:
In `@crates/engine_zmq_client/src/lib.rs`:
- Around line 38-42: Update the crate-root protocol re-exports to include
EngineLoad alongside EngineBatch, EngineOutput, and EngineProtocol, so the type
returned by Client::engine_load is publicly available consistently.

In `@crates/engine_zmq_client/src/protocol/tokenspeed/mod.rs`:
- Around line 109-112: Update tokenspeed’s `validate` method to inspect the
request’s sampling parameters and reject any request whose `n` is not 1,
preserving success for valid single-completion requests. Keep the documented
transport invariant in `sampling.rs` aligned with this enforcement.

In `@model_gateway/src/routers/grpc/zmq_client.rs`:
- Around line 90-109: Add unit tests in the existing test module for
ZmqEngineClient::connect covering both connect-time rejection branches:
TokenSpeed with engine_count greater than one and an unsupported runtime such as
RuntimeType::Sglang. Use unreachable endpoints and a short timeout to verify
each returns before the handshake, and assert the error messages identify the
single-engine and missing-engine-implementation conditions; run the
pr-test-analyzer agent to confirm coverage.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 7ff49401-9a52-482f-ad3f-17c662df0e33

📥 Commits

Reviewing files that changed from the base of the PR and between 244fd62 and 4443132.

📒 Files selected for processing (31)
  • bindings/python/src/lib.rs
  • bindings/python/src/smg/router.py
  • bindings/python/src/smg/router_args.py
  • bindings/python/src/smg/serve.py
  • bindings/python/tests/test_serve.py
  • crates/engine_zmq_client/src/connector.rs
  • crates/engine_zmq_client/src/lib.rs
  • crates/engine_zmq_client/src/mock_engine.rs
  • crates/engine_zmq_client/src/protocol/handshake.rs
  • crates/engine_zmq_client/src/protocol/mod.rs
  • crates/engine_zmq_client/src/protocol/tokenspeed/mod.rs
  • crates/engine_zmq_client/src/protocol/tokenspeed/output.rs
  • crates/engine_zmq_client/src/protocol/tokenspeed/request.rs
  • crates/engine_zmq_client/src/protocol/tokenspeed/sampling.rs
  • crates/engine_zmq_client/src/protocol/vllm/mod.rs
  • crates/engine_zmq_client/src/transport.rs
  • crates/protocols/src/worker.rs
  • crates/workflow/src/engine.rs
  • crates/workflow/tests/workflow_test.rs
  • model_gateway/src/config/builder.rs
  • model_gateway/src/config/types.rs
  • model_gateway/src/main.rs
  • model_gateway/src/routers/grpc/backend_client.rs
  • model_gateway/src/routers/grpc/common/stages/encode.rs
  • model_gateway/src/routers/grpc/zmq_client.rs
  • model_gateway/src/worker/builder.rs
  • model_gateway/src/worker/worker.rs
  • model_gateway/src/workflow/job_queue.rs
  • model_gateway/src/workflow/steps/local/create_worker.rs
  • model_gateway/src/workflow/steps/local/detect_backend.rs
  • model_gateway/src/workflow/steps/local/discover_metadata.rs

Comment on lines 242 to 253
let receiver = self.inner.registry.lock().register(request_id.clone())?;

let payload = encode_msgpack(&request)?;
// Text path carries no aux tensor frames.
let (payload, aux_frames) = P::encode_add(&request)?;
if let Err(error) = self
.inner
.send_to_engine(&engine_id, EngineCoreRequestType::Add, payload, Vec::new())
.send_to_engine(&engine_id, P::add_frame(), payload, aux_frames)
.await
{
// Roll back the registry entry so a failed send doesn't leak it.
self.inner.registry.lock().remove_all([&request_id]);
return Err(error);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

🔴 Important — Registry entry leaks when encode_add fails.

register runs at line 242, before P::encode_add at line 244. The rollback at line 251 only covers a send failure. If encode_add returns an error, submit returns early and leaves the request_id in the registry. The receiver is dropped, so no output can clear the entry, and every later submit with the same request_id fails with DuplicateRequestId until the client is dropped. Encode the payload before you register, or roll back on both error paths.

🐛 Proposed fix: encode before registering
         let request_id = P::request_id(&request).to_string();
+        let (payload, aux_frames) = P::encode_add(&request)?;
         let receiver = self.inner.registry.lock().register(request_id.clone())?;
 
-        let (payload, aux_frames) = P::encode_add(&request)?;
         if let Err(error) = self
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
let receiver = self.inner.registry.lock().register(request_id.clone())?;
let payload = encode_msgpack(&request)?;
// Text path carries no aux tensor frames.
let (payload, aux_frames) = P::encode_add(&request)?;
if let Err(error) = self
.inner
.send_to_engine(&engine_id, EngineCoreRequestType::Add, payload, Vec::new())
.send_to_engine(&engine_id, P::add_frame(), payload, aux_frames)
.await
{
// Roll back the registry entry so a failed send doesn't leak it.
self.inner.registry.lock().remove_all([&request_id]);
return Err(error);
}
let (payload, aux_frames) = P::encode_add(&request)?;
let receiver = self.inner.registry.lock().register(request_id.clone())?;
if let Err(error) = self
.inner
.send_to_engine(&engine_id, P::add_frame(), payload, aux_frames)
.await
{
// Roll back the registry entry so a failed send doesn't leak it.
self.inner.registry.lock().remove_all([&request_id]);
return Err(error);
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/engine_zmq_client/src/connector.rs` around lines 242 - 253, Update the
submit flow around register and P::encode_add so encoding occurs before the
registry entry is created, or otherwise remove the request_id entry when
encoding fails. Preserve the existing rollback for send_to_engine failures and
ensure successful requests still register before sending.

Comment on lines +628 to +634
let ns = IpcNamespace::new().unwrap();
let (handshake, input, output) = (
ns.handshake_endpoint(),
ns.input_endpoint(),
ns.output_endpoint(),
);
std::mem::forget(ns);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

🟡 Nit — Replace std::mem::forget(ns) with a scoped binding.

IpcNamespace owns the socket tempdir, as documented at lines 437-438. std::mem::forget leaks the temp directory and the ipc socket files on every run of this test. A plain binding keeps the namespace alive for the whole test and still cleans up on drop, which matches the connect() helper.

🧹 Proposed fix
-        let ns = IpcNamespace::new().unwrap();
-        let (handshake, input, output) = (
-            ns.handshake_endpoint(),
-            ns.input_endpoint(),
-            ns.output_endpoint(),
-        );
-        std::mem::forget(ns);
+        let _ns = IpcNamespace::new().unwrap();
+        let (handshake, input, output) = (
+            _ns.handshake_endpoint(),
+            _ns.input_endpoint(),
+            _ns.output_endpoint(),
+        );
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
let ns = IpcNamespace::new().unwrap();
let (handshake, input, output) = (
ns.handshake_endpoint(),
ns.input_endpoint(),
ns.output_endpoint(),
);
std::mem::forget(ns);
let _ns = IpcNamespace::new().unwrap();
let (handshake, input, output) = (
_ns.handshake_endpoint(),
_ns.input_endpoint(),
_ns.output_endpoint(),
);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/engine_zmq_client/src/connector.rs` around lines 628 - 634, Replace
the std::mem::forget(ns) call in the IpcNamespace setup with a scoped binding
that keeps ns alive for the full test while allowing it to be dropped normally
for cleanup. Preserve the endpoint values and ensure ns remains in scope through
all connection and assertion logic.

Comment on lines +993 to +1094
/// Regression test: instantly-skipped run_if steps feeding a dependent chain
/// must not trigger a spurious "Workflow deadlocked" failure. The skip signal
/// must be sent in the same tracker lock scope as the removal from `running`,
/// otherwise the scheduler can observe running == 0 with the completion unsent.
#[tokio::test]
async fn test_run_if_skip_chain_no_spurious_deadlock() {
for i in 0..50 {
let engine: WorkflowEngine<TestWorkflowData> = WorkflowEngine::new();

let workflow = WorkflowDefinition::new(
"run_if_skip_chain_workflow",
"Run If Skip Chain Deadlock Regression",
)
.add_step(
StepDefinition::new("root", "Root", Arc::new(AlwaysSucceedStep)).run_if(|_ctx| false),
)
.add_step(
StepDefinition::new("mid", "Mid", Arc::new(AlwaysSucceedStep))
.depends_on(&["root"])
.run_if(|_ctx| false),
)
.add_step(
StepDefinition::new("leaf", "Leaf", Arc::new(AlwaysSucceedStep)).depends_on(&["mid"]),
);

let workflow_id = workflow.id.clone();
engine.register_workflow(workflow).unwrap();

let instance_id = engine
.start_workflow(workflow_id, TestWorkflowData::default())
.await
.unwrap();

let result = engine
.wait_for_completion(instance_id, "skip-chain", Duration::from_secs(5))
.await;
assert!(result.is_ok(), "iteration {i} failed: {result:?}");
}
}

/// A -> B(run_if=false) -> C: the chain completes with B skipped and C executed.
#[tokio::test]
async fn test_run_if_false_mid_chain_completes() {
use tokio::time::sleep;

let engine: WorkflowEngine<TestWorkflowData> = WorkflowEngine::new();

let executed = Arc::new(AtomicU32::new(0));
let executed_clone = Arc::clone(&executed);

struct TrackingStep {
counter: Arc<AtomicU32>,
}

#[async_trait::async_trait]
impl StepExecutor<TestWorkflowData> for TrackingStep {
async fn execute(
&self,
_context: &mut WorkflowContext<TestWorkflowData>,
) -> WorkflowResult<StepResult> {
self.counter.fetch_add(1, Ordering::SeqCst);
Ok(StepResult::Success)
}
}

let workflow = WorkflowDefinition::new("run_if_mid_chain_workflow", "Run If Mid Chain Test")
.add_step(StepDefinition::new("a", "A", Arc::new(AlwaysSucceedStep)))
.add_step(
StepDefinition::new("b", "B", Arc::new(AlwaysSucceedStep))
.depends_on(&["a"])
.run_if(|_ctx| false),
)
.add_step(
StepDefinition::new("c", "C", Arc::new(TrackingStep { counter: executed }))
.depends_on(&["b"]),
);

let workflow_id = workflow.id.clone();
engine.register_workflow(workflow).unwrap();

let instance_id = engine
.start_workflow(workflow_id, TestWorkflowData::default())
.await
.unwrap();

// Poll instead of wait_for_completion so step states survive for inspection.
let mut state = engine.get_status(instance_id).await.unwrap();
for _ in 0..100 {
if state.status != WorkflowStatus::Running && state.status != WorkflowStatus::Pending {
break;
}
sleep(Duration::from_millis(50)).await;
state = engine.get_status(instance_id).await.unwrap();
}

assert_eq!(state.status, WorkflowStatus::Completed);
assert_eq!(
state.step_states.get(&StepId::new("b")).unwrap().status,
StepStatus::Skipped
);
assert_eq!(executed_clone.load(Ordering::SeqCst), 1);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

ast-grep outline crates/workflow --items all --match 'StateStore|WorkflowEngine'

rg -n --type rust -C 3 \
  'get_context.*Err|Err\(.*get_context|impl.*StateStore|Mock.*StateStore' \
  crates/workflow

rg -n --type rust -C 3 \
  'run_if.*context|context.*run_if|run_if' \
  crates/workflow/tests

Repository: smg-project/smg

Length of output: 12976


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== state.rs relevant section =="
sed -n '1,140p' crates/workflow/src/state.rs

echo
echo "== engine run_if/context references =="
rg -n --type rust -C 4 \
  'get_context|RunIf|run_if|skip|Skipped|deadlock|Workflow deadlocked' \
  crates/workflow/src crates/workflow/tests

echo
echo "== WorkflowContext/get_workflow_context definitions =="
rg -n --type rust -C 4 \
  'struct WorkflowContext|impl WorkflowContext|get_workflow_context|WorkflowContext<' \
  crates/workflow/src

Repository: smg-project/smg

Length of output: 50371


🟡 Nit: Add coverage for run_if context retrieval failure.

run_if fails the step when state_store.get_context() returns an error, but the tests only cover true/false/context-data conditions. Add a deterministic test using a custom StateStore that returns an error for get_context() and assert the workflow/failure status without deadlock.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/workflow/tests/workflow_test.rs` around lines 993 - 1094, Add a
deterministic regression test near test_run_if_false_mid_chain_completes using a
custom StateStore whose get_context() returns an error; configure a run_if step
to exercise that failure path, then assert the workflow reaches its expected
failure status and completes without a deadlock or timeout.

Source: Coding guidelines

Comment thread model_gateway/src/main.rs
Comment on lines +1401 to +1414
// `--backend` normally only steers the routing mode. Over ZMQ it
// additionally pins the startup workers' runtime: the shared EngineCore
// handshake carries no engine identity, so the wire protocol cannot be
// probed and must be declared. HTTP/gRPC keep auto-detection (None).
let startup_worker_runtime_type = if connection_mode == ConnectionMode::Zmq {
match self.backend {
Some(Backend::Vllm) => Some(RuntimeType::Vllm),
Some(Backend::Tokenspeed) => Some(RuntimeType::TokenSpeed),
_ => None,
}
} else {
None
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== repo files around model_gateway =="
git ls-files | rg '(^|/)model_gateway/src/main\.rs|detect_backend\.rs|parse_serve_args|types|bindings/python/src/lib\.rs' || true

echo
echo "== main.rs outline around relevant symbols =="
if [ -f model_gateway/src/main.rs ]; then
  wc -l model_gateway/src/main.rs
  ast-grep outline model_gateway/src/main.rs --match backend --view expanded || true
  ast-grep outline model_gateway/src/main.rs --match startup_worker_runtime_type --view expanded || true
  echo
  sed -n '60,115p' model_gateway/src/main.rs
  echo
  sed -n '1380,1425p' model_gateway/src/main.rs
  echo
  sed -n '1450,1475p' model_gateway/src/main.rs
  echo
  sed -n '1938,2015p' model_gateway/src/main.rs
fi

echo
echo "== backend/zmq validation symbols =="
rg -n "startup_worker_runtime_type|detect_backend|runtime_type|ConfigError::InvalidValue|Trtllm|Tokenspeed|Vllm|Backend::" -S .

Repository: smg-project/smg

Length of output: 50371


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== job_queue startup pin/default handling =="
sed -n '540,585p' model_gateway/src/workflow/job_queue.rs
sed -n '780,820p' model_gateway/src/workflow/job_queue.rs
echo

echo "== detect_backend runtime default =="
sed -n '1,120p' model_gateway/pkg/modules/local/detect_backend.rs 2>/dev/null || sed -n '1,140p' model_gateway/src/workflow/steps/local/detect_backend.rs
echo

echo "== Python launcher direct-backend/ZMQ restriction search =="
rg -n "tokenspeed|direct-backend|default_worker|worker_url|startup_worker_runtime_type|BackendType::|backend ==" bindings/python/src/smg/serve.py bindings/python/src/lib.rs clients/python/smg_client -S
echo

echo "== Rust binding startup_worker_runtime_type construction =="
sed -n '510,565p' bindings/python/src/lib.rs
sed -n '735,765p' bindings/python/src/lib.rs
echo

echo "== Go API types/search candidate files =="
git ls-files | rg 'bindings/golang|golang|go-sdk|sdk|grpc|model_gateway' | sed -n '1,200p'
rg -n "runtimeType|startupWorkerRuntimeType|runtime_type|BackendType|Tokenspeed|Trtllm|Sglang" bindings/golang crates clients examples -S 2>/dev/null | sed -n '1,220p' || true
echo

echo "== deterministic config conversion behavior for unsupported ZMQ backends =="
python3 - <<'PY'
from pathlib import Path
import re

main = Path("model_gateway/src/main.rs").read_text()
m = re.search(r'let startup_worker_runtime_type = if connection_mode == ConnectionMode::Zmq \{(?P<body>.*?)\n        \} else \{(?P<else>.*?)\n        \};', main, re.S)
print("found_startup_mapping:", bool(m))
if m:
    body = re.sub(r'\s+', ' ', m.group('body'))
    print("has_vllm_branch:", "Some(Backend::Vllm) => Some(RuntimeType::Vllm)" in body)
    print("has_tokenspeed_branch:", "Some(Backend::Tokenspeed) => Some(RuntimeType::TokenSpeed)" in body)
    print("has_else_default:", "_ => None" in body and "RuntimeType::External" not in body)
    print("has_immediate_error:", "return Err(ConfigError::InvalidValue" in body)
if Path("bindings/python/tests/test_serve.py").exists():
    txt = Path("bindings/python/tests/test_serve.py").read_text()
    print("python_test_has_trtllm_unsupported_zmq_error:", ("TrtllmWorkerLauncher" in txt) and ("ZMQ" in txt or "ipc" in txt) and ("ValueError" in txt or "TypeError" in txt or "only support" in txt.lower()))
PY

Repository: smg-project/smg

Length of output: 50371


🔴 Important — Reject unsupported ZMQ backends before leaving startup runtime unpinned.

startup_worker_runtime_type only maps Backend::Vllm and Backend::Tokenspeed for ZMQ; every other backend falls to _ => None. Direct CLI invocation can accept --backend trtllm --worker-urls ipc://..., and the unpinned ZMQ path later defaults to vLLM during backend detection. That turns an explicit unsupported backend request into a wire-protocol mismatch or hang instead of a startup error. Reject unsupported ZMQ backends immediately, and add a test covering --backend trtllm --worker-urls ipc://....

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/main.rs` around lines 1401 - 1414, Update the ZMQ branch
that initializes startup_worker_runtime_type to reject explicitly selected
backends other than Backend::Vllm and Backend::Tokenspeed before leaving the
runtime unpinned; return the existing startup error type with a clear
unsupported-backend message. Preserve None for non-ZMQ connections and
unspecified ZMQ backends, and add coverage for direct --backend trtllm with
--worker-urls ipc://....

Source: Coding guidelines

Comment thread model_gateway/src/main.rs
.mode(mode)
.policy(policy)
.connection_mode(connection_mode)
.startup_worker_runtime_type(startup_worker_runtime_type)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Description: Confirm startup_worker_runtime_type parity across bindings.
rg -n -C5 'startup_worker_runtime_type' bindings/python/src/lib.rs
fd -e go . bindings/golang 2>/dev/null | xargs -r rg -n -i 'runtime_type|RuntimeType' 2>/dev/null

Repository: smg-project/smg

Length of output: 1557


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== model_gateway/src/main.rs relevant section =="
rg -n -C8 'startup_worker_runtime_type|connection_mode|BackendType|RuntimeType|Tokenspeed' model_gateway/src/main.rs | sed -n '1,260p'

echo
echo "== Python bindings occurrences =="
rg -n -C8 'backend|BackendType|connection_mode|RuntimeType|Tokenspeed|startup_worker_runtime_type' bindings/python/src/lib.rs | sed -n '1,220p'

echo
echo "== Go SDK files =="
git ls-files 'bindings/golang/*' 'bindings/golang/**/*' 2>/dev/null | sed -n '1,120p'

echo
echo "== Go SDK runtime/config occurrences =="
fd -e go . bindings/golang 2>/dev/null | xargs -r rg -n -C6 'runtime_type|RuntimeType|StartupWorker|backend|Backend|ConnectionMode|connection_mode|Tokenspeed|tokenspeed|VLLM|vllm' 2>/dev/null || true

Repository: smg-project/smg

Length of output: 23287


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== go-specific backend/connection config symbols =="
rg -n -C5 -i 'backend|backend_type|connection_mode|worker_url|worker_urls|vllm|tokenspeed|runtimetype|startup_worker' bindings/golang -g '*.go' -g '*.rs' 2>/dev/null || true

echo
echo "== go binding constructor / config method definitions =="
ast-grep outline bindings/golang/src/lib.rs 2>/dev/null | sed -n '1,220p' || true
rg -n -C8 '#\[pyo3|class Config|to_router_config|RouterConfig|backend|connection_mode|worker_urls' bindings/golang/src/lib.rs bindings/golang/src/client.rs bindings/golang/client.go 2>/dev/null || true

echo
echo "== Go example config backend usage =="
rg -n -C5 -i 'Backend|backend|backend_type|WorkerUrls|worker_urls' bindings/golang/examples/oai_server/config config.go bindings/golang/examples/oai_server/main.go 2>/dev/null || true

Repository: smg-project/smg

Length of output: 10084


Add Golang binding support for startup_worker_runtime_type.

The Python binding applies the same ZMQ-only Vllm/Tokenspeed mapping, but the Go binding does not expose this config field. Add equivalent Go support so the public config parity applies consistently across bindings.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/main.rs` at line 1465, Add Go binding support for
startup_worker_runtime_type alongside the existing builder call, matching the
Python binding’s ZMQ-only Vllm/Tokenspeed mapping. Expose the field in the
public Go configuration API and propagate it through the relevant Go-to-Rust
conversion or builder path, preserving existing behavior for other runtime
types.

Source: Coding guidelines

Comment on lines +380 to +391
fn validate_zmq_handshake_override(
config: &WorkerSpec,
connection_mode: ConnectionMode,
) -> Result<(), String> {
if config.zmq_handshake_address.is_some() && connection_mode != ConnectionMode::Zmq {
return Err(format!(
"worker {} sets zmq_handshake_address but its connection mode is \
{connection_mode:?}: the field is only meaningful for ZMQ workers",
config.url
));
}
Ok(())

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🔴 Important: Reject non-TCP handshake overrides during registration.

validate_zmq_handshake_override accepts ipc:// and other non-TCP values for ZMQ workers. The field contract requires tcp://. Reject these values here so invalid worker configuration fails before the worker enters the registration and health-check flow.

Proposed fix
 fn validate_zmq_handshake_override(
     config: &WorkerSpec,
     connection_mode: ConnectionMode,
 ) -> Result<(), String> {
-    if config.zmq_handshake_address.is_some() && connection_mode != ConnectionMode::Zmq {
-        return Err(format!(
-            "worker {} sets zmq_handshake_address but its connection mode is \
-             {connection_mode:?}: the field is only meaningful for ZMQ workers",
-            config.url
-        ));
+    if let Some(address) = &config.zmq_handshake_address {
+        if connection_mode != ConnectionMode::Zmq {
+            return Err(format!(
+                "worker {} sets zmq_handshake_address but its connection mode is \
+                 {connection_mode:?}: the field is only meaningful for ZMQ workers",
+                config.url
+            ));
+        }
+        if !address.starts_with("tcp://") {
+            return Err(format!(
+                "worker {} sets invalid zmq_handshake_address {address:?}: \
+                 expected a tcp:// address",
+                config.url
+            ));
+        }
     }
     Ok(())
 }

As per coding guidelines, "Prioritize logic errors, production-breaking bugs, security vulnerabilities, missing error handling, broken cross-references, and incorrect defaults or configuration values."

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
fn validate_zmq_handshake_override(
config: &WorkerSpec,
connection_mode: ConnectionMode,
) -> Result<(), String> {
if config.zmq_handshake_address.is_some() && connection_mode != ConnectionMode::Zmq {
return Err(format!(
"worker {} sets zmq_handshake_address but its connection mode is \
{connection_mode:?}: the field is only meaningful for ZMQ workers",
config.url
));
}
Ok(())
fn validate_zmq_handshake_override(
config: &WorkerSpec,
connection_mode: ConnectionMode,
) -> Result<(), String> {
if let Some(address) = &config.zmq_handshake_address {
if connection_mode != ConnectionMode::Zmq {
return Err(format!(
"worker {} sets zmq_handshake_address but its connection mode is \
{connection_mode:?}: the field is only meaningful for ZMQ workers",
config.url
));
}
if !address.starts_with("tcp://") {
return Err(format!(
"worker {} sets invalid zmq_handshake_address {address:?}: \
expected a tcp:// address",
config.url
));
}
}
Ok(())
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/workflow/steps/local/create_worker.rs` around lines 380 -
391, Update validate_zmq_handshake_override to validate the configured
zmq_handshake_address scheme for ZMQ workers, accepting only tcp:// values and
returning an error for ipc:// or any other non-TCP value. Preserve the existing
error for overrides used with non-ZMQ connection modes and ensure invalid
configurations fail before registration.

Source: Coding guidelines

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes protocols Protocols crate changes python-bindings Python bindings changes tests Test changes workflow Workflow crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants