Skip to content

fix(runtime): support IPv6-only IP resolution - #13126

Merged
jthomson04 merged 8 commits into
mainfrom
jthomson04/fix-ipv6-ip-resolution
Sep 18, 2026
Merged

jthomson04 merged 8 commits into
mainfrom
jthomson04/fix-ipv6-ip-resolution

Conversation

@jthomson04

@jthomson04 jthomson04 commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Fix local address selection on IPv6-only hosts. Prefer usable non-loopback IPv4, then IPv6, then IPv4 loopback, then IPv6 loopback. On Unix, automatic selection uses one interface snapshot and excludes down interfaces; explicit interface lookup remains unchanged.
  • Accept IP literals, bracketed IPv6, wildcards, interface names, and aliases. Trim surrounding host whitespace for direct Rust callers. Canonicalize IPv4-mapped addresses before classification; mapped wildcards select a concrete advertised address through the IPv4 wildcard rules.
  • Keep bind and advertised addresses separate with typed SocketAddr values. Wildcards can switch families to select a usable address, and loopback fallback prefers the requested family. Publish addresses and actual ports served by the TCP and system-status listeners.
  • Preserve the public IpResolver and ServerOptions::interface APIs. The new list_up_afinet_netifas() method has a default for existing custom resolvers. Cache only successful non-loopback results from local_ip_for_advertise(); TCP servers resolve once at startup.
  • Document interface-state filtering, mapped-address handling, host whitespace and empty-value rules, and TCP versus QUIC binding. Keep port parsing unchanged and update the affected workspace lockfiles.

This change follows #7627 and #7632. Fixes #7619.

Limitation

Direct ZMQ publishers still bind an IPv4 wildcard. Automatic advertisement stays in that family and falls back to IPv4 loopback when no eligible non-loopback IPv4 address exists, so remote subscribers cannot connect. An explicit override must identify a reachable IPv4 address served by the listener. An IPv6 override does not enable IPv6 listening.

Validation

Local validation for f1b5aea3c0088b92c9caf6448cc21dcab23581f9:

  • 74 focused runtime tests passed: 15 resolver, 44 TCP-server, 10 direct-ZMQ, 2 environment-registry, 2 response-port, and 1 system-status test. Tests ran serially and include real IPv6 and mapped-wildcard binds, port connectivity, padded host values, down-interface filtering, and legacy resolver compatibility.
  • cargo clippy --locked --no-deps -p dynamo-runtime --all-targets -- -D warnings passed.
  • Root and Python binding formatting checks passed.
  • Docs lint passed with 0 errors; 8 warnings are outside the edited page.
  • fern check and fern docs broken-links passed after generating the ignored API reference pages.
  • git diff --check passed.

CI results are separate from these local checks.

@github-actions github-actions Bot added the fix label Aug 12, 2026
@jthomson04
jthomson04 marked this pull request as ready for review August 12, 2026 21:21
@jthomson04
jthomson04 requested a review from a team as a code owner August 12, 2026 21:21

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

Open in Devin Review

Comment thread lib/runtime/src/pipeline/network/tcp/server.rs Outdated
Comment thread lib/runtime/src/pipeline/network/tcp/server.rs Outdated
@coderabbitai

coderabbitai Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 83015aaf-1ff5-482a-843a-fd82ef714dee

📥 Commits

Reviewing files that changed from the base of the PR and between 975d6a1 and 86be9ef.

📒 Files selected for processing (2)
  • lib/runtime/src/pipeline/network/tcp/server.rs
  • lib/runtime/src/utils/ip_resolver.rs

Walkthrough

Changes

TCP IPv6 address handling

Layer / File(s) Summary
Local IP resolution flow
lib/runtime/src/pipeline/network/tcp/server.rs, lib/runtime/src/utils/ip_resolver.rs
A shared resolver retries IPv6 after IPv4 errors, preserves non-address errors, and supports loopback fallback. Tests cover strategy errors and IPv6 results.
Typed TCP address propagation
lib/runtime/src/pipeline/network/tcp/server.rs
TcpStreamServer stores IpAddr values and uses SocketAddr values for startup, listener binding, logging, and registration.
Resolver and registration validation
lib/runtime/src/pipeline/network/tcp/server.rs
Tests verify IPv6 fallback, IPv4 error preservation, and socket-address registration formatting.

Estimated code review effort: 3 (Moderate) | ~20 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The changes retry IPv6 after all IPv4 errors, preserve relevant failures, and use typed socket addresses, meeting issue #7619.
Out of Scope Changes check ✅ Passed All reported changes and regression tests support IPv6-only resolution, error handling, or valid TCP socket addressing.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Title check ✅ Passed The title clearly identifies the primary change: adding IPv6-only IP resolution support in the runtime.
Description check ✅ Passed The description provides detailed change scope, limitations, validation results, and issue references. It does not use the required template headings and omits the required Related Issues selection an…
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch

Comment @coderabbitai help to get the list of available commands.

@datadog-official

This comment has been minimized.

@jthomson04 jthomson04 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review — IPv6-only IP resolution

15 findings. 11 are inline below; 4 land in files this PR doesn't touch, so they're listed here.

The one to look at first: the headline fix likely doesn't work on Linux — see the comment on fallback_loopback. local_ipv6() returns LocalIpAddressNotFound on typical IPv6-only Linux hosts, so such a host falls through to fallback_loopback and advertises [::1], while the interface list it already fetched has the routable global address in it.

The branch type-checks clean (cargo check -p dynamo-runtime --lib --tests, exit 0).


Findings in files outside this diff

These are pre-existing, but this PR materially widens the conditions under which they fire — local_ip_for_advertise() now returns IPv6 in many more cases than before (IPv4 probe errors rather than only LocalIpAddressNotFound, IPv4 addresses newly classed unusable, and loopback-only hosts).

1. lib/runtime/src/transports/event_plane/mod.rs:412 — ZMQ binds IPv4 wildcard, advertises a possibly-IPv6 address.
ZmqPubTransport::bind("tcp://0.0.0.0:0") (L399) listens on IPv4 only, then L412-413 builds public_endpoint = format!("tcp://{}:{}", local_ip_for_advertise(), actual_port). On a host where the IPv4 probe now fails or yields an unusable address but IPv6 succeeds — the exact new retry path this PR adds — that becomes tcp://[2001:db8::5]:54321. Subscribers connecting to the discovered endpoint get ECONNREFUSED, because the publisher socket is bound to 0.0.0.0. Same mismatch on a loopback-only host (tcp://[::1]:54321).

2. lib/llm/src/local_model.rs:720 — self URL can name a family the listener doesn't serve.
DEFAULT_SYSTEM_HOST = "0.0.0.0" (config.rs:18) → spawn_system_status_server("0.0.0.0", port) listens on IPv4 only. self_host_base_url sees system_host == "0.0.0.0" and substitutes local_ip_for_advertise(). Result: http://[2001:db8::5]:8081 or http://[::1]:8081 pointing at an IPv4-only listener → connection refused for every consumer.

3. lib/runtime/src/component/endpoint.rs:287 — the request plane still advertises the raw 0.0.0.0 string.
The bind-vs-advertise separation this PR introduces lives only inside TcpStreamServer, so the identical 0.0.0.0 bug is still live one module over. With DYN_TCP_RPC_HOST=0.0.0.0, manager.rs:270 binds 0.0.0.0:port (correct) and endpoint.rs:287-305 publishes TransportType::Tcp("0.0.0.0:5678/<id>/<name>") to discovery. TcpRequestClient::parse_address parses it fine and TcpStream::connect("0.0.0.0:5678") resolves to localhost on the client's node, so remote callers silently talk to themselves or fail. The new resolve_host_or_interface/ResolvedHost machinery handles exactly this but is only reachable from TcpStreamServer::new_with_resolver; local_model.rs:720 also still carries a hand-rolled "0.0.0.0" | "::" | "[::]" special case. Three call sites, three diverging wildcard policies — worth routing all three through the new resolver.

4. lib/runtime/src/config/environment_names.rs:479 — DYN_TCP_RESPONSE_STREAM_HOST docs describe the old behavior.
The comment still reads "Host/interface for the TCP response stream server. If unset, the server auto-detects a routable local IP." The variable now also accepts plain IPv4/IPv6 literals, bracketed IPv6 ([2001:db8::1]), and wildcards (0.0.0.0, ::, [::]) with distinct bind-vs-advertise behavior — and it now hard-rejects values it previously accepted (alias names containing :, link-local/multicast/broadcast literals). Nothing else in the repo mentions the variable, so operators have no way to learn the new grammar or the new rejection rules. Relatedly, ServerOptions::interface (server.rs:78) is now misnamed — it's a host-or-interface field.

Comment thread lib/runtime/src/utils/ip_resolver.rs Outdated
Comment thread lib/runtime/src/utils/ip_resolver.rs Outdated
Comment thread lib/runtime/src/utils/ip_resolver.rs Outdated
Comment thread lib/runtime/src/utils/ip_resolver.rs Outdated
Comment thread lib/runtime/src/utils/ip_resolver.rs Outdated
Comment thread lib/runtime/src/utils/ip_resolver.rs Outdated
Comment thread lib/runtime/src/utils/ip_resolver.rs
Comment thread lib/runtime/src/utils/ip_resolver.rs Outdated
Comment thread lib/runtime/src/utils/ip_resolver.rs Outdated
Comment thread lib/runtime/src/pipeline/network/tcp/server.rs Outdated
@jthomson04
jthomson04 requested review from a team as code owners August 17, 2026 18:55
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Aug 17, 2026
@jthomson04

Copy link
Copy Markdown
Contributor Author

Addressed the four review-body findings in 697d854a4d:

  1. Direct ZMQ: derives the public endpoint from the actual bound SocketAddr and selects only a same-family advertisement address.
  2. System status: adds SystemStatusServerInfo::advertised_socket_addr() and uses it for the LLM self-host URL and Python binding.
  3. Request plane: resolves the raw DYN_TCP_RPC_HOST value when the shared server starts, stores the concrete advertised SocketAddr, and publishes that address after the real port is known.
  4. Documentation and compatibility: documents the accepted literal, wildcard, and interface grammar plus family switching and fallback caching; retains ServerOptions::interface with clear compatibility documentation.

Focused local validation passed: 9 resolver tests, 36 serial TCP server tests, 1 request-plane test, 1 event-plane test, 1 system-status test, and 2 no-default-feature dynamo-llm tests. Runtime and LLM clippy passed with warnings denied; root and Python Rust format checks, the Python binding check, Fern structure check, and Fern broken-link check also passed.

@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

@jthomson04
jthomson04 force-pushed the jthomson04/fix-ipv6-ip-resolution branch 2 times, most recently from bb01c01 to 4eb4612 Compare August 27, 2026 18:51

@dagil-nvidia dagil-nvidia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review of head 4eb4612: 1 must-fix, 5 consider. All gates pass, CI is green on this head, the resolver core is sound and well-tested, and every finding from the prior review rounds is verified addressed.

Must-fix

  1. lib/runtime/src/config/environment_names.rs:737 and docs/fern/pages/reference/components/runtime-configuration.mdx:142 - both changed surfaces claim "Enumeration errors and loopback fallbacks are retried and are not cached," but for the two TCP servers these docs sit under, a loopback fallback is never retried: both servers resolve through once-cells, so a server that comes up advertising loopback holds it for the process lifetime; only an Err leaves the cell empty. The retry sentence is true only of local_ip_for_advertise(), which neither env var governs. Suggested scoping: "Automatic selection for advertised URLs caches only successful non-loopback results and re-resolves after enumeration errors and loopback fallbacks. The TCP servers resolve once at startup; a loopback fallback persists until restart, and an enumeration error fails server startup." Mirror in the env comment.

Consider

  1. ip_resolver.rs:244 - the wildcard loopback fallback ignores the requested family: :: on a dual-loopback host binds 0.0.0.0 and advertises 127.0.0.1 even though ::1 is usable in the requested family; either try the requested family's loopback first or document the IPv4 preference.
  2. event_plane/mod.rs:454 - the direct ZMQ publisher still binds 0.0.0.0, so IPv6-only support does not extend to the direct ZMQ event plane (it now advertises a family-consistent but unreachable loopback); worth a scope note or follow-up.
  3. ip_resolver.rs:352 - tcp_rpc_host_from_env() has zero in-repo callers, is pub re-exported, and returns the raw env string; with the new grammar an external caller gets an interface name or wildcard as a URL host. Deprecate or route through the resolver.
  4. manager.rs:100 - DYN_TCP_RPC_HOST/DYN_TCP_RPC_PORT are read as raw literals and absent from the canonical env registry while this PR documents them publicly; register the constants.
  5. runtime-configuration.mdx:133 - "unspecified ... rejected" reads as contradicting the wildcard acceptance two sentences earlier, and "unscoped link-local rejected" implies scoped is accepted when it cannot even parse. Suggested: "Multicast, broadcast, and IPv6 link-local addresses are rejected; wildcards are accepted as bind addresses and are never advertised."

One adjacent pre-existing worth a one-line fix while in the area: lib/backend-common/src/rl.rs:67 carries a third copy of the old self-host-URL logic (substituting local_ip_for_advertise() on an unspecified bind), which can advertise a cross-family unreachable URL - the exact class this PR fixes in lib.rs and local_model.rs; info.advertised_socket_addr() closes it.

Verified sound: bracket round-trip through both TCP clients; the config grammar matches the docs including loud failures on malformed literals; the selection order and family-switch claims all confirmed; all 15 prior self-review findings and both bot findings addressed at head.

Comment-only review; severity calls stay with the human reviewers.

@jthomson04

Copy link
Copy Markdown
Contributor Author

@dagil-nvidia Addressed the review feedback in 33edfb486a:

  1. Scoped the cache and retry documentation to local_ip_for_advertise(). The TCP server docs now state that resolution occurs once at startup, loopback persists until restart, and enumeration errors fail startup.
  2. Wildcard fallback now tries loopback in the requested family first. The new regression test covers dual-loopback, cross-family fallback, and an empty inventory.
  3. Added a PR limitation that direct ZMQ still binds tcp://0.0.0.0; this PR only keeps its advertised address in the bound family.
  4. Deprecated tcp_rpc_host_from_env() without removing the public API.
  5. Added DYN_TCP_RPC_HOST and DYN_TCP_RPC_PORT to the canonical environment registry and replaced the raw string uses.
  6. Clarified that wildcards are accepted bind addresses and are never advertised, while multicast, broadcast, and IPv6 link-local addresses are rejected.

The adjacent RL self-host path now uses SystemStatusServerInfo::advertised_socket_addr().

Validation passed: 10 resolver tests, 2 environment registry tests, 41 serial TCP server tests, the focused system-status and RL tests, formatting, and warning-denied clippy for dynamo-runtime and dynamo-backend-common. Fern still stops on the 26 generated API reference pages missing from this checkout, which is the same repository-level gap seen on the parent head.

@TaekyungHeo

Copy link
Copy Markdown

Additional dual-stack cross-node reproduction from Aria ComputeLab SC qualification:

  • Interface ens6f0np0 carries IPv4 10.78.x.x and IPv6 link-local fe80::...%ens6f0np0.
  • Dynamo vLLM 1.1 and SGLang 1.1/1.3.1 intermittently select the link-local address when DYN_TCP_RESPONSE_STREAM_HOST=ens6f0np0, then fail with Failed to start TcpListender on fe80::...:0: Invalid argument (os error 22).
  • Supplying the IPv4 literal instead fails in released images with Interface not found: 10.78.x.x.
  • One vLLM 1.3.1 run succeeded on the same fabric while another reached the IPv6 failure, demonstrating nondeterministic selection rather than a deterministic interface policy.
  • Control rows on the same two-node Pyxis/TCP fabric pass: Native vLLM 4027060, Native SGLang 4027098, Native TensorRT-LLM 4026356, Dynamo TensorRT-LLM 4027434.

This matches this PR's need for one deterministic selection pass that prefers non-loopback IPv4 and accepts explicit literals. These are serving-protocol qualification canaries, not performance claims.

@jthomson04
jthomson04 force-pushed the jthomson04/fix-ipv6-ip-resolution branch from 33edfb4 to 48b2f2a Compare September 1, 2026 19:50
@copy-pr-bot

copy-pr-bot Bot commented Sep 1, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@harryskim harryskim left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs-only review, scoped to docs/fern/pages/reference/components/runtime-configuration.mdx. I checked each statement against the code on this head.

Accurate as written: the unset-selection order matches resolve_local_host/AddressCandidates; the wildcard family-switch matches resolve_wildcard; the link-local/multicast/broadcast rules match is_usable; the caching paragraph matches cached_local_ip_for_advertise; eth0:1 does resolve as an interface name (looks_like_ip_literal requires >= 2 colons); and the QUIC sentence holds, since TcpStreamServer::local_address() returns the advertised address that quic_response_server() reuses.

Two things I'd resolve before merge, both one-or-two-sentence fixes, flagged inline:

  1. The "accepts the same values" claim for DYN_TCP_RESPONSE_STREAM_HOST is not true for whitespace — DYN_TCP_RPC_HOST does not trim, and the trimming sentence was deleted from the response-stream entry even though the behavior still holds there.
  2. The DYN_EVENT_PLANE_HOST note understates a user-visible behavior change: on an IPv6-only host, automatic selection now advertises IPv4 loopback rather than the IPv6 address.

The rest of the inline notes are non-blocking follow-ups.

One caveat on verification: the PR description reports fern check and fern docs broken-links as blocked by absent generated API reference pages, so the docs build was not confirmed green for this change. The edits are balanced ParamField blocks, so I'd expect them to render, but that is inspection rather than a build.

Comment thread docs/fern/pages/reference/components/runtime-configuration.mdx Outdated
Comment thread docs/fern/pages/reference/components/runtime-configuration.mdx Outdated
Comment thread docs/fern/pages/reference/components/runtime-configuration.mdx Outdated
Comment thread docs/fern/pages/reference/components/runtime-configuration.mdx Outdated
Comment thread docs/fern/pages/reference/components/runtime-configuration.mdx Outdated
@jthomson04
jthomson04 force-pushed the jthomson04/fix-ipv6-ip-resolution branch from 4713b9f to 273929a Compare September 16, 2026 23:15
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
@jthomson04
jthomson04 force-pushed the jthomson04/fix-ipv6-ip-resolution branch from 273929a to 0d7c61e Compare September 17, 2026 19:34

@rmccorm4 rmccorm4 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Three P2 findings reproduced during local review of 0d7c61ee9b.

Comment thread lib/runtime/src/utils/ip_resolver.rs
Comment thread lib/runtime/src/utils/ip_resolver.rs
Comment thread lib/runtime/src/utils/ip_resolver.rs
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
@jthomson04
jthomson04 requested review from a team as code owners September 17, 2026 19:59
@jthomson04
jthomson04 requested a review from rmccorm4 September 17, 2026 20:22
@rmccorm4

Copy link
Copy Markdown
Contributor

Summarizing TLDR of the changes

Situation Before After
IPv6-only machine IPv4 detection could fail without finding the working IPv6 address. Checks interface addresses directly; prefers usable IPv4, then IPv6.
Interface has multiple addresses Could pick an unsuitable IPv6 link-local address. Uses a consistent preference order and rejects unsuitable addresses. Automatic selection also skips down interfaces.
Server listens on a wildcard such as 0.0.0.0 Some paths advertised the wildcard itself, which remote machines cannot use as a destination. Listens on the wildcard but advertises a concrete address.
Dynamically assigned port Different paths constructed addresses independently. Publishes the actual port assigned to the listener.

@dmitry-tokarev-nv dmitry-tokarev-nv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved

I verified the address handling by running it. Every place the change builds or splits an endpoint string round-trips for both address families, the new tests are real, and the two open items are low.

Verified: string round-trips, downstream consumers, and a simulated IPv6-only host

Built at head f1b5aea3c0 in a worktree, macOS ARM64, cargo test -p dynamo-runtime --no-default-features. Merge base a10a46a8a0, dated 2026-09-17, so the branch is current.

Endpoint strings, run end to end:

Built at Form for IPv6 Read back by Result
tcp/server.rs:214 advertised address [2001:db8::10]:1234 local_address(), TcpStreamConnectionInfo parses back to the same SocketAddr
component/endpoint.rs:346 transport [2001:db8::10]:1234/2a/generate TcpRequestClient::parse_address addr=[2001:db8::10]:1234 endpoint="2a/generate"
event_plane/mod.rs:87 ZMQ endpoint tcp://[2001:db8::20]:4321 ZMQ connect string family kept, port kept
rl.rs:93, local_model.rs:895, lib.rs:1475 http://[::1]:8080 URL consumers, admin.py brackets present
tcp/client.rs:368 TLS name from [2001:db8::10]:1234 ServerName::try_from Ok(IpAddress(V6(...)))

A zone identifier is refused and never silently reformatted: fe80::1%en0 and [2001:db8::1%eth0] both give InvalidLiteral. Link-local unicast is not a candidate in the first place.

Simulated IPv6-only host, because this machine has no IPv6-only interface. I gave TcpStreamServer a resolver whose whole inventory is IPv6, then bound and connected for real:

IPV6ONLY advertised=[::1]:61178
IPV6ONLY connected peer=Ok([::1]:61178)
EXPLICITLL  err="... interface has no usable IP address: eth0"   # only fe80::1 present
LINKLOCALONLY advertised=127.0.0.1:61180                          # automatic path, warns
EMPTYINV      advertised=127.0.0.1:61176                          # automatic path, warns
IPV6GLOBAL    err="... Can't assign requested address (os error 49)"

So an empty result or a link-local only interface does not fail startup. It warns and serves loopback. An enumeration error does fail startup, which tcp_server_preserves_interface_enumeration_error pins. The base tree reached the same 127.0.0.1 in those cases, so this is not new.

The new tests are real. Five mutations of the changed code, each turning tests red, with the unmutated tree green at 65 passed:

Mutation Tests that turned red
prefer IPv6 before IPv4 interface_alias_preserves_order..., mapped_interface_addresses...
drop the IPv6 brackets in format_host transient_fallbacks_are_not_cached_but_success_is
drop the IFF_UP filter system_inventory_keeps_only_up_ip_addresses
drop to_canonical in consider mapped_interface_addresses_use_ipv4_classification
drop trim in resolve_host_or_interface 5 tests, including two real binds in tcp::server

The lane is rust-tests (.), .github/workflows/pre-merge.yml:437, cargo test --locked --all-targets. It ran at head f1b5aea3c0 and passed, so the new tests executed.

Three event-plane ZMQ tests fail on this Mac. They fail the same way on the merge base, so they are not from this change.

Two open items, both low, both already stated in the PR body and the docs page. Neither blocks. Skipped on purpose: I did not run this on a real IPv6-only host, and I did not exercise the QUIC response listener under a wildcard host.

Comment thread lib/runtime/src/utils/ip_resolver.rs
Comment thread lib/runtime/src/utils/ip_resolver.rs
@jthomson04
jthomson04 merged commit 99875cb into main Sep 18, 2026
124 checks passed
@jthomson04
jthomson04 deleted the jthomson04/fix-ipv6-ip-resolution branch September 18, 2026 04:29
@A-Ishak

A-Ishak commented Sep 22, 2026

Copy link
Copy Markdown

pls release 🙏

aung-san-i added a commit to aung-san-i/dynamo that referenced this pull request Sep 28, 2026
* feat: KV DC Relay file based source mode (ai-dynamo#14807)

Add live-reloaded file sources for KV DC Relay namespace selection and expose readiness and source revisions through /engine/state.

Preserve applied membership on invalid updates, coalesce discovery refreshes, and isolate native integration tests in forked processes.

Signed-off-by: Nikita Sukharev <kaonael@gmail.com>

* feat(sglang): expose cross-encoder reranking through /v1/rerank (ai-dynamo#14032)

Signed-off-by: xianlubird <xianlubird@gmail.com>

* fix(profiler): explain inaccessible model paths during trust checks (ai-dynamo#14860)

Signed-off-by: hongkuanz <hongkuanz@nvidia.com>

* fix(sglang): sync discovery from native pause state (ai-dynamo#13951)

Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
Co-authored-by: Zero Rains <57100978+zeroRains@users.noreply.github.com>

* feat(recipes): add Solar Open2 250B NVFP4 aggregated and disaggregated recipes for B200 (ai-dynamo#14376)

Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com>

* refactor(agents): session_id reader from AgentContext + forward to vLLM (ai-dynamo#14428)

Signed-off-by: Karen Chung <karenc@nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* fix(discovery): allow served aliases for the same model source (ai-dynamo#14857)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix(router): reject unknown explicit worker targets (ai-dynamo#14858)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix(xpu): stabilize XPU test workers (ai-dynamo#14539)

Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>
Signed-off-by: VincyZhang <wenxin.zhang@intel.com>

* feat(mm-routing): add Nemotron 3 Nano Omni video routing (ai-dynamo#14653)

Signed-off-by: krishung5 <krish@nvidia.com>

* fix(sglang): validate diffusion input_reference and bound media fetches (ai-dynamo#14435)

The sglang image-diffusion and video-generation handlers passed the
client-supplied input_reference through to the generator's image_path after only
a non-empty check. Validate it first, and for remote references materialize it
locally before the generator sees it, so the generator is always handed a
trusted local path. This brings the sglang diffusion path in line with the
vLLM/omni and trtllm backends, which already validate the same field.

Behavior change: local I2I/I2V references now require DYN_MM_LOCAL_PATH to be
set to the allowed directory; previously any path was accepted.

common/http:

- validate_media_reference() returns a plain filesystem path for local
  references; local_media_reference() is an async context manager that fetches a
  remote one through fetch_bytes(policy=...), which revalidates every redirect
  hop, into a temp file removed on exit. data: is rejected -- a URI is not a path.
- fetch_bytes() gained max_bytes, streaming through collect_capped at an explicit
  read granularity so the cap is an allocation bound and not only a rejection: a
  128 MiB-decoded gzip body against the 64 MiB cap peaks at 68,032,217 bytes
  rather than the whole decompressed body. Content-Length is caller-controlled
  and absent when chunked, and aiohttp's read(n) returns at most n bytes, so
  neither a header check nor a single capped read suffices. Defaults to None,
  leaving existing callers unchanged.
- DYN_MM_MAX_FILE_SIZE_MB makes that cap operator-tunable, in megabytes, as the
  SGLang arg it replaces was. Read per call; empty, unparseable or non-positive
  falls back to 64 with a warning, so a malformed value neither takes the worker
  down nor reads as unlimited.
- Messages built from caller input are bounded via describe_media_source, moved
  from multimodal/media_source.py (it pulls in torch) into url_validator.py and
  re-exported from its old home; a no-op below 120 characters.
- HttpStatusError bounds its .message attribute, not only the rendered string:
  errors.rs::extract_http_like_error reads .status and .message off this class by
  name and forwards .message on a 4xx without calling str(). Backend exception
  text is bounded head-and-tail, since aiohttp renders the host before the errno.
- validate_local_path uses exc.strerror rather than the raw OSError, whose text
  repeats the filename, and now catches the ValueError that Path.resolve() raises
  on an embedded NUL so callers keep their 4xx-vs-5xx decision.

Rebased onto ai-dynamo#14563 (single aiohttp backend); the httpx-side half of the
max_bytes plumbing went with that backend.

Signed-off-by: nnshah1 <neelays@nvidia.com>
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(deps): upgrade fastokens to 0.3.2 (ai-dynamo#14798)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix(vllm): ship codec-free OpenCV for image inputs (ai-dynamo#14361)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: Anant Sharma <anants@nvidia.com>
Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com>

* docs: refresh community events

Automated refresh from the public Dynamo Google Calendar.

Generated by .github/workflows/community-events-refresh.yml.

Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com>

* ci: refresh the compliance baseline in auto-upgrade pipeline (ai-dynamo#14206)

Signed-off-by: Anant Sharma <anants@nvidia.com>

* feat(triton): honor KServe classification on tensor outputs (ai-dynamo#14783)

Signed-off-by: Yingge He <yinggeh@nvidia.com>

* docs(rl): stop the verl guide sending readers to a vLLM version it cannot run on (ai-dynamo#14571)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>

* feat(mocker): publish native KV events from the vLLM gRPC server (ai-dynamo#14737)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix(kv-router): release unowned radix branches after eviction (ai-dynamo#14878)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix: show correct backend versions in the install selectors (ai-dynamo#13599)

Signed-off-by: Anant Sharma <anants@nvidia.com>

* build(vllm): prepare v0.29.0 bump (ai-dynamo#14543)

Signed-off-by: Julien Darve <jdarve@NVIDIA.com>

* ci(xpu): validation PR for the re-applied XPU workflows and Dockerfile

Throwaway PR to prove the CI merged in #22 actually runs end to end on XPU
hardware. Adds only a comment to container/templates/vllm_runtime.Dockerfile,
which matches the `vllm` path filter (container/templates/vllm_*) and so makes
changed-files set vllm=true, which is what gates build-xpu and the
heterog-test-px-dn / heterog-test-pn-dx jobs.

What this exercises:
  - .github/workflows/pr-xpu.yaml            (push to pull-request/[0-9]+, needs the xpu label)
  - .github/workflows/pr-xpu-heterogeneous.yaml (push; its guard deliberately skips the label gate)
  - .github/workflows/epd-test-template.yml  (workflow_call, from the heterog jobs)
  - .github/scripts/test-filters.js          (the brace fix from #22)
  - container/templates/vllm_runtime.Dockerfile rendered and built for device=xpu

Not exercised: .github/workflows/xpu-heterogeneous-dispatch.yaml is
workflow_dispatch only and has to be run by hand from the Actions tab.

The marker comment must be removed before this branch is ever merged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(triton): Update Triton Base Image to 26.08 (ai-dynamo#14854)

Signed-off-by: J Wyman <jwyman@nvidia.com>
Co-authored-by: Rini Gupta <rinig@nvidia.com>

* fix(operator): normalize equivalent worker hash inputs (ai-dynamo#14721)

Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>

* test(sglang): exercise NIXL in embedding cache E/PD test (ai-dynamo#14795)

Signed-off-by: Sai Kiran Polisetty <spolisetty@nvidia.com>

* fix(sglang): stop the elastic-EP scale-up worker crash-looping at startup (ai-dynamo#14568)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com>

* fix(responses): honor tool_choice when parsing tool calls from text (ai-dynamo#14843)

Signed-off-by: xianlubird <xianlubird@gmail.com>

* ci: accept trusted full-CI request comments (ai-dynamo#14868)

Signed-off-by: Matej Kosec <mkosec@nvidia.com>

* docs: clarify EPP mode boundary and single-replica Dynamo mode fixes [DYN-4310] (ai-dynamo#14756)

Signed-off-by: Anna Tchernych <atchernych@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci(docs): move the generated-tables determinism gate out of link checking (ai-dynamo#14135)

Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* ci(docs): generate the Kubernetes API reference at publish time (ai-dynamo#14122)

Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(operator): discover pull secrets for init containers (ai-dynamo#14922)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* fix(sglang): stop an unusable mooncake backend crashing workers after model load (ai-dynamo#14461)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: glamr-agent <glamr-agent@users.noreply.github.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>

* fix(sglang): emit prefill handoff before completion in sidecar (ai-dynamo#14260)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* test(trtllm): enable fault tolerance coverage (ai-dynamo#14609)

Signed-off-by: tanmayv25 <tanmay2592@gmail.com>

* fix(frontend): evict async tokenizer executors when the tokenizer is retired (ai-dynamo#13368)

Signed-off-by: Peter Pan <Peter.Pan@daocloud.io>

* fix(llm): report KServe datatypes by their wire names, not protobuf variants (ai-dynamo#14957)

`ModelMetadata` reported each Triton-registered tensor's `datatype` using
`inference::DataType::as_str_name()`, which returns the `model_config.proto`
variant name (`TYPE_FP32`, `TYPE_STRING`, ...) instead of the KServe v2 wire
names (`FP32`, `BYTES`, ...). Every datatype was wrong, so spec-conforming
clients cannot parse any tensor the RPC describes. Adds `oip_name()` next to
`tensor::DataType::to_kserve` covering all fifteen proto variants (incl. FP16
and BF16) and mapping `TYPE_STRING → BYTES`.

Original PR by @ayaangazali: ai-dynamo#14770. Reissued under a signed commit to
unblock the copy-pr-bot signature gate; diff is byte-identical.

Closes ai-dynamo#14520.

Signed-off-by: ayaangazali <ayaangazali@users.noreply.github.com>
Signed-off-by: ayaangazali <ayaangazali.work@gmail.com>
Signed-off-by: Vinya Kestur <vinyak@nvidia.com>
Co-authored-by: ayaangazali <ayaangazali.work@gmail.com>

* docs(mm-routing): document video KV routing (ai-dynamo#14958)

Signed-off-by: krishung5 <krish@nvidia.com>

* fix(sidecar): honor worker namespace suffix (ai-dynamo#14955)

Signed-off-by: Biswa Panda <biswa.panda@gmail.com>

* fix(bindings): drain bridge tasks before interpreter finalization (ai-dynamo#14813)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: Tushar Sharma <tusharma@nvidia.com>

* fix(discovery): stop a Qwen3-VL worker from serving video with another worker's contract (ai-dynamo#14624)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>

* fix(gms): honor configured timeout during initial weights admission (ai-dynamo#14877)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>

* feat(kv-router): add construction-time indexer delegates (ai-dynamo#14945)

* fix(sglang): support min_tokens on tokenizer-free decode workers (ai-dynamo#14276)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: MatejKosec <mkosec@nvidia.com>

* feat(router): add SessionPrefixIndexer for session-block lineage (ai-dynamo#13807)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: Karen Chung <karenc@nvidia.com>
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
Co-authored-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Co-authored-by: Matej Kosec <mkosec@nvidia.com>

* fix(vllm): settle kvwarm stages through a per-step round on every attention-DP rank (ai-dynamo#14728)

Signed-off-by: Yiming Liu <yimingl@nvidia.com>

* feat(vllm): benchmark hybrid caches with random KDA state (ai-dynamo#14900)

Signed-off-by: hongkuanz <hongkuanz@nvidia.com>

* fix(runtime): fix QUIC reassembly and reduce response stalls (ai-dynamo#14876)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* feat(router): unify frontend and standalone selection core (ai-dynamo#14570)

Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
Signed-off-by: Thomas Montfort <tjmontfort12@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Thomas Montfort <tjmontfort12@gmail.com>

* fix(planner): keep control APIs responsive during Prometheus collection (ai-dynamo#14377)

Signed-off-by: xianlubird <xianlubird@gmail.com>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>

* fix(router): record SGLang prefill completion after stream ends (ai-dynamo#14968)

Signed-off-by: jain-ria <riajain@NVIDIA.com>

* fix(frontend): send inline media once on the TCP request plane (ai-dynamo#14801)

Signed-off-by: Sumit Mishra <sah299610@gmail.com>
Co-authored-by: Indrajit Bhosale <iamindrajitb@gmail.com>

* docs: refresh community events

Automated refresh from the public Dynamo Google Calendar.

Generated by .github/workflows/community-events-refresh.yml.

Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com>

* fix(vllm): initialize synchronizer in KV warmup capacity test (ai-dynamo#14984)

Signed-off-by: Alec Flowers <aflowers@nvidia.com>

* fix(recipes): make the Solar Open2 250B benchmark and docs link usable (ai-dynamo#14956)

Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com>

* feat(recipes): add K-EXAONE 2.0 750B-A37B NVFP4 vLLM recipes for B200 (ai-dynamo#14822)

Signed-off-by: Cheng Wang <chengwa@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: KVCR Resiliency Deployment Example (ai-dynamo#14695)

Add two-node DynamoGraphDeployment examples for process-local KVCR and
the KVCR memory service. Run one vLLM worker per GPU node, use stable
Grove ordinals for cache-owner slots, and request GPU-local RDMA
resources for engines and Guard services. Provide a deployment helper
for rendering and selecting either variant.

Run the KV state agent alongside vLLM for process-local host memory. In
memory-service mode, keep KVCR and the state agent in a separate
container so its Guard and shared-memory pool survive engine restarts.
Document that restarting the services sidecar invalidates the MVP
recovery contract and requires deployment-level replacement.

Add manifest coverage and an opt-in two-host lifecycle test. Kill the
source EngineCore, hold it offline, and verify that the promoted Guard
serves its preserved cache to the surviving target. Correlate response
equality and KVCR transfer metrics with transmit and receive counters
from the selected active HCA to prove RDMA transport.

Pin compatible KVCR and vLLM revisions and document the runtime,
discovery, compatibility-digest, and recovery prerequisites.

Signed-off-by: Adit Ranadive <aranadive@nvidia.com>

* feat(omni): add Nemotron Audex speech synthesis to /v1/audio/speech (ai-dynamo#12788)

Signed-off-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com>

* ci: allow glamr-agent to request CI on its own unsigned PRs (ai-dynamo#14964)

Signed-off-by: Matej Kosec <mkosec@nvidia.com>

* fix(vllm): isolate multimodal worker ports (ai-dynamo#14751)

Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>

* fix(runtime): reject invalid DYN_REQUEST_PLANE values (ai-dynamo#12612)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
Signed-off-by: Coding Agent <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: MatejKosec <mkosec@nvidia.com>

* fix(responses): preserve text instead of inferring tool calls (ai-dynamo#14846)

Signed-off-by: xianlubird <xianlubird@gmail.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>

* chore: temporarily increase frontend build time limit 45 --> 90 min (ai-dynamo#15019)

Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>

* test(operator): cover scoped CA injection ownership (ai-dynamo#14961)

Signed-off-by: Julien Mancuso <jmancuso@nvidia.com>

* feat(frontend): map semantic errors to HTTP responses (ai-dynamo#14396)

Signed-off-by: Biswa Panda <biswa.panda@gmail.com>

* docs: correct fault-tolerance architecture details (ai-dynamo#14880)

Signed-off-by: Elizabeth Thomas <email2eliza@gmail.com>

* build(deps): bump nats-server to v2.14.7 (ai-dynamo#14919)

Signed-off-by: Dan Gil <dagil@nvidia.com>

* build(deps): bump AISimulate to 0.12.0 (ai-dynamo#15012)

Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>

* remove oneAPI env for XPU detection

* feat(backends): expose native LoRA capacity in model registration (ai-dynamo#14754)

Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>
Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com>

* fix(planner): handle pending decisions in virtual connector wait (ai-dynamo#14841)

Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>

* feat(vllm): add sidecar LoRA lifecycle (ai-dynamo#13068)

Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>
Co-authored-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com>

* fix(vllm/omni): pass response_format into video EngineInputs (ai-dynamo#14667) (ai-dynamo#14844)

* chore: bump version to 1.6.0 post 1.5.0 branch cut (ai-dynamo#15009)

Signed-off-by: pvijayakrish <pvijayakrish@nvidia.com>
Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): Use `pytest --ignore` to Skip Tests Based on Framework (ai-dynamo#14815)

Signed-off-by: J Wyman <jwyman@nvidia.com>

* feat(sidecar): add e2e CI testing for sidecar launch scripts (ai-dynamo#14508)

Signed-off-by: tanmayv25 <tanmay2592@gmail.com>
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: Julien Darve <jdarve@NVIDIA.com>

* chore(xpu): upgrade vllm and omni to 0.29.0

Signed-off-by: wenxin.zhang <wenxin.zhang@intel.com>

* docs(operator): document the DGDR workload-creation trust boundary (ai-dynamo#14429)

Signed-off-by: nnshah1 <neelays@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* fix(xpu): use released vllm-omni prerelease

Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>

* test(efa): add the EFA disaggregated deploy test for sglang (ai-dynamo#13893)

Signed-off-by: Jie Hao <jihao@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(runtime): support IPv6-only IP resolution (ai-dynamo#13126)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* docs(fault-tolerance): clarify migration after shutdown grace expires (ai-dynamo#14872)

Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>

* feat(vllm-omni): preserve generated video audio (ai-dynamo#13707)

Signed-off-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>

* feat(vllm-omni): pass model-specific video parameters (ai-dynamo#13708)

Signed-off-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>

* feat(vllm-omni): qualify MiniMax-H3 T2VA on B200 (ai-dynamo#13589)

Signed-off-by: Guan Luo <gluo@nvidia.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>

* fix(vllm): remove obsolete Omni compatibility guard

Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>

* fix(vllm): retain Omni compatibility guard

Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>

* .github/workflows/pr-xpu-heterogeneous.yaml; pin GPU_TAG to latest

* .github/workflows/; add post-merge and nightly XPU heterogeneous CI

Extract the XPU heterogeneous P/D pipeline out of pr-xpu-heterogeneous.yaml
into xpu-heterogeneous-run.yml, a workflow_call reusable workflow, and call it
from three thin trigger workflows so all three merge phases run the identical
pipeline instead of drifting copies.

  xpu-heterogeneous-run.yml           new, reusable. guard, changed-files,
                                      build-xpu, build-nvidia, resolve-images
                                      and both heterog tests, unchanged, plus
                                      7 inputs.
  pr-xpu-heterogeneous.yaml           reduced to the pre-merge trigger, the
                                      slash-command gate and the reaction.
  post-merge-xpu-heterogeneous.yaml   new. push to main.
  nightly-xpu-heterogeneous.yaml      new file, but the cron is MOVED, not
                                      added: it is the 0 23 * * * schedule
                                      that was already in
                                      pr-xpu-heterogeneous.yaml.

No behaviour change per phase. force_all_tests replaces the old
  github.event_name == 'schedule' || github.event_name == 'issue_comment'
expression with the same truth table: pre-merge passes
github.event_name == 'issue_comment', nightly passes true. Post-merge also
passes true, because a push to main has no PR base for
.github/actions/changed-files to diff against, and post-merge exists to catch
what per-PR gating missed.

xpu-status-check stays a TOP-LEVEL job in each caller rather than moving into
the reusable workflow. A job contributed by a reusable workflow reports to the
Checks API as "run / xpu-status-check", so hosting it there would rename the
context and leave any branch protection rule requiring xpu-status-check waiting
forever on a check that no longer reports.

The concurrency mapping stays byte-identical across all four workflows that
touch this hardware, now including xpu-heterogeneous-dispatch.yaml. Three files
do NOT get three slots: the cluster, the dynamo-system namespace and the
onexpu-/onenvidia-rdma-kueue ResourceClaimTemplates are one global resource.
The reusable workflow deliberately carries no concurrency block of its own,
which would deadlock against the slot the caller's run already holds.

Parameterised gpu_tag, model, tensor_parallel and runner as inputs so the
callers can diverge; all default to the previously hardcoded values. Added
workflow_dispatch to the nightly, without which a schedule-only workflow cannot
be exercised before it reaches the default branch.

Verified: all files parse; the four concurrency mappings are byte-identical; the
reusable workflow declares no concurrency; every input each caller passes exists
and every required input is supplied; nesting is depth 3 of the 4 GitHub allows.
actionlint was not available to run, and will report queue:max as an unknown key
in all four files, a known false positive.

---------

Signed-off-by: Nikita Sukharev <kaonael@gmail.com>
Signed-off-by: xianlubird <xianlubird@gmail.com>
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com>
Signed-off-by: Karen Chung <karenc@nvidia.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>
Signed-off-by: VincyZhang <wenxin.zhang@intel.com>
Signed-off-by: krishung5 <krish@nvidia.com>
Signed-off-by: nnshah1 <neelays@nvidia.com>
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com>
Signed-off-by: Anant Sharma <anants@nvidia.com>
Signed-off-by: Yingge He <yinggeh@nvidia.com>
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: J Wyman <jwyman@nvidia.com>
Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>
Signed-off-by: Sai Kiran Polisetty <spolisetty@nvidia.com>
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
Signed-off-by: Anna Tchernych <atchernych@nvidia.com>
Signed-off-by: Dan Gil <dagil@nvidia.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: glamr-agent <glamr-agent@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Signed-off-by: tanmayv25 <tanmay2592@gmail.com>
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io>
Signed-off-by: ayaangazali <ayaangazali@users.noreply.github.com>
Signed-off-by: ayaangazali <ayaangazali.work@gmail.com>
Signed-off-by: Vinya Kestur <vinyak@nvidia.com>
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
Signed-off-by: Yiming Liu <yimingl@nvidia.com>
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
Signed-off-by: Thomas Montfort <tjmontfort12@gmail.com>
Signed-off-by: Sumit Mishra <sah299610@gmail.com>
Signed-off-by: Alec Flowers <aflowers@nvidia.com>
Signed-off-by: Cheng Wang <chengwa@nvidia.com>
Signed-off-by: Adit Ranadive <aranadive@nvidia.com>
Signed-off-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com>
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Signed-off-by: Coding Agent <svc-glamr@nvidia.com>
Signed-off-by: Julien Mancuso <jmancuso@nvidia.com>
Signed-off-by: Elizabeth Thomas <email2eliza@gmail.com>
Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>
Signed-off-by: pvijayakrish <pvijayakrish@nvidia.com>
Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
Signed-off-by: wenxin.zhang <wenxin.zhang@intel.com>
Signed-off-by: Jie Hao <jihao@nvidia.com>
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
Signed-off-by: Guan Luo <gluo@nvidia.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Nikita Sukharev <kaonael@gmail.com>
Co-authored-by: Xianlu Bird <xianlubird@gmail.com>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
Co-authored-by: Zero Rains <57100978+zeroRains@users.noreply.github.com>
Co-authored-by: snarravula-dl <snarravula@nvidia.com>
Co-authored-by: Karen Chung <karenc@nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: VincyZhang <wenxin.zhang@intel.com>
Co-authored-by: Kris Hung <krish@nvidia.com>
Co-authored-by: Neelay Shah <neelays@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: Anant Sharma <anants@nvidia.com>
Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com>
Co-authored-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com>
Co-authored-by: Yingge He <157551214+yinggeh@users.noreply.github.com>
Co-authored-by: JulienDarve <86800349+JulienDarve@users.noreply.github.com>
Co-authored-by: J Wyman <jwyman@nvidia.com>
Co-authored-by: Rini Gupta <rinig@nvidia.com>
Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com>
Co-authored-by: Sai Kiran Polisetty <spolisetty@nvidia.com>
Co-authored-by: MatejKosec <mkosec@nvidia.com>
Co-authored-by: atchernych <atchernych@nvidia.com>
Co-authored-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Bojiang Li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Connor Carpenter <connorcarpenter15@gmail.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Tanmay Verma <tanmayv@nvidia.com>
Co-authored-by: Peter Pan <peter.pan@daocloud.io>
Co-authored-by: Vinya Kestur Tumakuru Arun Kumar <vinyak@nvidia.com>
Co-authored-by: ayaangazali <ayaangazali.work@gmail.com>
Co-authored-by: Biswa Panda <biswa.panda@gmail.com>
Co-authored-by: Tushar Sharma <tusharma@nvidia.com>
Co-authored-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
Co-authored-by: Ryan Olson <ryanolson@users.noreply.github.com>
Co-authored-by: Yimingl_Nvidia <yimingl@nvidia.com>
Co-authored-by: Thomas Montfort <tjmontfort12@gmail.com>
Co-authored-by: Sumit884-byte <sah299610@gmail.com>
Co-authored-by: Indrajit Bhosale <iamindrajitb@gmail.com>
Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Co-authored-by: chw001 <chengwa@nvidia.com>
Co-authored-by: Adit Ranadive <aranadive@nvidia.com>
Co-authored-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com>
Co-authored-by: Keiven C <213854356+keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
Co-authored-by: Dmitry Tokarev <dtokarev@nvidia.com>
Co-authored-by: Julien Mancuso <161955438+julienmancuso@users.noreply.github.com>
Co-authored-by: Elizabeth Thomas <email2eliza@gmail.com>
Co-authored-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
Co-authored-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: Jasim Kareem <mj9034812@gmail.com>
Co-authored-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
Co-authored-by: Jie Hao <jihao@nvidia.com>
Co-authored-by: Jacky <18255193+kthui@users.noreply.github.com>
Co-authored-by: Qi Wang <qiwa@nvidia.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
@dagil-nvidia

Copy link
Copy Markdown
Collaborator

Thanks for the push on this. The IPv6 fix, plus the follow-up in #15324 that it needs, is too large to hand-port safely into 1.5.1, which ships today. Both are on main and will be in 1.6.0. Until then, the nightly containers have them.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation fix size/XXL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG]: Dynamo fails in IPv6-only environments — "ifa_prefixlen must be initialized"

8 participants