Skip to content

fix(bindings): drain bridge tasks before interpreter finalization - #14813

Merged
nv-tusharma merged 15 commits into
ai-dynamo:mainfrom
glamr-agent:dyn-4434-flaky-test-fetch-model-runtime-bridge-2a14ec73f4cb
Sep 16, 2026
Merged

nv-tusharma merged 15 commits into
ai-dynamo:mainfrom
glamr-agent:dyn-4434-flaky-test-fetch-model-runtime-bridge-2a14ec73f4cb

Conversation

@glamr-agent

@glamr-agent glamr-agent commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Reduce the observed fetch-model bridge shutdown race by giving existing Tokio tasks a bounded opportunity to finish before Python interpreter finalization.

Summary

  • Register an atexit hook that releases the GIL while polling the process runtime and any distinct, recorded PyO3 bridge runtime, deduplicating shared runtimes.
  • Route all binding future conversions and direct bridge-runtime access through shared recording helpers, so bridge-only processes receive the bounded exit drain without creating a Dynamo process runtime.
  • Centralize bridge runtime adoption for fetch-model setup, backend workers, and both DistributedRuntime constructors, including the context-first then detached path.

Details

The exit hook waits at most 250 ms and logs at info level if tasks remain. Processes with persistent service tasks can pay that entire additional shutdown delay. At the deadline, finalization proceeds with the same residual exposure as before the hook; this is a mitigation, not a guarantee that arbitrary pending Python cleanup is safe.

Worker::existing_process_runtime() reads the process runtime without creating worker threads during exit. Bridge adoption records whichever runtime PyO3 already selected, preserving a distinct context-first runtime. The detached constructor now uses that same adoption helper before wrapping the process runtime.

The test file is unchanged from main. The earlier port-reservation fixture and allocator test changes have been removed. Coverage for the combined context-first then detached-runtime order remains deferred to DYN-4434, with two review threads awaiting maintainer disposition.

Where should the reviewer start?

Review bridge_runtime, the future-conversion helpers, adopt_bridge_runtime, wait_for_bridge_tasks_at_exit, and both DistributedRuntime constructors in lib/bindings/python/rust/lib.rs, followed by lib/runtime/src/worker.rs.

Validation

Prior factory comments report measurements and local validation of earlier commits; those results do not validate the current head. Nursery validation uses remote CI only. For 1daee8fdbf44395d91613c209e84dff4df2ea23e, the September 16 CI snapshot has 65 successful check runs, 50 skipped, and one cancelled, with no queued or running checks. All four Rust test jobs, all four Clippy jobs, remote pre-commit, and the pre-merge gate passed. Full CI was authorized and ran, but planner / Compliance cpu, arm64 hit its 15-minute limit during image-package extraction before scanning. The targeted retry was rejected with HTTP 403 (repository admin rights required), so full CI is incomplete. This maintenance pass ran no local tests, builds, or benchmarks and added no tests.

Related Issues

🔗 This PR is linked to an issue:

  • Relates to DYN-4434. Do not automatically close the investigation: the bounded drain leaves residual shutdown-race exposure.

Summary by CodeRabbit

  • Improvements
    • Improved coordination between Python asynchronous operations and the application runtime.
    • Added more reliable cleanup of active asynchronous work when the process exits, with brief waiting for tasks to finish.
    • Improved compatibility with existing runtime configurations, including clearer warnings for mismatched runtime setups.
    • Standardized asynchronous behavior across services, model operations, streaming, scaling, and shutdown workflows without changing interfaces or error handling.

…ed port pool

test_fetch_model_runtime_bridge_orders[invalid-config-then-fetch] was
reported flaky with an environment_dependency category. The test body
took dynamo_dynamic_ports for every parametrization, so all six
scenarios reserved five host-wide ports (frontend, system, KV event,
FPM, NIXL) through the shared allocator, which probes candidates with a
real bind() and raises "Could not find N available ports after 100
retries" when the host is busy. Only fetch_then_backend_worker actually
binds a socket; the other five run with DYN_SYSTEM_PORT=-1.

Resolve the port lazily through request.getfixturevalue so only the
scenario that needs one touches the allocator, raise the child
subprocess budget from 20s to 60s (and the outer pytest timeout from
30s to 120s so the inner bound is what fires), and turn
subprocess.TimeoutExpired into a pytest.fail carrying the partial child
stdout and stderr so the next occurrence is classifiable instead of a
bare exception.

Adds a regression test asserting the invalid-config scenario still
passes with the allocator patched to raise, plus a negative control
asserting the backend scenario still fails without a port.

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
@glamr-agent
glamr-agent requested a review from a team as a code owner September 14, 2026 20:05
@copy-pr-bot

copy-pr-bot Bot commented Sep 14, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@glamr-agent
glamr-agent deployed to external_collaborator September 14, 2026 20:05 — with GitHub Actions Active
@glamr-agent
glamr-agent deployed to external_collaborator September 14, 2026 20:05 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi glamr-agent! Thank you for contributing to ai-dynamo/dynamo.

Just a reminder: The NVIDIA Test Github Validation CI runs an essential subset of the testing framework to quickly catch errors.Your PR reviewers may elect to test the changes comprehensively before approving your changes.

🚀

@github-actions github-actions Bot added test external-contribution Pull request is from an external contributor labels Sep 14, 2026
@glamr-agent

Copy link
Copy Markdown
Contributor Author
Automated evidence record — validation complete

Validation status: complete

Evidence summary: [3/3 validated]

AI review assessment (advisory, not an approval): sound. An AI agent judged the change logically sound from the code and the reported validation results. Repository CI and human reviewers decide whether to merge.

Validation result: pass — every command quoted in the change description reproduced here, including both rewritten paragraphs.

Evidence audit: complete [3/3 validated] — the command report below comes from recorded runs.

Commands and results [3/3 validated]

Generated from the commands recorded during this run.

Check 1

Builds the changed Dynamo source and confirms that Python can import its compiled extension.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

Check 2

Checks the changed files with the repository's fast lint and formatting commands.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

Check 3

Runs the relevant Python unit tests without requiring a GPU.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

No description provided.

@coderabbitai

coderabbitai Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: bbe5364f-6a0c-43a3-9a2f-7b741e39f807

📥 Commits

Reviewing files that changed from the base of the PR and between 95cadce and 1daee8f.

📒 Files selected for processing (14)
  • lib/bindings/python/rust/backend.rs
  • lib/bindings/python/rust/context.rs
  • lib/bindings/python/rust/http.rs
  • lib/bindings/python/rust/kserve_grpc.rs
  • lib/bindings/python/rust/lib.rs
  • lib/bindings/python/rust/llm/entrypoint.rs
  • lib/bindings/python/rust/llm/kv.rs
  • lib/bindings/python/rust/llm/kv/demand_driven.rs
  • lib/bindings/python/rust/llm/kv_dc_relay.rs
  • lib/bindings/python/rust/llm/kv_state_agent.rs
  • lib/bindings/python/rust/llm/lora.rs
  • lib/bindings/python/rust/llm/routed_engine.rs
  • lib/bindings/python/rust/planner.rs
  • lib/runtime/src/worker.rs

Included review availability: Your plan provides up to 12 included reviews per hour; 9 remain after this review.


Walkthrough

The Python bindings now use centralized future conversion and bridge-runtime helpers. Runtime adoption records process and distinct runtimes. An atexit hook drains active tasks for up to 250 ms.

Changes

Runtime bridge centralization

Layer / File(s) Summary
Runtime adoption and access
lib/runtime/src/worker.rs, lib/bindings/python/rust/backend.rs, lib/bindings/python/rust/lib.rs, lib/bindings/python/rust/llm/kv.rs
The runtime exposes a non-initializing process-runtime accessor. Worker creation, model fetching, distributed runtime creation, and KV routing use shared bridge runtimes.
Python exit drain
lib/bindings/python/rust/lib.rs
The binding records bridge runtimes, polls active tasks for up to 250 ms without holding the GIL, logs remaining tasks, and registers the drain with Python’s atexit module.
Shared Python async bridge
lib/bindings/python/rust/*.rs, lib/bindings/python/rust/llm/*.rs, lib/bindings/python/rust/llm/kv/*.rs
Asynchronous Python bindings use crate::future_into_py and crate::bridge_runtime() across services, model operations, endpoints, streams, KV operations, planner operations, and LLM operations.

Priority: ⬆️ High

Estimated code review effort: 4 (Complex) | ~45 minutes

Merge Risk: ⚪ Minimal · up to 1daee

The runtime tracking paths now include both normal and detached distributed-runtime initialization, so the previously identified exit-drain gap is addressed and no actionable merge-blocking risk remains.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 77.42% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 93 functions across 15 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the main change: draining bridge tasks before Python interpreter finalization.
Description check ✅ Passed The description includes the required overview, details, reviewer starting points, and related issue information. It also states validation status and limitations.
  • Fix all pre-merge checks with AI

Comment @coderabbitai help to get the list of available commands.

@glamr-agent

glamr-agent commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor Author

CI reached a terminal state on 3620eb9: 14 checks passed and 9 were skipped, with none failing or pending. The heavier jobs (rust-tests, rust-clippy, operator, and the docs checks) report skipping because the full suite runs on the pull-request/N branch that copy-pr-bot creates rather than on the fork branch; to schedule them, a maintainer needs to comment /ok to test 3620eb9 with this head SHA.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

factory: automated evidence record for the follow-up measurement run on this pull request. It supersedes the earlier automated evidence record on this request, which covered the original change and reported a passing validation.

Automated evidence record — validation incomplete

Validation status: incomplete

Evidence summary: [2/3 validated · 1 failing]

AI review assessment (advisory, not an approval): needs changes. An AI agent judged that the recorded measurement does not support the change as described. Repository CI and human reviewers decide whether to merge.

Validation result: fail — the runs are green but do not support the fix; at 1800 runs per arm the change shows no reduction in failures, and the child SIGSEGV reproduced on the changed arm with no port allocated.

Evidence audit: complete [2/3 validated · 1 failing] — the command report below comes from recorded runs.

Commands and results [2/3 validated · 1 failing]

Generated from the commands recorded during this run.

Check 1

Builds the changed Dynamo source and confirms that Python can import its compiled extension.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

Check 2

Checks the changed files with the repository's fast lint and formatting commands.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

Check 3

Runs the relevant Python unit tests without requiring a GPU.

Result: Passed (exit 0), but did not validate the change

Command:

Not shown because the exact command contained private run data.

Details:

The test runs completed green, but at 1800 runs per arm the changed test failed 14 times against 10 for unchanged main (p = 0.41), and the child SIGSEGV (returncode=-11) reproduced on the changed arm with system_port = 0, so these green runs do not support the claim that the flake is fixed.

@glamr-agent glamr-agent left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated AI review — advisory. The inline comments below are the findings. The full review is in a separate comment on this pull request.

Comment thread tests/runtime/test_fetch_model_runtime_bridge.py Outdated
Comment thread tests/runtime/test_fetch_model_runtime_bridge.py Outdated
Comment thread tests/runtime/test_fetch_model_runtime_bridge.py Outdated
@glamr-agent

glamr-agent commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor Author

factory: automated review

🤖 Automated AI review — advisory. An AI agent's judgment of
whether this change is logically sound based on the code and reported
validation results. This is not an approval. Repository CI and human
reviewers decide whether to merge.

Assessment: sound

Second pass. Last time I said the fix was right and the description was wrong:
it claimed the exit hook costs "no measurable delay" when a real frontend pays
its full 1 second bound on shutdown. The description has been corrected and the
code is untouched, so this now reads honestly. Nothing further from me.

The code did not move. git rev-parse HEAD is still
b07f875c1821d853f88799e0ffe7e9762b92635d, git status --short is empty, and
the diff over the branch hashes to the same bytes I reviewed before. Against
main, git diff --stat origin/main...HEAD is 3 files, 76 insertions(+), 2 deletions(-).

The three things I asked for are in.

  1. The exit cost is now stated as measured. A short-lived process whose tasks
    drain pays nothing — and those timings are labelled as a short-lived process,
    not a general claim. A process with long-lived service tasks waits the full
    bound: an observed frontend SIGTERM exit of 1.12s, with the hook's own
    message reporting alive_tasks=5 at the deadline. The consequence is in bold
    where a reviewer will see it: every long-running Dynamo process now pays about
    one extra second at shutdown.
  2. The scope note is there: because the wait is bounded, a process whose tasks
    never drain finalizes with tasks still alive (lib.rs:208-214), the same
    exposure it has today, and that is not a regression.
  3. The title and routing note are there — fix(bindings): drain bridge tasks before the interpreter finalizes, and a warning that touching
    lib/bindings/python/rust/ and lib/runtime/ pulls in different code owners
    than the tests-only version did. You can check who that is with
    python .github/codeowners/who_owns.py --codeowners CODEOWNERS --changed.

Two precisions on the new text, neither worth a round. The deadline line is
debug-level, so it was only visible in the runs where the log level was raised
— two of them, and it fired in both. Read strictly, "reported at its deadline on
every teardown" is scoped by the "with the log level raised" clause in front of
it and is true; it just rests on two observations. And "essentially all of the
1.12s is this wait" is an inference from the 1 second bound at
lib/bindings/python/rust/lib.rs:180 plus that deadline message, not a
subtraction of two measured numbers — nobody timed a frontend exit with the hook
disabled. A 1.12s total against a 1s bound leaves little room for anything
else, so it is the right expectation to give a reader.

What I checked myself, and you can run the same.
git diff origin/main...HEAD -- tests/runtime/test_fetch_model_runtime_bridge.py
is 15 insertions(+), 2 deletions(-) and is entirely the new system_port
fixture — the two port-allocator tests, their _break_port_allocation /
_exhausted_port_pool helpers, the importlib import and _partial_output are
gone, and the timeouts match main at pytest.mark.timeout(30) and
timeout=20. Every claim the description makes about the code holds against the
diff: the early return at lib.rs:199, the py.allow_threads at lib.rs:205
(load-bearing — an atexit callback holds the GIL and the tasks being waited on
need it to finish), the bounded num_alive_tasks() poll at lib.rs:206-216, and
the bare RT.get() at lib/runtime/src/worker.rs:154 that never builds a
runtime. The comment the description quotes is verbatim from lib.rs:178-179.

The evidence, which I audited rather than re-ran. Three arms at the same
time against the same compiled extension, 150,000 children each, driving the
test's own child program directly:

arm=hook-on    iterations=150000 workers=64
  returncode=0     ok           150000

arm=wrapped-on iterations=150000 workers=64
  returncode=0     ok           150000

arm=hook-off   iterations=150000 workers=64
  returncode=-11   SIGSEGV      126
  returncode=0     ok           149874
  non-zero exits that had already printed the sentinel: 126/126

hook-off differs from wrapped-on only by atexit._clear() after asserting
exactly one callback is registered, so the one callback removed is this change's
hook. At that rate the other 300,000 children would be expected to produce
roughly 250 crashes and produced none. The description's own smaller table
reports 10 in 40,000 at 32 children in flight rather than 64; that is the same
race with less concurrency, not a conflicting number.

All of these exited 0 on one host against this branch: maturin develop --uv
in lib/bindings/python then python3 -c "import dynamo, dynamo._core";
cargo fmt --all --check; cargo check --workspace --all-targets; cargo check --manifest-path lib/bindings/python/Cargo.toml --all-targets; cargo clippy --workspace --all-targets --no-deps -- -D warnings and the same clippy against
the bindings manifest; cargo test -p dynamo-runtime --lib -- --test-threads=1
(748 passed; 0 failed; 2 ignored); python -m pytest tests/runtime/test_fetch_model_runtime_bridge.py -v --tb=short -p no:randomly
(6 passed); pre-commit run --files over the three changed files. The bindings
crate is outside the Cargo workspace, so the two commands naming its manifest are
the ones that compile the hook. Full CI has not run on this head.

Smaller things, your call

  • lib/bindings/python/rust/lib.rs:209 logs the give-up at debug, which is why
    it took a raised log level to see. It fires at most once per process, at exit,
    so tracing::info! would cost nothing and would explain the extra second to
    whoever notices it.
  • lib/bindings/python/rust/lib.rs:180 sets a 1 second bound for a window
    measured as sub-millisecond. Free for a process that drains, paid in full by
    every process that does not. A shorter bound, or counting bridge tasks rather
    than every task on the process runtime, would close the same window for less. I
    would not hold the change for it, and the description already flags both of
    these as choices for you to judge.

Untested, and what would test it

No request completed through the engine. /v1/chat/completions returned HTTP
500 because components/src/dynamo/vllm/handlers.py:1099 lazily runs from vllm.exceptions import VLLMClientError and the vLLM 0.26.0 build on that host
does not define the name:

ImportError: cannot import name 'VLLMClientError' from 'vllm.exceptions'

That is independent of this change and fails the same way without it. What it
leaves untested is the exit hook in a worker that has actually served traffic —
its task mix at exit differs from an idle worker's. What would test it: the same
start / serve one completion / SIGTERM sequence on an engine build that defines
VLLMClientError, asserting exit status 0 and an exit time within the 1 second
bound; the repository's vLLM aggregated serving job covers that path if you run
full CI on this head. Separately, that import has no fallback, so on that engine
version the backend returns 500 for every request — worth its own issue,
outside this pull request.

Zero crashes in 300,000 children bounds the crash rate; it does not prove the
crash impossible. The measurement covers one scenario, one host, one CPython
version.

@glamr-agent glamr-agent changed the title test(runtime): stop the fetch-model bridge test flaking on the shared port pool test(runtime): reduce fetch-model bridge shared-port reservations Sep 15, 2026
@glamr-agent

Copy link
Copy Markdown
Contributor Author

nursery: @dynamo-ops please run full CI for 3620eb9

@glamr-agent

Copy link
Copy Markdown
Contributor Author

nursery: Handoff for this pass at 3620eb98107e368e0090b46e9929943c0d702665.

  • Corrected the title and description to claim reduced shared-port reservations and timeout diagnostics, removed the unsupported crash-fix claim, and corrected the related-issue and historical-validation sections. Resolved the inaccurate-claim and timeout-description threads. The child SIGSEGV remains unresolved; this PR must not close DYN-4434.
  • Left the optional fixture-placement nit open and skipped optional docstring additions. No source changes, commits, pushes, local tests, builds, or benchmarks. GitHub reports no merge conflict.
  • Latest checks: 15 successful lightweight check runs, nine skipped, successful CodeRabbit status, no failures or pending checks. Full CI has not started, so this is not a full validation pass. The once-per-head authorization request is here; it remains unanswered. Preserve this head and monitor that request rather than posting it again.
  • Requested dynamo-runtime-codeowners through the requested-reviewers API, but GitHub returned HTTP 404 (Not Found); no reviewer request landed. A maintainer with access needs to request the runtime team's review. No bot re-review was requested because there was no substantive code push.

Next: authorize and run full CI for this SHA, inspect any failures, and obtain runtime-owner review, including disposition of the remaining optional thread. This PR is not claimed ready to merge.

@dynamo-review-agent

Copy link
Copy Markdown

/ok-to-test 3620eb9

The child in tests/runtime/test_fetch_model_runtime_bridge.py printed its
success sentinel and then died with SIGSEGV, roughly once in every few
thousand runs. Every assertion had already passed, so the crash was in
process exit, not in anything the test asserts.

A captured backtrace names the faulting thread:

  tokio::runtime::task::core::Core<T,S>::poll
   -> pyo3_async_runtimes::tokio::TokioRuntime::spawn
   -> generic::future_into_py_with_locals::{{closure}}
   -> pyo3::marker::Python::with_gil
   -> pyo3_async_runtimes::generic::set_result
   -> pyo3_async_runtimes::call_soon_threadsafe
   -> PyAnyMethods::call -> Bound<PyTuple>::drop
   -> PyObject_GC_Del / PyObject_GC_UnTrack   <- SIGSEGV

future_into_py resolves the Python future through call_soon_threadsafe, so
the awaiting coroutine resumes while the Tokio worker is still inside
with_gil with an epilogue to run - it has yet to drop the argument tuple.
The main thread can leave asyncio.run and begin finalization in that
window, and the worker then touches the GC of a half-torn-down
interpreter. The process runtime lives in a static, so nothing ever joins
those threads.

Register an atexit hook at module import that waits, bounded, for the
process runtime to have no live tasks. It releases the GIL first: an
atexit callback holds it, and the tasks being waited on need it, so
waiting while holding it would deadlock against those very threads. It
returns immediately when no runtime was ever created.

Also narrow this pull request to Dynamo behavior. The two tests asserting
what the pytest port allocator does, their helpers, the raised timeouts
and the unused partial-output helper are dropped. The conditional port
acquisition stays, moved into a fixture that reads the scenario parameter
so an exhausted pool is a setup error rather than a test failure.

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
@glamr-agent
glamr-agent deployed to external_collaborator September 15, 2026 06:43 — with GitHub Actions Active
@glamr-agent

Copy link
Copy Markdown
Contributor Author

factory: automated evidence record for the current head 19be178. It supersedes the earlier automated evidence record on this request, which covered the previous head and reported an incomplete validation.

Automated evidence record — validation complete

Validation status: complete

Evidence summary: [5/5 validated]

AI review assessment (advisory, not an approval): sound. An AI agent judged the change logically sound from the code and the reported validation results. Repository CI and human reviewers decide whether to merge.

Validation result: pass. With the hook in place, 300,000 child processes across two arms produced 0 non-zero exits, where the hook-off arm's rate predicts about 252. Removing only the atexit registration from the same binary, on the same host, at the same time brought the crash back at 126 SIGSEGV in 150,000, and all 126 had already printed the child's success sentinel. Build, cargo fmt, cargo check and cargo clippy -D warnings over both the workspace and the bindings manifest, cargo test -p dynamo-runtime --lib, pre-commit run --files on the three changed files, and the six-case pytest file all exited 0. A real worker and frontend start, list their model on /v1/models with HTTP 200, and exit cleanly on SIGTERM in 1.12s and 7.69s.

Caveats a reviewer should read rather than skip:

  • The hook waits its full 1 second bound on any process that still has live service tasks. An earlier version of the change description understated this as "no measurable delay"; the current text states it plainly. A short-lived process that drains pays nothing: a four-probe timing run measured 0.106s–0.110s with the hook against 0.112s–0.125s with it cleared, and never reached the bound.
  • Token generation could not be exercised. This checkout imports VLLMClientError from an installed vLLM that does not define it, a pre-existing fault in code this change does not touch.

Change since the previous round: only the write-up. HEAD is unchanged apart from the working-notes commit, git status --short is empty, and the diff against 3620eb9 is byte-identical to the one already recorded. The cheap checks were re-run green against that unchanged build (maturin develop --uv plus an import that sees atexit._ncallbacks() go 0 to 1; pre-commit run --files on the three changed files; git diff --check, cargo fmt --all --check, and cargo check / cargo clippy -D warnings over both the workspace and the bindings manifest; 6 passed in 0.82s for the test file). A third serving smoke put the frontend's SIGTERM exit at 1.12s for the third time, with the hook's alive_tasks=5 deadline message again in the log and GPU memory back to 0 MiB.

Evidence audit: complete [5/5 validated] — the command report below comes from recorded runs.

Commands and results [5/5 validated]

Generated from the commands recorded during this run.

Check 1

Builds the changed Dynamo source and confirms that Python can import its compiled extension.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

Check 2

Checks the changed files with the repository's fast lint and formatting commands.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

Check 3

Checks the changed Rust crates with cargo check and Clippy.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

Check 4

Runs the relevant Python unit tests without requiring a GPU.

Result: Passed (exit 0)

Command:

/opt/dynamo/venv/bin/python exit_cost.py

Check 5

Starts Dynamo with vLLM on one GPU and sends a real request.

Result: Passed (exit 0)

Command:

Not shown because the exact command contained private run data.

@glamr-agent glamr-agent left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated AI review — advisory. The inline comments below are the findings. The full review is in a separate comment on this pull request.

Comment thread lib/bindings/python/rust/lib.rs Outdated
Comment thread lib/bindings/python/rust/lib.rs Outdated
@glamr-agent

Copy link
Copy Markdown
Contributor Author

factory: CI reached a terminal state on 19be1785b, the current head: 21 checks passed and 7 were skipped by path filters, with none failing and none pending. No check needs action. GitHub reports the merge as blocked, which reflects the outstanding maintainer review approval rather than any check result.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

factory: Answering the question about the failing pre-merge checks, and pushing the test removal that was asked for earlier in the same thread.

Why the checks were red on 9153d948. Nothing in the branch failed. rust-tests (lib/runtime/examples) never started. Its only annotation is:

The job was not started because it repeatedly failed to be acquired (5 attempts).

That job has no steps and no log. It failing cancelled the other three legs of the matrix: rust-tests (.), rust-tests (lib/bindings/python) and rust-tests (lib/bindings/kvbm) all ended with conclusion cancelled, each annotated:

The strategy configuration was canceled because "rust-tests.lib_runtime_examples" failed

pre-merge-status-check then failed at its Check all dependent jobs step, because one of the jobs it gates did not succeed. Everything that did run on that commit passed, including all four rust-clippy legs, which compile the same crates the cancelled test legs would have compiled. So the branch builds, and this is a runner-acquisition problem rather than a defect. The other side of that: the cancelled legs did not clear this code either, so the suite still needs a real run before anyone calls it green.

Test change dropped. Commit d4250b70c. tests/runtime/test_fetch_model_runtime_bridge.py is back to its main content, and git diff origin/main -- tests/runtime/test_fetch_model_runtime_bridge.py is now empty. The only change the file carried was a system_port fixture that narrowed how an existing test reserves a port from the shared pool, which is CI reliability rather than Dynamo functionality. The same commit deletes the comment above the bridge-runtime identity check in lib/bindings/python/rust/lib.rs, as review asked. The change is now confined to lib/bindings/python/rust/* and lib/runtime/src/worker.rs.

Checks run locally on the new head:

pre-commit run --files lib/bindings/python/rust/lib.rs tests/runtime/test_fetch_model_runtime_bridge.py
    passed
cargo fmt --manifest-path lib/bindings/python/Cargo.toml --check
    no changes
cargo check --manifest-path lib/bindings/python/Cargo.toml --all-targets
    exit 0 in 2m19s
pytest tests/runtime/test_fetch_model_runtime_bridge.py --collect-only -q
    6 tests collected, the same six ordering scenarios as main

Pushing this head starts Pre Merge again, which gives the cancelled rust-tests legs their first real attempt at this code. I will report the outcome here.

One caveat on the description: its Summary still has a bullet for the port reservation, and the reviewer-start line still names the test file. Both are stale now. The title and description were edited outside this run, so I have left them as they are rather than overwrite that edit.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

/devin review

@coderabbitai full review

@coderabbitai

coderabbitai Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

@glamr-agent I will perform a full review of #14813.

✅ Action performed

Full review finished.

The nested `if` only stayed under `clippy::collapsible_if` because a comment
sat between the two conditions. With that comment gone, clippy rejects the
shape, so fold the identity check into the `if let` as a let-chain.

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
@glamr-agent
glamr-agent deployed to external_collaborator September 15, 2026 17:39 — with GitHub Actions Active
@glamr-agent

Copy link
Copy Markdown
Contributor Author

factory: The last push broke rust-clippy (lib/bindings/python), and it was my doing. Fixed in 7709b1166, one file, lib/bindings/python/rust/lib.rs.

Deleting the comment above the bridge-runtime identity check, as review asked, left this:

if let Some(bridge) = BRIDGE_RUNTIME.get().copied() {
    if !runtimes.iter().any(|rt| std::ptr::eq(*rt, bridge)) {
        runtimes.push(bridge);
    }
}

clippy::collapsible_if does not fire when a comment sits between the two conditions, so the comment had been holding the lint off. With it gone the lint fires:

error: this `if` statement can be collapsed
   --> rust/lib.rs:264:5
    = note: `-D clippy::collapsible-if` implied by `-D warnings`

The fix folds the identity check into the if let as a let-chain, which is what the lint asks for and reads better than either earlier shape:

if let Some(bridge) = BRIDGE_RUNTIME.get().copied()
    && !runtimes.iter().any(|rt| std::ptr::eq(*rt, bridge))
{
    runtimes.push(bridge);
}

For the record on the other two red clippy legs: rust-clippy (.) and rust-clippy (lib/bindings/kvbm) were not independently broken. Both were annotated The strategy configuration was canceled because "rust-clippy.lib_bindings_python" failed, and the workspace leg's own Run Clippy Checks step had already succeeded before it was cancelled.

Checks run locally on 7709b1166:

cargo clippy --manifest-path lib/bindings/python/Cargo.toml --all-targets -- -D warnings
    clean
cargo clippy --workspace --all-targets -- -D warnings
    clean
cargo fmt --manifest-path lib/bindings/python/Cargo.toml --check
    no changes
pre-commit run --files lib/bindings/python/rust/lib.rs
    passed

One good result from the previous head: rust-tests (lib/runtime/examples) ran and passed in 3m25s on d4250b70c. That is the job whose runner could not be acquired before, so that failure was transient and is now cleared.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

factory: @dynamo-ops please run full CI for 7709b11

@dynamo-review-agent

Copy link
Copy Markdown

/ok-to-test 7709b11

@nv-tusharma

Copy link
Copy Markdown
Collaborator

/ok to test 7709b11

@nv-tusharma
nv-tusharma deployed to external_collaborator September 15, 2026 19:01 — with GitHub Actions Active
@glamr-agent

Copy link
Copy Markdown
Contributor Author

nursery: @dynamo-ops please run full CI for 405cba0

@dynamo-review-agent dynamo-review-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Previously reported defects still present:

  • Original discussion: The current checkout has no changed bridge-runtime test file, so the combined context-bridge, detached-runtime, and fetch ordering scenario remains untested as described in the open discussion.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

nursery: Handoff for this pass on 405cba027a371cd00f22cf11b240e4d88b14a0b6: corrected the description to remove stale claims about the removed port fixture. No code changes, commits, or pushes were needed; GitHub reports no merge conflicts. Earlier runtime-recording repairs are present.

Current-head CI has 18 successful checks, 7 skipped checks, and 2 running Rust test jobs, with no failures so far. CodeRabbit's commit status is successful. Full CI has not started for this head; the authorization request is posted above and remains unanswered. Pending jobs: workspace Rust tests, Python-binding Rust tests.

The runtime-codeowners review request through the requested_reviewers API returned HTTP 404 and did not land. A maintainer must route that review. The two context-first/detached coverage threads remain open for the existing DYN-4434 follow-up disposition; this pass added no tests or duplicate bot-review requests.

Next: authorize full CI for this exact head, inspect its results and the running jobs, route runtime-owner review, and decide the two deferred coverage threads. No local tests, builds, or benchmarks were run. This PR is not ready to merge.

@nv-tusharma

Copy link
Copy Markdown
Collaborator

/ok to test 405cba0

@glamr-agent

Copy link
Copy Markdown
Contributor Author

nursery: Handoff for this pass at 405cba027a371cd00f22cf11b240e4d88b14a0b6.

Read the complete current diff, refreshed paginated reviews/conversation and CI evidence, and verified the earlier runtime-recording repairs. No new actionable code defect or merge conflict was found; no edits, commits, pushes, or thread resolutions were made. No local tests, builds, benchmarks, or new tests.

Current-head results remain incomplete: 24 successful check runs, 33 skipped, seven running, none failed; CodeRabbit commit status is successful. All four Rust test jobs, four Clippy jobs, remote pre-commit, and the pre-merge gate passed. Full CI is already authorized and running on this exact SHA; no duplicate authorization or bot-review request was posted. Its changed-files log confirms core, frontend, and rust filters are active. Pending image builds:

Requested dynamo-runtime-codeowners through the upstream requested_reviewers API; it returned HTTP 404 (Not Found). Refreshed metadata confirms no reviewer request landed. This review-routing failure requires a maintainer; approval remains required.

The two coverage threads remain open for maintainer disposition and the existing DYN-4434 follow-up: context-first detached exit and combined context/detached/fetch order. The coverage is absent; this pass prohibits adding tests and does not claim those findings repaired. Existing replies already explain the deferral, so no duplicate thread replies were added.

Next: preserve this head while CI finishes, inspect any failures, have a maintainer route runtime-owner review, and decide the two deferred coverage threads. Skipped and running checks are not passes; this PR is not ready to merge.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

nursery: CI investigation for 405cba027a371cd00f22cf11b240e4d88b14a0b6: the Dynamo image job exceeded its one-hour execution limit. The main image build succeeded; the subsequent test-image build was cancelled during COPY steps. This caused dynamo-status-check to fail and downstream Dynamo tests, mypy, and rust-gpu jobs to be skipped. The logs show no compiler or test assertion failure. Available main-workflow comparisons are too old to establish whether this is also occurring on current main.

I attempted to rerun job 104561929996; GitHub denied it with HTTP 403, Must have admin rights to Repository. An authorized maintainer needs to rerun the timed-out job and its dependents. The runtime-owner review request also returned HTTP 404; refreshed metadata confirms it did not land, so a maintainer must route that review.

At this pass's final check, 54 check runs passed, 54 were skipped, one was cancelled, one failed, and two remain running: TRT-LLM tests and TRT-LLM two-GPU tests. All Rust tests, Clippy jobs, remote pre-commit, and the pre-merge gate passed. Full CI is already authorized for this SHA; no duplicate request was posted.

No new actionable code defect or merge conflict was found. No edits, commits, pushes, local tests, builds, benchmarks, or new tests in this pass. The two deferred context-first/detached coverage threads still need maintainer disposition under DYN-4434; neither was marked repaired. Preserve this head while CI and review are handled. This PR is not ready to merge.

@jthomson04 jthomson04 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved. The four P3 comments are non-blocking suggestions. Source review only; no tests were run and CI was not checked as part of this review.

Comment thread lib/bindings/python/rust/lib.rs Outdated
Comment thread lib/bindings/python/rust/lib.rs Outdated
Comment thread lib/bindings/python/rust/lib.rs
Comment thread lib/bindings/python/rust/lib.rs Outdated
…rage

Look the bridge runtime up once through the OnceLock instead of asking PyO3
on every async conversion, hold the at-most-two runtimes the exit hook drains
in a fixed array, call the plain std::thread::sleep import, and trim the
surrounding comments to the constraints they carry.

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
@glamr-agent
glamr-agent deployed to external_collaborator September 16, 2026 19:01 — with GitHub Actions Active
@glamr-agent

Copy link
Copy Markdown
Contributor Author

factory: Took all four P3 comments on lib/bindings/python/rust/lib.rs. One commit, 1daee8fdb, one file.

  • Cache the runtime lookup. bridge_runtime() is now BRIDGE_RUNTIME.get_or_init(pyo3_async_runtimes::tokio::get_runtime), so the lookup and the failing set happen once instead of on every conversion.
  • Stack storage. wait_for_bridge_tasks_at_exit holds [Option<&'static Runtime>; 2] and iterates it with .flatten(); the Vec and its allocation are gone.
  • Shorter comments. Kept the GIL requirement above py.allow_threads, which runtimes get drained and why they can differ, and what the timeout costs a process whose tasks never finish. Dropped the call_soon_threadsafe epilogue narrative and the sub-millisecond margin.
  • Sleep convention. use std::thread::sleep; at the top, sleep(BRIDGE_DRAIN_POLL) at the call site.

Checks, all from lib/bindings/python:

cargo check --lib                                                   ok
cargo clippy --all-targets --all-features -- -D warnings            ok
cargo fmt --manifest-path lib/bindings/python/Cargo.toml -- --check  ok
pre-commit run --files lib/bindings/python/rust/lib.rs              ok

Clippy earned its keep here: the first version of the cached lookup wrote *BRIDGE_RUNTIME.get_or_init(...) and explicit_auto_deref rejected the deref, so the pushed version has no *.

The title and description are left as they are: they were edited by hand after the factory last wrote them, so this run does not own them any more.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

/devin review

@coderabbitai full review

@coderabbitai

coderabbitai Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor
🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '== Pull request metadata =='
gh pr view 14813 --repo ai-dynamo/dynamo --json number,title,headRefOid,baseRefName,headRefName,mergeable,state,isDraft,reviewDecision,statusCheckRollup --jq '{number,title,headRefOid,baseRefName,headRefName,mergeable,state,isDraft,reviewDecision,checks: [.statusCheckRollup[]? | {name,conclusion,status}]}'

echo '== Diff summary and affected files =='
base="$(git merge-base "origin/main" HEAD)"
printf 'base=%s\nhead=%s\n' "$base" "$(git rev-parse HEAD)"
git diff --check "$base"...HEAD
git diff --stat "$base"...HEAD
git diff --name-only "$base"...HEAD

echo '== Rust implementation map =='
ast-grep outline lib/bindings/python/rust/lib.rs --items all --type function,struct,impl,static 2>/dev/null || true
ast-grep outline lib/bindings/python/rust/backend.rs --items all --type function,struct,impl 2>/dev/null || true
ast-grep outline lib/runtime/src/worker.rs --items all --type function,struct,impl 2>/dev/null || true

echo '== Current implementation sections =='
rg -n -C 8 -P '\b(BRIDGE_RUNTIME|bridge_runtime|record_bridge_runtime|adopt_bridge_runtime|wait_for_bridge_tasks_at_exit|register_at_exit|future_into_py|future_into_py_with_locals|ensure_fetch_model_runtime|DistributedRuntime::(new|detached)|existing_process_runtime)\b' \
  lib/bindings/python/rust/lib.rs lib/bindings/python/rust/backend.rs lib/runtime/src/worker.rs

echo '== Direct PyO3 Tokio bridge API references in bindings =='
rg -n -C 3 -P '\bpyo3_async_runtimes::tokio::(future_into_py|future_into_py_with_locals|get_runtime|init_with_runtime)\b' lib/bindings/python/rust

echo '== Full diff for implementation files =='
git diff --find-renames --unified=50 "$base"...HEAD -- \
  lib/bindings/python/rust/lib.rs \
  lib/bindings/python/rust/backend.rs \
  lib/bindings/python/rust/context.rs \
  lib/bindings/python/rust/http.rs \
  lib/bindings/python/rust/kserve_grpc.rs \
  lib/bindings/python/rust/llm \
  lib/bindings/python/rust/planner.rs \
  lib/runtime/src/worker.rs

Length of output: 226710


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '== Complete exit hook and registration =='
sed -n '170,330p' lib/bindings/python/rust/lib.rs

echo '== All PyO3 async runtime references and imports =='
rg -n -C 4 -P '\bpyo3_async_runtimes\b|use\s+.*\btokio\b' lib/bindings/python/rust

echo '== Runtime spawning and block_on calls in Python bindings =='
rg -n -C 4 -P '\.(spawn|block_on)\s*\(' lib/bindings/python/rust

echo '== All references to local wrappers =='
rg -n -C 2 -P '\b(crate::)?(future_into_py(?:_with_locals)?|bridge_runtime|adopt_bridge_runtime)\s*\(' lib/bindings/python/rust

echo '== Open review threads and reviews =='
gh api graphql -F owner=ai-dynamo -F name=dynamo -F number=14813 -f query='
query($owner: String!, $name: String!, $number: Int!) {
  repository(owner: $owner, name: $name) {
    pullRequest(number: $number) {
      reviews(last: 30) { nodes { author { login } state body url commit { oid } } }
      reviewThreads(first: 100) {
        nodes {
          isResolved
          isOutdated
          path
          line
          originalLine
          comments(first: 20) { nodes { author { login } body url } }
        }
      }
    }
  }
}' --jq '
.data.repository.pullRequest as $pr |
{reviews: $pr.reviews.nodes,
 openThreads: [
   $pr.reviewThreads.nodes[]
   | select(.isResolved | not)
   | {isOutdated,path,line,originalLine,comments: [.comments.nodes[] | {author: .author.login, body, url}]}
 ]}'

echo '== Current CI result summary =='
gh pr view 14813 --repo ai-dynamo/dynamo --json headRefOid,statusCheckRollup --jq '{head: .headRefOid, checks: [.statusCheckRollup[]? | {name,status,conclusion}]}'

Length of output: 121682


@glamr-agent I completed the full review of the current head, 1daee8f.

I found no new actionable defects.

The bridge-runtime helpers now cover direct future conversion and direct bridge runtime access. The exit hook records and drains the process runtime and the distinct PyO3 runtime. It releases the GIL during the bounded wait and does not initialize a runtime at interpreter exit.

I used the recorded maintainer guidance for this PR. I did not repeat the deferred context-first DistributedRuntime::detached() test request. That coverage belongs to DYN-4434.

git diff --check passed. The Rust test and Clippy jobs are still running, so CI is not complete.


🧠 Learnings used
Learnt from: glamr-agent
Repo: ai-dynamo/dynamo PR: 14813
File: tests/runtime/test_fetch_model_runtime_bridge.py:274-287
Timestamp: 2026-09-15T17:27:47.225Z
Learning: For Dynamo pull request `#14813`, the maintainer requested no tests that cover CI reliability. Coverage for the combined `trigger_context_bridge()` then `DistributedRuntime.detached()` runtime-adoption order belongs in follow-up DYN-4434, not in this pull request.

Learnt from: glamr-agent
Repo: ai-dynamo/dynamo PR: 14813
File: tests/runtime/test_fetch_model_runtime_bridge.py:277-280
Timestamp: 2026-09-15T17:27:43.678Z
Learning: For ai-dynamo/dynamo PR `#14813`, the maintainer requested that the PR not add tests for CI reliability. Commit `d4250b70c` restored `tests/runtime/test_fetch_model_runtime_bridge.py` to its `main` content. The missing context-first-then-`DistributedRuntime.detached()` interpreter-exit regression scenario is a valid coverage gap, but it must be handled in a follow-up associated with `DYN-4434`, not in this PR.

Learnt from: glamr-agent
Repo: ai-dynamo/dynamo PR: 14813
File: lib/bindings/python/rust/lib.rs:254-267
Timestamp: 2026-09-15T08:06:39.608Z
Learning: In `lib/bindings/python/rust/lib.rs`, `wait_for_bridge_tasks_at_exit` must use a bounded wait. Frontend service tasks can remain alive indefinitely at interpreter exit, so waiting for `num_alive_tasks()` to reach zero can hang shutdown. Tasks created by `pyo3_async_runtimes::tokio::future_into_py` have no cancellation or join handles available to the Dynamo Python bindings, and the static Tokio runtimes cannot use consuming shutdown APIs. At the timeout, continuing interpreter finalization restores pre-hook behavior; the bounded drain only reduces the race window.
✅ Action performed

Full review finished.

@glamr-agent

Copy link
Copy Markdown
Contributor Author

factory: @dynamo-ops please run full CI for 1daee8f

@nv-tusharma

Copy link
Copy Markdown
Collaborator

/ok to test 1daee8f

@nv-tusharma
nv-tusharma enabled auto-merge (squash) September 16, 2026 19:25
@glamr-agent

Copy link
Copy Markdown
Contributor Author

nursery: Handoff at 1daee8fdbf44395d91613c209e84dff4df2ea23e: refreshed the paginated review and CI evidence, read the complete diff, and corrected the description's stale head and CI-authorization status. No new actionable code defect or merge conflict was found. No source edits, commits, pushes, local tests, builds, benchmarks, or new tests.

Current-head CI: 66 successful check runs, 50 skipped, one cancelled, none pending; CodeRabbit status is successful. All four Rust test jobs, all four Clippy jobs, remote pre-commit, and the pre-merge gate passed. Full CI ran but is incomplete: planner / Compliance cpu, arm64 exceeded its 15-minute limit during target-image package extraction. BuildKit bootstrap succeeded, but context loading stalled and extraction did not complete; scanning and attribution generation never ran. No source failure was shown. The returned main-branch comparison runs were from March and lacked this job, so they do not establish whether the timeout is recurring.

Attempted a targeted retry through POST /repos/ai-dynamo/dynamo/actions/jobs/104946087778/rerun; GitHub rejected it with HTTP 403: Must have admin rights to Repository. A maintainer must retry that job and its dependents, then inspect the results on this same head. No duplicate full-CI authorization request was posted.

Requested dynamo-runtime-codeowners through the PR's requested_reviewers API; GitHub returned HTTP 404, so the request did not land. Existing approver jthomson04 was not re-requested. No duplicate bot-review request was posted. The two deferred coverage threads (4013792778 and 4013977254) remain open for maintainer disposition under DYN-4434; no threads were resolved in this pass.

Next: a maintainer retries the timed-out CI job, routes any outstanding runtime-owner review, and decides the deferred coverage threads. Merge readiness is not established.

@nv-tusharma
nv-tusharma merged commit 0cb120e into ai-dynamo:main Sep 16, 2026
117 of 118 checks passed
aung-san-i added a commit to aung-san-i/dynamo that referenced this pull request Sep 28, 2026
* feat: KV DC Relay file based source mode (ai-dynamo#14807)

Add live-reloaded file sources for KV DC Relay namespace selection and expose readiness and source revisions through /engine/state.

Preserve applied membership on invalid updates, coalesce discovery refreshes, and isolate native integration tests in forked processes.

Signed-off-by: Nikita Sukharev <kaonael@gmail.com>

* feat(sglang): expose cross-encoder reranking through /v1/rerank (ai-dynamo#14032)

Signed-off-by: xianlubird <xianlubird@gmail.com>

* fix(profiler): explain inaccessible model paths during trust checks (ai-dynamo#14860)

Signed-off-by: hongkuanz <hongkuanz@nvidia.com>

* fix(sglang): sync discovery from native pause state (ai-dynamo#13951)

Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
Co-authored-by: Zero Rains <57100978+zeroRains@users.noreply.github.com>

* feat(recipes): add Solar Open2 250B NVFP4 aggregated and disaggregated recipes for B200 (ai-dynamo#14376)

Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com>

* refactor(agents): session_id reader from AgentContext + forward to vLLM (ai-dynamo#14428)

Signed-off-by: Karen Chung <karenc@nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* fix(discovery): allow served aliases for the same model source (ai-dynamo#14857)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix(router): reject unknown explicit worker targets (ai-dynamo#14858)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix(xpu): stabilize XPU test workers (ai-dynamo#14539)

Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>
Signed-off-by: VincyZhang <wenxin.zhang@intel.com>

* feat(mm-routing): add Nemotron 3 Nano Omni video routing (ai-dynamo#14653)

Signed-off-by: krishung5 <krish@nvidia.com>

* fix(sglang): validate diffusion input_reference and bound media fetches (ai-dynamo#14435)

The sglang image-diffusion and video-generation handlers passed the
client-supplied input_reference through to the generator's image_path after only
a non-empty check. Validate it first, and for remote references materialize it
locally before the generator sees it, so the generator is always handed a
trusted local path. This brings the sglang diffusion path in line with the
vLLM/omni and trtllm backends, which already validate the same field.

Behavior change: local I2I/I2V references now require DYN_MM_LOCAL_PATH to be
set to the allowed directory; previously any path was accepted.

common/http:

- validate_media_reference() returns a plain filesystem path for local
  references; local_media_reference() is an async context manager that fetches a
  remote one through fetch_bytes(policy=...), which revalidates every redirect
  hop, into a temp file removed on exit. data: is rejected -- a URI is not a path.
- fetch_bytes() gained max_bytes, streaming through collect_capped at an explicit
  read granularity so the cap is an allocation bound and not only a rejection: a
  128 MiB-decoded gzip body against the 64 MiB cap peaks at 68,032,217 bytes
  rather than the whole decompressed body. Content-Length is caller-controlled
  and absent when chunked, and aiohttp's read(n) returns at most n bytes, so
  neither a header check nor a single capped read suffices. Defaults to None,
  leaving existing callers unchanged.
- DYN_MM_MAX_FILE_SIZE_MB makes that cap operator-tunable, in megabytes, as the
  SGLang arg it replaces was. Read per call; empty, unparseable or non-positive
  falls back to 64 with a warning, so a malformed value neither takes the worker
  down nor reads as unlimited.
- Messages built from caller input are bounded via describe_media_source, moved
  from multimodal/media_source.py (it pulls in torch) into url_validator.py and
  re-exported from its old home; a no-op below 120 characters.
- HttpStatusError bounds its .message attribute, not only the rendered string:
  errors.rs::extract_http_like_error reads .status and .message off this class by
  name and forwards .message on a 4xx without calling str(). Backend exception
  text is bounded head-and-tail, since aiohttp renders the host before the errno.
- validate_local_path uses exc.strerror rather than the raw OSError, whose text
  repeats the filename, and now catches the ValueError that Path.resolve() raises
  on an embedded NUL so callers keep their 4xx-vs-5xx decision.

Rebased onto ai-dynamo#14563 (single aiohttp backend); the httpx-side half of the
max_bytes plumbing went with that backend.

Signed-off-by: nnshah1 <neelays@nvidia.com>
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(deps): upgrade fastokens to 0.3.2 (ai-dynamo#14798)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix(vllm): ship codec-free OpenCV for image inputs (ai-dynamo#14361)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: Anant Sharma <anants@nvidia.com>
Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com>

* docs: refresh community events

Automated refresh from the public Dynamo Google Calendar.

Generated by .github/workflows/community-events-refresh.yml.

Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com>

* ci: refresh the compliance baseline in auto-upgrade pipeline (ai-dynamo#14206)

Signed-off-by: Anant Sharma <anants@nvidia.com>

* feat(triton): honor KServe classification on tensor outputs (ai-dynamo#14783)

Signed-off-by: Yingge He <yinggeh@nvidia.com>

* docs(rl): stop the verl guide sending readers to a vLLM version it cannot run on (ai-dynamo#14571)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>

* feat(mocker): publish native KV events from the vLLM gRPC server (ai-dynamo#14737)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix(kv-router): release unowned radix branches after eviction (ai-dynamo#14878)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* fix: show correct backend versions in the install selectors (ai-dynamo#13599)

Signed-off-by: Anant Sharma <anants@nvidia.com>

* build(vllm): prepare v0.29.0 bump (ai-dynamo#14543)

Signed-off-by: Julien Darve <jdarve@NVIDIA.com>

* ci(xpu): validation PR for the re-applied XPU workflows and Dockerfile

Throwaway PR to prove the CI merged in #22 actually runs end to end on XPU
hardware. Adds only a comment to container/templates/vllm_runtime.Dockerfile,
which matches the `vllm` path filter (container/templates/vllm_*) and so makes
changed-files set vllm=true, which is what gates build-xpu and the
heterog-test-px-dn / heterog-test-pn-dx jobs.

What this exercises:
  - .github/workflows/pr-xpu.yaml            (push to pull-request/[0-9]+, needs the xpu label)
  - .github/workflows/pr-xpu-heterogeneous.yaml (push; its guard deliberately skips the label gate)
  - .github/workflows/epd-test-template.yml  (workflow_call, from the heterog jobs)
  - .github/scripts/test-filters.js          (the brace fix from #22)
  - container/templates/vllm_runtime.Dockerfile rendered and built for device=xpu

Not exercised: .github/workflows/xpu-heterogeneous-dispatch.yaml is
workflow_dispatch only and has to be run by hand from the Actions tab.

The marker comment must be removed before this branch is ever merged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(triton): Update Triton Base Image to 26.08 (ai-dynamo#14854)

Signed-off-by: J Wyman <jwyman@nvidia.com>
Co-authored-by: Rini Gupta <rinig@nvidia.com>

* fix(operator): normalize equivalent worker hash inputs (ai-dynamo#14721)

Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>

* test(sglang): exercise NIXL in embedding cache E/PD test (ai-dynamo#14795)

Signed-off-by: Sai Kiran Polisetty <spolisetty@nvidia.com>

* fix(sglang): stop the elastic-EP scale-up worker crash-looping at startup (ai-dynamo#14568)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com>

* fix(responses): honor tool_choice when parsing tool calls from text (ai-dynamo#14843)

Signed-off-by: xianlubird <xianlubird@gmail.com>

* ci: accept trusted full-CI request comments (ai-dynamo#14868)

Signed-off-by: Matej Kosec <mkosec@nvidia.com>

* docs: clarify EPP mode boundary and single-replica Dynamo mode fixes [DYN-4310] (ai-dynamo#14756)

Signed-off-by: Anna Tchernych <atchernych@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci(docs): move the generated-tables determinism gate out of link checking (ai-dynamo#14135)

Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* ci(docs): generate the Kubernetes API reference at publish time (ai-dynamo#14122)

Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(operator): discover pull secrets for init containers (ai-dynamo#14922)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* fix(sglang): stop an unusable mooncake backend crashing workers after model load (ai-dynamo#14461)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: glamr-agent <glamr-agent@users.noreply.github.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>

* fix(sglang): emit prefill handoff before completion in sidecar (ai-dynamo#14260)

Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>

* test(trtllm): enable fault tolerance coverage (ai-dynamo#14609)

Signed-off-by: tanmayv25 <tanmay2592@gmail.com>

* fix(frontend): evict async tokenizer executors when the tokenizer is retired (ai-dynamo#13368)

Signed-off-by: Peter Pan <Peter.Pan@daocloud.io>

* fix(llm): report KServe datatypes by their wire names, not protobuf variants (ai-dynamo#14957)

`ModelMetadata` reported each Triton-registered tensor's `datatype` using
`inference::DataType::as_str_name()`, which returns the `model_config.proto`
variant name (`TYPE_FP32`, `TYPE_STRING`, ...) instead of the KServe v2 wire
names (`FP32`, `BYTES`, ...). Every datatype was wrong, so spec-conforming
clients cannot parse any tensor the RPC describes. Adds `oip_name()` next to
`tensor::DataType::to_kserve` covering all fifteen proto variants (incl. FP16
and BF16) and mapping `TYPE_STRING → BYTES`.

Original PR by @ayaangazali: ai-dynamo#14770. Reissued under a signed commit to
unblock the copy-pr-bot signature gate; diff is byte-identical.

Closes ai-dynamo#14520.

Signed-off-by: ayaangazali <ayaangazali@users.noreply.github.com>
Signed-off-by: ayaangazali <ayaangazali.work@gmail.com>
Signed-off-by: Vinya Kestur <vinyak@nvidia.com>
Co-authored-by: ayaangazali <ayaangazali.work@gmail.com>

* docs(mm-routing): document video KV routing (ai-dynamo#14958)

Signed-off-by: krishung5 <krish@nvidia.com>

* fix(sidecar): honor worker namespace suffix (ai-dynamo#14955)

Signed-off-by: Biswa Panda <biswa.panda@gmail.com>

* fix(bindings): drain bridge tasks before interpreter finalization (ai-dynamo#14813)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: Tushar Sharma <tusharma@nvidia.com>

* fix(discovery): stop a Qwen3-VL worker from serving video with another worker's contract (ai-dynamo#14624)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>

* fix(gms): honor configured timeout during initial weights admission (ai-dynamo#14877)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>

* feat(kv-router): add construction-time indexer delegates (ai-dynamo#14945)

* fix(sglang): support min_tokens on tokenizer-free decode workers (ai-dynamo#14276)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: MatejKosec <mkosec@nvidia.com>

* feat(router): add SessionPrefixIndexer for session-block lineage (ai-dynamo#13807)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: Karen Chung <karenc@nvidia.com>
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
Co-authored-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Co-authored-by: Matej Kosec <mkosec@nvidia.com>

* fix(vllm): settle kvwarm stages through a per-step round on every attention-DP rank (ai-dynamo#14728)

Signed-off-by: Yiming Liu <yimingl@nvidia.com>

* feat(vllm): benchmark hybrid caches with random KDA state (ai-dynamo#14900)

Signed-off-by: hongkuanz <hongkuanz@nvidia.com>

* fix(runtime): fix QUIC reassembly and reduce response stalls (ai-dynamo#14876)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* feat(router): unify frontend and standalone selection core (ai-dynamo#14570)

Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
Signed-off-by: Thomas Montfort <tjmontfort12@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Thomas Montfort <tjmontfort12@gmail.com>

* fix(planner): keep control APIs responsive during Prometheus collection (ai-dynamo#14377)

Signed-off-by: xianlubird <xianlubird@gmail.com>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>

* fix(router): record SGLang prefill completion after stream ends (ai-dynamo#14968)

Signed-off-by: jain-ria <riajain@NVIDIA.com>

* fix(frontend): send inline media once on the TCP request plane (ai-dynamo#14801)

Signed-off-by: Sumit Mishra <sah299610@gmail.com>
Co-authored-by: Indrajit Bhosale <iamindrajitb@gmail.com>

* docs: refresh community events

Automated refresh from the public Dynamo Google Calendar.

Generated by .github/workflows/community-events-refresh.yml.

Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com>

* fix(vllm): initialize synchronizer in KV warmup capacity test (ai-dynamo#14984)

Signed-off-by: Alec Flowers <aflowers@nvidia.com>

* fix(recipes): make the Solar Open2 250B benchmark and docs link usable (ai-dynamo#14956)

Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com>

* feat(recipes): add K-EXAONE 2.0 750B-A37B NVFP4 vLLM recipes for B200 (ai-dynamo#14822)

Signed-off-by: Cheng Wang <chengwa@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: KVCR Resiliency Deployment Example (ai-dynamo#14695)

Add two-node DynamoGraphDeployment examples for process-local KVCR and
the KVCR memory service. Run one vLLM worker per GPU node, use stable
Grove ordinals for cache-owner slots, and request GPU-local RDMA
resources for engines and Guard services. Provide a deployment helper
for rendering and selecting either variant.

Run the KV state agent alongside vLLM for process-local host memory. In
memory-service mode, keep KVCR and the state agent in a separate
container so its Guard and shared-memory pool survive engine restarts.
Document that restarting the services sidecar invalidates the MVP
recovery contract and requires deployment-level replacement.

Add manifest coverage and an opt-in two-host lifecycle test. Kill the
source EngineCore, hold it offline, and verify that the promoted Guard
serves its preserved cache to the surviving target. Correlate response
equality and KVCR transfer metrics with transmit and receive counters
from the selected active HCA to prove RDMA transport.

Pin compatible KVCR and vLLM revisions and document the runtime,
discovery, compatibility-digest, and recovery prerequisites.

Signed-off-by: Adit Ranadive <aranadive@nvidia.com>

* feat(omni): add Nemotron Audex speech synthesis to /v1/audio/speech (ai-dynamo#12788)

Signed-off-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com>

* ci: allow glamr-agent to request CI on its own unsigned PRs (ai-dynamo#14964)

Signed-off-by: Matej Kosec <mkosec@nvidia.com>

* fix(vllm): isolate multimodal worker ports (ai-dynamo#14751)

Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>

* fix(runtime): reject invalid DYN_REQUEST_PLANE values (ai-dynamo#12612)

Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
Signed-off-by: Coding Agent <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: MatejKosec <mkosec@nvidia.com>

* fix(responses): preserve text instead of inferring tool calls (ai-dynamo#14846)

Signed-off-by: xianlubird <xianlubird@gmail.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>

* chore: temporarily increase frontend build time limit 45 --> 90 min (ai-dynamo#15019)

Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>

* test(operator): cover scoped CA injection ownership (ai-dynamo#14961)

Signed-off-by: Julien Mancuso <jmancuso@nvidia.com>

* feat(frontend): map semantic errors to HTTP responses (ai-dynamo#14396)

Signed-off-by: Biswa Panda <biswa.panda@gmail.com>

* docs: correct fault-tolerance architecture details (ai-dynamo#14880)

Signed-off-by: Elizabeth Thomas <email2eliza@gmail.com>

* build(deps): bump nats-server to v2.14.7 (ai-dynamo#14919)

Signed-off-by: Dan Gil <dagil@nvidia.com>

* build(deps): bump AISimulate to 0.12.0 (ai-dynamo#15012)

Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>

* remove oneAPI env for XPU detection

* feat(backends): expose native LoRA capacity in model registration (ai-dynamo#14754)

Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>
Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com>

* fix(planner): handle pending decisions in virtual connector wait (ai-dynamo#14841)

Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>

* feat(vllm): add sidecar LoRA lifecycle (ai-dynamo#13068)

Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>
Co-authored-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com>

* fix(vllm/omni): pass response_format into video EngineInputs (ai-dynamo#14667) (ai-dynamo#14844)

* chore: bump version to 1.6.0 post 1.5.0 branch cut (ai-dynamo#15009)

Signed-off-by: pvijayakrish <pvijayakrish@nvidia.com>
Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): Use `pytest --ignore` to Skip Tests Based on Framework (ai-dynamo#14815)

Signed-off-by: J Wyman <jwyman@nvidia.com>

* feat(sidecar): add e2e CI testing for sidecar launch scripts (ai-dynamo#14508)

Signed-off-by: tanmayv25 <tanmay2592@gmail.com>
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: Julien Darve <jdarve@NVIDIA.com>

* chore(xpu): upgrade vllm and omni to 0.29.0

Signed-off-by: wenxin.zhang <wenxin.zhang@intel.com>

* docs(operator): document the DGDR workload-creation trust boundary (ai-dynamo#14429)

Signed-off-by: nnshah1 <neelays@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* fix(xpu): use released vllm-omni prerelease

Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>

* test(efa): add the EFA disaggregated deploy test for sglang (ai-dynamo#13893)

Signed-off-by: Jie Hao <jihao@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(runtime): support IPv6-only IP resolution (ai-dynamo#13126)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

* docs(fault-tolerance): clarify migration after shutdown grace expires (ai-dynamo#14872)

Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>

* feat(vllm-omni): preserve generated video audio (ai-dynamo#13707)

Signed-off-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>

* feat(vllm-omni): pass model-specific video parameters (ai-dynamo#13708)

Signed-off-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>

* feat(vllm-omni): qualify MiniMax-H3 T2VA on B200 (ai-dynamo#13589)

Signed-off-by: Guan Luo <gluo@nvidia.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>

* fix(vllm): remove obsolete Omni compatibility guard

Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>

* fix(vllm): retain Omni compatibility guard

Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>

* .github/workflows/pr-xpu-heterogeneous.yaml; pin GPU_TAG to latest

* .github/workflows/; add post-merge and nightly XPU heterogeneous CI

Extract the XPU heterogeneous P/D pipeline out of pr-xpu-heterogeneous.yaml
into xpu-heterogeneous-run.yml, a workflow_call reusable workflow, and call it
from three thin trigger workflows so all three merge phases run the identical
pipeline instead of drifting copies.

  xpu-heterogeneous-run.yml           new, reusable. guard, changed-files,
                                      build-xpu, build-nvidia, resolve-images
                                      and both heterog tests, unchanged, plus
                                      7 inputs.
  pr-xpu-heterogeneous.yaml           reduced to the pre-merge trigger, the
                                      slash-command gate and the reaction.
  post-merge-xpu-heterogeneous.yaml   new. push to main.
  nightly-xpu-heterogeneous.yaml      new file, but the cron is MOVED, not
                                      added: it is the 0 23 * * * schedule
                                      that was already in
                                      pr-xpu-heterogeneous.yaml.

No behaviour change per phase. force_all_tests replaces the old
  github.event_name == 'schedule' || github.event_name == 'issue_comment'
expression with the same truth table: pre-merge passes
github.event_name == 'issue_comment', nightly passes true. Post-merge also
passes true, because a push to main has no PR base for
.github/actions/changed-files to diff against, and post-merge exists to catch
what per-PR gating missed.

xpu-status-check stays a TOP-LEVEL job in each caller rather than moving into
the reusable workflow. A job contributed by a reusable workflow reports to the
Checks API as "run / xpu-status-check", so hosting it there would rename the
context and leave any branch protection rule requiring xpu-status-check waiting
forever on a check that no longer reports.

The concurrency mapping stays byte-identical across all four workflows that
touch this hardware, now including xpu-heterogeneous-dispatch.yaml. Three files
do NOT get three slots: the cluster, the dynamo-system namespace and the
onexpu-/onenvidia-rdma-kueue ResourceClaimTemplates are one global resource.
The reusable workflow deliberately carries no concurrency block of its own,
which would deadlock against the slot the caller's run already holds.

Parameterised gpu_tag, model, tensor_parallel and runner as inputs so the
callers can diverge; all default to the previously hardcoded values. Added
workflow_dispatch to the nightly, without which a schedule-only workflow cannot
be exercised before it reaches the default branch.

Verified: all files parse; the four concurrency mappings are byte-identical; the
reusable workflow declares no concurrency; every input each caller passes exists
and every required input is supplied; nesting is depth 3 of the 4 GitHub allows.
actionlint was not available to run, and will report queue:max as an unknown key
in all four files, a known false positive.

---------

Signed-off-by: Nikita Sukharev <kaonael@gmail.com>
Signed-off-by: xianlubird <xianlubird@gmail.com>
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com>
Signed-off-by: Karen Chung <karenc@nvidia.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com>
Signed-off-by: VincyZhang <wenxin.zhang@intel.com>
Signed-off-by: krishung5 <krish@nvidia.com>
Signed-off-by: nnshah1 <neelays@nvidia.com>
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: GLAMR <svc-glamr@nvidia.com>
Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com>
Signed-off-by: Anant Sharma <anants@nvidia.com>
Signed-off-by: Yingge He <yinggeh@nvidia.com>
Signed-off-by: Julien Darve <jdarve@NVIDIA.com>
Signed-off-by: J Wyman <jwyman@nvidia.com>
Signed-off-by: bzsuni <bingzhe.sun@daocloud.io>
Signed-off-by: Sai Kiran Polisetty <spolisetty@nvidia.com>
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
Signed-off-by: Anna Tchernych <atchernych@nvidia.com>
Signed-off-by: Dan Gil <dagil@nvidia.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: glamr-agent <glamr-agent@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Signed-off-by: tanmayv25 <tanmay2592@gmail.com>
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io>
Signed-off-by: ayaangazali <ayaangazali@users.noreply.github.com>
Signed-off-by: ayaangazali <ayaangazali.work@gmail.com>
Signed-off-by: Vinya Kestur <vinyak@nvidia.com>
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
Signed-off-by: Yiming Liu <yimingl@nvidia.com>
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
Signed-off-by: Thomas Montfort <tjmontfort12@gmail.com>
Signed-off-by: Sumit Mishra <sah299610@gmail.com>
Signed-off-by: Alec Flowers <aflowers@nvidia.com>
Signed-off-by: Cheng Wang <chengwa@nvidia.com>
Signed-off-by: Adit Ranadive <aranadive@nvidia.com>
Signed-off-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com>
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Signed-off-by: Coding Agent <svc-glamr@nvidia.com>
Signed-off-by: Julien Mancuso <jmancuso@nvidia.com>
Signed-off-by: Elizabeth Thomas <email2eliza@gmail.com>
Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>
Signed-off-by: pvijayakrish <pvijayakrish@nvidia.com>
Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
Signed-off-by: wenxin.zhang <wenxin.zhang@intel.com>
Signed-off-by: Jie Hao <jihao@nvidia.com>
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
Signed-off-by: Guan Luo <gluo@nvidia.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Nikita Sukharev <kaonael@gmail.com>
Co-authored-by: Xianlu Bird <xianlubird@gmail.com>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
Co-authored-by: Zero Rains <57100978+zeroRains@users.noreply.github.com>
Co-authored-by: snarravula-dl <snarravula@nvidia.com>
Co-authored-by: Karen Chung <karenc@nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: VincyZhang <wenxin.zhang@intel.com>
Co-authored-by: Kris Hung <krish@nvidia.com>
Co-authored-by: Neelay Shah <neelays@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: GLAMR <svc-glamr@nvidia.com>
Co-authored-by: Anant Sharma <anants@nvidia.com>
Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com>
Co-authored-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com>
Co-authored-by: Yingge He <157551214+yinggeh@users.noreply.github.com>
Co-authored-by: JulienDarve <86800349+JulienDarve@users.noreply.github.com>
Co-authored-by: J Wyman <jwyman@nvidia.com>
Co-authored-by: Rini Gupta <rinig@nvidia.com>
Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com>
Co-authored-by: Sai Kiran Polisetty <spolisetty@nvidia.com>
Co-authored-by: MatejKosec <mkosec@nvidia.com>
Co-authored-by: atchernych <atchernych@nvidia.com>
Co-authored-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Bojiang Li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Connor Carpenter <connorcarpenter15@gmail.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Tanmay Verma <tanmayv@nvidia.com>
Co-authored-by: Peter Pan <peter.pan@daocloud.io>
Co-authored-by: Vinya Kestur Tumakuru Arun Kumar <vinyak@nvidia.com>
Co-authored-by: ayaangazali <ayaangazali.work@gmail.com>
Co-authored-by: Biswa Panda <biswa.panda@gmail.com>
Co-authored-by: Tushar Sharma <tusharma@nvidia.com>
Co-authored-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
Co-authored-by: Ryan Olson <ryanolson@users.noreply.github.com>
Co-authored-by: Yimingl_Nvidia <yimingl@nvidia.com>
Co-authored-by: Thomas Montfort <tjmontfort12@gmail.com>
Co-authored-by: Sumit884-byte <sah299610@gmail.com>
Co-authored-by: Indrajit Bhosale <iamindrajitb@gmail.com>
Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Co-authored-by: chw001 <chengwa@nvidia.com>
Co-authored-by: Adit Ranadive <aranadive@nvidia.com>
Co-authored-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com>
Co-authored-by: Keiven C <213854356+keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
Co-authored-by: Dmitry Tokarev <dtokarev@nvidia.com>
Co-authored-by: Julien Mancuso <161955438+julienmancuso@users.noreply.github.com>
Co-authored-by: Elizabeth Thomas <email2eliza@gmail.com>
Co-authored-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
Co-authored-by: Julien Darve <jdarve@NVIDIA.com>
Co-authored-by: Jasim Kareem <mj9034812@gmail.com>
Co-authored-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
Co-authored-by: Jie Hao <jihao@nvidia.com>
Co-authored-by: Jacky <18255193+kthui@users.noreply.github.com>
Co-authored-by: Qi Wang <qiwa@nvidia.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>

This branch was successfully deployed

1 active deployment
external_collaborator — 1daee8fd Deployed Sep 16, 2026 by glamr-agent via ok-to-test #18836
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

external-contribution Pull request is from an external contributor fix size/L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants