Repository navigation
fix(bindings): drain bridge tasks before interpreter finalization - #14813
Conversation
…ed port pool test_fetch_model_runtime_bridge_orders[invalid-config-then-fetch] was reported flaky with an environment_dependency category. The test body took dynamo_dynamic_ports for every parametrization, so all six scenarios reserved five host-wide ports (frontend, system, KV event, FPM, NIXL) through the shared allocator, which probes candidates with a real bind() and raises "Could not find N available ports after 100 retries" when the host is busy. Only fetch_then_backend_worker actually binds a socket; the other five run with DYN_SYSTEM_PORT=-1. Resolve the port lazily through request.getfixturevalue so only the scenario that needs one touches the allocator, raise the child subprocess budget from 20s to 60s (and the outer pytest timeout from 30s to 120s so the inner bound is what fires), and turn subprocess.TimeoutExpired into a pytest.fail carrying the partial child stdout and stderr so the next occurrence is classifiable instead of a bare exception. Adds a regression test asserting the invalid-config scenario still passes with the allocator patched to raise, plus a negative control asserting the backend scenario still fails without a port. Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
|
👋 Hi glamr-agent! Thank you for contributing to ai-dynamo/dynamo. Just a reminder: The 🚀 |
Automated evidence record — validation completeValidation status: complete Evidence summary: [3/3 validated] AI review assessment (advisory, not an approval): sound. An AI agent judged the change logically sound from the code and the reported validation results. Repository CI and human reviewers decide whether to merge. Validation result: pass — every command quoted in the change description reproduced here, including both rewritten paragraphs. Evidence audit: complete [3/3 validated] — the command report below comes from recorded runs. Commands and results [3/3 validated]Generated from the commands recorded during this run. Check 1Builds the changed Dynamo source and confirms that Python can import its compiled extension. Result: Passed ( Command: Not shown because the exact command contained private run data. Check 2Checks the changed files with the repository's fast lint and formatting commands. Result: Passed ( Command: Not shown because the exact command contained private run data. Check 3Runs the relevant Python unit tests without requiring a GPU. Result: Passed ( Command: Not shown because the exact command contained private run data. |
|
No description provided. |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (14)
Included review availability: Your plan provides up to 12 included reviews per hour; 9 remain after this review. WalkthroughThe Python bindings now use centralized future conversion and bridge-runtime helpers. Runtime adoption records process and distinct runtimes. An ChangesRuntime bridge centralization
Priority: ⬆️ High Estimated code review effort: 4 (Complex) | ~45 minutes Merge Risk: ⚪ Minimal · up to The runtime tracking paths now include both normal and detached distributed-runtime initialization, so the previously identified exit-drain gap is addressed and no actionable merge-blocking risk remains. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Comment |
|
CI reached a terminal state on |
|
factory: automated evidence record for the follow-up measurement run on this pull request. It supersedes the earlier automated evidence record on this request, which covered the original change and reported a passing validation. Automated evidence record — validation incompleteValidation status: incomplete Evidence summary: [2/3 validated · 1 failing] AI review assessment (advisory, not an approval): needs changes. An AI agent judged that the recorded measurement does not support the change as described. Repository CI and human reviewers decide whether to merge. Validation result: fail — the runs are green but do not support the fix; at 1800 runs per arm the change shows no reduction in failures, and the child Evidence audit: complete [2/3 validated · 1 failing] — the command report below comes from recorded runs. Commands and results [2/3 validated · 1 failing]Generated from the commands recorded during this run. Check 1Builds the changed Dynamo source and confirms that Python can import its compiled extension. Result: Passed ( Command: Not shown because the exact command contained private run data. Check 2Checks the changed files with the repository's fast lint and formatting commands. Result: Passed ( Command: Not shown because the exact command contained private run data. Check 3Runs the relevant Python unit tests without requiring a GPU. Result: Passed ( Command: Not shown because the exact command contained private run data. Details: The test runs completed green, but at 1800 runs per arm the changed test failed 14 times against 10 for unchanged |
glamr-agent
left a comment
There was a problem hiding this comment.
Automated AI review — advisory. The inline comments below are the findings. The full review is in a separate comment on this pull request.
|
factory: automated review
Assessment: soundSecond pass. Last time I said the fix was right and the description was wrong: The code did not move. The three things I asked for are in.
Two precisions on the new text, neither worth a round. The deadline line is What I checked myself, and you can run the same. The evidence, which I audited rather than re-ran. Three arms at the same
All of these exited Smaller things, your call
Untested, and what would test itNo request completed through the engine. That is independent of this change and fails the same way without it. What it Zero crashes in 300,000 children bounds the crash rate; it does not prove the |
|
nursery: @dynamo-ops please run full CI for 3620eb9 |
|
nursery: Handoff for this pass at
Next: authorize and run full CI for this SHA, inspect any failures, and obtain runtime-owner review, including disposition of the remaining optional thread. This PR is not claimed ready to merge. |
|
/ok-to-test 3620eb9 |
The child in tests/runtime/test_fetch_model_runtime_bridge.py printed its
success sentinel and then died with SIGSEGV, roughly once in every few
thousand runs. Every assertion had already passed, so the crash was in
process exit, not in anything the test asserts.
A captured backtrace names the faulting thread:
tokio::runtime::task::core::Core<T,S>::poll
-> pyo3_async_runtimes::tokio::TokioRuntime::spawn
-> generic::future_into_py_with_locals::{{closure}}
-> pyo3::marker::Python::with_gil
-> pyo3_async_runtimes::generic::set_result
-> pyo3_async_runtimes::call_soon_threadsafe
-> PyAnyMethods::call -> Bound<PyTuple>::drop
-> PyObject_GC_Del / PyObject_GC_UnTrack <- SIGSEGV
future_into_py resolves the Python future through call_soon_threadsafe, so
the awaiting coroutine resumes while the Tokio worker is still inside
with_gil with an epilogue to run - it has yet to drop the argument tuple.
The main thread can leave asyncio.run and begin finalization in that
window, and the worker then touches the GC of a half-torn-down
interpreter. The process runtime lives in a static, so nothing ever joins
those threads.
Register an atexit hook at module import that waits, bounded, for the
process runtime to have no live tasks. It releases the GIL first: an
atexit callback holds it, and the tasks being waited on need it, so
waiting while holding it would deadlock against those very threads. It
returns immediately when no runtime was ever created.
Also narrow this pull request to Dynamo behavior. The two tests asserting
what the pytest port allocator does, their helpers, the raised timeouts
and the unused partial-output helper are dropped. The conditional port
acquisition stays, moved into a fixture that reads the scenario parameter
so an exhausted pool is a setup error rather than a test failure.
Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
|
factory: automated evidence record for the current head Automated evidence record — validation completeValidation status: complete Evidence summary: [5/5 validated] AI review assessment (advisory, not an approval): sound. An AI agent judged the change logically sound from the code and the reported validation results. Repository CI and human reviewers decide whether to merge. Validation result: pass. With the hook in place, 300,000 child processes across two arms produced 0 non-zero exits, where the Caveats a reviewer should read rather than skip:
Change since the previous round: only the write-up. Evidence audit: complete [5/5 validated] — the command report below comes from recorded runs. Commands and results [5/5 validated]Generated from the commands recorded during this run. Check 1Builds the changed Dynamo source and confirms that Python can import its compiled extension. Result: Passed ( Command: Not shown because the exact command contained private run data. Check 2Checks the changed files with the repository's fast lint and formatting commands. Result: Passed ( Command: Not shown because the exact command contained private run data. Check 3Checks the changed Rust crates with Result: Passed ( Command: Not shown because the exact command contained private run data. Check 4Runs the relevant Python unit tests without requiring a GPU. Result: Passed ( Command:
Check 5Starts Dynamo with vLLM on one GPU and sends a real request. Result: Passed ( Command: Not shown because the exact command contained private run data. |
glamr-agent
left a comment
There was a problem hiding this comment.
Automated AI review — advisory. The inline comments below are the findings. The full review is in a separate comment on this pull request.
|
factory: CI reached a terminal state on |
|
factory: Answering the question about the failing pre-merge checks, and pushing the test removal that was asked for earlier in the same thread. Why the checks were red on That job has no steps and no log. It failing cancelled the other three legs of the matrix:
Test change dropped. Commit Checks run locally on the new head: Pushing this head starts One caveat on the description: its Summary still has a bullet for the port reservation, and the reviewer-start line still names the test file. Both are stale now. The title and description were edited outside this run, so I have left them as they are rather than overwrite that edit. |
|
/devin review @coderabbitai full review |
|
✅ Action performedFull review finished. |
The nested `if` only stayed under `clippy::collapsible_if` because a comment sat between the two conditions. With that comment gone, clippy rejects the shape, so fold the identity check into the `if let` as a let-chain. Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
|
factory: The last push broke Deleting the comment above the bridge-runtime identity check, as review asked, left this: if let Some(bridge) = BRIDGE_RUNTIME.get().copied() {
if !runtimes.iter().any(|rt| std::ptr::eq(*rt, bridge)) {
runtimes.push(bridge);
}
}
The fix folds the identity check into the if let Some(bridge) = BRIDGE_RUNTIME.get().copied()
&& !runtimes.iter().any(|rt| std::ptr::eq(*rt, bridge))
{
runtimes.push(bridge);
}For the record on the other two red clippy legs: Checks run locally on One good result from the previous head: |
|
factory: @dynamo-ops please run full CI for 7709b11 |
|
/ok-to-test 7709b11 |
|
/ok to test 7709b11 |
|
nursery: @dynamo-ops please run full CI for 405cba0 |
There was a problem hiding this comment.
Previously reported defects still present:
- Original discussion: The current checkout has no changed bridge-runtime test file, so the combined context-bridge, detached-runtime, and fetch ordering scenario remains untested as described in the open discussion.
|
nursery: Handoff for this pass on Current-head CI has 18 successful checks, 7 skipped checks, and 2 running Rust test jobs, with no failures so far. CodeRabbit's commit status is successful. Full CI has not started for this head; the authorization request is posted above and remains unanswered. Pending jobs: workspace Rust tests, Python-binding Rust tests. The runtime-codeowners review request through the requested_reviewers API returned HTTP 404 and did not land. A maintainer must route that review. The two context-first/detached coverage threads remain open for the existing DYN-4434 follow-up disposition; this pass added no tests or duplicate bot-review requests. Next: authorize full CI for this exact head, inspect its results and the running jobs, route runtime-owner review, and decide the two deferred coverage threads. No local tests, builds, or benchmarks were run. This PR is not ready to merge. |
|
/ok to test 405cba0 |
|
nursery: Handoff for this pass at Read the complete current diff, refreshed paginated reviews/conversation and CI evidence, and verified the earlier runtime-recording repairs. No new actionable code defect or merge conflict was found; no edits, commits, pushes, or thread resolutions were made. No local tests, builds, benchmarks, or new tests. Current-head results remain incomplete: 24 successful check runs, 33 skipped, seven running, none failed; CodeRabbit commit status is successful. All four Rust test jobs, four Clippy jobs, remote pre-commit, and the pre-merge gate passed. Full CI is already authorized and running on this exact SHA; no duplicate authorization or bot-review request was posted. Its changed-files log confirms
Requested The two coverage threads remain open for maintainer disposition and the existing DYN-4434 follow-up: context-first detached exit and combined context/detached/fetch order. The coverage is absent; this pass prohibits adding tests and does not claim those findings repaired. Existing replies already explain the deferral, so no duplicate thread replies were added. Next: preserve this head while CI finishes, inspect any failures, have a maintainer route runtime-owner review, and decide the two deferred coverage threads. Skipped and running checks are not passes; this PR is not ready to merge. |
|
nursery: CI investigation for I attempted to rerun job At this pass's final check, 54 check runs passed, 54 were skipped, one was cancelled, one failed, and two remain running: TRT-LLM tests and TRT-LLM two-GPU tests. All Rust tests, Clippy jobs, remote pre-commit, and the pre-merge gate passed. Full CI is already authorized for this SHA; no duplicate request was posted. No new actionable code defect or merge conflict was found. No edits, commits, pushes, local tests, builds, benchmarks, or new tests in this pass. The two deferred context-first/detached coverage threads still need maintainer disposition under DYN-4434; neither was marked repaired. Preserve this head while CI and review are handled. This PR is not ready to merge. |
jthomson04
left a comment
There was a problem hiding this comment.
Approved. The four P3 comments are non-blocking suggestions. Source review only; no tests were run and CI was not checked as part of this review.
…rage Look the bridge runtime up once through the OnceLock instead of asking PyO3 on every async conversion, hold the at-most-two runtimes the exit hook drains in a fixed array, call the plain std::thread::sleep import, and trim the surrounding comments to the constraints they carry. Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com>
|
factory: Took all four P3 comments on
Checks, all from Clippy earned its keep here: the first version of the cached lookup wrote The title and description are left as they are: they were edited by hand after the factory last wrote them, so this run does not own them any more. |
|
/devin review @coderabbitai full review |
🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
echo '== Pull request metadata =='
gh pr view 14813 --repo ai-dynamo/dynamo --json number,title,headRefOid,baseRefName,headRefName,mergeable,state,isDraft,reviewDecision,statusCheckRollup --jq '{number,title,headRefOid,baseRefName,headRefName,mergeable,state,isDraft,reviewDecision,checks: [.statusCheckRollup[]? | {name,conclusion,status}]}'
echo '== Diff summary and affected files =='
base="$(git merge-base "origin/main" HEAD)"
printf 'base=%s\nhead=%s\n' "$base" "$(git rev-parse HEAD)"
git diff --check "$base"...HEAD
git diff --stat "$base"...HEAD
git diff --name-only "$base"...HEAD
echo '== Rust implementation map =='
ast-grep outline lib/bindings/python/rust/lib.rs --items all --type function,struct,impl,static 2>/dev/null || true
ast-grep outline lib/bindings/python/rust/backend.rs --items all --type function,struct,impl 2>/dev/null || true
ast-grep outline lib/runtime/src/worker.rs --items all --type function,struct,impl 2>/dev/null || true
echo '== Current implementation sections =='
rg -n -C 8 -P '\b(BRIDGE_RUNTIME|bridge_runtime|record_bridge_runtime|adopt_bridge_runtime|wait_for_bridge_tasks_at_exit|register_at_exit|future_into_py|future_into_py_with_locals|ensure_fetch_model_runtime|DistributedRuntime::(new|detached)|existing_process_runtime)\b' \
lib/bindings/python/rust/lib.rs lib/bindings/python/rust/backend.rs lib/runtime/src/worker.rs
echo '== Direct PyO3 Tokio bridge API references in bindings =='
rg -n -C 3 -P '\bpyo3_async_runtimes::tokio::(future_into_py|future_into_py_with_locals|get_runtime|init_with_runtime)\b' lib/bindings/python/rust
echo '== Full diff for implementation files =='
git diff --find-renames --unified=50 "$base"...HEAD -- \
lib/bindings/python/rust/lib.rs \
lib/bindings/python/rust/backend.rs \
lib/bindings/python/rust/context.rs \
lib/bindings/python/rust/http.rs \
lib/bindings/python/rust/kserve_grpc.rs \
lib/bindings/python/rust/llm \
lib/bindings/python/rust/planner.rs \
lib/runtime/src/worker.rsLength of output: 226710 🏁 Script executed: #!/bin/bash
set -euo pipefail
echo '== Complete exit hook and registration =='
sed -n '170,330p' lib/bindings/python/rust/lib.rs
echo '== All PyO3 async runtime references and imports =='
rg -n -C 4 -P '\bpyo3_async_runtimes\b|use\s+.*\btokio\b' lib/bindings/python/rust
echo '== Runtime spawning and block_on calls in Python bindings =='
rg -n -C 4 -P '\.(spawn|block_on)\s*\(' lib/bindings/python/rust
echo '== All references to local wrappers =='
rg -n -C 2 -P '\b(crate::)?(future_into_py(?:_with_locals)?|bridge_runtime|adopt_bridge_runtime)\s*\(' lib/bindings/python/rust
echo '== Open review threads and reviews =='
gh api graphql -F owner=ai-dynamo -F name=dynamo -F number=14813 -f query='
query($owner: String!, $name: String!, $number: Int!) {
repository(owner: $owner, name: $name) {
pullRequest(number: $number) {
reviews(last: 30) { nodes { author { login } state body url commit { oid } } }
reviewThreads(first: 100) {
nodes {
isResolved
isOutdated
path
line
originalLine
comments(first: 20) { nodes { author { login } body url } }
}
}
}
}
}' --jq '
.data.repository.pullRequest as $pr |
{reviews: $pr.reviews.nodes,
openThreads: [
$pr.reviewThreads.nodes[]
| select(.isResolved | not)
| {isOutdated,path,line,originalLine,comments: [.comments.nodes[] | {author: .author.login, body, url}]}
]}'
echo '== Current CI result summary =='
gh pr view 14813 --repo ai-dynamo/dynamo --json headRefOid,statusCheckRollup --jq '{head: .headRefOid, checks: [.statusCheckRollup[]? | {name,status,conclusion}]}'Length of output: 121682
I found no new actionable defects. The bridge-runtime helpers now cover direct future conversion and direct bridge runtime access. The exit hook records and drains the process runtime and the distinct PyO3 runtime. It releases the GIL during the bounded wait and does not initialize a runtime at interpreter exit. I used the recorded maintainer guidance for this PR. I did not repeat the deferred context-first
🧠 Learnings used✅ Action performedFull review finished. |
|
factory: @dynamo-ops please run full CI for 1daee8f |
|
/ok to test 1daee8f |
|
nursery: Handoff at Current-head CI: 66 successful check runs, 50 skipped, one cancelled, none pending; CodeRabbit status is successful. All four Rust test jobs, all four Clippy jobs, remote pre-commit, and the pre-merge gate passed. Full CI ran but is incomplete: planner / Compliance cpu, arm64 exceeded its 15-minute limit during target-image package extraction. BuildKit bootstrap succeeded, but context loading stalled and extraction did not complete; scanning and attribution generation never ran. No source failure was shown. The returned main-branch comparison runs were from March and lacked this job, so they do not establish whether the timeout is recurring. Attempted a targeted retry through Requested Next: a maintainer retries the timed-out CI job, routes any outstanding runtime-owner review, and decides the deferred coverage threads. Merge readiness is not established. |
* feat: KV DC Relay file based source mode (ai-dynamo#14807) Add live-reloaded file sources for KV DC Relay namespace selection and expose readiness and source revisions through /engine/state. Preserve applied membership on invalid updates, coalesce discovery refreshes, and isolate native integration tests in forked processes. Signed-off-by: Nikita Sukharev <kaonael@gmail.com> * feat(sglang): expose cross-encoder reranking through /v1/rerank (ai-dynamo#14032) Signed-off-by: xianlubird <xianlubird@gmail.com> * fix(profiler): explain inaccessible model paths during trust checks (ai-dynamo#14860) Signed-off-by: hongkuanz <hongkuanz@nvidia.com> * fix(sglang): sync discovery from native pause state (ai-dynamo#13951) Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com> Co-authored-by: Zero Rains <57100978+zeroRains@users.noreply.github.com> * feat(recipes): add Solar Open2 250B NVFP4 aggregated and disaggregated recipes for B200 (ai-dynamo#14376) Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com> * refactor(agents): session_id reader from AgentContext + forward to vLLM (ai-dynamo#14428) Signed-off-by: Karen Chung <karenc@nvidia.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> * fix(discovery): allow served aliases for the same model source (ai-dynamo#14857) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix(router): reject unknown explicit worker targets (ai-dynamo#14858) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix(xpu): stabilize XPU test workers (ai-dynamo#14539) Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> Signed-off-by: VincyZhang <wenxin.zhang@intel.com> * feat(mm-routing): add Nemotron 3 Nano Omni video routing (ai-dynamo#14653) Signed-off-by: krishung5 <krish@nvidia.com> * fix(sglang): validate diffusion input_reference and bound media fetches (ai-dynamo#14435) The sglang image-diffusion and video-generation handlers passed the client-supplied input_reference through to the generator's image_path after only a non-empty check. Validate it first, and for remote references materialize it locally before the generator sees it, so the generator is always handed a trusted local path. This brings the sglang diffusion path in line with the vLLM/omni and trtllm backends, which already validate the same field. Behavior change: local I2I/I2V references now require DYN_MM_LOCAL_PATH to be set to the allowed directory; previously any path was accepted. common/http: - validate_media_reference() returns a plain filesystem path for local references; local_media_reference() is an async context manager that fetches a remote one through fetch_bytes(policy=...), which revalidates every redirect hop, into a temp file removed on exit. data: is rejected -- a URI is not a path. - fetch_bytes() gained max_bytes, streaming through collect_capped at an explicit read granularity so the cap is an allocation bound and not only a rejection: a 128 MiB-decoded gzip body against the 64 MiB cap peaks at 68,032,217 bytes rather than the whole decompressed body. Content-Length is caller-controlled and absent when chunked, and aiohttp's read(n) returns at most n bytes, so neither a header check nor a single capped read suffices. Defaults to None, leaving existing callers unchanged. - DYN_MM_MAX_FILE_SIZE_MB makes that cap operator-tunable, in megabytes, as the SGLang arg it replaces was. Read per call; empty, unparseable or non-positive falls back to 64 with a warning, so a malformed value neither takes the worker down nor reads as unlimited. - Messages built from caller input are bounded via describe_media_source, moved from multimodal/media_source.py (it pulls in torch) into url_validator.py and re-exported from its old home; a no-op below 120 characters. - HttpStatusError bounds its .message attribute, not only the rendered string: errors.rs::extract_http_like_error reads .status and .message off this class by name and forwards .message on a 4xx without calling str(). Backend exception text is bounded head-and-tail, since aiohttp renders the host before the errno. - validate_local_path uses exc.strerror rather than the raw OSError, whose text repeats the filename, and now catches the ValueError that Path.resolve() raises on an embedded NUL so callers keep their 4xx-vs-5xx decision. Rebased onto ai-dynamo#14563 (single aiohttp backend); the httpx-side half of the max_bytes plumbing went with that backend. Signed-off-by: nnshah1 <neelays@nvidia.com> Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com> Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(deps): upgrade fastokens to 0.3.2 (ai-dynamo#14798) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix(vllm): ship codec-free OpenCV for image inputs (ai-dynamo#14361) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: Anant Sharma <anants@nvidia.com> Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com> * docs: refresh community events Automated refresh from the public Dynamo Google Calendar. Generated by .github/workflows/community-events-refresh.yml. Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com> * ci: refresh the compliance baseline in auto-upgrade pipeline (ai-dynamo#14206) Signed-off-by: Anant Sharma <anants@nvidia.com> * feat(triton): honor KServe classification on tensor outputs (ai-dynamo#14783) Signed-off-by: Yingge He <yinggeh@nvidia.com> * docs(rl): stop the verl guide sending readers to a vLLM version it cannot run on (ai-dynamo#14571) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> * feat(mocker): publish native KV events from the vLLM gRPC server (ai-dynamo#14737) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix(kv-router): release unowned radix branches after eviction (ai-dynamo#14878) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix: show correct backend versions in the install selectors (ai-dynamo#13599) Signed-off-by: Anant Sharma <anants@nvidia.com> * build(vllm): prepare v0.29.0 bump (ai-dynamo#14543) Signed-off-by: Julien Darve <jdarve@NVIDIA.com> * ci(xpu): validation PR for the re-applied XPU workflows and Dockerfile Throwaway PR to prove the CI merged in #22 actually runs end to end on XPU hardware. Adds only a comment to container/templates/vllm_runtime.Dockerfile, which matches the `vllm` path filter (container/templates/vllm_*) and so makes changed-files set vllm=true, which is what gates build-xpu and the heterog-test-px-dn / heterog-test-pn-dx jobs. What this exercises: - .github/workflows/pr-xpu.yaml (push to pull-request/[0-9]+, needs the xpu label) - .github/workflows/pr-xpu-heterogeneous.yaml (push; its guard deliberately skips the label gate) - .github/workflows/epd-test-template.yml (workflow_call, from the heterog jobs) - .github/scripts/test-filters.js (the brace fix from #22) - container/templates/vllm_runtime.Dockerfile rendered and built for device=xpu Not exercised: .github/workflows/xpu-heterogeneous-dispatch.yaml is workflow_dispatch only and has to be run by hand from the Actions tab. The marker comment must be removed before this branch is ever merged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(triton): Update Triton Base Image to 26.08 (ai-dynamo#14854) Signed-off-by: J Wyman <jwyman@nvidia.com> Co-authored-by: Rini Gupta <rinig@nvidia.com> * fix(operator): normalize equivalent worker hash inputs (ai-dynamo#14721) Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> * test(sglang): exercise NIXL in embedding cache E/PD test (ai-dynamo#14795) Signed-off-by: Sai Kiran Polisetty <spolisetty@nvidia.com> * fix(sglang): stop the elastic-EP scale-up worker crash-looping at startup (ai-dynamo#14568) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com> * fix(responses): honor tool_choice when parsing tool calls from text (ai-dynamo#14843) Signed-off-by: xianlubird <xianlubird@gmail.com> * ci: accept trusted full-CI request comments (ai-dynamo#14868) Signed-off-by: Matej Kosec <mkosec@nvidia.com> * docs: clarify EPP mode boundary and single-replica Dynamo mode fixes [DYN-4310] (ai-dynamo#14756) Signed-off-by: Anna Tchernych <atchernych@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci(docs): move the generated-tables determinism gate out of link checking (ai-dynamo#14135) Signed-off-by: Dan Gil <dagil@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * ci(docs): generate the Kubernetes API reference at publish time (ai-dynamo#14122) Signed-off-by: Dan Gil <dagil@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix(operator): discover pull secrets for init containers (ai-dynamo#14922) Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> * fix(sglang): stop an unusable mooncake backend crashing workers after model load (ai-dynamo#14461) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: glamr-agent <glamr-agent@users.noreply.github.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> * fix(sglang): emit prefill handoff before completion in sidecar (ai-dynamo#14260) Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Connor Carpenter <connorc@nvidia.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> * test(trtllm): enable fault tolerance coverage (ai-dynamo#14609) Signed-off-by: tanmayv25 <tanmay2592@gmail.com> * fix(frontend): evict async tokenizer executors when the tokenizer is retired (ai-dynamo#13368) Signed-off-by: Peter Pan <Peter.Pan@daocloud.io> * fix(llm): report KServe datatypes by their wire names, not protobuf variants (ai-dynamo#14957) `ModelMetadata` reported each Triton-registered tensor's `datatype` using `inference::DataType::as_str_name()`, which returns the `model_config.proto` variant name (`TYPE_FP32`, `TYPE_STRING`, ...) instead of the KServe v2 wire names (`FP32`, `BYTES`, ...). Every datatype was wrong, so spec-conforming clients cannot parse any tensor the RPC describes. Adds `oip_name()` next to `tensor::DataType::to_kserve` covering all fifteen proto variants (incl. FP16 and BF16) and mapping `TYPE_STRING → BYTES`. Original PR by @ayaangazali: ai-dynamo#14770. Reissued under a signed commit to unblock the copy-pr-bot signature gate; diff is byte-identical. Closes ai-dynamo#14520. Signed-off-by: ayaangazali <ayaangazali@users.noreply.github.com> Signed-off-by: ayaangazali <ayaangazali.work@gmail.com> Signed-off-by: Vinya Kestur <vinyak@nvidia.com> Co-authored-by: ayaangazali <ayaangazali.work@gmail.com> * docs(mm-routing): document video KV routing (ai-dynamo#14958) Signed-off-by: krishung5 <krish@nvidia.com> * fix(sidecar): honor worker namespace suffix (ai-dynamo#14955) Signed-off-by: Biswa Panda <biswa.panda@gmail.com> * fix(bindings): drain bridge tasks before interpreter finalization (ai-dynamo#14813) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: Tushar Sharma <tusharma@nvidia.com> * fix(discovery): stop a Qwen3-VL worker from serving video with another worker's contract (ai-dynamo#14624) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> * fix(gms): honor configured timeout during initial weights admission (ai-dynamo#14877) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: Schwinn Saereesitthipitak <schwinns@nvidia.com> * feat(kv-router): add construction-time indexer delegates (ai-dynamo#14945) * fix(sglang): support min_tokens on tokenizer-free decode workers (ai-dynamo#14276) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: MatejKosec <mkosec@nvidia.com> * feat(router): add SessionPrefixIndexer for session-block lineage (ai-dynamo#13807) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: Karen Chung <karenc@nvidia.com> Signed-off-by: Matej Kosec <mkosec@nvidia.com> Co-authored-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Co-authored-by: Matej Kosec <mkosec@nvidia.com> * fix(vllm): settle kvwarm stages through a per-step round on every attention-DP rank (ai-dynamo#14728) Signed-off-by: Yiming Liu <yimingl@nvidia.com> * feat(vllm): benchmark hybrid caches with random KDA state (ai-dynamo#14900) Signed-off-by: hongkuanz <hongkuanz@nvidia.com> * fix(runtime): fix QUIC reassembly and reduce response stalls (ai-dynamo#14876) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * feat(router): unify frontend and standalone selection core (ai-dynamo#14570) Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com> Signed-off-by: Thomas Montfort <tjmontfort12@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Thomas Montfort <tjmontfort12@gmail.com> * fix(planner): keep control APIs responsive during Prometheus collection (ai-dynamo#14377) Signed-off-by: xianlubird <xianlubird@gmail.com> Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com> * fix(router): record SGLang prefill completion after stream ends (ai-dynamo#14968) Signed-off-by: jain-ria <riajain@NVIDIA.com> * fix(frontend): send inline media once on the TCP request plane (ai-dynamo#14801) Signed-off-by: Sumit Mishra <sah299610@gmail.com> Co-authored-by: Indrajit Bhosale <iamindrajitb@gmail.com> * docs: refresh community events Automated refresh from the public Dynamo Google Calendar. Generated by .github/workflows/community-events-refresh.yml. Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com> * fix(vllm): initialize synchronizer in KV warmup capacity test (ai-dynamo#14984) Signed-off-by: Alec Flowers <aflowers@nvidia.com> * fix(recipes): make the Solar Open2 250B benchmark and docs link usable (ai-dynamo#14956) Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com> * feat(recipes): add K-EXAONE 2.0 750B-A37B NVFP4 vLLM recipes for B200 (ai-dynamo#14822) Signed-off-by: Cheng Wang <chengwa@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: KVCR Resiliency Deployment Example (ai-dynamo#14695) Add two-node DynamoGraphDeployment examples for process-local KVCR and the KVCR memory service. Run one vLLM worker per GPU node, use stable Grove ordinals for cache-owner slots, and request GPU-local RDMA resources for engines and Guard services. Provide a deployment helper for rendering and selecting either variant. Run the KV state agent alongside vLLM for process-local host memory. In memory-service mode, keep KVCR and the state agent in a separate container so its Guard and shared-memory pool survive engine restarts. Document that restarting the services sidecar invalidates the MVP recovery contract and requires deployment-level replacement. Add manifest coverage and an opt-in two-host lifecycle test. Kill the source EngineCore, hold it offline, and verify that the promoted Guard serves its preserved cache to the surviving target. Correlate response equality and KVCR transfer metrics with transmit and receive counters from the selected active HCA to prove RDMA transport. Pin compatible KVCR and vLLM revisions and document the runtime, discovery, compatibility-digest, and recovery prerequisites. Signed-off-by: Adit Ranadive <aranadive@nvidia.com> * feat(omni): add Nemotron Audex speech synthesis to /v1/audio/speech (ai-dynamo#12788) Signed-off-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com> * ci: allow glamr-agent to request CI on its own unsigned PRs (ai-dynamo#14964) Signed-off-by: Matej Kosec <mkosec@nvidia.com> * fix(vllm): isolate multimodal worker ports (ai-dynamo#14751) Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com> Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com> * fix(runtime): reject invalid DYN_REQUEST_PLANE values (ai-dynamo#12612) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: Matej Kosec <mkosec@nvidia.com> Signed-off-by: Coding Agent <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: MatejKosec <mkosec@nvidia.com> * fix(responses): preserve text instead of inferring tool calls (ai-dynamo#14846) Signed-off-by: xianlubird <xianlubird@gmail.com> Co-authored-by: Ryan McCormick <rmccormick@nvidia.com> * chore: temporarily increase frontend build time limit 45 --> 90 min (ai-dynamo#15019) Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com> * test(operator): cover scoped CA injection ownership (ai-dynamo#14961) Signed-off-by: Julien Mancuso <jmancuso@nvidia.com> * feat(frontend): map semantic errors to HTTP responses (ai-dynamo#14396) Signed-off-by: Biswa Panda <biswa.panda@gmail.com> * docs: correct fault-tolerance architecture details (ai-dynamo#14880) Signed-off-by: Elizabeth Thomas <email2eliza@gmail.com> * build(deps): bump nats-server to v2.14.7 (ai-dynamo#14919) Signed-off-by: Dan Gil <dagil@nvidia.com> * build(deps): bump AISimulate to 0.12.0 (ai-dynamo#15012) Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com> * remove oneAPI env for XPU detection * feat(backends): expose native LoRA capacity in model registration (ai-dynamo#14754) Signed-off-by: Julien Darve <jdarve@NVIDIA.com> Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com> * fix(planner): handle pending decisions in virtual connector wait (ai-dynamo#14841) Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com> * feat(vllm): add sidecar LoRA lifecycle (ai-dynamo#13068) Signed-off-by: Julien Darve <jdarve@NVIDIA.com> Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> Co-authored-by: Julien Darve <jdarve@NVIDIA.com> Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com> * fix(vllm/omni): pass response_format into video EngineInputs (ai-dynamo#14667) (ai-dynamo#14844) * chore: bump version to 1.6.0 post 1.5.0 branch cut (ai-dynamo#15009) Signed-off-by: pvijayakrish <pvijayakrish@nvidia.com> Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ci): Use `pytest --ignore` to Skip Tests Based on Framework (ai-dynamo#14815) Signed-off-by: J Wyman <jwyman@nvidia.com> * feat(sidecar): add e2e CI testing for sidecar launch scripts (ai-dynamo#14508) Signed-off-by: tanmayv25 <tanmay2592@gmail.com> Signed-off-by: Julien Darve <jdarve@NVIDIA.com> Co-authored-by: Julien Darve <jdarve@NVIDIA.com> * chore(xpu): upgrade vllm and omni to 0.29.0 Signed-off-by: wenxin.zhang <wenxin.zhang@intel.com> * docs(operator): document the DGDR workload-creation trust boundary (ai-dynamo#14429) Signed-off-by: nnshah1 <neelays@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * fix(xpu): use released vllm-omni prerelease Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> * test(efa): add the EFA disaggregated deploy test for sglang (ai-dynamo#13893) Signed-off-by: Jie Hao <jihao@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(runtime): support IPv6-only IP resolution (ai-dynamo#13126) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * docs(fault-tolerance): clarify migration after shutdown grace expires (ai-dynamo#14872) Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com> * feat(vllm-omni): preserve generated video audio (ai-dynamo#13707) Signed-off-by: Guan Luo <gluo@nvidia.com> Co-authored-by: Guan Luo <gluo@nvidia.com> * feat(vllm-omni): pass model-specific video parameters (ai-dynamo#13708) Signed-off-by: Guan Luo <gluo@nvidia.com> Co-authored-by: Guan Luo <gluo@nvidia.com> * feat(vllm-omni): qualify MiniMax-H3 T2VA on B200 (ai-dynamo#13589) Signed-off-by: Guan Luo <gluo@nvidia.com> Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com> Co-authored-by: Guan Luo <gluo@nvidia.com> Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com> Co-authored-by: Ryan McCormick <rmccormick@nvidia.com> * fix(vllm): remove obsolete Omni compatibility guard Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> * fix(vllm): retain Omni compatibility guard Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> * .github/workflows/pr-xpu-heterogeneous.yaml; pin GPU_TAG to latest * .github/workflows/; add post-merge and nightly XPU heterogeneous CI Extract the XPU heterogeneous P/D pipeline out of pr-xpu-heterogeneous.yaml into xpu-heterogeneous-run.yml, a workflow_call reusable workflow, and call it from three thin trigger workflows so all three merge phases run the identical pipeline instead of drifting copies. xpu-heterogeneous-run.yml new, reusable. guard, changed-files, build-xpu, build-nvidia, resolve-images and both heterog tests, unchanged, plus 7 inputs. pr-xpu-heterogeneous.yaml reduced to the pre-merge trigger, the slash-command gate and the reaction. post-merge-xpu-heterogeneous.yaml new. push to main. nightly-xpu-heterogeneous.yaml new file, but the cron is MOVED, not added: it is the 0 23 * * * schedule that was already in pr-xpu-heterogeneous.yaml. No behaviour change per phase. force_all_tests replaces the old github.event_name == 'schedule' || github.event_name == 'issue_comment' expression with the same truth table: pre-merge passes github.event_name == 'issue_comment', nightly passes true. Post-merge also passes true, because a push to main has no PR base for .github/actions/changed-files to diff against, and post-merge exists to catch what per-PR gating missed. xpu-status-check stays a TOP-LEVEL job in each caller rather than moving into the reusable workflow. A job contributed by a reusable workflow reports to the Checks API as "run / xpu-status-check", so hosting it there would rename the context and leave any branch protection rule requiring xpu-status-check waiting forever on a check that no longer reports. The concurrency mapping stays byte-identical across all four workflows that touch this hardware, now including xpu-heterogeneous-dispatch.yaml. Three files do NOT get three slots: the cluster, the dynamo-system namespace and the onexpu-/onenvidia-rdma-kueue ResourceClaimTemplates are one global resource. The reusable workflow deliberately carries no concurrency block of its own, which would deadlock against the slot the caller's run already holds. Parameterised gpu_tag, model, tensor_parallel and runner as inputs so the callers can diverge; all default to the previously hardcoded values. Added workflow_dispatch to the nightly, without which a schedule-only workflow cannot be exercised before it reaches the default branch. Verified: all files parse; the four concurrency mappings are byte-identical; the reusable workflow declares no concurrency; every input each caller passes exists and every required input is supplied; nesting is depth 3 of the 4 GitHub allows. actionlint was not available to run, and will report queue:max as an unknown key in all four files, a known false positive. --------- Signed-off-by: Nikita Sukharev <kaonael@gmail.com> Signed-off-by: xianlubird <xianlubird@gmail.com> Signed-off-by: hongkuanz <hongkuanz@nvidia.com> Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com> Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com> Signed-off-by: Karen Chung <karenc@nvidia.com> Signed-off-by: jthomson04 <jwillthomson19@gmail.com> Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> Signed-off-by: VincyZhang <wenxin.zhang@intel.com> Signed-off-by: krishung5 <krish@nvidia.com> Signed-off-by: nnshah1 <neelays@nvidia.com> Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com> Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com> Signed-off-by: Anant Sharma <anants@nvidia.com> Signed-off-by: Yingge He <yinggeh@nvidia.com> Signed-off-by: Julien Darve <jdarve@NVIDIA.com> Signed-off-by: J Wyman <jwyman@nvidia.com> Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> Signed-off-by: Sai Kiran Polisetty <spolisetty@nvidia.com> Signed-off-by: Matej Kosec <mkosec@nvidia.com> Signed-off-by: Anna Tchernych <atchernych@nvidia.com> Signed-off-by: Dan Gil <dagil@nvidia.com> Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Signed-off-by: glamr-agent <glamr-agent@users.noreply.github.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Signed-off-by: tanmayv25 <tanmay2592@gmail.com> Signed-off-by: Peter Pan <Peter.Pan@daocloud.io> Signed-off-by: ayaangazali <ayaangazali@users.noreply.github.com> Signed-off-by: ayaangazali <ayaangazali.work@gmail.com> Signed-off-by: Vinya Kestur <vinyak@nvidia.com> Signed-off-by: Biswa Panda <biswa.panda@gmail.com> Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com> Signed-off-by: Yiming Liu <yimingl@nvidia.com> Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com> Signed-off-by: Thomas Montfort <tjmontfort12@gmail.com> Signed-off-by: Sumit Mishra <sah299610@gmail.com> Signed-off-by: Alec Flowers <aflowers@nvidia.com> Signed-off-by: Cheng Wang <chengwa@nvidia.com> Signed-off-by: Adit Ranadive <aranadive@nvidia.com> Signed-off-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com> Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com> Signed-off-by: Coding Agent <svc-glamr@nvidia.com> Signed-off-by: Julien Mancuso <jmancuso@nvidia.com> Signed-off-by: Elizabeth Thomas <email2eliza@gmail.com> Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com> Signed-off-by: pvijayakrish <pvijayakrish@nvidia.com> Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com> Signed-off-by: wenxin.zhang <wenxin.zhang@intel.com> Signed-off-by: Jie Hao <jihao@nvidia.com> Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com> Signed-off-by: Guan Luo <gluo@nvidia.com> Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com> Co-authored-by: Nikita Sukharev <kaonael@gmail.com> Co-authored-by: Xianlu Bird <xianlubird@gmail.com> Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com> Co-authored-by: William Arnold <7565007+Aphoh@users.noreply.github.com> Co-authored-by: Zero Rains <57100978+zeroRains@users.noreply.github.com> Co-authored-by: snarravula-dl <snarravula@nvidia.com> Co-authored-by: Karen Chung <karenc@nvidia.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: jthomson04 <jwillthomson19@gmail.com> Co-authored-by: VincyZhang <wenxin.zhang@intel.com> Co-authored-by: Kris Hung <krish@nvidia.com> Co-authored-by: Neelay Shah <neelays@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: Anant Sharma <anants@nvidia.com> Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com> Co-authored-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com> Co-authored-by: Yingge He <157551214+yinggeh@users.noreply.github.com> Co-authored-by: JulienDarve <86800349+JulienDarve@users.noreply.github.com> Co-authored-by: J Wyman <jwyman@nvidia.com> Co-authored-by: Rini Gupta <rinig@nvidia.com> Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com> Co-authored-by: Sai Kiran Polisetty <spolisetty@nvidia.com> Co-authored-by: MatejKosec <mkosec@nvidia.com> Co-authored-by: atchernych <atchernych@nvidia.com> Co-authored-by: Dan Gil <dagil@nvidia.com> Co-authored-by: Bojiang Li <327132355+bojiang-li@users.noreply.github.com> Co-authored-by: Connor Carpenter <connorcarpenter15@gmail.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Connor Carpenter <connorc@nvidia.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Tanmay Verma <tanmayv@nvidia.com> Co-authored-by: Peter Pan <peter.pan@daocloud.io> Co-authored-by: Vinya Kestur Tumakuru Arun Kumar <vinyak@nvidia.com> Co-authored-by: ayaangazali <ayaangazali.work@gmail.com> Co-authored-by: Biswa Panda <biswa.panda@gmail.com> Co-authored-by: Tushar Sharma <tusharma@nvidia.com> Co-authored-by: Schwinn Saereesitthipitak <schwinns@nvidia.com> Co-authored-by: Ryan Olson <ryanolson@users.noreply.github.com> Co-authored-by: Yimingl_Nvidia <yimingl@nvidia.com> Co-authored-by: Thomas Montfort <tjmontfort12@gmail.com> Co-authored-by: Sumit884-byte <sah299610@gmail.com> Co-authored-by: Indrajit Bhosale <iamindrajitb@gmail.com> Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com> Co-authored-by: chw001 <chengwa@nvidia.com> Co-authored-by: Adit Ranadive <aranadive@nvidia.com> Co-authored-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com> Co-authored-by: Keiven C <213854356+keivenchang@users.noreply.github.com> Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com> Co-authored-by: Ryan McCormick <rmccormick@nvidia.com> Co-authored-by: Dmitry Tokarev <dtokarev@nvidia.com> Co-authored-by: Julien Mancuso <161955438+julienmancuso@users.noreply.github.com> Co-authored-by: Elizabeth Thomas <email2eliza@gmail.com> Co-authored-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com> Co-authored-by: Julien Darve <jdarve@NVIDIA.com> Co-authored-by: Jasim Kareem <mj9034812@gmail.com> Co-authored-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com> Co-authored-by: Jie Hao <jihao@nvidia.com> Co-authored-by: Jacky <18255193+kthui@users.noreply.github.com> Co-authored-by: Qi Wang <qiwa@nvidia.com> Co-authored-by: Guan Luo <gluo@nvidia.com> Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Overview
Reduce the observed fetch-model bridge shutdown race by giving existing Tokio tasks a bounded opportunity to finish before Python interpreter finalization.
Summary
atexithook that releases the GIL while polling the process runtime and any distinct, recorded PyO3 bridge runtime, deduplicating shared runtimes.DistributedRuntimeconstructors, including the context-first then detached path.Details
The exit hook waits at most 250 ms and logs at info level if tasks remain. Processes with persistent service tasks can pay that entire additional shutdown delay. At the deadline, finalization proceeds with the same residual exposure as before the hook; this is a mitigation, not a guarantee that arbitrary pending Python cleanup is safe.
Worker::existing_process_runtime()reads the process runtime without creating worker threads during exit. Bridge adoption records whichever runtime PyO3 already selected, preserving a distinct context-first runtime. The detached constructor now uses that same adoption helper before wrapping the process runtime.The test file is unchanged from
main. The earlier port-reservation fixture and allocator test changes have been removed. Coverage for the combined context-first then detached-runtime order remains deferred to DYN-4434, with two review threads awaiting maintainer disposition.Where should the reviewer start?
Review
bridge_runtime, the future-conversion helpers,adopt_bridge_runtime,wait_for_bridge_tasks_at_exit, and bothDistributedRuntimeconstructors inlib/bindings/python/rust/lib.rs, followed bylib/runtime/src/worker.rs.Validation
Prior factory comments report measurements and local validation of earlier commits; those results do not validate the current head. Nursery validation uses remote CI only. For
1daee8fdbf44395d91613c209e84dff4df2ea23e, the September 16 CI snapshot has 65 successful check runs, 50 skipped, and one cancelled, with no queued or running checks. All four Rust test jobs, all four Clippy jobs, remote pre-commit, and the pre-merge gate passed. Full CI was authorized and ran, butplanner / Compliance cpu, arm64hit its 15-minute limit during image-package extraction before scanning. The targeted retry was rejected with HTTP 403 (repository admin rights required), so full CI is incomplete. This maintenance pass ran no local tests, builds, or benchmarks and added no tests.Related Issues
🔗 This PR is linked to an issue:
Summary by CodeRabbit