Repository navigation
feat(model): add shared NGC fetching support with vLLM/SGLang integration - #15441
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: ai-dynamo/dynamo/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (10)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. WalkthroughThe change adds NGC provider selection for model fetching and cache lookup. SGLang and vLLM resolve NGC model URIs to local paths and pass those paths into their model setup flows. ChangesNGC model resolution
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~25 minutes Merge Risk: ⚪ Minimal · up to No actionable merge-blocking issue was identified. NGC cache-layout compatibility remains unverified, and authenticated downloads and two-worker transfers still need validation. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Comment ✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
|
2c82022 to
991eaba
Compare
|
/ok to test 5ce7124 |
❌ Dynamo PR CI failed — run 37377438887 (attempt 3) on
|
| Other | Jobs |
|---|---|
| dynamo-runtime | ❌ 2 ⏹️ 1 ✅ 5 |
| frontend | ⏹️ 1 |
Failure details
2 jobs failed (amd64 and arm64) with the same 2 tests in tests/fault_tolerance/cancellation/test_utils.py timing out after 30s waiting for the HTTP response. The 2 cancelled jobs (dynamo-runtime / rust-gpu, frontend / Build multi-arch cpu) were cancelled within minutes of the arm64 failure, so they look like fail-fast cancels.
❌ dynamo-runtime / test / parallel cuda13.0, amd64: test_drained_stream_* requests.exceptions.Timeout (30s)
Job: dynamo-runtime / test / parallel cuda13.0, amd64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112071996083 -R ai-dynamo/dynamo --log-failed
______________ test_drained_stream_can_require_generated_content _______________
tests/fault_tolerance/cancellation/test_utils.py:50: in test_drained_stream_can_require_generated_content
read_streaming_responses(
tests/fault_tolerance/cancellation/utils.py:309: in read_streaming_responses
response_raw = cancellable_req.get_response()
drain = True
require_content = True
tests/fault_tolerance/cancellation/utils.py:188: in get_response
raise requests.exceptions.Timeout(
E requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
FAILED tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_can_require_generated_content - requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
FAILED tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_accepts_generated_content - requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
=== 2 failed, 789 passed, 23 skipped, 5401 deselected in 1139.89s (0:18:59) ====
test_drained_stream_can_require_generated_content and test_drained_stream_accepts_generated_content both fail in CancellableRequest.get_response() (utils.py:188), which got no HTTP response within 30s (drain=True, require_content=True). Both failed again on every auto-retry. A later Upload test logs error about an * in the artifact path test_unsafe_or_nonexact_protobuf_pins[protobuf==6.33.*] is a follow-on step error, not the cause.
❌ dynamo-runtime / test / parallel cuda13.0, arm64: same test_drained_stream_* Timeout (30s)
Job: dynamo-runtime / test / parallel cuda13.0, arm64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112071996268 -R ai-dynamo/dynamo --log-failed
tests/fault_tolerance/cancellation/utils.py:188: in get_response
raise requests.exceptions.Timeout(
E requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
FAILED tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_can_require_generated_content - requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
FAILED tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_accepts_generated_content - requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
=== 2 failed, 789 passed, 23 skipped, 5401 deselected in 1472.27s (0:24:32) ====
The same 2 tests fail the same way as on amd64: no HTTP response from CancellableRequest.get_response() within 30s. All other tests passed.
For agents
{"pr": 15441, "run_id": 37377438887, "run_attempt": 3, "head_sha": "6875f908142108c41b68cc42215f0de45c2bdef6",
"failures": [
{"job": "dynamo-runtime / test / parallel cuda13.0, amd64", "job_id": 112071996083, "failed_step": "Run CPU-only tests (parallelized)", "signature": "requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response", "tests": ["tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_can_require_generated_content", "tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_accepts_generated_content"], "log_cmd": "gh run view --job 112071996083 -R ai-dynamo/dynamo --log-failed"},
{"job": "dynamo-runtime / test / parallel cuda13.0, arm64", "job_id": 112071996268, "failed_step": "Run CPU-only tests (parallelized)", "signature": "requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response", "tests": ["tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_can_require_generated_content", "tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_accepts_generated_content"], "log_cmd": "gh run view --job 112071996268 -R ai-dynamo/dynamo --log-failed"}
]}Posted automatically by Devin for run 37377438887. Updated on every full-CI run of this PR.
be886a0 to
b9240ee
Compare
…tion Route NGC sources through ModelExpress and reuse complete cached models. Resolve local paths before vLLM and SGLang engine initialization while preserving existing Hugging Face P2P acquisition paths. Signed-off-by: sunil1511 <sunilupare45@gmail.com>
Signed-off-by: sunil1511 <sunilupare45@gmail.com>
Signed-off-by: sunil1511 <sunilupare45@gmail.com>
…e to all consumers Signed-off-by: sunil1511 <sunilupare45@gmail.com>
b9240ee to
48af879
Compare
…overage Signed-off-by: sunil1511 <sunilupare45@gmail.com>
|
/ok to test 6875f90 |
|
/ok to test 228ded5 |
Merges origin/main at ce45db9. Four files conflicted, all resolved by keeping both sides: - instrumented_scheduler.py imports: this branch's importlib.metadata version import and main's itertools chain (vLLM 0.31.0 bump, #15643). - worker_factory.py imports from instrumented_scheduler: main's InstrumentedScheduler (#12545) and this branch's benchmark_content_point_key. - test_vllm_worker_factory.py imports: main's Config (#15441), ENV_FPM_WORKER_ID, InstrumentedScheduler and FPM_SET_WORKER_ID_METHOD_NAME (#12545), next to this branch's FpmBenchmarkWorkerExtension, ENV_FPM_BENCHMARK_OUTPUT_PATH, benchmark_content_point_key and the probe and restore timeouts. - test_vllm_instrumented_scheduler.py: both sides appended tests at the end of the file. This branch's engine provenance and measurement tests come first, then main's FPM worker_id propagation tests (#12545). Main's vLLM 0.31.0 bump deleted test_vllm_kv_cache_metadata_compat.py, which the realseed_prefix_cache fixture's comment named as the importorskip probe that its function-local KVCacheManager import protects. The comment now names test_vllm_dcp_kv_events.py, which still probes vllm.v1.core.kv_cache_manager. No code change for vLLM 0.31.0: the vLLM interfaces this branch uses (collective_rpc by method name, worker extension mixing, cudagraph_metrics, CUDAGraphStat, the stat loggers, the fields the engine probe and the provenance capture read) are unchanged from 0.30.0. Signed-off-by: Yiming Liu <yimingl@nvidia.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Overview / Summary
Add shared NGC model fetching through ModelExpress and integrate it into vLLM and SGLang worker startup.
This builds on ai-dynamo/modelexpress#232. NGC sources are resolved to local directories before engine initialization, while existing Hugging Face ModelExpress acquisition and P2P paths are preserved.
Details
ngc://sources through ModelExpress’s NGC provider.--served-model-namefor NGC models in SGLang.Existing Hugging Face behavior
The new full-download and local-path rewrite behavior applies only to NGC sources.
HF models retain their source identifiers. With ModelExpress loaders enabled, vLLM and SGLang continue to skip Dynamo’s full weight prefetch, and model registration continues to request metadata only.
Scope and NGC P2P limitation
The fetching implementation is shared, but this PR integrates it only into the standard vLLM and SGLang worker startup paths. TensorRT-LLM and other entry points require separate integration and validation.
NGC weights are downloaded or found in the local cache before the ModelExpress weight loader runs. Consequently, peer loading cannot avoid the initial NGC download on a cold-cache worker.
P2P is not explicitly disabled, but peer matching currently depends on the resolved local path and other compatibility fields. Optimized NGC P2P startup requires follow-up work on metadata-only preparation, stable source identity, and NGC-aware fallback.
Where should the reviewer start?
lib/llm/src/hub.rscomponents/src/dynamo/vllm/args.pyandcomponents/src/dynamo/sglang/args.pycomponents/src/dynamo/vllm/main.pyandcomponents/src/dynamo/vllm/worker_factory.pyValidation
Related Issues
Summary by CodeRabbit
ngc://across supported inference engines.