[NVBUG-6448152][test] TEST ONLY retry pre-admission native 85665f5f - #16960
[NVBUG-6448152][test] TEST ONLY retry pre-admission native 85665f5f#16960chienchunhung wants to merge 10 commits into
Conversation
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Yihan Wang <yihwang@nvidia.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> (cherry picked from commit 6fc7f33)
…sensus factor Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> (cherry picked from commit 9c8bf59)
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> (cherry picked from commit 4b182f1)
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #62236 [ run ] triggered by Bot. Commit: |
|
PR_Github #62236 [ run ] completed with state
|
Terminal outcome: censored before pytestThe single pre-admission discriminator run reached the exact requested GB300 stage at head Before pytest collection or workload startup, the disaggregated-server process failed while importing This is a runtime/package compatibility censor, not a performance result. There is no JUnit result, request count, protocol-mode evidence, shutdown evidence, or output-token metric. Consequently:
This TEST ONLY draft is being closed unmerged. Its diagnostic branch is preserved. |
|
The import blocker was traced to the historical TensorRT 10.15/PyTorch 26.02 dependency input being installed inside the frozen TensorRT 10.16.1 runtime. The dependency-compatible replacement discriminator ports the exact upstream DLFW 26.04 requirements update used by the completed comparator, without changing product source or workload behavior. Its single targeted GB300 launch is active; this closed branch remains preserved. |
TEST ONLY
Retry of the exact pre-admission historical-source/current-runtime discriminator after the prior attempt was censored in prerequisite compilation. This draft is diagnostic only and is not intended for merge.
85665f5fd3, exact native tree8ada9734721bcd127722b367152eeb5bb22fa4ca, parenteaf5693b3aee9acea0d322dc8be997c4c0975280.9ff00c869a36055fcf2ebf4de4950a0471a3039f.mainappears only through no-tree CI-ancestry merges; it contributes no product content to the tested tree.2128f376c81a96a070037f6d139abaf6c9b37092. It pins DeepEP compilation to the vendored NVSHMEM headers, removes a Clang-only flag after the sub-build switches to GCC, and normalizes CUDA virtual-memory aggregate initialization. It changes no admission, scheduling, transceiver, workload, or benchmark behavior.The merged asynchronous-consensus change is not under re-review here; this draft only localizes the separate source-throughput regression.