[NVBUG-6448152][test] isolate gen status caller guard - #17220
[NVBUG-6448152][test] isolate gen status caller guard#17220chienchunhung wants to merge 12 commits into
Conversation
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Yihan Wang <yihwang@nvidia.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> (cherry picked from commit 6fc7f33)
…sensus factor Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> (cherry picked from commit 9c8bf59)
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> (cherry picked from commit 4b182f1)
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #63548 [ run ] triggered by Bot. Commit: |
|
PR_Github #63548 [ run ] completed with state
|
|
Terminal TEST ONLY result: the exact three-node GB300 workload completed 512/512 requests with zero client failures, b_is_valid=true, clean CTX/GEN/disaggregated shutdown, and 800.66 output tok/s (13,611.26 total tok/s). The exact slow baseline was 799.73 output tok/s, so the treatment changed throughput by only +0.12%. The caller guard was fully exercised: every CTX rank skipped all 643 empty loop-entry checks; every GEN rank emitted both skip and active-call evidence, skipping 194,534–241,284 of 244,601 checks while retaining active-transfer calls. Therefore restoring the pre-PR #15356 PyExecutor caller guard does not recover throughput and rules empty/idle status-poll cadence down as a sufficient or dominant cause for this workload. The CI failure is the expected performance-regression threshold, not a functional failure. Closing unmerged; the exact diagnostic branch is preserved and must not be rerun unchanged. |
Caution
TEST ONLY — never merge. This draft is a single-factor performance diagnostic for NVBUG 6448152.
Question
Does the unconditional per-loop generation-transfer status call introduced by PR #15356 cause the localized output-throughput regression by issuing empty/idle status work?
The adjacent comparison already localizes the regression to PR #15356 as a commit unit: parent 1515.84 output tok/s versus child 799.73. Prior admission diagnostics did not recover throughput, and the complete first-gate replay did not change an admission decision on its measured trace: terminal first-gate evidence.
Frozen causal identity
dd4a7ac2992f2b02dce381ce4e630c4ba41994363835a8758ca0a20a69a6c13275ce767d2bfc2e067f36854604a2f9aee4bd8826b9f30788b0124ea6bed490ecfc567777ebbdbc8b3721d6b60fa9b39f73efcee5922de49560ca39ba7938a944e5d13bd3f71a36a24d3a718087eb20b23cb45d4e88031abf1e2161b613f9190ea1689c476c4221c717dc850b0d73e4de167bd9d52f51ae62The pre-main treatment is preserved at
codex/backup-nvbug-6448152-gen-status-caller-guard-pre-main-20260803-3835a87.Single treatment
With the TEST ONLY environment switch enabled,
_check_disagg_gen_transfer_statusrestores only the pre-PR #15356 entry predicate outside KV-capacity warmup:atLeastRequestNum=0call.It deliberately does not restore the old
atLeastRequestNum=1selection. The immediate receive-side poll and admission-progress poll are unchanged. The bounded C++ callee, packed terminal consensus, ready-ID gathers, both admission gates, scheduler, runtime, image, model, dataset, and topology remain unchanged.Warmup preserves the baseline unconditional call and is excluded from diagnostic counters.
Exercised evidence
ID-free startup, first-decision, and shutdown markers report topology and per-loop counts on every worker:
The exact topology makes the diagnostic safe and distinguishes the expected effects:
Frozen workload and acceptance
Run only the existing GB300 three-node DeepSeek-R1 128K/8K context-first Python/NIXL selector: CTX PP4, GEN DEP8, 256 concurrency, two rounds (512 total requests), multiplier 1, and the current rendered image.
Required before interpreting throughput:
b_is_valid=true;skipped_loop_calls > 0on every CTX rank;Interpretation against the exact 799.73 tok/s slow baseline:
Local validation
2683unique entries validated) because the configured older hook runner cannot parsestr | None.transformerspackage.