Repository navigation
Conversation
There was a problem hiding this comment.
Exact-head self-review at 2cd136f311c02f26d7680cc57a72eac7bf8019b2
Reviewed the complete four-file diff and the READY/MERGE/NIGHTLY MiniJinja semantic renders.
- READY/MERGE add only the exact heavy-file ignore; all other Model Executor and PersonaPlex selection remains intact.
- Nightly has one serial, nonblocking 90-minute owner with nonzero collection evidence, durations, JUnit, complete log, and summary artifacts.
- The existing runner preserves piped pytest failures and exit 5; no allow-no-tests override, retry, xdist, template edit, or CUDA YAML change is present.
- Contract tests cover source routing, rendered soft-fail/dependency behavior, and unchanged CUDA ownership.
2cd136f to
64a6bd5
Compare
Assisted-by: GPT-5.6 Sol Signed-off-by: haic0 <149741444+haic0@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: andyluo7 <andy.luo@amd.com>
64a6bd5 to
51a0d26
Compare
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR was classified as CI work. CI owner: @yenuo26 @congw729 @NickCao Routing: @yenuo26 via semantic router, CI owner, CODEOWNERS; @congw729 via CODEOWNERS; @NickCao via CODEOWNERS @haic0, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
andyluo7
left a comment
There was a problem hiding this comment.
Reviewed the exact four-file CI routing change at 51a0d26. READY/MERGE retain the ten fast PersonaPlex nodes and deselect only the four long stress cases; the dedicated nightly lane collected and passed those four cases in 1:25:28 with artifacts uploaded. AMD build 13573, CUDA build 16986, GitHub Actions, DCO, and docs are green. The unrelated R2-01 GPU coverage failures were soft-failed and outside this patch.
Purpose
Reduce the AMD
mi300_1Model Executor critical-path time by moving only the four long-running PersonaPlex temporal oracle cases from READY/MERGE into one serial, nonblocking nightly job.Historical observations from AMD builds #13509 and #13511 put the four stress cases at roughly 57-84 minutes total. Hash sharding uses node identity/count rather than measured duration, so these loop-heavy cases disproportionately extend Model Executor shards 1/3 and 3/3.
Coverage and behavior
core_model and cpu and not omni, and their existing three-way Model Executor sharding.--deselectnode prefixes:test_temporal_streaming_step_matches_legacy_end_to_endtest_mimi_streaming_step_matches_legacy_end_to_endmi300_1, and isNonBlocking.timeoutof 110 minutes, with TERM plus a two-minute kill-after buffer for orderly artifact finalization and teardown.Scope
This changes only three AMD pipeline YAML files and the AMD pipeline contract test. It does not change CUDA YAML, production code, test bodies, test markers, assertions, templates, retries, or model behavior.
Exact revision
f3391da7c58ee8dc990829aab82166edf865d19b51a0d262fd44b63dd178121eed9920582e9a5e9fhaic0:ci/amd-personaplex-temporal-nightly-20261008Validation
git diff --check: passed.slowmarkers;Exact-head CI qualification
The 1h25m28s nightly runtime confirms that a 90-minute outer timeout would not leave safe setup/artifact/teardown headroom; the new 110-minute inner and 120-minute outer limits are warranted.
Model evaluation: N/A; this is CI-only routing with no model, output, accuracy, or serving change.