Skip to content

[CI/Build][ROCm] Move PersonaPlex temporal stress to nightly - #8670

Merged
andyluo7 merged 1 commit into
vllm-project:mainfrom
haic0:ci/amd-personaplex-temporal-nightly-20261008
Oct 9, 2026
Merged

andyluo7 merged 1 commit into
vllm-project:mainfrom
haic0:ci/amd-personaplex-temporal-nightly-20261008

Conversation

@haic0

@haic0 haic0 commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Reduce the AMD mi300_1 Model Executor critical-path time by moving only the four long-running PersonaPlex temporal oracle cases from READY/MERGE into one serial, nonblocking nightly job.

Historical observations from AMD builds #13509 and #13511 put the four stress cases at roughly 57-84 minutes total. Hash sharding uses node identity/count rather than measured duration, so these loop-heavy cases disproportionately extend Model Executor shards 1/3 and 3/3.

Coverage and behavior

  • AMD READY and MERGE retain the original marker expression, core_model and cpu and not omni, and their existing three-way Model Executor sharding.
  • They add exactly two --deselect node prefixes:
    • test_temporal_streaming_step_matches_legacy_end_to_end
    • test_mimi_streaming_step_matches_legacy_end_to_end
  • Pytest's node-prefix selection means those prefixes move exactly four stress nodes: one temporal case and the three parametrized Mimi cases.
  • The remaining ten fast RoPE/RingKV nodes in the same file stay in the blocking READY/MERGE shards.
  • AMD nightly explicitly selects the same two prefixes, runs all four stress nodes serially on mi300_1, and is NonBlocking.
  • The nightly step has a 120-minute Buildkite timeout and an inner timeout of 110 minutes, with TERM plus a two-minute kill-after buffer for orderly artifact finalization and teardown.
  • Collection output, JUnit XML, the full pytest log, and a tail summary are uploaded.
  • No retries, xdist, empty-collection override, or failure-masking fallback are added.
  • CUDA READY/MERGE remain unfiltered: no ignore, deselection, or slow marker excludes any of the file's 14 nodes.
  • The Omni Processor smoke/nightly routing merged in #8571 is preserved.

Scope

This changes only three AMD pipeline YAML files and the AMD pipeline contract test. It does not change CUDA YAML, production code, test bodies, test markers, assertions, templates, retries, or model behavior.

Exact revision

  • Base: f3391da7c58ee8dc990829aab82166edf865d19b
  • Head: 51a0d262fd44b63dd178121eed9920582e9a5e9f
  • Branch: haic0:ci/amd-personaplex-temporal-nightly-20261008
  • Diff: 4 files, 197 insertions, 3 deletions

Validation

  • Focused AMD pipeline contract: 22 passed.
  • Broader AMD/Buildkite contract suite: 71 passed.
  • Changed-file pre-commit: all applicable hooks passed, including YAML, Ruff, Ruff format, typos, mypy 3.10, test marks, SPDX, forbidden imports, and Buildkite schema validation.
  • git diff --check: passed.
  • Contract coverage verifies:
    • READY/MERGE use exactly the two intended deselection prefixes while retaining the original marker expression and three shards;
    • nightly selects the same prefixes, has the 120/110-minute timeout structure, remains nonblocking, and publishes all expected artifacts;
    • CUDA has no ignore/deselect filter and the PersonaPlex source has no slow markers;
    • pipe failures and empty collections remain fatal through the existing AMD runner/template behavior.

Exact-head CI qualification

  • AMD build #13573: aggregate passed.
    • Model Executor shard 1/3: 1344 passed, 3 skipped in 9m15s pytest time (10m16s job elapsed).
    • Model Executor shard 2/3: 1380 passed, 1 skipped in 7m25s pytest time (8m21s job elapsed).
    • Model Executor shard 3/3: 1414 passed in 6m43s pytest time (7m32s job elapsed).
    • Live collection/execution retained the ten fast PersonaPlex nodes across shards 2/3 and 3/3 and excluded the four stress nodes from all three blocking shards.
    • PersonaPlex nightly: 4 passed in 5128.25s (1h25m28s pytest time; 1h26m32s job elapsed).
    • Nightly uploaded all four requested artifacts successfully: collection output, JUnit XML, full log, and summary.
    • The only nonzero AMD job was the pre-existing nonblocking R2-01 GPU coverage lane: eight unrelated MOSS-TTS/Breeze failures, with 1095 passed and 368 skipped. It did not exercise the changed PersonaPlex routing.
  • CUDA build #16986: all 19 jobs passed, including the unfiltered Model Executor coverage.
  • GitHub pre-commit, Python 3.11/3.12 wheel builds, DCO, and Read the Docs: passed.

The 1h25m28s nightly runtime confirms that a 90-minute outer timeout would not leave safe setup/artifact/teardown headroom; the new 110-minute inner and 120-minute outer limits are warranted.

Model evaluation: N/A; this is CI-only routing with no model, output, accuracy, or serving change.

@haic0 haic0 left a comment •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Exact-head self-review at 2cd136f311c02f26d7680cc57a72eac7bf8019b2

Reviewed the complete four-file diff and the READY/MERGE/NIGHTLY MiniJinja semantic renders.

  • READY/MERGE add only the exact heavy-file ignore; all other Model Executor and PersonaPlex selection remains intact.
  • Nightly has one serial, nonblocking 90-minute owner with nonzero collection evidence, durations, JUnit, complete log, and summary artifacts.
  • The existing runner preserves piped pytest failures and exit 5; no allow-no-tests override, retry, xdist, template edit, or CUDA YAML change is present.
  • Contract tests cover source routing, rendered soft-fail/dependency behavior, and unchanged CUDA ownership.

@andyluo7
andyluo7 force-pushed the ci/amd-personaplex-temporal-nightly-20261008 branch from 2cd136f to 64a6bd5 Compare October 9, 2026 15:47
Assisted-by: GPT-5.6 Sol
Signed-off-by: haic0 <149741444+haic0@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7
andyluo7 force-pushed the ci/amd-personaplex-temporal-nightly-20261008 branch from 64a6bd5 to 51a0d26 Compare October 9, 2026 15:53
@andyluo7 andyluo7 added ready label to trigger buildkite CI ROCm PR related to AMD hardware nightly-test label to trigger buildkite nightly test CI CI/CD codes related to changes to CI/CD amd-test Used to trigger AMD CI separately. cuda-test Used to trigger vllm-omni cuda CI separately. labels Oct 9, 2026
@andyluo7
andyluo7 marked this pull request as ready for review October 9, 2026 15:54
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR was classified as CI work.

CI owner: @yenuo26 @congw729 @NickCao

Routing: @yenuo26 via semantic router, CI owner, CODEOWNERS; @congw729 via CODEOWNERS; @NickCao via CODEOWNERS

@haic0, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@andyluo7 andyluo7 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the exact four-file CI routing change at 51a0d26. READY/MERGE retain the ten fast PersonaPlex nodes and deselect only the four long stress cases; the dedicated nightly lane collected and passed those four cases in 1:25:28 with artifacts uploaded. AMD build 13573, CUDA build 16986, GitHub Actions, DCO, and docs are green. The unrelated R2-01 GPU coverage failures were soft-failed and outside this patch.

@andyluo7
andyluo7 merged commit b65107d into vllm-project:main Oct 9, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd-test Used to trigger AMD CI separately. CI/CD codes related to changes to CI/CD cuda-test Used to trigger vllm-omni cuda CI separately. nightly-test label to trigger buildkite nightly test CI ready label to trigger buildkite CI ROCm PR related to AMD hardware

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants