Skip to content

ci(recipes): benchmark DeepSeek V4 Pro disaggregated nightly - #15393

Merged
lavanyavijayk merged 16 commits into
mainfrom
ci/deepseek-v4-pro-disagg-nightly
Oct 1, 2026
Merged

lavanyavijayk merged 16 commits into
mainfrom
ci/deepseek-v4-pro-disagg-nightly

Conversation

@lavanyavijayk

@lavanyavijayk lavanyavijayk commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add a DeepSeek-V4-Pro vLLM disaggregated GB200 nightly benchmark, following the aggregated setup in PR #15325. The overlay uses the existing 1P/1D deployment and its 8K input / 1K output perf recipe at concurrency 512.

Where should the reviewer start?

  • .github/ci/deepseek-v4-pro-disagg/: Review the cache, scheduling, and perf patches.
  • .github/workflows/deepseek-v4-pro-disagg-nightly.yml: Review image pinning, deployment, benchmark verification, artifacts, and cleanup.
  • .github/workflows/nightly-ci.yml: Review the nightly job wiring.

Validation

  • Kustomize rendered successfully; the DGD has Frontend, prefill, and decode services using shared-model-cache.
  • The target cluster accepted the Job, DGD, and ComputeDomain in a server-side dry run. It warned that the existing nvidia.com/v1alpha1 DGD API is deprecated.
  • Workflow YAML parsed, all eight shell steps passed bash -n, and git diff --check passed.
  • A live deployment and benchmark have not run. GitHub requires this new manual workflow to exist on the default branch before workflow_dispatch can run it from a branch.

Related Issues

🚫 This PR is NOT linked to an issue:

  • Confirmed — no related issue

Summary by CodeRabbit

  • New Features
    • Added nightly performance validation for the DeepSeek V4 Pro disaggregated deployment, including deployment readiness checks and benchmark results.
    • Added a manually triggered preflight option to validate configuration without deploying.
    • Performance runs now capture diagnostic and benchmark artifacts, and report whether request-volume, error-rate, and output-length checks passed.

@github-actions github-actions Bot added ci Issues/PRs that reference CI build/test actions labels Sep 29, 2026
@lavanyavijayk
lavanyavijayk force-pushed the ci/deepseek-v4-pro-disagg-nightly branch from 3f42ade to ee2a278 Compare September 30, 2026 14:00
@lavanyavijayk
lavanyavijayk marked this pull request as ready for review September 30, 2026 16:05
@lavanyavijayk
lavanyavijayk requested a review from a team as a code owner September 30, 2026 16:05
@coderabbitai

coderabbitai Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: ai-dynamo/dynamo/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: ecd3ffa9-0d55-433b-b584-e2b9c65a8fb4

📥 Commits

Reviewing files that changed from the base of the PR and between 0c0ce6d and 39dde44.

📒 Files selected for processing (5)
  • .github/ci/deepseek-v4-pro-disagg/kustomization.yaml
  • .github/ci/deepseek-v4-pro-disagg/patch-deploy.yaml
  • .github/ci/deepseek-v4-pro-disagg/patch-perf.yaml
  • .github/workflows/deepseek-v4-pro-disagg-nightly.yml
  • .github/workflows/nightly-ci.yml

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


Walkthrough

Adds Kustomize resources and a reusable workflow for DeepSeek V4 Pro disaggregated deployment and performance testing. Nightly CI invokes the workflow, which can validate manifests in preflight mode or deploy, benchmark, check results, collect artifacts, and clean up.

Changes

Disaggregated nightly performance

Layer / File(s) Summary
Deployment and benchmark resources
.github/ci/deepseek-v4-pro-disagg/*
Adds a Kustomize overlay with deployment and benchmark resources. The deployment configures Frontend, prefill, and decode with shared model-cache access and scheduling settings. The benchmark Job sets its scheduling constraints and AIPerf parameters.
Workflow setup and preflight
.github/workflows/deepseek-v4-pro-disagg-nightly.yml
Adds reusable and manual workflow inputs. The workflow resolves an ARM64 image digest, renders manifests, validates the PVC and manifests, and prepares per-run credentials.
Deployment and benchmark execution
.github/workflows/deepseek-v4-pro-disagg-nightly.yml
Outside preflight mode, creates the deployment, waits for readiness, runs the benchmark Job, and polls for completion.
Results, artifacts, and cleanup
.github/workflows/deepseek-v4-pro-disagg-nightly.yml
Collects resource, log, event, and performance artifacts. Checks AIPerf results, deletes per-run resources and secrets, and uploads artifacts.
Nightly CI integration
.github/workflows/nightly-ci.yml
Adds a test-gated job that calls the reusable workflow. The Slack notification job now waits for it.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Merge Risk: ⚪ Minimal · up to 39dde

The new nightly benchmark is mergeable after normal checks. Configuration and cleanup contracts are consistent; live deployment and benchmark performance remain unvalidated.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the main change: adding a nightly benchmark for the DeepSeek V4 Pro disaggregated deployment.
Description check ✅ Passed The description explains the benchmark, identifies key files for review, documents validation results and limitations, and includes the required no-related-issue confirmation. It does not use the temp…
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@dmitry-tokarev-nv dmitry-tokarev-nv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I approve at 39dde44c2e. No P0 or P1 is open. Two P2 findings are on inline threads, and two P3 findings are below. main requires resolved threads, so the two threads hold the merge until you fix or resolve them.

  • [P2] deepseek-v4-pro-disagg-nightly.yml:49: in one nightly, this job and the aggregated job get the same RUN_KEY. The thread has the measurement and a suggestion.
  • [P2] deepseek-v4-pro-disagg-nightly.yml:183: the Verify step can read the warmup export and fail a good run. The thread has the measurement and a suggestion.
  • [P3] deepseek-v4-pro-disagg/kustomization.yaml:10 pins the Job image to docker.io/library/python@sha256:44ff437b…, a linux/amd64-only image. The aggregated Job uses python:3.12-slim, because #15325 removed the same pin in a06a5e7c4a. Please remove it here too, or add a comment that the amd64 nodeSelector in patch-perf.yaml must stay.
  • [P3] The Validation section of the description is out of date. It says "A live deployment and benchmark have not run." But dispatch run 36646643297 deployed this overlay at a405ac850c and passed the benchmark. Run 36641543850 failed at ISL 8192. Please update that section, because the passing run is the best evidence for this PR.
What I measured, offline, with no contact to a cluster.
  • I ran the Render step of a405ac850c offline with kubectl 1.36.3 (kustomize v5.8.1). I replaced only the image host and digest. The result is byte-identical to the rendered.yaml, deploy.yaml, and perf.yaml artifacts of run 36646643297. With the same run key, the head renders the same deploy.yaml. Its Job differs only in backoffLimit: 0, activeDeadlineSeconds: 7200, and the cleanup trap. A merge with the current main renders the same as the head.
  • In run 36646643297, 1,536 requests at ISL 8064 passed with 0 errors in 191.66 s. The Job took about 6 minutes, and the head gives it a 7,200 s deadline. At ISL 8192, run 36641543850 got 1,536 errors.
  • The fixes from #15325 are all here. I ran the Job as root and the Collect step as the frontend user (uid 1000) on a Docker volume. Each control fails without its fix. The Job leaves no inputs.json on the volume, and the artifact has none. If that delete fails, the file stays on the volume, but the artifact still has none. A failed benchmark keeps its exit code (3 and 5) through the trap. After a lost run, the next preflight deletes the two Secrets and the three fixed names of that run.
  • The job uses the runner label, needs, condition, and concurrency group of the aggregated job, and both actions are pinned by SHA. Apart from RUN_KEY, the object names, the Secret label, and the artifact name differ from the aggregated job.
  • At earlier commits, the server-side dry runs of preflight-only runs 36639830031 and 36641433494 accepted the DGD, the ComputeDomain, the Job, and both Secrets. A dry run does not test GPU capacity, scheduling, image pulls, readiness, or the deletes, which run only in a full run.

Comment thread .github/workflows/deepseek-v4-pro-disagg-nightly.yml Outdated
Comment thread .github/workflows/deepseek-v4-pro-disagg-nightly.yml Outdated
Comment thread .github/ci/deepseek-v4-pro-disagg/kustomization.yaml Outdated
Comment thread .github/ci/deepseek-v4-pro-disagg/patch-deploy.yaml Outdated
Comment thread .github/workflows/nightly-ci.yml Outdated

@dmitry-tokarev-nv dmitry-tokarev-nv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I approve at dcf127f089. No P0, P1, or P2 is open, and one P3 from my last review is still open. You applied both of my suggestions, so part of this review covers my own code.

  • [P2] Fixed: the shared RUN_KEY. I replied on its thread with the measurement and resolved it.
  • [P2] Fixed: the Verify step that read the warmup export. I replied on its thread with the measurement and resolved it.
  • [P3] Fixed: the python@sha256 pin is gone. The Job now uses python:3.12-slim from the recipe, the same image as the aggregated Job.
  • [P3] Open: the Validation section of the description still says "A live deployment and benchmark have not run." Run 36646643297 deployed this overlay at a405ac850c and passed.
What I measured at dcf127f, offline, with no contact to a cluster.
  • I rendered the overlay at 39dde44c2e and at dcf127f089 with the same run key. The DGD and the ComputeDomain are byte-identical, so the YAML anchors in patch-deploy.yaml change nothing in the output. The only change is the Job image, from docker.io/library/python@sha256:44ff437b… to python:3.12-slim. The rendered files hold no anchors or aliases.
  • The new name: lines change only the names that GitHub shows. In nightly run 36688149628, the aggregated job showed as DeepSeek V4 Pro aggregated perf / perf. The step in notify-slack.yml turns that name into :failed: perf, so with the old names both jobs print :failed: perf. With the new names, the same step prints :failed: DeepSeek V4 Pro aggregated perf and :failed: DeepSeek V4 Pro disaggregated perf. I found no other reader of these names in .github/. Apart from the job name, the aggregated workflow did not change.
  • sara4dev asked for the upstream image, YAML anchors, and one recipes group. The pin is gone, the anchors render the same objects, and both caller jobs are now named recipes. I did not look at the result in the GitHub UI. Their three threads are still open.
  • A merge with main at 1a26bf9cb7 conflicts only in the needs list of notify-slack, where main added sidecar-trtllm-test. With both entries kept, the render does not change, and actionlint 1.7.12 reports no new findings.

Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
@lavanyavijayk
lavanyavijayk force-pushed the ci/deepseek-v4-pro-disagg-nightly branch from dcf127f to 58c36db Compare October 1, 2026 02:04
@devin-ai-integration

devin-ai-integration Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

⏹️ Dynamo PR CI cancelled — run 36882120463 (attempt 1) on fde2a8a7ce

✅ 55 passed · ❌ 0 failed · ⏹️ 1 cancelled · ⏭️ 27 skipped (not needed for this change)

Gate checks: ❌ backend-status-check · ✅ deploy-status-check · ✅ dynamo-status-check

Each cell counts that group's GitHub Actions jobs by result (✅ passed, ❌ failed, ⏰ timed out, ⏹️ cancelled); ⏭️ = all skipped, ➖ = no such job.

Framework Build 1-GPU amd64 1-GPU arm64 Multi-GPU amd64 Deploy Snapshot
vLLM ✅ 3 ✅ 1 ✅ 1 ⏹️ 1 ✅ 4 ⏭️
SGLang ✅ 3 ✅ 1 ✅ 1 ✅ 1 ✅ 2 ⏭️
TRT-LLM ✅ 3 ✅ 1 ✅ 1 ✅ 1 ✅ 2 ⏭️
Other Jobs
changed-files ✅ 1
deploy-operator ✅ 1
DGDR Deploy Test ✅ 4
dynamo-runtime ✅ 8
frontend ✅ 2
frontend (amd64) ✅ 1
frontend (arm64) ✅ 1
Helm Chart Tests ✅ 1
Operator ✅ 1
Operator Integration ✅ 1
planner ✅ 6
Power Agent ✅ 1
triton-runtime ✅ 2
⏹️ 1 cancelled jobs
⏭️ 3 other components not run (skipped by change detection or an upstream result)

allure-report, dynamo-sidecar, sidecar-runtime

Failure details

No jobs failed. The 1 cancelled job (vllm-runtime / 2-GPU Test cuda13.0, amd64) was not a fail-fast cancel: its self-hosted runner received a shutdown signal mid-test (infrastructure), which also turned backend-status-check red; rerunning failed jobs is likely enough.

Job Summary
⏹️ vllm-runtime / 2-GPU Test cuda13.0, amd64 Runner shutdown signal during Run GPU tests (sequential); infrastructure, rerun failed jobs.
⏹️ vllm-runtime / 2-GPU Test cuda13.0, amd64: runner received a shutdown signal

Failed step: Run GPU tests (sequential) · Logs: gh api repos/ai-dynamo/dynamo/actions/jobs/110448255052/logs

tests/test_predownload_models.py::test_predownload_models[predownload_models_vllm_gpu2] PASSED [ 10%]
tests/serve/test_vllm.py::test_serve_deployment[agg-router-3]
##[error]Process completed with exit code 130.
##[error]The runner has received a shutdown signal. This can happen when the runner service is stopped, or a manually started runner is canceled.
##[error]Executing the custom container implementation failed. Please contact your self hosted runner administrator.

The self-hosted runner was shut down about a minute into the sequential GPU tests, while test_serve_deployment[agg-router-3] was running; no test failure was reported. This is an infrastructure signature (the rest of the run kept going for ~30 more minutes, so it was not a fail-fast cancel), and rerunning failed jobs is likely enough.

For agents
{"pr": 15393, "run_id": 36882120463, "run_attempt": 1, "head_sha": "fde2a8a7ce008c4ba749b6e83a4d2d53543e1c63", "failures": [], "cancelled": [{"job": "vllm-runtime / 2-GPU Test cuda13.0, amd64", "job_id": 110448255052, "failed_step": "Run GPU tests (sequential)", "signature": "The runner has received a shutdown signal (exit code 130)", "tests": [], "log_cmd": "gh api repos/ai-dynamo/dynamo/actions/jobs/110448255052/logs"}]}
Previous runs
  • ✅ run 36804087340 (attempt 2) on 58c36dbf1a: Passed (56 passed, 0 failed, 0 cancelled).
  • ❌ run 36804087340 (attempt 1) on 58c36dbf1a: 1 job failed: 1 SGLang tool-calling test (test_chained_tool_use_search_then_calculate[rust_parsers]) fails because the first step ends with finish_reason='length' instead of tool_calls, after 3 automatic retries. No jobs were cancelled.

Posted automatically by Devin for run 36882120463. Updated on every full-CI run of this PR.

@dmitry-tokarev-nv dmitry-tokarev-nv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I approve at 58c36dbf1a. No P0 or P1 is open. One P2 is on an inline thread, and two P3 findings are below. main requires resolved threads, so the P2 thread holds the merge until you fix or resolve it.

  • [P2] deepseek-v4-pro-disagg-nightly.yml:34: a dispatch of either DeepSeek workflow cancels the nightly DeepSeek job that waits for the other one. The thread has the evidence.
  • [P3] nightly-ci.yml:541 and nightly-ci.yml:551: when either DeepSeek job fails, Slack prints :failed: recipes > perf, so the alert does not name the recipe. Please make the first or the last part of the two job names differ. This corrects my last review.
  • [P3] Still open from my last review: the Validation section of the description says "A live deployment and benchmark have not run." Run 36646643297 deployed this overlay and passed.
The Slack names, measured with the real notifier step.
  • GitHub joins the names of nested jobs with " / ". In nightly 36688149628, one such job showed as dynamo-runtime / image / Build multi-arch cuda13.0. No nightly ran at this head yet. By the same pattern, the two DeepSeek jobs get recipes / DeepSeek V4 Pro aggregated perf / perf and recipes / DeepSeek V4 Pro disaggregated perf / perf.
  • For a name with more than two parts, the "Get Failed jobs" step of notify-slack.yml keeps only the first and the last part. I ran that step offline on the real job list of nightly 36688149628, with a stub for curl, and changed only the DeepSeek names:
Names Slack lines
This head, both jobs fail :failed: recipes > perf, two times
main, the aggregated job fails :failed: DeepSeek V4 Pro aggregated perf > perf
dcf127f089, two parts, both jobs fail :failed: DeepSeek V4 Pro aggregated perf and :failed: DeepSeek V4 Pro disaggregated perf
name: ${{ inputs.recipe }} on the perf job of shared-recipe-perf.yml, both jobs fail :failed: recipes > deepseek-v4-pro-agg and :failed: recipes > deepseek-v4-pro-disagg
  • My last review said that the new names print two different lines. That was true at dcf127f089, where each name had two parts. It is not true at this head.
  • The run page shows the full name of each job, so only the Slack line loses the recipe.
  • The GitHub docs allow the inputs context in jobs.<job_id>.name. That example also changes the Kimi K2.5 line to :failed: Kimi K2.5 aggregated functional smoke > kimi-k25-agg. I did not run it on GitHub.
What I measured at 58c36db, offline, with no contact to a cluster.
  • The rebase onto 7b1da11015 kept both sides of each conflict. The needs list of notify-slack has the entries from main and deepseek-v4-pro-disagg-perf. deepseek-v4-pro-nightly.yml is the file from main plus your name: line.
  • I ran the real Configure and Render steps for kimi-k25-agg and deepseek-v4-pro-agg at 7b1da11015 and at this head. rendered.yaml, deploy.yaml, perf.yaml, and prerequisites.yaml are byte-identical. In GITHUB_ENV, only WORKER= became WORKERS=, and no other step reads WORKER. Each of 4 one-line mutants of the workflow changed the render.
  • With a stub kubectl, both callers gave the same results at both commits. I compared the step outputs, the kubectl calls, the deletes, the leftovers, and the artifacts. The cases were a full run, a lost run and its recovery, a preflight-only run, and a Job with no output.
  • The disaggregated DGD has exactly Frontend, prefill, and decode. Both workers get hf-ci-deepseek-v4-pro-disagg, and all three services get acr-ci-deepseek-v4-pro-disagg and the image by digest. No hf-token-secret or acr-token-secret remains. A mutant with workers=prefill,decod added a fourth service, so this check can see an extra service.
  • In the disaggregated Job, command is [/bin/sh, -c, <script>], and the cleanup trap goes in front of the script, as in the aggregated Job. Apart from the labels, the Secret names, and ARTIFACT_ROOT, the objects match the render of the old workflow at b841fc246d.
  • Configure has its own deepseek-v4-pro-disagg arm, and Render names the recipe in its elif. Verify AIPerf results takes the branch for the DeepSeek recipes and expects 3 × 512 requests, the same count that the Job sends. Deploy, Run AIPerf, Collect, and Clean up do not branch on RECIPE.
  • Verify AIPerf results passed on the real output of run 36646643297. When I copied the warmup export over the top-level file, the step failed. When I removed only the top-level file, it also failed. It failed on 1,535 requests, one error, no output, CONCURRENCIES set to "256 512", and the wrong container name.
  • Our two earlier P2 findings stay fixed. In one nightly, the keys are ci-deepseek-v4-pro-agg-111-1 and ci-deepseek-v4-pro-disagg-111-1. With both jobs live, each Clean up deleted only its own DGD, Job, ComputeDomain, and two Secrets, in both orders. Each artifact held only its own concurrencies.
  • The control mutant with one shared key deleted the objects of the other job and mixed the artifacts. After a lost run of either recipe, the other recipe passed. The next nightly of the lost recipe then removed only its own leftovers.
  • The two DeepSeek recipes and Kimi K2.5 share no fixed name. I compared the DGD, Job, ComputeDomain, claim template, Secrets, ConfigMap, frontend endpoint, and artifact root.
  • Kimi K2.5 and the aggregated recipe each ask for 8 GPUs on 2 GB200 nodes. The disaggregated recipe asks for 16 GPUs on 4 nodes. All three use the same two node pools. I did not query the cluster, so I cannot say whether all three fit at once.

Comment thread .github/workflows/deepseek-v4-pro-disagg-nightly.yml
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>
Signed-off-by: Lavanya <lvijayakrish@nvidia.com>

@dmitry-tokarev-nv dmitry-tokarev-nv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I approve at 075a9f0084. No P0 or P1 is open. One new P2 is inline, and two earlier P3 findings are below. main requires resolved threads, so the P2 thread holds the merge until you fix or resolve it.

  • [P2] shared-recipe-perf.yml:233: the new --previous loop saves no log for either DeepSeek recipe. The thread has the evidence and a suggestion.
  • [P3] Still open from my last review: nightly-ci.yml:541 and nightly-ci.yml:551 name both caller jobs recipes. When a DeepSeek job fails, Slack prints :failed: recipes > perf and does not name the recipe.
  • [P3] Still open: the Validation section of the description says "A live deployment and benchmark have not run." Run 36646643297 deployed this overlay and passed.
  • Fixed: my P2 on the shared concurrency group. 17ce31b9f1 adds queue: max to both DeepSeek workflows. You resolved that thread, and I replied there with the evidence.
What I measured at 075a9f0, offline, with no contact to a cluster.
  • Startup probe. Kustomize merges failureThreshold: 900 into the recipe probe, so each worker keeps httpGet /health:9090, period 10 s, and timeout 10 s. The operator replaces its default probe with this whole probe on the leader pod (graph.go:1844). It removes all probes from the second pod of each worker (backend_vllm.go:77). I ran GenerateBasePodSpec on the rendered DGD at this head and at main 5d6d44beef. The leader probes allow 900 × 10 s = 150 min, which is less than the 180-minute Deploy step. The deadline of the bench Job (7,200 s) starts after Deploy, and the job limit is 960 min. A control with only failureThreshold reached the pod spec with no handler, so the full recipe probe keeps the result valid.
  • vLLM. VLLM_ENGINE_READY_TIMEOUT_S stays 5400. In vLLM v0.30.0, that wait starts only after every engine finishes loading (core_client.py:709, utils.py:1240). It does not stop a long first load.
  • Deploy timeout. The GitHub contexts table allows inputs in jobs.<job_id>.steps.timeout-minutes, and the step maximum is 360. @actions/expressions 0.3.61 returns 105 for kimi-k25-agg and deepseek-v4-pro-agg, and 180 for deepseek-v4-pro-disagg. actionlint 1.7.12 accepts the line.
  • Collect. I ran the step at 58c36dbf1a, at this head, and with the suggestion, for all three recipes. test and GNU tar ran in the vLLM runtime image as uid 1000. The #15361 failed-Deploy case exits 2 at 58c36dbf1a and 0 now. A finished benchmark copies the same files as before (2, 21, and 14 files, without inputs.json), and Verify passes. If the Job writes nothing, Verify still fails the run.
  • NIXL telemetry. From 17ce31b9f1 to this head, each worker in the rendered DGD loses only NIXL_TELEMETRY_ENABLE=1. After the operator merge, both workers get the operator default NIXL_TELEMETRY_ENABLE=n, as in run 36646643297.
  • Kimi and aggregated callers. rendered.yaml, deploy.yaml, perf.yaml, and prerequisites.yaml are byte-identical to 58c36dbf1a. A full run, a lost run and its recovery, and a preflight-only run give the same steps, deletes, leftovers, and artifacts. The only new kubectl calls are in Collect.
  • main at 3af83c2056 merges without a conflict and changes no file that these renders read. Its only change to a file of this PR is one comment in nightly-ci.yml.

Comment thread .github/workflows/shared-recipe-perf.yml Outdated
lavanyavijayk and others added 3 commits October 1, 2026 08:09
@lavanyavijayk
lavanyavijayk enabled auto-merge (squash) October 1, 2026 16:52
@lavanyavijayk
lavanyavijayk force-pushed the ci/deepseek-v4-pro-disagg-nightly branch from 8ce574c to a625c72 Compare October 1, 2026 18:11
@lavanyavijayk
lavanyavijayk merged commit dc02e97 into main Oct 1, 2026
240 of 248 checks passed
@lavanyavijayk
lavanyavijayk deleted the ci/deepseek-v4-pro-disagg-nightly branch October 1, 2026 19:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

actions ci Issues/PRs that reference CI build/test size/L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants