Skip to content

ci(recipes): use one load point for aggregated nightly tests - #15571

Merged
sara4dev merged 4 commits into
mainfrom
codex/cap-deepseek-agg-perf-concurrency
Oct 6, 2026
Merged

sara4dev merged 4 commits into
mainfrom
codex/cap-deepseek-agg-perf-concurrency

Conversation

@sara4dev

@sara4dev sara4dev commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Use one configured load point for each aggregated nightly benchmark: DeepSeek V4 Pro at concurrency 128 and Kimi K2.5 at concurrency 20. DeepSeek removes its multi-point sweep after the October 2 run crashed at concurrency 256; its validator requires 384 measured requests, zero errors and positive output.

Kimi preserves its existing trace of 20 conversations with 10 sequential turns each. Cap concurrency at 20 and require all 200 requests, zero errors, positive output and no cancellation. Select exactly one matching export using the rendered concurrency. Its 120-second duration is a maximum; trace timing and the existing ramp can yield lower achieved concurrency or earlier completion.

Validation

  • Rendered all three CI overlays with Kustomize. Before/after comparison confirms only Kimi's concurrency value changed in its rendered resources; the trace, benchmark script, duration and deployment are preserved. DeepSeek remains at 128 aggregated / 512 disaggregated.
  • Executed the extracted result validator against 19 fixtures covering complete/partial counts, errors, empty output, cancellation, missing/duplicate/wrong-concurrency exports, malformed settings, warmup exclusion and unchanged DeepSeek acceptance. The yq extraction was stubbed.
  • Workflow shell syntax, applicable pre-commit hooks and git diff --check passed. No GPU benchmark was run.

@sara4dev
sara4dev requested a review from a team as a code owner October 2, 2026 16:29
@copy-pr-bot

copy-pr-bot Bot commented Oct 2, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the ci Issues/PRs that reference CI build/test label Oct 2, 2026
@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: ai-dynamo/dynamo/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 75f4750a-bd48-4b99-93e8-898544f78530

📥 Commits

Reviewing files that changed from the base of the PR and between bf923d9 and a49e465.

📒 Files selected for processing (1)
  • .github/ci/deepseek-v4-pro-agg/patch-perf.yaml

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.


Walkthrough

The patch performance benchmark now uses concurrency levels 1, 8, 64, and 128. It no longer uses levels 256, 512, or 1024.

Changes

Benchmark concurrency settings

Layer / File(s) Summary
Update benchmark concurrency levels
.github/ci/deepseek-v4-pro-agg/patch-perf.yaml
The concurrency list retains 1, 8, 64, and 128, and removes 256, 512, and 1024.

Priority: ⬇️ Low

Estimated code review effort: 1 (Trivial) | ~3 minutes

Merge Risk: ⚪ Minimal · up to a49e4

The benchmark retains the intended concurrency levels and is ready to merge.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description does not match the changeset: it describes a single DeepSeek load point and Kimi changes, while the stated objective retains four DeepSeek concurrency levels and does not mention a Kim… Revise the description to state that DeepSeek retains concurrency levels 1, 8, 64, and 128 and removes 256, 512, and 1024. Remove unsupported claims about a single DeepSeek load point and Kimi changes. Add the required Related Issues sectio…
✅ Passed checks (4 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title relates to the aggregated nightly benchmark configuration, but “one load point” is inaccurate because the change retains four concurrency levels: 1, 8, 64, and 128.
Full details: Description check

Explanation

The description does not match the changeset: it describes a single DeepSeek load point and Kimi changes, while the stated objective retains four DeepSeek concurrency levels and does not mention a Kimi change. It also omits the required Related Issues section.

Resolution

Revise the description to state that DeepSeek retains concurrency levels 1, 8, 64, and 128 and removes 256, 512, and 1024. Remove unsupported claims about a single DeepSeek load point and Kimi changes. Add the required Related Issues section and select either the linked-issue option or the no-related-issue confirmation.

  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@sara4dev sara4dev changed the title ci(recipes): cap DeepSeek aggregated perf concurrency at 128 ci(recipes): use one DeepSeek aggregated perf concurrency Oct 2, 2026
@sara4dev sara4dev changed the title ci(recipes): use one DeepSeek aggregated perf concurrency ci(recipes): use one load point for aggregated nightly tests Oct 2, 2026
@pull-request-size pull-request-size Bot added size/S and removed size/XS labels Oct 2, 2026
@sara4dev
sara4dev enabled auto-merge (squash) October 2, 2026 21:31
Comment thread .github/ci/deepseek-v4-pro-agg/patch-perf.yaml

@dmitry-tokarev-nv dmitry-tokarev-nv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved at db33213f835df2c5bd4577b2214840aa679dd850. Two P2 findings are open in inline threads. There is no P3.

  • [P2] .github/ci/deepseek-v4-pro-agg/patch-perf.yaml:27: the removed load points are where an untracked vLLM worker crash happens. The thread asks for an issue.
  • [P2] .github/ci/kimi-k25-agg/patch-perf.yaml:24: the Kimi trace allows at most 20 requests at once, so the job never reaches 128.
  • Open item for #15537: it changes the same "Verify AIPerf results" step, and the two heads conflict there. To resolve it, keep the concurrencies= line and the regex from #15571, and add || "$RECIPE" == kimi-k25-disagg from #15537 to the condition. With only the #15537 side, no line sets concurrencies, and every DeepSeek result passes. With only the #15571 side, every Kimi disaggregated result fails.
What I measured.
  • The merge commit adds no change of its own. A new merge of c4c7342054 with main at d9ea8c475b gives the committed tree 343283eed5.
  • At the head, kustomize renders Kimi with CONCURRENCIES=128 and BENCHMARK_DURATION=120. It renders DeepSeek aggregated with CONCURRENCIES=128, and DeepSeek disaggregated keeps 512.
  • I extracted the "Verify AIPerf results" step at the base and at the head. I ran each copy against 34 fixtures with bash --noprofile --norc -e -o pipefail, the shell that the job log shows. The fixtures use the real Oct 2 exports, the real perf.yaml files, and the real yq.
Fixtures Base Head
Kimi, valid c128 profile (3 fixtures, one with 1 of 200 requests) fail pass
Kimi, CONCURRENCIES=128 with only a c1 profile pass fail
Kimi, CONCURRENCIES is 1,128 or 1 128 pass fail
Kimi, CONCURRENCIES missing or empty, with a c1 profile pass fail
Kimi, valid c1 profile pass pass
Kimi, 14 bad profiles: missing or duplicate profile, wrong concurrency, errors, cancelled, zero or missing request count, zero output, empty file fail fail
Kimi, CONCURRENCIES missing or *, with a c128 profile fail fail
DeepSeek, 4 valid cases, among them c128 with 384 requests pass pass
DeepSeek, the real Oct 2 sweep, wrong count, errors, missing profile, missing CONCURRENCIES fail fail

Each failure at the head stops at the expected command. Two mutants of the head step fail the fixtures. With == 1 changed to >= 1, the two-profile fixtures pass. With the old _trace_c1_ glob, the valid c128 fixtures fail.

  • actionlint 1.7.12 with shellcheck 0.11.0 gives the same single error at the base and at the head: the runner label prod-deploy-tester-v1 is not in actionlint.yaml.
  • CI on the head: every source-branch job that ran passed, and the rust-tests, rust-clippy, Recipe Check, and docs jobs were skipped. pull-request/15571 does not exist, so the PR and deploy lanes did not run. Only the nightly or a manual dispatch runs the changed step.

Comment thread .github/ci/deepseek-v4-pro-agg/patch-perf.yaml
Comment thread .github/ci/kimi-k25-agg/patch-perf.yaml Outdated
@pull-request-size pull-request-size Bot added size/L and removed size/S labels Oct 3, 2026
@sara4dev sara4dev changed the title ci(recipes): use one load point for aggregated nightly tests ci(recipes): standardize synthetic nightly benchmarks Oct 3, 2026
@sara4dev
sara4dev force-pushed the codex/cap-deepseek-agg-perf-concurrency branch from 1c31c39 to db33213 Compare October 3, 2026 18:31
@pull-request-size pull-request-size Bot added size/S and removed size/L labels Oct 3, 2026
@sara4dev sara4dev changed the title ci(recipes): standardize synthetic nightly benchmarks ci(recipes): use one load point for aggregated nightly tests Oct 3, 2026

@dynamo-review-agent dynamo-review-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Previously reported defects still present:

  • Original discussion: The DeepSeek aggregated nightly still drops the c256 and higher load points by setting CONCURRENCIES to only 128, and the current file has no linked issue or adjacent comment documenting the vLLM worker crash that the removed c256 point exposed.

@sara4dev
sara4dev force-pushed the codex/cap-deepseek-agg-perf-concurrency branch from 2e01cdd to 5970c35 Compare October 5, 2026 01:37

@dynamo-review-agent dynamo-review-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Previously reported defects still present:

  • Original discussion: The DeepSeek aggregated nightly still sets CONCURRENCIES to only 128, removing the c256 and higher load points where the previously reported vLLM worker crash was observed. The current patch-perf.yaml has no adjacent comment or linked issue documenting that unresolved crash, so the earlier tracking concern remains present.
  • Original discussion: The DeepSeek aggregated nightly still sets CONCURRENCIES to only 128, and the file still has no adjacent comment or linked issue documenting the vLLM worker crash exposed by the removed c256-and-higher load points.
  • Original discussion: The DeepSeek aggregated nightly still sets CONCURRENCIES to only 128, dropping the c256 and higher load points where the vLLM worker crash was observed, and the current file still has no adjacent comment or linked issue documenting that unresolved crash.
  • Original discussion: The DeepSeek aggregated nightly still sets CONCURRENCIES to only 128, removing the c256 and higher load points where the previously reported vLLM worker crash was observed. I do not see an adjacent comment or linked issue documenting that unresolved crash, so the previously raised defect remains present.
  • Original discussion: The DeepSeek aggregated nightly still sets CONCURRENCIES to only 128, so the c256 and higher load points that exposed the vLLM worker crash remain removed, and the current file still has no adjacent comment or linked issue documenting that unresolved crash.

@devin-ai-integration

devin-ai-integration Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

❌ Dynamo PR CI failed — run 37370105482 (attempt 2) on 3ba5ced536

Gate checks: ✅ backend-status-check · ✅ deploy-status-check · ❌ dynamo-status-check

Other Jobs
dynamo-runtime ❌ 2 ✅ 6

Failure details

2 jobs failed: on both amd64 and arm64, the same 2 tests in tests/fault_tolerance/cancellation/test_utils.py fail because the streaming HTTP request gets no response within 30s (requests.exceptions.Timeout). The other 789 tests in each job passed.

❌ dynamo-runtime / test / parallel cuda13.0, amd64: 2 cancellation drained-stream tests time out (30s HTTP response)

Job: dynamo-runtime / test / parallel cuda13.0, amd64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112000906111 -R ai-dynamo/dynamo --log-failed

tests/fault_tolerance/cancellation/test_utils.py:50: in test_drained_stream_can_require_generated_content
    read_streaming_responses(
tests/fault_tolerance/cancellation/utils.py:309: in read_streaming_responses
    response_raw = cancellable_req.get_response()
tests/fault_tolerance/cancellation/utils.py:188: in get_response
    raise requests.exceptions.Timeout(
E   requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
FAILED tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_can_require_generated_content - requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
FAILED tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_accepts_generated_content - requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
=== 2 failed, 789 passed, 23 skipped, 5398 deselected in 1192.66s (0:19:52) ====

test_drained_stream_can_require_generated_content and test_drained_stream_accepts_generated_content both fail the same way, also after 3 auto-retries each: CancellableRequest.get_response() gets no HTTP response within 30s while read_streaming_responses drains the stream (drain=True, require_content=True). The later upload-artifact error (asterisk in the test_unsafe_or_nonexact_protobuf_pins[protobuf==6.33.*] log path) and the stage-verification error are follow-on noise, not the cause.

❌ dynamo-runtime / test / parallel cuda13.0, arm64: 2 cancellation drained-stream tests time out (30s HTTP response)

Job: dynamo-runtime / test / parallel cuda13.0, arm64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112000906131 -R ai-dynamo/dynamo --log-failed

tests/fault_tolerance/cancellation/test_utils.py:50: in test_drained_stream_can_require_generated_content
    read_streaming_responses(
tests/fault_tolerance/cancellation/utils.py:309: in read_streaming_responses
    response_raw = cancellable_req.get_response()
tests/fault_tolerance/cancellation/utils.py:188: in get_response
    raise requests.exceptions.Timeout(
E   requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
FAILED tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_can_require_generated_content - requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
FAILED tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_accepts_generated_content - requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response
=== 2 failed, 789 passed, 23 skipped, 5398 deselected in 1434.44s (0:23:54) ====

test_drained_stream_can_require_generated_content and test_drained_stream_accepts_generated_content both fail the same way, also after 3 auto-retries each: CancellableRequest.get_response() gets no HTTP response within 30s while read_streaming_responses drains the stream (drain=True, require_content=True). The later upload-artifact error (asterisk in the test_unsafe_or_nonexact_protobuf_pins[protobuf==6.33.*] log path) and the stage-verification error are follow-on noise, not the cause.

For agents
{"pr": 15571, "run_id": 37370105482, "run_attempt": 2, "head_sha": "3ba5ced536ab74853734d28b66d46e93a9bcea12", "failures": [{"job": "dynamo-runtime / test / parallel cuda13.0, amd64", "job_id": 112000906111, "failed_step": "Run CPU-only tests (parallelized)", "signature": "requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response", "tests": ["tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_can_require_generated_content", "tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_accepts_generated_content"], "log_cmd": "gh run view --job 112000906111 -R ai-dynamo/dynamo --log-failed"}, {"job": "dynamo-runtime / test / parallel cuda13.0, arm64", "job_id": 112000906131, "failed_step": "Run CPU-only tests (parallelized)", "signature": "requests.exceptions.Timeout: Timed out after 30.0s waiting for the HTTP response", "tests": ["tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_can_require_generated_content", "tests/fault_tolerance/cancellation/test_utils.py::test_drained_stream_accepts_generated_content"], "log_cmd": "gh run view --job 112000906131 -R ai-dynamo/dynamo --log-failed"}]}

Posted automatically by Devin for run 37370105482. Updated on every full-CI run of this PR.

@dmitry-tokarev-nv dmitry-tokarev-nv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved at 3ba5ced536ab74853734d28b66d46e93a9bcea12. One P2 of mine stays open. There is no P3.

  • [P2] .github/ci/deepseek-v4-pro-agg/patch-perf.yaml:27: no issue tracks the vLLM worker crash, and on Oct 4 it hit c128, the load point that this PR keeps.
  • Fixed: .github/ci/kimi-k25-agg/patch-perf.yaml:24. Kimi now uses concurrency 20, and the Verify AIPerf results step needs all 200 requests.
  • Open item for #15537: the two PRs still conflict in the Verify AIPerf results step. Keep the block of this PR and add || "$RECIPE" == kimi-k25-disagg. With only the #15537 side, no line sets concurrencies, and every DeepSeek result passes.
  • Not run yet: the changed step and settings run only in the nightly recipe jobs or in a manual dispatch.
What I measured.
  • Since db33213f83, the diff of this PR changed in three lines. Kimi CONCURRENCIES "128" became "20", and .request_count.avg > 0 became == 200, with a comment. The recommitted commits have the same patches as before, and the three merges of main add no change of their own. A merge with main at def3b79b15 keeps the three files of this PR unchanged.
  • The conflict with #15537 occurs in both merge orders, at the #15537 heads c4f65da29c and 1477b7b4d1.
  • I ran the step with bash --noprofile --norc -e -o pipefail against 55 fixtures: the 36 of my first round, 14 new Kimi fixtures, and the real DeepSeek c128 exports of 5 nights.
Script Result
This head Same as at db33213f83, except that Kimi profiles with 1, 199, or 201 requests now fail
Union with #15537 Passes each valid fixture and fails each bad one, 55 of 55. Kimi disaggregated also needs 200 requests, and it reads the same trace file
Only the #15537 side Passes 6 DeepSeek fixtures that must fail: the real Oct 2 sweep, the real Oct 4 c128 export, a wrong count, errors, a missing profile, and a missing CONCURRENCIES
This head with > 0 in place of == 200 Passes the Kimi fixtures with 1, 199, or 201 requests
  • CI on the head: PR run 37370105482 (attempt 2) fails dynamo-runtime / test / parallel cuda13.0 on amd64 and arm64, in 2 cancellation tests. The same 2 tests fail on main at d15ec1dda0, the merge base, in post-merge job 111954651344. This diff does not touch them. On the source branch, every check that ran passed.

Signed-off-by: Saravana Periyasamy <saperiyasamy@nvidia.com>
Signed-off-by: Saravana Periyasamy <saperiyasamy@nvidia.com>
Signed-off-by: Saravana Periyasamy <saperiyasamy@nvidia.com>
Signed-off-by: Saravana Periyasamy <saperiyasamy@nvidia.com>
@sara4dev
sara4dev force-pushed the codex/cap-deepseek-agg-perf-concurrency branch from 3ba5ced to da2797f Compare October 6, 2026 01:59
@sara4dev

sara4dev commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test da2797f

@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown

/ok to test da2797f

@sara4dev, there was an error processing your request: E2

See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/

@sara4dev

sara4dev commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test da2797f

@sara4dev
sara4dev merged commit 6ebaf07 into main Oct 6, 2026
200 of 202 checks passed
@sara4dev
sara4dev deleted the codex/cap-deepseek-agg-perf-concurrency branch October 6, 2026 03:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

actions ci Issues/PRs that reference CI build/test size/S

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants