Skip to content

feat(mesh-ci): report p90 latency and AIPerf raw metrics in benchmark summary and support kimi-k3 agentic ci - #2177

Merged
zhuyuhua-v merged 4 commits into
mainfrom
zwan/k3-pd-ci-benchmark
Sep 10, 2026
Merged

zhuyuhua-v merged 4 commits into
mainfrom
zwan/k3-pd-ci-benchmark

Conversation

@wanzhenchn

Copy link
Copy Markdown
Contributor

CI

Benchmark matrix — .github/benchmark/models_atomesh.yaml

  • All 11 concurrency points from recipes/Agentic-Kimi-K3.md at (recipe) add Kimi-K3 AgentX recipe #2131 (30c2f1b): 1/2/4/8/12/14/16/32/40/56/64 across the three serving bands.
  • Verified field by field against the recipe's per-concurrency table: DCP, spec tokens, acceptance length, LMCache size, ReplaySSM, max-num-seqs, batched tokens, GPU utilization, graph_max.
  • CUDA graphs: --cudagraph-mode FULL --level 3 over the dense range [2 … 2*CONC*(1+spec)].
  • ReplaySSM off on C1/C2/C4, on for C8–C16, off from C32 up. AITER_REUSE_IDENTICAL_COMM_GROUPS only on C56/C64.
  • LMCache 128 GiB per rank on C8–C40, 192 GiB on C56/C64, with explicit ATOM_NUMA_NODE=0,0,0,0,1,1,1,1 and auto-binding off.
  • Throughput band drops the CPU state-offload tier: it keeps no draft model and no ReplaySSM, so OFFLOAD_STATE_CPU_SIZE was taking 32 GiB per rank from the paged KV for a state pool nothing replays from.
  • --state-checkpoint-interval-tokens -1 and ATOM_STATE_CHECKPOINT_DEMAND=0 both differ from the ATOM defaults; omitting either reproduces a different configuration than the recipe's published numbers.

Launcher plumbing — .github/scripts/atomesh/pd_server_atom.sh, pd_submit.sh, pd_slurm_job.sh

  • CUDA-graph capture sizes derived from concurrency and spec width instead of enumerated per case.
  • cudagraph_max_num_seqs available as a per-case override — C14 pins the window to 32, since 2*CONC gives 28 and the server dies during graph warmup.
  • DCP, speculative-decoding and state-checkpoint parameters propagated per role; STATE_CHECKPOINT_ allowed through the env forwarding list that gates what reaches the container.
  • server_args self-validating: the export mapping records which keys it reads and the submit step fails on any key it never consumed, so a config key can no longer be dropped without a trace.
  • GSM8K joins SWE-bench in getting its own service phase, so the accuracy server restarts with its own arguments rather than reusing the benchmark server.

Result reporting — .github/scripts/atomesh/process_result.py, aiperf_console_report.py, .github/workflows/atomesh-benchmark.yaml, .github/dashboard/index.html

  • Mean / p90 / p99 for TTFT, TPOT and E2E plus TP / DCP / Spec, replacing single averaged numbers that hid the tail dominating agentic traces; p90 exported at the source and carried on every dashboard point.
  • New aiperf_console_report.py reproduces AIPerf's console tables — prefill/decode throughput, tokens in flight, CO-aware latency — otherwise untracked; no-op for fixed ISL/OSL runs.
  • Each case's log path surfaced in the job summary; performance table and log-path table both ordered by (benchmark_model_name, max_concurrency).
  • Dashboard details popover builds parallelism and percentile fields on the markdown-summary row path, and no longer reports a hardcoded expert-parallelism value it never measured.

wanzhenchn and others added 4 commits September 9, 2026 11:18
… summary

A single averaged number per latency metric hides the tail that dominates
agentic traces, and nothing in the pipeline produced p90 at all.

- Export p90 at the source: p90_ttft_ms / p90_itl_ms / p90_e2el_ms from AIPerf,
  and --metric-percentiles=90,99 for the vLLM path, which emitted p99 only.
- Replace the summary's TTFT/TPOT/E2E columns with explicit mean / p90 / p99,
  and add TP / DCP / Spec; carry the p90 fields on every dashboard point too.
- Name the interactivity definition in its column header (P90 E2E Normalized
  vs 1 / median_tpot_s), dropping the Intvty Def column. A mixed aggregate
  summary cannot name one in the header, so it footnotes both.
- Reproduce AIPerf's console tables under the summary via
  aiperf_console_report.py: prefill/decode throughput, tokens in flight and
  CO-aware latency are otherwise untracked. No-op for fixed ISL/OSL runs.
- Fix the dashboard details popover: the markdown-summary row path never built
  the parallelism or percentile fields, "eagle32" now reads "eagle3-2", and
  Expert Parallelism no longer reports a hardcoded 1 it never measured.

Co-authored-by: Cursor <cursoragent@cursor.com>
- Add TP8/DCP8 concurrency recipes with DSpark, LMCache, and state offload.
- Pin the matching AIPerf revision and recipe-level options.
- Add GSM8K evaluation for the non-DSpark configuration.
- Propagate DCP, speculative decoding, and state checkpoint parameters.
- Isolate agentic benchmark and accuracy evaluation service phases.

fix(ci): align Kimi-K3 sweep and benchmark details

- Remove the unsupported Kimi-K3 c16 case.
- Export prefill/decode DCP and speculative decoding metadata in benchmark results.
- Show DCP and dspark/mtp configuration in dashboard point details.
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every eligible PR before approval:

  • ✅ Pre Checkin: Black, Ruff, catalog schema validation, non-GPU unit tests

Heavy model tests:

  • ✅ Run after the PR is approved and Pre Checkin passes
  • ✅ Run immediately when an approval review is submitted
  • ✅ Can be requested before approval with labels
Label Tests
ci:full Run all heavy PR model tests: native ATOM, vLLM, and SGLang
ci:atom Run native ATOM model accuracy tests
ci:vllm Run ATOM vLLM OOT model accuracy tests
ci:sglang Run ATOM SGLang model accuracy tests

Heavy jobs are skipped when the PR is not approved and no matching ci:* label is present.
Add labels via the sidebar or gh pr edit 2177 --add-label <label>

@zhuyuhua-v
zhuyuhua-v merged commit 27802a3 into main Sep 10, 2026
51 of 68 checks passed
@zhuyuhua-v
zhuyuhua-v deleted the zwan/k3-pd-ci-benchmark branch September 10, 2026 02:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants