Repository navigation
[CI/Build] Fix Qwen3-Omni Talker benchmark workload - #8057
Conversation
Signed-off-by: psv666 <2693925048@qq.com>
|
This PR appears to belong to: docs/design/module/benchmarking.md. Module owners: @alex-jw-brooks @Bounty-hunter @yenuo26 Routing: @alex-jw-brooks via module of the changed files, CODEOWNERS; @Bounty-hunter via module of the changed files, CODEOWNERS; @yenuo26 via CI owner, CODEOWNERS @psv666, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Signed-off-by: psv666 <2693925048@qq.com>
Omni ReviewBot: supersededThe CI failure noted on |
Omni ReviewBot routing recordAssigned Strict under experiment |
| # fixed stage workloads without parsing logs or storing audio payloads. | ||
| request_stage_metrics = [getattr(output, "stage_metrics", None) for output in outputs] | ||
| if any(snapshot is not None for snapshot in request_stage_metrics): | ||
| result["request_stage_metrics"] = request_stage_metrics |
There was a problem hiding this comment.
If it is only added here, it will merely be written into the result file and will not appear in the benchmark's printed output. I'm not sure whether this matches your expectation.
There was a problem hiding this comment.
Intentional: the runner reads the saved result JSON, so the check doesn't depend on stdout, and printing 128 per-request snapshots would flood the log. I added a one-line summary after the check passes (Fixed stage workload OK: 128 requests, num_tokens_out stage 0=900, stage 1=1536), and trimmed the persisted snapshot to num_tokens_out/finish_reason/audio_frames/audio_duration_s per stage.
| for stage, expected in fixed_lengths.items(): | ||
| metrics = snapshot.get(stage) if isinstance(snapshot, dict) else None | ||
| actual = metrics.get("num_tokens_out") if isinstance(metrics, dict) else None | ||
| assert actual == expected, ( |
There was a problem hiding this comment.
If you want to add a validation here that the talker needs to reach max_token, then shouldn't the thinker also add a validation that it reaches max_tokens after ignore_eos takes effect?
There was a problem hiding this comment.
Agreed. Thinker now has min_tokens == max_tokens == 900, so stage 0 enters the same check. Behavior is unchanged (Thinker has no extra stop tokens and was already 900 on all 276 local requests), but it is now enforced. Added reject cases for a missing/short Thinker.
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
Scan:
| Category | Result |
|---|---|
| Tests / verification | 9 finding(s) below |
| Security | no finding reported |
| Docs / comments | no finding reported |
| Behavior / compatibility | no finding reported |
| Correctness | no finding reported |
Validated:
- [claim-verified] 13 omitted Code2Wav tail cloned from deploy defaults: serving_chat.py:1301-1303 appends clone_sampling_params(default_params_list[idx]).
- [claim-verified] H100 arrays 0.169..0.4595 are still in the JSON; no L20X substitution at this head.
- [claim-refuted] README:46 Do not merge this workload change with those old baselines vs JSON:66-74 still shipping them as baseline.H100.mean_audio_rtf
- [validated] pipeline.py:62 Talker sampling_constraints stop_token_ids=[2150]; omni_base.py:380 merge keeps them; min_tokens==1536 is the actual pin.
- [validated] stats.py:668-669 + definitions.py:137 snapshot keys are str(stage_id) and num_tokens_out
- [validated] sibling test_qwen3_omni_async_chunk.json and test_qwen3_omni_multi_replicas.json have no sampling_params_list; production defaults unchanged.
The primary leftover is that the new Talker 1536 pin ships beside live historical H100 mean_audio_rtf targets the README itself says not to merge—clear or quarantine that slot (major). Two concrete follow-throughs belong in this PR: add Thinker min_tokens=900 so the new length assert actually covers the 900-token overhead the RTF story assumes (minor), and rerun the L20X five-point sweep with the restored max(2, concurrency) warmups before treating those numbers as evidence for this head (minor). The sibling multi_replicas random row is still variable-length Talker and needs the same extra_body or an explicit out-of-scope note (minor). RTD redness, README relocation, Mix {} persistence, and the tighter failed-assert are dropped as already-gated, polish, or no-issue.
Checked, no defect found:
tests/dfx/perf/tests/test_runner_metadata.py:708— Accept extra_body/extra-body plus the reject snapshot table correctly exercise _assert_fixed_stage_workloads for exact 1536 and missing/wrong lengths.
Verdict: REQUEST CHANGES
Findings
- [P2] Residual Thinker length is still unpinned: extra_body stage 0 is only max_token… — ``
Residual Thinker length is still unpinned: extra_body stage 0 is only max_tokens 900 + ignore_eos (no min_tokens) in tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:51-55, and tests/dfx/perf/scripts/run_benchmark.py:272 only asserts stages where min_tokens==max_tokens, so stage 0 is skipped.
Evidence: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:51-55 {"temperature": 0.0, "max_tokens": 900, "ignore_eos": true} — Thinker/stage 0 has no min_tokens. tests/dfx/perf/scripts/run_benchmark.py:272 and sampling.get("min_tokens") == sampling["max_tokens"] — only those stages enter fixed_lengths, so stage 0 is skipped. tests/dfx/perf/tests/test_runner_metadata.py:715-720 accepts {"1": {"num_tokens_out": 1536}} snapshots with stage 0 as {"max_tokens": 900} only.
- *[P2] Resolved for order: asyncio.gather(tasks) at patch.py:3469 returns outputs in… —
test_patch.py:2106
Resolved for order: asyncio.gather(*tasks) at patch.py:3469 returns outputs in input order, and test_patch.py:2106-2107 pins that after forcing request 1 to finish before 0. Residual: the stub injects stage_metrics=snapshots[index] (test_patch.py:2055,2069) and never proves a live server num_tokens_out.
Evidence: test_patch.py:2055 snapshots = [{"1": {"num_tokens_out": 1536}}, None, {"1": {"num_tokens_out": 486}}] — injected fixture, not a server response. test_patch.py:2069 stage_metrics=snapshots[index], — monkeypatched request_func writes those values. test_patch.py:2106-2107 assert completion_order.index(1) < completion_order.index(0) / assert result["request_stage_metrics"] == snapshots — pins gather order after out-of-order completion. patch.py:3469 outputs: list[MixRequestFuncOutput] = await asyncio.gather(*tasks) — gather keeps input/task order (unchanged gather contract; this PR only consumes outputs at 3584).
- [P2] Do not persist request_stage_metrics for empty Mix snapshots —
patch.py:3585
Do not persist request_stage_metrics for empty Mix snapshots. openai-chat-omni inits stage_metrics={} at patch.py:1862 before any SSE merge, and the new persist gate at patch.py:3585 uses any(snapshot is not None), so {} is treated as present. Every openai-chat-omni benchmark() result therefore gains request_stage_metrics as a list of empty dicts, not only the no_async_chunk extra_body path. Plain-vLLM RequestFuncOutput lacks the attribute, so getattr(..., None) omits the key.
Evidence: unchanged by this diff, present in the PR-time tree: vllm_omni/benchmarks/patch/patch.py:1862 output.stage_metrics = {} — Mix chat async_request_openai_chat_omni_completions (ASYNC_REQUEST_FUNCS['openai-chat-omni']) inits empty dict unconditionally. In this diff: vllm_omni/benchmarks/patch/patch.py:3584 request_stage_metrics = [getattr(output, "stage_metrics", None) for output in outputs]; vllm_omni/benchmarks/patch/patch.py:3585 if any(snapshot is not None for snapshot in request_stage_metrics): — {} is not None, so empty Mix inits persist on every openai-chat-omni path. Unchanged dataclass default: vllm_omni/benchmarks/patch/patch.py:883 stage_metrics: dict[str, dict] | None = None — without the chat overwrite, getattr would stay None and omit the key, matching plain-vLLM.
- [P2] run_benchmark.py:292
not result.get("failed", 0)is inert when the key is mis… —metrics.py:292
run_benchmark.py:292not result.get("failed", 0)is inert when the key is missing and redundant on the MultiModals path: metrics.py incrementscompletedonly on success and setsfailed=len(failed_outputs), while patch.py gathers only formal requests intooutputs, socompleted==num_promptalready impliesfailed==0. The non-MM result dict omitsfailedentirely._assert_fixed_stage_workloadsstill returns immediately unless some stage has min_tokens==max_tokens. Require the key ("failed" in result and result["failed"]==0) or drop the conjunct.
Evidence: Changed line: run_benchmark.py:292 assert result["completed"] == num_prompt and not result.get("failed", 0), "Request failures exist" — a missing failed becomes 0, so the new conjunct never fires. Unchanged by this diff, present in the PR-time tree: metrics.py:1022 completed += 1 (only inside if outputs[i].success); metrics.py:1072 failed_outputs = [output for output in outputs if not output.success]; metrics.py:1152 failed=len(failed_outputs),; patch.py:3469 outputs: list[MixRequestFuncOutput] = await asyncio.gather(*tasks) (formal requests only; warmups discarded at patch.py:3320 _ = await asyncio.gather(*warmup_tasks)); patch.py:3521-3522 "completed": mm_metrics.completed, / "failed": mm_metrics.failed, so completed+failed==len(outputs)==num_prompt. Unchanged non-MM branch patch.py:3571-3581 sets "completed": metrics.completed and no failed key. Unchanged run_benchmark.py:274-275 if not fixed_lengths: return.
- [P2] This PR pins Talker min_tokens==max_tokens==1536 only on tests/dfx/perf/tests/t… —
tests/dfx/perf/tests/test_qwen3_omni_multi_replicas.json
This PR pins Talker min_tokens==max_tokens==1536 only on tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:48-62. The unchanged sibling random row at tests/dfx/perf/tests/test_qwen3_omni_multi_replicas.json:28-36 still uses openai-chat-omni, random_input_len 2500, random_output_len 900, ignore_eos, and no sampling_params_list, so that replica sweep still measures variable-length Talker RTF. tests/dfx/perf/README.md documents only the no_async_chunk random sweep and never mentions multi_replicas. If #7915 incomparability is why 1536 is required, share the same extra_body or document that multi_replicas is intentionally left variable.
Evidence: unchanged by this diff, present in the PR-time tree: tests/dfx/perf/tests/test_qwen3_omni_multi_replicas.json:28-36 "dataset_name": "random", / "backend": "openai-chat-omni", / "random_input_len": 2500, / "random_output_len": 900, / "ignore_eos": true, / "percentile-metrics": "ttft,tpot,itl,e2el,audio_rtf,audio_ttfp,audio_duration", / "baseline": { — no extra_body or sampling_params_list on this random row. Changed by this diff: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:48-62 "extra_body": { / "return_stage_metrics": true, / "sampling_params_list": [ / { / "temperature": 0.0, / "max_tokens": 900, / "ignore_eos": true / }, / { / "temperature": 0.9, / "top_k": 50, / "repetition_penalty": 1.05, / "seed": 0, / "min_tokens": 1536, / "max_tokens": 1536 — only dfx/perf JSON with min_tokens. Changed by this diff: tests/dfx/perf/README.md:3 The \random` sweep in `tests/test_qwen3_omni_no_async_chunk.json` measures` — README never mentions multi_replicas.
| } | ||
| ] | ||
| }, | ||
| "baseline": { |
There was a problem hiding this comment.
[P1] This diff pins Talker to min_tokens==max_tokens==1536 on the existing random ro…
This diff pins Talker to min_tokens==max_tokens==1536 on the existing random row (JSON:48-62) while leaving baseline.H100.mean_audio_rtf as the historical variable-length list [0.169, 0.2041, 0.2382, 0.3177, 0.4595] (JSON:66-74). The in-diff README:46 calls those five values old variable-length baselines and says "Do not merge this workload change with those old baselines," yet CUDA/NPU nightly already load this same JSON. Unchanged conftest.py:757-762 always attaches that block to result["baseline"] for each sweep step, so nightly artifacts carry new fixed-length RTF beside the old H100 targets. Clear or quarantine the H100 slot until fixed-workload H100 medians exist. NPU also inherits _assert_fixed_stage_workloads (run_benchmark.py:263) from this extra_body with no A3 snapshot in the file.
Evidence: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:61-62 "min_tokens": 1536, / "max_tokens": 1536 — Talker fixed to 1536 on the existing random row; JSON:66-73 "baseline": { / "H100": { / "mean_audio_rtf": [ / 0.169, / 0.2041, / 0.2382, / 0.3177, / 0.4595 — historical variable-length H100 list left live. tests/dfx/perf/README.md:46 are available. **Do not merge this workload change with those old baselines.** — PR's own merge gate. Unchanged by this diff, present in the PR-time tree: tests/dfx/conftest.py:757-762 if baseline_config: / # Keep every hardware bucket; resolve list metrics to this sweep step. / result["baseline"] = resolve_baseline_for_sweep( / baseline_config, / sweep_index=sweep_index, / ) — runner always writes that JSON baseline onto the result. Unchanged by this diff, present in the PR-time tree: .buildkite/cuda/test-nightly.yml:81 - pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json -m "H100 and B200 and cards_2" and .buildkite/npu/test-npu-nightly.yml:22 pytest -s -v tests/dfx/perf/scripts/run_benchmark.py -m npu --test-config-file tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json — both nightlies load this file. tests/dfx/perf/scripts/run_benchmark.py:274-280 if not fixed_lengths: / return / `` / snapshots = result.get("request_stage_metrics") / ` assert isinstance(snapshots, list) and len(snapshots) == num_prompts, (` / ` "Fixed stage workload: missing per-request stage metrics"` / ` )` — shared length assert fires for this extra_body on NPU too; JSON has no A3 baseline slot.
Suggestion: "baseline": {}
There was a problem hiding this comment.
Agreed this must not merge with the variable-length H100 values. @yenuo26 which do you prefer: (A) run the H100 calibration described in tests/dfx/perf/qwen3_omni_fixed_workload.md and I'll put the medians into this PR, or (B) clear the H100 slot ("baseline": {}) here and recalibrate in a follow-up issue? Nightly doesn't pass --assert-baseline, but the old values would make the dashboard report a false ~15% RTF regression.
| @@ -0,0 +1,83 @@ | |||
| # Qwen3-Omni fixed-workload performance test | |||
There was a problem hiding this comment.
[P2] This diff adds tests/dfx/perf/README.md as the suite-root README, but title and…
This diff adds tests/dfx/perf/README.md as the suite-root README, but title and body only document the Qwen3-Omni no-async-chunk fixed-length sweep in tests/test_qwen3_omni_no_async_chunk.json. Unchanged sibling tests/dfx/reliability/README.md:1 indexes the whole reliability directory, and the same perf tree also holds Wan, Hunyuan, MiniCPM, Cosmos3 and other JSON configs. Move this note off the suite root (for example next to that JSON) so the perf directory index is not a single-model calibration page.
Evidence: tests/dfx/perf/README.md:1 # Qwen3-Omni fixed-workload performance test; tests/dfx/perf/README.md:3-4 The \random` sweep in `tests/test_qwen3_omni_no_async_chunk.json` measures/2500 configured input tokens, 900 Thinker output tokens, and exactly 1536 Talker— new file, entire page stays on that one JSON and H100 calibration. Unchanged by this diff, present in the PR-time tree: tests/dfx/reliability/README.md:1# L5(b) Reliability (`tests/dfx/reliability`); tests/dfx/reliability/README.md:3 This directory contains reliability fault-injection tests for key models.` Also present and unchanged: tests/dfx/perf/tests/test_wan22_i2v_vllm_omni.json, test_hunyuan_image3_it2i.json, test_minicpmo_4_5.json, test_cosmos3_vllm_omni.json and other sibling perf configs under the same suite root.
Suggestion: # Qwen3-Omni no-async-chunk fixed-workload notes
This page is not the tests/dfx/perf suite index.
| "sampling_params_list": [ | ||
| { | ||
| "temperature": 0.0, | ||
| "max_tokens": 900, |
There was a problem hiding this comment.
[P2] Thinker sampling_params_list[0] (JSON:53) sets only max_tokens=900 and ignore_e…
Thinker sampling_params_list[0] (JSON:53) sets only max_tokens=900 and ignore_eos. _assert_fixed_stage_workloads keeps a stage only when min_tokens==max_tokens (run_benchmark.py:272), so nightly never checks stage "0". The accept-path unit test uses this same Thinker-only-max_tokens shape and snapshots only stage "1" (test_runner_metadata.py:720). A short Thinker still passes. README calibration requires Thinker 900 and notes ignore_eos does not lock length. Add min_tokens=900 so stage "0" enters fixed_lengths.
Evidence: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:51-55 "sampling_params_list": [ { "temperature": 0.0, "max_tokens": 900, "ignore_eos": true }, — Thinker has no min_tokens. tests/dfx/perf/scripts/run_benchmark.py:272 and sampling.get("min_tokens") == sampling["max_tokens"] — stage 0 is dropped because get("min_tokens") is None. tests/dfx/perf/tests/test_runner_metadata.py:720 {extra_body_key: {"sampling_params_list": [{"max_tokens": 900}, {"min_tokens": 1536, "max_tokens": 1536}]}} with snapshots only for stage "1" is accepted. tests/dfx/perf/tests/test_runner_metadata.py:756 {"sampling_params_list": [{"max_tokens": 900}]} is treated as variable and does not require snapshots. tests/dfx/perf/README.md:22 ignore_eos alone does not disable explicit stop tokens. tests/dfx/perf/README.md:59 Check every Thinker length (900),
Suggestion: {
"temperature": 0.0,
"max_tokens": 900,
"min_tokens": 900,
"ignore_eos": true
},
Omni ReviewBot: finding feedback[p1] This diff pins Talker to min_tokens==max_tokens==1536 on the existing random ro… — This diff pins Talker to min_tokens==max_tokens==1536 on the existing random row (JSON:48-62) while leaving baseline.H100.mean_audio_rtf as the historical variable-length list [0.169, 0.2041, 0.2382, 0.3177, 0.4595] (JSON:66-74). The in-diff README:46 calls those five values old variable-length baselines and says "Do not merge this workload change with those old baselines," yet CUDA/NPU nightly already load this same JSON. Unchanged conftest.py:757-762 always attaches that block to result["baseline"] for each sweep step, so nightly artifacts carry new fixed-length RTF beside the old H100 targets. Clear or quarantine the H100 slot until fixed-workload H100 medians exist. NPU also inherits _assert_fixed_stage_workloads (run_benchmark.py:263) from this extra_body with no A3 snapshot in the file. Evidence: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:61-62 Suggestion: "baseline": {} This finding appears in the bot's COMMENT review, but GitHub could not place it as an inline diff comment. If you are the PR author and disagree, react 👎 to this comment. The disagreement will be shown to the maintainer; it does not approve or merge the PR. |
- Pin Thinker min_tokens=900 so the runner also verifies stage 0 length. - Persist only num_tokens_out/finish_reason/audio frames/duration per stage, and skip empty chat-omni snapshots, instead of full per-token snapshots. - Print a one-line summary after the fixed-length check passes. - Drop the no-op failed-count conjunct in assert_result. - Move the calibration notes off the perf suite root and note that multi_replicas is out of scope. Signed-off-by: psv666 <2693925048@qq.com>
|
Self-review / follow-up in d7f0f55:
|
…weep The historical H100 mean_audio_rtf values were measured with a variable-length Talker workload and are not comparable with the fixed 900/1536 workload. Leave the random row without a baseline until fixed-workload H100 medians are calibrated in a follow-up. Signed-off-by: psv666 <2693925048@qq.com>
Move the H100 calibration procedure to vllm-project#8100 and keep a one-line rationale in the perf config, matching other perf configs. No behavior change. Signed-off-by: psv666 <2693925048@qq.com>
|
Follow-up in beb9650: removed |
…er-perf Signed-off-by: psv666 <2693925048@qq.com> # Conflicts: # tests/dfx/perf/scripts/run_benchmark.py
Head branch was pushed to by a user without write access
Merge current main and rebuild the Qwen3-Omni baseline file from its updated configuration. Keep the random workload baseline empty after vllm-project#8057; the five archived A3 RTF values no longer describe that workload. Preserve all remaining 54 A3 values, including the five approved MiniCPM-o Aug 20-25 means. Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Purpose
The Qwen3-Omni non-async random performance test fixes Thinker output at 900 tokens
but lets Talker stop naturally. Different audio lengths change both the workload
and the fixed Thinker overhead per generated audio second, making RTF comparisons
hard to interpret. Related to #7915.
This change makes the five random concurrency points request exactly 1536 Talker
output tokens through the existing
sampling_params_listAPI. It preserves theTalker temperature/top-k/repetition penalty, fixes the request seed, and explicitly
sets benchmark seed and input-length range. Warmups retain the existing
max(2, concurrency)request count. Code2Wav retains server defaults. CUDA and NPU share this configuration; async/random-mm tests andproduction/model/API defaults are unchanged.
Benchmark results retain ordered
request_stage_metricsfor formal requests,including missing entries, without audio bodies. The runner rejects failed
requests, missing metrics and mismatched output counts for stages configured with
min_tokens == max_tokens. Tests cover exact lengths, premature termination,missing metrics, compatibility with nonfixed workloads, warmup exclusion and
out-of-order completion.
H100 baseline status: the random row's baseline is cleared (
"baseline": {}),as agreed with the CI owner. The five historical values were measured with the
previous variable-length workload and are not comparable, so they are removed
rather than left as false regression targets. Fixed-workload H100 medians will be
added in follow-up #8100. Local L20X results are not used as H100
baselines. The random-mm rows keep their existing baselines.
This addresses benchmark comparability; it does not establish that the historical
audio-shortening behavior in #7915 has been fixed.
Test Plan
vLLM Version: 0.30.0
vLLM-Omni Commit: upstream base
75570ee3a6572a7dc7a9463ada35a9a11c367993, plus commit491193eff401a2034b7fca119ee5f4a519d200c9.Local environment: 2 NVIDIA L20X GPUs (~140 GiB each), driver 570.133.20,
PyTorch 2.13.0+cu129, Transformers 5.14.1, flashinfer-python 0.6.18.post1,
NumPy 2.3.5, aiohttp 3.14.1. Model:
Qwen/Qwen3-Omni-30B-A3B-Instruct, cached snapshot26291f793822fb6be9555850f06dfe95f2d7e695. These runs used a shared host and uvenvironments, not a pinned H100 CI container. The environment reported an existing
flashinfer-cubin 0.6.13 / flashinfer-python version mismatch; this same runtime was
used for both local workloads.
Validation scope: the L20X measurements below used two warmups per point,
before follow-up commit
535201079restoredmax(2, concurrency). That follow-uppassed 43 CPU runner tests, a five-point runner argument check and pre-commit;
the GPU sweep has not been rerun with the restored warmup counts; only the warmup count changed and warmups are excluded from metrics, so the effect is expected to be negligible.
H100 recalibration and NPU runtime validation have not been performed because
those devices are not available locally. Follow-up validation can be arranged through CI or contributors
with access to the required hardware. The calibration procedure is tracked in #8100.
Reproduction commands for the recorded GPU measurements
Requires the repository uv environment and the cached model snapshot above.
On a shared NVIDIA host, run the server within a two-GPU reservation (
gpu run).From the checkout, start the server:
After
/healthis ready, run the fixed-workload sweep from another terminal:The command above reproduces the recorded two-warmup measurements. To run the
current PR configuration, replace
--num-warmups 2with--num-warmups "$((c < 2 ? 2 : c))".For the original workload, use unmodified base
75570ee3a, omit--extra-body,and use warmups
max(2, c). The original seed/range defaults are already 0/0.0.The original-main experiment exported existing formal-request stage metrics with
an external client observer after timing stopped; source and requests were unchanged.
Unit validation (CPU tests; no model weights required):
CI-style marker selection for the changed benchmark test:
Test Result
mypy, Markdown, SPDX and test markers).
Talker 1536 tokens on every request, Talker finish reason
length, nonemptyaudio. This covers all five concurrency points plus a concurrency-8 repeat.
on every request, 193 Talker natural stops and 51 stops at the default 4096-token
limit. One fresh service covered all five points.
Fixed-workload results (two service starts, A then B; PCM sample counts are at
24 kHz and exclude warmups):
Run A used GPUs 6/7; run B used GPUs 4/7 and temporary diagnostic logs. Run A stopped
after c=8 because an extra diagnostic assertion required exact waveform-length
equality; request and token checks passed. Run B recorded the length distribution
while retaining request/token checks. Temporary model instrumentation is removed
from this diff. These are functional measurements, not three-round calibration.
Original main/original-workload results (GPUs 6/7; warmups 2/4/8/16/32):
These tables use different output lengths and warmup settings. They demonstrate
why workload control is necessary and do not measure a code speedup/regression.
Fixed-length audio ranged from 2,943,870 to 2,947,200 samples
(122.66125-122.8 seconds, approximately 0.113% spread). Diagnostic logs located the
small variation in Code2Wav batching/tail cropping: 1536 Talker steps yielded 1535
valid codec rows; a split into 1026+509 rows can return a different tail length
when a shorter segment is cropped from a padded batch. For example,
1,969,920 + 976,170 = 2,946,090 samples, matching the observed +0.0925 seconds.
This does not explain the historical large shortening, and no model fix is included.
Natural-termination and audio-quality tests remain separate.