test[notask]: file-driven TTS-GGML mobile benchmark selection (test-groups + perf-tests) - #2918
Merged
Merged
Conversation
Adopt the same mobile sharding mechanism the llm-llamacpp addon uses: drop a test-groups.json (auto-detected by the mobile CI action) listing the functional tests, and a perf-tests.json listing the benchmarks. The two benchmark runners (runRtfBenchmarkTest, runStreamingBenchmarkTest) are deliberately left out of the groups, so normal mobile runs never execute them and they cannot produce a 0/0 result. They run only in perf-only mode (QVAC_PERF_ONLY=true) via perf-tests.json. This removes the need for any intentional-skip handling in the test harness. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SBoRDUf9ZQeE5LFH5rwnWs
A benchmark run (run_rtf_benchmarks=true) now builds the device-farm test-groups from packages/tts-ggml/test/mobile/perf-tests.json, so only the benchmark runners execute. Normal runs skip this step (empty output) and the action auto-detects test-groups.json (functional tests, benchmarks excluded). Net: perf-tests.json is the single editable source of truth for which mobile benchmarks run — add a benchmark by adding its function name there, no workflow changes. Mirrors the file-driven convention used by llm-llamacpp. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SBoRDUf9ZQeE5LFH5rwnWs
Contributor
Review StatusCurrent Status: ❌ PENDING Pending reviews: Needs 1 Management or Team Lead, and 1 more from Management, Team Lead, or Member. |
GustavoA1604
approved these changes
Jun 29, 2026
ogad-tether
approved these changes
Jun 29, 2026
This was referenced Jun 29, 2026
ogad-tether
added a commit
to ogad-tether/qvac
that referenced
this pull request
Jul 1, 2026
…d mobile group That test loads Chatterbox variants on the GPU; on Adreno (Samsung S25 Ultra) it hits a PRE-EXISTING ggml-opencl SIGSEGV in ggml_backend_opencl_buffer_set_tensor (Q4_0 SOA upload, clEnqueueWriteBuffer) at model load — unrelated to this PR and present since before it (proven by a base A/B: the pre-tetherto#71 tts-cpp and the pre-tetherto#2918 commit crash identically; Pixel 9 / Mali-Vulkan passes). It was only added to the Android group by tetherto#2918 and merged before its mobile E2E ran, so it slipped through. Keep it on iOS (Metal) — it passes there and exercises this PR's q8-KV-on-Metal fix. Re-add to Android once the upstream ggml-opencl Adreno upload bug is fixed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🎯 What problem does this PR solve?
runRtfBenchmarkTest,runStreamingBenchmarkTest) are heavy, env-gated tests. On a normal mobile run they register zero sub-tests, and under the0/0 = FAILrule (PR testing whispercpp #47) that turned the mobile integration job red.📝 How does it solve it?
Adopt the same file-driven mobile sharding the
llm-llamacppaddon already uses — so the benchmarks simply don't run on normal runs, and the selection lives in editable files (no workflow edits to change it):test/mobile/test-groups.json— lists the functional tests. The mobile CI action auto-detects this file and runs only these (Mocha--grep). The two benchmarks are deliberately excluded → they never run on a normal run → no0/0.test/mobile/perf-tests.json— the single editable list of mobile benchmarks. Add/remove a benchmark by editing this one file.integration-mobile-test-tts-ggml.yml— onrun_rtf_benchmarks=true(the dedicatedbenchmark-rtf-tts-ggml.ymlpath), a step builds the device-farmtest-groupsfromperf-tests.jsonso only the benchmarks run. Normal runs skip the step and auto-detect the functional groups.No test-harness (
qvac-test-addon-mobile) change is needed:mainalready treats0/0as FAIL, and with this change the benchmarks never produce a0/0.🧪 How was it tested?
Two CI verification runs on this branch (test harness =
main):run_rtf_benchmarks=trueviabenchmark-rtf-tts-ggml.yml— ✅ success: https://github.com/tetherto/qvac/actions/runs/28277742131total: 2), both PASSED on iPhone 16 Pro, iPhone 17 Pro, Pixel 9 Pro XL, and Samsung S25 Ultra.perf-report-tts-ggml-iOS/-Android) — metrics captured.total: 10, passed: 10, failed: 0. The 10 functional tests ran and passed; both benchmarks absent, no0/0.runChatterboxKvCacheGpuTest, unrelated to this change).