fix(tests): rebalance portable serial shards from current CI timings - #164
Merged
Merged
Conversation
Shard 1's embedded duration hints drifted far below reality (watch-triage 4.4min hinted vs 11.7min measured), packing ~30min of work into shard 1 while shard 3 carried ~16min. Refresh the hint table from four green CI runs so all five shards pack to ~21min.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Shard 1 of the portable serial lane kept landing on its 30-minute timeout with all tests passing. The cause was stale balance weights, not suite size: the embedded duration hints in bin/fm-test-run.sh still described the suite as of early September, while the real suite had grown underneath them.
The single biggest drift: tests/fm-watch-triage.test.sh is hinted at 4.4 min but measures 11.7 min. tests/fm-procevent.test.sh (1.2 min hinted, 3.7 min measured), tests/fm-secondmate-safety.test.sh (1.0 hinted, 2.7 measured), and 28 scripts with no hint at all filled out the rest of the gap.
Fix: refreshed portable_serial_weight_hints from the slowest duration_ms per script in the fm-test-timing-portable-serial artifacts of four green CI runs (35665191329, 35664891424, 35663110890, 35661755328), retaining the 5121 ms native-Windows measurement for fm-pi-windows-shell-invocation.test.sh, which the portable shards gate-skip. No timeout raised, no shard count change: the suite totals 105.8 min of slowest-sum script time, so five shards at ~21.2 min each keep ~8 min of headroom under the 30-minute hang tripwire.
Per-shard script-time sums, slowest-of-4-runs per shard before (old packing) vs slowest-sum after (new packing):
Spread narrowed from ~13 min to ~0. Before numbers are measured sums from the four runs above (job setup adds ~1-2 min on top, which is what pushed shard 1 over the cap); after numbers are the new packing's slowest-sums. serial_unhinted goes 28 -> 0.
Validation: bin/fm-test-run.sh --check-coverage passes, tests/fm-test-run.test.sh 42/42, tests/fm-ci-workflow.test.sh 6/6, tests/fm-lint-workflows.test.sh 16/16, actionlint and fm-doc-audience-check clean.