Conversation
…6b2) 13 of 14 failed merge_group runs (2026-09-25 15:00-20:30 PT) were two files crossing the FIXED 300 s per-file wall (0 assertion failures): test_kanban_home_session.py and test_execution_flag_detection.py. Each kill ejected the queue entry and cancelled the builds behind it. - run_tests_parallel.py: per-file budget = clamp(3 x p90, floor, cap), floor = --file-timeout (300), cap = max(floor, 900). Basis = p90 of the file's recent samples (new test_durations_history.json) plus its last cached duration. --generate-slices stamps `file_timeouts` (path=secs, only files above the floor) into each slice row; the test job passes it via the new --file-timeouts flag. - tests.yml: generate restores the history cache (own key/path, LPT cache untouched); save-durations appends each main run's durations, keeping the newest 20 per file. - ci.yaml: merge_group runs 16 slices like pull_request (was 8). Verified: tests/test_run_tests_parallel_file_budget.py 19 passed (+ arm and runner routing suites, 86 total); mutants (runner ignores budgets / generate skips stamping / basis ignores history) turn 4 and 2 tests RED. Replayed over 399 real slice-duration artifacts: 3x last-sample would still have killed 2 runs; 3x p90(last 20) killed 0.
… (t_bf25c6b2) A file with no duration sample is not known to fit in 300 s; a false kill ejects a merge-queue entry and cancels every build behind it. The budget for an unmeasured file is now max(floor, 900). - generate stamps unmeasured files at the cap; with no duration data at all it stamps a single '*=<cap>' default entry, not one per file. - --file-timeouts default is None. When the flag is passed (CI always passes it, possibly ''), the stamp is authoritative: unlisted file -> '*' entry, else floor. The test jobs have no local cache, so falling through to it would have made every file unmeasured. Verified: tests/test_run_tests_parallel_file_budget.py 23 passed. Mutants: unmeasured->floor 5 RED; passed empty stamp not authoritative 1 RED. ruff clean.
|
🤖 merged-by: apollo · lane: unspecified · gate: BYPASS: FleetReview daily budget exhausted (spent=600 paged=true), no terminal record possible today; Ace directed admin-merge of the queue-bottleneck fixes (21:06 PT) · why: GUARD for merge-queue ejections: per-file timeout budgets measured at 3x p90 instead of a flat 300s ceiling, 16 merge_group slices. Companion to root-cause #1183 (merged f94ff55). 0 conflicts vs main. · red-ci allowed [Review label gate / Review label gate,All required checks pass (required, ABSENT on head)]: slice 10/16 red = tests/discord/test_restart_backfill.py::test_backfill_first_boot_no_state_scans_nothing, not in diff; 17/17 green 2x on branch hermetic HOME; 3rd sighting today of this file flaking under xdist load (also hit #1182); Review label gate red = ci-reviewed label just applied |
This comment has been minimized.
This comment has been minimized.
…files (t_bf25c6b2) 5c755ba gave every unmeasured file the 900 s cap on BOTH paths. The runner's own kill self-tests (test_run_tests_parallel_timeout_verdict.py) run it with --file-timeout 8 on an uncached hanging probe, so the probe got 900 s, and the outer file hit its own 300 s budget: CI slice 4/16 TIMED OUT on run 36217271474. The cap for unmeasured files now applies only on the stamped path, which is what CI always runs. That is where the queue fix is needed. The local no-stamp path keeps the caller's floor. Verified: file_budget + timeout_verdict 35 passed; noop_guard + run_tests_parallel 19 passed / 1 skipped; ruff clean.
|
Closing as superseded (rebase shard t_b88390a3). The fixed-300s per-file kill this PR budgets around no longer exists on main:
|
Kanban: t_bf25c6b2 (related: t_9c66df3d makes the two slow files fast + retry; t_bd6d58c3 kanban_db import cost)
Why
13 of 14 failed
merge_groupruns on 2026-09-25 (15:00-20:30 PT) were a FIXED 300 s per-file timeout on two files, with 0 assertion failures:tests/hermes_cli/test_kanban_home_session.pyandtests/tools/test_execution_flag_detection.py. Each kill ejected its queue entry and cancelled the builds behind it (70 success / 30 failure / 69 cancelled).Measured over 3,200 real
test-durations-slice-*artifacts (2026-09-25 12:00Z to 09-26 03:47Z): those files ran 75-300 s on single attempts, so a fixed 300 s is a coin flip.What
scripts/run_tests_parallel.py:clamp(3 x p90, floor, cap). The floor is--file-timeout(default 300 s). The cap ismax(floor, 900 s). The basis is the p90 of the file's last 20 main-branch samples plus its last cached duration.--generate-slicesstampsfile_timeouts(path=secs:..., only for files above the floor) into each slice row. The test job passes it with the new--file-timeouts. A file with no measured duration gets the 900 s cap (not proven to fit 300). Unmeasured files are stamped individually. With no duration data at all, the stamp is a single*=900default. A passed stamp (CI always passes one, possibly empty) is authoritative, and unlisted files get the floor. The local no-stamp path (no--file-timeouts) keeps the caller's--file-timeoutfor unmeasured files, so the runner's kill self-tests (--file-timeout 8) still kill.test_durations_history.jsonlives in its own cache entry (test-durations-history-*).generaterestores it.save-durationsappends each main run through--merge-duration-history. The existing LPT cache entry and its path are untouched.Replay (same 3,200 artifacts, chronological, per file-run)
All 6 residual runs are
test_local_env_blocklist.py(normally 15-31 s, recorded 309-318 s) ortest_kanban_home_cards.pyonce. These are hang-then-quick-retry sums, not slow files. The budget deliberately does not widen for them, and the existing one-shot retry already rescues them.Caveat: recorded values above the timeout are retry sums (attempt 1 killed + attempt 2), so they are lower bounds on the first attempt's wall time.
Verification
tests/test_run_tests_parallel_file_budget.py: 24 passed (5 new for unmeasured files at the cap /*default / authoritative empty stamp). With the arm and runner routing suites, 86 passed.actionlintshows only the 2 pre-existing ci.yaml:134-135sparse-checkoutfindings, which are identical on base.Rollout note
The history cache is empty until the first
push: mainsave-durationsafter merge. Until then, budgets fall back to 3 x the single cached LPT duration. Files missing from that cache get 900 s. Nothing gets less than today's 300 s.