ci(tests): restore upstream slice-matrix — Python wall fix (4,847 files vs 30-min budget) - #16
Conversation
૮ >ﻌ< ა ci reviewran on f7de6f2 — fix(tests): overlay triage — dedupe fireworks shadow, drop-i ❌ Job failuresPython tests / Run tests slice 1/8 · View jobJob Python tests / Run tests slice 1/8 failed. Python tests / Run tests slice 2/8 · View jobJob Python tests / Run tests slice 2/8 failed. Python tests / Run tests slice 5/8 · View jobJob Python tests / Run tests slice 5/8 failed. Python tests / Run tests slice 6/8 · View jobJob Python tests / Run tests slice 6/8 failed. Python tests / Run tests slice 7/8 · View jobJob Python tests / Run tests slice 7/8 failed. Python tests / Run tests slice 8/8 · View jobJob Python tests / Run tests slice 8/8 failed. ℹ️ InfoCI-sensitive file review · View jobPR touches sensitive files, but the Sensitive files changed: debug infoCI timingsCI timings · View report · View jobWall time 17m18s vs 31m35s (-45.2%). 6 job(s) slower, 11 faster, 1 unchanged.
|
Label audit:
|
| Item | Status |
|---|---|
| Branch provenance | ci/test-slices cut at fork main 467624c (post-#15 merge); single commit a44ebf58 — tests.yml restored verbatim from upstream parent-of-removal fce30d818 |
| Content integrity | contents-API read-back byte-identical to the staged source; single-file commit asserted (commit touched only tests.yml); YAML parse OK (PyYAML) before push |
| Design provenance | upstream's own pre-August slice matrix (removal: 10f99bc15, 2026-08-22, "run the work lanes on larger runners"); July upstream runs green on this design at ubuntu-latest (slices 1/8–8/8 at 68391930d) |
| Machinery alive | run_tests_parallel.py --generate-slices rc=0 at fork main; 8 slices balanced 606×7+605 = 4,847 files; stdlib-only imports |
| Root cause | unsliced 4,847-file suite vs 30-min job budget on 4-core — receipted every-head cancelled at ~30m40s; filed upstream as NousResearch#120148 |
| What validates here | classify_changes.py runs all lanes on .github/ changes — this PR's run exercises the full matrix including the restored workflow itself |
| Expected noise | JS may red intermittently (receipted flaky virtualHistoryOffsetCache, upstream NousResearch#120138 — non-blocking by classification); nix expected green (hermes_platform fix is in the base); per-slice wall time is the metric of interest — cold duration cache means blind LPT on this first run |
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: POWERFULMOVES/PMOVES-hermes-agent/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…d in schema probe fixture, allowlist pmoves_bootstrap bare_which row
CI verdict — slice-matrix restore: the wall is dead; first real Python verdicts; 7 run-1 failures triaged (3 fork-owned, 4 upstream-class across 3 files)The fix worked. 8 slices, Fork-owned defects — fixing on this branch (each push re-runs all slices)
Under probe (upstream content; upstream tip runs the suite green)
Disclosed trade-off, on the record: this restores upstream's own July design, but the parallelism profile differs from today's upstream (8 workers × 4-core vs 1 × 96-core). Timing-sensitive tests get a fresh verdict on every slice run; the two timing candidates (4, 6) will re-verdict on the fix pushes before this merges. No merge until the three fork-owned fixes land green. Run 2 (
|
Summary
Restores upstream's pre-August slice-matrix design for
Python tests / Run tests— one file (.github/workflows/tests.yml, ~140-line diff), restored verbatim from upstream's own history (fce30d818, parent of the 2026-08-22 removal commit10f99bc15"ci: run the work lanes on larger runners and merge the split jobs"). Zero editorial changes: this is what upstream itself ran in July, when this exact suite partitioned 8 ways ran green on identicalubuntu-latestrunners.Why
Upstream removed slicing when it moved the suite to a 96-core runner class. This fork cannot service that runner class (receipted in the label audits of PRs #9–#15): our
ubuntu-latestis a 4-core host, and the unsliced suite (4,847 test files) cannot finish inside the 30-minute job budget — every CI head since the v2026.9.21 absorb diedcompleted/cancelledat ~30m40s. Full write-up filed upstream: NousResearch#120148.The change
generatejob (LPT slicing viascripts/run_tests_parallel.py --generate-slices, seeded from the duration cache), the 8-way slice matrix (run_tests.sh --files '<matrix.slice.files>'), and thesave-durationsmerge job. All machinery is still alive in the repo's runner scripts (stdlib-only) — receipted at fork main467624c.merge-multiplesame-path race).ci.yamlneeds no edit: it calls this workflow withoutwith:, andslice_countdefaults to 8.classify_changes.pyruns the full lane set on any.github/change — this PR validates everything automatically.Evidence before push
467624c: rc=0 → 8 slices, balanced 606×7 + 605 files.68391930d:Run tests slice 1/8…8/8all success on this design.Validation plan
This PR's own run is the empirical verdict. Key metric: per-slice wall time on 4-core. If slices approach the 30-minute per-job budget,
slice_countis a one-line tunable (workflow input default;ci.yamlgains awith:) — iterate by push.