Repository navigation
ci: self-calibrating warm-distance compile estimates - #14932
Conversation
The warm-distance model's compile estimates were hand-fitted once and never fed back: over 243 owned admissions the tiers were off by a median 58 s (bias -28 s), and a near compile from a kept build ran about 90 s against the near tier's 140 s while one from a seed ran about 155 s. - fit() adds tiers_by_start: the p50/p90 compile per tier and start kind (kept, seed, ...) with counts; predict() uses a cell from 5 compiles, else the tier (models without it predict as before). Admission records the prediction for its own start kind. - `refit` refits only the tiers and cells from the last 14 days, keeping hot files, start_classes and job_seconds; `drift` flags a p50 with at least 20 compiles that moved more than 20%; `backtest` replays admissions in time order against a model refit from those before each. - warm_model_refit.py runs that daily beside ci-dash on mini-6 and, only on drift, opens a pull request from ci/warm-model-refit (never main), or leaves the patch and summary when it has no token. - collect read nothing from a mini with no per-root log: the minis' zsh failed the whole command on the unmatched glob. It uses find now. - The committed model is refit: near 105 s (kept 91, seed 157), far 206 s, rebuild 373 s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
Warning Review limit reachedNext included review available in 1 minute. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (5)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
|
Merge receipt for |
f5c179f iOS: fix the test failures that keep iOS CI red on main (manaflow-ai#14803) 8685bf5 Hold update relaunch while agents are mid-turn (manaflow-ai#14969) dc90332 Keep CLI socket-discovery tests off the host's real cmux (manaflow-ai#14919) dd3c91b docs: shorten root agent instructions and link existing procedures (manaflow-ai#14998) 8c744df Rename edits inline or in the palette, never in an alert (manaflow-ai#14986) 9ae4383 Calmer chrome motion: appear instantly, fade out only, no overshoot (manaflow-ai#14984) 6510f56 Write opencode config JSON without escaping slashes (cmux 7140) (manaflow-ai#14805) ab5e7da ci: stop catch-up merges from failing the CLA check (manaflow-ai#14913) 52c8f41 Add cmux session move for Claude sessions (manaflow-ai#14959) 36785b1 Hide decorative Settings sidebar icons from VoiceOver (manaflow-ai#14989) 4c7158c Label the sound preview button and fix mistranslated action verbs (manaflow-ai#14983) e704a77 Bound untracked paths stored in last-turn diff baselines (manaflow-ai#14980) f073df1 Fix remote Files sidebar for names that change under NFD (manaflow-ai#14978) 5c68499 Bump bonsplit: mouse wheel scrolls the overflowed tab strip (manaflow-ai#14985) 9466dcb Keep agent resume bindings through the update-relaunch save (manaflow-ai#14971) ef8b037 docs: take release notes from a Changelog section in each PR instead of CHANGELOG.md edits (manaflow-ai#14934) 6eddfd7 ci: skip the delta diff when main moved further than the pull request (manaflow-ai#14987) fefcec7 ci: attribute red PR runs to the machine or the code, re-run machine failures once (manaflow-ai#14977) c185deb Accept file drops on remote tmux mirror panes (manaflow-ai#14981) 90773c7 test: make CmuxSidebarGit probe waits event-driven (manaflow-ai#14973) 1f2dbfe ci: skip the scheduled Blacksmith cache warmers while owned pools serve PRs (manaflow-ai#14827) 2850651 docs: add a guide to customizing cmux's look (manaflow-ai#14850) b66e365 Resolve a separate sidebar's content against its own backdrop (manaflow-ai#14841) 88a9360 UI tests: one labelled frame per action, built in CI; scripts/ui-test (manaflow-ai#14966) 20cfa78 fix(omo): resolve relative file refs in the shadow config without double-loading OpenCode config (manaflow-ai#14935) f0e964c ci: make the aggregate app-host product the default, layers opt-in (manaflow-ai#14975) 52dce98 ci: run and register the machine-failure test (manaflow-ai#14972) 7bf48bc ci: route compile admission by kept-build distance across minis (manaflow-ai#14949) 44fa3f5 Offer cmux in Open With for Markdown, source, and text files (manaflow-ai#14968) 45c2d66 Replay the Claude session id of agents in cmux ssh (cmux-tui) panes (manaflow-ai#14906) b4c1b31 Label icon-only chrome buttons and localize project panel text (manaflow-ai#14926) 14a6909 seed prefetch: keep the seed adopt would pick, of any seeded width (manaflow-ai#14944) 19e73d2 ci: self-calibrating warm-distance compile estimates (manaflow-ai#14932) fa98d86 ci: redispatch focused runs the Mac failed before any test started (manaflow-ai#14963)
Recommendation 3 of the owned-fleet warmth audit: make the compile estimates correct themselves.
Why
scripts/ci/warm-distance-model.json's tier p50s (near 140 / far 267 / rebuild 401 s) were fitted by hand once, and nothing fed the actual compile seconds back. Every admission recordspredicted_secondsagainstcompile_secondsin each mini'sadmissions.jsonl. Over the 243 scored owned compiles in those logs (2026-09-26 03:03 to 09-27 09:08 UTC), the model's median absolute error is 58 s, with a bias of −28 s, and only 22% of near compiles land within 25%. The tiers miss the start kind. A near compile from a kept build ran ~91 s and one from a seed ran ~157 s, but both were predicted at 140 s.What
fit()addstiers_by_start. It stores p50/p90/mean/n and the app-rebuilt count per tier × start kind (kept,seed, ...), with counts.predict(feature, model, start)uses the cell once it has ≥5 compiles, and otherwise falls back to the tier. A p50 that is missing, zero, negative, NaN or a bool is ignored. Models withouttiers_by_startpredict exactly as before, so the schema is additive andversionstays 1.admissionnow records the prediction for its own start kind. Medians throughout, so one hung 3,000 s compile moves nothing (tested).warm_distance.py refit. It refits only the tiers and cells from the last 14 days of admissions under the committed near threshold and hot files. Hot files,start_classesandjob_secondsare left as committed, because they need git history or glaeda's job log, and the picker work ownsstart_classes. Before fitting it dedupes admissions and drops implausible compiles (<5 s or >2 h). A tier with fewer than 5 rows keeps its committed entry.drift()flags any tier or cell p50 with n ≥ 20 that moved more than 20% from what the committed model predicts for it.backtestis the time-ordered replay below.scripts/ci/warm_model_refit.pyis the daily job, meant for mini-6. It reads admissions (ci-dash'sestimates.jsonl, whose compact linesread_rowsnow accepts, or--hostsover SSH), refits, and stops unless something drifted. On drift it writeswarm-distance-model.json, agit apply-ablerefit.patch,summary.md(the drift, the cells, errors before and after, and the replay) andstatus.json. With--token-fileit force-resetsci/warm-model-refit, and only that branch (the name is hard-coded and asserted not to be main), to the checkout's HEAD. It commits the model there through the contents API and opens or updates one PR against main. It never writes main. Without a token, the patch and summary in the output directory are the result.warm_distance.py collectread nothing from any mini without a per-rootcmux-ci-*/admissions.jsonl. The minis' login shell is zsh, and an unmatched glob fails the wholecat. It returned 32 of the 251 lines on the fleet. It usesfindnow.refiton the 243 admissions). Tiers: near 140→105, far 267→206, rebuild 401→373. Cells: near kept 91 / seed 157, far kept 121 / seed 247, rebuild kept 355 / seed 442. Onlytiers,tiers_by_start,misclassified,rows,fitted_atand the newcalibratedchange.start_classesandjob_secondsare untouched.tiers[*].p50, so it keeps working unchanged. Hook: cost a root's kept build with the model's kept-start cell teamleaderleo/glaeda#1304 makes it cost a root's kept build with the(tier, kept)cell under the same ≥5 rule, with a parity test.Backtest: committed model vs self-calibrated, time-ordered replay
Data: every mini's
admissions.jsonl, read over SSH (read-only) and deduped, 243 successful compiles. "Model" is the committed tier p50s, which predicted all of these. The self-calibrated prediction for each compile isrefit()from only the admissions before it: its cell once that has 5 earlier compiles, else the (refit) tier. Within 25% means |actual − predicted| ≤ 25% of predicted.Refit before every admission (continuous):
Refit once a day (the job's cadence), scored on the 65 compiles of 09-27, each predicted by the refit from 09-26's admissions (out of sample):
Where the refit runs, and why
CI_SEED_TRUSTED_POOL, cmux15): holds only its own mini's log (cmux15 has none), and owned runners run PR code.estimates.jsonl, 30 days, the fields the refit needs) and runs no GitHub runners (hq#595). A token there is out of reach of PR code, so the refit runs there as_cidashfrom a checkout of main.GITHUB_TOKENcannot open PRs here (the repo's "Allow GitHub Actions to create and approve pull requests" is off). The job therefore takes an explicit token, and without one it leaves the patch and summary as its artifact.Setup on mini-6
manaflow-ai/cmuxterm-hq#785 installs the refit with ci-dash's own
install.sh, and hq-autodeploy runs it. It covers the_cidashsparse https checkout of this repo's main and the daily LaunchDaemonai.manaflow.ci-dash-warm-refit, and it only syncs until this PR is on main. After both PRs merge:fleet push-probeso ci-dash ingests admission lines. It needs no sudo;/estimates.jsoncurrently reports n = 0.state/warm-model-refit.token(_cidash, 600; the command is in the ci-dash README). Without it, the job only writesrefit.patchandsummary.md. A PAT-opened PR triggers CI, where aGITHUB_TOKENone would not.Tests
tests/test_ci_warm_distance.py(29 tests) covers:git apply --checkagainst the tree), and PR calls that touch onlyci/warm-model-refitand refuse maintest_seed_decide,test_ci_pr_runner_pool,test_ci_owned_build_stateandtest_ci_admission_placementpass.The edits in
warm_distance.pystay inpredict/fit/the model schema, plus new refit/backtest functions andcollect.route_admission, the candidate scoring andstart_classesare untouched.🤖 Generated with Claude Code