Skip to content

ci: self-calibrating warm-distance compile estimates - #14932

Merged
teamleaderleo merged 1 commit into
mainfrom
ci/warm-model-selfcal
Sep 27, 2026
Merged

teamleaderleo merged 1 commit into
mainfrom
ci/warm-model-selfcal

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Recommendation 3 of the owned-fleet warmth audit: make the compile estimates correct themselves.

Why

scripts/ci/warm-distance-model.json's tier p50s (near 140 / far 267 / rebuild 401 s) were fitted by hand once, and nothing fed the actual compile seconds back. Every admission records predicted_seconds against compile_seconds in each mini's admissions.jsonl. Over the 243 scored owned compiles in those logs (2026-09-26 03:03 to 09-27 09:08 UTC), the model's median absolute error is 58 s, with a bias of −28 s, and only 22% of near compiles land within 25%. The tiers miss the start kind. A near compile from a kept build ran ~91 s and one from a seed ran ~157 s, but both were predicted at 140 s.

What

  • fit() adds tiers_by_start. It stores p50/p90/mean/n and the app-rebuilt count per tier × start kind (kept, seed, ...), with counts. predict(feature, model, start) uses the cell once it has ≥5 compiles, and otherwise falls back to the tier. A p50 that is missing, zero, negative, NaN or a bool is ignored. Models without tiers_by_start predict exactly as before, so the schema is additive and version stays 1. admission now records the prediction for its own start kind. Medians throughout, so one hung 3,000 s compile moves nothing (tested).
  • warm_distance.py refit. It refits only the tiers and cells from the last 14 days of admissions under the committed near threshold and hot files. Hot files, start_classes and job_seconds are left as committed, because they need git history or glaeda's job log, and the picker work owns start_classes. Before fitting it dedupes admissions and drops implausible compiles (<5 s or >2 h). A tier with fewer than 5 rows keeps its committed entry. drift() flags any tier or cell p50 with n ≥ 20 that moved more than 20% from what the committed model predicts for it. backtest is the time-ordered replay below.
  • scripts/ci/warm_model_refit.py is the daily job, meant for mini-6. It reads admissions (ci-dash's estimates.jsonl, whose compact lines read_rows now accepts, or --hosts over SSH), refits, and stops unless something drifted. On drift it writes warm-distance-model.json, a git apply-able refit.patch, summary.md (the drift, the cells, errors before and after, and the replay) and status.json. With --token-file it force-resets ci/warm-model-refit, and only that branch (the name is hard-coded and asserted not to be main), to the checkout's HEAD. It commits the model there through the contents API and opens or updates one PR against main. It never writes main. Without a token, the patch and summary in the output directory are the result.
  • Collector bug fixed. warm_distance.py collect read nothing from any mini without a per-root cmux-ci-*/admissions.jsonl. The minis' login shell is zsh, and an unmatched glob fails the whole cat. It returned 32 of the 251 lines on the fleet. It uses find now.
  • The committed model is refit with exactly what the job would propose (refit on the 243 admissions). Tiers: near 140→105, far 267→206, rebuild 401→373. Cells: near kept 91 / seed 157, far kept 121 / seed 247, rebuild kept 355 / seed 442. Only tiers, tiers_by_start, misclassified, rows, fitted_at and the new calibrated change. start_classes and job_seconds are untouched.
  • glaeda's hook reads only tiers[*].p50, so it keeps working unchanged. Hook: cost a root's kept build with the model's kept-start cell teamleaderleo/glaeda#1304 makes it cost a root's kept build with the (tier, kept) cell under the same ≥5 rule, with a parity test.

Backtest: committed model vs self-calibrated, time-ordered replay

Data: every mini's admissions.jsonl, read over SSH (read-only) and deduped, 243 successful compiles. "Model" is the committed tier p50s, which predicted all of these. The self-calibrated prediction for each compile is refit() from only the admissions before it: its cell once that has 5 earlier compiles, else the (refit) tier. Within 25% means |actual − predicted| ≤ 25% of predicted.

Refit before every admission (continuous):

compiles n model MAE s model bias s model within 25% self-calibrated MAE s bias s within 25%
all 243 57.8 -28.1 53% 43.0 -7.1 60%
tier near 98 54.5 -31.2 22% 43.2 -9.5 38%
tier far 22 69.0 -56.7 41% 60.2 -40.5 50%
tier rebuild 123 58.6 -27.3 80% 40.9 -1.0 80%
start kept 130 57.8 -52.2 55% 33.5 -7.1 65%
start seed 113 57.8 22.5 51% 64.2 -8.3 55%
near from kept 44 55.2 -49.9 14% 29.4 -5.0 41%
near from seed 54 53.4 17.9 30% 71.1 -23.8 35%
far from kept 10 116.4 -116.4 20% 84.2 -60.2 40%
far from seed 12 43.4 -22.7 58% 40.5 -12.7 58%
rebuild from kept 76 57.0 -48.7 83% 31.5 -4.7 83%
rebuild from seed 47 62.0 41.0 74% 52.0 6.7 77%

Refit once a day (the job's cadence), scored on the 65 compiles of 09-27, each predicted by the refit from 09-26's admissions (out of sample):

compiles n model MAE s model bias s model within 25% refit daily MAE s bias s within 25%
all 65 55.6 -28.0 52% 38.4 12.4 63%
tier near 27 46.7 -22.4 22% 36.7 16.5 44%
tier far 2 118.4 -118.4 0% 64.7 -64.7 0%
tier rebuild 36 60.0 -29.2 78% 40.8 14.5 81%
start kept 37 54.0 -40.4 62% 31.9 16.5 70%
start seed 28 59.1 17.6 39% 61.6 -14.0 54%
near from kept 13 53.3 -36.2 23% 25.1 16.5 46%
near from seed 14 45.4 21.9 21% 48.4 2.5 43%
far from kept 1 145.6 -145.6 0% 58.3 -58.3 0%
far from seed 1 91.1 -91.1 0% 71.1 -71.1 0%
rebuild from kept 23 56.5 -40.4 87% 31.9 16.6 87%
rebuild from seed 13 68.1 32.7 62% 71.5 -8.3 69%
  • Kept starts gain the most. Near-from-kept MAE falls 55→29 s and within-25% rises 14%→41%. Rebuild-from-kept falls 57→32 s.
  • Near-from-seed does not improve. That cell is wide (50–450 s), and its first rows (09-26 morning) ran high, so early cells over-predicted. Recency windows and shrinkage toward the tier did no better (tried windows 10–50 and blends k = 5/10: near-from-seed MAE stayed 58–69 s). Its full-window p50 (157 s) sits within 20% of the old 140 s, so the drift gate would not have proposed it alone.
  • Far is thin. Far has only 22 compiles, 10 of them from kept builds, so the far cells are the least settled. The kept cell still predicts from 5.
  • Limited history. The admission logs start on 09-26, so the daily replay has one refit boundary. The job's summary will carry the same table with more days as they accrue.

Where the refit runs, and why

  • GitHub-hosted runners: cannot reach the minis or the tailnet.
  • An owned trusted runner (CI_SEED_TRUSTED_POOL, cmux15): holds only its own mini's log (cmux15 has none), and owned runners run PR code.
  • ci-dash (mini-6): already collects every mini's admission lines (estimates.jsonl, 30 days, the fields the refit needs) and runs no GitHub runners (hq#595). A token there is out of reach of PR code, so the refit runs there as _cidash from a checkout of main.
  • Opening PRs: GITHUB_TOKEN cannot open PRs here (the repo's "Allow GitHub Actions to create and approve pull requests" is off). The job therefore takes an explicit token, and without one it leaves the patch and summary as its artifact.
  • Untrusted inputs: admission lines are written by jobs, pull requests' included. The refit uses medians and plausible times only, and a person merges every change.

Setup on mini-6

manaflow-ai/cmuxterm-hq#785 installs the refit with ci-dash's own install.sh, and hq-autodeploy runs it. It covers the _cidash sparse https checkout of this repo's main and the daily LaunchDaemon ai.manaflow.ci-dash-warm-refit, and it only syncs until this PR is on main. After both PRs merge:

  1. Probe. Run fleet push-probe so ci-dash ingests admission lines. It needs no sudo; /estimates.json currently reports n = 0.
  2. Optional token. For automatic PRs, create a fine-grained token limited to manaflow-ai/cmux with Contents: write and Pull requests: write, and install it as state/warm-model-refit.token (_cidash, 600; the command is in the ci-dash README). Without it, the job only writes refit.patch and summary.md. A PAT-opened PR triggers CI, where a GITHUB_TOKEN one would not.

Tests

tests/test_ci_warm_distance.py (29 tests) covers:

  • fit cells and counts, and robustness to an outlier
  • predict with a full cell, a sparse cell, bad cells and old models
  • the admission recording the kept cell
  • refit's window, dedupe, implausible rows and kept-as-committed parts
  • the drift threshold and minimum n
  • backtest time order, with no future rows
  • ci-dash line normalization and the collect command
  • the refit job: no-drift, drift patch/summary (git apply --check against the tree), and PR calls that touch only ci/warm-model-refit and refuse main

test_seed_decide, test_ci_pr_runner_pool, test_ci_owned_build_state and test_ci_admission_placement pass.

The edits in warm_distance.py stay in predict/fit/the model schema, plus new refit/backtest functions and collect. route_admission, the candidate scoring and start_classes are untouched.

🤖 Generated with Claude Code

The warm-distance model's compile estimates were hand-fitted once and never
fed back: over 243 owned admissions the tiers were off by a median 58 s
(bias -28 s), and a near compile from a kept build ran about 90 s against
the near tier's 140 s while one from a seed ran about 155 s.

- fit() adds tiers_by_start: the p50/p90 compile per tier and start kind
  (kept, seed, ...) with counts; predict() uses a cell from 5 compiles, else
  the tier (models without it predict as before). Admission records the
  prediction for its own start kind.
- `refit` refits only the tiers and cells from the last 14 days, keeping hot
  files, start_classes and job_seconds; `drift` flags a p50 with at least 20
  compiles that moved more than 20%; `backtest` replays admissions in time
  order against a model refit from those before each.
- warm_model_refit.py runs that daily beside ci-dash on mini-6 and, only on
  drift, opens a pull request from ci/warm-model-refit (never main), or
  leaves the patch and summary when it has no token.
- collect read nothing from a mini with no per-root log: the minis' zsh
  failed the whole command on the unmatched glob. It uses find now.
- The committed model is refit: near 105 s (kept 91, seed 157), far 206 s,
  rebuild 373 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cursor

cursor Bot commented Sep 27, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@coderabbitai

coderabbitai Bot commented Sep 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 1 minute.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 859655e3-aab1-4d25-8e10-fb88407befd5

📥 Commits

Reviewing files that changed from the base of the PR and between 6431ac2 and b484034.

📒 Files selected for processing (5)
  • docs/ci-runners.md
  • scripts/ci/warm-distance-model.json
  • scripts/ci/warm_distance.py
  • scripts/ci/warm_model_refit.py
  • tests/test_ci_warm_distance.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@teamleaderleo
teamleaderleo merged commit 19e73d2 into main Sep 27, 2026
64 of 65 checks passed
@teamleaderleo
teamleaderleo deleted the ci/warm-model-selfcal branch September 27, 2026 11:14
@github-actions

Copy link
Copy Markdown
Contributor

Merge receipt for b4840343b9: every check was green at merge (13 verified; 17 skipped by policy). Full suite runs on main after merge.

rustybret pushed a commit to rustybret/bmux that referenced this pull request Sep 27, 2026
f5c179f iOS: fix the test failures that keep iOS CI red on main (manaflow-ai#14803)
8685bf5 Hold update relaunch while agents are mid-turn (manaflow-ai#14969)
dc90332 Keep CLI socket-discovery tests off the host's real cmux (manaflow-ai#14919)
dd3c91b docs: shorten root agent instructions and link existing procedures (manaflow-ai#14998)
8c744df Rename edits inline or in the palette, never in an alert (manaflow-ai#14986)
9ae4383 Calmer chrome motion: appear instantly, fade out only, no overshoot (manaflow-ai#14984)
6510f56 Write opencode config JSON without escaping slashes (cmux 7140) (manaflow-ai#14805)
ab5e7da ci: stop catch-up merges from failing the CLA check (manaflow-ai#14913)
52c8f41 Add cmux session move for Claude sessions (manaflow-ai#14959)
36785b1 Hide decorative Settings sidebar icons from VoiceOver (manaflow-ai#14989)
4c7158c Label the sound preview button and fix mistranslated action verbs (manaflow-ai#14983)
e704a77 Bound untracked paths stored in last-turn diff baselines (manaflow-ai#14980)
f073df1 Fix remote Files sidebar for names that change under NFD (manaflow-ai#14978)
5c68499 Bump bonsplit: mouse wheel scrolls the overflowed tab strip (manaflow-ai#14985)
9466dcb Keep agent resume bindings through the update-relaunch save (manaflow-ai#14971)
ef8b037 docs: take release notes from a Changelog section in each PR instead of CHANGELOG.md edits (manaflow-ai#14934)
6eddfd7 ci: skip the delta diff when main moved further than the pull request (manaflow-ai#14987)
fefcec7 ci: attribute red PR runs to the machine or the code, re-run machine failures once (manaflow-ai#14977)
c185deb Accept file drops on remote tmux mirror panes (manaflow-ai#14981)
90773c7 test: make CmuxSidebarGit probe waits event-driven (manaflow-ai#14973)
1f2dbfe ci: skip the scheduled Blacksmith cache warmers while owned pools serve PRs (manaflow-ai#14827)
2850651 docs: add a guide to customizing cmux's look (manaflow-ai#14850)
b66e365 Resolve a separate sidebar's content against its own backdrop (manaflow-ai#14841)
88a9360 UI tests: one labelled frame per action, built in CI; scripts/ui-test (manaflow-ai#14966)
20cfa78 fix(omo): resolve relative file refs in the shadow config without double-loading OpenCode config (manaflow-ai#14935)
f0e964c ci: make the aggregate app-host product the default, layers opt-in (manaflow-ai#14975)
52dce98 ci: run and register the machine-failure test (manaflow-ai#14972)
7bf48bc ci: route compile admission by kept-build distance across minis (manaflow-ai#14949)
44fa3f5 Offer cmux in Open With for Markdown, source, and text files (manaflow-ai#14968)
45c2d66 Replay the Claude session id of agents in cmux ssh (cmux-tui) panes (manaflow-ai#14906)
b4c1b31 Label icon-only chrome buttons and localize project panel text (manaflow-ai#14926)
14a6909 seed prefetch: keep the seed adopt would pick, of any seeded width (manaflow-ai#14944)
19e73d2 ci: self-calibrating warm-distance compile estimates (manaflow-ai#14932)
fa98d86 ci: redispatch focused runs the Mac failed before any test started (manaflow-ai#14963)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant