Skip to content

Finish the destroy work where a Cloud machine is first found gone - #15359

Merged
teamleaderleo merged 5 commits into
manaflow-ai:mainfrom
teamleaderleo:fix/vm-provider-gone-destroy
Sep 28, 2026
Merged

teamleaderleo merged 5 commits into
manaflow-ai:mainfrom
teamleaderleo:fix/vm-provider-gone-destroy

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Three code paths retire a Cloud machine's row when the provider reports it no longer has the machine: the status read (getVm, behind GET /api/vm/[id]), the resume preflight that every access operation runs (attach, exec, ssh, open-port, fork, resize), and the stats read. Each one wrote the row's new status and stopped there. The reconcile cron performs the same transition and also revokes the machine's model-plane route tokens and records a vm.destroyed ledger event.

That gap could not be repaired later. destroyed is terminal: findUserVm filters destroyed rows out, so a DELETE /api/vm/[id] for that machine answers 404 without running destroyVm's revoke or its ledger write, and reconciliationCandidates also filters them out, so the cron never looks at the row again. Whichever of the three paths noticed the machine first left a machine that is absent from the destroy ledger for good, with its route-token rows still marked unrevoked.

After this change all three go through one applyObservedProviderStatus, which performs the status write and, when that write lands on destroyed, does the same revoke and ledger event the cron does. The status and stats routes hand their workflows the model-plane revoker they already build for DELETE. The provider_status_read, provider_status_access and provider_status_stats reasons join the analytics allowlist, so the event says which read noticed instead of degrading to unknown.

A provider 404 does not always mean destroyed

The helper derives the row's new status itself, from observedDbStatus, rather than accepting one from the caller. That is the part worth reading closely, because getting it wrong is worse than the bug this PR set out to fix.

A 404 from the provider on a machine with a persistent home volume means the compute is gone and the volume is not. observedDbStatus, which getVm and the reconcile cron have always used, maps that case to paused rather than to a terminal status. The preflight and stats paths used to hardcode destroyed, disagreeing with both of them, which terminalized such a row and would have billed a vm.destroyed for a machine the provider never destroyed. Nothing revisits a terminal row, so neither could be taken back. An earlier revision of this PR kept that hardcoding and added the ledger write on top of it, which would have made the damage permanent and billable instead of merely permanent; the review caught it.

reopenBaseIfProviderDeleted is the one caller that overrides the mapping, through an explicit forceStatus, and the reason is in a comment at the call site: its row is a Base's active generation, and beginBaseOpen only allocates a replacement once this row has stopped being a machine the Base could open. Leaving it paused would hand the same dead provider id back on every later open, forever. The home volume is not lost by that: beginBaseOpen retains the old generation rather than deleting it, which is where a normal reset leaves it too.

What this does and does not fix

The ledger consequence is the substantive one. vm.destroyed is what cloud_vm_destroyed product analytics is built from, including machine lifetime, so machines whose disappearance was noticed by a read rather than the cron were missing from it entirely.

The credential side is narrower than it looks, and worth stating plainly rather than overselling: a route token for a row that reached destroyed was already unusable. authenticateVmAuthorization joins the token to cloud_vms and requires status in ('provisioning','running','paused'), and no write can move a row out of destroyed, so the missing revoke left an inaccurate revoked_at IS NULL row rather than a working credential. The revoke is now accurate, and it still matters for the interval before the row goes terminal. No billing or money path depends on the missing event; billing stop keys on the row's own status and destroyed_at, which these paths did already write.

Two consequences of routing the preflight through the shared mapping are worth naming, because both are behavior changes and neither is an improvement on its own.

Route tokens now outlive the observation on a volume-backed machine. authenticateRouteToken and authenticateVmAuthorization accept provisioning, running and paused, so a row that lands on paused keeps usable route tokens for the rest of their 30-day lifetime. The old hardcoded destroyed at this one call site killed them immediately. The preflight still threads no revoker, deliberately: revoking here would put one of its eight call sites at odds with getVm and the cron, which leave tokens alive for exactly the same observation, and a sleeping machine is supposed to keep its tokens. A volume-backed row that must lose its tokens has to be destroyed, by the user or by account deletion, both of which revoke.

The preflight was also the last automatic path that retired a volume-backed row whose sandbox the provider had removed. Such rows now stay paused: the cron's reconciliationCandidates sees the observed status already matches and reports unchanged, and they keep counting against a Go plan's saved-machine allowance. A user can still delete them. I left this as is because the alternative is to keep one entrypoint disagreeing with the other two about what a 404 means, which is the bug at the top of this description; the durable fix is a recovery or retirement path for volume-backed rows whose compute is gone, which is larger than this PR.

Validation

Run from web/, against a real PostgreSQL 16 with the repo's migrations applied, so the CMUX_DB_TEST=1 database cases execute rather than skip.

Red, at the tests-only commit ecaf0697f14, bun test tests/vm-workflows.test.ts:

(fail) status read that observes a gone machine > revokes the model plane and records vm.destroyed, like the cron reconcile does
error: expect(received).toEqual(expected)
- [
-   "00000000-0000-4000-8000-000000000150",
- ]
+ []

(fail) status read that observes a gone machine > an access preflight that retires the row records vm.destroyed too
error: expect(received).toEqual(expected)
- [
-   "vm.destroyed",
- ]
+ []

Red again for the home-volume defect the review found, with the tests of the final commit against the code of the one before it (fa6b69c58cf), bun test tests/vm-stats-not-found.test.ts:

(fail) stats provider missing classification > a machine with a home volume is paused rather than destroyed, and bills no destroy
error: expect(received).toEqual(expected)

@@ -3,3 +3,3 @@
    "providerVmId": "vm-fixture",
-   "status": "paused",
+   "status": "destroyed",
  }

 7 pass
 1 fail

Green at the head of this PR: bun test tests/vm-workflows.test.ts is 148 pass, 0 fail across 148 tests, with the database cases executing; bun test tests/vm-stats-not-found.test.ts is 8 pass, 0 fail.

Whole-suite, because the previous revision of this PR shipped two regressions in files I had not run. bun test over all of web/ at this branch is 3695 pass, 304 fail, 4000 tests across 370 files. At the merge base it is 3679 pass, 305 fail, 3985 tests across 369 files. Those ~300 failures are an artifact of running 370 files in parallel against one database, where the dbTest files truncate shared tables out from under each other; the number is what matters and the comparison is the point. Diffing the two runs by failing test name, no test fails on this branch that does not already fail at the merge base.

Also from web/: bun x tsc --noEmit clean; bun run lint:complexity passes with the baseline untouched at 42 grandfathered findings and no new entry.

No fleet dogfood. cmux#8029 disabled Vercel branch previews while keeping main deployments, so an unmerged change to this app has no preview URL an app build could talk to, and a fleet build would exercise main rather than this branch. The database-backed runs above are the substitute.

Changelog

Fixed: a Cloud machine the provider had already removed could be recorded as gone without its destroy being logged or its model-plane tokens marked revoked, and a sleeping machine with a saved home directory could be wrongly recorded as destroyed.

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 3 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 2f0486b9-3c24-4478-8024-c8c21e393bf9

📥 Commits

Reviewing files that changed from the base of the PR and between d2877b2 and d02fb73.

📒 Files selected for processing (7)
  • web/app/api/vm/[id]/route.ts
  • web/app/api/vm/[id]/stats/route.ts
  • web/services/vms/productAnalytics.ts
  • web/services/vms/workflows.ts
  • web/tests/vm-route-auth.test.ts
  • web/tests/vm-stats-not-found.test.ts
  • web/tests/vm-workflows.test.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

CI failure attribution

CI passes on d02fb731de (run 36456146272 attempt 1).

Written by scripts/ci/classify_failures.py (ci-failure-attribution.yml); signatures are its SIGNATURES table. A machine verdict is the runner's fault, not this PR's.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Cross-model review (Codex gpt-5.6-sol)

  • web/services/vms/workflows.ts:2669-2686 — the workflow commits status = destroyed before a separate recordUsageEvent, then explicitly swallows any event-write failure. A crash in that gap or one transient insert failure permanently loses vm.destroyed: later reads/retries exclude destroyed rows (repository.ts:3289-3304), reconciliation excludes them (:2737-2746), and the conditional status writer cannot win again (:3053-3072). The added tests cover a successful insert and a lost status race, not failure after the successful transition. Commit the terminal transition and ledger fact in one transaction, or atomically enqueue an idempotent outbox/pending record; add a stateful failure-then-reconciliation test that produces exactly one destroy event.

teamleaderleo and others added 3 commits September 28, 2026 07:55
A status read and an access preflight both move the row to `destroyed`
themselves when the provider no longer has the machine. Once either does,
`destroyVm` can never see the row again (its lookup skips destroyed rows)
and the provider-status cron skips it too, so the model-plane revoke and
the `vm.destroyed` usage event that the cron performs for the same
transition have to happen at these write sites as well.

Fails today: both retire the row and record nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Three places retire a Cloud machine's row after the provider says it no
longer has the machine: the status read, an access operation's resume
preflight, and the stats read. Each wrote `status = destroyed` and stopped
there, while the reconcile cron doing the same transition also revoked the
machine's model-plane tokens and recorded a `vm.destroyed` ledger event.

That difference was permanent, not a race to lose. `destroyed` is terminal:
`findUserVm` hides such a row from every destroy request and
`reconciliationCandidates` drops it from the cron, so whichever of the three
got there first left a machine that never appears as destroyed in the ledger
and never has its route tokens marked revoked, with nothing able to finish
the job afterwards.

All three now go through one `applyObservedProviderStatus` that performs the
write and, when the write lands on `destroyed`, the same revoke and ledger
event as the cron. The status route hands `getVm` the model-plane revoker it
already builds for delete. The new `provider_status_*` destroy reasons join
the analytics allowlist so the event says which read noticed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review found that the previous commit took the row's new status from each
caller. That let the access preflight and the stats read hardcode
"destroyed" for a provider 404, including for a machine with a persistent
home volume, which observedDbStatus maps to "paused" because the compute
is gone and the machine is not. Those rows were terminalized and billed a
vm.destroyed that never happened, and nothing revisits a terminal row to
take either back.

The status is now derived inside applyObservedProviderStatus, so every
entrypoint agrees about what a 404 means. reopenBaseIfProviderDeleted is
the one caller that must override it, and passes forceStatus with the
reason: its row is a Base's active generation, and leaving it paused would
hand the same dead provider id back on every later open.

Also threads the model-plane revoker through getVmStats from its route,
covers the stats entrypoint including the home-volume case, and fixes the
two test-file regressions the previous commit shipped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo
teamleaderleo force-pushed the fix/vm-provider-gone-destroy branch from 758ef73 to c7da06f Compare September 28, 2026 14:58
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Review: a review subagent went over 758ef732721 against main, with a real PostgreSQL so the CMUX_DB_TEST=1 cases ran instead of skipping. Verdict was request changes, with three blocking items. Two of them it executed rather than argued: the PR turned two currently-green suites red, and against the database it newly wrote permanent, false vm.destroyed ledger rows for exactly the machines its own code calls asleep rather than destroyed.

Fixed:

  • B3, the substantive one. The helper took the row's new status from each caller, and the preflight and stats paths passed destroyed unconditionally. On a machine with a persistent home volume observedDbStatus maps a provider 404 to paused, because the compute is gone and the machine is not. Those rows were being terminalized and billed a destroy that had not happened, and nothing revisits a terminal row to take either back. The status is now derived inside applyObservedProviderStatus, so every entrypoint agrees about what a 404 means. This is the reviewer's own recommended shape, and it closes S1 with it.
  • S1. reopenBaseIfProviderDeleted was still hand-rolling the same write plus revoke plus ledger sequence beside the new shared helper. It goes through the helper now, with forceStatus as the single deliberate override and the reason at the call site: its row is a Base's active generation, and leaving it paused would hand the same dead provider id back on every later open.
  • B1. tests/vm-stats-not-found.test.ts drives a Proxy repository that throws on any method it does not define, so adding a recordUsageEvent call to the stats path made that suite throw. It now stubs the method and captures the events.
  • B2. tests/vm-route-auth.test.ts pinned getVm's argument with an exact object, so threading the revoker through the status route broke it. Switched to expect.objectContaining, the form this file already uses for createVm. Note for anyone tempted by the obvious alternative: expect.any is not in this repo's bun:test typings and fails tsc --noEmit even though it runs.
  • S2. The stats entrypoint had no coverage at all. It has a case now, including the home-volume one, which is red against the previous head with exactly the paused versus destroyed diff, quoted in the description.
  • S3. getVmStats now takes the model-plane revoker and app/api/vm/[id]/stats/route.ts passes it, so the stats path revokes like the status path does.
  • N1. Corrected the goneMachine fixture comment that implied a 404 always means destroyed.
  • The review also found four false or overstated claims in the description. The description is rewritten: the "deliberately unchanged" paragraph described behavior this PR now changes, and the "no database was available" line was simply wrong on this host.

Left:

  • The access preflight still threads no revoker, and that is now argued in the description rather than asserted: seven call sites, against credentials that are already inert on a row in the state this write produces. If someone disagrees, the change is mechanical and the helper is ready for it.
  • The ~300 whole-suite failures against a single local database are not addressed here. They are pre-existing and are a harness artifact, not a defect in the code: 370 test files run in parallel and the dbTest files truncate shared tables out from under each other. I compared against a merge-base run by failing test name; nothing fails on this branch that does not already fail at the base. Worth a separate look by someone who owns the harness.

Root cause of the two regressions, stated plainly because it is the useful part: the previous revision's validation was one test file. This one ran the whole suite, twice, with a baseline to compare against.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Second review round, against c7da06fb9f8. It came back REQUEST CHANGES, and it was right. Fixed in 38ee5a03b69.

Correction to my previous comment on this PR. Its "Left" section said the access preflight threads no revoker because the credentials it would mark are "already inert ... which is exactly the state this write puts it in". That is false, and it is false precisely for the case this PR introduces. authenticateRouteToken and authenticateVmAuthorization (web/services/coderouter/repository.ts:268, :293) both accept status in ('provisioning', 'running', 'paused'). A volume-backed machine reaching that branch now lands on paused, so its route tokens stay valid for the rest of their 30-day lifetime. The same comment also said "seven call sites"; there are eight. Both claims were in the PR description too. Description and code comment are rewritten; the description now states the token-lifetime change as a change rather than burying it.

Review: blocking (1)

  • The comment at workflows.ts:2991-2994 justified omitting the revoker with a claim the PR's own headline change falsifies, with a 30-day consequence. Not a cosmetic problem: it would have been read as a proof that the state is inert.

Fixed

  • Replaced that comment. It now says what is true: paused is a live status for credentials, tokens stay valid on purpose as they do for any sleeping machine, and the revoker is omitted because revoking would put one of eight call sites at odds with getVm and the reconcile cron, which leave tokens alive for the same observation. I kept the behavior rather than threading a revoker, because agreeing with the other two readers of the same provider state is the entire point of the change.
  • Dropped the stale "mark the row so the next fleet refresh drops it" sentence, which stopped being true for paused rows in this PR.
  • "seven call sites" corrected to eight everywhere.
  • Removed the claim that a later attach can still resurrect such a row, in the code comment, the helper's doc comment and the test comments. The reviewer checked and there is no implemented recovery path: drivers/freestyle.ts maps a 404 to destroyed and resume() throws on a removed id. The comments now say what the code does (the row is not terminal) instead of predicting what will happen.
  • usageEventSource typed VmDestroySource instead of string, so a source outside the analytics allowlist fails to compile rather than degrading to unknown at runtime.

Left, disclosed in the description rather than fixed

  • This PR removes the last automatic path that retired a volume-backed row whose sandbox the provider removed. Those rows now stay paused, are reported unchanged by the cron, and occupy a Go-plan saved-machine slot until the user deletes them. Fixing it means either keeping one entrypoint disagreeing with the other two, which is the bug this PR exists to fix, or building a recovery/retirement path for volume-backed rows, which is a separate change.
  • destroyVm records vm.destroyed after an unguarded markDestroyed, so two concurrent destroys can write two ledger rows. Pre-existing through the cron and not introduced here; PostHog dedupes on insertId: cloud_vm_destroyed:<vmId> and no billing path reads the event. Noted for whoever next touches destroyVm.

Verified at 38ee5a03b69, from web/, against a real PostgreSQL 16 with migrations applied so the CMUX_DB_TEST=1 cases execute: bun test tests/vm-workflows.test.ts 148 pass / 0 fail; tests/vm-stats-not-found.test.ts 8 pass / 0 fail; tests/vm-route-auth.test.ts 89 pass / 0 fail. bun x tsc --noEmit clean. bun run lint:complexity passes, baseline untouched at 42 grandfathered findings.

The four mechanical questions I asked the reviewer all checked out: no wrong row status at any of the five call sites of the shared helper, forceStatus is safe because beginBaseOpen retains the old generation rather than deleting it and no reaper deletes the volume, and no duplicate ledger row, because markProviderObservedStatus is a conditional update that returns a row only to the caller that wins the transition.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Subagent review at c7da06f (second round): request changes, one blocking item (a code comment justified omitting the revoker with a false claim). Addressed in 38ee5a0, with the paused token behavior disclosed in the description. First round (758ef73) blockers B1 to B3 and S1 to S3 were addressed in c7da06f. 38ee5a0 is the fix for that review (comment and type changes plus one test tweak), so no re-review. Landing: rewriting the author of 38ee5a0 to the linked identity so CLA Assistant can pass, merging main after #15414, then auto-merge.

teamleaderleo and others added 2 commits September 28, 2026 13:02
The access preflight carried a comment claiming a revoke was pointless
there because the credentials were already inert. That is false for the
case this branch introduces: authenticateRouteToken and
authenticateVmAuthorization both accept `paused`, so route tokens on a
volume-backed machine stay valid after this write, for the rest of their
30-day lifetime.

Keep the behavior, which is what getVm and the reconcile cron already do
for the same observation, and replace the comment with what is true. Also
drop the stale "next fleet refresh drops it" line, correct "seven call
sites" to eight, stop asserting a resurrection path that is not
implemented, and type usageEventSource as VmDestroySource so an unknown
source cannot silently degrade in PostHog.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@teamleaderleo
teamleaderleo force-pushed the fix/vm-provider-gone-destroy branch from 38ee5a0 to d02fb73 Compare September 28, 2026 17:09
@teamleaderleo
teamleaderleo enabled auto-merge (squash) September 28, 2026 17:09
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

recheck

@teamleaderleo
teamleaderleo merged commit 5663c13 into manaflow-ai:main Sep 28, 2026
62 checks passed
@github-actions

Copy link
Copy Markdown
Contributor

Merge receipt for d02fb731de: every check was green at merge (21 verified; 16 skipped by policy). Full suite runs on main after merge.

rustybret pushed a commit to rustybret/bmux that referenced this pull request Sep 28, 2026
762c3ed Recover a Cloud machine graph stuck on an equal-cursor conflict (manaflow-ai#15328)
524ebff ci: replay the fuzz regressions on sidebar, split and window changes (manaflow-ai#15412)
818d475 Let a user's Cloud open dial even right after a background link failure (manaflow-ai#15291)
97491a7 Let the Cloud toolbar name the machine-list failure it has (manaflow-ai#15236)
0abac32 PR media: adopt CI's build only, start when CI completes, run for every app PR (manaflow-ai#15418)
7f08715 ci(seed): keep the trusted seed on the Mac before the R2 upload (manaflow-ai#15411)
5663c13 Finish the destroy work where a Cloud machine is first found gone (manaflow-ai#15359)

# Conflicts:
#	.github/workflows/ci.yml
#	.github/workflows/pr-media.yml
#	.github/workflows/seed-derived-data.yml
#	.github/workflows/test-e2e.yml
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant