Repository navigation
fix(ops): judge the canary install by the SHA on disk, not npm's exit code - #10699
Conversation
… code `npm install -g <tarball>` on the .17 gateway writes the whole package and then fails renaming the old tree into its staging directory (ENOTEMPTY, exit 217). The canary read that non-zero exit as "install failed", aborted before the restart, and discarded npm's stderr through execFileSync throwing — so on 2026-08-18 the deploy stopped half-done twice, each time leaving new files on disk under an old running process, with no clue in the log. The exit code is not trustworthy in either direction: the 2026-08-14 outage installed a package built from the wrong branch and exited 0. classifyInstallOutcome() therefore decides on the BUILD_SHA read back from the installed package, and fails closed when it is absent or does not match — a zero exit with the wrong artifact is still a failure. npm reuses the same staging directory name, so the orphan blocks the next install with the same error; orphanStagingDirFromStderr() surfaces the exact path. It is not removed automatically — that is an rm -rf under /usr/lib, not something a deploy script should decide on its own. Refs #10429
Review: logic is sound — one wiring gap found and fixedI applied the PR to the current The gap —
|
|
The 7 red checks are inherited base-red, not this PR — proven, not assumed. Both this PR and # fail on the exact same 7 jobs despite having no file in common. Reproduced on the pure base tip The failures are in areas this diff does not touch: Antigravity cloudcode envelope + public models, GLM provider import (4 tests), free-model catalog qwen-web ids, i18n key drift (#6695), Tracked under #9985. |
… code (diegosouzapw#10699) `npm install -g <tarball>` on the .17 gateway writes the whole package and then fails renaming the old tree into its staging directory (ENOTEMPTY, exit 217). The canary read that non-zero exit as "install failed", aborted before the restart, and discarded npm's stderr through execFileSync throwing — so on 2026-08-18 the deploy stopped half-done twice, each time leaving new files on disk under an old running process, with no clue in the log. The exit code is not trustworthy in either direction: the 2026-08-14 outage installed a package built from the wrong branch and exited 0. classifyInstallOutcome() therefore decides on the BUILD_SHA read back from the installed package, and fails closed when it is absent or does not match — a zero exit with the wrong artifact is still a failure. npm reuses the same staging directory name, so the orphan blocks the next install with the same error; orphanStagingDirFromStderr() surfaces the exact path. It is not removed automatically — that is an rm -rf under /usr/lib, not something a deploy script should decide on its own. Refs diegosouzapw#10429 Co-authored-by: Xiangzhe <bakryun0718@proton.me>
… code (diegosouzapw#10699) `npm install -g <tarball>` on the .17 gateway writes the whole package and then fails renaming the old tree into its staging directory (ENOTEMPTY, exit 217). The canary read that non-zero exit as "install failed", aborted before the restart, and discarded npm's stderr through execFileSync throwing — so on 2026-08-18 the deploy stopped half-done twice, each time leaving new files on disk under an old running process, with no clue in the log. The exit code is not trustworthy in either direction: the 2026-08-14 outage installed a package built from the wrong branch and exited 0. classifyInstallOutcome() therefore decides on the BUILD_SHA read back from the installed package, and fails closed when it is absent or does not match — a zero exit with the wrong artifact is still a failure. npm reuses the same staging directory name, so the orphan blocks the next install with the same error; orphanStagingDirFromStderr() surfaces the exact path. It is not removed automatically — that is an rm -rf under /usr/lib, not something a deploy script should decide on its own. Refs diegosouzapw#10429 Co-authored-by: Xiangzhe <bakryun0718@proton.me>
… code (diegosouzapw#10699) `npm install -g <tarball>` on the .17 gateway writes the whole package and then fails renaming the old tree into its staging directory (ENOTEMPTY, exit 217). The canary read that non-zero exit as "install failed", aborted before the restart, and discarded npm's stderr through execFileSync throwing — so on 2026-08-18 the deploy stopped half-done twice, each time leaving new files on disk under an old running process, with no clue in the log. The exit code is not trustworthy in either direction: the 2026-08-14 outage installed a package built from the wrong branch and exited 0. classifyInstallOutcome() therefore decides on the BUILD_SHA read back from the installed package, and fails closed when it is absent or does not match — a zero exit with the wrong artifact is still a failure. npm reuses the same staging directory name, so the orphan blocks the next install with the same error; orphanStagingDirFromStderr() surfaces the exact path. It is not removed automatically — that is an rm -rf under /usr/lib, not something a deploy script should decide on its own. Refs diegosouzapw#10429 Co-authored-by: Xiangzhe <bakryun0718@proton.me>
Refs #10429.
Problem
npm install -g <tarball>on the .17 gateway writes the whole package and then fails renamingthe old tree into its staging directory:
Exit status 217 — but
dist/BUILD_SHA, the package version and every dependency are the newones. The canary read the non-zero exit as "install failed", aborted before the restart, and
execFileSyncthrew away npm's stderr, so the log said onlyCommand failed: ssh … npm install.This stopped the deploy half-done twice on 2026-08-18, each time leaving the host with the
new files on disk under the old running process — the exact split state a deploy script exists
to prevent. Both times it took a manual SSH round-trip to discover the install had actually
succeeded.
Why not just tolerate ENOTEMPTY
Because the exit code is untrustworthy in both directions. The 2026-08-14 outage (#10427) went
the other way:
npm install -gexited 0 while shipping a package built from the wrongbranch, and everything downstream believed it.
So neither "non-zero = failed" nor "zero = fine" holds.
classifyInstallOutcome()decides on theBUILD_SHA read back from the installed package:
installed-with-cleanup-failure→ continue, warninstalledfailedfailedfailed(fails closed, same rule as the provenance gate)Staging directory
npm reuses the same staging name, so the orphan blocks the next install with the identical
error — that is why it recurred.
orphanStagingDirFromStderr()extracts the exact path and therun warns with it.
It is deliberately not removed automatically: that is an
rm -rfunder/usr/lib, and adeploy script should not decide that on its own.
Validation
tests/unit/canary-install-outcome-10429.test.ts— 7 cases, written first and confirmed failing(
does not provide an export named 'classifyInstallOutcome'), including the real 2026-08-18stderr verbatim and the inverse trap (zero exit + wrong artifact must still fail).
Sibling suites green:
deploy-canary-10429,build-sha-provenance-10427(25 tests total).typecheck:coreclean. CLI exercised with--dry-run, which also caught a wiring bug before itshipped:
plandoes not exposebuildSha, so the expected SHA had to come fromreadBuildSha()— left as-is it would have failed every install.