Repository navigation
runner_unit: ExitType=main stops a finished incarnation; ExitType=cgroup holds it forever - #11662
Merged
Merged
Conversation
…oup holds it forever lifecycle_intent_cannot_orphan had its ExitType arms inverted. Under ExitType=cgroup a surviving process keeps the unit running after MainPID exits, so no stop is issued and KillMode=control-group never executes: srv2-02 held a cancelled job's claim_executor (16 GB) for 2.7 days, and srv1-03 / srv1-07 were held 16 days by processes a successful job left. Measured on srv1 (systemd 255) 2026-09-18 in transient units: ExitType=main reaps the child whether it honors or ignores SIGTERM; ExitType=cgroup leaves the unit active with MainPID=0 and the child alive in both cases. - predicate renamed lifecycle_intent_stops_finished_incarnation; its input is whether the main process exits when the incarnation ends (run.sh does, except its return-code-2 relaunch loop), not whether the listener is MainPID - declared intent: ExitType=main + KillMode=control-group - witnesses refuse both contracts the fleet has run (KillMode=process and the ExitType=cgroup drop-in); microVM admission refuses the latter and is stated as necessary, not sufficient, pending the sanitation gate Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… admits Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
briansrls
added this pull request to the merge queue
Sep 19, 2026
This was referenced Sep 19, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
lifecycle_intent_cannot_orphaningunbc.runner_unittreatedExitType=cgroupas the safe setting. It is the opposite: a single surviving process keeps the unitactive, withMainPID=0, and no stop is ever issued, soKillMode=control-groupnever runs.claim_executorchild under it, and the slot was held for 2.7 days with 16 GB.docker eventsandttydprocesses that a successful job left behind. The survivors have been killed and the slots have restarted.The measurement
On srv1, systemd 255, in transient units (no runner slot touched). Settings:
KillMode=control-group,SendSIGKILL=yes,TimeoutStopSec=4s. The main process forks a child and exits 0.An explicit
systemctl stopunderExitType=cgroupreaps the ignoring child in 4s. SoExitTypedecides whether a stop is issued, andKillModedecides whether that stop reaches every process in the cgroup.Change
lifecycle_intent_stops_finished_incarnation. Its input is nowmain_exits_when_incarnation_ends, replacinglistener_is_main_process, which was the wrong question.run.shexits when a job finishes, except in its return-code-2 relaunch loop, and that gap is named.ExitType=main+KillMode=control-group. Note that the host converger will report every slot asnot-effective:control-group:cgroupuntil the new drop-in is installed. That is the intended signal.KillMode=processand theExitType=cgroupdrop-in. The incident configuration is a RED in the microVM admission witness.product.fabric.sanitationCellReadiness, which microVM admission will consume.Not in this PR (same program)
CellReadiness: stop, then force-stop, recursive readback, and quarantine, gated before admission.claim_executorin the step, and have it exit cleanly on SIGINT/SIGTERM.Evidence status
I could not run the witnesses outside CI. The remote runner refused on page thrashing while resolving, so CI's witness lane is the first execution of them.
🤖 Generated with Claude Code