Skip to content

Studio: stop leaking child processes on an abnormal exit - #8170

Merged
danielhanchen merged 61 commits into
mainfrom
studio/no-orphaned-child-processes
Aug 9, 2026
Merged

danielhanchen merged 61 commits into
mainfrom
studio/no-orphaned-child-processes

Conversation

@danielhanchen

Copy link
Copy Markdown
Member

A Windows user reported that shutting Studio down to update it left errors behind and the update would not run until they killed a stray python by hand. The chain is real and reproduces on a Windows runner:

  1. every sandboxed tool call runs under a shell wrapper (cmd /c or bash -c)
  2. _capture_process_group returned None on Windows and _kill_process_tree fell back to proc.kill(), so only the wrapper died and the payload, usually the venv's own python, kept running
  3. ensure_managed_environment_is_idle then refuses the update while any process image under <studio home>\unsloth_studio is alive: The managed Studio environment is in use by python.exe (PID 9008). Stop that process, then retry the update.

Changes

Windows process trees. taskkill /T /F in _kill_process_tree, and _capture_process_group captures the wrapper pid (tagged) rather than returning None, so a payload that outlives its wrapper is still reachable from _killpg_captured.

Console close. Closing the console window raises CTRL_CLOSE_EVENT, which Python never turns into a signal, so neither the signal handler nor atexit ran and the cooperative cleanup was skipped completely. _install_windows_console_handler runs _graceful_shutdown for close, logoff and shutdown events, and passes Ctrl+C and Ctrl+Break through to the existing signal path rather than shutting down twice.

Desktop updater. suspend_for_update_installer cleared KILL_ON_JOB_CLOSE and nothing ever restored it, so from that point the app ran with no reaper. It now drains the job first, since clearing the flag removes the last backstop and anything the cooperative stop missed would outlive the app and hold the venv open against the installer. A new resume_desktop_update_cleanup command re-arms the flag when the install fails or is cancelled, called from the finally in use-tauri-update.

macOS. No PR_SET_PDEATHSIG, no job objects, so a crash or a force quit left llama-server and the other sidecars running. Tracked children are now recorded under the studio home and swept at the next startup, before anything new spawns. llama-server is tracked through adopt_pid/forget_pid alongside its existing pidfile. Ownership is decided by start-time identity rather than pid alone, because a pid is recycled quickly enough on a busy machine to make a dead owner look alive.

Visibility. _install_windows_job failed silently in four places. The outcome is recorded and logged, and readable through windows_job_status().

Also fixes a real bug found while testing this: the liveness probe used os.kill(pid, 0), which on Windows is TerminateProcess(handle, 0), so it killed the process it was asking about. It uses OpenProcess plus WaitForSingleObject now, treating ACCESS_DENIED as alive.

Tests

test_orphaned_children.py and test_child_lifetime_boundary.py, driving real processes rather than mocks. They cover the shell payload dying with its wrapper, the update gate firing on a single orphan, the console handler, the job status, the record and its sweep, a live owner never being reaped, and the liveness probe not killing its subject. The boundary tests pin both sides on Windows: with the job in force a grandchild is reaped, without it the grandchild survives, which is what the updater drain exists to prevent.

Validated on ubuntu-latest, windows-latest and macos-14, plus cargo check --all-targets and the three windows_job tests on a real Windows toolchain. End to end against an installed Studio: kill the server outright, and where macOS previously left llama-server running, the next launch logs Reaped 1 orphaned child process(es) left by a previous Studio and nothing survives.

danielhanchen and others added 2 commits August 8, 2026 11:51
A Windows user could not update Studio until they killed a stray python by hand:
the tool sandbox runs its payload under a shell wrapper, the kill path reaped
only the wrapper, and `unsloth studio update` then refused to run because a
process still held the managed environment.

- taskkill /T on Windows, so a tool payload cannot outlive its wrapper
- a console-close handler, since CTRL_CLOSE_EVENT never becomes a Python signal
  and the graceful shutdown was skipped entirely when the window was closed
- the desktop updater drains the app job before standing down crash cleanup, and
  re-arms it when the install never happens
- children are recorded on disk and swept at the next startup, which is the only
  reaper macOS has after a crash or a force quit
- the job status is logged instead of failing silently

Also fixes a liveness probe that used os.kill(pid, 0); on Windows that is
TerminateProcess, so it killed the process it was asking about.
chatgpt-codex-connector[bot]

This comment was marked as resolved.

…d per owner, drop the job drain

- the console handler ran _signal_handler on the thread Windows creates for the
  event, where signal.signal raises, so closing the window did no cleanup at all
- bound that work to the ~5s Windows allows before it kills the process
- one record file per owner pid: two Studios can share a home, and a single file
  let the second erase the first's children
- add a Windows process identity (creation time) and refuse to signal a pid that
  cannot be verified
- drop the whole-job drain: it would also terminate the WebView2 hosts, and
  cleanup_child_processes already taskkills the backend tree
A record that is not an object, or whose children are not dicts, raised out of
the startup sweep and would have stopped Studio from starting. Pair the Linux
start time with the command name as well, since start time alone has 10ms
granularity.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5185a542af

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

if sys.platform == "darwin":
assert survived, "macOS now reaps orphans -- update this repro"
else:
assert not survived

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Do not assert Linux reaps shell grandchildren

On Linux this new contrast test fails for the same shell-wrapped shape it builds: child_popen_kwargs() installs PR_SET_PDEATHSIG only in the direct bash wrapper, and when _run_case() SIGKILLs the parent, bash dies but the Python payload it spawned remains alive. Update the expectation or make the Linux path actually bind/kill the whole shell tree; otherwise the new test suite encodes a guarantee the implementation does not provide.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Checked this one and it does not reproduce: the test passes on Linux. _get_shell_cmd produces bash -c '"python" -c "..."', a single simple command, so bash execs it rather than forking. Measured it: the Popen pid and the payload's os.getpid() are the same process, and PR_SET_PDEATHSIG survives execve, so there is no grandchild for the guarantee to miss. The assertion is right as written.

@danielhanchen

Copy link
Copy Markdown
Member Author

Follow-up verification, plus a fix pushed to this branch.

What happened before, and what happens now

Before After
Shutting Studio down to update, on Windows A venv python survived, and unsloth studio update refused to run until it was killed by hand The tree is killed with taskkill /T /F, so the shell wrapper and everything under it goes
Closing the console window Windows raises CTRL_CLOSE_EVENT, which Python never turns into a signal, so no signal handler and no atexit ran and cleanup was skipped entirely A console handler runs the graceful shutdown, bounded to fit inside the roughly five seconds Windows allows before it kills the process
Studio crashes or is force quit macOS has neither PR_SET_PDEATHSIG nor job objects, so every sidecar kept running with nothing to clean them up Children are recorded as they are adopted, and the next start sweeps anything left by a Studio that is gone
The updater suspend_for_update_installer() turns off kill-on-close and nothing turned it back on The cleanup is resumed from the frontend finally, so the guarantee comes back whether or not the update proceeds

Real or not

Real, and both sides were proved on a Windows runner rather than argued: with kill-on-close in force the child did not survive the parent, and with it disabled (which is exactly what the updater path does) it did. windows_job.rs was the only place that turned it off.

Does merging break anything

No new hardware path, and none of this is GPU or device specific. The risk in a reaper is killing something it should not, so that is what the tests are aimed at:

  • A child is only signalled when its recorded start-time identity still matches, so a recycled pid is never touched. There is a test that deliberately points a stale record at an unrelated live process and asserts it survives.
  • Two Studios can share a home on different ports. Each keeps its own record file, and a start never touches a running Studio's children. A Studio that crashed while a sibling was running is still cleaned up.
  • On Windows there is no identity source in the record, so the sweep declines to signal rather than guess, and the job object covers that case already.
  • The Windows liveness probe uses OpenProcess and WaitForSingleObject, not os.kill(pid, 0), which on Windows is TerminateProcess and would kill the process it is asking about. There is a test that the probe leaves its subject running.
  • Old installs upgrading in place: there is no run/children directory yet, so the sweep returns nothing and startup proceeds normally.

Fix in this push

A record that was not an object, or whose children were not dicts, raised out of the startup sweep. Since the sweep runs before the server binds, a record truncated by a power cut or written by an older build would have stopped Studio from starting. Type guards added, with the malformed shapes as test cases. The Linux identity now pairs the start time with the command name as well, since start time alone has 10ms granularity.

Testing

  • 51 simulations: an end-to-end orphan reproduction (spawn a sidecar, SIGKILL the parent, confirm the orphan survives, then confirm the next start reaps it), repeated sweeps being idempotent, corrupt and legacy records, the macOS identity path driven for real, and the Windows ctypes paths driven through an injected kernel32 so the branch logic is checked rather than assumed.
  • Backend suite: 18918 passed, with a failure set byte-identical to unmodified main on this machine.
  • Staging CI on Windows, macOS and Linux, plus Tauri, UI, API, GGUF and startup profile: all green. The one red workflow fails the same way on staging main with no PR applied.

danielhanchen and others added 3 commits August 8, 2026 13:59
Ctrl+C on Windows raised UnboundLocalError inside the console callback, where
the BOOL result is then undefined, so the event could be reported as handled and
Studio would not stop. The updater no longer re-arms kill-on-close before
relaunching, which would have made the old process kill the replacement it just
started, and it resets the exit-cleanup guard so a retry after a failed
installer still reaps the backend. The RAG embedder and cloudflared are recorded
like the other sidecars, a llama-server that survived a failed kill stays
recorded, the Windows kill path checks the captured creation time before
taskkill, and record writes are serialised.
chatgpt-codex-connector[bot]

This comment was marked as resolved.

danielhanchen and others added 2 commits August 8, 2026 14:44
The delayed Windows kill skips a captured pid it cannot verify, since the job
object still takes the tree at exit. An owner whose identity cannot be read
counts as live rather than gone, so a momentary ps failure no longer costs a
running Studio its sidecars, and that lookup pins TZ so a timezone change does
not read as a different process. A child that outlived terminate_all keeps its
record instead of losing the only handle on it, with zombies told apart from
survivors. The whole relaunch handoff is inside the recovery scope, so any path
that leaves this process running re-arms cleanup.
chatgpt-codex-connector[bot]

This comment was marked as resolved.

terminate_all now applies the same test the startup sweep does: a pid whose
identity cannot be read is left alone and kept in the record for the next launch
to retry, rather than signalled on the chance it is still ours. The startup
reaper re-checks liveness before dropping a record, so a kill that did not take
stays reapable, and the breadcrumb unlink happens under the record lock so a
concurrent adopt cannot have its record deleted from under it.
@danielhanchen

Copy link
Copy Markdown
Member Author

@codex review

1 similar comment
@danielhanchen

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6755e699d4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread studio/backend/run.py
# crashed left its sidecars running. Sweep before spawning anything: a
# leftover holds VRAM, a port, and the files an update has to replace.
try:
reaped = reap_recorded_children()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reap the backend after an uncatchable desktop exit

When the macOS/Linux desktop app is force-killed with SIGKILL, the new sweep never fixes the leak: studio/src-tauri/src/process.rs only makes the Python backend a process-group leader, so killing Tauri does not terminate that backend, and this sweep runs inside the still-alive backend rather than the replacement desktop process. Even if another backend starts, utils/process_lifetime.py::_reap_one_record explicitly skips records whose Python owner is alive, leaving the backend and its llama/cloudflared children consuming the port and GPU after the Force Quit scenario this change is intended to cover. Bind the backend itself to the desktop parent's lifetime or have desktop startup identify and terminate the abandoned owned backend.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Checked this and I do not think it is a gap. A backend outliving a force-quit desktop is deliberate: the app writes owner metadata (token, backend_pid, port) to disk, and preflight.rs probes and verifies that on the next launch, then adopt_verified_backend reclaims it. Binding the backend to the desktop's lifetime would break adoption, which exists so a desktop restart does not have to reload the model. The sweep skipping records whose owner is alive is the same deliberate rule that stops a second Studio killing a running one's sidecars, and there is a test for it. The residual is a backend running until the next launch reclaims it, which is the design rather than something this PR left behind.

Linux startup probes prctl with the read-only PR_GET_PDEATHSIG before reporting
the parent-death signal as in force, so a seccomp or container policy that
blocks it is reported as such rather than as a guarantee nothing keeps. The
backstop sweep takes its snapshot under the lock the writes already hold.
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 8, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 8, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 8, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 8, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 9, 2026
chatgpt-codex-connector[bot]

This comment was marked as resolved.

danielhanchen and others added 2 commits August 9, 2026 10:25
The leader-only fallback runs when the job object was unavailable, which is
exactly when the record is the only handle on those workers. Both callers
read the dead leader as the tree being gone and dropped it, so anything that
survived became unreachable. The tree kill now reports whether it took, and a
failure keeps the pid tracked and its record on disk for the next launch.
chatgpt-codex-connector[bot]

This comment was marked as resolved.

The installer announced its validation server as stopped as soon as the group
leader exited, which drops the record while a child that ignored the SIGTERM
is still holding the GPU. It now waits for the group to empty, escalates to
SIGKILL, and only announces the stop once nothing is left.

An installer timeout killed the installer alone and left the announced server
for a sweep that never runs while this process lives; those children are now
terminated with it.

The diffusion group id is kept from the spawn, so a shim that exited before
the kill path (a failed health check, a crash before a reload) no longer
leaves its visual server with nothing able to reach it.
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 9, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 9, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 9, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 9, 2026
chatgpt-codex-connector[bot]

This comment was marked as resolved.

danielhanchen and others added 5 commits August 9, 2026 11:12
# Conflicts:
#	studio/src-tauri/src/main.rs
The timeout path took them; a nonzero exit or a stream that ended mid-line
left them running. This process stays up after an update failure, and its own
live record shields those pids from a sweep that would not run anyway, so the
cleanup now happens in the finally that covers every way out.
chatgpt-codex-connector[bot]

This comment was marked as resolved.

terminate_pid falls back to the recorded process group when the leader has
already exited, keeps the record when a Windows tree kill could not be
confirmed, drains the announced children under a lock so the watchdog and the
reader thread cannot race, and the installer arms the parent-death signal on
the validation server it puts in a session of its own.
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 9, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Aug 9, 2026
chatgpt-codex-connector[bot]

This comment was marked as resolved.

…dentity before a single-pid kill

The desktop stops this backend by signalling its process group and force-kills
it five seconds later, so a runner in a session of its own survives a backend
that is slow to shut down and keeps the GPU until the next launch sweeps it.
Put it back in that group, as the component installer already is, and reach the
visual server by walking the runner's children instead of killpg. The cached
group id goes with it: a pid is reusable once nothing holds the number as a
process group any more, so an id kept past its group eventually names a stranger.

terminate_pid signalled on the pid alone. An announced validation server can
exit without the line that clears it, so run the same identity test terminate_all
does before either termination branch.
chatgpt-codex-connector[bot]

This comment was marked as resolved.

The server was started in a session of its own everywhere, but only Linux can
pair that with a parent-death signal. On macOS it left the group Studio
force-kills while the only record of it is the announcement the backend has yet
to read, so a kill in that window orphaned it with nothing able to find it.

It stays in the inherited group there, and the kill path only reaches for killpg
when the server actually leads a group, so a shared group is never signalled.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. You're on a roll.

Reviewed commit: aefd7e8a68

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant