Repository navigation
fix(gateway,cron): make shutdown drain see in-flight cron work (#60432) - #60711
Merged
Merged
Conversation
/update and other shutdown paths only waited on gateway session agents, so active cron tool work was killed immediately in final-cleanup while the scheduler could still mark the job successful (#60432).
Cron jobs run through cron/scheduler.py's own ThreadPoolExecutor via a
standalone AIAgent (run_job/run_one_job), entirely outside
GatewayRunner._running_agents -- the dict _drain_active_agents() and
every other active-work check on that class reads. A gateway shutdown
(/update, /restart, and SIGUSR1 all funnel through the same stop())
could log active_at_start=0 and immediately kill tool subprocesses
while a cron job's terminal command was still running, with no wait
and no indication anything was interrupted.
Real-world impact (from the issue): a scheduled daily briefing cron
job was in flight during /update, its tool subprocess got killed
by the unconditional shutdown cleanup, and the job was never marked
failed -- it simply never completed or delivered, with no error
surfaced anywhere. A repro with a 30-minute `sleep` cron job in flight
during /update reproduced the same pattern: subprocess killed at
+0.22s of drain (active_at_start=0), the job's agent thread continued
in-process and produced a plausible-looking final response from the
truncated tool output, and the scheduler marked the run successful.
Root cause is layered, not a single line:
1. GatewayRunner._drain_active_agents() only waits on _running_agents.
Cron work was invisible to it, so drain returned instantly whenever
the only active work was a cron job.
2. Even with visibility, the shutdown's final tool-subprocess kill
(process_registry.kill_all()) is a global, unconditional sweep with
no per-job targeting -- a long-running cron job that outlives the
drain timeout still gets its subprocess killed.
3. cron/scheduler.py had no way to detect that a job's tool subprocess
was killed out from under it mid-run; the agent thread kept going
and its eventual (often degraded but plausible-looking) response
got reported as a normal successful completion.
Fix, three parts:
- cron/scheduler.py: expose get_running_job_ids() (thread-safe
snapshot of the existing _running_job_ids set, already used to
prevent double-dispatch) so the gateway can read cron's in-flight
state without reaching into private module internals.
- gateway/run.py: GatewayRunner._active_cron_job_count() reads that
snapshot. _drain_active_agents() now waits on
(_running_agents OR active cron jobs), so a cron-only workload gets
the same bounded wait chat sessions already get instead of an
instant active_at_start=0. Shutdown drain logging gains
cron_active_at_start/cron_active_now fields alongside the existing
ones (unchanged, for compat).
- cron/scheduler.py: mark_running_jobs_interrupted(reason), called by
gateway/run.py's _kill_tool_subprocesses() right after
process_registry.kill_all(), marks every job still in
_running_job_ids at that instant as failed/interrupted via the
existing mark_job_run() -- and records the job IDs in
_interrupted_job_ids BEFORE writing, so run_one_job()'s own
eventual completion for the same run (racing in its own thread)
checks that flag and skips its normal write instead of clobbering
the interrupted status with a false "ok" produced from the
now-truncated tool output. This does not attempt to correlate a
killed PID to a specific job ID (process_registry tracks PIDs, not
job IDs) -- any job still dispatched at the moment of a forced kill
is treated as interrupted, matching the existing coarser precedent
set by _interrupt_running_agents(), which interrupts every entry in
_running_agents on a drain timeout without per-agent correlation
either.
Deliberately out of scope (flagged in the issue as a separate,
lower-priority concern): startup-time reconciliation of cron runs that
started but never reached a terminal status.
Testing:
- tests/cron/test_shutdown_interrupt.py (12 tests): get_running_job_ids
snapshot semantics, mark_running_jobs_interrupted marking/no-op/
partial-failure behavior, and -- the core race guard -- run_one_job
skipping its own last_status write (both the success path and the
exception path) when the shutdown path already marked the run
interrupted, with a control test proving ordinary un-interrupted
completions are unaffected.
- tests/gateway/test_cron_active_work_drain.py (9 tests):
_active_cron_job_count reading cron state and failing closed (0) if
the cron module is unavailable; _drain_active_agents waiting for an
in-flight cron job the same way it waits for chat sessions, timing
out if the job outruns the window, and leaving existing chat-session
drain behavior unchanged; a full runner.stop() integration test
(drain-timeout path) proving mark_running_jobs_interrupted actually
fires with the right job ID when a tool subprocess is force-killed,
plus a no-op control when nothing cron-related is in flight.
- tests/gateway/test_shutdown_cache_cleanup.py: added
_active_cron_job_count() to that file's hand-rolled _FakeGateway test
double, which stop() now calls -- without it those 8 pre-existing
tests AttributeError (caught by fail-then-pass below, not a
production bug).
Fail-then-pass: reverted gateway/run.py + cron/scheduler.py, all 21
new tests fail (fixture/attribute errors -- the feature doesn't exist
yet); restored, all 21 pass.
Regression check: ran the full plausibly-affected surface --
tests/gateway/{test_gateway_shutdown,test_restart_drain,
test_restart_notification,test_restart_redelivery_dedup,
test_restart_resume_pending,test_restart_service_detection,
test_shutdown_cache_cleanup,test_stuck_loop,test_clean_shutdown_marker,
test_external_drain_control,test_session_state_cleanup,
test_update_command,test_update_streaming}.py plus tests/cron/ (944
tests) -- against a clean upstream/main checkout and against this
branch. Diffed the two FAILED lists: identical, 20 pre-existing
failures on both sides (Windows-locale/cp1252 file-encoding issues and
Unix-permission-bit assertions that don't apply on this Windows dev
box), zero new failures, zero fixed-by-accident. The 8
test_shutdown_cache_cleanup.py failures found mid-development were
from the _FakeGateway gap above, fixed in the same commit and
confirmed clean on the final rerun (diff against baseline: exit 0).
Fixes #60432
Follow-up to the previous commit on #60432. The status-write guard (_consume_interrupted_flag, checked right before mark_job_run) closes the false-success bookkeeping gap, but run_one_job delivers its result BEFORE that check: delivery happens right after run_job() returns, mark_job_run happens at the very end. A job whose tool subprocess was killed mid-flight can still produce a plausible-looking final_response from the truncated output, and that response would reach the user via _deliver_result before the interrupted flag was ever consulted -- correct status in jobs.json, wrong message already sent. Adds _is_interrupted(), a non-destructive peek at the same _interrupted_job_ids set (_consume_interrupted_flag stays as the consuming, authoritative check right before the status write -- this needed a peek instead since the flag has to still be visible there). Checked right after save_job_output, before the deliver_content decision: if the run looked successful but was flagged interrupted, force success=False with an explicit interruption message. This routes delivery through the existing _summarize_cron_failure_for_delivery path (the same one a real failure already uses) instead of the raw final_response, so the user gets an honest "this run was interrupted" instead of a truncated/misleading result. Testing: 4 new tests in tests/cron/test_shutdown_interrupt.py -- _is_interrupted peek semantics (false/true/does-not-clear, as opposed to the consuming _consume_interrupted_flag), and the delivery-gate test itself, which mocks run_job to return a normal-looking success with a "plausible final response" while the job is pre-marked interrupted, and asserts _deliver_result receives the failure summary ("This run was interrupted.") instead, with the summarizer's error argument confirmed to mention the interruption. Fail-then-pass: reverted cron/scheduler.py only, the 4 new tests fail (3 on the missing _is_interrupted attribute, 1 -- the delivery-gate test -- on _summarize_cron_failure_for_delivery never being called, i.e. the raw response would have gone out); restored, all 16 tests in the file pass. Regression: tests/cron/ (683 tests) + test_cron_active_work_drain.py + test_gateway_shutdown.py + test_shutdown_cache_cleanup.py -- 11 pre-existing failures (Unix file-permission-bit and path-tilde assertions that don't apply on this Windows dev box), matching the same set already established as pre-existing in the prior commit's regression check. Zero new failures. Continues #60432
Related: maintainer salvage of the union of #60612 (@HexLab98, drain-visibility) and #60631 (@JoaoMarcos44, superset with interrupted-status + delivery gate), fixing #60432. Note there are now TWO teknium1 salvages of #60432 in the same window — this PR (#60711, salvages both sources) and the earlier #60708 (04:40, salvages #60631 only). A maintainer should pick one. Also relates to #60690 (a P2 different-mechanism |
This was referenced Jul 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Gateway shutdown (
/update,/restart, SIGUSR1) now sees in-flight cron work: the drain waits for running cron jobs (they execute on the scheduler's own thread pool, invisible to_running_agents), and a job whose tool subprocess gets force-killed anyway can never be recorded — or delivered — as a success. Fixes #60432.Salvages the union of PR #60612 by @HexLab98 (earliest, drain-visibility + tests) and PR #60631 by @JoaoMarcos44 (superset: drain-visibility + interrupted-status guard + delivery gate), both authorships preserved.
Changes
cron/scheduler.py:get_running_job_ids()thread-safe snapshot;mark_running_jobs_interrupted()called by the gateway right afterprocess_registry.kill_all();_is_interrupted()peek forces the failure/honest-delivery path inrun_one_job;_consume_interrupted_flag()stops the job's own racing thread from clobbering the interrupted status with a false "ok".gateway/run.py:_active_cron_job_count()(import-guarded for minimal test doubles) folded into_drain_active_agents()wait/timeout logic; shutdown forensics log gainscron_at_start/cron_now; kill path marks in-flight jobs interrupted.cron_jobs_in_flight()helper so there is a single read surface; retargeted its drain tests atget_running_job_ids().tests/gateway/test_update_cron_drain.py(drain waits / times out / kill_all ordering),tests/gateway/test_cron_active_work_drain.py,tests/cron/test_shutdown_interrupt.py(17 tests on the interrupt primitives + race guard).Validation
/updatewith cron job mid-tool-callactive_at_start=0, drain 0.00s, subprocess killedE2E-verified with real imports (isolated HERMES_HOME): snapshot, interrupt marking, peek/consume semantics,
_active_cron_job_count()on a bare GatewayRunner. Targeted suites green includingtest_restart_drain.py.Infographic