Skip to content

fix(gateway): drain timeout=0 must still wait for in-flight cron/API work - #82282

Closed
JonthanaHanh wants to merge 1 commit into
NousResearch:mainfrom
JonthanaHanh:fix/drain-cron-zero-timeout
Closed

JonthanaHanh wants to merge 1 commit into
NousResearch:mainfrom
JonthanaHanh:fix/drain-cron-zero-timeout

Conversation

@JonthanaHanh

Copy link
Copy Markdown

Summary

When restart_drain_timeout is 0 (the default), the drain phase returns immediately with timed_out=True — even when cron jobs or API-server runs are the only active work. Cron jobs live outside _running_agents and have no pre-stop protection via restart_after_turn_timeout, so the zero-budget drain kills them with no grace period.

Root Cause

In _drain_active_agents() (gateway/run.py:9381), the check if timeout <= 0: return snapshot, True fires even when only background work (cron/API) is active. The early-return at line 9376 (not self._running_agents and cron_count == 0 and api_count == 0) correctly falls through when cron is running, but the zero-timeout check immediately kills it.

The design intent of restart_drain_timeout=0 is "don't wait for chat sessions" (they are protected by the pre-stop restart_after_turn_timeout). But cron jobs have no pre-stop protection — they run on the scheduler's thread pool, entirely outside _running_agents.

Fix

Apply a minimum drain budget (30s) when the only active work is background (cron/API) and the configured timeout is 0. Chat sessions still get the immediate drain behavior when timeout=0.

Changes

  • gateway/run.py: Add _MIN_DRAIN_FOR_BACKGROUND_WORK constant and use it when timeout <= 0 with active cron/API work.
  • tests/gateway/test_update_cron_drain.py: Add regression test test_drain_zero_timeout_still_waits_for_cron_jobs.

Fixes #82161

…work

When ``restart_drain_timeout`` is 0 (the default), the drain phase returns
immediately with ``timed_out=True`` — even when cron jobs or API-server
runs are the only active work.  Cron jobs live outside ``_running_agents``
and have no pre-stop protection via ``restart_after_turn_timeout``, so the
zero-budget drain kills them with no grace period.

Apply a minimum drain budget (30s) when the only active work is background
(cron/API) and the configured timeout is 0.  Chat sessions still get the
immediate drain behavior when timeout=0 (they are protected by the
pre-stop ``restart_after_turn_timeout`` wait).

Fixes NousResearch#82161
@teknium1

Copy link
Copy Markdown
Collaborator

Closing with credit: the timeout <= 0 short-circuit you identified is exactly the bug, and the fix merged via #86684 using the configurable per-class drain budget from #82224 (agent.cron_drain_timeout, default 30s, watchdog-clamped) rather than a fixed constant — it supersedes the _MIN_DRAIN_FOR_BACKGROUND_WORK approach here. Thanks @JonthanaHanh for the precise root-cause analysis!

@teknium1 teknium1 closed this Aug 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Gateway drain exits after 0.00s with in-flight cron job, killing it mid-run

3 participants