Skip to content

fix(delegate): nested orchestrators get their workers' results back — delegate_task exempt from the 420 s tool deadline; summary budget uses current prompt, not the session sum - #103486

Merged
teknium1 merged 2 commits into
mainfrom
fix/nested-delegate-deadline-and-summary-budget
Sep 6, 2026
Merged

teknium1 merged 2 commits into
mainfrom
fix/nested-delegate-deadline-and-summary-budget

Conversation

@teknium1

@teknium1 teknium1 commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator

Status: Merged into main on 2026-09-06 as 3d5831fa59062ee94330ecb19fb4f4bc717a299a. Tracking checklist: #103563. Validation below is the recorded pre-merge evidence, not a new live run.

Nested orchestrators now actually receive their workers' results: delegate_task is exempt from the 420 s sequential-tool deadline, and the per-child summary budget is computed from the parent's current prompt size instead of the cumulative session counter.

Symptom (from the 1,393-agent refactor run). A subagent that fans out (depth ≥ 1) calls delegate_task synchronously by design, since it needs its workers' results inside its own turn. Every batch longer than seven minutes came back as Error executing tool 'delegate_task': timed out after 420.0s while the children kept running as orphans. 332 such timeouts across 234 orchestrator sessions; only 89 nested delegate_task calls in the entire run returned a real result. What the orchestrators did next: 1,526 reads of cache/delegation/live/deleg_*/task-*.log, 551 action=list calls, 242 h of explicit sleep, 388 h of wall time after the first timeout, $4,034.69 in lifetime cost across the affected orchestrator sessions (21,797 calls), including productive work; this is not a measured recovery-only cost or savings estimate. Separately, 114 of 121 inspected root summaries were truncated to the 2,000-char floor, so the root often planned from stubs.

Root causes.

  • agent/tool_executor.py: _run_sequential_tool_execution_middleware applied _resolve_sequential_tool_timeout() (default _DEFAULT_CONCURRENT_TOOL_TIMEOUT_S = 420) to every non-clarify tool. delegate_task was not exempt, so a call that legitimately blocks for a 30-minute batch was abandoned at 7 minutes; the batch's own liveness machinery (per-child heartbeats, stale monitor, delegation.child_timeout_seconds) never got to be the arbiter.
  • tools/delegate_tool_results.py::_parent_summary_char_budget: headroom = context_length − parent.session_prompt_tokens − reserve. session_prompt_tokens is the running sum of prompt tokens over every call in the session (agent/turn_usage.py), so after a few hundred calls it exceeds any window, headroom goes negative, and the budget returns _MIN_SUMMARY_CHARS.

Change.

  • _SEQUENTIAL_DEADLINE_EXEMPT_TOOLS = {"delegate_task"}; the sequential runner passes None as the deadline for those. Interrupts still work (the poll loop runs regardless of deadline). All other tools unchanged.
  • Budget first reads parent._last_prompt_size_tokens (the aggregator’s own prompt before MoA usage folding), with current-turn usage as fallback, instead of the session sum. Unknown usage returns None for the static ceiling; it is not treated as an empty context.

Behavior. A nested orchestrator's delegate_task blocks until its batch finishes and returns the consolidated results; the generic sequential-tool deadline no longer abandons that call; interruption and delegation’s own liveness controls still apply. Child summaries are trimmed only when the parent is genuinely near its window.

Validation Result
Live A/B: depth-1 orchestrator (glm-5.3-flash via Nous) dispatches a leaf running sleep 75, sequential deadline set to 40 s in a temp HERMES_HOME main: RESULT: Error executing tool 'delegate_task': timed out after 40.0s, leaf result lost (wall 52 s). Branch: orchestrator blocked 161 s and returned LEAF_DONE_MARKER from the leaf.
tests/tools/ + tests/agent/ (1,046 files) 14,929 passed; 9 failed of which 6 are pre-existing on origin/main (same 4 files, identical counts) and 3 were the pre-fix shape of my own test, since rewritten
Targeted: sequential-interrupt, delegate, timeout-cleanup, cost-footer, output-schema, summary-budget, tool-executor suites 126 passed, 0 failed
ruff, footguns, git diff --check clean

Tests added (3): delegate_task in the exempt set and the set is narrow; a long-lived parent (session sum 25M) gets the same summary budget as a fresh parent with the same 30k current prompt, and it exceeds the floor.

Not in this PR (same lane, separate PRs): per-child completion delivery instead of batch-join (233 h of finished results sat undelivered waiting for stragglers); a delegate_task(action="wait") so an orchestrator never needs sleep+tail; the tree-wide concurrency budget.

Independent review (round 2)

An independent /review found two holes in the summary-budget half (P2): a parent with no usage row yet was treated as 0 tokens used, so a 190K/200K prompt received a 384K-char dynamic budget instead of ~4K; and under MoA the folded usage includes advisor prompts that are not in the parent's context, over-stating the prompt size and wrongly truncating summaries.

Fix: the budget returns None (static ceiling only) when nothing is known, never "zero context"; turn_usage records the aggregator's pre-fold prompt_tokens as _last_prompt_size_tokens and the budget reads that first. Tests added: unknown usage → None; MoA-folded and unfolded parents with the same real prompt get the same budget. The deadline exemption was verified by the reviewer through actual dispatch and needed no change.

Infographic

nested-delegate-results

… no 420 s deadline on delegate_task, summary budget uses the current prompt not the session sum

Two defects in the same path, both measured on the 1,393-agent refactor run.

1. A nested orchestrator (depth > 0) runs delegate_task synchronously by
   design: it needs its workers' results inside its own turn. But the
   sequential tool runner put every tool call under the generic 420 s
   deadline, and delegate_task was not exempt, so every batch longer than
   seven minutes returned "timed out after 420.0s" while the children kept
   running as orphans. 332 such timeouts in 234 orchestrator sessions; only
   89 nested delegate_task calls in the whole run ever returned a real result.
   The orchestrators then spent 388 h of wall time polling: 1,526 reads of
   the live transcript files, 551 list actions, 242 h of explicit sleep,
   about $4k of API turns. delegate_task is now exempt from the sequential
   deadline (the batch owns its liveness: per-child heartbeats, the stale
   monitor, delegation.child_timeout_seconds).

   Live A/B, depth-1 orchestrator dispatching a 75 s leaf with the deadline
   set to 40 s (glm-5.3-flash via Nous): main -> "Error executing tool
   'delegate_task': timed out after 40.0s", leaf result lost; branch ->
   orchestrator blocked 161 s and returned the leaf's LEAF_DONE_MARKER.

2. _parent_summary_char_budget computed the parent's context headroom from
   session_prompt_tokens, which is the running SUM of prompt tokens over
   every API call in the session. After a few hundred calls it exceeds any
   window, headroom goes negative, and every child summary collapses to the
   2,000-char floor with the full text spilled to disk. All 1,393 child
   summaries in the run were truncated this way; the orchestrator planned
   from stubs. The budget now reads the last call's prompt_tokens from
   _last_turn_usage.

Tests: delegate_task is in the exempt set and the set is narrow; budget for
a long-lived parent equals the budget for a fresh parent with the same
current prompt, and exceeds the floor.
@github-actions

github-actions Bot commented Sep 5, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 0913885 — fix(delegation): summary headroom uses the aggregator's own

⚠️ Warnings

OSV vulnerability scan · View job

28 known vulnerabilities found in pinned dependencies.

How to fix:

Review the findings in the Security tab. Update the affected dependencies if a patched version is available.


debug info

CI timings

CI timings · View report · View job

Wall time 4m43s vs 4m11s (+12.7%). 5 job(s) slower, 6 faster, 3 unchanged.

  • OS-specific tests / Windows-only tests: +20.0s
  • Python lints / Windows footguns (blocking): -14.0s
  • OSV scan / Emit review status: -11.0s
  • Python lints / ruff enforcement (blocking): +8.0s
  • Check contributors / check-attribution: -8.0s

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint tool/delegate Subagent delegation labels Sep 5, 2026
teknium1 added a commit that referenced this pull request Sep 5, 2026
…inalizer has not started yet

Relay's producer is pumped by the consumer thread's own event loop (ManagedLlmStream.__next__ -> run_until_complete), so the finalizer can only START inside a consumer next(). The harness blocked the consumer in _count_chunk waiting for the finalizer to start; when the loop had not reached it yet by the final chunk, that wait could never be satisfied and expired at 5 s. Reproduced ~1 run in 6 locally with a thread dump (consumer parked in Event.wait, no other thread anywhere near Relay); it took two unrelated PRs red in CI the same day (#103476, #103486).

The hook now forces the ordering only when the finalizer has already started (that is the race under test), releases it otherwise, reports whether the race was forced, and the two tests repeat the stream until it was, asserting the invariant on every run. 15/15 green; with 74de0fd's agent/ change reverted the tool-call test still fails 3/3, so it keeps guarding what it pins.
@teknium1

teknium1 commented Sep 5, 2026

Copy link
Copy Markdown
Collaborator Author

Part of the forensic post-mortem of the #102117 run; tracking issue with sources, method and the full measured table: #103563

…ze; unknown usage means the static ceiling, never zero context

Independent review found two holes in the first fix. A parent with no usage
row yet was treated as 0 tokens used, so a 190K/200K prompt received a
384K-char dynamic summary budget instead of ~4K; the budget now returns None
(static ceiling only) when nothing is known. And under MoA the folded usage
includes advisor prompts that are not in the parent's context, over-stating
the prompt size and wrongly truncating summaries; turn_usage now records the
aggregator's pre-fold prompt_tokens as _last_prompt_size_tokens and the
budget reads that first.

Tests (2 new): unknown usage -> None; MoA-folded and unfolded parents with
the same real prompt get the same budget.
teknium1 added a commit that referenced this pull request Sep 6, 2026
…inalizer has not started yet

Relay's producer is pumped by the consumer thread's own event loop (ManagedLlmStream.__next__ -> run_until_complete), so the finalizer can only START inside a consumer next(). The harness blocked the consumer in _count_chunk waiting for the finalizer to start; when the loop had not reached it yet by the final chunk, that wait could never be satisfied and expired at 5 s. Reproduced ~1 run in 6 locally with a thread dump (consumer parked in Event.wait, no other thread anywhere near Relay); it took two unrelated PRs red in CI the same day (#103476, #103486).

The hook now forces the ordering only when the finalizer has already started (that is the race under test), releases it otherwise, reports whether the race was forced, and the two tests repeat the stream until it was, asserting the invariant on every run. 15/15 green; with 74de0fd's agent/ change reverted the tool-call test still fails 3/3, so it keeps guarding what it pins.
@teknium1
teknium1 merged commit 3d5831f into main Sep 6, 2026
37 checks passed
@teknium1
teknium1 deleted the fix/nested-delegate-deadline-and-summary-budget branch September 6, 2026 19:02
ppazosp added a commit to useomnia/hermes-agent that referenced this pull request Sep 18, 2026
* fix(skills): make omitted instructions explicit and recoverable

Adapt NousResearch#98736 (2fce577) to the fork without its upstream-only repeat-view cache. Preserve linked-file selectors and recover complete sections through both per-result and aggregate budgets.

Co-authored-by: Mira Solari <268252643+mira-solari@users.noreply.github.com>

* fix(delegation): preserve worker context and deliver complete artifacts

Adapt the current-prompt budget correction from upstream NousResearch#103486, cache-path mapping from NousResearch#103667, and tasks-only schema from NousResearch#96424. Retain the legacy call interface, preserve shared batch context, and transfer only active-profile delegation artifacts into the paired Toolbox using existing file APIs.

* fix(execute-code): deliver large RPC results without replaying tools

Use the existing file transport or bounded shell chunks, publish atomically, and retain dispatched results through delivery retries. Fail explicitly after exhausted delivery instead of executing the same request again.

* docs(delegation): explain remote transcript refresh behavior

* test(execute-code): assert transferred bytes instead of shell command order

---------

Co-authored-by: Mira Solari <268252643+mira-solari@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists tool/delegate Subagent delegation type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants