fix(delegate): nested orchestrators get their workers' results back — delegate_task exempt from the 420 s tool deadline; summary budget uses current prompt, not the session sum - #103486
Merged
Conversation
… no 420 s deadline on delegate_task, summary budget uses the current prompt not the session sum Two defects in the same path, both measured on the 1,393-agent refactor run. 1. A nested orchestrator (depth > 0) runs delegate_task synchronously by design: it needs its workers' results inside its own turn. But the sequential tool runner put every tool call under the generic 420 s deadline, and delegate_task was not exempt, so every batch longer than seven minutes returned "timed out after 420.0s" while the children kept running as orphans. 332 such timeouts in 234 orchestrator sessions; only 89 nested delegate_task calls in the whole run ever returned a real result. The orchestrators then spent 388 h of wall time polling: 1,526 reads of the live transcript files, 551 list actions, 242 h of explicit sleep, about $4k of API turns. delegate_task is now exempt from the sequential deadline (the batch owns its liveness: per-child heartbeats, the stale monitor, delegation.child_timeout_seconds). Live A/B, depth-1 orchestrator dispatching a 75 s leaf with the deadline set to 40 s (glm-5.3-flash via Nous): main -> "Error executing tool 'delegate_task': timed out after 40.0s", leaf result lost; branch -> orchestrator blocked 161 s and returned the leaf's LEAF_DONE_MARKER. 2. _parent_summary_char_budget computed the parent's context headroom from session_prompt_tokens, which is the running SUM of prompt tokens over every API call in the session. After a few hundred calls it exceeds any window, headroom goes negative, and every child summary collapses to the 2,000-char floor with the full text spilled to disk. All 1,393 child summaries in the run were truncated this way; the orchestrator planned from stubs. The budget now reads the last call's prompt_tokens from _last_turn_usage. Tests: delegate_task is in the exempt set and the set is narrow; budget for a long-lived parent equals the budget for a fresh parent with the same current prompt, and exceeds the floor.
૮ >ﻌ< ა ci reviewran on 0913885 — fix(delegation): summary headroom uses the aggregator's own
|
teknium1
added a commit
that referenced
this pull request
Sep 5, 2026
…inalizer has not started yet Relay's producer is pumped by the consumer thread's own event loop (ManagedLlmStream.__next__ -> run_until_complete), so the finalizer can only START inside a consumer next(). The harness blocked the consumer in _count_chunk waiting for the finalizer to start; when the loop had not reached it yet by the final chunk, that wait could never be satisfied and expired at 5 s. Reproduced ~1 run in 6 locally with a thread dump (consumer parked in Event.wait, no other thread anywhere near Relay); it took two unrelated PRs red in CI the same day (#103476, #103486). The hook now forces the ordering only when the finalizer has already started (that is the race under test), releases it otherwise, reports whether the race was forced, and the two tests repeat the stream until it was, asserting the invariant on every run. 15/15 green; with 74de0fd's agent/ change reverted the tool-call test still fails 3/3, so it keeps guarding what it pins.
This was referenced Sep 5, 2026
Collaborator
Author
…ze; unknown usage means the static ceiling, never zero context Independent review found two holes in the first fix. A parent with no usage row yet was treated as 0 tokens used, so a 190K/200K prompt received a 384K-char dynamic summary budget instead of ~4K; the budget now returns None (static ceiling only) when nothing is known. And under MoA the folded usage includes advisor prompts that are not in the parent's context, over-stating the prompt size and wrongly truncating summaries; turn_usage now records the aggregator's pre-fold prompt_tokens as _last_prompt_size_tokens and the budget reads that first. Tests (2 new): unknown usage -> None; MoA-folded and unfolded parents with the same real prompt get the same budget.
13 tasks
teknium1
added a commit
that referenced
this pull request
Sep 6, 2026
…inalizer has not started yet Relay's producer is pumped by the consumer thread's own event loop (ManagedLlmStream.__next__ -> run_until_complete), so the finalizer can only START inside a consumer next(). The harness blocked the consumer in _count_chunk waiting for the finalizer to start; when the loop had not reached it yet by the final chunk, that wait could never be satisfied and expired at 5 s. Reproduced ~1 run in 6 locally with a thread dump (consumer parked in Event.wait, no other thread anywhere near Relay); it took two unrelated PRs red in CI the same day (#103476, #103486). The hook now forces the ordering only when the finalizer has already started (that is the race under test), releases it otherwise, reports whether the race was forced, and the two tests repeat the stream until it was, asserting the invariant on every run. 15/15 green; with 74de0fd's agent/ change reverted the tool-call test still fails 3/3, so it keeps guarding what it pins.
6 tasks
ppazosp
added a commit
to useomnia/hermes-agent
that referenced
this pull request
Sep 18, 2026
* fix(skills): make omitted instructions explicit and recoverable Adapt NousResearch#98736 (2fce577) to the fork without its upstream-only repeat-view cache. Preserve linked-file selectors and recover complete sections through both per-result and aggregate budgets. Co-authored-by: Mira Solari <268252643+mira-solari@users.noreply.github.com> * fix(delegation): preserve worker context and deliver complete artifacts Adapt the current-prompt budget correction from upstream NousResearch#103486, cache-path mapping from NousResearch#103667, and tasks-only schema from NousResearch#96424. Retain the legacy call interface, preserve shared batch context, and transfer only active-profile delegation artifacts into the paired Toolbox using existing file APIs. * fix(execute-code): deliver large RPC results without replaying tools Use the existing file transport or bounded shell chunks, publish atomically, and retain dispatched results through delivery retries. Fail explicitly after exhausted delivery instead of executing the same request again. * docs(delegation): explain remote transcript refresh behavior * test(execute-code): assert transferred bytes instead of shell command order --------- Co-authored-by: Mira Solari <268252643+mira-solari@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Status: Merged into
mainon 2026-09-06 as3d5831fa59062ee94330ecb19fb4f4bc717a299a. Tracking checklist: #103563. Validation below is the recorded pre-merge evidence, not a new live run.Nested orchestrators now actually receive their workers' results:
delegate_taskis exempt from the 420 s sequential-tool deadline, and the per-child summary budget is computed from the parent's current prompt size instead of the cumulative session counter.Symptom (from the 1,393-agent refactor run). A subagent that fans out (depth ≥ 1) calls
delegate_tasksynchronously by design, since it needs its workers' results inside its own turn. Every batch longer than seven minutes came back asError executing tool 'delegate_task': timed out after 420.0swhile the children kept running as orphans. 332 such timeouts across 234 orchestrator sessions; only 89 nesteddelegate_taskcalls in the entire run returned a real result. What the orchestrators did next: 1,526 reads ofcache/delegation/live/deleg_*/task-*.log, 551action=listcalls, 242 h of explicitsleep, 388 h of wall time after the first timeout, $4,034.69 in lifetime cost across the affected orchestrator sessions (21,797 calls), including productive work; this is not a measured recovery-only cost or savings estimate. Separately, 114 of 121 inspected root summaries were truncated to the 2,000-char floor, so the root often planned from stubs.Root causes.
agent/tool_executor.py:_run_sequential_tool_execution_middlewareapplied_resolve_sequential_tool_timeout()(default_DEFAULT_CONCURRENT_TOOL_TIMEOUT_S = 420) to every non-clarifytool.delegate_taskwas not exempt, so a call that legitimately blocks for a 30-minute batch was abandoned at 7 minutes; the batch's own liveness machinery (per-child heartbeats, stale monitor,delegation.child_timeout_seconds) never got to be the arbiter.tools/delegate_tool_results.py::_parent_summary_char_budget: headroom =context_length − parent.session_prompt_tokens − reserve.session_prompt_tokensis the running sum of prompt tokens over every call in the session (agent/turn_usage.py), so after a few hundred calls it exceeds any window, headroom goes negative, and the budget returns_MIN_SUMMARY_CHARS.Change.
_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS = {"delegate_task"}; the sequential runner passesNoneas the deadline for those. Interrupts still work (the poll loop runs regardless of deadline). All other tools unchanged.parent._last_prompt_size_tokens(the aggregator’s own prompt before MoA usage folding), with current-turn usage as fallback, instead of the session sum. Unknown usage returnsNonefor the static ceiling; it is not treated as an empty context.Behavior. A nested orchestrator's
delegate_taskblocks until its batch finishes and returns the consolidated results; the generic sequential-tool deadline no longer abandons that call; interruption and delegation’s own liveness controls still apply. Child summaries are trimmed only when the parent is genuinely near its window.sleep 75, sequential deadline set to 40 s in a temp HERMES_HOMEmain:RESULT: Error executing tool 'delegate_task': timed out after 40.0s, leaf result lost (wall 52 s). Branch: orchestrator blocked 161 s and returnedLEAF_DONE_MARKERfrom the leaf.tests/tools/+tests/agent/(1,046 files)origin/main(same 4 files, identical counts) and 3 were the pre-fix shape of my own test, since rewrittenruff, footguns,git diff --checkTests added (3):
delegate_taskin the exempt set and the set is narrow; a long-lived parent (session sum 25M) gets the same summary budget as a fresh parent with the same 30k current prompt, and it exceeds the floor.Not in this PR (same lane, separate PRs): per-child completion delivery instead of batch-join (233 h of finished results sat undelivered waiting for stragglers); a
delegate_task(action="wait")so an orchestrator never needssleep+tail; the tree-wide concurrency budget.Independent review (round 2)
An independent
/reviewfound two holes in the summary-budget half (P2): a parent with no usage row yet was treated as 0 tokens used, so a 190K/200K prompt received a 384K-char dynamic budget instead of ~4K; and under MoA the folded usage includes advisor prompts that are not in the parent's context, over-stating the prompt size and wrongly truncating summaries.Fix: the budget returns
None(static ceiling only) when nothing is known, never "zero context";turn_usagerecords the aggregator's pre-foldprompt_tokensas_last_prompt_size_tokensand the budget reads that first. Tests added: unknown usage →None; MoA-folded and unfolded parents with the same real prompt get the same budget. The deadline exemption was verified by the reviewer through actual dispatch and needed no change.Infographic