Skip to content

feat(a2a): return a task id for long-running peer work - #91688

Open
Adolanium wants to merge 2 commits into
NousResearch:mainfrom
Adolanium:feat/a2a-async-task-id
Open

Adolanium wants to merge 2 commits into
NousResearch:mainfrom
Adolanium:feat/a2a-async-task-id

Conversation

@Adolanium

@Adolanium Adolanium commented Aug 21, 2026 •

Copy link
Copy Markdown

What does this PR do?

Add nonblocking A2A SendMessage support. Callers can request a working task immediately and poll with a2a_get_task. Routed profile agents are included. Blocking calls and streaming keep their current behavior. Background completion preserves profile context and stops after a 24-hour ceiling.

On the forwarded-profile path, blocking SendMessage uses subprocess.run with the route timeout. When that timeout fires, Python kills the child and the answer is discarded. returnImmediately on that path waits up to 24 hours, then records the result for GetTask.

Repeated polling records each peer/tenant task outcome once. A later reply, including another clarification after new user input, remains visible in history and metrics. The implementation uses the current active-task watchdog and tool registration APIs.

Follow-up from #91687 (comment). Session wake into the originating Hermes session stays out of this PR.

Related Issue

Fixes #91687

Type of Change

  • ✨ New feature (non-breaking change that adds functionality)
  • 📝 Documentation update
  • ✅ Tests

Changes Made

  • Honor configuration.returnImmediately and legacy configuration.blocking: false on inbound SendMessage
  • Return a working task id at once and record the real result in a background waiter
  • Skip the 5-minute orphan watchdog while a waiter is still live
  • Bound that waiter to 24 hours (_BACKGROUND_WAIT_SECONDS)
  • On forwarded profiles, the nonblocking path uses that 24-hour child wait so a 240s or 300s route timeout does not kill hermes chat
  • Add return_immediately on a2a_call and a new a2a_get_task poll tool
  • Cache the peer Agent Card for 60s on call and get_task
  • Dispatch one turn per context. Later same-context tasks wait in the adapter, so the gateway busy queue never merges two tasks into one turn (reported by @Neomail2)
  • Resolve each final by its reply anchor (the task id), not the oldest pending task in the context. A late final can no longer answer a sibling task

How to Test

  1. scripts/run_tests.sh tests/plugins/test_a2a_async_tasks.py tests/plugins/test_a2a_phase23.py tests/plugins/test_a2a_plugin.py tests/plugins/test_a2a_tools_gate.py tests/hermes_cli/test_deferred_platform_client_tools.py
  2. On current main: 181 passed, 1 skipped (a Linux-only shebang test, skipped on Windows). The other tests that load the A2A plugin and the async-delegation tests also pass: 86 passed, 2 skipped.
  3. Loopback HTTP tests cover concurrent polling, repeated clarifications, and tenant separation. Completion tests check local and routed profile history and audit isolation.
  4. A loopback test sends three returnImmediately tasks into one context through the real handle_message and checks that each gets its own reply. It fails without the per-context dispatch.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run pytest tests/ -q and all tests pass
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: Windows 11

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — or N/A
  • I've updated cli-config.yaml.example if I added/changed config keys — or N/A
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — or N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide — or N/A
  • I've updated tool descriptions/schemas if I changed tool behavior — or N/A

🤖 Generated with Claude Code

@alt-glitch alt-glitch added type/feature New feature or request comp/plugins Plugin system and bundled plugins P3 Low — cosmetic, nice to have labels Aug 21, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference; please use your judgment.

  1. plugins/platforms/a2a/adapter.py:~962–995 (_wait_in_background) — one daemon thread per background task, blocked on fut.result() with no timeout at all, while the watchdog now skips any task id with a live waiter — why it matters: a hung gateway turn that never resolves its future produces an immortal combination — a blocked thread, a permanently non-terminal task, and zero watchdog recourse; many long jobs over days of uptime accumulate zombies — suggestion: bound the background wait (fut.result(timeout=<generous ceiling>), e.g. hours, then finalize as failed) or use a shared bounded executor; alternatively let the watchdog fail live tasks after a much larger second threshold.

  2. plugins/platforms/a2a/tools.py:~357–366 (a2a_get_task) — every poll re-fetches the Agent Card (_fetch_card) before the JSON-RPC call — why it matters: a model polling every few seconds turns card fetches into most of the traffic to the peer, and each fetch is another failure surface — suggestion: reuse the existing per-peer card cache (with TTL) that other call sites use, or cache within the tool for the process lifetime keyed by base URL.

  3. plugins/platforms/a2a/tools.py:~229–232 (_send_task) — for in-progress states the reply is deliberately not persisted and no metric fires; the terminal outcome lands in the store only if someone actually calls a2a_get_task until completion — why it matters: an unpolled job leaves no trace in a2a_history and undercounts inbound totals, so history quietly diverges from what actually happened — suggestion: record the accepted task (state=working) immediately with a marker, and update it on poll; or document the gap in DESIGN.md so it's a known property rather than an oversight.

  4. Nit (tools.py:~371–376): GetTask params send both id and taskId — reasonable peer-compat hedge, but a one-line comment saying which spec version expects which would stop the next reader from 'fixing' it; same for unwrap_send_message_response being applied to a GetTask result (works because both share the Task envelope — say so).

Overall: strong feature work — the inbound returnImmediately path (flag parsing incl. legacy blocking:false, background waiter, cancel-race-safe finalize), the watchdog skip, the new outbound poll tool, schema/registration tests, and complete docs all hang together; item 1 is the one real resource-lifetime question I'd settle before merge.

— reviewer-a · automated agent review (Hermes week-review)

@Adolanium

Copy link
Copy Markdown
Author

Yeah, item 1 was the one worth fixing.

The background waiter now stops after 24 hours (_BACKGROUND_WAIT_SECONDS) and records the task as failed. The watchdog still skips live waiters so a 40 minute job is not killed at 5 minutes, but a hung gateway turn cannot pin a daemon thread forever.

Item 2: there was no existing per-peer card cache. I added a 60s TTL for a2a_call / a2a_get_task. a2a_discover still fetches fresh.

Item 3 is intentional. Unpolled outbound jobs write the prompt at send time and only write the reply when a2a_get_task sees a terminal state. Noted in DESIGN.md.

Item 4 comments are in. v1 GetTask uses id, older peers used taskId. Unwrap is reused because SendMessage and GetTask both return a Task envelope.

@Adolanium

Copy link
Copy Markdown
Author

I updated this on the refactored main and pushed 100c0e1 as one commit. I fixed duplicate history and metrics from repeated polls, preserved profile context in background completion, and included routed profiles in the nonblocking path. I also covered repeated clarifications after new input and kept streaming behavior intact.

144 targeted Python tests passed through scripts/run_tests.sh across A2A async tasks, protocol/adapter/client behavior, schemas, and deferred tool registration. One Linux-only shebang test was skipped on Windows. Loopback HTTP tests cover concurrent polling, repeated clarifications, and tenant separation; completion tests check local/routed profile history and audit isolation.

@Adolanium

Copy link
Copy Markdown
Author

Follow-up from #91687 (comment)

The production report is the forwarded-profile kill path. This PR already covers that for returnImmediately. Background forwarded work uses a 24-hour child wait, so a 240s or 300s route timeout does not kill the profile process. Blocking SendMessage and streaming stay as they are.

Session wake into the originating Hermes session after it has moved on or restarted stays out of this PR.

This branch last moved on 2026-09-14. Next step is rebase onto current main, then re-run the targeted A2A tests.

P3 is the cosmetic bucket. This is dropped completed work. Maintainers, please move this PR and #91687 to P2.

@Neomail2

Copy link
Copy Markdown

Hi @Adolanium, thanks for #91688 — returning a working task for returnImmediately and polling it with GetTask is exactly the shape A2A §3.2.2 describes, and we'd rather build on it than open a competing PR.

While testing the branch at 100c0e1 we hit a reply-ownership problem that becomes much more visible once callers stop blocking on the socket:

  1. send() still resolves the oldest pending task of the context (_resolve_oldest_for_context), even though the gateway passes the owning task id as reply_to (the inbound message_id is the task id). With tasks A and B in one context, finishing B first stores B's answer on A and leaves B working. A late or unanchored final of an already-finished task can also settle the only remaining sibling.
  2. A turn that starts a background delegate_task and returns a short acknowledgement is recorded as the final answer (on_processing_complete SUCCESS -> COMPLETED). The real result, delivered by the completion wake, then has nothing to settle. The same applies to a wake that starts a second wave of delegation.
  3. Two A2A tasks of one context whose background completions (notify_on_complete processes or async delegations) finish in the same window are coalesced into one wake anchored on the first task.
  4. Each detached request starts its own 24h waiter thread with no admission cap.

We prepared a small complement on top of your head (not on main, so it stays mergeable into your branch; your commit and authorship are untouched):

  • finals resolve only the task named by reply_to / metadata.reply_to_message_id in the same context; no FIFO, no sole-pending fallback; a late final of a finished task is kept as a trace under its own id;
  • the handler return value is recorded, so a turn that yields (returns None) or that newly dispatched a background delegation keeps the task working; only fresh delegation receipts of the current execution count, so a later wave cannot be held by old receipts; streamed finals still settle (the A2A: reply text lost when gateway streams the session — message/send and message/stream return TASK_STATE_COMPLETED with empty reply #116944 stash);
  • async_delegation records the triggering message_id for A2A turns only, and exposes delegation_ids_for_message / release_reply_receipts; the async-delegation group key and the process-completion batch key add the triggering message_id for platform == "a2a" only, so one wake carries one task anchor. Other platforms keep their current grouping and batching unchanged;
  • A2A_MAX_INFLIGHT (default 8, clamp 1..64) is checked before audit, persistence or dispatch, for local and routed-profile tasks, so waiter threads are bounded; if a waiter thread cannot be started, the task is finalized FAILED immediately and its slot is released (no leaked slot).

Related open PRs we checked, so this does not duplicate them:

Evidence (local run, not CI: existing venv with Python 3.11.16 and pytest 9.1.1, not scripts/run_tests.sh; hermetic env -i, loopback only, no credentials):

  • new tests/plugins/test_a2a_task_ownership.py (16 tests): on your unmodified head 12 of its original 13 tests fail (reverse order, unanchored late final, admission, yield/stream, A2A process batch, returnImmediately + 0/1/2 delegation waves); with the complement all 16 pass (3 repeated runs). The loopback test drives the real HTTP server, handle_message, async_delegation and the real GatewayRunner._async_delegation_watcher; only the model turns and child bodies are simulated;
  • test_a2a_task_ownership.py + test_a2a_plugin.py + test_a2a_tools_gate.py: 124 passed, 1 skipped, 1 failed. The failure, test_a2a_tools_gate::test_all_five_tools_carry_the_gate (expects 5 tools, the branch registers 6 with a2a_get_task), fails identically on the unmodified head;
  • in test_a2a_plugin.py three tests that encoded FIFO now pass the task anchor, which is what the gateway does; one is renamed ..._resolve_by_task_anchor.

Not covered / limits: full suite and GitHub CI not run; TaskStore stays in memory (a restart still loses active tasks); no exactly-once claim for SendMessage (§3.3.1 only makes idempotency optional); a notify_on_complete process that finishes before the turn returns is only seen through process liveness; executions of one task are assumed to be serialized by the gateway session. A2A_MAX_INFLIGHT follows the plugin's existing env-var style (A2A_REPLY_TIMEOUT, ...); happy to move it to config if preferred.

The patch is ~250 changed lines in 4 files plus one new test file. Would you prefer we send it as a PR against your branch, or would you rather cherry-pick it? If maintainers prefer it on main instead, we can rebase — just say which.

@Neomail2

Copy link
Copy Markdown

Follow-up: the complement described above is available as a single commit on top of your head 100c0e1, so you can review or cherry-pick it without anything else changing:

No PR opened yet; happy to open one against feat/a2a-async-task-id if you prefer that to a cherry-pick.

Adolanium added a commit to Adolanium/hermes-agent that referenced this pull request Sep 28, 2026
With returnImmediately, a peer can have several tasks in flight on one
context. The adapter dispatched them all into the gateway, whose busy queue
merges queued text for a session: three tasks became two turns, task 2 got
task 3's answer, and task 3 stayed WORKING until the 24h waiter ceiling.
send() also resolved the oldest pending task of the context, so a late
final could answer a sibling.

The adapter now dispatches one turn per context and hands the context on
when the in-flight task is popped. send() resolves the task named by the
final's reply anchor (the gateway anchors on the inbound message id, which
is the task id), falling back to the in-flight task only when no anchor is
present.

Reported by Neomail2 on NousResearch#91688.

Co-authored-by: Neomail2 <120675713+Neomail2@users.noreply.github.com>
@Adolanium

Copy link
Copy Markdown
Author

@Neomail2 thank you for this. You tested the branch properly, wrote it up clearly, and checked it against the other open PRs. That saved me a lot of time, and the ownership bug is real.

I pushed 8b52861 to this branch. It takes part of the complement and fixes one case it doesn't cover.

What I found reproducing it. I sent three returnImmediately tasks into one context and let the real handle_message process them. On 100c0e1, tasks 2 and 3 arrive while task 1 is running. The gateway busy queue merges their text into one pending turn. Task 2 gets task 3's answer, and task 3 stays WORKING until the 24h ceiling. Your commit gives the same result: anchoring fixes which task gets the final, but the merged turn only ever has one anchor.

What landed (adapter only):

  • The adapter now dispatches one turn per context. Later tasks in that context wait in the adapter until the in-flight task is popped. After that, the busy queue never holds two A2A tasks to merge, and interrupt mode can't cut one task off with another.
  • send() resolves the task named by reply_to / metadata.reply_to_message_id. It falls back to the context's in-flight task only when a final has no anchor. A late final anchored on a finished task settles nothing, which covers your points 1 and 3 on the reply side.
  • Your test_a2a_plugin.py hunks are included as-is (anchored finals, the FIFO test renamed to ..._resolve_by_task_anchor). The commit credits you as co-author.
  • Two new tests fail on 100c0e1 and pass now. One is the 3-task loopback run through the real handle_message; the other checks that a late anchored final can't answer a sibling. The A2A suite passes: 201 passed, 1 skipped. I also fixed test_all_five_tools_carry_the_gate, which this PR broke by adding a2a_get_task. It now checks that every tool carries the gate instead of checking the count.

What I left out, and why:

  • The delegation-receipt and wave tracking (points 2 and 3). It adds platform == "a2a" branches in tools/async_delegation.py and gateway/run_notifications.py. The plugin rules here say a plugin must not special-case itself in core; if it needs a capability, the generic surface gets widened. It also reads process_registry private fields and wraps _message_handler in a property. The problem itself is real: a turn that starts a background delegation returns an ACK, and the ACK becomes the final. I'd rather fix that in a focused follow-up with a generic "this turn yielded to a background continuation" signal from the gateway that any adapter can use. If you want to take that on, I'll review it.
  • A2A_MAX_INFLIGHT. Per-peer rate limiting already bounds how fast waiters can be created, and non-secret settings are meant to go in config.yaml, not new env vars. If we still want a cap, it should be a platforms.a2a.extra setting in a follow-up.
  • The build_source(..., message_id=task_id) hunk. fix(a2a): carry the inbound task id into the session source #109522 already has it, and nothing on this branch depends on it now.

Thanks again. The ownership part only got fixed because you tested this carefully.

@alt-glitch alt-glitch added the needs-decision Awaiting maintainer decision before any implementation label Sep 28, 2026
Adolanium and others added 2 commits September 28, 2026 22:24
SendMessage with returnImmediately now returns a working task at once.
A background waiter records the real result so GetTask can poll it.
The watchdog skips tasks whose gateway turn is still running.
Background waiters stop after 24 hours so a hung turn cannot wait forever.

a2a_call accepts return_immediately. New a2a_get_task polls by id.
Call and get_task reuse a 60s Agent Card cache. Discover still fetches fresh.
With returnImmediately, a peer can have several tasks in flight on one
context. The adapter dispatched them all into the gateway, whose busy queue
merges queued text for a session: three tasks became two turns, task 2 got
task 3's answer, and task 3 stayed WORKING until the 24h waiter ceiling.
send() also resolved the oldest pending task of the context, so a late
final could answer a sibling.

The adapter now dispatches one turn per context and hands the context on
when the in-flight task is popped. send() resolves the task named by the
final's reply anchor (the gateway anchors on the inbound message id, which
is the task id), falling back to the in-flight task only when no anchor is
present.

Reported by Neomail2 on NousResearch#91688.

Co-authored-by: Neomail2 <120675713+Neomail2@users.noreply.github.com>
@Adolanium
Adolanium force-pushed the feat/a2a-async-task-id branch from 8b52861 to 3c8b24e Compare September 28, 2026 19:27
@Neomail2

Copy link
Copy Markdown

Thanks for taking the anchored-final/same-context part into #91688 and crediting @Neomail2. We looked at the generic continuation signal you suggested.

A small local spike can expose an opt-in, adapter-neutral continuation_pending / continuation_of in gateway event metadata when a turn dispatches a background delegation. The limited path exercised did not show transcript or existing-adapter changes; those behaviors still need broader validation. This is not a PR-ready fix: the spike covers the non-streaming handler path only; streaming, queued-first, failure/cancellation, and consumption by the A2A adapter still need a defined contract and tests. In particular, a yielded ACK must leave its task WORKING, and a later synthetic final must settle only its own root task.

Would you prefer a follow-up against your #91688 branch that includes both the generic gateway signal and its A2A consumer, or a separate generic core PR after #91688 merges? We will not submit the spike as-is; we want to match the interface and base you prefer before completing the missing paths.

@Adolanium

Copy link
Copy Markdown
Author

Thanks for checking those paths! I'd keep this as a separate follow-up PR with both the generic gateway signal and the A2A consumer. Feel free to build against the current #91688 branch, then rebase onto main once it lands. No need to wait for the merge to keep working on it.

The important part is that a delegation ACK leaves the original task WORKING, including across further delegation, and the eventual final only completes that task. Please cover streaming, queued-first execution, cancellation/failure, and late finals. The core interface should stay opt-in and adapter-neutral, with existing adapters keeping their current behavior.

We should also make the current limitation explicit before #91688 merges. If nested background delegation is part of the supported flow there, the consumer fix needs to be in place before we call that flow supported.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/plugins Plugin system and bundled plugins needs-decision Awaiting maintainer decision before any implementation P3 Low — cosmetic, nice to have type/feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: A2A long jobs should return a task id

4 participants