Skip to content

fix(a2a): keep late replies and refuse a shared port - #130744

Open
deadczarvc wants to merge 2 commits into
NousResearch:mainfrom
deadczarvc:fix/a2a-late-reply-exclusive-port
Open

deadczarvc wants to merge 2 commits into
NousResearch:mainfrom
deadczarvc:fix/a2a-late-reply-exclusive-port

Conversation

@deadczarvc

Copy link
Copy Markdown

What does this PR do?

Two bugs in the A2A platform adapter that we hit on a Windows host with several agents talking over A2A.

1. A reply that comes after the server stops waiting is lost. When a blocking message/send runs longer than
A2A_REPLY_TIMEOUT, or a message/stream client disconnects, the task is marked TASK_STATE_FAILED. The agent
keeps working, the gateway logs the finished response, but the task is already terminal, so GetTask keeps
returning FAILED and the reply is gone. The docs said the opposite ("a late reply is stored, not discarded").

With this change the task is detached instead of failed: it stays TASK_STATE_WORKING, the caller gets that task
back, and a daemon thread records the reply when it arrives, so GetTask, ListTasks and push notifications see it.
A task that never gets a reply still fails, at the existing orphan ceiling (24 h).

2. Two profiles can share one A2A port on Windows. The server is a stdlib ThreadingHTTPServer, which sets
SO_REUSEADDR. On Windows that lets a second socket bind a port that is already listening; with multiplex profiles
we saw two gateways on 9902 and requests split between them. The server now binds with SO_EXCLUSIVEADDRUSE on
Windows, so the second bind fails loudly (the existing bind_failed path). Elsewhere the stdlib behaviour is kept:
there SO_REUSEADDR only lets a restart reuse a port in TIME_WAIT. A restart right after closing a server with
TIME_WAIT connections still binds on Windows (checked).

Related Issue

Related to #91687 (long A2A jobs). This PR does not add non-blocking sends (#91688, #94880, #103453 cover that);
it only stops losing the work when the server gives up waiting. The question this raises for the spec (a blocking
send MUST wait for a terminal state, a server cannot always keep waiting) is in
a2aproject/A2A#2277.

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)

Changes Made

  • plugins/platforms/a2a/adapter.py: _detach() / _finish() replace "fail on timeout or disconnect" in
    message/send and message/stream; _ExclusiveHTTPServer for the listener.
  • tests/plugins/test_a2a_plugin.py: the timeout test now expects WORKING and then the late reply; a ceiling test;
    a test that a second server on the same port fails to bind.
  • tests/plugins/test_a2a_phase23.py: the stream-disconnect test now expects the task to stay WORKING, then complete
    with the late reply and release the active request.
  • website/docs/user-guide/messaging/a2a.md, plugins/platforms/a2a/README.md: what happens past
    A2A_REPLY_TIMEOUT.

Behaviour change to note: a blocking caller whose request outlives A2A_REPLY_TIMEOUT now gets a WORKING task
instead of a FAILED one, and has to follow it with GetTask.

How to Test

  1. pytest tests/plugins/test_a2a_plugin.py tests/plugins/test_a2a_phase23.py tests/plugins/test_a2a_tools_gate.py plugins/platforms/a2a -m ""
    (the round-trip tests are marked integration, so plain pytest skips them): 160 passed, 1 skipped.
  2. Without the adapter change, the four new or updated tests fail.
  3. We run this on a live gateway: a reply after the timeout lands in the task; a second profile on a used port fails
    to start its A2A platform.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix
  • I've run pytest tests/ -q and all tests pass — I ran the A2A suites above (all pass); not the whole tree
  • I've added tests for my changes
  • I've tested on my platform: Windows 11

Documentation & Housekeeping

  • I've updated relevant documentation
  • I've updated cli-config.yaml.example if I added/changed config keys — N/A
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — N/A
  • I've considered cross-platform impact — the exclusive bind is Windows-only; other platforms keep stdlib behaviour

When the request thread stops waiting (A2A_REPLY_TIMEOUT on a blocking send, or
a stream client that disconnected), the task used to be marked FAILED and the
reply that arrived later was dropped: GetTask kept returning FAILED. The task is
now detached instead: it stays WORKING, the caller gets that task, and a daemon
thread records the reply when it arrives (bounded by the orphan ceiling, after
which it fails as before).

On Windows, stdlib SO_REUSEADDR let a second gateway profile bind a listening
A2A port, and requests split between the two. The server now binds with
SO_EXCLUSIVEADDRUSE there; elsewhere the stdlib behaviour is kept.
@alt-glitch alt-glitch added type/bug Something isn't working P3 Low — cosmetic, nice to have comp/plugins Plugin system and bundled plugins platform/windows Native Windows-specific behavior or breakage sweeper:risk-platform-windows Sweeper risk: may break or behave differently on native Windows labels Oct 1, 2026

@marcusrfaust marcusrfaust left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review (marcusrfaust): LGTM. Late-reply detach keeps task WORKING instead of failing on timeout/disconnect, bounded by 24h orphan ceiling. Exclusive-port fix Windows-only, safe elsewhere. Tests updated to assert late reply lands.

kvnloo commented Oct 1, 2026

Copy link
Copy Markdown

A2A lifecycle review: detaching instead of terminal-failing at the request timeout preserves ownership of the in-flight agent work. The task remains WORKING, the daemon finalizer owns the eventual completion/failure, and a stream disconnect does not discard a reply that arrives later. The separate 24h ceiling prevents detached tasks from becoming immortal. On Windows, replacing SO_REUSEADDR with SO_EXCLUSIVEADDRUSE only for the listener addresses the shared-port split-brain without changing POSIX TIME_WAIT restart behavior. I don't see a correctness blocker in these two focused fixes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/plugins Plugin system and bundled plugins P3 Low — cosmetic, nice to have platform/windows Native Windows-specific behavior or breakage sweeper:risk-platform-windows Sweeper risk: may break or behave differently on native Windows type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants