Skip to content

fix(a2a): preserve inbound message in task writeback - #4

Merged
RuniThomsen merged 1 commit into
mainfrom
hermes/727-receiver-writeback-s687
Aug 22, 2026
Merged

RuniThomsen merged 1 commit into
mainfrom
hermes/727-receiver-writeback-s687

Conversation

@RuniThomsen

Copy link
Copy Markdown
Collaborator

Summary

  • persist the inbound A2A Message as the receiver-owned Task.history child
  • preserve referenceTaskIds, governed metadata, extensions, parts, and correlation IDs
  • stamp the receiver-minted taskId without mutating the caller request
  • return the write-back record in both non-holding SendMessage receipts and the first SendStreamingMessage Task frame
  • drop caller-only top-level keys so strict Task ProtoJSON remains valid

Why

NousResearch#727 requires one seat dispatch to produce/attach a visible task row without a second write. The sender-side coupling needs the receiver's first receipt to carry the authoritative Task plus the original Message as its child. Hermes previously minted the Task but discarded that Message from Task.history.

Verification

  • TDD RED observed for TaskStore history, non-holding receipt, and streaming first frame
  • 157 passed, 18 deselected for test_a2a_plugin.py + test_a2a_phase23.py
  • focused streaming integration: 1 passed
  • py_compile and git diff --check pass
  • broad integration has one unrelated anti-loop expectation that already fails on exact base d048b78480 because that test still expects holding completed instead of the new non-holding working receipt

@RuniThomsen

Copy link
Copy Markdown
Collaborator Author

CI classification: the fork workflows failed before execution. Each failing detector/OSV job reports an empty runner name, zero steps, and no logs; all code/lint/test jobs were skipped. This is infrastructure red, not a candidate test failure.

Executed evidence on exact base d048b78480:

@RuniThomsen
RuniThomsen merged commit ea6dd04 into main Aug 22, 2026
20 of 27 checks passed
@RuniThomsen
RuniThomsen deleted the hermes/727-receiver-writeback-s687 branch August 22, 2026 16:02
RuniThomsen added a commit that referenced this pull request Aug 23, 2026
…s687

fix(a2a): preserve inbound message in task writeback
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
… fingerprint never authorizes a signal

NousResearch#111617 review (andrexibiza P1 #3/#4, kvnloo nit):

- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
  is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
  a reboot, and that counter does not, so an unrelated process on a later boot with the
  same PID and the same tick value passed _start_times_agree(). The fingerprint is now
  "<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
  the witness the drain marker already uses); both halves must match. Integer values on
  rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
  row and falls back to bare PID existence - a new spawn silently recreated the NousResearch#89614/
  NousResearch#99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
  claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
  by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
  once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
  the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
  and must not outlive its pid.

Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.

Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
…as a died link

The exit-notify wrap reported every receive-loop exception at ERROR with a
traceback, including the ConnectionClosedOK that follows our own CLOSE frame in
disconnect(). Gate on the adapter's _running flag (published through the WS
thread-local next to on_link_up): a live link's death stays ERROR, an
intentional shutdown logs at DEBUG. Live pass side-effect #4 on NousResearch#113662.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant