docs(handoff): STREAM/ real-time anonymous handoff protocol + perpetual wakeup - #39
Conversation
…al wakeup Adds two coordination layers to keep the multi-agent loop running while the user is asleep: 1. handoffs/STREAM/ — markdown-based pub/sub between Claude and Codex (and any other client: KiloCode/Cursor/Windsurf/VSCode+Copilot). Files: - PROTOCOL.md — message format, types, polling cadence, conflict rules - STATE.md — live snapshot, both sides update - CLAUDE_INBOX.md — messages for Claude - CODEX_INBOX.md — messages for Codex (bootstrapped with 7 messages) - LEDGER.md — append-only audit trail - GATE_GAP_QUEUE.md — 18 missing gates (6 P0, 8 P1, 4 P2) - ENHANCEMENT_QUEUE.md — 9 unfinished work items - WATCHDOG.md — heartbeat + auto-reassign + backup spec - CLIENT_ADAPTERS.md — drop-in instructions per client Actual scripts (validate/watchdog/backup/archive) live in HermesProof at scripts/stream-*.mjs (Node, zero deps). 2. handoffs/HANDOFF_TO_CODEX_PERPETUAL_WAKEUP.md — supersedes the §3 "Stop after that." condition in HANDOFF_TO_CODEX_OVERNIGHT_AUTOPILOT.md. Codex no longer exits on queue drain; instead it idle-polls STREAM/ continuously until the user wakes up and explicitly says stop. This addresses the overnight stall where Codex correctly hit the original exit condition after 6 tasks but the project still had ~27 unfinished items. No code changes; pure docs + handoff infrastructure. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
Warning Rate limit exceeded
To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (10)
✨ Finishing Touches🧪 Generate unit tests (beta)
Review rate limit: 0/5 reviews remaining, refill in 4 minutes and 54 seconds. Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces the "STREAM" protocol, a markdown-based coordination layer designed for continuous, anonymous collaboration between AI agents during user downtime. It establishes a structured system of inboxes, state snapshots, and prioritized work queues (GATE_GAP and ENHANCEMENT) to manage tasks across repositories. Feedback highlights several critical synchronization and concurrency risks, including potential clock skew issues with timestamp-based tie-breaking, the lack of file-level locking for state updates, and race conditions inherent in in-place status edits. Additionally, there are inconsistencies between the intended watchdog cadence and the GitHub Action schedule, as well as ambiguity regarding the terminal nature of "expired" versus "stuck" message statuses.
| Two builders can't lock the same files (HermesProof prevents it). But two | ||
| agents could still try to **claim the same logical task**. Resolution: | ||
|
|
||
| 1. First TASK_CLAIMED message wins (timestamp tiebreak, lexicographic on author ID otherwise). |
There was a problem hiding this comment.
| Update STATE.md anytime you change material state. Don't wait — stale STATE.md | ||
| is worse than no STATE.md. |
| workspaces). Cross-repo coordination uses **mirror messages**: post in BOTH | ||
| inboxes with the same correlation ID. |
There was a problem hiding this comment.
| ## Bookkeeping | ||
|
|
||
| - 2026-05-03 11:30Z — initial seed by Claude SCRIBE | ||
| - claim convention: edit `status:` field in place, post TASK_CLAIMED in other side's inbox |
There was a problem hiding this comment.
|
|
||
| | Job | Cadence | Runs in | What it does | | ||
| |---|---|---|---| | ||
| | **W1 — heartbeat-check** | every 1 min | local node script + GH Action cron | Flags any role that hasn't written to STATE.md in >15 min as `IDLE`. Flags any `in_progress` message past expiry as `STUCK`. | |
There was a problem hiding this comment.
| A correlation is **STUCK** if: | ||
| - Has >3 messages all `open` or `acknowledged` (none `resolved`) | ||
| - AND last message in correlation aged >20 min with no follow-up | ||
| - OR a message's `expires:` field has passed and `status` is not `resolved`/`expired` |
There was a problem hiding this comment.
Summary
Two coordination layers so the multi-agent overnight loop never stalls again:
This is the response to tonight's stall: Codex correctly hit the original exit condition after 6 tasks but the project had ~27 unfinished items. Now there's no exit condition until the user wakes.
The actual scripts (validate / watchdog / backup / archive) ship in HermesProof PR #20.
Test plan
🤖 Generated with Claude Code