chore: rolling promotion dev -> main - #1059
Conversation
Add first-class trace_id (UUID) and parent_event_id (BIGINT FK) columns to genie_runtime_events for distributed tracing. Includes partial index on trace_id, query support via traceId filter, and promotes traceId from event-router data JSONB to the dedicated column.
Add prune-events subcommand with --older-than and --dry-run options for on-demand cleanup of old runtime events. Document all retention policies in docs/RETENTION.md. Includes integration tests.
Wire getEventMetrics() into the metrics CLI to expose events_emitted, events_failed, last_emit_duration_ms, and circuit_state in both text and JSON output. Also log warnings in runtime-emit handler instead of silently swallowing errors.
resolveTargetTeams() was broadcasting executor events to ALL active teams regardless of whether the erroring agent belonged to them. Test fixtures creating executors in shared PG (e.g. err-agent in protocol-router.test.ts) would flood every live agent session with spurious [executor.error] messages. Now checks event.agentId against team.members and team.leader before routing. Events without an agentId still broadcast to all active teams (existing behavior). Closes #1048
The publish-next job in ci.yml runs on every push to dev. Its purpose is to short-circuit and let version.yml (which runs after CI success) do the real version bump and publish. The skip check previously compared the local package.json version only against `npm view @automagik/genie@next` — i.e. only the @next dist-tag. This missed cross-tag collisions: after a dev→main merge, version.yml bumps dev to a new version (e.g. 4.260404.2), publishes it as @latest, and commits the bump back to dev. The next dev push then reads that version from package.json. Because the @next tag still points at the previous version (4.260404.1), the skip check does NOT match, and bun publish then fails with 403 Forbidden because the version already exists in the registry under @latest. Symptom: every dev-push CI since the last main merge shows Publish @next in failure state. Silent because Publish @next is not a required check. The deviation also produces an inverted npm state where @next lags @latest indefinitely (as of this commit: @next=4.260404.1, @latest=4.260405.1). Fix: query by exact version (`npm view <pkg>@<version>`) so any tag collision short-circuits the step. version.yml will still derive a fresh version and publish it on the same CI run, so @next stays current. Verified locally against the live registry: LOCAL=4.260405.1 → skip (already on @latest) LOCAL=99.99.99 → publish (unpublished)
Implements Group 3 of the unified-executor-layer wish: the omni-bridge must start without PG and gracefully degrade on runtime failures so a dropped PG connection never drops a user reply. - Add `pgAvailable: boolean` on OmniBridge and expose it on BridgeStatus - New private `probePg()` runs at start() after NATS connects; never throws, logs warn on failure, and sets pgAvailable=false for degraded mode - New private `safePgCall<T>(op, fn, fallback, ctx?)` helper — single entry point for every downstream PG call. Try-once semantics, 2s runtime timeout, fallback return on error, flips pgAvailable=false on connection-level errors - Inject hooks: `BridgeConfig.pgProvider` and `BridgeConfig.natsConnectFn` for hermetic unit tests (default to lib/db.getConnection() and nats.connect) - Helper utilities: `withTimeout` and `isPgConnectionError` classify failures by postgres.js code and common message fragments - Tests in src/services/__tests__/omni-bridge.test.ts cover: degraded startup (provider throws, SELECT 1 fails), mid-run connection loss flips the flag, non-connection errors keep the flag true, fast-path short-circuit when PG was never healthy, and the happy path forwards the fn result. Downstream groups (4, 5, 6, 7) will wire their PG writes through safePgCall.
… poisoning Root cause of "my agent exists on disk but genie spawn can't find it": startAgentSync() in serve.ts silently returned null when findWorkspace() returned null, disabling the entire auto-discovery subsystem (initial sync + file watcher) for the whole serve lifetime. Users saw new agents in workspace/agents/ never register and received no diagnostic anywhere. The reason findWorkspace() returned null in practice: test runs persist their tmpdir workspace paths to ~/.genie/config.json via saveWorkspaceRoot(). Once a serve boots from outside the real workspace, walk-up fails, the fallback hits the stale tmp path (now removed), existsSync returns false, returns null. Fixes: - workspace.ts: saveWorkspaceRoot() refuses to persist paths under os.tmpdir(), so test runs can't poison the global config. - workspace.ts: loadWorkspaceRoot() self-heals — if the saved path no longer has a .genie/workspace.json marker, scrub it from config instead of returning a stale fallback forever. - workspace.ts: GENIE_HOME is now resolved lazily via genieHome() so tests can override it via process.env.GENIE_HOME per-test. - serve.ts: startAgentSync() logs loudly when workspace is unresolved, when sync has per-agent errors, and when the watcher fails to start. Also includes the resolved workspace in the success log so "which workspace did serve attach to" is visible at a glance. - workspace.test.ts: two new tests covering tmp-path guard and stale-path self-heal; tests now isolate GENIE_HOME so they don't touch the developer's real ~/.genie/config.json.
fix(ci): detect cross-tag version collisions in publish-next skip check
…kspace-failure fix(agent-sync): surface workspace resolution failures, stop tmp path poisoning
The `genie brain update` command was intercepted before reaching executeBrainCommand(), preventing brain's own update subcommand from running. Rename the package-update trigger to `upgrade` so `genie brain update` delegates to the brain module as intended.
…T env override (#1043) WSL2's slow filesystem causes pgserve to exceed the hardcoded 15s timeout. Default is now 30s, overridable via GENIE_PGSERVE_TIMEOUT env var.
`genie init agent` always used the workspace root's agents/ directory, ignoring the user's current working directory. Now: - If CWD is inside an agents/ directory, create the agent there - Otherwise, fall back to workspace root's agents/ - New --dir <path> option for explicit override
…low-query test
Closes the three gaps flagged in WISH Post-Audit Decision 3 before Wave 2
dispatches. Groups 4/5/6/7 depend on these changes to start.
1. `safePgCall` is now `public` on `OmniBridge` (Decision 2). Downstream
executors hold an `OmniBridge` reference and call the method directly,
e.g. `bridge.safePgCall('op', fn, fallback, { chatId })`. Tests updated
to drop the `(bridge as any)` workaround.
2. `probePg()` now classifies startup errors (Decision 3):
- Connection-level (ECONNREFUSED, connection terminated, …) → degrade
gracefully, log warn, `pgAvailable=false`. Dev without PG still works.
- Anything else (schema mismatch, missing relation, permission denied)
→ fail-fast with an actionable error pointing at the migration
command. Matches the wish's PG Error Handling Strategy table row for
"Migration missing / schema mismatch" — silent corruption is worse
than noisy startup failure.
3. New unit tests in `src/services/__tests__/omni-bridge.test.ts`:
- Fail-fast path: SELECT 1 rejects with "relation 'sessions' does not
exist" → `start()` throws with "PG schema mismatch … migrate" hint.
- Fail-fast variant: provider throws a non-connection error with
code 42501 (permission denied) → `start()` throws with hint.
- Slow-query fallback: fn delays 2.5s (beyond the 2s runtime budget)
→ `safePgCall` returns fallback, `pgAvailable` STAYS true (slow is
not a connection loss), elapsed time confirms the budget held.
Validation:
bun run typecheck → clean
bun run lint → 0 errors
bun test src/services/__tests__/omni-bridge.test.ts → 9 pass / 31 expects
…path in serve warning Follow-up to #1056 addressing two review comments from gemini-code-assist. 1. isTempPath() now uses realpathSync to canonicalize both tmpdir() and the candidate root before prefix-comparing. On macOS /tmp is a symlink to /private/tmp and tmpdir() returns /var/folders/..., so the previous path.resolve()-only check would miss any literal /tmp/... path and bypass the guard against poisoning the global workspaceRoot config. Fails CLOSED on canonicalization errors (treats as temp → refuses to persist) since the whole purpose of this guard is to be conservative. 2. startAgentSync() now builds its "no workspace found" warning using genieHome() (newly exported) instead of hardcoding ~/.genie/config.json, so users who set GENIE_HOME see the actual path in the fix guidance. Also adds: - A regression test using a real symlink within the test's tmpdir to verify canonicalization works end-to-end (skips gracefully on CI environments that disallow symlinks). - A sanity test that genieHome() honors the GENIE_HOME env override.
…w-up)
Two bot reviewers (Gemini and Codex) independently flagged three defects
in `cwd.includes('/agents')`:
1. Substring match — falsely matches `/tmp/agents-backup`,
`/repo/agentship`, `/home/agentsmith`, or any path containing the
literal `agents` anywhere, causing `genie init agent` to scaffold
OUTSIDE the workspace.
2. Windows portability — hardcoded forward slash fails on systems where
paths use backslashes.
3. Nesting bug — when CWD is inside an existing agent directory
(e.g. `.../agents/my-agent`), the function returned CWD, causing the
new agent to be scaffolded as a nested child instead of a sibling.
Fix: walk workspace-relative path segments (via `path.relative` and
`path.sep`) and match the first exact `agents` segment, returning the
path up to and including that segment. Escapes outside the workspace
(rel starts with `..`) fall back to `<wsRoot>/agents`.
…solution fix(workspace): canonicalize tmp paths, dynamic config path in serve warning
fix: CLI core bugs — task stage, brain paths, init CWD, pgserve timeout
feat: runtime events hardening — trace_id, circuit breaker, metrics, prune CLI
Group 5 — SDK sessions now produce the same sessions + session_content rows that session-capture.ts (filewatch) produces for tmux sessions. New files: - src/lib/audit-events.ts: shared AuditEventType union + recordAuditEvent helper (writes to audit_events via safePgCall, used by Group 5 & 7) - src/services/executors/sdk-session-capture.ts: startSession, recordTurn, updateTurnCount, endSession — all writes through SafePgCallFn - src/services/executors/__tests__/sdk-session-capture.test.ts: 15 tests Integration in claude-sdk.ts _processDelivery(): - deliver.start / deliver.end audit events bracketing each query - User + assistant turns recorded in session_content per delivery - PG session row created lazily on first delivery, ended on shutdown - All writes skip silently when safePgCall is null (degraded mode)
…e rows When multiple PG rows share the same role (e.g. a dir:omni directory entry and stale runtime rows from prior genie spawn sessions), resolve() picked whichever row PG returned first from an unordered LIMIT 1 query. If PG returned a stale runtime row (empty metadata) instead of the dir: row (full metadata from AGENTS.md frontmatter), the agent appeared unregistered — missing dir, model, color, description, promptMode. This was a coin-flip bug: some agents worked (PG happened to return the dir: row) while others silently lost all their configuration. The omni agent had 3 rows in PG and consistently resolved to a stale one. Fix: add ORDER BY (CASE WHEN id LIKE 'dir:%' THEN 0 ELSE 1 END) to both resolve() and ls() queries, so directory-managed rows always win over runtime-spawned rows regardless of started_at ordering.
…sdk test mocks - Add _sql parameter to happySafePgCall/degradedSafePgCall to match SafePgCallFn type - Add mockSql tagged-template stub so audit-events callbacks don't crash - Add missing findLatestByMetadata mock to executor-registry mock - Add mock parameter types to createAndLinkExecutorMock for call destructuring - Restore SDK query mock in concurrent delivery afterEach to prevent test pollution
Add bridge restart recovery for in-flight Claude sessions: - Migration 026: partial JSONB index on executors(agent_id, source, chat_id) for fast omni-sourced executor lookups (WHERE ended_at IS NULL) - executor-registry: add findLatestByMetadata, relinkExecutorToAgent, updateClaudeSessionId helpers - claude-sdk spawn(): before creating fresh executor, look up existing live executor via findLatestByMetadata and reuse it (with session ID) - claude-sdk deliver(): detect resume rejection (SDK returns different session ID), persist new session ID, write session.resume_rejected audit - Audit events: session.resumed, session.created_fresh, session.resume_rejected - 11 new tests covering all resume paths in claude-sdk-resume.test.ts
…tory-rows fix(directory): resolve() and ls() prefer dir: rows over stale runtime rows
Remove .genie/wishes/ from .gitignore so wish plans and audit artifacts are version-controlled alongside the code they describe.
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughBumps package/plugin versions; adds tracing and a circuit-breaker with metrics for runtime events; implements safe‑PG degraded mode and executor registry integration in the Omni bridge and executors; adds source-based executor filtering to CLI; retention docs and prune command; many DB migrations and tests. Changes
Sequence Diagram(s)sequenceDiagram
participant NATS as NATS
participant Bridge as OmniBridge
participant Exec as Executor (SDK/Code)
participant DB as PostgreSQL
rect rgba(100,200,100,0.5)
Note over NATS,DB: Spawn + registry registration (safe-PG)
NATS->>Bridge: omni.turn.open / spawn message (env, chatId)
Bridge->>Bridge: probePg() -> pgAvailable true
Bridge->>Exec: executor.setSafePgCall(safePgCall)
Bridge->>Exec: spawn(agentName, chatId, env)
Exec->>DB: safePgCall(registerInWorldA) -> create/relink executor
DB-->>Exec: executorId / claudeSessionId
Exec-->>Bridge: spawn complete (session ready)
end
sequenceDiagram
participant App as Publisher
participant CB as EventCircuitBreaker
participant DB as PostgreSQL
participant Metrics as In-process Metrics
rect rgba(100,200,100,0.5)
Note over App,DB: Successful publish
App->>CB: publishRuntimeEvent(event with traceId)
CB->>CB: check closed state
CB->>DB: INSERT runtime event (trace_id,parent_event_id)
DB-->>CB: success
CB->>Metrics: eventsEmitted++, reset failCount
CB-->>App: success
end
rect rgba(200,100,100,0.5)
Note over App,DB: Failure path -> circuit open
App->>CB: publishRuntimeEvent(event)
CB->>DB: INSERT runtime event
DB-->>CB: error
CB->>Metrics: eventsFailed++, failCount++
CB->>CB: open if threshold reached
alt circuit opened
CB-->>App: throw "circuit breaker open — PG event write skipped"
else
CB-->>App: return error (write failed)
end
end
Estimated code review effort🎯 5 (Critical) | ⏱️ ~150 minutes Possibly related issues
Possibly related PRs
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Both types are only used internally — no external imports exist. Fixes knip dead-code warnings on PR #1073.
feat(genie-app): v2-ui 4-wave delivery + wish rescope + flake fix
feat(events): real-time stream command with LISTEN/NOTIFY + StreamTable
Root cause: three interrelated issues caused CI-only failures: 1. mock.restore() in afterAll clears all process-global mock.module registrations, breaking whichever test file runs second 2. mockClear() only clears call counts but not mockImplementation() overrides set by other test files' nested describe blocks 3. Concurrent delivery afterEach set queryMock via dynamic import without session_id in result events Fix: - Add resetAllMocks() to _sdk-mocks.ts that does mockReset() + mockImplementation() for every shared mock function - Use resetAllMocks() in both test files' beforeEach blocks - Remove mock.restore() from both files' afterAll - Use shared queryMock directly instead of dynamic imports
fix(test): prevent cross-file mock leak in SDK executor tests
bunx tauri fails to resolve the package, and older npm-based Tauri CLI has a broken bundle_dmg.sh that uses AppleScript (fails without Finder permissions). cargo tauri uses the native bundler — matches khal-os's proven approach.
Audit all 80+ wishes against dev codebase. 17 had stale statuses: SHIPPED (14): - unified-omni-bridge (PRs #1063, #1065) - fix-omni-bridge-hardening (PR #1065) - unified-executor-layer (PR #1062) - auto-orchestrate, fix-depends-parser, parallel-execution - task-projects, test-pg-ram-isolation, task-auto-close-on-merge - worktree-out-of-repo, docs-overhaul, genie-hacks-community-docs - multi-agent-session-isolation, session-auto-create OBSOLETE (3): - genie-omni-marriage (superseded by smaller wishes) - fix-session-uuid-resume (replaced by --continue by name) - qa-dev-to-main (time-bound QA from March 20)
chore: wish housekeeping — mark 14 shipped, 3 obsolete
…build @khal-os/sdk leaks server-only imports (WorkOS AuthKit for Next.js) into the client bundle. Tauri/Vite builds fail because these modules are not resolvable in a non-Next.js context. Externalize them since the desktop app does not use WorkOS auth.
Tauri's bundle_dmg.sh was failing silently. khal-os builds DMG successfully with the same Tauri version (2.10.3) by including macOS minimumSystemVersion, icon references, and category in the bundle config. Aligning genie's tauri.conf.json to match.
Tauri's bundle_dmg.sh fails on macOS 26.3 due to a relative path bug in the bundler (current_dir mismatch when invoking the script). The .app builds fine — only the DMG packaging step breaks. Split make tauri into: - make tauri-app: builds .app via cargo tauri build --bundles app - make tauri-dmg: packages .app into .dmg via hdiutil create - make tauri: runs both (the full pipeline) hdiutil is native macOS, no dependencies, no AppleScript permissions, works in CI headless environments.
Add waitForExecutorReady() that uses PG LISTEN/NOTIFY on the genie_executor_state channel to detect when an executor reaches 'running' or 'idle' state, with polling fallback every 2s. - spawn-command.ts: new waitForExecutorReady() with _pgDeps injection - executor-registry.ts: emit executor.ready audit event on running - protocol-router.ts: try PG readiness before falling back to tmux - Graceful degradation: falls back to tmux scraping when PG unavailable - Mark daily-metrics-agent wish as SHIPPED
Stage .app + /Applications symlink into a temp dir before creating the DMG so users see the standard macOS drag-to-install layout.
…vents feat: add PG-based readiness detection for spawned agents
Rolling Promotion PR
Auto-maintained rolling promotion PR from
devtomain.What is in this promotion
Major Features
Bug Fixes
Known limitations (not blocking)