Release 4.260507.7 — done --report mandatory + 484-commit dev sweep - #1704
Conversation
…ilent-fail fix(spawn): #1600 — surface silent-fail spawn pipeline (validation + audit events)
…ing spawns ROOT CAUSE of the long-running #1589/#1600 phantom dispatch saga, found empirically on .25 after PR #1601's audit observability shipped: `runWorkDispatch` called `process.exit(1)` on three error paths (wish not found, group not found, startGroup failed). When `autoOrchestrateCommand` runs `Promise.allSettled([runWorkDispatch × N])` for a parallel wave, ANY group's process.exit terminates the entire node process — killing every sibling spawn mid-flight before: - handleWorkerSpawn's worker.spawn audit event fires - launchTmuxSpawn reaches the new validateSpawnedPane check - any of the new worker.spawn.failed / worker.spawn.ok events fire This explains why .24 + .25 still phantom-dispatched despite the `wish.dispatch.work` event firing correctly: the dispatch event landed during the brief window between Group N's PG state mutation and Group N's spawn pipeline being killed by Group N+1's exit(). The fix: throw `new Error(...)` on each runWorkDispatch error path. Both callers handle the throw correctly: - `workDispatchCommand` (single-group CLI): the outer commander handler in registerDispatchCommands already wraps in try/catch + process.exit, so single-group semantics are preserved. - `autoOrchestrateCommand` (wave): Promise.allSettled collects rejections into the `failed` array, my Round 2 (#1601) wish.dispatch.failed event fires per group, and stderr summary prints with the [ErrorClass] prefix. Empirical confirmation: Before: `genie work tui-bottom-bar-opentui` exits silently after Group 5's dependency error kills the process. No worker.spawn / .ok / .failed events fire for in-flight Group 1. After (this fix): Group 5's startGroup throws; Promise.allSettled collects the rejection. Group 1's spawn pipeline runs to completion, emitting either worker.spawn.ok (success) or worker.spawn.failed (real spawn failure surfaced). Tests: - 1 new regression-guard in dispatch.test.ts (#1600 Group 4): asserts zero EXECUTABLE process.exit calls in runWorkDispatch (comment text mentioning the OLD behavior is excluded), and >=3 throw-new-Error statements covering wish-not-found, group-not-found, startGroup-failed. - Total: 86 pass / 0 fail in dispatch.test.ts. - bun run typecheck: clean. Followup to #1599 + #1601. Should close the root-cause iteration on the long-running #1589/#1600 phantom-dispatch saga. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…h on phantom Round 4 of the #1589/#1600 phantom-dispatch saga. Felipe identified in parallel with #1602 (process.exit-kills-siblings) that there's ALSO a silent-fail surface in the auto-resume path: `handleWorkerSpawn` calls `resolveTeamAndResume` which calls `findDeadResumable` to locate a stale dead row matching the requested role+team. If found, it calls `resumeAgent(deadResumable)` and returns the agent id as `resumed`. The caller then short-circuits at `if (resumed) return resumed;` (agents.ts:2356). Failure mode: `resumeAgent` calls `createResumeTmuxPaneOrExit` which calls `createTmuxPane`. tmux split-window returns a paneId atomically, but the script invoked inside that pane (the resume command) can fail to exec — the pane closes immediately. The rest of `resumeAgent` operates on a ghost, recordAuditEvent('resumed') fires, the caller short-circuits, and no actual worker exists. This is the same failure mode as #1601 (validateSpawnedPane) but on the OTHER spawn path that wasn't instrumented yet. Fix: 1. **Post-resume validation** in `resumeAgent` — after createTmuxPane returns paneId, capture pane PID and verify (a) pane is in `tmux list-panes -a`, (b) PID is alive (process.kill(pid, 0) ESRCH check). Throws `ResumePaneVanishedError` (typed, mirrors SpawnPaneVanishedError from #1601) if either check fails. 2. **Resume observability** — three new audit events: - `worker.resume.attempted` before validation - `worker.resume.completed` after validation passes - `worker.resume.skipped` on validation failure (with reason) 3. **Fall-through to fresh spawn** in `resolveTeamAndResume` — wraps resumeAgent in try/catch; on ResumePaneVanishedError, marks the stale executor as 'terminated' (so next dispatch doesn't loop on the same dead row) and returns without `resumed` so handleWorkerSpawn proceeds with a fresh spawn. Tests: - 3 new regression-guard tests in dispatch.test.ts: - ResumePaneVanishedError class shape - resumeAgent emits all 3 resume events + uses validation primitives - resolveTeamAndResume catches typed error and falls through - Total: 88 pass / 0 fail (was 85 + 3 new) - bun run typecheck: clean - bun run lint: 0 errors, 18 warnings (pre-existing only) Companion to #1602. Together with #1599 + #1601 + #1602 this completes the 4-PR root-cause iteration for the long-running #1589/#1600 phantom-dispatch. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…exit-kills-siblings fix(dispatch): runWorkDispatch must throw not exit — preserves sibling spawns in wave dispatch
…e-row fix(spawn): auto-resume validates resumed pane is alive — fall through on phantom (Round 4)
…cal pgserve is detected Closes the wish's "shared backbone" loop for genie. Previously, `genie install` registered genie-serve under pm2 with NO env block — meaning genie-serve always fell back to spawning its own embedded pgserve on `:19644` regardless of whether canonical pgserve was registered. Operators had to hand-edit `~/.genie/genie-serve.config.cjs` to add an env block (caught live on this server during the canonical migration). What changes ------------ - `tryPgservePort()` (new) — probes `pgserve port` to discover the canonical port. Uses the real subcommand (NOT `--version`, which doesn't exist in pgserve@^2.1.0 and false-negatived in `omni doctor --fix` historically). - `buildGenieDatabaseUrl(port)` (new) — composes the canonical URL using the `genie` database (auto-provisioned by pgserve on first connection, mirrors omni's pattern). - `buildEcosystemConfigSource(geniePath, databaseUrl?)` — accepts an optional `databaseUrl` and bakes it into an `env` block when present. Omits the env block entirely when absent (so genie-serve falls back to its embedded auto-spawn path without an empty env clobbering an operator's shell-set DATABASE_URL). - `buildPm2StartArgs(geniePath, databaseUrl?)` and `writeEcosystemConfig(geniePath, databaseUrl?)` — thread the URL through. - `installCommand` — when `tryPgserveInstall()` succeeds AND `tryPgservePort()` returns a valid port, derives the canonical URL and passes it through. Logs the URL on success so operators can see the wire that was made. - The "already installed" early-return now hints `pm2 delete genie-serve && genie install` as the way to refresh env on URL change. Tests ----- - `omits env block when no databaseUrl provided (legacy fallback path)` - `bakes DATABASE_URL into env block when canonical pgserve url provided` - 14/14 tests pass (was 12, +2 new env-wiring tests). - Typecheck green. Linter: no new warnings on changed files. Validated on khal-os -------------------- After hand-editing the ecosystem config to include the env block, genie-serve connects to canonical pgserve via TCP `:8432` and the embedded `:19644` postgres no longer spawns. This PR codifies that hand-edit so future installs are correct out of the box.
Mirrors `omni doctor`'s `pgserve-canonical` check on the genie side
so operators see the same shared-backbone visibility from both halves
of the canonical-stack. Surfaces three signals:
✓ pgserve binary canonical port 8432
✓ pgserve under pm2 online — shared backbone for genie-serve + omni-api
…or, when something is off:
! pgserve binary not on PATH (or `pgserve port` failed)
Install canonical pgserve: bun add -g pgserve@^2.1.0
! pgserve under pm2 binary present but not registered under pm2
Register canonical pgserve: pgserve install
Probes
------
- Binary detection via `pgserve port` (NOT `--version` — that flag
doesn't exist in pgserve@^2.1.0 and false-negatived in historical
doctor implementations; same lesson surfaced in `omni doctor --fix`
on 2026-04-30).
- Pm2 registration + online status via `pgserve status --json`.
Severity
--------
Both checks are WARN, never FAIL. Genie can auto-spawn its own daemon
as a fallback for fingerprint-routed CLI commands, so a missing
canonical pgserve doesn't break local development — it just means
genie isn't sharing the backbone with omni and other automagik
services on this host.
Tests
-----
Live-validated on canonical-pgserve-running host: both checks PASS,
section renders cleanly, exit-code unchanged.
…onical fix(install): bake DATABASE_URL env into ecosystem config when canonical pgserve is detected
…ical feat(doctor): add Pgserve (canonical backbone) check section
Add end-to-end integration test wiring HeartbeatPublisher against a fake
in-memory NATS bus and a fake omni-side TurnMonitor. Uses a virtual
clock with injected setInterval/clearInterval to compress 200s of
simulated time into ~300ms wall-clock.
Covers the two acceptance criteria from the wish:
1. Busy session: 200s with heartbeats → ZERO turn.nudge events.
2. Idle session: 200s without heartbeats → exactly ONE nudge at
120 ± 5s (regression guard).
Plus a busy-then-settled scenario that proves bridge stop/wire-down
hands off cleanly to the existing nudge path.
Wish: omni-activity-heartbeat (group 3)
Genie worktree is wish-omni-heartbeat (no slash, matches team name); omni side ships on wish/omni-activity-heartbeat. Update the WISH.md header so reviewers do not flag the mismatch. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
feat(heartbeat): publish omni.agent.heartbeat.* during busy sessions
…-roadmap docs(wishes): add Genie observability roadmap
Removes the dynamic-import fallback path (`await import('pgserve')`)
in `src/lib/db.ts` that lived behind every binary lookup in
`findPgserveDaemonCommand`. The fallback only fired when ALL FOUR
binary discovery paths (local node_modules walker, require.resolve,
bun-global, PATH) failed — which never happens on a host with
canonical pgserve. It was dead code carrying a runtime npm dependency.
What's removed
--------------
- `importPgserveSdk()` and `tryEnsureDaemonWithSdk()` functions
- `resolveLocalPgserveEntry()` (only consumed by importPgserveSdk)
- `PgserveSdk`, `PgserveSdkDaemonState` interfaces
- `import { pathToFileURL } from 'node:url'` (only used by SDK path)
- `require.resolve('pgserve/bin/pgserve-wrapper.cjs')` blocks in BOTH
`findPgserveDaemonCommand()` and `findPgserveBin()` — these only
worked when pgserve was a declared runtime dep, which it no longer is
- `pgserve` line from `package.json` dependencies (-1 package on
install per `bun install` output)
What stays
----------
- `findLocalPgserveRoot()` — walks `node_modules/pgserve` from
`import.meta.dir` WITHOUT using `require.resolve('pgserve')`. Still
works in monorepo / yarn-link / npm-install scenarios where pgserve
happens to be co-located. No package.json declaration needed.
- `findPgserveDaemonCommand()` and `findPgserveBin()` — both keep
their bun-global + PATH fallbacks. The canonical install path
(`bun add -g pgserve@^2.1.0`) flows through bun-global; container
/ system installs flow through PATH.
- `--external pgserve` in the `build` script — no-op safety net since
the bundle no longer references the package.
Compatibility
-------------
Every host with canonical pgserve (PR pgserve#57 / pgserve#62) keeps
working — the binary is found via bun-global or PATH. The error path
("pgserve binary not found") now points operators at
`bun add -g pgserve@^2.1.0` instead of `bun add pgserve` since
canonical install is the new norm.
Tests
-----
- `src/lib/db.test.ts` updated:
* Replaced SDK-order assertions with negative-locks (NOT-contains)
so we don't reintroduce `importPgserveSdk` / `tryEnsureDaemonWithSdk`
/ `await import('pgserve')` / `resolveLocalPgserveEntry`.
* Reordered fallback assertion: local → bun-global → PATH (no
`require.resolve` step in between).
* Updated test that previously asserted the SDK fallback chain.
- 66/66 db tests pass.
- 4527/4528 total tests pass (the 1 fail is a pre-existing
`docs/_internal/state-machine.mdx` issue unrelated to this change).
- `knip` clean (added `pgserve` to `ignoreBinaries` since CI workflows
still reference the binary via shell).
- typecheck green, build green.
Closes the npm-runtime-dep gap on the genie side. Pairs with omni's
embedded-mode removal (still pending — different PR).
After the chore/drop-pgserve-sdk-runtime-dep PR removed pgserve from
genie's runtime npm deps, ./node_modules/.bin/pgserve no longer exists
post bun install. Two CI jobs were depending on it:
1. pgserve-v2-smoke (Unix socket round-trip) — explicitly invoked
./node_modules/.bin/pgserve daemon ... and timed out waiting for
the libpq socket because the binary was missing.
2. pg-tests matrix shards 1-4 — bootstrap a pgserve daemon via the
bun preload test-setup, which goes through findPgserveDaemonCommand.
That walker no longer has require.resolve('pgserve/...') (removed
in the same PR), so it fell through to bun-global / PATH lookups
which had no pgserve installed on the runner.
Fix: install pgserve globally (`bun add -g pgserve@^2.1.0`) in both
jobs. This matches the canonical install path operators use in
production and aligns the CI environment with how genie actually runs
on real hosts.
Validated locally: the chore/drop-pgserve-sdk-runtime-dep branch's full
test suite passes (4527/4528, the 1 fail is pre-existing and unrelated)
when pgserve is installed via bun-global.
Adds `process.stdout.isTTY === false` to `isTuiDisabled()` so piped/redirected invocations don't try to attach the TUI. Operator-reflex `--no-tui` becomes a manual override only. Repro before fix: `genie | head` attempts TUI attach in non-TTY context. Repro after fix: pipe-mode short-circuits cleanly via existing tui-skip path. Trace: /tmp/trace-genie-no-tui-pollution.md (out-of-tree) Closes: (none — no GH issue filed for this DX papercut)
fix(tui): auto-disable TUI bootstrap when stdout is not a TTY
…ntime-dep chore(db): drop pgserve SDK fallback + runtime npm dep
Group 1 of the fix-agent-session-linkage wish — make the bug measurable before changing production behavior. Two failing tests pin the two distinct symptoms found on the live DBs: 1. Tmux JSONL path — ensureSession() in session-capture.ts early-returns when the row already exists, so an orphaned row inserted before the executor became known is never upgraded. 2. SDK path — startSession() uses ON CONFLICT (id) DO NOTHING, dropping linkage when ingestion has already created an orphan for the same id. session-link-repair.ts exposes pure-read diagnostics (diagnoseSessionLinks / sampleLinkableOrphanSessions / findAmbiguousExecutorSessions) — no mutation, suitable to feed a future genie sessions repair-links --dry-run. REPORT.md records the local baseline (6 linkable orphans, 1855 sessions with NULL executor_id, 70 735 / 71 052 tool_events with empty-string attribution). Remote ssh felipe baseline is blocked from this worktree (no key for the tailscale alias) — flagged for team-lead. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…th version The verify banner emitted `✖ Server: vX (mismatch)` on EVERY actual upgrade. The diagnostic JSON literally captured the cause: update.cliVersion: "4.260507.1" (compile-time const of running CLI) update.plugin.version: "4.260507.2" (post-bun-swap disk truth) update.latestVersion: "4.260507.2" (registry target) → verify.kind: "version-mismatch" `bun add -g @automagik/genie@next` atomically swaps the package on disk WHILE THE CLI THAT INVOKED `genie update` IS STILL RUNNING. The running CLI's `VERSION` is frozen at module-import time (the OLD value). Comparing that against post-update disk truth always fails on a real bump. The verify probe introduced in 109c9b8 traded the previous tautology for a guaranteed false-positive. Truthful verify --------------- Drop the CLI-vs-disk comparison entirely. The only question that matters post-pm2-restart is: "does the daemon's running inode match disk?" The kernel `/proc/<pid>/cwd` `(deleted)` marker answers that directly. type VerifyResult = | { kind: 'ok'; version: string|null; pid: number|null } | { kind: 'health-unreachable'; endpoint: string } | { kind: 'daemon-stale-inode'; diskVersion: string|null; pid: number; cwd: string } | { kind: 'auth-invalid' } | { kind: 'skipped'; reason: VerifySkipReason }; `version-mismatch` is removed. `decideVerify` no longer takes `cliVersion`. Banner becomes a single line on the happy path: ✔ Genie v4.260507.2 (pid 851758, healthy) pm2 rename: "genie-serve" → "Genie" ----------------------------------- - PM2_PROCESS_NAME = 'Genie' — capital G matches the project brand and no longer blends with the lowercase `genie` CLI invocations operators see in the same `pm2 list` output. - LEGACY_PM2_PROCESS_NAMES = ['genie-serve'] — auto-migration list. `genie install` and `restartServeIfStale` discover entries by canonical OR legacy name. First post-rename install/update deletes the legacy entry and registers the canonical one. - PM2_LOG_PREFIX = 'genie-serve' — pinned independently so existing log-rotation rules referencing `genie-serve-{out,error}.log` keep working across the rename. Version field in pm2 listing ---------------------------- The N/A in the version column happened because pm2 walks the SCRIPT directory's package.json — `~/.bun/bin/genie` resolves to `.../@automagik/genie/dist/genie.js` and pm2 looks at `dist/`, which has no package.json. Fixed by: 1. `readGenieVersionFromDisk()` walks up from the resolved binary path until it finds a `@automagik/genie` package.json (mirrors the resolver in `src/lib/version.ts` — but reads at WRITE time, not import time, so it tracks bun's package swaps). 2. `buildEcosystemConfigSource` bakes the resolved version into the ecosystem config's `version` field. 3. `regenerateEcosystemConfig()` re-writes the config on every update, and `restartServeIfStale` does `pm2 startOrReload <config>` instead of plain `pm2 restart` — so the new `version` value lands in pm2 metadata without manual delete-and-recreate. Now `pm2 list` shows: │ Genie │ default │ 4.260507.2 │ fork │ 851758 │ ... │ online │ ... │ Validation: 129/129 unit tests pass, typecheck clean, lint clean (the serve.test.ts complexity warning is pre-existing). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…coped per reviewer Lands the wish doc that scaffolds PR-A (#1634) and PR-B (#1636/#1637/#1638/ #1640/#1642), plus the 2026-05-07 PR-C draft + reviewer FIX-FIRST corrections. Why this is a separate docs commit: - The wish file was authored 2026-05-04 but only ever sat in a stash; never committed despite shipping work referencing it. This commit lands the reference document for completed + pending work in one place. - PR-C as originally drafted had three invalid premises against live 4.260507.1 (G3 amendment already implemented at scheduler-daemon.ts:1296; G9 line is on stderr not stdout; G10 design assumes binary-spawn that the HTTP probe doesn't do). Reviewer corrections folded in. - Only G8 (kill-path shadow+UUID dedup) survives intact — file path corrected to src/term-commands/agents.ts:2817 (handleWorkerKill). - G9 reframed as stderr-noise reduction (DEBUG=pgserve gating). - G10 deferred pending /trace into update.ts:362. QA dogfooding-72h artifacts (AUDIT.md, QA-PLAN.md) document the 72-h fix-audit sweep that surfaced the bugs and triggered the wish update. Refs: #1677 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
#1677 (G8) Killing a `dir:<name>` shadow OR its paired UUID twin now removes both halves of the logical agent in one atomic transaction. Today's behavior left the other half alive in `genie ls`, forcing operators into a second kill to clean up the visual zombie (proven on 2026-05-07: 7× `dir:codex-*` kills left 7× UUID twins in `error` state). - New helper `killAgentWithDedup` in `src/term-commands/agents.ts` issues a single transactional cascade: `dir:` kill → all UUID twins for the same (name, team); UUID kill → `dir:` shadow when no other UUIDs share the name. - New audit event `agent.kill.dedup_paired { matched, paired }` fires once per cascade for forensic traceability. - `--keep-paired` escape hatch preserves today's single-row behavior for the rare case an operator wants the surviving half to study. - Tests at `src/term-commands/agents.test.ts -t "kill dedup paired"` cover: kill-dir cascade, kill-UUID cascade, both `--keep-paired` variants, and cross-team isolation under migration 061's unique constraint. Closes #1677 (G8 of wish cli-noise-and-hygiene-cleanup PR-C). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…closes #1677 (G9) Every CLI invocation that touches the DB printed `[pgserve] connected to <db>` on stderr. The line is on stderr (so JSON-on-stdout pipelines still work) but clutters every operator terminal. Gate it behind `DEBUG=pgserve`, matching the G1 pg-seed pattern; default-mode operator terminals stay quiet, debug recovery still works. - `src/lib/db.ts:maybePrintBanner` now requires `process.env.DEBUG?.includes('pgserve')` before emitting. Other audit-worthy stderr writes in the same file (retention warnings, pgserve cwd-pin failures, GENIE_PROFILE_DB instrumentation) get explicit `// emit-discipline: ok — <reason>` markers. - New `_resetBannerForTest` export keeps the module-level `bannerPrinted` flag testable without touching production paths. - `tools/lint/emit-discipline-connection.ts` adds a CI gate that flags any new informational `process.stderr.write` / `console.error` in connection/bootstrap modules without an exemption marker. Wired into `bun run check:fast` via `scripts/lint-emit-discipline.ts`. - Tests at `src/lib/db.test.ts -t "no default stderr emit on connect"` pin the gating contract: default mode silent, `DEBUG=pgserve` (and comma-list variants) recover the line, plus a defense-in-depth source-string check that fails if the gate is ever removed. Live verified on dist/genie.js: ./dist/genie.js ls --json 2>&1 1>/dev/null # silent DEBUG=pgserve ./dist/genie.js ls --json 2>&1 1>/dev/null # one banner line Closes #1677 (G9 of wish cli-noise-and-hygiene-cleanup PR-C). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…paired lookup Two findings from Codex review on PR #1685: - **P1 (high):** killing a UUID owned by team B was deleting `dir:<name>` even when the dir shadow belonged to team A. The dir-shadow delete now requires `team IS NOT DISTINCT FROM ${team}` so a UUID kill in one team can never orphan another team's directory identity. New regression test: `UUID kill respects team scope when dir shadow lives in another team`. - **P2 (medium):** the dir-kill paired lookup matched the dir row itself when legacy shadows carry `custom_name = <name>` alongside the `dir:` prefix — emitting a false "paired row(s) also removed" message and a spurious `agent.kill.dedup_paired` audit event for what is really a single-row delete. The lookup now excludes the matched row and any other `dir:%` ids. New regression test: `dir kill with no UUID twins reports zero paired even when dir.custom_name is set`. Also fixes the pgserve v2 smoke step in CI: G9 silenced the `[pgserve] connected` banner by default, so the smoke needs to opt in via `DEBUG=pgserve` to keep asserting the connection round-trip without re-introducing operator stderr noise. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
fix(update): truthful verify probe + rename pm2 service to "Genie" with version
fix(cli-hygiene): kill-path dedup + pgserve stderr gate (G8 + G9, closes #1677)
GitHub deprecated Node.js 20 actions on 2026-09-19 (Node 20 removed from runners on 2026-09-16, force-default to Node 24 on 2026-06-02). The release workflow has been emitting deprecation annotations on every run. Bumps: - actions/checkout v4 -> v5 (8 occurrences across 7 workflows) - actions/setup-node v4 -> v5 (1 occurrence in version.yml) - slsa-framework/slsa-github-generator/.github/workflows/generator_generic_slsa3.yml v2.0.0 -> v2.1.0 (release.yml) oven-sh/setup-bun@v2 already runs on the runner's default Node and is unaffected by this deprecation. No source changes; no functional changes; workflow YAML only. Validation fires on the next push to main. Refs: https://github.blog/changelog/2025-09-19-deprecation-of-node-20-on-github-actions-runners/ Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
fix(ci): bump GitHub Actions to Node.js 24 LTS
Spawned Claude Code agents had no AskUserQuestion in their permissions.allow list, so calling the tool routed through the team-lead approval queue. Felipe saw "Waiting for team lead approval" popups for a tool whose entire purpose is to ask the operator a question. Root cause: AskUserQuestion was not seeded in any of the four allow-list sources used by genie at spawn time. Fix sites: 1. src/lib/claude-settings.ts — ensureClaudeSettingsSafe() now seeds GENIE_BASELINE_ALLOWED_TOOLS into ~/.claude/settings.json on every call (idempotent). Path resolution moved to call-time so test isolation via process.env.HOME works under Bun. 2. src/lib/providers/claude-sdk-permissions.ts — PRESET_READ_ONLY and PRESET_CHAT_ONLY now include AskUserQuestion (PRESET_FULL already covers it via the '*' wildcard). 3. src/lib/provider-adapters.ts — buildSettingsObject() always emits permissions.allow with the baseline merged into any explicit allow list. Existing deny rules are preserved verbatim. 4. src/lib/team-lead-command.ts — buildTeamLeadCommand() now emits --settings carrying the baseline so the team-lead doesn't route its own user-prompt UI through its own approval queue. PreToolUse hook handlers (brain-inject, runtime-emit-tool, session-sync-tool) already pass through unknown tools unchanged — no deny path exists for AskUserQuestion in src/hooks/. Test plan: - src/lib/claude-settings.test.ts (new) — 5 tests covering fresh install, append-to-existing, dedup, key preservation, idempotency. - src/lib/provider-adapters.test.ts — updated 3 existing tests + added 3 regression tests (baseline always present, no duplicate when explicit, empty config still emits baseline). - src/lib/team-lead-command.test.ts — added 2 tests asserting --settings emits AskUserQuestion in permissions.allow. - src/lib/providers/__tests__/claude-sdk-permissions.test.ts — updated 2 preset tests to expect baseline. Validation: bun run typecheck clean; lint clean (only pre-existing warnings on dev for db.test.ts/serve.test.ts unrelated to this change); 97 tests pass across the touched files.
…ssion PR #1692 introduced `process.env.HOME ?? homedir()` in ensureClaudeSettingsSafe so the test could pivot HOME via a tmpdir. That worked locally but broke CI: team-lead-command.test.ts has a long-standing afterEach that hardcodes `process.env.HOME = '/home/genie'` (Felipe's workstation HOME). On CI runners that path doesn't exist, so when claude- native-teams.test.ts later called ensureNativeTeam → ensureClaudeSettingsSafe, mkdirSync('/home/genie/.claude') threw EACCES and cascaded into 31 failures across claude-native-teams + claude-sdk + claude-sdk-omni-executor. Fix: - Revert ensureClaudeSettingsSafe to use the cached CLAUDE_SETTINGS_FILE / homedir() — no HOME pivoting in production code. - Export ensureBaselineAllowedTools as a pure helper so tests can exercise the merge logic against an in-memory object instead of fighting the cached module path. - Rewrite claude-settings.test.ts to test ensureBaselineAllowedTools directly: 6 cases (fresh, append, dedup, preserve unrelated keys, drop non-strings, idempotent) + 2 invariant assertions on the constant. No production behavior change vs PR #1692 — the file-I/O wrapper still seeds the baseline through ensureBaselineAllowedTools, and the spawned-agent / team-lead / SDK preset code paths are untouched.
…default-allow fix(permissions): allow AskUserQuestion by default — closes #1688
Why: when an agent closes a turn (or a team-lead marks a wish group done), the only audit-trail entry today is "Agent <uuid> killed" with no actor, no rationale, no summary. When auto-cleanup cascades on wish completion the parent/orchestrator sees N agents vanish with zero context. The "Run \`genie team done\` to clean up" message in notifyWaveCompletion is also misleading — the next line in doneCommand calls autoCleanupTeam() unconditionally, so the message implies cleanup is pending while the team is already disbanded. Change: add -r/--report <message> as a required option on both \`genie done [ref]\` and \`genie wish done <ref>\`. Validation lives in the action handlers (not Commander's requiredOption) so we emit a multi-line friendly hint with examples instead of Commander's generic missing-option error. The report flows through: - turnClose() reason for agent-session closes (already supported, was unused for outcome=done) - notifyWaveCompletion mailbox message so the orchestrator sees WHAT shipped, not just WHICH groups closed - console output of doneCommand for terminal observers Wave/wish-complete notification now also names auto-cleanup honestly: "Team will be auto-cleaned. Run \`genie team done\` to confirm or override." instead of implying nothing has happened. This does not change the auto-cleanup behavior itself — that's a separate over-reach (kills team members unrelated to the completed wish, including workspace primary agents) worth a follow-up that scopes killTeamMembers to the wish_slug of the completed wish. Tests: 14 done.test.ts cases pass (12 existing + 2 new for the mandatory-report path). Bundle builds clean.
The previous wording ("one-line summary of what you did") understated
what the report is for. The report is the orchestrator's primary view
into a closing turn — the only summary anyone reading later will see
without replaying the transcript. A one-liner is almost never enough.
- CLI hint now asks for a structured handoff: goal attempted, what
shipped, verified vs unverified, what's left or deferred, surprises.
- Help text and wish-done error mirror the same structure.
- Multi-line reports render as fenced "--- Handoff ---" blocks in
both console output and the wave/wish-complete mailbox so the
structure survives renderers and the indented "Report:" prefix
doesn't mangle line breaks.
- Help text explicitly tells the user to pass via heredoc for
multi-line reports — the natural path for a real session summary.
No API changes; tests still pass (14/14).
feat(done): make --report mandatory on every close
Brings main's session-id writer hotfix (#1698) and recover-orphans CLI (#1699) onto dev so the next dev → main PR triggers Version workflow's @latest npm publish (gated on '/dev' in commit message). Conflict resolutions: - src/genie-commands/session.ts: kept BOTH _deps injection from #1698 AND findOrCreateAgent UUID identity from wish #175 G3. Hotfix's claudeSessionId plumbing into createAndLinkExecutor preserved. - src/genie.ts: additive — recover-orphans subcommand registered. - src/lib/agent-directory.ts, executor-registry.ts, protocol-router.ts: surrounding context kept consistent with both branches' direction. - src/__tests__/agent-team-inheritance.test.ts: adapted seedTemplate helper to post-migration-061 UUID-id + name lookup schema. Carries main's other in-flight fixes: - migrations 054 + 055 (subagent team inheritance, auto_resume default) - agent-team-inheritance test fixture (132 LOC) - release.yml + 044 test refinements
|
| GitGuardian id | GitGuardian status | Secret | Commit | Filename | |
|---|---|---|---|---|---|
| 32492444 | Triggered | GitHub Personal Access Token | 1baaf6e | src/hooks/tests/redaction.test.ts | View secret |
🛠 Guidelines to remediate hardcoded secrets
- Understand the implications of revoking this secret by investigating where it is used in your code.
- Replace and store your secret safely. Learn here the best practices.
- Revoke and rotate this secret.
- If possible, rewrite git history. Rewriting git history is not a trivial act. You might completely break other contributing developers' workflow and you risk accidentally deleting legitimate data.
To avoid such incidents in the future consider
- following these best practices for managing and storing secrets including API keys and other credentials
- install secret detection on pre-commit to catch secret before it leaves your machine and ease remediation.
🦉 GitGuardian detects secrets in your source code to help developers and security teams secure the modern development process. You are seeing this because you or someone else with access to this repository has authorized GitGuardian to scan your pull request.
|
Important Review skippedToo many files! This PR contains 291 files, which is 141 over the limit of 150. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: ⛔ Files ignored due to path filters (9)
📒 Files selected for processing (291)
You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a4aad263c1
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // The DB name is per-app fingerprint; genie defaults to 'postgres' when | ||
| // the host hasn't run autopg create-app — install sets the real name. | ||
| // For the migration we use the pgserve discovery URL. | ||
| return `postgresql://postgres:postgres@127.0.0.1:${CANONICAL_PORT}/postgres`; |
There was a problem hiding this comment.
Preserve Genie database name in migration URL
Migration 001-pm2-env-databaseurl-bake hardcodes .../postgres when setting DATABASE_URL, but the installer’s canonical connection string uses .../genie (src/genie-commands/install.ts line 239). On hosts where this migration runs (legacy pm2 env missing), it can silently repoint genie-serve to a different database, causing existing Genie state to disappear from the service view and potentially running against the wrong schema/data set.
Useful? React with 👍 / 👎.
| function findLegacyEmbedded(): ListeningPg | undefined { | ||
| if (process.env.GENIE_KEEP_LEGACY_PG === '1') return undefined; | ||
| if (!canonicalReachable()) return undefined; | ||
| return listListeningPgserve().find((p) => p.port !== CANONICAL_PORT); |
There was a problem hiding this comment.
Restrict legacy pg kill target to Genie-managed instances
findLegacyEmbedded treats any local postgres listener on a non-8432 port as “legacy embedded” and later terminates that PID. Because there is no check that the process belongs to Genie/pgserve-managed data dirs, running genie migrate on a host with another legitimate PostgreSQL service (e.g. local dev DB on 5432) can kill an unrelated database process.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Code Review
This pull request implements a comprehensive set of updates, including new feature designs ("wishes"), execution reports, and significant infrastructure hardening. Key highlights include the introduction of a host-migration framework, the normalization of observability signals, and the unification of update and output primitives across the Genie and Omni CLIs. Additionally, the PR lays the groundwork for a major v5 release by defining a CDN-based distribution strategy and a data-safe cutover process. Review feedback correctly points out that error handling in the backend observability endpoints should include logging to aid debugging and suggests merging duplicated sections in the changelog to maintain a standard format.
| ## v4.x.x — Host Migrations | ||
|
|
||
| **Added:** `genie migrate` CLI verb — versioned, applied-once host-state migrations that detect and fix drift between current host state and current code expectations (pm2 env blocks, embedded pgserve fantasmas, config drifts). Mirrors the DB-migrations pattern but for the HOST itself. | ||
|
|
||
| **Added:** npm `postinstall` hook (`scripts/postinstall-migrations.js`) auto-runs `genie migrate --quiet` after `bun add -g @automagik/genie@latest`. Soft-fails so package install never breaks; manual `genie migrate` remains the explicit escape hatch. | ||
|
|
||
| **Initial migrations shipped:** | ||
| - `001-pm2-env-databaseurl-bake` — re-applies the bake-DATABASE_URL fix (commit 5567e202) on hosts with pm2 genie-serve env missing the variable | ||
| - `002-kill-embedded-pgserve-legacy` — stops legacy embedded pgserve listening on non-canonical ports when canonical 8432 is healthy | ||
|
|
||
| **Contract:** Users upgrading to genie@>=4.260503.x get host-state migrations applied transparently via postinstall. Manual `genie migrate` remains as the explicit escape hatch for forced re-runs. | ||
|
|
||
| **Override:** Set `GENIE_SKIP_MIGRATIONS=1` to bypass the hook (CI / containers / install-only flows). | ||
|
|
||
| **Tracking:** Applied migrations recorded in `~/.genie/migrations.json` (atomic write, file-based to avoid PG dependency during early-boot self-heal). | ||
|
|
||
| # Changelog | ||
|
|
||
| ## Unreleased | ||
|
|
||
| ### Fixed | ||
|
|
||
| - TUI startup no longer crashes with opaque `output: [null, null, null]` when | ||
| an existing `genie-tui` session has unexpected layout. `startTuiTmuxServer` | ||
| now probes with `has-session` first, recovers corrupt sessions via | ||
| `kill-session` + fresh create (logging the original cause to | ||
| `~/.genie/logs/tui-crash.log`), and surfaces tmux's actual stderr | ||
| (e.g. `duplicate session: genie-tui`) in any error that does bubble up. | ||
| Wish: `genie-tui-startup-resilience`. | ||
|
|
||
| ### Breaking — pgserve canonical cutover (consumer-only, pm2-supervised) | ||
|
|
||
| - **Genie no longer spawns pgserve.** The pre-canonical genie was a daemon | ||
| *owner*: `getOrStartDaemon` would spawn `pgserve daemon` as a detached | ||
| child, `selfHealPostgres` would `pkill -9` postgres backends to recover | ||
| from stuck state, and `genie serve start` treated pgserve startup as | ||
| part of its boot sequence. Canonical `pgserve@^2` is a pm2-supervised | ||
| singleton (`pgserve install` registers it) — every `pkill -9` from the | ||
| old self-heal triggered an immediate pm2 respawn, producing the | ||
| "Could not kill stale postgres processes" + "pgserve v2 daemon exited | ||
| before binding" fight-with-pm2 cycle that motivated this cutover. | ||
| - **`getOrStartDaemon` → `requirePgserveDaemon`.** Probe-only: succeeds | ||
| when the canonical socket is reachable, throws a pm2-recovery hint | ||
| (`pm2 status` / `pm2 restart pgserve` / `pgserve install`) otherwise. | ||
| The pre-cutover `getOrStartDaemon` symbol is **removed** in this | ||
| release (a deprecation alias was considered but the project's | ||
| `dead-code` (knip) gate doesn't honour `@deprecated`; downstream | ||
| callers should rename to `requirePgserveDaemon` — the new contract | ||
| is documented above and matches the throw-on-unreachable behaviour | ||
| the deleted Mode B/C paths intermittently produced anyway). | ||
| - **`genie install` is now fatal on canonical pgserve failure.** No more | ||
| warn-and-continue. Operators see a copy-paste recovery hint: | ||
| ``` | ||
| Error: canonical pgserve registration failed (<reason>). | ||
| Genie depends on pm2-supervised pgserve. To proceed: | ||
| bun add -g pgserve@^2 | ||
| pgserve install | ||
| genie install | ||
| ``` | ||
| - **`genie doctor --fix` no longer pkills postgres processes.** The old | ||
| `killStalePostgres` step is replaced by a hint-only | ||
| `printPgserveRecoveryHint` that prints pm2 commands and exits. | ||
| Operators run them manually if needed. | ||
| - **`genie serve start` uses `requirePgserveReady` (probe-only).** On | ||
| success: `pgserve daemon ready (canonical, pm2-supervised) on | ||
| <socket>`. On failure: clear pm2-recovery hint + sets | ||
| `GENIE_PG_NO_AUTOSTART=1` so subsequent code doesn't loop on the same | ||
| failure. | ||
| - **Deleted from `src/lib/db.ts`** (`~745 LOC` removed): | ||
| `startPgserveDaemonOnce`, `evictOrphanDataDirHolder`, | ||
| `detectOrphanDataDirLock`, `terminatePgserveTree`, `signalPgserveTree`, | ||
| `signalPgserveDaemonPid`, `recoverUnresponsivePgserveDaemon`, | ||
| `isLikelyPgserveDaemonProcess`, `cleanPartialDaemonState`, | ||
| `removeStalePgserveSocketArtifacts`, `unlinkIfPresent`, | ||
| `waitForDaemonSocket`, `formatPgserveDaemonCommand`, | ||
| `spawnPgserveDirect`, `startPgserveOnPort`, `findPgserveBin`, | ||
| `findPgserveDaemonCommand`, `findLocalPgserveRoot`, | ||
| `resolvePgservePackageCommand`, `findBunRuntime`, | ||
| `selfHealPostgres`, `waitForDaemonPort`, `throwDaemonTimeout`, | ||
| the `PgserveDaemonCommand` interface, and the | ||
| `lastAutoStartOutcome`/`lastAutoStartPid` tracking. | ||
| - **Migration for pre-canonical operators:** | ||
| ```bash | ||
| # 1. Install canonical pgserve (pm2-supervised singleton) | ||
| bun add -g pgserve@^2 | ||
| pgserve install # registers under pm2; auto-detects host | ||
| # If you have existing data at ~/.genie/data/pgserve and want to keep | ||
| # it, point pgserve install at that data dir BEFORE the cutover: | ||
| pgserve install --data ~/.genie/data/pgserve | ||
|
|
||
| # 2. Re-run genie install (fails fatally if pgserve isn't ready) | ||
| genie install | ||
|
|
||
| # 3. Verify | ||
| genie doctor # all [ok] for pgserve preconditions | ||
| ``` | ||
|
|
||
| ### Breaking — pgserve v2 (Unix socket, auto-fingerprint, no credentials) | ||
|
|
||
| - **Switched to pgserve v2 daemon model.** Genie now connects to pgserve | ||
| via the well-known Unix control socket at | ||
| `$XDG_RUNTIME_DIR/pgserve/.s.PGSQL.5432` (fallback `/tmp/pgserve/...`) | ||
| instead of TCP loopback. The pgserve v2 daemon authenticates the peer | ||
| via `SO_PEERCRED`, derives a stable fingerprint from the nearest | ||
| ancestor `package.json` (`sha256(realpath + name + uid)[:12]`), and | ||
| routes the connection to that fingerprint's own | ||
| `app_<sanitized-name>_<12hex>` database. As a consumer, genie no | ||
| longer specifies a port, user, or password. | ||
| - **`PGHOST`/`PGPORT`/`PGUSER`/`PGPASSWORD` env vars removed.** The only | ||
| variable still set when shelling out to `pg_dump` / `psql` is `PGHOST`, | ||
| pointing at the v2 socket directory. Auth happens at the kernel layer. | ||
| - **`pgserve.persist: true` declared in `package.json`.** Genie holds | ||
| long-lived state (wishes, agents, events). The persist flag opts the | ||
| database out of pgserve v2's default 24h TTL reaper, so a restarted | ||
| daemon doesn't drop the wishes table after a quiet weekend. | ||
| - **Visible fingerprint banner on boot.** First successful connection in | ||
| a process prints `[pgserve] connected to <db>` to stderr so | ||
| developers can see the routed database name (e.g. | ||
| `app_genie_a1b2c3d4e5f6`). Set `GENIE_NO_BANNER=1` to suppress. | ||
| - **Migration note for self-hosted deployments.** pgserve v2 expects an | ||
| externally-supervised daemon (PM2 / systemd snippets in pgserve | ||
| README). The legacy `genie serve` headless spawn path remains as a | ||
| TCP fallback for environments that haven't adopted the daemon yet — | ||
| set `GENIE_PG_FORCE_TCP=1` to opt back into TCP loopback. | ||
| - **`pgserve@2.0.0` consumed from npm.** The temporary local file pin | ||
| (`file:../pgserve`) has been swapped for the published `^2.0.0` | ||
| range now that pgserve v2 is on the registry. Genie tracks the | ||
| daemon's released artifact rather than a sibling working copy. | ||
|
|
||
| ### Breaking — design system |
There was a problem hiding this comment.
The new changelog entries have been prepended to the file, resulting in a duplicated structure with two # Changelog and ## Unreleased sections. To maintain a clean and standard changelog format, the new entries should be merged into the existing structure under the ## Unreleased section.
For example, the ## v4.x.x — Host Migrations section could become a subsection under ### Added.
| FROM agents a | ||
| ORDER BY a.team, a.custom_name | ||
| `, | ||
| listAgentObservability({ includeHarness: true }).catch(() => []), |
There was a problem hiding this comment.
While catching the error to prevent the endpoint from failing is a good defensive measure, silently swallowing the error can make debugging difficult if issues arise with listAgentObservability. It would be beneficial to log the error to provide visibility into potential problems.
listAgentObservability({ includeHarness: true }).catch((err) => {
console.error('Failed to fetch agent observability data:', err);
return [];
}),| WHERE agent = ${agent.custom_name ?? params.agent_id} | ||
| ORDER BY id DESC LIMIT 50 | ||
| `, | ||
| getAgentObservability(params.agent_id).catch(() => null), |
There was a problem hiding this comment.
Similar to the agents.list handler, silently swallowing errors from getAgentObservability can make debugging difficult. It is better to log the error for visibility while still returning null to maintain endpoint stability.
getAgentObservability(params.agent_id).catch((err) => {
console.error('Failed to fetch observability data for agent ' + params.agent_id + ':', err);
return null;
}),| }> { | ||
| const [agents, observability] = await Promise.all([ | ||
| listAgents(), | ||
| listAgentObservability({ includeHarness: true }).catch(() => []), |
There was a problem hiding this comment.
Silently swallowing errors from listAgentObservability can hide underlying issues. It is better to log the error for debugging purposes while still returning an empty array to ensure the function remains resilient.
listAgentObservability({ includeHarness: true }).catch((err) => {
console.error('Failed to list agent observability data:', err);
return [];
}),| } | ||
|
|
||
| return { agent, executor }; | ||
| const observability = await getAgentObservability(id).catch(() => null); |
There was a problem hiding this comment.
Errors from getAgentObservability are silently swallowed here. For better diagnostics, it is advisable to log the error before returning null.
const observability = await getAgentObservability(id).catch((err) => {
console.error('Failed to get observability data for agent ' + id + ':', err);
return null;
});
Release: dev → main
Promote dev (
a4aad263c— 4.260507.7) to main (e3d839c10— currently 4.260507.x post #1699).main is now an ancestor of dev (Felipe's
58228e35d chore(merge): main into dev — sync hotfixes #1698 + #1699 forwardalready pulled main hotfixes forward), so this PR is a clean fast-forward — no conflicts.Scale
plugins/genie/.orphaned_at— internal sync marker), 219 modifiedHeadlines (most recent first)
genie doneandgenie wish donenow requires a structured handoff message (goal, what shipped, verified, left, surprises). Lands as--- Handoff ---blocks in events + wave/wish-complete mailbox so closures stop being silent. Plus honest auto-cleanup wording.…plus ~440 more commits (auto-version bumps, fixes, doc updates).
Notable schema
Test plan
bun run build— clean on dev tipbun dist/genie.js --help—done,recover-orphans, all wave/wish primitives presentdone --reportround-trip on a real wish + observability dashboards still ingest cleanlyNote
This supersedes the previously-closed Release 4.260429.13 PR (#1446) — version bumped, scope grew while #1446 was open.