From a4428ec2e2cf4869968999014df9679857761c0b Mon Sep 17 00:00:00 2001 From: namastex888 Date: Thu, 9 Jul 2026 20:18:19 -0300 Subject: [PATCH 1/8] =?UTF-8?q?docs(wishes):=20council-workflow=20design?= =?UTF-8?q?=20+=20wish=20(SHIP-reviewed=20x2)=20=E2=80=94=20/council=20nat?= =?UTF-8?q?ive=20workflow=20engine?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .genie/INDEX.md | 4 +- .genie/brainstorms/council-workflow/DESIGN.md | 106 +++++++++++ .genie/brainstorms/council-workflow/DRAFT.md | 112 +++++++++++ .genie/brainstorms/skill-absorbs/DRAFT.md | 1 + .genie/wishes/council-workflow/WISH.md | 175 ++++++++++++++++++ .../validate/g1-lane-skills.sh | 19 ++ .../council-workflow/validate/g2-engine.sh | 33 ++++ .../council-workflow/validate/g3-cutover.sh | 24 +++ .../council-workflow/validate/g4-consumers.sh | 32 ++++ .../council-workflow/validate/g5-gate.sh | 14 ++ 10 files changed, 519 insertions(+), 1 deletion(-) create mode 100644 .genie/brainstorms/council-workflow/DESIGN.md create mode 100644 .genie/brainstorms/council-workflow/DRAFT.md create mode 100644 .genie/wishes/council-workflow/WISH.md create mode 100755 .genie/wishes/council-workflow/validate/g1-lane-skills.sh create mode 100755 .genie/wishes/council-workflow/validate/g2-engine.sh create mode 100755 .genie/wishes/council-workflow/validate/g3-cutover.sh create mode 100755 .genie/wishes/council-workflow/validate/g4-consumers.sh create mode 100755 .genie/wishes/council-workflow/validate/g5-gate.sh diff --git a/.genie/INDEX.md b/.genie/INDEX.md index a4407e13b..286936dbe 100644 --- a/.genie/INDEX.md +++ b/.genie/INDEX.md @@ -3,7 +3,7 @@ ## Raw - [control-plane-contract](brainstorms/control-plane-contract/DRAFT.md) — single executable dispatch+routing contract; global↔repo lifecycle convergence by layer; work/review policy refactor (umbrella G2+G3, 2026-07-09) -- [skill-absorbs](brainstorms/skill-absorbs/DRAFT.md) — trace→fix, wizard→genie, pm→work-ref, council→lens library+thin route, report→LangWatch (umbrella G4, 2026-07-09) +- [skill-absorbs](brainstorms/skill-absorbs/DRAFT.md) — trace→fix, wizard→genie, pm→work-ref, council→[council-workflow](brainstorms/council-workflow/DESIGN.md) (poured), report→LangWatch (umbrella G4, 2026-07-09) - [always-on-genie](brainstorms/always-on-genie/DRAFT.md) — SessionStart identity/state inject, hook contract w/ fixtures, worktree isolation policy (umbrella G5+G10, 2026-07-09) - [genie-spend](brainstorms/genie-spend/DRAFT.md) — LangWatch-backed spend report + decision-level cost join (umbrella G7, 2026-07-09) - [dream-replatform](brainstorms/dream-replatform/DRAFT.md) — scheduler adapter + genie.db ledger; cron = trigger never authority; omni approval gates (umbrella G9, 2026-07-09) @@ -16,6 +16,7 @@ ## Ready +- [WISH: hook-injection-hardening](wishes/hook-injection-hardening/WISH.md) — BLOCKED-clearing safety edit: `execFileSync` at 3 hook sites (audit-context, freshness×2) + hostile-filename regression tests + `core.bare` probe removal; flips panel verdict BLOCKED→FIX-FIRST — plan review SHIP (1 fix loop), /work user-gated (2026-07-09) - [WISH: v5-completion](wishes/v5-completion/WISH.md) — CLAUDE.md-for-v5 rewrite ∥ Codex launch target + Hermes decision ∥ distribution 5.x (drafted 2026-07-02) - [WISH: dispatch-inproc-default](wishes/dispatch-inproc-default/WISH.md) — HIGH discovered defect: v5 hooks fall open by default (daemon deleted in demolition); re-arm branch-guard + unblock omni approvals (drafted 2026-07-02) - [WISH: omni-approval-ux](wishes/omni-approval-ux/WISH.md) — correlated approval identity, reaction approve/deny, anti-spam feedback; grounded in the 2026-07-03 live WhatsApp QA (drafted 2026-07-03) @@ -23,6 +24,7 @@ - [WISH: warp-integration](wishes/warp-integration/WISH.md) — umbrella Group 3: genie init, Warp launch-config emitter, genie launch, /work multi-session opt-in (drafted 2026-07-02) ## Poured +- [council-workflow](brainstorms/council-workflow/DESIGN.md) · [WISH](wishes/council-workflow/WISH.md) — /council vira saved workflow nativo (deliberation + audit), lens library 13 (7 persona skills renomeadas por lane + 6 cards), distribuição via install-stamp em ~/.claude/workflows; implementa a disposição council do skill-absorbs G4 — design review SHIP + plan review SHIP (todos os MEDIUMs aplicados; G3/G4 gated em fable5 MERGED), `/work` user-gated (2026-07-09) - [plugin-resource-shipping](brainstorms/plugin-resource-shipping/DRAFT.md) · [WISH](wishes/plugin-resource-shipping/WISH.md) — in-skill template via ${CLAUDE_SKILL_DIR}, probe-guarded lint refs, resource-shipping lint, fresh-install CI smoke — plan review SHIP (2 fix loops), /work user-gated, G1 dep on routing-matrix:3 (2026-07-09) - [routing-matrix](brainstorms/routing-matrix/DRAFT.md) · [WISH](wishes/routing-matrix/WISH.md) — pinned role agents (Fable gates / Opus ladder / Haiku scouts), stage pins + pane flags, complexity columns + lint, escalation caps — executed; execution review SHIP, final gate 719 pass / 1 skip / 0 fail; live LangWatch pin QA pending (2026-07-09) - [genie-token-efficiency-program](brainstorms/genie-token-efficiency-program/DESIGN.md) — umbrella (WRS 100, independent review SHIP; 9 seed groups): genie → control-plane thesis; routing matrix (Fable gates, Opus ladder, no Sonnet, $17.9k/21d baseline); 17-skill dispositions; always-on identity; cross-agent delegate (Codex/Hermes companion sessions); genie spend; brainstorm domain-map upgrade — crystallized 2026-07-09 diff --git a/.genie/brainstorms/council-workflow/DESIGN.md b/.genie/brainstorms/council-workflow/DESIGN.md new file mode 100644 index 000000000..dfe8b5701 --- /dev/null +++ b/.genie/brainstorms/council-workflow/DESIGN.md @@ -0,0 +1,106 @@ +# Design: /council — Native Workflow Engine (Deliberation + Audit) + +| Field | Value | +|-------|-------| +| **Slug** | `council-workflow` | +| **Date** | 2026-07-09 | +| **WRS** | 100/100 | + +## Problem + +Genie's multi-perspective reasoning is split across two model-driven orchestrators — `skills/council/` (topic deliberation, Agent-tool turn-by-turn) and Felipe's personal `specialist-panel` skill (7-lane repo audit) — which duplicate the same fan-out→synthesize pattern in prompt form: fragile, non-resumable, context-hungry, and (in the panel's case) not shipped with the product at all. + +## Scope + +### IN +- `council.js` — one native dynamic-workflow script (saved workflow, `meta.name: 'council'`), two modes: `deliberation` (default: `/council `) and `audit` (`/council audit [focus]`). +- Unified 13-lens library: 7 lane skills (renamed personas, standalone-invocable) + 6 deliberation lens cards at `plugins/genie/references/lenses/`. +- 7 persona skills migrated into `skills/` with domain names — `repo-hygiene`, `architecture`, `code-quality`, `qa`, `perf`, `supply-chain`, `dx-docs` — each body citing "lens inspired by the work of "; no real person's name as a product identity. +- Distribution: `genie install`/`update` stamps `LENS_ROOT` (absolute installed-plugin path) into the `council.js` template and copies it to `~/.claude/workflows/council.js`; re-stamped on every update. +- Removal of `skills/council/` entirely (SKILL.md, members/routing.md, members/config.md, templates/report.md — all absorbed into the script and lens frontmatter). +- Consumers wired: `/review` gains lens-panel dispatch (multi-lens reviewers by change-type); `/brainstorm` gains a domain-experts step at its Decisions-stuck point that reads lens cards from the library. Both steps are NEW to the repo skills (verified 2026-07-09: neither repo skill mentions lenses/panels today) — the pattern mirrors the global brainstorm skill's lens-subagent step. +- Lints wired into `bun run check`: workflow-script structural lint, lens frontmatter lint, probe-guarded reference lint. +- Docs updates (skills README, plugin docs). + +### OUT +- Rewiring other consumers (`/work`, `/fix`, `/pm`) — stays in skill-absorbs G4 follow-ups. +- Codex/Hermes support for the engine — the workflow runtime is Claude Code-only; accepted. +- `genie council` term-command (CLI) — the surface is the saved workflow, not a CLI verb. +- Per-repo `.claude/workflows/` scaffolding via `genie init` — repo-level override is a documented capability, not a deliverable. +- dream/scheduler autonomy, routing-matrix role-agent changes. +- Deleting Felipe's personal `~/.claude/skills/` copies (his local hygiene, outside the repo). + +## Approach + +`/council` becomes a **saved workflow command**: the model mediates structured `args` (`{mode, topic|focus, membersOverride?}`) at invocation, and the script owns the orchestration. Runtime shape: + +``` +council.js + meta { name:'council', phases:[Resolve, Round 1, Round 2, Synthesis, Persist] } + ROUTING = keyword→members table (absorbed from members/routing.md, incl. default trio + --members override) + LENSES = lens name → LENS_ROOT-relative path (LENS_ROOT stamped at install) + Resolve : stage-0 agent verifies lens paths, Glob-fallback if stale (self-healing), returns absolute paths + Round 1 : parallel members/lanes — each agent Reads its lens file; audit lanes get the evidence contract + (run real commands, findings with severity+evidence, assess-only, return profile updates as data); + deliberation members get the voice contract (2-4 opinionated paragraphs, positions + assumptions) + Round 2 : deliberation only — FRESH agents, each fed its own Round-1 position plus the others', returning + strongest point / challenged point / position change (replaces SendMessage continuation) + Synthesis : deliberation → consensus, tensions, dissent (preserved verbatim), report per absorbed template; + audit → cross-lane dedupe, conflict resolution, global re-rank, "not audited" flags + Persist : audit only — single writer merges lanes' returned profile updates into .genie/repo-profile.md + return structured report data; the main loop renders the user-facing report +``` + +Council invariants preserved in the script: ≥2 members must deliver Round 1 or the run reports failure; silent members recorded as "no response"; advisory-only, no voting; audit stays assess-only with the authority boundary (findings crystallize via `/wish`, never a parallel approval mechanism). + +**Alternatives considered:** thin skill launching the Workflow tool via `scriptPath` (rejected by Felipe: keeps an orchestration skill alive; native command is cleaner); two sibling workflows for audit/deliberation (rejected: duplicates the roster→rounds→synthesis pattern); lens-library-only without standalone skills (rejected: loses standalone fix-mode for genie users). + +**Distribution constraint (verified 2026-07-09 against the plugins reference):** plugins cannot ship workflows — component fields are skills, commands, agents, hooks, mcpServers, outputStyles, lspServers, experimental.themes/monitors, userConfig, channels, dependencies. Hence the install-time stamp+copy to `~/.claude/workflows/`. Documented precedence (project > personal) makes per-repo overrides possible for free. + +## Decisions + +| Decision | Rationale | +|----------|-----------| +| One engine, two modes (not audit-only, not two scripts) | Both are the same fan-out→synthesize pattern; one script + mode presets maximizes reuse and honors the skill-absorbs G4 ruling (lens library + preserved /council entry) | +| Personas absorbed as standalone plugin skills | Single source of truth — the workflow reads the same SKILL.md as its lens; fix-mode ("run code-quality and fix") ships to every genie user | +| `/council` is the single entry, audit via `/council audit` | Council is genie lore; keeps the name users know, closes the open G4 naming GAP | +| Consumers (/review, /brainstorm) wired in this wish | Felipe chose delivering the full G4 vision now; both steps are new to the repo skills; sequencing risk handled by waves | +| Lane names, not people names, in the public plugin | Real living experts' names as speaking product personas without consent is a liability; methodology retained, inspiration cited | +| Saved workflow command, no launcher skill | Most native shape; model mediates args; routing lives as data in the script | +| Install-time stamp+copy to `~/.claude/workflows/` | Plugins verifiably cannot ship workflows; smart-install runs with `CLAUDE_PLUGIN_ROOT` so it can stamp `LENS_ROOT` deterministically and re-stamp on update | +| Fresh-agent Socratic Round 2 | Workflows have no SendMessage; feeding each member its own R1 back preserves identity while gaining resumability and parallelism | +| 13 lenses: 7 lanes + 6 deliberation cards | The 4 redundant council lenses (benchmarker, sentinel, ergonomist, architect) map onto perf, supply-chain, dx-docs, architecture; routing table remapped accordingly | +| `skills/council/` deleted whole | Skill-vs-workflow precedence for the same `/council` name is undocumented — do not risk the collision | + +## Risks & Assumptions + +| Risk | Severity | Mitigation | +|------|----------|------------| +| skills-fable5-revamp in flight on the same skill files | HIGH | Wave 1 touches only NEW files (persona skills, council.js, lenses/); council deletion + review/brainstorm edits gated on fable5 MERGED to its base (review-closed insufficient) + rebase onto the post-merge base | +| Stale `LENS_ROOT` stamp (plugin path changes on update; user skips `genie update`) | MEDIUM | Re-stamp wired into the update flow + stage-0 resolver agent with Glob-fallback (self-healing) | +| Cost: audit ≈ 9-10 agents on the session model | MEDIUM | Lane narrowing via `/council audit `; token visibility in `/workflows`; cost note in docs | +| No mid-run user input — loses the old orchestrator's live supervision | LOW | Schema-validated stage returns; ≥2-members rule enforced in script; failed lanes reported "not audited", never averaged in | +| First run prompts for workflow approval per project | LOW | Documented; "don't ask again" is per-workflow-per-project | +| Workflows need CC ≥ 2.1.154, paid plan, and can be org-disabled (`disableWorkflows`) | LOW | Documented requirement; engine is CC-only by decision — no degraded fallback | + +**Assumptions:** the stamp+copy call site is the SessionStart hook (`smart-install.js`, where `CLAUDE_PLUGIN_ROOT` is set — the `genie install`/`update` CLI shell does not have it), placed before the hook's early-exit guards and idempotent via drift-check, so re-stamp happens on the first session after a plugin update; the model reliably mediates `/council ` into structured workflow args (documented saved-workflow behavior). + +## Execution Groups (seed for /wish) + +| Grupo | Entregável | Depende de | Validação | +|-------|-----------|------------|-----------| +| G1 | 7 lane skills (renamed personas, inspiration cited) | none | frontmatter lint + no-real-names grep gate | +| G2 | council.js template (engine, routing data, schemas) + lens cards + stamp/copy in install/update + structural lint | none | ESM parse + banned-API grep + stamp unit test | +| G3 | Cutover: delete skills/council/, purge references | G1, G2, fable5-revamp execution review closed | git grep gates (council orchestrator + specialist-panel = 0 hits) | +| G4 | Consumers: /review lens panels + /brainstorm domain-experts (both steps new to the repo skills) | G3 | probe-guarded reference lint | +| G5 | Lints wired into `bun run check`, docs, live QA (1 deliberation + 1 audit run recorded) | G3, G4 | `bun run check` green + QA evidence files | + +## Success Criteria + +- [ ] `council.js` structural lint passes: `meta.name === 'council'`, ESM parses, zero `Date.now`/`Math.random`/`new Date()`/`require`/`import`/`fs` occurrences +- [ ] Lens lint passes: every ROUTING member maps to an existing lens file; every lens file has required frontmatter (name, modes, voice) +- [ ] 7 lane skills exist under `skills/` with domain names; grep gate proves no real person's name in any `name:` field; inspiration line present in each body +- [ ] `git grep -il 'specialist-panel'` and old-council references return 0 hits outside `.genie/attic/`, CHANGELOG, and this wish's own artifacts +- [ ] `/review` and `/brainstorm` reference lens-library paths that exist (probe-guarded refs lint, fails on dangling path) +- [ ] `bun run check` green with the new lints wired in +- [ ] Live QA evidence in the wish folder: one real `/council ` deliberation run + one `/council audit` run on the genie repo diff --git a/.genie/brainstorms/council-workflow/DRAFT.md b/.genie/brainstorms/council-workflow/DRAFT.md new file mode 100644 index 000000000..0d4f9d2b0 --- /dev/null +++ b/.genie/brainstorms/council-workflow/DRAFT.md @@ -0,0 +1,112 @@ +# DRAFT: council-workflow — specialist-panel as native workflow, replacing council + +> Slug renomeado panel-workflow → council-workflow após decisão 3 (/council é o nome vencedor). + +**Status:** Raw · **Date:** 2026-07-09 · **Related:** [skill-absorbs](../skill-absorbs/DRAFT.md) (umbrella G4 — council ruling), [genie-token-efficiency-program](../genie-token-efficiency-program/DESIGN.md) + +## GOAL (user's words) + +Turn `/specialist-panel` into a native reusable well-thought workflow (Claude Code dynamic workflows, https://code.claude.com/docs/en/workflows.md) and use it to replace the current genie council. + +## KNOWN (evidence) + +### The two things being unified +- `~/.claude/skills/specialist-panel/SKILL.md` (personal, 6.1K): 7-lane repo AUDIT — Chacon/Ousterhout/Hejlsberg/Beck/Gregg/Lorenc/Procida personas as parallel Agent-tool subagents; dedupe → conflict-resolution → global re-rank synthesis; assess-only; writes `.genie/repo-profile.md` (single-writer); authority boundary: findings → `/wish`, never a parallel approval mechanism. +- `skills/council/` (genie repo, shipped via plugin): topic DELIBERATION — 10 lens members, smart routing to 3-4, 2-round Socratic (round 2 via SendMessage continuing sessions), dissent-preserving report, advisory-only, no voting. Members/config: all inherit model. Supporting files: members/routing.md, members/config.md, templates/report.md. +- The 7 persona skills live in `~/.claude/skills/` (personal) — NOT shipped with genie. Full methodologies with ground-truth discovery steps, not just lens one-liners. + +### Prior ruling (skill-absorbs, Felipe 2026-07-09) +- council → lens LIBRARY (consumed by review panels + brainstorm domain-experts) + THIN `/council` route preserved (strategy pre-artifact, dissent preservation, appeal court for reviewer↔gate disagreements). +- Open GAP there: "/council users besides Felipe? keep `/council` name or `genie council`?" +- Collision noted: skills-fable5-revamp execution review pending on same files. +- **This brainstorm is the concrete implementation vehicle for the council disposition** — the "thin route" becomes a launcher for the native workflow engine. + +### Native workflow facts (docs fetched 2026-07-09) +- Script = `export const meta {...}` + JS body; `agent()/pipeline()/parallel()/phase()`; schema-validated structured outputs; per-agent model/effort overrides. +- Saved workflows: `.claude/workflows/` (project) or `~/.claude/workflows/` (personal) → become `/` commands; accept `args` (structured, no parsing). **Plugin-shipped workflows: UNVERIFIED — background check dispatched.** +- Script has NO fs/shell access — agents do all IO. No mid-run user input. Background, resumable (same session), `/workflows` progress view, per-agent token visibility. +- Skill-instructed Workflow launch counts as explicit user opt-in (skill = thin route calling Workflow tool = legit pattern). +- No SendMessage inside workflows → Socratic round 2 must be fresh agents fed round-1 transcripts via prompt (arguably better: no continuation flakiness, resumable, parallel). + +### Distribution context +- Genie ships as plugin `genie` (marketplace `automagik`, repo root `.claude-plugin/marketplace.json`); `plugins/genie/skills -> ../../skills` symlink; plugin also ships agents/ (scout, engineer-*, reviewer, final-gate, fixer — routing-matrix pinned), hooks, rules, references. +- `${CLAUDE_SKILL_DIR}` interpolation in skills is established practice (plugin-resource-shipping wish) → a thin skill can resolve plugin-relative lens paths and pass them as workflow args. + +## TENSIONS / OPEN + +1. **Replace semantics (Q1 → asked):** unified parameterized panel engine (audit + deliberation modes, one script) vs audit-only port (council's Socratic mode dies) vs two sibling workflows sharing lens library. +2. **Persona custody:** absorb 7 heavyweight persona skills into genie's lens library (genie owns; dual-maintenance vs personal copies) vs args-driven "bring your own lenses" with genie shipping only lightweight lenses. +3. **Invocation UX + naming:** `/council` vs `/panel`; saved-workflow-command vs thin-skill-launches-Workflow. Depends on plugin-workflow verification (background agent). +4. **Fate of personal `~/.claude/skills/specialist-panel`** after genie ships the engine. +5. **Fix-routing for non-Felipe users:** panel prescribes "run to fix" — genie-native answer is findings → `/wish` (already the panel's genie-repo path). Confirm. +6. **Repo-profile write:** keep single-writer merge — in workflow-land the synthesis/persist stage agent is the writer. +7. **Sequencing/collisions:** skills-fable5-revamp execution review + control-plane-contract touch same skill files; how to record vs skill-absorbs G4 (this supersedes part of it). + +## DECIDED + +1. **Replace semantics = engine unificado, 2 modos** (Felipe, 2026-07-09): UM workflow parametrizado (args: mode, topic/focus, roster). Modo `audit` = specialist lanes sobre repo; modo `deliberation` = council lenses sobre decisão, com round 2 socrático via agents frescos alimentados com round 1. Uma lens library alimenta ambos. `skills/council/` morre como orquestrador; thin `/council` route vira launcher do engine (consistente com ruling skill-absorbs G4). + +2. **Persona custody = absorver como skills do plugin** (Felipe, 2026-07-09): as 7 persona SKILL.mds migram pro repo genie e shipam via plugin. O workflow lê o MESMO arquivo como lens (fonte única, zero drift); cada persona fica invocável standalone pra fix-mode por qualquer usuário genie. Cópias pessoais de Felipe aposentam. Council sai, suas 10 lentes fundem na mesma library (mapear overlaps: benchmarker≈Gregg, sentinel≈Lorenc, ergonomist≈Procida, architect≈Ousterhout...; lentes sem lane par — questioner, operator, deployer, measurer, tracer, simplifier — viram lens cards de deliberação). + +3. **Entrada única = `/council`, absorve audit** (Felipe, 2026-07-09): council é O nome do engine. `/council ` → deliberation; `/council audit [focus]` → audit (7 lanes). O nome specialist-panel aposenta. Fecha o GAP do skill-absorbs ("keep /council name?" → SIM, e vira a entrada de tudo). Uma rota, um workflow script, dois modos por preset. + +4. **Escopo = incluir consumidores** (Felipe, 2026-07-09): o wish TAMBÉM rewira `/review` (panels multi-lens por change-type) e `/brainstorm` (o passo "dispatch 2-3 lens subagents" passa a ler lens cards da library). Entrega a visão G4 completa. Consequência: colisão com skills-fable5-revamp vira dependência de sequenciamento explícita. +5. **Naming público = renomear por lane, inspiração citada** (Felipe, 2026-07-09): skills shipam como `repo-hygiene`, `architecture`, `code-quality`, `qa`, `perf`, `supply-chain`, `dx-docs`; corpo cita "lens inspired by the work of ". Zero nome de pessoa real como identidade de produto. Cópias pessoais de Felipe podem manter os nomes antigos até aposentar. + +### Resolvidos por consequência (não perguntados) +- Fix-routing p/ usuários genie: personas shipam como skills standalone (decisão 2) → "run and fix" funciona pra todos; em repos genie a authority boundary continua findings→`/wish`. +- Repo-profile: mantém single-writer — o stage de synthesis/persist do workflow é o único escritor de `.genie/repo-profile.md`. +- Skills pessoais de Felipe (specialist-panel + 7 personas): aposentam após o ship (hygiene local dele, fora do escopo do repo). +- Lens library unificada: 7 persona SKILL.mds (modes: both — voz + metodologia) + 6 lentes council-only (questioner, simplifier, operator, deployer, measurer, tracer; cards leves). As 4 lentes redundantes do council (benchmarker≈perf, sentinel≈supply-chain, ergonomist≈dx-docs, architect≈architecture) morrem; routing table remapeada pra apontar pros persona skills. 13 lentes totais. +- Round 2 socrático em workflow: agents FRESCOS por round (round 2 recebe a própria posição round-1 do membro + as dos outros via prompt). Sem SendMessage — mais resumível/cacheável; perde continuidade de sessão (aceito). +- Audit mode: round1 lanes → synthesis (dedupe/conflito/re-rank global) → persist profile. Deliberation: round1 → round2 → synthesis (consenso/tensões/dissent). templates/report.md absorvido no prompt de synthesis. +- Council engine passa a exigir Claude Code (workflow runtime) — alvo Codex/Hermes fica OUT (o council antigo via Agent tool já era CC-only na prática). + +6. **Launch = saved workflow como comando** (Felipe, 2026-07-09): `/council` é o próprio saved workflow (meta name 'council'); o modelo media args estruturado na invocação; routing table vive como dado JS no script. SEM skill intermediária de orquestração. + +### Verificação de distribuição (2026-07-09, plugins-reference fetch direto) +- **Plugins NÃO shipam workflows.** Component path fields completos: skills, commands, agents, hooks, mcpServers, outputStyles, lspServers, experimental.themes, experimental.monitors, userConfig, channels, dependencies. Zero menção a workflows na referência inteira. +- **Mecanismo resolvido:** `genie install`/`update` (smart-install.js roda com `CLAUDE_PLUGIN_ROOT` — linha 19) estampa `LENS_ROOT` (path absoluto do plugin instalado) num template `council.js` e copia pra `~/.claude/workflows/council.js`. Disponível em todo projeto do usuário; re-stamp a cada update — necessário porque o path do plugin MUDA em update (docs: "This path changes when the plugin updates"). +- **Self-healing:** o script não tem fs/env; um stage-0 resolver agent (barato) verifica os lens paths estampados e localiza via Glob se stale. +- **Colisão de nome evitada:** `skills/council/` morre INTEIRO — precedência skill-vs-workflow pro mesmo nome `/council` é indocumentada, não arriscar. Os 6 lens cards de deliberação vão pra `plugins/genie/references/lenses/` (references/ já existe no plugin). +- Precedência documentada: project workflow > personal — um repo pode dar override no council com `.claude/workflows/council.js` próprio (feature, não bug). +- **Cross-check cc-guide (2026-07-09):** relatório independente confirmou tudo (sem workflows em plugin components; interpolação `${CLAUDE_PLUGIN_ROOT}`/`${CLAUDE_PLUGIN_DATA}`/`${CLAUDE_PROJECT_DIR}`/`${CLAUDE_SKILL_DIR}`; args model-mediated). Fato novo: `${CLAUDE_PLUGIN_DATA}` persiste através de updates — alternativa de mitigação pro stamp-staleness (copiar lenses pra data-dir estável e estampar esse path), registrada mas NÃO adotada: re-stamp no update + stage-0 resolver já cobre, e a cópia data-dir criaria seu próprio drift. + +## SCOPE (draft) +**IN:** `council.js` saved workflow (engine 2 modos, routing como dado JS), lens library unificada (13: 7 lane skills + 6 cards em references/lenses/), 7 persona skills renomeadas por lane, stamp+copy no install/update (smart-install), morte de `skills/council/` inteiro (routing/config/report absorvidos), rewiring `/review` panels + `/brainstorm` domain-experts, lints (lens frontmatter, refs probe-guarded, workflow script structural), docs. +**OUT:** rewiring de outros consumidores (work/fix/pm), Codex/Hermes support pro engine, autonomia dream/scheduler, mudanças no routing-matrix dos role agents, deleção das skills pessoais de Felipe (hygiene local), CLI `genie council` (term-command), `.claude/workflows/` scaffold por repo via `genie init` (override por repo fica como feature documentada, não entregue). + +## RISKS (draft) +| Risco | Sev | Mitigação | +|---|---|---| +| fable5-revamp execution review pendente nos MESMOS arquivos de skill | HIGH | Waves: começar por arquivos novos (personas, council.js, lenses/); edits em review/brainstorm + deleção do council só após aquele review fechar (dependência explícita no wish) | +| Stamp de LENS_ROOT stale (plugin path muda em update; user não re-roda install) | MED | Re-stamp no fluxo de update + stage-0 resolver agent com Glob-fallback (self-healing) | +| Custo: audit ≈ 9-10 agents no modelo da sessão | MED | Lane narrowing (`/council audit `), nota de custo no report/docs, tokens visíveis em /workflows | +| Sem input mid-run (workflow) — perde supervisão do orquestrador | LOW | Schema-validated returns por stage; regra ≥2 membros do council preservada no script | +| Primeira execução pede aprovação do workflow por projeto | LOW | Documentar; "don't ask again" per-project | +| Workflows exigem CC ≥ 2.1.154 + plano pago; org pode desabilitar (`disableWorkflows`) | LOW | Documentar requisito; sem fallback degradado (decisão: engine é CC-only) | + +## CRITERIA (draft — fail-hard) +- [ ] Lint estrutural do council.js: meta name 'council', zero Date.now/Math.random/new Date()/require/fs/import, parse ESM ok +- [ ] Lint da lens library: toda entrada da routing table aponta pra lens file existente; todo lens tem frontmatter obrigatório +- [ ] 7 lane skills existem com nomes de domínio; grep-gate: nenhum nome de pessoa real em `name:`; linha de inspiração presente +- [ ] git grep -i 'specialist-panel' → 0 hits fora de attic/CHANGELOG; members/config.md + routing.md do council antigo removidos/absorvidos +- [ ] /review e /brainstorm referenciam paths da library que existem (probe-guarded refs) +- [ ] bun run check verde +- [ ] QA vivo (manual): 1 run deliberation + 1 run audit no repo genie, evidência no wish + +## WAVES (seed p/ wish) +| Wave | Grupos | Nota | +|---|---|---| +| 1 | G1 persona skills (dirs novos) ∥ G2 council.js engine + lens cards + stamp no install (files novos) | zero colisão com fable5-review | +| 2 | G3 cutover: morte de skills/council/ + purge de referências | depende G1+G2 + fable5 execution review fechado | +| 3 | G4 consumidores (/review, /brainstorm) | depende G3 | +| 4 | G5 lints wired no check + docs + QA vivo | depende G3 (lints) / G4 (docs) | + +## REVIEW GATE (2026-07-09) + +- **Design review: SHIP** (reviewer independente; 0 CRITICAL / 0 HIGH remanescente — o HIGH de consumer-wording foi corrigido mid-review e re-verificado). 4 MEDIUMs aplicados no plano na hora: g4 gate estrutural, stamp site pinado no SessionStart hook (antes dos early-exits, idempotente), banned-API gate re-sourçado no spec do Workflow tool, gate G3 endurecido pra fable5 MERGED + rebase. LOWs: wording "rewired" corrigido; denylist de sobrenomes explicitada; `meta.phases` NÃO é defeito — o spec do Workflow tool documenta `phases` como campo opcional do meta (workflows.md só mostra o exemplo mínimo). +- Reviewer confirmou sound: plugins-sem-workflows (verbatim), runtime facts, council/panel/personas como descritos, layout do plugin, mapping 10→13, DAG válido, Wave 1 genuinamente new-files-only, role agents corretamente OUT. +- **Plan review (WISH): SHIP** (mesmo reviewer; 0 CRITICAL/HIGH). MEDIUM novo aplicado: gate do G2 tornado self-contained (parse ESM via bun build, cards-only) e integridade 13-lens completa movida pro lint a partir do G3 — preserva Wave 1 paralela. LOWs aplicados: g5 prova comportamentalmente que o lint roda DENTRO do check (grep no output); ban amplo de `new Date(` mantido de propósito (documentado no script). Status do wish: DRAFT — reviews SHIP, `/work` user-gated. + +## WRS: ██████████ 100/100 — Problem ✅ | Scope ✅ | Decisions ✅ | Risks ✅ | Criteria ✅ → crystallized em DESIGN.md (2026-07-09) diff --git a/.genie/brainstorms/skill-absorbs/DRAFT.md b/.genie/brainstorms/skill-absorbs/DRAFT.md index ab586eac3..f206c31b6 100644 --- a/.genie/brainstorms/skill-absorbs/DRAFT.md +++ b/.genie/brainstorms/skill-absorbs/DRAFT.md @@ -12,6 +12,7 @@ - wizard → ABSORB into genie router (first-run branch). Router scope-fenced: intent→route→precondition→delegate only; no absorbed-logic dumping ground. - pm → ABSORB: triage guidance → work reference; autopilot → dream successor (with human-gate policy). - council → lens LIBRARY (consumed by review panels + brainstorm domain-experts) + THIN /council route preserved (strategy pre-artifact, dissent preservation, appeal court for reviewer↔gate disagreements). + - **2026-07-09 UPDATE — implemented via [council-workflow](../council-workflow/DESIGN.md) (poured):** /council = native saved workflow (deliberation + audit modes); lens library = 7 lane skills (renamed personas) + 6 deliberation cards; "thin route" superseded — the route IS the workflow command. GAP "/council name?" closed (name kept). Consumer rewiring (/review panels + /brainstorm domain-experts) moved INTO that wish's scope. - report → REFACTOR: observability re-sourced to LangWatch + portable local markdown/JSON fallback; keep gh-issue intake. - refine → re-scoped to cross-LLM prompt adapter (OWN track: cross-agent-delegate). - learn → PARKED at .genie/attic/skills/learn/ (brain transfer pending; not this wish). diff --git a/.genie/wishes/council-workflow/WISH.md b/.genie/wishes/council-workflow/WISH.md new file mode 100644 index 000000000..b6a75a591 --- /dev/null +++ b/.genie/wishes/council-workflow/WISH.md @@ -0,0 +1,175 @@ +# Wish: /council — Native Workflow Engine (Deliberation + Audit) + +| Field | Value | +|-------|-------| +| **Status** | DRAFT — design review SHIP + plan review SHIP (2026-07-09, same independent reviewer, 1 fix pass each); `/work` user-gated | +| **Slug** | `council-workflow` | +| **Date** | 2026-07-09 | +| **Author** | Felipe (planned with Fable 5) | +| **Appetite** | medium (≈1 week) | +| **Branch** | `wish/council-workflow` | +| **Design** | [DESIGN.md](../../brainstorms/council-workflow/DESIGN.md) | + +## Summary + +Genie's multi-perspective reasoning is split across two model-driven orchestrators — `skills/council/` (deliberation) and the personal `specialist-panel` skill (7-lane audit) — duplicating the same fan-out→synthesize pattern in fragile, non-resumable prompt form. This wish replaces both with one native dynamic-workflow command: `/council ` deliberates, `/council audit [focus]` audits, backed by a unified 13-lens library (7 persona skills renamed by lane + 6 deliberation cards) shipped with the genie plugin and distributed by install-time stamp+copy to `~/.claude/workflows/council.js` (plugins verifiably cannot ship workflows). `/review` and `/brainstorm` gain lens-library consumption (panel dispatch and domain-experts steps, both new to the repo skills). + +## Scope + +### IN +- `plugins/genie/workflows/council.js` — workflow template: `meta.name 'council'`, phases Resolve → Round 1 → Round 2 (deliberation) → Synthesis → Persist (audit), ROUTING keyword table + LENSES map as JS data, schema-validated stage returns, ≥2-members rule, fresh-agent Socratic round 2, single-writer `.genie/repo-profile.md` persist. +- 7 standalone lane skills under `skills/`: `repo-hygiene`, `architecture`, `code-quality`, `qa`, `perf`, `supply-chain`, `dx-docs` — persona methodologies migrated, "inspired by the work of " cited, no real person's name as identity. +- 6 deliberation lens cards at `plugins/genie/references/lenses/` (questioner, simplifier, operator, deployer, measurer, tracer) with frontmatter (name, modes, voice). +- Install/update stamp+copy: `LENS_ROOT` (absolute installed-plugin path) stamped into the template, result written to `~/.claude/workflows/council.js`; re-stamped on update. +- Cutover: delete `skills/council/` entirely (SKILL.md, members/, templates/ absorbed into script + lens frontmatter); purge stale references. +- Consumers: `/review` gains lens-panel dispatch (multi-lens reviewers by change-type); `/brainstorm` lens-subagent step reads library cards. +- Lints wired into `bun run check`: council.js structural lint (ESM parse, banned APIs, meta, ROUTING↔LENSES integrity), lens frontmatter lint, lens-reference existence check for review/brainstorm. +- Docs: skills/README.md decision table, plugin docs, workflow requirements note (CC ≥ 2.1.154, paid plan, org `disableWorkflows`). + +### OUT +- Rewiring `/work`, `/fix`, `/pm` to consume lenses (skill-absorbs G4 follow-up). +- Codex/Hermes support for the engine (workflow runtime is Claude Code-only — accepted). +- `genie council` term-command; per-repo `.claude/workflows/` scaffolding via `genie init` (repo override stays a documented capability). +- dream/scheduler autonomy; routing-matrix role-agent changes. +- Deleting Felipe's personal `~/.claude/skills/` copies (local hygiene, outside the repo). + +## Decisions + +| # | Decision | Rationale | +|---|----------|-----------| +| 1 | One engine, two modes | Both surfaces are the same fan-out→synthesize pattern; one script + mode presets maximizes reuse and honors the skill-absorbs G4 ruling | +| 2 | Personas absorbed as standalone plugin skills | Workflow reads the same SKILL.md as its lens — single source of truth; fix-mode ships to every genie user | +| 3 | `/council` is the single entry (`audit` subcommand) | Council is genie lore; closes the G4 naming GAP | +| 4 | Lane names, not people names | Real experts' names as speaking product personas without consent is a liability; inspiration cited in body | +| 5 | Saved workflow command, no launcher skill | Most native shape; args model-mediated; routing lives as data in the script | +| 6 | Install-time stamp+copy to `~/.claude/workflows/` | Plugins cannot ship workflows (verified against plugins reference 2026-07-09); smart-install runs with `CLAUDE_PLUGIN_ROOT` so `LENS_ROOT` stamping is deterministic and update-safe | +| 7 | Fresh-agent Socratic round 2 | Workflows have no SendMessage; feeding each member its own R1 back preserves identity, gains resumability | +| 8 | `skills/council/` deleted whole | Skill-vs-workflow precedence for one name is undocumented — avoid the collision | +| 9 | Consumers wired now (not deferred) | Felipe chose delivering the full G4 vision in this wish; both steps are new to the repo skills; sequencing handled by waves | + +## Success Criteria + +- [ ] `bun run lint:council-workflow` passes (exercised from G3 onward, once G1's lanes exist, and permanently via `bun run check`): template parses as ESM, `meta.name === 'council'`, zero banned APIs — `Date.now`, `Math.random`, `new Date(` (the Workflow tool spec: the determinism trio throws in scripts because it would break resume; `new Date(` banned broader-than-spec deliberately) plus `require(`, `import `, `process.`, `fs.` (workflows.md + tool spec: scripts are self-contained, no filesystem or Node.js API access) — every ROUTING member resolves to an existing lens file on disk (all 13), every lens file has required frontmatter +- [ ] All 7 lane skills exist with domain names; no real person's name in any `name:` field under `skills/` (grep denylist: chacon, ousterhout, hejlsberg, beck, gregg, lorenc, procida); inspiration line present in each +- [ ] Stamp unit test green: installer function replaces `LENS_ROOT` placeholder with an absolute path and output lands at the expected `~/.claude/workflows/council.js` target (tmpdir-isolated via `GENIE_HOME`-style env override) +- [ ] `skills/council/` no longer exists; `git grep -il 'specialist-panel'` returns 0 hits outside `.genie/attic/`, `CHANGELOG.md`, and this wish's artifacts +- [ ] `skills/review/SKILL.md` and `skills/brainstorm/SKILL.md` reference lens-library paths that exist on disk (lint fails on dangling path) +- [ ] `bun run check` green with new lints wired in +- [ ] Live QA evidence committed: one real `/council ` deliberation run + one `/council audit` run on the genie repo, reports saved under `.genie/wishes/council-workflow/qa/` + +## Execution Strategy + +| Wave | Groups | Notes | +|------|--------|-------| +| 1 | G1 (lane skills) ∥ G2 (engine + lenses + stamp) | New files only — zero collision with the pending skills-fable5-revamp execution review | +| 2 | G3 (cutover) | Gated on G1+G2 AND skills-fable5-revamp MERGED to its base (review-closed insufficient); rebase `wish/council-workflow` onto the post-merge base first | +| 3 | G4 (consumers) | Edits review/brainstorm skills — behind the same fable5 merged-gate | +| 4 | G5 (lints in check + docs + live QA) | Final gate | + +--- + +## Execution Groups + +### Group 1: Lane skills — personas migrated and renamed +**Goal:** Ship the 7 specialist personas as standalone genie plugin skills under domain names. + +**Deliverables:** +1. `skills/{repo-hygiene,architecture,code-quality,qa,perf,supply-chain,dx-docs}/SKILL.md` — methodology preserved from the personal persona skills, frontmatter `name:` = lane name, body cites "inspired by the work of ", assess/fix contract kept per skill. + +**Acceptance Criteria:** +- [ ] All 7 SKILL.md files exist with lane-named frontmatter +- [ ] No real person's name in any `name:` field under `skills/` +- [ ] Each body contains an inspiration attribution line + +**Validation:** +```bash +cd "$(git rev-parse --show-toplevel)" +bash .genie/wishes/council-workflow/validate/g1-lane-skills.sh +``` + +**depends-on:** none + +### Group 2: Engine — council.js template, lens cards, install stamp +**Goal:** The workflow engine exists as a plugin-shipped template with its lens data and reaches `~/.claude/workflows/` through the install/update flow. + +**Deliverables:** +1. `plugins/genie/workflows/council.js` — meta + Resolve/Round 1/Round 2/Synthesis/Persist phases, ROUTING table (absorbed from `members/routing.md`, incl. default trio + members override), LENSES map (`LENS_ROOT`-relative), deliberation and audit stage contracts, ≥2-members failure rule. +2. `plugins/genie/references/lenses/{questioner,simplifier,operator,deployer,measurer,tracer}.md` — frontmatter: name, modes, voice. +3. Stamp function as `src/lib/council-workflow-stamp.ts` (+ colocated `council-workflow-stamp.test.ts`, tmpdir-isolated): replaces the `LENS_ROOT` placeholder with the absolute installed-plugin path and writes `~/.claude/workflows/council.js`. **Call site pinned: the SessionStart hook (`plugins/genie/scripts/smart-install.js`, where `CLAUDE_PLUGIN_ROOT` is set), placed BEFORE its early-exit guards (deps-present, `GENIE_WORKER=1`) and idempotent via drift-check (rewrite only when stamped `LENS_ROOT` ≠ current `CLAUDE_PLUGIN_ROOT` or template hash changed). Re-stamp is therefore driven by the first session start after `claude plugin update` — the `genie update` CLI does not own it.** +4. `scripts/council-workflow-lint.ts` + `package.json` script `lint:council-workflow`. + +**Acceptance Criteria:** +- [ ] Template parses as ESM (self-contained `bun build --no-bundle` check); `meta.name === 'council'`; zero banned APIs +- [ ] Every ROUTING member maps to a known lens NAME; the 6 deliberation cards exist with required frontmatter (full on-disk 13-lens resolution — including G1's lane skills — is asserted by `lint:council-workflow` from G3 onward, keeping G2 validatable in parallel with G1) +- [ ] Stamp unit test green in tmpdir isolation + +**Validation:** +```bash +cd "$(git rev-parse --show-toplevel)" +bash .genie/wishes/council-workflow/validate/g2-engine.sh +``` + +**depends-on:** none + +### Group 3: Cutover — old council dies +**Goal:** Remove the model-driven council orchestrator and every stale reference to it or to specialist-panel. + +**Deliverables:** +1. `skills/council/` deleted (SKILL.md, members/routing.md, members/config.md, templates/report.md). +2. References purged across skills/, plugins/, docs (report template + routing absorbed into the script in G2). +3. skills/README.md council row updated to point at the workflow. + +**Acceptance Criteria:** +- [ ] `skills/council/` gone; no references to `members/routing.md`, `members/config.md`, or the council skill remain in live surfaces +- [ ] `bun run lint:council-workflow` green — first point where full 13-lens on-disk integrity is assertable (G1+G2 both done) +- [ ] `git grep -il 'specialist-panel'` → 0 hits outside `.genie/attic/`, `CHANGELOG.md`, this wish's artifacts +- [ ] **Process gate:** skills-fable5-revamp MERGED to its base branch (review-closed is not enough — its Wave 1 rewrote these same skill files in place), and `wish/council-workflow` rebased onto the post-merge base before this group starts (checked by the worker, recorded in the group log) + +**Validation:** +```bash +cd "$(git rev-parse --show-toplevel)" +bash .genie/wishes/council-workflow/validate/g3-cutover.sh +``` + +**depends-on:** Group 1, Group 2 + +### Group 4: Consumers — /review panels + /brainstorm domain-experts +**Goal:** The lens library is consumed beyond the council: review can convene multi-lens panels, brainstorm's lens-subagent step reads library cards. + +**Deliverables:** +1. `skills/review/SKILL.md` — lens-panel dispatch section (change-type → lens cards; e.g. auth-touching diff adds the supply-chain lens reviewer). +2. `skills/brainstorm/SKILL.md` — gains a Decisions-stuck domain-experts step that dispatches 2-3 lens subagents reading `references/lenses/` cards (step is NEW to the repo skill; pattern mirrors the global brainstorm skill's lens-subagent step). + +**Acceptance Criteria:** +- [ ] `skills/review/SKILL.md` has a lens-panel section with change-type → lens mapping (structural grep markers: `lens panel`, `change-type`) +- [ ] `skills/brainstorm/SKILL.md` has a domain-experts step (structural grep marker: `domain-expert`) +- [ ] Both skills reference lens-library paths (cards and lane skills) that exist on disk +- [ ] No duplicated lens definitions inline in either skill — no `voice:`/`modes:` frontmatter blocks outside the library (single source: the library) + +**Validation:** +```bash +cd "$(git rev-parse --show-toplevel)" +bash .genie/wishes/council-workflow/validate/g4-consumers.sh +``` + +**depends-on:** Group 3 + +### Group 5: Gate — lints wired, docs, live QA +**Goal:** The new lints run in the standard gate, docs reflect the new surface, and both modes are proven live. + +**Deliverables:** +1. `lint:council-workflow` wired into `bun run check` (package.json). +2. Docs: skills/README.md, plugin README/docs notes, workflow requirements (CC ≥ 2.1.154, paid plans, `disableWorkflows`). +3. Live QA evidence: `.genie/wishes/council-workflow/qa/deliberation-run.md` + `qa/audit-run.md` (real runs on the genie repo, with `/workflows` token totals captured). + +**Acceptance Criteria:** +- [ ] `bun run check` green AND its output proves `lint:council-workflow` actually ran (behavioral wiring check, not just script existence) +- [ ] Both QA evidence files present with real run output + +**Validation:** +```bash +cd "$(git rev-parse --show-toplevel)" +bash .genie/wishes/council-workflow/validate/g5-gate.sh +``` + +**depends-on:** Group 3, Group 4 diff --git a/.genie/wishes/council-workflow/validate/g1-lane-skills.sh b/.genie/wishes/council-workflow/validate/g1-lane-skills.sh new file mode 100755 index 000000000..b902e3d34 --- /dev/null +++ b/.genie/wishes/council-workflow/validate/g1-lane-skills.sh @@ -0,0 +1,19 @@ +#!/usr/bin/env bash +set -euo pipefail +cd "$(git rev-parse --show-toplevel)" + +for lane in repo-hygiene architecture code-quality qa perf supply-chain dx-docs; do + f="skills/$lane/SKILL.md" + test -f "$f" || { echo "FAIL: missing $f"; exit 1; } + grep -Eq "^name: ${lane}$" "$f" || { echo "FAIL: frontmatter name != ${lane} in $f"; exit 1; } + grep -qi 'inspired by' "$f" || { echo "FAIL: no inspiration attribution in $f"; exit 1; } +done + +for expert in chacon ousterhout hejlsberg beck gregg lorenc procida; do + if grep -rEiq "^name:.*${expert}" skills/; then + echo "FAIL: real person's name '${expert}' used as a skill identity (name: field) under skills/" + exit 1 + fi +done + +echo "G1 PASS" diff --git a/.genie/wishes/council-workflow/validate/g2-engine.sh b/.genie/wishes/council-workflow/validate/g2-engine.sh new file mode 100755 index 000000000..2604548cf --- /dev/null +++ b/.genie/wishes/council-workflow/validate/g2-engine.sh @@ -0,0 +1,33 @@ +#!/usr/bin/env bash +set -euo pipefail +cd "$(git rev-parse --show-toplevel)" + +t="plugins/genie/workflows/council.js" +test -f "$t" || { echo "FAIL: missing $t"; exit 1; } +grep -q "name: 'council'" "$t" || { echo "FAIL: meta.name !== 'council' in $t"; exit 1; } + +# Deliberately broader than the spec: ALL `new Date(` is banned (not just argless) — +# timestamps travel via args; a seeded date in the script is a smell, and failing closed is cheap. +for banned in 'Date\.now' 'Math\.random' 'new Date\(' 'require\(' '^import ' 'process\.' '[^a-zA-Z.]fs\.'; do + if grep -Eq "$banned" "$t"; then echo "FAIL: banned API matching '$banned' in $t"; exit 1; fi +done + +# ESM parse check, self-contained (no cross-group deps): transpile fails on syntax errors. +tmp="$(mktemp -d)" +trap 'rm -rf "$tmp"' EXIT +bun build "$t" --no-bundle --outdir "$tmp" > /dev/null || { echo "FAIL: $t does not parse as ESM"; exit 1; } + +for card in questioner simplifier operator deployer measurer tracer; do + c="plugins/genie/references/lenses/$card.md" + test -f "$c" || { echo "FAIL: missing lens card $c"; exit 1; } + grep -q '^name: ' "$c" || { echo "FAIL: no name frontmatter in $c"; exit 1; } + grep -q '^modes: ' "$c" || { echo "FAIL: no modes frontmatter in $c"; exit 1; } + grep -q '^voice: ' "$c" || { echo "FAIL: no voice frontmatter in $c"; exit 1; } +done + +bun test src/lib/council-workflow-stamp.test.ts + +# NOTE: full 13-lens integrity (ROUTING -> files on disk, incl. G1's lane skills) lives in +# `bun run lint:council-workflow`, exercised from G3 onward and permanently via `bun run check` — +# NOT here, so G2 stays validatable in parallel with G1 (Wave 1). +echo "G2 PASS" diff --git a/.genie/wishes/council-workflow/validate/g3-cutover.sh b/.genie/wishes/council-workflow/validate/g3-cutover.sh new file mode 100755 index 000000000..bb53e509f --- /dev/null +++ b/.genie/wishes/council-workflow/validate/g3-cutover.sh @@ -0,0 +1,24 @@ +#!/usr/bin/env bash +set -euo pipefail +cd "$(git rev-parse --show-toplevel)" + +test ! -e skills/council || { echo "FAIL: skills/council still exists"; exit 1; } + +hits=$(git grep -il 'specialist-panel' -- ':!.genie/attic' ':!CHANGELOG.md' ':!.genie/wishes/council-workflow' ':!.genie/brainstorms' || true) +if [ -n "$hits" ]; then + echo "FAIL: stale specialist-panel references:" + echo "$hits" + exit 1 +fi + +refs=$(git grep -l 'members/routing\.md\|members/config\.md\|council/templates/report\.md' -- skills plugins 2>/dev/null || true) +if [ -n "$refs" ]; then + echo "FAIL: stale council-internals references:" + echo "$refs" + exit 1 +fi + +# G1+G2 are both done by now — first point where full 13-lens integrity is assertable. +bun run lint:council-workflow + +echo "G3 PASS" diff --git a/.genie/wishes/council-workflow/validate/g4-consumers.sh b/.genie/wishes/council-workflow/validate/g4-consumers.sh new file mode 100755 index 000000000..167cfa501 --- /dev/null +++ b/.genie/wishes/council-workflow/validate/g4-consumers.sh @@ -0,0 +1,32 @@ +#!/usr/bin/env bash +set -euo pipefail +cd "$(git rev-parse --show-toplevel)" + +# --- /review: lens-panel dispatch must exist structurally, not just as a path mention +grep -qiE 'lens[ -]panel' skills/review/SKILL.md || { echo "FAIL: review skill lacks a lens-panel section"; exit 1; } +grep -qiE 'change[ -]type' skills/review/SKILL.md || { echo "FAIL: review skill lacks change-type -> lens mapping"; exit 1; } +grep -q 'references/lenses' skills/review/SKILL.md || { echo "FAIL: review skill lacks lens-library reference"; exit 1; } + +# --- /brainstorm: domain-experts lens-subagent step must exist structurally +grep -qiE 'domain[ -]experts?' skills/brainstorm/SKILL.md || { echo "FAIL: brainstorm skill lacks a domain-experts step"; exit 1; } +grep -q 'references/lenses' skills/brainstorm/SKILL.md || { echo "FAIL: brainstorm skill lacks lens-library reference"; exit 1; } + +for f in skills/review/SKILL.md skills/brainstorm/SKILL.md; do + # every cited lens-card path resolves + while IFS= read -r p; do + test -e "plugins/genie/$p" || { echo "FAIL: dangling lens ref '$p' in $f"; exit 1; } + done < <(grep -oE 'references/lenses/[a-z-]+\.md' "$f" | sort -u) + + # every cited lane-skill lens path resolves + while IFS= read -r p; do + test -e "$p" || { echo "FAIL: dangling lane-skill ref '$p' in $f"; exit 1; } + done < <(grep -oE 'skills/(repo-hygiene|architecture|code-quality|qa|perf|supply-chain|dx-docs)/SKILL\.md' "$f" | sort -u) + + # no inline lens definitions (lens cards carry voice:/modes: frontmatter — skills must reference, not embed) + if grep -qE '^(voice|modes): ' "$f"; then + echo "FAIL: inline lens definition (voice:/modes: block) embedded in $f — reference the library instead" + exit 1 + fi +done + +echo "G4 PASS" diff --git a/.genie/wishes/council-workflow/validate/g5-gate.sh b/.genie/wishes/council-workflow/validate/g5-gate.sh new file mode 100755 index 000000000..0b0f011c5 --- /dev/null +++ b/.genie/wishes/council-workflow/validate/g5-gate.sh @@ -0,0 +1,14 @@ +#!/usr/bin/env bash +set -euo pipefail +cd "$(git rev-parse --show-toplevel)" + +grep -q '"lint:council-workflow"' package.json || { echo "FAIL: lint:council-workflow script missing from package.json"; exit 1; } + +# Behavioral wiring proof: check must actually RUN the council lint, not merely coexist with it. +out="$(bun run check 2>&1)" || { echo "$out"; echo "FAIL: bun run check failed"; exit 1; } +echo "$out" | grep -q 'council-workflow' || { echo "FAIL: lint:council-workflow did not run as part of bun run check"; exit 1; } + +test -s .genie/wishes/council-workflow/qa/deliberation-run.md || { echo "FAIL: missing live QA evidence (deliberation)"; exit 1; } +test -s .genie/wishes/council-workflow/qa/audit-run.md || { echo "FAIL: missing live QA evidence (audit)"; exit 1; } + +echo "G5 PASS" From 0b222e17acd7a168c33bf9c993ebfba6b8476438 Mon Sep 17 00:00:00 2001 From: namastex888 Date: Thu, 9 Jul 2026 20:37:39 -0300 Subject: [PATCH 2/8] feat(skills): ship 7 specialist lane skills (council-workflow G1) --- skills/architecture/SKILL.md | 48 +++++++++++++++++++++++++++++++++++ skills/code-quality/SKILL.md | 48 +++++++++++++++++++++++++++++++++++ skills/dx-docs/SKILL.md | 49 ++++++++++++++++++++++++++++++++++++ skills/perf/SKILL.md | 46 +++++++++++++++++++++++++++++++++ skills/qa/SKILL.md | 48 +++++++++++++++++++++++++++++++++++ skills/repo-hygiene/SKILL.md | 49 ++++++++++++++++++++++++++++++++++++ skills/supply-chain/SKILL.md | 47 ++++++++++++++++++++++++++++++++++ 7 files changed, 335 insertions(+) create mode 100644 skills/architecture/SKILL.md create mode 100644 skills/code-quality/SKILL.md create mode 100644 skills/dx-docs/SKILL.md create mode 100644 skills/perf/SKILL.md create mode 100644 skills/qa/SKILL.md create mode 100644 skills/repo-hygiene/SKILL.md create mode 100644 skills/supply-chain/SKILL.md diff --git a/skills/architecture/SKILL.md b/skills/architecture/SKILL.md new file mode 100644 index 000000000..8e8b29d82 --- /dev/null +++ b/skills/architecture/SKILL.md @@ -0,0 +1,48 @@ +--- +name: architecture +description: Use when reviewing architecture in any codebase — module boundaries, stated design contracts, abstraction depth, error-handling design. Assess by default, apply changes on request; complexity is dependencies plus obscurity, and deep modules win. +--- + +# Architecture Review + +## Lens + +This lane treats complexity as anything that makes a system hard to understand or modify — it accumulates as dependencies and obscurity. The unit of judgment is the module: deep modules (simple interface, substantial implementation) are good; shallow modules (interface as complicated as what they hide) are architecture debt. Information leakage, pass-through methods, and temporal decomposition are the smells to hunt. Prize "define errors out of existence" and design-it-twice thinking. + +This lane's lens is inspired by the work of John Ousterhout, author of *A Philosophy of Software Design*. + +## Mandate + +Assess and report by default. Apply changes only when the invocation explicitly asks. Every finding must cite the concrete interface, import, or branch that embodies it — no vibes. Findings outside this lane (failing gates, security holes, missing tests) get a one-line handoff to the relevant lane skill under `skills/`. When you have enough information to judge, judge; recommend one design, not a survey. + +## Discover the Ground Truth First + +Architecture is judged against the repo's own stated intent, then against first principles. Before scoring anything, collect: `CLAUDE.md` / `AGENTS.md` architecture sections, ADRs or design docs, any documented invariants ("X must never import Y", "state lives in Z"), and the real module graph traced from the entry points via imports. A repo's deliberate constraints (zero-daemon designs, intentionally-duplicated modules, forbidden cross-imports) are the design under review — the defect is a *violated* contract or a contract the code has outgrown, not the contract's existence. + +**Genie-framework repos**: `.genie/` documents (wishes, brainstorms) often record the intended design and its acceptance criteria — read the relevant wish before judging the code it produced. + +**Repo profile — recall, verify, persist.** Before deriving from scratch, recall a stored profile for this repo: a memory/brain store if one is available this session, else a well-known file (in genie-framework repos, `.genie/repo-profile.md`). For this lane the profile records the module map, documented invariants, key interfaces and their depth verdicts. Recalled anchors are hypotheses — re-verify each invariant you rely on against current code and report drift as a finding. After the audit, persist what discovery learned: update rather than duplicate, delete what proved wrong. + +## Workflow + +1. **Map the module graph.** Trace imports from the entry points; identify layers, cycles, and upward imports. Done when you have the real dependency picture, not the README's. +2. **Verify the stated contracts.** For each documented invariant found in discovery, read the code that must uphold it. Done when each is confirmed intact or broken with file:line evidence. +3. **Depth-score the key interfaces.** For the repo's central abstractions: interface surface vs implementation hidden, leakage of internals to callers, pass-throughs. Done when each has a deep/shallow verdict with the specific signature that decides it. +4. **Hunt the classic smells**: information leakage (two modules that must change together), temporal decomposition (modules named after steps, not capabilities), exceptions where errors could be defined away, configuration knobs exporting decisions the module should make. Done when each smell has a concrete instance or the category is declared clean. +5. **Rank by change amplification** — how many places must be touched when the underlying decision changes — and report. + +## Grounded Reporting + +Every structural claim traces to code read this session, cited file:line. A design opinion is only a defect if you can name the modification scenario it makes expensive; interfaces judged without reading their implementation are labeled as such. + +## Output Format + +Lead with a one-sentence verdict on architectural health. Then findings ranked by change-amplification risk, each with evidence, the modification scenario it hurts, and one recommended structural move. Explicitly list contracts verified intact — a review that only lists problems hides where the design is strong. Cross-lane handoffs last. In a genie-framework repo, tag severities in `/review` vocabulary (SHIP / FIX-FIRST / BLOCKED) and note which findings warrant a refactor wish via `/wish` rather than opportunistic edits. + +## Pitfalls + +- Deliberate separation is not duplication to consolidate — when the docs say two modules must not share code, the finding would be a cross-import, not the existence of two modules. +- A documented, empirically-forced exception to a clean rule is design; judge how it is encapsulated, not that it exists. +- Do not recommend extracting helpers from a readable linear workflow just to lower a complexity score — indirection with one caller is the opposite of a deep module. Respect the repo's own complexity-budget policy if it has one. +- Architectural constraints like "no resident daemon" or "state never in files" are usually load-bearing product decisions; proposing their reversal is a scope change to surface, not a finding to assert. +- Read the design doc or wish behind a subsystem before judging it — code that looks odd often implements a stated requirement. diff --git a/skills/code-quality/SKILL.md b/skills/code-quality/SKILL.md new file mode 100644 index 000000000..9fa1c05a9 --- /dev/null +++ b/skills/code-quality/SKILL.md @@ -0,0 +1,48 @@ +--- +name: code-quality +description: Use when auditing code quality in any codebase — discover and run the repo's real gates (typecheck, lint, dead-code, complexity), judge type discipline and duplication. Assess by default, apply changes on request; the compiler is the first reviewer. +--- + +# Code Quality Review + +## Lens + +This lane treats the type system as the cheapest, fastest reviewer on the team: a codebase's quality is measured by how much of its correctness the compiler can prove. Escape hatches — `any`, unchecked casts, suppression comments, `unsafe`, `# type: ignore` — are places where the team chose not to know. Gates exist to be run, not admired: a quality review that doesn't execute the toolchain is an opinion. + +This lane's lens is inspired by the work of Anders Hejlsberg — architect of Turbo Pascal, Delphi, C#, and TypeScript. + +## Mandate + +Assess and report by default. Apply changes only when the invocation explicitly asks. Never assess from reading alone when a gate exists — run it and report its actual output. Findings outside this lane (architecture judgment, test gaps, performance) get a one-line handoff to the relevant lane skill under `skills/`. When you have enough information to act, act. + +## Discover the Ground Truth First + +Every repo defines its own gates; find them before running anything. Read the package manifest scripts, `Makefile`/`justfile`, CI workflows, and `CLAUDE.md`/`AGENTS.md` for: the full check command, the individual typecheck / lint / dead-code / complexity commands, the formatter contract, and — critically — **documented known false positives and complexity-budget policies**. A repo that says "tool X flags Y, it's pre-existing" has told you what not to report. Note which language(s) and type systems are in play and their idiomatic escape hatches. + +**Genie-framework repos**: check `.genie/` for quality-related wishes (e.g. a complexity-budget or refactor wish with a hotspot ledger) — new violations are drift against that ledger, not fresh discoveries. + +**Repo profile — recall, verify, persist.** Before deriving from scratch, recall a stored profile for this repo: a memory/brain store if one is available this session, else a well-known file (in genie-framework repos, `.genie/repo-profile.md`). For this lane the profile records the gate commands, known false positives, complexity-budget policy, and ledger locations. Recalled gate commands are hypotheses — they must still exist and run; report drift as a finding. After the audit, persist what discovery learned: update rather than duplicate, delete what proved wrong. + +## Workflow + +1. **Run the gates individually** (typecheck, lint, dead-code, complexity — whatever discovery found), so one failure doesn't mask the rest. Done when each has an exit code and captured output. +2. **Audit type discipline at the boundaries.** Grep for the language's escape hatches; check compiler strictness config. For each hit: boundary where validation belongs (fine if runtime-validated) or interior hole. Done when every escape hatch has a verdict. +3. **Reconcile against the repo's own ledgers.** Compare current warnings to any documented hotspot list, baseline file, or suppression policy — undocumented new violations are drift; suppressions without a substantive reason are violations. Done when ledger and reality are reconciled. +4. **Hunt duplication** with at least two cited sites and one proposed home per instance — but respect documented deliberate non-sharing between modules. Done when each candidate is a finding or dismissed. +5. **Rank and report**: gate failures first, then type holes by blast radius, then ledger drift, then duplication. + +## Grounded Reporting + +Every gate claim quotes the command, exit code, and relevant output from this session; a gate not run (e.g. tests, owned by the QA lane) is named as not run. Never report "gates pass" from memory or from documentation. + +## Output Format + +Lead with a one-sentence verdict: which gates pass, which fail. Then findings ranked by severity, each with evidence, the correctness risk in plain language, and the exact edit you'd make on ask. Distinguish "gate is red" (fact) from "discipline is eroding" (trend with examples). In a genie-framework repo, tag severities in `/review` vocabulary (SHIP / FIX-FIRST / BLOCKED); systemic findings (a hotspot ledger growing, strictness never enabled) belong in a wish via `/wish`, not a drive-by fix list. + +## Pitfalls + +- Reporting a repo's documented false positives as findings is itself a finding against you — discovery exists to prevent exactly this. +- Complexity ceilings are usually warn-level budgets for linear workflows, not targets; do not demand extraction of a readable linear flow into single-caller helpers. +- Deliberately-unshared parallel modules (documented in the repo) are contract, not duplication — do not propose the shared-utils layer their docs forbid. +- Lint rules often carry test-directory relaxations; check the override before flagging test files. +- An escape hatch at a validated system boundary (user input, external API) is correct usage; only interior holes where the compiler was silenced without runtime backing are findings. diff --git a/skills/dx-docs/SKILL.md b/skills/dx-docs/SKILL.md new file mode 100644 index 000000000..21d6c9bde --- /dev/null +++ b/skills/dx-docs/SKILL.md @@ -0,0 +1,49 @@ +--- +name: dx-docs +description: Use when auditing DX, docs, and delivery in any codebase — the 30-minute-contributor test, docs-vs-reality drift, onboarding friction, error-message quality. Assess by default, fix docs on request; docs are judged by use, and every failure is a misfiled or missing Diátaxis quadrant. +--- + +# DX, Docs & Delivery Review + +## Lens + +This lane treats documentation as four different things — tutorials (learning-oriented), how-to guides (task-oriented), reference (information-oriented), explanation (understanding-oriented) — and nearly every documentation failure is one quadrant's content misfiled in another, or a quadrant missing entirely. Documentation is judged by use, not by existence: a doc that cannot be followed is worse than no doc, because it costs trust. Developer experience is documentation's runtime — error messages, help text, and onboarding friction are docs delivered at the moment of need. + +This lane's lens is inspired by the work of Daniele Procida, creator of the Diátaxis framework. + +## Mandate + +Assess and report by default. Apply doc fixes only when the invocation explicitly asks — and only through the repo's documented docs workflow if it has one (submodules, docs repos, review gates). Findings outside this lane get a one-line handoff to the relevant lane skill under `skills/`. Judgments come from *using* the docs and the product, never from reading them approvingly. + +## Discover the Ground Truth First + +Map the docs estate before judging it: where docs live (in-repo, submodule, separate site), which are public vs internal, what the contribution/onboarding path claims to be (README, CONTRIBUTING, `CLAUDE.md`/`AGENTS.md`), and what the product's real interface is — for a CLI, the live `--help` output of every command; for an API, the actual routes/signatures; for a library, the exported surface. The live interface is the truth; every doc, README table, and agent-context file is a claim to diff against it. Note the repo's stated DX bar (e.g. a 30-minute-contributor promise) — hold it to its own standard. + +**Genie-framework repos**: the lifecycle skills (brainstorm → wish → work → review, plus their kin) are part of the user-facing surface. Their SKILL.md descriptions, inputs, and outputs must chain coherently — does `/wish` consume what `/brainstorm` produces, does `/review` validate what `/work` emits — and match what the docs claim about them. + +**Repo profile — recall, verify, persist.** Before deriving from scratch, recall a stored profile for this repo: a memory/brain store if one is available this session, else a well-known file (in genie-framework repos, `.genie/repo-profile.md`). For this lane the profile records the docs topology, the live-interface inventory, past stumble logs, and open drift findings. Recalled drift may have been fixed since — re-check each entry against the live interface before reporting, and report new drift as a finding. After the audit, persist what discovery learned: update rather than duplicate, delete what proved wrong. + +## Workflow + +1. **Run the contributor test.** Follow the written onboarding path verbatim — clone/install through the first passing check — with no insider shortcuts, logging every divergence between docs and reality with a timestamp. Done when you have a stumble log and a pass/fail against the repo's stated (or a 30-minute default) bar. +2. **Diff docs against the live interface.** Enumerate the real commands/routes/exports; diff names, flags, and described behavior against every doc that mentions them. Done when each drift instance quotes both sides. +3. **Diátaxis-classify the docs tree.** Assign each page a quadrant; flag misfiled content (reference dumps inside how-tos, explanation blocking a tutorial path) and name missing quadrants (is there any true tutorial?). Check audience leakage across public/internal boundaries. Done when the tree has a quadrant map with gaps named. +4. **Trace the workflow chain** (in genie-framework repos: the lifecycle skills; elsewhere: the documented contributor workflow). Verify each handoff's stated inputs/outputs against actual behavior. Done when each link is confirmed coherent or flagged with mismatched quotes. +5. **Sample error messages.** Run 4–6 realistic failure invocations; record exit code, stderr, and text; grade each on what failed / why / what to do next. Done when each sample has a grade and quote. +6. **Rank**: onboarding blockers first (they cost every new contributor), then drift (it costs trust), then misfiling, then message polish. + +## Grounded Reporting + +Every drift claim quotes both sides; every stumble names the exact step and what actually happened; skipped steps (e.g. no fresh clone was feasible) are stated, with affected conclusions marked partial. + +## Output Format + +Lead with a one-sentence verdict: did the repo pass its contributor test, and what is the worst drift. Then findings ranked as above with evidence and concrete fixes — routed through the repo's docs workflow where one exists. Include what works well; a review that only lists friction misleads. In a genie-framework repo, tag severities in `/review` vocabulary (SHIP / FIX-FIRST / BLOCKED) and offer to crystallize a docs-overhaul into a wish via `/wish`. + +## Pitfalls + +- Never evaluate docs by reading them — a page can read beautifully and be unfollowable; every "docs are good" claim must trace to a followed procedure. +- Internal-only docs deliberately excluded from a public site are design, not gaps — check the exclusion mechanism before reporting "missing" pages. +- If docs live in a submodule or separate repo, a fix recommendation that says "edit and commit here" strands changes — name the real workflow in the fix. +- Agent-context files (CLAUDE.md and kin) drifting from the product is real drift, but its fix lands in this repo, not the docs pipeline — route the two drift classes separately. +- Terse is not bad: an error message answering what/why/next in one line beats a paragraph. Grade on the three questions, not on length. diff --git a/skills/perf/SKILL.md b/skills/perf/SKILL.md new file mode 100644 index 000000000..4c8c0c0a4 --- /dev/null +++ b/skills/perf/SKILL.md @@ -0,0 +1,46 @@ +--- +name: perf +description: Use when auditing performance in any codebase — cold starts, hot paths, dependency weight, storage query patterns. Assess by default, optimize on request; measure, don't guess, and every number carries the command that produced it. +--- + +# Performance Review + +## Lens + +This lane begins performance work with measurement of the running system, never with intuition about the code. Every claim carries the command that produced it and the number it produced. The USE method frames each resource — utilization, saturation, errors. The most expensive performance bug is the one "fixed" without measuring before and after. + +This lane's lens is inspired by the work of Brendan Gregg — author of *Systems Performance*, inventor of flame graphs and the USE method. + +## Mandate + +Assess and report by default. Apply optimizations only when the invocation explicitly asks — and then only with a before/after measurement pair. Never report an estimate where a measurement is obtainable this session. Findings outside this lane get a one-line handoff to the relevant lane skill under `skills/`. When you have enough numbers to conclude, conclude. + +## Discover the Ground Truth First + +Before measuring anything, establish what the product *is* and which latency its users actually feel: a CLI pays cold start per invocation (and per hook event, if it's invoked by hooks — the hook timeout is then the hard ceiling); a server pays per-request latency and saturation; a batch tool pays throughput. Read the entry points, the build config (bundling, minification, what's inlined), the manifest for dependency weight, and `CLAUDE.md`/`AGENTS.md` for stated performance constraints and deliberate tradeoffs (fork-per-event models, zero-daemon rules, chosen storage engines). Identify the shipped artifact users run — measure that, not the dev-mode path. Never carry numbers forward from documentation; a documented size or timing is a claim to re-measure. + +**Repo profile — recall, verify, persist.** Before deriving from scratch, recall a stored profile for this repo: a memory/brain store if one is available this session, else a well-known file (in genie-framework repos, `.genie/repo-profile.md`). For this lane the profile records the headline paths, hard ceilings, and baseline numbers with the commands that produced them. Baselines are the one profile entry you never trust — re-measure the headline path every run and report the delta against the stored baseline; that delta is often the most valuable finding. After the audit, persist the new numbers: update rather than duplicate, delete what proved wrong. + +## Workflow + +1. **Measure the headline path.** Time the user-felt path on the shipped artifact — `hyperfine` or a 10+-run loop; report median and spread, with first-run (cold cache) noted separately. Compare against any hard ceiling discovery found (hook timeouts, SLOs). Done when you have medians with exact commands. +2. **Profile startup/dispatch weight.** Identify what executes before useful work begins (eager imports, top-level side effects, artifact parse cost); check whether heavy dependencies load eagerly on paths that don't need them. Done when pre-work cost is characterized with evidence. +3. **Weigh the artifact and dependencies.** Measure the shipped size; attribute weight where feasible. Done when you know — from step 1's numbers, not assumption — whether size materially drives the headline path. +4. **Read the storage patterns.** Review hot-path queries and schema: per-row queries in loops, missing indexes vs actual WHERE/ORDER BY clauses, missing transactions around multi-statement writes, recomputation that grows with data size. Where a suspicion is testable, seed a throwaway store in an isolated tmpdir and time it. Done when each smell is confirmed with a citation (and a number where obtainable) or dismissed. +5. **Rank by user-felt impact** — frequency × cost: a path paid on every invocation outranks a slow rarely-used command. Done when the report is ordered by that product. + +## Grounded Reporting + +Every number was produced by a command this session and is quoted with that command; anything else is either inferred (calculation shown) or explicitly unmeasured (with the command that would measure it). + +## Output Format + +Lead with a one-sentence verdict anchored on the headline number vs its ceiling. Then findings ranked by user-felt impact, each with the measurement, the mechanism in plain language, and a recommended change whose expected effect is stated testably. Close with what was not measured and how to measure it. In a genie-framework repo, tag severities in `/review` vocabulary (SHIP / FIX-FIRST / BLOCKED); optimization campaigns bigger than one change belong in a wish via `/wish` with the baseline numbers as its acceptance criteria. + +## Pitfalls + +- Measure what users execute (the installed binary, the built bundle, the deployed server), never the dev-mode or source-interpreted path. +- First-run timings include OS cache warming — report the median of many runs and the first-run outlier separately; the outlier is often the realistic cold-start story. +- Retry/conflict patterns that implement correctness (claim conflicts, optimistic-lock retries) are design, not contention to optimize away — distinguish them from genuine saturation before flagging. +- Reversing an architectural constraint (adding a daemon to a zero-daemon design, adding a cache layer the docs forbid) is a cross-lane architecture proposal — hand it off with your numbers attached; the numbers are your contribution, the decision is not yours. +- Check the build config before recommending "enable minification/optimization" — partial minification or debug symbols are often deliberate. diff --git a/skills/qa/SKILL.md b/skills/qa/SKILL.md new file mode 100644 index 000000000..be3fe93b3 --- /dev/null +++ b/skills/qa/SKILL.md @@ -0,0 +1,48 @@ +--- +name: qa +description: Use when auditing test quality in any codebase — run the real suite, map coverage topology, rank untested behaviors by risk. Assess by default, write tests on request; tests are a spec, and the question is what change no test would catch. +--- + +# Quality Engineering Review + +## Lens + +This lane treats tests as a specification and a fear-reduction device — "test until fear turns to boredom." A suite's value is not its count but its topology: whether the behaviors that would hurt most are the ones pinned down. A test that never watched its subject fail proves nothing; a regression that broke once must be owned by a test forever. Coverage percentage is a proxy; the real question is "what change could I make that no test would catch?" + +This lane's lens is inspired by the work of Kent Beck, creator of test-driven development and the xUnit lineage. + +## Mandate + +Assess and report by default. Apply changes (writing tests, fixing flake) only when the invocation explicitly asks. Product bugs uncovered along the way, type holes, and performance cliffs get a one-line handoff to the relevant lane skill under `skills/`. When you have enough information to act, act. + +## Discover the Ground Truth First + +Find how this repo actually tests before judging: the framework and runner command (manifest scripts, CI workflows, `CLAUDE.md`/`AGENTS.md`), the test-file convention (colocated, mirrored tree, separate dir), the isolation patterns the repo has established (tmpdir fixtures, env-var redirection of global state, real-resource-vs-mock policy), and any named regression tests guarding past incidents. The repo's own testing doctrine — e.g. "real git repos, not mocks" or "tests drive the shipped bundle" — is the standard to hold it to. Then identify the product's highest-blast-radius behaviors from what it actually does (the entry points, the state it mutates, the money/data/permissions it touches). + +**Genie-framework repos**: wishes in `.genie/` carry acceptance criteria — the suite should own them; an accepted wish whose criteria no test exercises is a first-class gap. + +**Repo profile — recall, verify, persist.** Before deriving from scratch, recall a stored profile for this repo: a memory/brain store if one is available this session, else a well-known file (in genie-framework repos, `.genie/repo-profile.md`). For this lane the profile records the runner command, test conventions, isolation patterns, the blast-radius behavior list, and previously confirmed gaps. Recalled entries are hypotheses — a "confirmed gap" may have been closed since; re-check before reporting, and report drift as a finding. After the audit, persist what discovery learned: update rather than duplicate, delete what proved wrong. + +## Workflow + +1. **Run the suite** with real output captured: pass/fail/skip counts, duration, a second run if flake is suspected. Done when you have numbers, not assumptions. +2. **Map the topology.** Pair source modules with their tests per the repo's convention; tag every module tested / untested / partial. Done when the map is complete. +3. **Build the failure inventory.** For each high-blast-radius behavior: which test owns it? Read the owning test — does it exercise the failure mode or just the happy path, and would it fail for the right reason? Done when each behavior maps to a test file:line or a named gap. +4. **Judge quality, not presence.** Sample 3–5 test files: behavior vs implementation-detail assertions, state cleanup, whether CLI/API tests check exit codes and error output, whether concurrency tests genuinely race. Done when suite quality is characterized with cited examples. +5. **Rank the gaps** by blast radius × likelihood of change; sketch the failing test for each top gap (setup, action, assertion — including the repo's isolation pattern). Done when each sketch is executable on ask. + +## Grounded Reporting + +Only test results produced this session, with actual counts. A gap is "confirmed" only after searching for the test and reading near-misses; coverage judged from filenames alone is labeled "apparent." + +## Output Format + +Lead with a one-sentence verdict: suite state (numbers) plus the single scariest untested behavior. Then the ranked gap list with evidence-of-absence and test sketches, then suite-quality observations with examples, then cross-lane handoffs. In a genie-framework repo, tag severities in `/review` vocabulary (SHIP / FIX-FIRST / BLOCKED) and offer to turn the top gaps into a wish via `/wish` — the sketches become its acceptance criteria. + +## Pitfalls + +- A colocated test file is not coverage — read it before crediting; happy-path-only files leave the failure modes unowned. +- Respect the repo's realism choices: if the doctrine is real databases and real repos in tmpdirs, recommending mocks "for speed" reverses a deliberate decision. +- A regression test wired to a built artifact can fail because the artifact is stale — check the build before reporting a product regression. +- Concurrency tests asserting exactly-one-winner conflicts are testing correctness, not exhibiting flake. +- Every test sketch must include the repo's isolation pattern (tmpdir, env redirection) — a sketch that would touch the user's real global state is a defective recommendation. diff --git a/skills/repo-hygiene/SKILL.md b/skills/repo-hygiene/SKILL.md new file mode 100644 index 000000000..43b349e09 --- /dev/null +++ b/skills/repo-hygiene/SKILL.md @@ -0,0 +1,49 @@ +--- +name: repo-hygiene +description: Use when auditing repo hygiene in any codebase — file layout, git history, config sprawl, ignore contracts, open-source readiness. Assess by default, fix on request; treats the repository as a product whose users are contributors. +--- + +# Repo Hygiene Review + +## Lens + +This lane audits a repository as a product whose users are contributors — its layout, history, and configuration either invite people in or quietly turn them away. Judge the repo the way its next outside contributor will experience it: clone it, look around, read the log. Commit history is documentation; branching rules are UX; every config file is a promise that must still be true. + +This lane's lens is inspired by the work of Scott Chacon — GitHub co-founder, author of *Pro Git*, builder of GitButler. + +## Mandate + +Assess and report by default. Apply changes only when the invocation explicitly asks (e.g. "fix", "clean up", "apply"). When you spot a finding outside this lane (architecture, security, tests), name it in one line as a handoff to the relevant lane skill under `skills/` — do not investigate it yourself. When you have enough information to act, act; do not re-derive settled facts or survey options you will not pursue. + +## Discover the Ground Truth First + +Never judge against generic convention when the repo states its own. Before any verdict, read what exists of: `CLAUDE.md` / `AGENTS.md`, `README`, `CONTRIBUTING`, the package manifest, `.gitignore`, git hook tooling (husky, pre-commit, commitlint or equivalents), and CI config. These define the repo's *intended* contracts — your job is to find where reality has drifted from them, and where a contract is missing entirely. Deliberate tradeoffs documented there (bot commits, generated files kept on purpose, submodule workflows) are design, not defects. + +**Genie-framework repos**: if `.genie/` exists, its contract is: `wishes/`, `brainstorms/`, and `INDEX.md` are git-tracked; `genie.db` (and WAL/SHM siblings) must be ignored. Verify with `git check-ignore` and `git ls-files .genie/`. + +**Repo profile — recall, verify, persist.** Before deriving from scratch, recall a stored profile for this repo: a memory/brain store if one is available this session, else a well-known file (in genie-framework repos, `.genie/repo-profile.md`). For this lane the profile records the ignore contracts, config-to-enforcement map, commit conventions, and documented tradeoffs. Recalled anchors are hypotheses, not truth — spot-check them against current code and report drift as a finding. After the audit, persist what discovery learned back to the store: update rather than duplicate, delete what proved wrong. + +## Workflow + +1. **Walk the tree as a stranger.** `git ls-files` at top level plus `ls` for untracked clutter. Flag stray root files, tracked generated files, and ignore-contract violations both ways. Done when every top-level entry has a verdict: earns its place / sprawl / misplaced. +2. **Audit the ignore contracts.** `git check-ignore -v` against local-state and build-artifact paths; `git status --porcelain` for leakage. Done when each contract from discovery is confirmed or broken with evidence. +3. **Read the history.** `git log --oneline -50`: commit-convention conformance, bot-to-human ratio, whether human messages explain *why*; sample `git log --stat` for accidental large binaries or secrets. Done when history quality fits one sentence with examples. +4. **Census the configs.** For each config file, name what enforces it (a script, a hook, CI) — an unenforced config is sprawl; a hook that doesn't exist or isn't executable is a broken promise. Done when every config maps to an enforcement point or is flagged. +5. **Open-source readiness pass.** LICENSE present and consistent with the manifest; README answers what/install/first-command; contribution path stated; no internal URLs or credentials in tracked files. Done when a hypothetical public flip has a punch list. +6. **Rank and report** per the output format. + +## Grounded Reporting + +Every claim traces to a command output from this session; anything unchecked is stated as unchecked, not implied covered. Failed or erroring checks are reported with their output. + +## Output Format + +Lead with a one-sentence verdict on overall hygiene. Then findings ranked by cost-to-the-next-contributor, each with evidence (command + result or file path), why it matters, and the concrete action — precise enough to execute verbatim on ask. Close with cross-lane handoffs. In a genie-framework repo, tag severities in `/review`-compatible vocabulary (SHIP / FIX-FIRST / BLOCKED) and offer — without starting it — to crystallize the top findings into a wish via `/wish`. + +## Pitfalls + +- A documented tradeoff is not a defect: automated version-bump commits, intentional symlinks, submodule-managed directories, and deliberately tracked artifacts are only findings if they contradict what the repo says about itself. +- Bot commit noise is judged by whether it drowns out human history, not by its existence. +- Do not judge build-artifact tracking by generic convention; check what the release workflow actually consumes before calling it misplaced. +- Framework state directories (like genie's `.genie/`) mix tracked docs and ignored databases on purpose — verify against the framework's contract, not against "dotdirs shouldn't be tracked." +- Verify a config is genuinely dead (nothing loads it, no script or CI references it) before calling it sprawl. diff --git a/skills/supply-chain/SKILL.md b/skills/supply-chain/SKILL.md new file mode 100644 index 000000000..5a36a6ad3 --- /dev/null +++ b/skills/supply-chain/SKILL.md @@ -0,0 +1,47 @@ +--- +name: supply-chain +description: Use when auditing security and supply chain in any codebase — trust boundaries, credential handling, injection surfaces, update/release integrity, CI permissions, dependency pinning. Assess by default, harden on request; provenance or it didn't happen. +--- + +# Security & Supply-Chain Review + +## Lens + +This lane holds that every artifact a system trusts — a release binary, a dependency, a CI token, an inbound message — needs verifiable provenance, and "we downloaded it over HTTPS" is not provenance. Trust boundaries are enumerated, not assumed; the interesting question at each one is "what does an attacker who controls this input get?" Least privilege is the default; every credential and CI permission must justify its scope. + +This lane's lens is inspired by the work of Dan Lorenc, creator of Sigstore and founder of Chainguard. + +## Mandate + +Assess and report by default; this is a defensive audit of the user's own repo. Apply hardening only when the invocation explicitly asks. Do not build exploit tooling — demonstrating a finding means citing the code path and describing the impact, not weaponizing it. Findings outside this lane get a one-line handoff to the relevant lane skill under `skills/`. When the evidence supports a conclusion, state it. + +## Discover the Ground Truth First + +Enumerate before auditing. From the code, CI config, and `CLAUDE.md`/`AGENTS.md`, list: every point where external input enters (network listeners, webhook/hook stdin, message queues, downloaded artifacts, CLI args crossing privilege levels), every credential at rest (env vars, key files, tokens) and its handling, the update/release chain (how users get new versions, what verifies them), the CI surface (workflows, triggers, permissions, secrets, third-party actions), and the dependency posture (lockfile, count, where the build runs). Also collect the repo's *stated* security decisions — fail-closed contracts, documented trust delegations (e.g. "approval authority = membership in channel X"), known accepted risks — the audit judges whether they hold and whether their scope has silently widened, not whether you'd have chosen them. + +**Repo profile — recall, verify, persist.** Before deriving from scratch, recall a stored profile for this repo: a memory/brain store if one is available this session, else a well-known file (in genie-framework repos, `.genie/repo-profile.md`). For this lane the profile records the boundary map, credential inventory, trust delegations, and the previously verified-safe list. The verified-safe list is the dangerous entry — code changes since the last audit can invalidate it, so re-verify any safe-listed boundary the current diff touches and report scope drift as a finding. After the audit, persist what discovery learned: update rather than duplicate, delete what proved wrong. + +## Workflow + +1. **Confirm the boundary map.** Each entry point names its input source and what it can reach (paths, shell, DB writes, network). Done when nothing external enters unmapped. +2. **Audit the highest-exposure inbound path** (the one that runs most often or with most privilege — often a hook/webhook handler or message consumer): input validation at the boundary, fail-open vs fail-closed behavior on malformed input, side effects reachable from attacker-shaped payloads. Done when each handler has a verdict with file:line. +3. **Audit credentials and the update chain.** Key generation/storage/permissions; what signatures actually authenticate; secrets never logged, committed, or echoed in errors; update chain: pinned source? checksum or signature verification? time-of-check gaps? State plainly what a compromised update source gets. Done when each has a confirmed answer, not an assumption. +4. **Audit CI.** Per-workflow least-privilege `permissions:`, dangerous triggers (`pull_request_target`), secret exposure to forks, submodule checkout trust, actions pinned by SHA vs tag. Done when each workflow has a verdict. +5. **Sweep injection surfaces.** Grep for shell construction, path joins from user strings, and query string-building; trace each tainted variable to its origin. Done when each hit is confirmed parameterized/safe or flagged with the taint path. +6. **Rank by impact × exposure**: the update chain and always-running inbound handlers outrank local-only issues. + +## Grounded Reporting + +Every finding cites file:line read this session; every "verified safe" names what was checked. Findings are graded confirmed (path traced end to end), plausible (suspicious, taint not fully traced — with what remains), or not-assessed — a boundary is never safe because it "looks like" it validates. + +## Output Format + +Lead with a one-sentence verdict naming the most serious confirmed finding, or stating the audited surfaces are clean. Then findings ranked by impact × exposure, each with the trust boundary, evidence, plain-language impact, and the concrete hardening action. Include the verified-safe list — an audit that only reports holes hides its coverage. In a genie-framework repo, tag severities in `/review` vocabulary (SHIP / FIX-FIRST / BLOCKED); hardening campaigns become a wish via `/wish`, and anything actively exploitable is BLOCKED regardless of effort to fix. + +## Pitfalls + +- A documented trust delegation (fail-closed carve-outs, approval-by-channel-membership) is a decision to scope-check, not a hole to report — the finding is silent scope widening, not the delegation's existence. +- Before reporting "errors are swallowed" or "fails open," find the test that locks the fail-closed behavior and run or read it; bots misread fail-closed envelopes constantly. +- Do not report theoretical issues on inputs that never cross a privilege boundary — a user's own CLI args writing the user's own files is not a finding. +- The user's own local state files are not a secret store to flag; the audit target is what *external* input can write into them. +- Severity inflation destroys audit credibility: label something critical only when you can state the attacker, the input, and the concrete impact in one sentence. From 12270b21be801402d654c0eb6815c464184e5330 Mon Sep 17 00:00:00 2001 From: namastex888 Date: Thu, 9 Jul 2026 21:32:20 -0300 Subject: [PATCH 3/8] =?UTF-8?q?feat(workflows):=20/council=20native=20engi?= =?UTF-8?q?ne=20=E2=80=94=20template,=20lenses,=20stamp,=20lint=20(council?= =?UTF-8?q?-workflow=20G2)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .genie/wishes/council-workflow/WISH.md | 20 +- .../council-workflow/validate/g2-engine.sh | 10 +- biome.json | 1 + package.json | 1 + plugins/genie/references/lenses/deployer.md | 12 + plugins/genie/references/lenses/measurer.md | 12 + plugins/genie/references/lenses/operator.md | 12 + plugins/genie/references/lenses/questioner.md | 12 + plugins/genie/references/lenses/simplifier.md | 12 + plugins/genie/references/lenses/tracer.md | 12 + plugins/genie/scripts/council-stamp.cjs | 53 ++ plugins/genie/scripts/smart-install.js | 25 + plugins/genie/workflows/council.js | 679 ++++++++++++++++++ scripts/council-workflow-lint.ts | 272 +++++++ src/lib/council-workflow-stamp.test.ts | 102 +++ 15 files changed, 1223 insertions(+), 12 deletions(-) create mode 100644 plugins/genie/references/lenses/deployer.md create mode 100644 plugins/genie/references/lenses/measurer.md create mode 100644 plugins/genie/references/lenses/operator.md create mode 100644 plugins/genie/references/lenses/questioner.md create mode 100644 plugins/genie/references/lenses/simplifier.md create mode 100644 plugins/genie/references/lenses/tracer.md create mode 100644 plugins/genie/scripts/council-stamp.cjs create mode 100644 plugins/genie/workflows/council.js create mode 100644 scripts/council-workflow-lint.ts create mode 100644 src/lib/council-workflow-stamp.test.ts diff --git a/.genie/wishes/council-workflow/WISH.md b/.genie/wishes/council-workflow/WISH.md index b6a75a591..31df077fd 100644 --- a/.genie/wishes/council-workflow/WISH.md +++ b/.genie/wishes/council-workflow/WISH.md @@ -49,7 +49,7 @@ Genie's multi-perspective reasoning is split across two model-driven orchestrato ## Success Criteria -- [ ] `bun run lint:council-workflow` passes (exercised from G3 onward, once G1's lanes exist, and permanently via `bun run check`): template parses as ESM, `meta.name === 'council'`, zero banned APIs — `Date.now`, `Math.random`, `new Date(` (the Workflow tool spec: the determinism trio throws in scripts because it would break resume; `new Date(` banned broader-than-spec deliberately) plus `require(`, `import `, `process.`, `fs.` (workflows.md + tool spec: scripts are self-contained, no filesystem or Node.js API access) — every ROUTING member resolves to an existing lens file on disk (all 13), every lens file has required frontmatter +- [ ] `bun run lint:council-workflow` passes (exercised from G3 onward, once G1's lanes exist, and permanently via `bun run check`): template parses as a workflow async-body — the runtime shape: sole `export const meta` + top-level await/return, `export default` banned (module-legal ESM is the wrong contract; biome ignores `plugins/genie/workflows` for the same reason) — `meta.name === 'council'`, zero banned APIs — `Date.now`, `Math.random`, `new Date(` (the Workflow tool spec: the determinism trio throws in scripts because it would break resume; `new Date(` banned broader-than-spec deliberately) plus `require(`, `import `, `process.`, `fs.` (workflows.md + tool spec: scripts are self-contained, no filesystem or Node.js API access) — every ROUTING member resolves to an existing lens file on disk (all 13), every lens file has required frontmatter - [ ] All 7 lane skills exist with domain names; no real person's name in any `name:` field under `skills/` (grep denylist: chacon, ousterhout, hejlsberg, beck, gregg, lorenc, procida); inspiration line present in each - [ ] Stamp unit test green: installer function replaces `LENS_ROOT` placeholder with an absolute path and output lands at the expected `~/.claude/workflows/council.js` target (tmpdir-isolated via `GENIE_HOME`-style env override) - [ ] `skills/council/` no longer exists; `git grep -il 'specialist-panel'` returns 0 hits outside `.genie/attic/`, `CHANGELOG.md`, and this wish's artifacts @@ -77,9 +77,11 @@ Genie's multi-perspective reasoning is split across two model-driven orchestrato 1. `skills/{repo-hygiene,architecture,code-quality,qa,perf,supply-chain,dx-docs}/SKILL.md` — methodology preserved from the personal persona skills, frontmatter `name:` = lane name, body cites "inspired by the work of ", assess/fix contract kept per skill. **Acceptance Criteria:** -- [ ] All 7 SKILL.md files exist with lane-named frontmatter -- [ ] No real person's name in any `name:` field under `skills/` -- [ ] Each body contains an inspiration attribution line +- [x] All 7 SKILL.md files exist with lane-named frontmatter +- [x] No real person's name in any `name:` field under `skills/` +- [x] Each body contains an inspiration attribution line + +**Status:** DONE (2026-07-09) — gate `G1 PASS` (orchestrator-run), execution review SHIP (0 gaps ≥MEDIUM; LOW: handoff pointer was normalized across all 7 files — "specialist skill"→"lane skill under skills/" — an improvement over the 2 reported; NIT: unquoted YAML descriptions, house-cosmetic). Methodology bodies byte-identical to sources. Commit `0b222e17`. **Validation:** ```bash @@ -95,13 +97,15 @@ bash .genie/wishes/council-workflow/validate/g1-lane-skills.sh **Deliverables:** 1. `plugins/genie/workflows/council.js` — meta + Resolve/Round 1/Round 2/Synthesis/Persist phases, ROUTING table (absorbed from `members/routing.md`, incl. default trio + members override), LENSES map (`LENS_ROOT`-relative), deliberation and audit stage contracts, ≥2-members failure rule. 2. `plugins/genie/references/lenses/{questioner,simplifier,operator,deployer,measurer,tracer}.md` — frontmatter: name, modes, voice. -3. Stamp function as `src/lib/council-workflow-stamp.ts` (+ colocated `council-workflow-stamp.test.ts`, tmpdir-isolated): replaces the `LENS_ROOT` placeholder with the absolute installed-plugin path and writes `~/.claude/workflows/council.js`. **Call site pinned: the SessionStart hook (`plugins/genie/scripts/smart-install.js`, where `CLAUDE_PLUGIN_ROOT` is set), placed BEFORE its early-exit guards (deps-present, `GENIE_WORKER=1`) and idempotent via drift-check (rewrite only when stamped `LENS_ROOT` ≠ current `CLAUDE_PLUGIN_ROOT` or template hash changed). Re-stamp is therefore driven by the first session start after `claude plugin update` — the `genie update` CLI does not own it.** +3. Stamp function as `plugins/genie/scripts/council-stamp.cjs` (pure, dependency-injectable; loaded via `createRequire` from the ESM SessionStart hook — it must ship with the plugin, not the bundled CLI; deviation from the originally drafted `src/lib/*.ts` path, gate-tested via `src/lib/council-workflow-stamp.test.ts` importing the `.cjs`): replaces the `LENS_ROOT` placeholder with the absolute installed-plugin path and writes `~/.claude/workflows/council.js`. **Call site pinned: the SessionStart hook (`plugins/genie/scripts/smart-install.js`, where `CLAUDE_PLUGIN_ROOT` is set), placed BEFORE its early-exit guards (deps-present, `GENIE_WORKER=1`) and idempotent via drift-check (rewrite only when stamped `LENS_ROOT` ≠ current `CLAUDE_PLUGIN_ROOT` or template hash changed). Re-stamp is therefore driven by the first session start after `claude plugin update` — the `genie update` CLI does not own it.** 4. `scripts/council-workflow-lint.ts` + `package.json` script `lint:council-workflow`. **Acceptance Criteria:** -- [ ] Template parses as ESM (self-contained `bun build --no-bundle` check); `meta.name === 'council'`; zero banned APIs -- [ ] Every ROUTING member maps to a known lens NAME; the 6 deliberation cards exist with required frontmatter (full on-disk 13-lens resolution — including G1's lane skills — is asserted by `lint:council-workflow` from G3 onward, keeping G2 validatable in parallel with G1) -- [ ] Stamp unit test green in tmpdir isolation +- [x] Template parses as a workflow async-body (self-contained `lint:council-workflow --parse-only` check — runtime shape, not module-legal ESM); `meta.name === 'council'`; zero banned APIs; `export default` banned +- [x] Every ROUTING member maps to a known lens NAME; the 6 deliberation cards exist with required frontmatter (full on-disk 13-lens resolution — including G1's lane skills — is asserted by `lint:council-workflow` from G3 onward, keeping G2 validatable in parallel with G1) +- [x] Stamp unit test green in tmpdir isolation + +**Status:** DONE (2026-07-10) — gate `G2 PASS` (orchestrator-run), execution review FIX-FIRST → fixer → re-review SHIP (loop 1). CRITICAL fixed: engine unwrapped from `export default` to the runtime's body-style (sole `export const meta`, top-level await/return; ground-truth anchor: official Anthropic plugin workflows share this shape). HIGH fixed: parse checks validate the runtime shape (`--parse-only` mode, `export default` banned with verified negative proofs). MEDIUM fixed: silent lanes + unconvened lenses surface in synthesis and the Not Fully Audited section. Deviations (review-adjudicated): stamp lives at `plugins/genie/scripts/council-stamp.cjs`; `biome.json` ignores `plugins/genie/workflows` (biome can't parse body-style; necessary + minimally scoped). Residual NITs accepted: ROUTING row 4 drops `api` (deterministic-scorer double-match), `.cjs` outside biome globs (unit-tested). Full lint 13/13, biome 117 files clean, typecheck 0, stamp test 6/6. **Validation:** ```bash diff --git a/.genie/wishes/council-workflow/validate/g2-engine.sh b/.genie/wishes/council-workflow/validate/g2-engine.sh index 2604548cf..056dd5d0d 100755 --- a/.genie/wishes/council-workflow/validate/g2-engine.sh +++ b/.genie/wishes/council-workflow/validate/g2-engine.sh @@ -12,10 +12,12 @@ for banned in 'Date\.now' 'Math\.random' 'new Date\(' 'require\(' '^import ' 'pr if grep -Eq "$banned" "$t"; then echo "FAIL: banned API matching '$banned' in $t"; exit 1; fi done -# ESM parse check, self-contained (no cross-group deps): transpile fails on syntax errors. -tmp="$(mktemp -d)" -trap 'rm -rf "$tmp"' EXIT -bun build "$t" --no-bundle --outdir "$tmp" > /dev/null || { echo "FAIL: $t does not parse as ESM"; exit 1; } +# Parse check against the workflow RUNTIME shape, NOT module-legal ESM. The dynamic-workflow +# runtime runs the script as an async function body (top-level await/return are the contract) +# after extracting `export const meta` statically — so a plain ESM parse would REJECT the correct +# shape. --parse-only strips meta, forbids any other export (incl. `export default`), and +# transpiles the remainder as an async body; it fails on any syntax error. +bun scripts/council-workflow-lint.ts --parse-only "$t" || { echo "FAIL: $t does not parse as a workflow body"; exit 1; } for card in questioner simplifier operator deployer measurer tracer; do c="plugins/genie/references/lenses/$card.md" diff --git a/biome.json b/biome.json index 0296ca168..77baedc73 100644 --- a/biome.json +++ b/biome.json @@ -52,6 +52,7 @@ ".worktrees", ".claude", "plugins/genie/scripts", + "plugins/genie/workflows", ".docs-vendor", "docs" ] diff --git a/package.json b/package.json index 22d314257..583fa9321 100644 --- a/package.json +++ b/package.json @@ -25,6 +25,7 @@ "wishes:lint": "bun run scripts/wishes-lint.ts", "skills:audit": "bun run scripts/skills-audit.ts", "lint:complexity-budget": "bun run scripts/complexity-budget.ts", + "lint:council-workflow": "bun scripts/council-workflow-lint.ts", "check": "bun run typecheck && bun run lint && bun run dead-code && bun run skills:lint && bun run wishes:lint && bun test", "check:fast": "bun run typecheck && bun run lint && bun run dead-code && bun run skills:lint && bun run wishes:lint", "verify:release": "scripts/verify-release.sh" diff --git a/plugins/genie/references/lenses/deployer.md b/plugins/genie/references/lenses/deployer.md new file mode 100644 index 000000000..044cc0c99 --- /dev/null +++ b/plugins/genie/references/lenses/deployer.md @@ -0,0 +1,12 @@ +--- +name: deployer +modes: deliberation +voice: "Zero-config with infinite scale." +--- + +The deployer asks how this ships and how it rolls back before how it works. + +- Pushes for zero-config defaults and one obvious, documented install path. +- Treats every required manual step as a future outage waiting for the wrong operator. +- Wants the same artifact to scale from one machine to many without a rewrite. +- Names the deployment and rollback story explicitly, so it is a decision and not an afterthought. diff --git a/plugins/genie/references/lenses/measurer.md b/plugins/genie/references/lenses/measurer.md new file mode 100644 index 000000000..3ce508cc3 --- /dev/null +++ b/plugins/genie/references/lenses/measurer.md @@ -0,0 +1,12 @@ +--- +name: measurer +modes: deliberation +voice: "Measure, don't guess." +--- + +The measurer refuses claims that arrive without a number and the command that produced it. + +- Asks what signal would confirm or refute each position before the council commits to it. +- Wants the metric defined before the feature, not bolted on after it ships. +- Treats "it feels faster" or "it should scale" as a hypothesis, never as evidence. +- Names the measurement each proposal still owes, so decisions rest on data instead of confidence. diff --git a/plugins/genie/references/lenses/operator.md b/plugins/genie/references/lenses/operator.md new file mode 100644 index 000000000..6aa98b1bf --- /dev/null +++ b/plugins/genie/references/lenses/operator.md @@ -0,0 +1,12 @@ +--- +name: operator +modes: deliberation +voice: "No one wants to run your code." +--- + +The operator judges a design by the 3am pager, not the demo. + +- Asks who actually runs this, how it fails, and what a tired human does when it does. +- Prefers boring, observable, restartable behavior over clever fragility. +- Flags anything that quietly assumes the happy path holds in production. +- Wants failure modes named up front, not discovered during the first incident. diff --git a/plugins/genie/references/lenses/questioner.md b/plugins/genie/references/lenses/questioner.md new file mode 100644 index 000000000..8d6119640 --- /dev/null +++ b/plugins/genie/references/lenses/questioner.md @@ -0,0 +1,12 @@ +--- +name: questioner +modes: deliberation +voice: "Why? Is there a simpler way?" +--- + +The questioner challenges assumptions before accepting any framing. + +- Opens by asking what problem is actually being solved, and whether it is the real problem or a proxy for it. +- Names the load-bearing assumption inside every proposal and asks what breaks if it turns out false. +- Prefers the simplest thing that could work, and treats added machinery as debt until it is justified. +- Separates "we decided this" from "we assumed this", and drags the second into the open where the council can see it. diff --git a/plugins/genie/references/lenses/simplifier.md b/plugins/genie/references/lenses/simplifier.md new file mode 100644 index 000000000..a9f6114d7 --- /dev/null +++ b/plugins/genie/references/lenses/simplifier.md @@ -0,0 +1,12 @@ +--- +name: simplifier +modes: deliberation +voice: "Delete code. Ship features." +--- + +The simplifier measures progress in concepts removed, not lines added. + +- Asks which parts can be deleted, merged, or defaulted away before anything new is introduced. +- Treats every option, flag, and abstraction as a carrying cost paid on every future read. +- Pushes for the smaller design and lets real, observed demand justify the larger one. +- Names the specific complexity each proposal adds, so the council can weigh it against the benefit. diff --git a/plugins/genie/references/lenses/tracer.md b/plugins/genie/references/lenses/tracer.md new file mode 100644 index 000000000..dc576f027 --- /dev/null +++ b/plugins/genie/references/lenses/tracer.md @@ -0,0 +1,12 @@ +--- +name: tracer +modes: deliberation +voice: "You will debug this in production." +--- + +The tracer assumes the incident is inevitable and asks how you will find the cause at 3am. + +- Wants high-cardinality context on the request path, not just aggregate logs and dashboards. +- Asks what state you would need at the moment of failure and whether anything captures it. +- Values a design you can reason about under pressure over one that is merely elegant on paper. +- Names the debugging story for each proposal, so observability is designed in rather than retrofitted. diff --git a/plugins/genie/scripts/council-stamp.cjs b/plugins/genie/scripts/council-stamp.cjs new file mode 100644 index 000000000..1109b1ac9 --- /dev/null +++ b/plugins/genie/scripts/council-stamp.cjs @@ -0,0 +1,53 @@ +'use strict'; + +/** + * council-stamp: stamp the /council workflow template into ~/.claude/workflows. + * + * Plugins cannot ship Claude Code workflows directly, so the template lives in + * the plugin (plugins/genie/workflows/council.js) with a `__GENIE_LENS_ROOT__` + * placeholder, and the SessionStart hook (smart-install.js) calls this on every + * start to write the stamped file to ~/.claude/workflows/council.js. + * + * Pure and dependency-injectable: all paths are arguments, so the unit test can + * drive it entirely inside a tmpdir. CommonJS (.cjs) so it is requireable from + * the ESM smart-install.js via createRequire, and from bun:test. + */ + +const fs = require('node:fs'); +const path = require('node:path'); + +const PLACEHOLDER = '__GENIE_LENS_ROOT__'; +const TARGET_NAME = 'council.js'; + +/** + * Stamp the template's LENS_ROOT placeholder with the absolute plugin path and + * write it to /council.js. + * + * Idempotent: the stamped bytes are a pure function of (template, pluginRoot), + * so an unchanged template and an unchanged root produce output identical to + * what is already on disk — in that case we skip the write. Any drift (template + * updated, root changed, or the target hand-edited) makes the bytes differ and + * we rewrite, which is also self-healing. + * + * @param {{templatePath: string, pluginRoot: string, targetDir: string}} opts + * @returns {{action: 'written'|'skipped', targetPath: string}} + */ +function stampCouncilWorkflow({ templatePath, pluginRoot, targetDir } = {}) { + if (!templatePath || !pluginRoot || !targetDir) { + throw new Error('stampCouncilWorkflow requires templatePath, pluginRoot, and targetDir'); + } + + const template = fs.readFileSync(templatePath, 'utf8'); + const stamped = template.split(PLACEHOLDER).join(pluginRoot); + const targetPath = path.join(targetDir, TARGET_NAME); + + if (fs.existsSync(targetPath) && fs.readFileSync(targetPath, 'utf8') === stamped) { + return { action: 'skipped', targetPath }; + } + + fs.mkdirSync(targetDir, { recursive: true }); + fs.writeFileSync(targetPath, stamped, 'utf8'); + return { action: 'written', targetPath }; +} + +module.exports = { stampCouncilWorkflow, PLACEHOLDER, TARGET_NAME }; diff --git a/plugins/genie/scripts/smart-install.js b/plugins/genie/scripts/smart-install.js index 05910f09f..f40a3fc19 100644 --- a/plugins/genie/scripts/smart-install.js +++ b/plugins/genie/scripts/smart-install.js @@ -13,9 +13,14 @@ import { execSync, spawnSync } from 'node:child_process'; * - Version marker management */ import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs'; +import { createRequire } from 'node:module'; import { homedir } from 'node:os'; import { join } from 'node:path'; +// This file is ESM (plugin package.json is type:module), so load the CommonJS +// council-stamp helper through createRequire rather than a bare require. +const requireCjs = createRequire(import.meta.url); + const ROOT = process.env.CLAUDE_PLUGIN_ROOT || join(homedir(), '.claude', 'plugins', 'genie'); const GENIE_DIR = join(homedir(), '.genie'); const MARKER = join(GENIE_DIR, '.install-version'); @@ -387,6 +392,26 @@ function adviseGenieCliInstall() { // Main execution try { + // Stamp the /council workflow template into ~/.claude/workflows on every session + // start. This runs BEFORE the early-exit guards below (worker fast-path and + // deps-already-present) so a plugin update re-stamps LENS_ROOT even on machines + // that would otherwise skip all install work. The stamp is idempotent — it only + // writes when the stamped output would differ — and is fully sandboxed in + // try/catch so it can never break session start. + try { + const { stampCouncilWorkflow } = requireCjs('./council-stamp.cjs'); + const stampResult = stampCouncilWorkflow({ + templatePath: join(ROOT, 'workflows', 'council.js'), + pluginRoot: ROOT, + targetDir: join(homedir(), '.claude', 'workflows'), + }); + if (stampResult.action === 'written') { + console.error(`Stamped /council workflow to ${stampResult.targetPath}`); + } + } catch (e) { + console.error(`Warning: could not stamp /council workflow: ${e.message}`); + } + // Workers inherit parent's deps — skip all checks to reduce spawn latency (#712) if (process.env.GENIE_WORKER === '1') { process.exit(0); diff --git a/plugins/genie/workflows/council.js b/plugins/genie/workflows/council.js new file mode 100644 index 000000000..553b086a0 --- /dev/null +++ b/plugins/genie/workflows/council.js @@ -0,0 +1,679 @@ +export const meta = { + name: 'council', + description: 'Convene a lens council: deliberate a decision, or audit the repo across specialist lanes.', + phases: [ + { title: 'Resolve' }, + { title: 'Round 1' }, + { title: 'Round 2' }, + { title: 'Synthesis' }, + { title: 'Persist' }, + ], +}; + +// The /council workflow engine. Two modes over one shared lens library: +// deliberation — Resolve -> Round 1 -> Round 2 -> Synthesis (advisory, writes nothing) +// audit — Resolve -> Round 1 -> Synthesis -> Persist (assess-only; single writer merges the profile) +// +// This template ships inside the genie plugin. LENS_ROOT is a placeholder that +// the install-time stamp (scripts/council-stamp.cjs) rewrites to the absolute +// installed-plugin path before the file lands in ~/.claude/workflows/council.js. +// All file access happens inside spawned agents (Read/Glob/Bash/Write); the +// script itself touches no filesystem and no Node globals, so runs stay resumable. + +const LENS_ROOT = '__GENIE_LENS_ROOT__'; + +// Lens name -> path relative to LENS_ROOT. Seven audit lanes (persona skills, +// reached through the plugin's `skills` symlink) plus six deliberation cards. +const LENSES = { + 'repo-hygiene': 'skills/repo-hygiene/SKILL.md', + architecture: 'skills/architecture/SKILL.md', + 'code-quality': 'skills/code-quality/SKILL.md', + qa: 'skills/qa/SKILL.md', + perf: 'skills/perf/SKILL.md', + 'supply-chain': 'skills/supply-chain/SKILL.md', + 'dx-docs': 'skills/dx-docs/SKILL.md', + questioner: 'references/lenses/questioner.md', + simplifier: 'references/lenses/simplifier.md', + operator: 'references/lenses/operator.md', + deployer: 'references/lenses/deployer.md', + measurer: 'references/lenses/measurer.md', + tracer: 'references/lenses/tracer.md', +}; + +// Keyword routing for deliberation, absorbed from the old council members/routing.md. +// The four retired council lenses are remapped onto lanes: benchmarker -> perf, +// sentinel -> supply-chain, ergonomist -> dx-docs, architect -> architecture. +const ROUTING = [ + { + keywords: ['architecture', 'design', 'system', 'interface', 'api'], + members: ['questioner', 'architecture', 'simplifier', 'perf'], + }, + { + keywords: ['performance', 'latency', 'throughput', 'scale'], + members: ['perf', 'questioner', 'architecture', 'measurer'], + }, + { keywords: ['security', 'auth', 'secret', 'blast radius'], members: ['questioner', 'supply-chain', 'simplifier'] }, + { keywords: ['endpoint', 'dx', 'developer', 'sdk'], members: ['questioner', 'simplifier', 'dx-docs', 'deployer'] }, + { + keywords: ['ops', 'deploy', 'infra', 'ci/cd', 'monitoring'], + members: ['operator', 'deployer', 'tracer', 'measurer'], + }, + { keywords: ['debug', 'trace', 'observability', 'logging'], members: ['tracer', 'measurer', 'perf'] }, + { keywords: ['plan', 'scope', 'wish', 'feature'], members: ['questioner', 'simplifier', 'architecture', 'dx-docs'] }, +]; + +// Default trio when no keyword matches: wrong-problem, over-engineering, short-term thinking. +const DEFAULT_TRIO = ['questioner', 'simplifier', 'architecture']; + +// Audit default roster: the seven specialist lanes. +const AUDIT_ROSTER = ['repo-hygiene', 'architecture', 'code-quality', 'qa', 'perf', 'supply-chain', 'dx-docs']; + +const SEVERITY = ['critical', 'high', 'medium', 'low']; + +const RESOLVE_SCHEMA = { + type: 'object', + required: ['resolved', 'missing'], + properties: { + resolved: { + type: 'array', + items: { + type: 'object', + required: ['name', 'absPath'], + properties: { name: { type: 'string' }, absPath: { type: 'string' } }, + }, + }, + missing: { + type: 'array', + items: { + type: 'object', + required: ['name', 'relPath'], + properties: { name: { type: 'string' }, relPath: { type: 'string' } }, + }, + }, + }, +}; + +const R1_DELIBERATION_SCHEMA = { + type: 'object', + required: ['member', 'position', 'assumptions'], + properties: { + member: { type: 'string' }, + position: { type: 'string', description: '2-4 opinionated paragraphs' }, + assumptions: { type: 'array', items: { type: 'string' } }, + }, +}; + +const FINDING_SCHEMA = { + type: 'object', + required: ['severity', 'summary', 'evidence', 'action'], + properties: { + severity: { type: 'string', enum: SEVERITY }, + summary: { type: 'string' }, + evidence: { type: 'string', description: 'file:line or command + its output' }, + action: { type: 'string' }, + }, +}; + +const PROFILE_UPDATE_SCHEMA = { + type: 'object', + required: ['anchor', 'change', 'note'], + properties: { + anchor: { type: 'string' }, + change: { type: 'string', enum: ['new', 'changed', 'invalidated'] }, + note: { type: 'string' }, + }, +}; + +const R1_AUDIT_SCHEMA = { + type: 'object', + required: ['lane', 'verdict', 'findings', 'verifiedSound', 'couldNotVerify', 'profileUpdates'], + properties: { + lane: { type: 'string' }, + verdict: { type: 'string', description: 'one-sentence lane verdict' }, + findings: { type: 'array', items: FINDING_SCHEMA }, + verifiedSound: { type: 'string' }, + couldNotVerify: { type: 'string' }, + profileUpdates: { type: 'array', items: PROFILE_UPDATE_SCHEMA }, + }, +}; + +const R2_SCHEMA = { + type: 'object', + required: ['member', 'strongestOther', 'challenge', 'changed', 'evolution'], + properties: { + member: { type: 'string' }, + strongestOther: { type: 'string' }, + challenge: { type: 'string' }, + changed: { type: 'boolean' }, + evolution: { type: 'string' }, + }, +}; + +const SYNTH_DELIBERATION_SCHEMA = { + type: 'object', + required: ['executiveSummary', 'consensus', 'tensions', 'evolution', 'recommendations', 'dissent'], + properties: { + executiveSummary: { type: 'string' }, + consensus: { type: 'array', items: { type: 'string' } }, + tensions: { type: 'array', items: { type: 'string' } }, + evolution: { type: 'string' }, + recommendations: { + type: 'array', + items: { + type: 'object', + required: ['priority', 'recommendation', 'rationale', 'risk'], + properties: { + priority: { type: 'string', enum: ['P0', 'P1', 'P2'] }, + recommendation: { type: 'string' }, + rationale: { type: 'string' }, + risk: { type: 'string' }, + }, + }, + }, + dissent: { type: 'string', description: 'minority views preserved verbatim' }, + }, +}; + +const SYNTH_AUDIT_SCHEMA = { + type: 'object', + required: ['verdict', 'laneVerdicts', 'topFindings', 'notFullyAudited'], + properties: { + verdict: { type: 'string', description: 'one-sentence overall verdict' }, + laneVerdicts: { + type: 'array', + items: { + type: 'object', + required: ['lane', 'verdict'], + properties: { lane: { type: 'string' }, verdict: { type: 'string' } }, + }, + }, + topFindings: { + type: 'array', + maxItems: 10, + items: { + type: 'object', + required: ['lanes', 'severity', 'summary', 'evidence', 'action'], + properties: { + lanes: { type: 'array', items: { type: 'string' } }, + severity: { type: 'string', enum: SEVERITY }, + summary: { type: 'string' }, + evidence: { type: 'string' }, + action: { type: 'string' }, + }, + }, + }, + notFullyAudited: { type: 'array', items: { type: 'string' } }, + }, +}; + +const PERSIST_SCHEMA = { + type: 'object', + required: ['written', 'path'], + properties: { + written: { type: 'boolean' }, + path: { type: 'string' }, + note: { type: 'string' }, + }, +}; + +function failure(error, detail) { + return detail === undefined ? { ok: false, error } : { ok: false, error, detail }; +} + +function dedupe(names) { + return [...new Set(names)]; +} + +// Best-match routing: the row with the most keyword hits wins; ties break by order. +function classifyMembers(topic) { + const hay = topic.toLowerCase(); + let best = null; + let bestScore = 0; + for (const row of ROUTING) { + let score = 0; + for (const kw of row.keywords) { + if (hay.includes(kw)) score += 1; + } + if (score > bestScore) { + bestScore = score; + best = row; + } + } + return best ? best.members.slice() : DEFAULT_TRIO.slice(); +} + +// A `members` override bypasses routing. Unknown names are logged and ignored; +// if nothing valid survives, fall back to the mode's default roster. +function selectRoster(mode, topic, membersArg) { + if (Array.isArray(membersArg) && membersArg.length) { + const valid = []; + for (const raw of membersArg) { + const name = typeof raw === 'string' ? raw.trim() : ''; + if (name && Object.hasOwn(LENSES, name)) { + valid.push(name); + } else if (name) { + log(`Ignoring unknown member "${name}" — not in the lens library.`); + } + } + if (valid.length) return dedupe(valid); + log('Members override had no known lenses — using the default roster.'); + } + if (mode === 'audit') return AUDIT_ROSTER.slice(); + return classifyMembers(topic); +} + +function rosterPaths(names) { + return names.map((name) => ({ name, relPath: LENSES[name] })); +} + +function collectProfileUpdates(responded) { + const out = []; + for (const entry of responded) { + const ups = entry.response && Array.isArray(entry.response.profileUpdates) ? entry.response.profileUpdates : []; + for (const u of ups) out.push({ lane: entry.member, anchor: u.anchor, change: u.change, note: u.note }); + } + return out; +} + +function resolvePrompt(wanted) { + const lines = wanted.map((w) => `- ${w.name}: ${w.relPath}`).join('\n'); + return [ + 'You are the council stage-0 lens resolver. Keep effort low: verify files, do not read or apply them.', + '', + `LENS_ROOT (the installed genie plugin directory): ${LENS_ROOT}`, + '', + 'Roster to resolve (lens name -> path relative to LENS_ROOT):', + lines, + '', + 'For each entry, confirm the lens file exists on disk:', + '1. First try LENS_ROOT joined with the relative path (use Read or Glob).', + '2. If NONE of them exist under LENS_ROOT, the stamped root is stale. Glob your home', + ' ".claude" tree for the marker "**/plugins/**/references/lenses/questioner.md", take the', + ' plugin root (the directory two levels above "references/lenses"), and re-resolve every', + ' relative path against that discovered root instead.', + '', + 'Return the schema object only: resolved = [{name, absPath}] for files that exist (absolute', + 'paths), missing = [{name, relPath}] for files you could not find anywhere. A missing lens is', + 'reported, never silently dropped.', + ].join('\n'); +} + +function round1DeliberationPrompt(member, topic) { + return [ + `You are the council's ${member.name} lens. Read your lens file at ${member.absPath} and adopt its voice and method.`, + '', + 'Council topic:', + topic, + '', + 'Deliver an opinionated Round 1 position: 2-4 paragraphs that take clear positions and name the', + 'assumptions behind them. Ground each claim in your expertise or cite the evidence you would need.', + "Do not retreat into neutrality — the council's value is distinct, committed perspectives.", + '', + `Return the schema object: member = "${member.name}", position = your 2-4 paragraphs, assumptions = the named assumptions.`, + ].join('\n'); +} + +function round1AuditPrompt(member, focus) { + const target = focus ? focus : 'the whole repository at the current working directory'; + return [ + `You are the ${member.name} audit lane. Read your lens file at ${member.absPath} and follow its`, + 'methodology exactly, including its ground-truth discovery step.', + '', + `Audit target: ${target}.`, + '', + 'You are strictly assess-only for this run: make NO edits and do NOT write the repo profile —', + "return your profile updates as data instead. Run the lens's real commands and measurements", + '(Bash/Read/Glob). Rank findings by severity. For each finding give: severity', + '(critical/high/medium/low), a one-sentence summary, evidence (file:line or a command plus its', + 'output), and a recommended action. Also give a one-sentence lane verdict, what you verified as', + 'sound, and what you could not verify.', + '', + `Return the schema object: lane = "${member.name}", verdict, findings[], verifiedSound,`, + 'couldNotVerify, and profileUpdates[] (knowledge anchors you would change — as data; you write nothing).', + ].join('\n'); +} + +function round2Prompt(self, others, topic) { + const otherBlocks = others.map((o) => `### ${o.member}\n${o.response.position}`).join('\n\n'); + return [ + `You are the council's ${self.member} lens in Round 2 of a deliberation. You already delivered a`, + "Round 1 position; the other members' positions follow, attributed by name.", + '', + 'Council topic:', + topic, + '', + 'YOUR Round 1 position:', + self.response.position, + '', + "OTHER MEMBERS' Round 1 positions:", + otherBlocks, + '', + 'Reply, staying in your lens voice, with: (1) the single strongest point another member made,', + '(2) at least one point you challenge or refine, and (3) whether your position changed and why.', + '', + `Return the schema object: member = "${self.member}", strongestOther, challenge, changed (boolean), evolution.`, + ].join('\n'); +} + +function synthesisDeliberationPrompt(topic, responded, round2) { + const r2by = {}; + for (const r of round2) { + if (r.response) r2by[r.member] = r.response; + } + const blocks = responded + .map((x) => { + const r2 = r2by[x.member]; + const r2text = r2 + ? `${r2.evolution} (challenge: ${r2.challenge}; strongest other: ${r2.strongestOther})` + : 'no Round 2 response'; + return `### ${x.member}\nRound 1:\n${x.response.position}\n\nRound 2:\n${r2text}`; + }) + .join('\n\n'); + return [ + 'You are the council synthesizer. You did not deliberate; you integrate the members into one', + 'advisory report. Advisory only — no voting, no verdict, no gate-keeping language. Preserve', + "minority views in the members' own words.", + '', + 'Council topic:', + topic, + '', + 'Members and their two rounds:', + blocks, + '', + 'Produce: an executive summary (2-3 sentences); the points of consensus; the key tensions and', + 'unresolved disagreements; how thinking evolved between rounds; prioritized recommendations', + '(P0/P1/P2, each with rationale and the risk if ignored); and a Dissent section that preserves', + 'minority views verbatim.', + '', + 'Return the schema object: executiveSummary, consensus[], tensions[], evolution, recommendations[], dissent.', + ].join('\n'); +} + +function synthesisAuditPrompt(focus, responded, silentRound1, notConvened) { + const target = focus ? focus : 'the whole repository'; + const blocks = responded + .map((x) => { + const r = x.response; + return [ + `### ${x.member} (verdict: ${r.verdict})`, + `Findings: ${JSON.stringify(r.findings)}`, + `Verified sound: ${r.verifiedSound}`, + `Could not verify: ${r.couldNotVerify}`, + ].join('\n'); + }) + .join('\n\n'); + const silent = silentRound1.length ? silentRound1.join(', ') : '(none)'; + const absent = notConvened.length ? notConvened.join(', ') : '(none)'; + return [ + 'You are the panel synthesizer. Integrate the lanes into ONE cross-cutting report. Advisory only:', + 'approved findings route to /wish — there is no approval mechanism here.', + '', + `Audit target: ${target}.`, + '', + 'Lane outputs:', + blocks, + '', + `Lanes that resolved but returned NOTHING this run: ${silent}.`, + `Lenses that never convened (lens file did not resolve): ${absent}.`, + 'These lanes were NOT audited: never average them into a finding or verdict, and never imply the', + 'area they cover is clean. They are surfaced separately in the report, so do NOT list them in', + 'notFullyAudited[] — reserve that array for lanes that DID respond but whose coverage is partial.', + '', + 'Do the synthesis work: (1) DEDUPE — when the same fact appears in multiple lanes, merge it into', + 'one finding that names every lane it came from; (2) RESOLVE CONFLICTS — when lanes disagree,', + 'present both judgments with their evidence and recommend one side, saying why; (3) RE-RANK', + "GLOBALLY by real risk across lanes, not per-lane labels — a lane's high may be the panel's", + 'medium; (4) keep only the top findings (at most 10); (5) flag any responding lane whose audit', + 'was not fully completed — never silently average it in.', + '', + 'Return the schema object: verdict (one sentence), laneVerdicts[], topFindings[] (<=10, each', + 'naming its lane(s), severity, summary, evidence, action), notFullyAudited[].', + ].join('\n'); +} + +function persistPrompt(updates) { + return [ + 'You are the panel single profile writer. Merge the knowledge updates below into the repo profile', + 'at /.genie/repo-profile.md, where is the toplevel of the current git repository', + '(resolve it with `git rev-parse --show-toplevel`). Create the file if it does not exist. You are', + 'the ONLY writer this run, so there is no conflict to reconcile — apply every update.', + '', + 'These are knowledge anchors (new / changed / invalidated), NOT fixes. Do not modify any source', + 'file; write only the profile.', + '', + 'Updates:', + JSON.stringify(updates, null, 2), + '', + 'Merge rules: a "new" anchor is added; a "changed" anchor replaces the prior note for that anchor;', + 'an "invalidated" anchor is removed. Keep the profile concise and de-duplicated.', + '', + 'Return the schema object: written (boolean), path (the profile path you wrote), note (one line on what changed).', + ].join('\n'); +} + +function renderDeliberation(topic, convened, notConvened, synth) { + const composition = convened.map((c) => `- ${c.name}`).join('\n'); + const consensus = synth.consensus.length ? synth.consensus.map((c) => `- ${c}`).join('\n') : '- (none recorded)'; + const tensions = synth.tensions.length ? synth.tensions.map((t) => `- ${t}`).join('\n') : '- (none recorded)'; + const recs = synth.recommendations + .map((r) => `| ${r.priority} | ${r.recommendation} | ${r.rationale} | ${r.risk} |`) + .join('\n'); + const missing = notConvened.length ? `\n\n_Not convened (lens file missing): ${notConvened.join(', ')}_` : ''; + return [ + `# Council Report: ${topic}`, + '', + '## Executive Summary', + synth.executiveSummary, + '', + '## Council Composition', + composition, + '', + '## Consensus', + consensus, + '', + '## Tensions', + tensions, + '', + '## Evolution', + synth.evolution, + '', + '## Recommendations', + '| Priority | Recommendation | Rationale | Risk if ignored |', + '| --- | --- | --- | --- |', + recs, + '', + '## Dissent', + synth.dissent, + `${missing}`, + ].join('\n'); +} + +function renderAudit(focus, notConvened, silentRound1, synth, persist) { + const target = focus ? focus : 'whole repository'; + const laneVerdicts = synth.laneVerdicts.map((l) => `- **${l.lane}** — ${l.verdict}`).join('\n'); + const findings = synth.topFindings + .map( + (f, i) => + `${i + 1}. **[${f.severity}]** (${f.lanes.join(', ')}) ${f.summary}\n - Evidence: ${f.evidence}\n - Action: ${f.action}`, + ) + .join('\n'); + // A lane that returned nothing, or a lens that never convened, is reported "not audited" here — + // never silently dropped and never averaged in. Merge the synthesizer's partial-coverage list + // with the code-known silent lanes and unresolved lenses. + const notAuditedEntries = [ + ...synth.notFullyAudited, + ...silentRound1.map((n) => `${n} (resolved but returned nothing)`), + ...notConvened.map((n) => `${n} (lens file missing)`), + ]; + const notAudited = notAuditedEntries.length + ? notAuditedEntries.map((n) => `- ${n}`).join('\n') + : '- (all convened lanes fully audited)'; + const profile = persist?.written + ? `Profile updated at ${persist.path}${persist.note ? ` — ${persist.note}` : ''}.` + : 'No profile updates persisted.'; + return [ + `# Council Audit: ${target}`, + '', + '## Verdict', + synth.verdict, + '', + '## Lane Verdicts', + laneVerdicts, + '', + '## Top Findings', + findings, + '', + '## Not Fully Audited', + notAudited, + '', + '## Profile', + profile, + '', + '_Findings are advisory. Route the ones you approve into a wish via /wish._', + ].join('\n'); +} + +if (!args || typeof args !== 'object') { + return failure('No input received. Try /council to deliberate, or /council audit [focus] to audit.'); +} + +const mode = args.mode === 'audit' ? 'audit' : 'deliberation'; +const topic = typeof args.topic === 'string' ? args.topic.trim() : ''; +const focus = typeof args.focus === 'string' ? args.focus.trim() : ''; + +if (mode === 'deliberation' && !topic) { + return failure('No topic to deliberate. Try /council .'); +} + +const roster = selectRoster(mode, topic, args.members); +if (!roster.length) { + return failure('No lenses selected. Provide --members, or a topic that routes to a lens.'); +} +log(`Mode: ${mode}. Roster: ${roster.join(', ')}.`); + +// Phase 1 — Resolve lens files on disk (fail-open: a stale LENS_ROOT is rediscovered by the agent). +phase('Resolve'); +const resolveResult = await agent(resolvePrompt(rosterPaths(roster)), { + label: 'resolve-lenses', + phase: 'Resolve', + effort: 'low', + schema: RESOLVE_SCHEMA, +}); +if (!resolveResult || !Array.isArray(resolveResult.resolved) || resolveResult.resolved.length < 2) { + return failure('Fewer than two lenses resolved on disk — cannot convene the council.', { + requested: roster, + resolveResult, + }); +} +const convened = resolveResult.resolved; +const notConvened = (resolveResult.missing || []).map((m) => m.name); +if (notConvened.length) log(`Not convened (lens file missing): ${notConvened.join(', ')}.`); + +// Phase 2 — Round 1: every convened lens delivers in parallel. +phase('Round 1'); +const round1Raw = await parallel( + convened.map( + (member) => () => + agent(mode === 'audit' ? round1AuditPrompt(member, focus) : round1DeliberationPrompt(member, topic), { + label: `round1-${member.name}`, + phase: 'Round 1', + schema: mode === 'audit' ? R1_AUDIT_SCHEMA : R1_DELIBERATION_SCHEMA, + }), + ), +); +const round1 = convened.map((member, i) => ({ + member: member.name, + absPath: member.absPath, + response: round1Raw[i], +})); +const responded = round1.filter((x) => x.response); +const silentRound1 = round1.filter((x) => !x.response).map((x) => x.member); +if (silentRound1.length) log(`No Round 1 response: ${silentRound1.join(', ')}.`); +if (responded.length < 2) { + return failure('Fewer than two lenses delivered Round 1 — the council cannot proceed.', { + silent: silentRound1, + notConvened, + }); +} + +// Phase 3 — Round 2: deliberation only, one fresh agent per responding member. +phase('Round 2'); +let round2 = []; +if (mode === 'deliberation') { + const round2Raw = await parallel( + responded.map( + (self) => () => + agent( + round2Prompt( + self, + responded.filter((o) => o.member !== self.member), + topic, + ), + { + label: `round2-${self.member}`, + phase: 'Round 2', + schema: R2_SCHEMA, + }, + ), + ), + ); + round2 = responded.map((self, i) => ({ member: self.member, response: round2Raw[i] })); + const silentRound2 = round2.filter((r) => !r.response).map((r) => r.member); + if (silentRound2.length) log(`No Round 2 response: ${silentRound2.join(', ')}.`); +} else { + log('Round 2 skipped — audit mode is single-round.'); +} + +// Phase 4 — Synthesis: one agent integrates everything into an advisory report. +phase('Synthesis'); +const synthesis = await agent( + mode === 'audit' + ? synthesisAuditPrompt(focus, responded, silentRound1, notConvened) + : synthesisDeliberationPrompt(topic, responded, round2), + { + label: 'synthesis', + phase: 'Synthesis', + schema: mode === 'audit' ? SYNTH_AUDIT_SCHEMA : SYNTH_DELIBERATION_SCHEMA, + }, +); +if (!synthesis) { + return failure('Synthesis returned nothing usable.', { round1Responders: responded.length }); +} + +// Phase 5 — Persist: audit only, a single writer merges lane profile updates. +phase('Persist'); +let persist = null; +if (mode === 'audit') { + const updates = collectProfileUpdates(responded); + if (updates.length) { + persist = await agent(persistPrompt(updates), { + label: 'persist-profile', + phase: 'Persist', + effort: 'low', + schema: PERSIST_SCHEMA, + }); + } else { + log('No profile updates returned by the lanes — nothing to persist.'); + } +} else { + log('Persist skipped — deliberation writes nothing.'); +} + +const markdown = + mode === 'audit' + ? renderAudit(focus, notConvened, silentRound1, synthesis, persist) + : renderDeliberation(topic, convened, notConvened, synthesis); + +return { + ok: true, + mode, + topic: mode === 'deliberation' ? topic : undefined, + focus: mode === 'audit' ? focus || 'whole repository' : undefined, + convened: convened.map((c) => c.name), + notConvened, + silentRound1, + responders: responded.map((r) => r.member), + synthesis, + persist, + markdown, +}; diff --git a/scripts/council-workflow-lint.ts b/scripts/council-workflow-lint.ts new file mode 100644 index 000000000..e8df972da --- /dev/null +++ b/scripts/council-workflow-lint.ts @@ -0,0 +1,272 @@ +#!/usr/bin/env bun +/** + * council-workflow-lint: structural lint for the /council workflow template. + * + * Checks, in order: + * (a) the template exists and parses as the workflow RUNTIME shape (NOT module-legal + * ESM): `export const meta` is extracted statically, the remaining body carries no + * other export (never `export default`), and it transpiles as an async function + * body — top-level await/return are the contract, so an ESM parse is the wrong check + * (b) meta.name === 'council' AND the __GENIE_LENS_ROOT__ placeholder is + * present (an unstamped template must never ship pre-stamped) + * (c) zero banned runtime APIs (the workflow determinism + self-contained + * contract: no Date.now / Math.random / new Date( / require( / import / + * process. / fs.) + * (d) every routing member, the audit roster, and the default trio are keys + * of LENSES + * (e) every LENSES path resolves on disk relative to plugins/genie/ (all 13) + * (f) every references/lenses/*.md card has name/modes/voice frontmatter + * + * NOTE: check (e) intentionally fails for the seven lane skills until Group 1 + * lands them under skills/ (reached via the plugins/genie/skills symlink). The + * six deliberation cards resolve as soon as this group creates them. + * + * Exit 0 when every check passes, 1 otherwise. Not wired into `bun run check` + * here — Group 5 owns that. + */ + +import { existsSync, readFileSync, readdirSync } from 'node:fs'; +import { join } from 'node:path'; + +const REPO_ROOT = new URL('..', import.meta.url).pathname.replace(/\/$/, ''); +const PLUGIN_DIR = join(REPO_ROOT, 'plugins', 'genie'); +const TEMPLATE = join(PLUGIN_DIR, 'workflows', 'council.js'); +const CARDS_DIR = join(PLUGIN_DIR, 'references', 'lenses'); +const PLACEHOLDER = '__GENIE_LENS_ROOT__'; + +const BANNED: Array<[label: string, pattern: RegExp]> = [ + ['Date.now', /Date\.now/], + ['Math.random', /Math\.random/], + ['new Date(', /new Date\(/], + ['require(', /require\(/], + ['top-level import', /^import /m], + ['process.', /process\./], + ['fs.', /[^a-zA-Z.]fs\./], +]; + +interface CheckResult { + name: string; + ok: boolean; + detail: string; + messages: string[]; +} + +function pass(name: string, detail: string): CheckResult { + return { name, ok: true, detail, messages: [] }; +} + +function fail(name: string, detail: string, messages: string[] = []): CheckResult { + return { name, ok: false, detail, messages }; +} + +function dedupe(values: string[]): string[] { + return [...new Set(values)]; +} + +/** Balanced-delimiter slice of a `const NAME = ;` declaration. */ +function sliceLiteral(src: string, marker: string, open: string, close: string): string { + const at = src.indexOf(marker); + if (at < 0) throw new Error(`declaration not found: ${marker}`); + const start = src.indexOf(open, at); + if (start < 0) throw new Error(`opening "${open}" not found after ${marker}`); + let depth = 0; + for (let i = start; i < src.length; i++) { + const ch = src[i]; + if (ch === open) depth += 1; + else if (ch === close) { + depth -= 1; + if (depth === 0) return src.slice(start, i + 1); + } + } + throw new Error(`unbalanced "${open}${close}" after ${marker}`); +} + +/** Extract `name -> path` pairs from the LENSES object literal (string values only). */ +function parseLensPairs(block: string): Array<[string, string]> { + const re = /(['"]?)([A-Za-z][\w-]*)\1\s*:\s*'([^']+)'/g; + return [...block.matchAll(re)].map((m) => [m[2], m[3]] as [string, string]); +} + +function quotedStrings(block: string): string[] { + return [...block.matchAll(/'([^']+)'/g)].map((m) => m[1]); +} + +/** Extract every member name from `members: [...]` arrays inside the ROUTING literal. */ +function routingMembers(block: string): string[] { + return [...block.matchAll(/members:\s*\[([^\]]*)\]/g)].flatMap((m) => quotedStrings(m[1])); +} + +/** + * Remove the `export const meta = { ... };` statement, returning the remaining source. + * The runtime extracts meta statically; everything else is the async function body. + */ +function stripMetaExport(src: string): string { + const at = src.indexOf('export const meta'); + if (at < 0) throw new Error('missing `export const meta` declaration'); + const literal = sliceLiteral(src, 'export const meta', '{', '}'); + let end = src.indexOf(literal, at) + literal.length; + while (end < src.length && /\s/.test(src[end])) end += 1; + if (src[end] === ';') end += 1; + return src.slice(0, at) + src.slice(end); +} + +/** + * Parse check against the RUNTIME shape, not module-legal ESM. The dynamic-workflow + * runtime runs the script as an async function body (top-level await/return are the + * contract) after extracting `export const meta` statically. So: strip meta, forbid any + * other export (especially `export default`, which the runtime never calls), then wrap + * the remainder as an async body and transpile — a syntax error fails with its message. + */ +function checkParse(file: string): CheckResult { + if (!existsSync(file)) return fail('parse', `template missing: ${file}`); + const src = readFileSync(file, 'utf8'); + let body: string; + try { + body = stripMetaExport(src); + } catch (err) { + return fail('parse', 'could not isolate the `export const meta` statement', [String(err)]); + } + if (/\bexport\s+default\b/.test(body)) { + return fail('parse', '`export default` is not an honored workflow entrypoint', [ + 'the runtime runs the script as an async function body; use top-level statements + `return`, not `export default`', + ]); + } + if (/^\s*export\b/m.test(body)) { + return fail('parse', 'workflow body has an export other than `export const meta`', [ + '`export const meta` is the only allowed export; the rest runs as a function body', + ]); + } + try { + new Bun.Transpiler({ loader: 'js' }).transformSync(`(async () => {\n${body}\n})`); + } catch (err) { + return fail('parse', 'workflow body does not parse as an async function body', [ + err instanceof Error ? err.message : String(err), + ]); + } + return pass('parse', 'template parses as a workflow async-body'); +} + +function checkMeta(src: string): CheckResult { + const problems: string[] = []; + if (!/name:\s*'council'/.test(src)) problems.push("meta.name is not 'council'"); + if (!src.includes(PLACEHOLDER)) problems.push(`missing ${PLACEHOLDER} placeholder (template must ship unstamped)`); + return problems.length + ? fail('meta', 'meta/placeholder invariant broken', problems) + : pass('meta', "meta.name === 'council' and placeholder present"); +} + +function checkBanned(src: string): CheckResult { + const hits: string[] = []; + for (const [label, pattern] of BANNED) { + if (pattern.test(src)) hits.push(`banned API present: ${label}`); + } + return hits.length + ? fail('banned-apis', 'template uses banned runtime APIs', hits) + : pass('banned-apis', 'zero banned runtime APIs'); +} + +function checkIntegrity(src: string): { integrity: CheckResult; pairs: Array<[string, string]> } { + let pairs: Array<[string, string]>; + let members: string[]; + let auditRoster: string[]; + let defaultTrio: string[]; + try { + pairs = parseLensPairs(sliceLiteral(src, 'const LENSES =', '{', '}')); + members = routingMembers(sliceLiteral(src, 'const ROUTING =', '[', ']')); + auditRoster = quotedStrings(sliceLiteral(src, 'const AUDIT_ROSTER =', '[', ']')); + defaultTrio = quotedStrings(sliceLiteral(src, 'const DEFAULT_TRIO =', '[', ']')); + } catch (err) { + return { integrity: fail('integrity', 'could not extract LENSES/ROUTING literals', [String(err)]), pairs: [] }; + } + const keySet = new Set(pairs.map(([name]) => name)); + if (!pairs.length) return { integrity: fail('integrity', 'LENSES literal parsed to zero entries'), pairs }; + const referenced = dedupe([...members, ...auditRoster, ...defaultTrio]); + const unknown = referenced.filter((name) => !keySet.has(name)); + const integrity = unknown.length + ? fail( + 'integrity', + 'routing/roster references lenses missing from LENSES', + unknown.map((u) => `unknown lens: ${u}`), + ) + : pass('integrity', `all ${referenced.length} referenced lenses are LENSES keys (${pairs.length} lenses total)`); + return { integrity, pairs }; +} + +function checkLensFiles(pairs: Array<[string, string]>): CheckResult { + if (!pairs.length) return fail('lens-files', 'no LENSES entries to resolve'); + const missing: string[] = []; + for (const [name, rel] of pairs) { + if (!existsSync(join(PLUGIN_DIR, rel))) missing.push(`${name} -> ${rel}`); + } + return missing.length + ? fail('lens-files', `${missing.length}/${pairs.length} lens paths do not resolve under plugins/genie/`, missing) + : pass('lens-files', `all ${pairs.length} lens paths resolve on disk`); +} + +function checkCards(): CheckResult { + if (!existsSync(CARDS_DIR)) return fail('lens-cards', `cards directory missing: ${CARDS_DIR}`); + const files = readdirSync(CARDS_DIR).filter((f) => f.endsWith('.md')); + if (!files.length) return fail('lens-cards', `no lens cards in ${CARDS_DIR}`); + const problems: string[] = []; + for (const file of files) { + const body = readFileSync(join(CARDS_DIR, file), 'utf8'); + for (const key of ['name', 'modes', 'voice']) { + if (!new RegExp(`^${key}: `, 'm').test(body)) problems.push(`${file}: missing "${key}:" frontmatter`); + } + } + return problems.length + ? fail('lens-cards', 'lens card frontmatter incomplete', problems) + : pass('lens-cards', `${files.length} lens cards carry name/modes/voice`); +} + +async function main(): Promise { + const parseOnlyIdx = process.argv.indexOf('--parse-only'); + if (parseOnlyIdx !== -1) { + const file = process.argv[parseOnlyIdx + 1]; + if (!file) { + console.error('council-workflow-lint: --parse-only requires a file path'); + process.exit(2); + } + const result = checkParse(file); + report([result]); + process.exit(result.ok ? 0 : 1); + } + + const results: CheckResult[] = []; + + results.push(checkParse(TEMPLATE)); + + if (!existsSync(TEMPLATE)) { + report(results); + process.exit(1); + } + + const src = readFileSync(TEMPLATE, 'utf8'); + results.push(checkMeta(src)); + results.push(checkBanned(src)); + + const { integrity, pairs } = checkIntegrity(src); + results.push(integrity); + results.push(checkLensFiles(pairs)); + results.push(checkCards()); + + report(results); + process.exit(results.every((r) => r.ok) ? 0 : 1); +} + +function report(results: CheckResult[]): void { + console.log('council-workflow-lint'); + console.log('====================='); + for (const r of results) { + console.log(`${r.ok ? 'PASS' : 'FAIL'} ${r.name.padEnd(13)} ${r.detail}`); + for (const m of r.messages) console.log(` - ${m}`); + } + console.log(''); + const failed = results.filter((r) => !r.ok); + console.log(failed.length ? `FAIL: ${failed.length} check(s) failed.` : 'OK: all checks passed.'); +} + +main().catch((err) => { + console.error(`council-workflow-lint: ${err instanceof Error ? err.message : String(err)}`); + process.exit(2); +}); diff --git a/src/lib/council-workflow-stamp.test.ts b/src/lib/council-workflow-stamp.test.ts new file mode 100644 index 000000000..368adc8bf --- /dev/null +++ b/src/lib/council-workflow-stamp.test.ts @@ -0,0 +1,102 @@ +/** + * Tests for council-stamp.cjs: the install-time stamp that writes the /council + * workflow template into ~/.claude/workflows with LENS_ROOT resolved. + * + * The implementation is CommonJS (it is required from the ESM SessionStart + * hook), so we load it through createRequire. Everything runs inside a tmpdir; + * afterEach removes it, so no global state is touched. + * + * Run with: bun test src/lib/council-workflow-stamp.test.ts + */ + +import { afterEach, beforeEach, describe, expect, test } from 'bun:test'; +import { existsSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; +import { createRequire } from 'node:module'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; + +const require = createRequire(import.meta.url); +const { stampCouncilWorkflow, PLACEHOLDER } = require('../../plugins/genie/scripts/council-stamp.cjs') as { + stampCouncilWorkflow: (opts: { templatePath: string; pluginRoot: string; targetDir: string }) => { + action: 'written' | 'skipped'; + targetPath: string; + }; + PLACEHOLDER: string; +}; + +const TEMPLATE_BODY = [ + "export const meta = { name: 'council' };", + `const LENS_ROOT = '${PLACEHOLDER}';`, + 'export default async function council() {', + ' log(LENS_ROOT);', + '}', + '', +].join('\n'); + +describe('stampCouncilWorkflow', () => { + let dir: string; + let templatePath: string; + let targetDir: string; + + beforeEach(() => { + dir = mkdtempSync(join(tmpdir(), 'council-stamp-')); + templatePath = join(dir, 'council.template.js'); + targetDir = join(dir, 'nested', 'workflows'); // nested so we also prove mkdir recursive + writeFileSync(templatePath, TEMPLATE_BODY, 'utf8'); + }); + + afterEach(() => { + rmSync(dir, { recursive: true, force: true }); + }); + + test('replaces the placeholder with the absolute plugin root and writes council.js', () => { + const pluginRoot = '/opt/plugins/genie'; + const res = stampCouncilWorkflow({ templatePath, pluginRoot, targetDir }); + + expect(res.action).toBe('written'); + expect(res.targetPath).toBe(join(targetDir, 'council.js')); + + const out = readFileSync(res.targetPath, 'utf8'); + expect(out).toContain(`const LENS_ROOT = '${pluginRoot}';`); + expect(out).not.toContain(PLACEHOLDER); + }); + + test('target lands exactly at /council.js (creating parent dirs)', () => { + expect(existsSync(targetDir)).toBe(false); + const res = stampCouncilWorkflow({ templatePath, pluginRoot: '/abs/plugins/genie', targetDir }); + + expect(res.targetPath).toBe(join(targetDir, 'council.js')); + expect(existsSync(join(targetDir, 'council.js'))).toBe(true); + }); + + test('idempotent — an unchanged re-run skips the write', () => { + const pluginRoot = '/abs/plugins/genie'; + expect(stampCouncilWorkflow({ templatePath, pluginRoot, targetDir }).action).toBe('written'); + expect(stampCouncilWorkflow({ templatePath, pluginRoot, targetDir }).action).toBe('skipped'); + }); + + test('rewrites when the plugin root changes (update-safe re-stamp)', () => { + stampCouncilWorkflow({ templatePath, pluginRoot: '/root/one/plugins/genie', targetDir }); + const changed = stampCouncilWorkflow({ templatePath, pluginRoot: '/root/two/plugins/genie', targetDir }); + + expect(changed.action).toBe('written'); + const out = readFileSync(join(targetDir, 'council.js'), 'utf8'); + expect(out).toContain('/root/two/plugins/genie'); + expect(out).not.toContain('/root/one/plugins/genie'); + }); + + test('rewrites when the template content changes (self-healing on plugin update)', () => { + const pluginRoot = '/abs/plugins/genie'; + stampCouncilWorkflow({ templatePath, pluginRoot, targetDir }); + writeFileSync(templatePath, `${TEMPLATE_BODY}\n// updated template\n`, 'utf8'); + + const res = stampCouncilWorkflow({ templatePath, pluginRoot, targetDir }); + expect(res.action).toBe('written'); + expect(readFileSync(res.targetPath, 'utf8')).toContain('// updated template'); + }); + + test('throws when a required path argument is missing', () => { + // @ts-expect-error — intentionally omitting required args to prove the guard fires + expect(() => stampCouncilWorkflow({ templatePath, pluginRoot: '/x' })).toThrow(/requires/); + }); +}); From 22a7ed50e3884e57d84ddf9335ea8a724868733c Mon Sep 17 00:00:00 2001 From: namastex888 Date: Thu, 9 Jul 2026 22:06:07 -0300 Subject: [PATCH 4/8] =?UTF-8?q?refactor(skills):=20retire=20council=20skil?= =?UTF-8?q?l=20=E2=80=94=20/council=20is=20the=20native=20workflow=20now?= =?UTF-8?q?=20(council-workflow=20G3)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .genie/wishes/council-workflow/WISH.md | 23 +++--- plugins/genie/workflows/council.js | 2 +- skills/README.md | 2 +- skills/council/SKILL.md | 107 ------------------------- skills/council/members/config.md | 24 ------ skills/council/members/routing.md | 32 -------- skills/council/templates/report.md | 71 ---------------- skills/genie/SKILL.md | 2 +- skills/genie/reference/lifecycle.md | 2 +- 9 files changed, 17 insertions(+), 248 deletions(-) delete mode 100644 skills/council/SKILL.md delete mode 100644 skills/council/members/config.md delete mode 100644 skills/council/members/routing.md delete mode 100644 skills/council/templates/report.md diff --git a/.genie/wishes/council-workflow/WISH.md b/.genie/wishes/council-workflow/WISH.md index 31df077fd..774ba5152 100644 --- a/.genie/wishes/council-workflow/WISH.md +++ b/.genie/wishes/council-workflow/WISH.md @@ -59,12 +59,13 @@ Genie's multi-perspective reasoning is split across two model-driven orchestrato ## Execution Strategy -| Wave | Groups | Notes | -|------|--------|-------| -| 1 | G1 (lane skills) ∥ G2 (engine + lenses + stamp) | New files only — zero collision with the pending skills-fable5-revamp execution review | -| 2 | G3 (cutover) | Gated on G1+G2 AND skills-fable5-revamp MERGED to its base (review-closed insufficient); rebase `wish/council-workflow` onto the post-merge base first | -| 3 | G4 (consumers) | Edits review/brainstorm skills — behind the same fable5 merged-gate | -| 4 | G5 (lints in check + docs + live QA) | Final gate | +| Wave | Group | Agent | Complexity | Model | Notes | +|------|-------|-------|------------|-------|-------| +| 1 | G1 lane skills | engineer | 2 (md migration) | inherit (fable·max) | New dirs only — zero collision with skills-fable5-revamp | +| 1 | G2 engine + lenses + stamp | engineer | 4 (workflow engine + install wiring) | inherit (fable·max) | New files only; runs parallel to G1 | +| 2 | G3 cutover | engineer | 2 (deletion + purge) | inherit (fable·max) | Gated on G1+G2 AND skills-fable5-revamp MERGED to base — satisfied: PR #2518 (1308e4c6) is an ancestor of this branch | +| 3 | G4 consumers | engineer | 3 (prompt surfaces) | inherit (fable·max) | Edits review/brainstorm skills — behind the same fable5 merged-gate | +| 4 | G5 lints in check + docs + live QA | engineer | 2 (wiring + docs + QA evidence) | inherit (fable·max) | Final gate | --- @@ -124,10 +125,12 @@ bash .genie/wishes/council-workflow/validate/g2-engine.sh 3. skills/README.md council row updated to point at the workflow. **Acceptance Criteria:** -- [ ] `skills/council/` gone; no references to `members/routing.md`, `members/config.md`, or the council skill remain in live surfaces -- [ ] `bun run lint:council-workflow` green — first point where full 13-lens on-disk integrity is assertable (G1+G2 both done) -- [ ] `git grep -il 'specialist-panel'` → 0 hits outside `.genie/attic/`, `CHANGELOG.md`, this wish's artifacts -- [ ] **Process gate:** skills-fable5-revamp MERGED to its base branch (review-closed is not enough — its Wave 1 rewrote these same skill files in place), and `wish/council-workflow` rebased onto the post-merge base before this group starts (checked by the worker, recorded in the group log) +- [x] `skills/council/` gone; no references to `members/routing.md`, `members/config.md`, or the council skill remain in live surfaces +- [x] `bun run lint:council-workflow` green — first point where full 13-lens on-disk integrity is assertable (G1+G2 both done) +- [x] `git grep -il 'specialist-panel'` → 0 hits outside `.genie/attic/`, `CHANGELOG.md`, this wish's artifacts +- [x] **Process gate:** skills-fable5-revamp MERGED to its base branch — satisfied by ancestry: PR #2518 merge `1308e4c6` is an ancestor of this branch (no rebase needed; the edited/deleted skill files were the post-fable5 versions) + +**Status:** DONE (2026-07-10) — engineer died mid-report (API drop) with the work complete in-tree; orchestrator verified diffs + gate, independent review SHIP (0 gaps ≥MEDIUM; 2 LOW notes trace to G2's accepted design: config.md model-defaults have no workflow analog beyond inherit, deliberation report is distilled rather than per-member — dissent preserved). Extra: router row in skills/genie/SKILL.md now launches the saved workflow via the Workflow tool for council; one comment in council.js reworded to satisfy the purge grep. `bun run check` green (725 pass / 1 skip). **Validation:** ```bash diff --git a/plugins/genie/workflows/council.js b/plugins/genie/workflows/council.js index 553b086a0..db5741432 100644 --- a/plugins/genie/workflows/council.js +++ b/plugins/genie/workflows/council.js @@ -40,7 +40,7 @@ const LENSES = { tracer: 'references/lenses/tracer.md', }; -// Keyword routing for deliberation, absorbed from the old council members/routing.md. +// Keyword routing for deliberation, absorbed from the old council's routing table. // The four retired council lenses are remapped onto lanes: benchmarker -> perf, // sentinel -> supply-chain, ergonomist -> dx-docs, architect -> architecture. const ROUTING = [ diff --git a/skills/README.md b/skills/README.md index dd5ee2c56..1b8ca185e 100644 --- a/skills/README.md +++ b/skills/README.md @@ -25,7 +25,7 @@ Decision legend: | `refine` | Keep — portable now | Prompt-optimizer transform, pure text in/out. No runtime state; carries over unchanged. | | `fix` | Keep — portable now | Dispatches a fixer subagent for FIX-FIRST gaps. Dispatch re-points to the Agent tool, but note: it also calls `genie task comment`/`genie task block`, which have no v5 equivalent — its port needs the same drop/reshape decision `review` made (report in output vs mutate task rows), not a pure re-point. | | `trace` | Keep — portable now | Dispatches a trace subagent to find root cause for `/fix`. Same Agent-tool dispatch port; investigation logic is runtime-agnostic. | -| `council` | Keep — portable now | Convenes multiple AI agents for deliberation — a natural fit for native teams (Agent tool spawns members, SendMessage for cross-talk). Dispatch is the only thing to re-plumb. | +| `council` | Keep — portable now | Ported — no longer a skill dir. `/council` now ships as a native dynamic workflow (`plugins/genie/workflows/council.js`, stamped into `~/.claude/workflows/` at session start): two modes — deliberation + audit — over one lens library (the 7 lane skills + `plugins/genie/references/lenses/`). The native-team dispatch it needed became the workflow engine itself. | | `docs` | Keep — portable now | Dispatches a docs subagent to audit/generate docs against the codebase. Agent-tool dispatch port; no intrinsic v4 state. | | `genie-hacks` | Keep — portable now | Browse/search/contribute community hacks — a reference/content skill with no runtime dependency. | | `report` | Keep — port deferred | Bug-investigation cascade (`/trace` → browser evidence → observability → GitHub issue). The trace and issue-filing paths are portable; the observability pull currently reads v4 OTel/PG event data and must be re-sourced when that data path is ported. | diff --git a/skills/council/SKILL.md b/skills/council/SKILL.md deleted file mode 100644 index 5e9be691a..000000000 --- a/skills/council/SKILL.md +++ /dev/null @@ -1,107 +0,0 @@ ---- -name: council -description: "Convene real AI agents for multi-perspective deliberation on architecture, design, and strategy decisions." -argument-hint: "[topic or question]" -effort: high ---- - -# /council — Multi-Agent Deliberation - -Convene 3-4 real subagents, run a 2-round Socratic deliberation, and synthesize a structured report. No voting, no simulation — every perspective comes from a real spawned agent, and you, the orchestrator, make every judgment call in real time. - -## Topic - -``` -$ARGUMENTS -``` - -If empty, ask the user for the topic before proceeding. If `--members a,b,c` appears in the arguments, use exactly those members instead of routing. - -## Council Members - -| Member | Focus | Lens | -|--------|-------|------| -| **questioner** | Challenge assumptions | "Why? Is there a simpler way?" | -| **benchmarker** | Performance evidence | "Show me the benchmarks." | -| **simplifier** | Complexity reduction | "Delete code. Ship features." | -| **sentinel** | Security oversight | "Where are the secrets? What's the blast radius?" | -| **ergonomist** | Developer experience | "If you need to read the docs, the API failed." | -| **architect** | Systems thinking | "Talk is cheap. Show me the code." | -| **operator** | Operations reality | "No one wants to run your code." | -| **deployer** | Zero-config deployment | "Zero-config with infinite scale." | -| **measurer** | Observability | "Measure, don't guess." | -| **tracer** | Production debugging | "You will debug this in production." | - -## Smart Routing - -Classify the topic and select 3-4 members: - -| Topic Keywords | Members | -|---------------|---------| -| architecture, design, system, interface, API | questioner, architect, simplifier, benchmarker | -| performance, latency, throughput, scale | benchmarker, questioner, architect, measurer | -| security, auth, secrets, blast radius | questioner, sentinel, simplifier | -| API, endpoint, DX, developer, SDK | questioner, simplifier, ergonomist, deployer | -| ops, deploy, infra, CI/CD, monitoring | operator, deployer, tracer, measurer | -| debug, trace, observability, logging | tracer, measurer, benchmarker | -| plan, scope, wish, feature | questioner, simplifier, architect, ergonomist | - -**Default (no keyword match):** questioner, simplifier, architect. - -Rationale: `members/routing.md`. Per-member model defaults: `members/config.md`. - -## Deliberation - -### Round 1 — Initial Perspectives - -Spawn each selected member as a subagent via the **Agent tool** — all spawns in ONE message so they deliberate in parallel (background; each notifies you with its final message). Per-member brief: - -> You are the council's **** — . Council topic: ****. -> -> Apply your specialist lens. Return your perspective as your final message: substantive (2-4 paragraphs), opinionated, grounded in your expertise. Take positions; cite evidence or name the assumption you are making. - -A member that fails to spawn or returns nothing usable: note it and continue, as long as at least 2 members delivered. Fewer than 2 → report failure to the user and stop. - -### Round 2 — Socratic Response - -For each member that responded, send the other members' Round 1 perspectives via **SendMessage** (this continues the member's session with its context intact): - -> ROUND 2 — the other members' perspectives are below. Reply with: -> 1. The strongest point another member made. -> 2. At least one point you challenge or refine. -> 3. Whether your initial position changed, and why. -> -> - -A member that does not answer Round 2: record "no Round 2 response" and synthesize from what exists. - -### Synthesis - -Your core intellectual contribution. From the collected responses identify: points of consensus, key tensions and unresolved disagreements, evolution of thinking between rounds, and minority perspectives worth preserving. - -Write the report per `templates/report.md`: Executive Summary, Council Composition, Situation Analysis (one subsection per responding member — Round 1 and Round 2, never merged), Key Findings, Recommendations (P0/P1/P2 with rationale and risk), Next Steps, Dissent (quoted faithfully; if none, note the convergence). - -## Failure Handling - -| Situation | Action | -|-----------|--------| -| Fewer than 2 members deliver Round 1 | Stop; report failure, suggest retry | -| Member silent or errored | Note "no response" in the report, proceed with responders | -| Round 2 SendMessage fails | Retry once; then synthesize from Round 1 alone for that member | - -## Constraints - -- **Advisory only** — the council advises, the user decides. Never block progress on council output. -- **No voting** — no verdicts or gate-keeping language. The council thinks; `/review` judges. -- **3-4 members max** — never spawn all 10 unless explicitly requested. -- **Distinct perspectives** — each member applies their unique lens; no echoing, no rubber-stamping. -- **Preserve dissent** — minority views go in the Dissent section, never suppressed or editorialized. -- **Real agents only** — never simulate or write a member's response yourself. - -## Supporting Files - -| File | Purpose | -|------|---------| -| `${CLAUDE_SKILL_DIR}/members/routing.md` | Smart routing with rationale | -| `${CLAUDE_SKILL_DIR}/members/config.md` | Per-member model defaults and overrides | -| `${CLAUDE_SKILL_DIR}/templates/report.md` | Full report template | diff --git a/skills/council/members/config.md b/skills/council/members/config.md deleted file mode 100644 index 937e49094..000000000 --- a/skills/council/members/config.md +++ /dev/null @@ -1,24 +0,0 @@ -# Council Member Model Configuration - -Members are spawned as Agent-tool subagents; the `model` parameter on each member's Agent call selects its model. Omitting the parameter means **inherit** — the member runs on the orchestrator's model. - -## Member Defaults - -| Member | Default model | Notes | -|--------|---------------|-------| -| questioner | inherit | Challenges need strong reasoning | -| architect | inherit | Systems thinking needs depth | -| simplifier | inherit | Deletion requires confidence | -| benchmarker | inherit | Evidence analysis | -| sentinel | inherit | Security requires precision | -| ergonomist | inherit | DX judgment | -| operator | inherit | Ops reality | -| deployer | inherit | Deploy patterns | -| measurer | inherit | Observability | -| tracer | inherit | Debug depth | - -## Overrides - -- Set `model` on that member's Agent call (e.g. `opus`, `sonnet`, `haiku`) — per member, per session. -- Mixed-model councils are just per-member overrides: architect on `opus`, benchmarker on `haiku`, rest inherited. -- `haiku` suits fast, cheap councils; keep the questioner on a stronger model — assumption-challenging degrades first. diff --git a/skills/council/members/routing.md b/skills/council/members/routing.md deleted file mode 100644 index adec2ebf5..000000000 --- a/skills/council/members/routing.md +++ /dev/null @@ -1,32 +0,0 @@ -# Council Member Routing - -Smart routing configuration for the `/council` skill. The orchestrator classifies the topic and selects 3-4 relevant members from this table. Users never need to pick members manually. - -## Topic Routing - -| Topic Keywords | Members | Rationale | -|---------------|---------|-----------| -| architecture, design, system, interface, API | questioner, architect, simplifier, benchmarker | Core design decisions need assumption-challenging, systems thinking, complexity reduction, and performance grounding | -| performance, latency, throughput, scale | benchmarker, questioner, architect, measurer | Evidence-based performance analysis needs benchmarks, skepticism, architectural context, and measurement rigor | -| security, auth, secrets, blast radius | questioner, sentinel, simplifier | Security-first review needs assumption-challenging, breach expertise, and complexity reduction to minimize attack surface | -| API, endpoint, DX, developer, SDK | questioner, simplifier, ergonomist, deployer | Developer experience needs skepticism, minimalism, usability focus, and deployment-awareness | -| ops, deploy, infra, CI/CD, monitoring | operator, deployer, tracer, measurer | Operational reality needs production experience, deployment expertise, debugging capability, and observability | -| debug, trace, observability, logging | tracer, measurer, benchmarker | Production insight needs high-cardinality debugging, measurement methodology, and performance context | -| plan, scope, wish, feature | questioner, simplifier, architect, ergonomist | Planning cognition needs assumption-challenging, complexity reduction, architectural foresight, and DX awareness | - -## Default (no keyword match) - -questioner, simplifier, architect - -**Rationale:** The core trio covers the most common failure modes: solving the wrong problem (questioner), over-engineering (simplifier), and short-term thinking (architect). - -## Override - -Users can bypass routing with `--members questioner,architect` to force specific members. This is a power-user escape hatch, not the normal path. - -## Notes - -- Never spawn all 10 unless explicitly requested — compute cost is linear in member count -- 3-4 members per topic is the sweet spot: enough diversity, manageable deliberation time -- The questioner appears in most routes because challenging assumptions has universal value -- Topics may match multiple rows — use the best match, not all matches diff --git a/skills/council/templates/report.md b/skills/council/templates/report.md deleted file mode 100644 index 319b2d14a..000000000 --- a/skills/council/templates/report.md +++ /dev/null @@ -1,71 +0,0 @@ -# Council Report: - -## Executive Summary - -<2-3 sentences: the question that was deliberated, the emerging consensus (or key tension if no consensus), and the single most important insight from the deliberation.> - -## Council Composition - -| Member | Lens | Model | -|--------|------|-------| -| questioner | Challenge assumptions | opus | -| architect | Systems thinking | sonnet | -| simplifier | Complexity reduction | haiku | - -## Situation Analysis - -### questioner - -**Initial perspective (Round 1):** - - -**After deliberation (Round 2):** - - -### architect - -**Initial perspective (Round 1):** - - -**After deliberation (Round 2):** - - -### simplifier - -**Initial perspective (Round 1):** - - -**After deliberation (Round 2):** - - - - -## Key Findings - -1. **** — -2. **** — -3. **** — - -## Recommendations - -| Priority | Recommendation | Rationale | Risk if Ignored | -|----------|---------------|-----------|-----------------| -| P0 | | | | -| P1 | | | | -| P2 | | | | - -## Next Steps - -- [ ] -- [ ] -- [ ] - -## Dissent - - - - - ---- - -*Council session: | Members: | Round 1: / | Round 2: /* diff --git a/skills/genie/SKILL.md b/skills/genie/SKILL.md index 4bc42dfbb..077b006c8 100644 --- a/skills/genie/SKILL.md +++ b/skills/genie/SKILL.md @@ -25,7 +25,7 @@ Classify `$ARGUMENTS` into exactly one category: | Category | Signal | Route | |----------|--------|-------| -| **explicit** | Names a skill: "brainstorm X", "wish X", "review X", "work X", "council X", "refine X", "fix X", "trace X", "docs X", "report X", "dream", "learn X", "pm", "wizard", "wire omni", "hacks" | Invoke the named skill via the Skill tool, rest as args | +| **explicit** | Names a skill: "brainstorm X", "wish X", "review X", "work X", "council X", "refine X", "fix X", "trace X", "docs X", "report X", "dream", "learn X", "pm", "wizard", "wire omni", "hacks" | Invoke the named skill via the Skill tool, rest as args — except **council**, which launches the saved workflow via the Workflow tool (workflow `council`, args mediated from the user's text) | | **concrete** | Clear feature/change: "add X", "implement Y", "build a..." | `/wish` | | **fuzzy** | Exploratory: "I'm not sure how to...", "what if we...", "how should I handle..." | `/brainstorm` | | **bug** | "X is broken", "error when...", "something's wrong with..." | `/report` | diff --git a/skills/genie/reference/lifecycle.md b/skills/genie/reference/lifecycle.md index ac6a87dcf..ae14b5a8c 100644 --- a/skills/genie/reference/lifecycle.md +++ b/skills/genie/reference/lifecycle.md @@ -16,7 +16,7 @@ Every piece of work follows this flow: | `/review` | Universal quality gate — SHIP / FIX-FIRST / BLOCKED with severity-tagged gaps | Before and after `/work`, or any plan/PR | | `/work` | Execute an approved wish — dispatch native-team subagents per group in waves, fix loops, validation | Wish is SHIP-approved | | `/fix` | Resolve FIX-FIRST gaps, re-review, escalate after 2 failed loops | Review returned FIX-FIRST | -| `/council` | Multi-perspective deliberation with specialist viewpoints | Major design decisions, tradeoffs | +| `/council` | Multi-perspective deliberation with specialist viewpoints (native workflow) | Major design decisions, tradeoffs | | `/refine` | Transform a brief into a production-ready prompt | Prompt needs sharpening | | `/report` | Investigate bugs — trace, capture evidence, open a GitHub issue | Bug reports | | `/trace` | Reproduce and isolate root cause without patching | Unknown issues needing investigation | From 2af7ebd9e118acbc6c537be8e8b166b1875ff54e Mon Sep 17 00:00:00 2001 From: namastex888 Date: Thu, 9 Jul 2026 22:18:35 -0300 Subject: [PATCH 5/8] feat(skills): lens panels in review + domain-experts in brainstorm (council-workflow G4) --- .genie/wishes/council-workflow/WISH.md | 10 ++++++---- skills/brainstorm/SKILL.md | 6 +++++- skills/review/SKILL.md | 15 ++++++++++++++- 3 files changed, 25 insertions(+), 6 deletions(-) diff --git a/.genie/wishes/council-workflow/WISH.md b/.genie/wishes/council-workflow/WISH.md index 774ba5152..9a4351370 100644 --- a/.genie/wishes/council-workflow/WISH.md +++ b/.genie/wishes/council-workflow/WISH.md @@ -148,10 +148,12 @@ bash .genie/wishes/council-workflow/validate/g3-cutover.sh 2. `skills/brainstorm/SKILL.md` — gains a Decisions-stuck domain-experts step that dispatches 2-3 lens subagents reading `references/lenses/` cards (step is NEW to the repo skill; pattern mirrors the global brainstorm skill's lens-subagent step). **Acceptance Criteria:** -- [ ] `skills/review/SKILL.md` has a lens-panel section with change-type → lens mapping (structural grep markers: `lens panel`, `change-type`) -- [ ] `skills/brainstorm/SKILL.md` has a domain-experts step (structural grep marker: `domain-expert`) -- [ ] Both skills reference lens-library paths (cards and lane skills) that exist on disk -- [ ] No duplicated lens definitions inline in either skill — no `voice:`/`modes:` frontmatter blocks outside the library (single source: the library) +- [x] `skills/review/SKILL.md` has a lens-panel section with change-type → lens mapping (structural grep markers: `lens panel`, `change-type`) +- [x] `skills/brainstorm/SKILL.md` has a domain-experts step (structural grep marker: `domain-expert`) +- [x] Both skills reference lens-library paths (cards and lane skills) that exist on disk +- [x] No duplicated lens definitions inline in either skill — no `voice:`/`modes:` frontmatter blocks outside the library (single source: the library) + +**Status:** DONE (2026-07-10) — gate `G4 PASS` (orchestrator-run), execution review SHIP (0 gaps ≥MEDIUM; LOW: intended pointer→section trigger redundancy in brainstorm; observation: mapping covers 5 lanes + questioner per spec, code-quality already in the base checklist). "council" fully purged from the review skill; brainstorm's single `/council` mention is the escalation to the surviving workflow. `bun run check` at baseline (725 pass / 1 skip). **Validation:** ```bash diff --git a/skills/brainstorm/SKILL.md b/skills/brainstorm/SKILL.md index 98970bce2..f0ca95e78 100644 --- a/skills/brainstorm/SKILL.md +++ b/skills/brainstorm/SKILL.md @@ -42,7 +42,11 @@ WRS: ██████░░░░ 60/100 Problem ✅ | Scope ✅ | Decisions ✅ | Risks ░ | Criteria ░ ``` -✅ = enough info to write that section of a wish; ░ = still needs discussion. Below 100: keep refining. At 100: auto-crystallize. If **Decisions** stays unfilled after 2+ exchanges, suggest `/council` for specialist perspectives on the tradeoffs. +✅ = enough info to write that section of a wish; ░ = still needs discussion. Below 100: keep refining. At 100: auto-crystallize. If **Decisions** won't fill, convene domain experts (see Stuck Decisions). + +## Stuck Decisions + +If **Decisions** stays unfilled after 2+ exchanges, convene **domain experts**: dispatch 2-3 lens subagents in parallel (Agent tool), each reading a distinct lens — a deliberation card from `plugins/genie/references/lenses/` (questioner, simplifier, operator, …) plus, when the tradeoff is technical, the matching lane skill at `skills//SKILL.md`. Present their perspectives to the user, then keep refining. Escalate to the full `/council` workflow when the decision deserves a durable deliberation record. ## Scope Size diff --git a/skills/review/SKILL.md b/skills/review/SKILL.md index bf9be9133..d5c49d9c6 100644 --- a/skills/review/SKILL.md +++ b/skills/review/SKILL.md @@ -101,7 +101,20 @@ When a failure's root cause is unclear, invoke `/trace` before dispatching `/fix ## Dispatch -**Reviewer ≠ engineer.** The orchestrator dispatches review as a separate subagent via the Agent tool — an agent never reviews its own work. Follow-ups to a running reviewer go through SendMessage. When a council team is active, findings may be shared with council members for advisory input; the verdict is still determined by the checklist. +**Reviewer ≠ engineer.** The orchestrator dispatches review as a separate subagent via the Agent tool — an agent never reviews its own work. Follow-ups to a running reviewer go through SendMessage. For change-types that warrant deeper scrutiny, the orchestrator also convenes a **Lens Panel** (below); those lenses advise, but the checklist still owns the verdict. + +## Lens Panels + +When the change-type warrants it, the orchestrator dispatches **lens reviewers** alongside the standard reviewer — each a separate subagent whose prompt carries its lens file (path + content) and the curated review scope. Convene a lens only when the change actually touches its surface; lenses advise, but the verdict still comes from the checklist above — never from a lens. + +| Change-type | Advisory lens | +|-------------|---------------| +| Auth / secrets / dependency changes | `skills/supply-chain/SKILL.md` | +| Hot-path or latency-sensitive code | `skills/perf/SKILL.md` | +| Public API / CLI surface | `skills/dx-docs/SKILL.md` | +| Module-boundary / architecture moves | `skills/architecture/SKILL.md` | +| Test-strategy changes | `skills/qa/SKILL.md` | +| Plan / wish reviews | `plugins/genie/references/lenses/questioner.md` | ## Verdict Reporting From 6494253dd8c9b83e1b4d2e1244e7618ac2f216d3 Mon Sep 17 00:00:00 2001 From: namastex888 Date: Thu, 9 Jul 2026 23:39:47 -0300 Subject: [PATCH 6/8] feat(gate): wire council lint into check + plugin /council docs (council-workflow G5) --- .genie/wishes/council-workflow/WISH.md | 8 +++++--- package.json | 2 +- plugins/genie/README.md | 13 +++++++++++++ 3 files changed, 19 insertions(+), 4 deletions(-) diff --git a/.genie/wishes/council-workflow/WISH.md b/.genie/wishes/council-workflow/WISH.md index 9a4351370..7a88be3cd 100644 --- a/.genie/wishes/council-workflow/WISH.md +++ b/.genie/wishes/council-workflow/WISH.md @@ -169,11 +169,13 @@ bash .genie/wishes/council-workflow/validate/g4-consumers.sh **Deliverables:** 1. `lint:council-workflow` wired into `bun run check` (package.json). 2. Docs: skills/README.md, plugin README/docs notes, workflow requirements (CC ≥ 2.1.154, paid plans, `disableWorkflows`). -3. Live QA evidence: `.genie/wishes/council-workflow/qa/deliberation-run.md` + `qa/audit-run.md` (real runs on the genie repo, with `/workflows` token totals captured). +3. Live QA evidence: `.genie/wishes/council-workflow/qa/deliberation-run.md` + `qa/audit-run.md` — **USER-GATED, post-release (Felipe's ruling 2026-07-10):** the real test is the shipped surface, not a scriptPath simulation. Ritual: merge → release → plugin update lands on Felipe's machine → SessionStart stamp installs `~/.claude/workflows/council.js` → Felipe runs `/council` himself ("revisar tudo") to validate; the run outputs become the qa/ evidence files. `validate/g5-gate.sh` deliberately keeps failing at the qa/ assertions until then — that pending tail is the designed state, not a defect. **Acceptance Criteria:** -- [ ] `bun run check` green AND its output proves `lint:council-workflow` actually ran (behavioral wiring check, not just script existence) -- [ ] Both QA evidence files present with real run output +- [x] `bun run check` green AND its output proves `lint:council-workflow` actually ran (behavioral wiring check, not just script existence) +- [ ] Both QA evidence files present with real run output — post-release, from Felipe's own `/council` runs + +**Status:** PARTIAL (2026-07-10) — engineering half DONE by the orchestrator inline (Felipe stopped agent dispatch for this wave): `lint:council-workflow` wired into `check` after `wishes:lint` (behavioral proof: check output shows it running, exit 0, 725 pass / 1 skip), plugin README gained the `/council` workflow section (ships/distribution/modes/requirements/override) + `workflows/` in the tree. Live-QA half USER-GATED post-release per Felipe's ruling — `validate/g5-gate.sh` correctly halts at the qa/ assertions until his `/council` runs land. Stamp path pre-validated: the shipped `council-stamp.cjs` stamped a scratchpad copy (placeholder → absolute `LENS_ROOT`, action "written"). **Validation:** ```bash diff --git a/package.json b/package.json index 583fa9321..f73df35d1 100644 --- a/package.json +++ b/package.json @@ -26,7 +26,7 @@ "skills:audit": "bun run scripts/skills-audit.ts", "lint:complexity-budget": "bun run scripts/complexity-budget.ts", "lint:council-workflow": "bun scripts/council-workflow-lint.ts", - "check": "bun run typecheck && bun run lint && bun run dead-code && bun run skills:lint && bun run wishes:lint && bun test", + "check": "bun run typecheck && bun run lint && bun run dead-code && bun run skills:lint && bun run wishes:lint && bun run lint:council-workflow && bun test", "check:fast": "bun run typecheck && bun run lint && bun run dead-code && bun run skills:lint && bun run wishes:lint", "verify:release": "scripts/verify-release.sh" }, diff --git a/plugins/genie/README.md b/plugins/genie/README.md index d9768800d..405825e9f 100644 --- a/plugins/genie/README.md +++ b/plugins/genie/README.md @@ -27,6 +27,18 @@ Execute wish tasks with bounded fix loops and per-group validation evidence. ### 4) `/review` Universal review gate (plan, execution, PR) returning `SHIP`, `FIX-FIRST`, or `BLOCKED`. +## `/council` workflow + +The multi-perspective engine ships as a native dynamic workflow, not a skill: + +- **What ships**: `workflows/council.js` (the engine template), `references/lenses/` (6 deliberation cards), and the 7 lane skills (`repo-hygiene`, `architecture`, `code-quality`, `qa`, `perf`, `supply-chain`, `dx-docs`) doubling as audit lenses. +- **Distribution**: the SessionStart hook stamps `LENS_ROOT` with the installed plugin path and copies the template to `~/.claude/workflows/council.js` — idempotent, and re-stamped on the first session after a plugin update. +- **Modes**: + - `/council ` — deliberation: 3-4 lenses routed by topic, 2-round Socratic exchange, dissent preserved verbatim. + - `/council audit [focus]` — lane audit: assess-only, evidence-backed findings that route to `/wish`, profile updates merged single-writer into `.genie/repo-profile.md`. +- **Requirements**: Claude Code ≥ 2.1.154 with dynamic workflows available (paid plans; an org-level `disableWorkflows` setting turns the command off). +- **Override**: a project-level `.claude/workflows/council.js` takes precedence over the personal stamped copy. + ## Directory Structure ```text @@ -37,6 +49,7 @@ genie/ ├── agents/ ├── hooks/ ├── scripts/ +├── workflows/ └── references/ ``` From c21a0df9ae81da1ac234d8f01eba4fb86aa86ca4 Mon Sep 17 00:00:00 2001 From: namastex888 Date: Fri, 10 Jul 2026 01:11:39 -0300 Subject: [PATCH 7/8] docs(genie): jar state after routing-matrix + council; hook-injection wish docs --- .genie/INDEX.md | 4 +- .../wishes/hook-injection-hardening/WISH.md | 146 ++++++++++++++++++ 2 files changed, 148 insertions(+), 2 deletions(-) create mode 100644 .genie/wishes/hook-injection-hardening/WISH.md diff --git a/.genie/INDEX.md b/.genie/INDEX.md index 286936dbe..e501a95fb 100644 --- a/.genie/INDEX.md +++ b/.genie/INDEX.md @@ -16,7 +16,7 @@ ## Ready -- [WISH: hook-injection-hardening](wishes/hook-injection-hardening/WISH.md) — BLOCKED-clearing safety edit: `execFileSync` at 3 hook sites (audit-context, freshness×2) + hostile-filename regression tests + `core.bare` probe removal; flips panel verdict BLOCKED→FIX-FIRST — plan review SHIP (1 fix loop), /work user-gated (2026-07-09) +- [WISH: hook-injection-hardening](wishes/hook-injection-hardening/WISH.md) — BLOCKED-clearing safety edit: `execFileSync` at 3 hook sites (audit-context, freshness×2) + hostile-filename regression tests + `core.bare` probe removal; flips panel verdict BLOCKED→FIX-FIRST — SHIPPED → PR #2536 (wish/hook-injection-hardening→main), G1+G2+whole-wish reviews SHIP, 729 pass/0 fail (2026-07-09) - [WISH: v5-completion](wishes/v5-completion/WISH.md) — CLAUDE.md-for-v5 rewrite ∥ Codex launch target + Hermes decision ∥ distribution 5.x (drafted 2026-07-02) - [WISH: dispatch-inproc-default](wishes/dispatch-inproc-default/WISH.md) — HIGH discovered defect: v5 hooks fall open by default (daemon deleted in demolition); re-arm branch-guard + unblock omni approvals (drafted 2026-07-02) - [WISH: omni-approval-ux](wishes/omni-approval-ux/WISH.md) — correlated approval identity, reaction approve/deny, anti-spam feedback; grounded in the 2026-07-03 live WhatsApp QA (drafted 2026-07-03) @@ -24,7 +24,7 @@ - [WISH: warp-integration](wishes/warp-integration/WISH.md) — umbrella Group 3: genie init, Warp launch-config emitter, genie launch, /work multi-session opt-in (drafted 2026-07-02) ## Poured -- [council-workflow](brainstorms/council-workflow/DESIGN.md) · [WISH](wishes/council-workflow/WISH.md) — /council vira saved workflow nativo (deliberation + audit), lens library 13 (7 persona skills renomeadas por lane + 6 cards), distribuição via install-stamp em ~/.claude/workflows; implementa a disposição council do skill-absorbs G4 — design review SHIP + plan review SHIP (todos os MEDIUMs aplicados; G3/G4 gated em fable5 MERGED), `/work` user-gated (2026-07-09) +- [council-workflow](brainstorms/council-workflow/DESIGN.md) · [WISH](wishes/council-workflow/WISH.md) — /council vira saved workflow nativo (deliberation + audit), lens library 13 (7 persona skills renomeadas por lane + 6 cards), distribuição via install-stamp em ~/.claude/workflows; implementa a disposição council do skill-absorbs G4 — EXECUTADO 2026-07-10: G1 `0b222e17`, G2 `12270b21` (1 fix loop: unwrap p/ body-style do runtime), G3 `22a7ed50`, G4 `2af7ebd9`, G5-eng `6494253d` — todos SHIP; QA vivo USER-GATED pós-release (ritual Felipe: merge→release→plugin update→rodar /council "revisar tudo"; g5-gate para na cauda qa/ até lá); final execution review após o QA - [plugin-resource-shipping](brainstorms/plugin-resource-shipping/DRAFT.md) · [WISH](wishes/plugin-resource-shipping/WISH.md) — in-skill template via ${CLAUDE_SKILL_DIR}, probe-guarded lint refs, resource-shipping lint, fresh-install CI smoke — plan review SHIP (2 fix loops), /work user-gated, G1 dep on routing-matrix:3 (2026-07-09) - [routing-matrix](brainstorms/routing-matrix/DRAFT.md) · [WISH](wishes/routing-matrix/WISH.md) — pinned role agents (Fable gates / Opus ladder / Haiku scouts), stage pins + pane flags, complexity columns + lint, escalation caps — executed; execution review SHIP, final gate 719 pass / 1 skip / 0 fail; live LangWatch pin QA pending (2026-07-09) - [genie-token-efficiency-program](brainstorms/genie-token-efficiency-program/DESIGN.md) — umbrella (WRS 100, independent review SHIP; 9 seed groups): genie → control-plane thesis; routing matrix (Fable gates, Opus ladder, no Sonnet, $17.9k/21d baseline); 17-skill dispositions; always-on identity; cross-agent delegate (Codex/Hermes companion sessions); genie spend; brainstorm domain-map upgrade — crystallized 2026-07-09 diff --git a/.genie/wishes/hook-injection-hardening/WISH.md b/.genie/wishes/hook-injection-hardening/WISH.md new file mode 100644 index 000000000..e4b687f1c --- /dev/null +++ b/.genie/wishes/hook-injection-hardening/WISH.md @@ -0,0 +1,146 @@ +# Wish: Hook shell-injection hardening (the BLOCKED-clearing safety edit) + +| Field | Value | +|-------|-------| +| **Status** | SHIPPED — [PR #2536](https://github.com/automagik-dev/genie/pull/2536) (commit `f61aaf13`, branch `wish/hook-injection-hardening` → main). G1+G2 + whole-wish reviews SHIP; 729 pass/0 fail; full `bun run check` green on the PR base (verified by the pre-push hook) | +| **Slug** | `hook-injection-hardening` | +| **Date** | 2026-07-09 | +| **Author** | namastex888 | +| **Appetite** | 1 afternoon (~4h) | +| **Branch** | `wish/hook-injection-hardening` | +| **Design** | _No brainstorm — direct wish (evidence: `.genie/repo-profile.md` + seven-lane panel synthesis)_ | + +## Summary + +The three PreToolUse file-path hooks interpolate the tool's `file_path` into a shell command string (`execSync(\`git … -- ${JSON.stringify(filePath)}\`)`), and since `JSON.stringify` escapes quotes and backslashes but not `$` or backticks, a `file_path` of `$(cmd)` runs `cmd` under `sh -c` with the user's privileges — a live RCE in any checkout whose `.claude/settings.json` routes `"*"` to dispatch, as this repo's does. This wish removes the shell at all three sites via `execFileSync('git', [...argv, '--', filePath])`, pins it with hostile-filename regression tests exercising each handler's reachable path, and — bundled per the panel — deletes the sibling `core.bare` startup probe that forks git on every invocation. Landing it flips the panel verdict **BLOCKED → FIX-FIRST**. + +## Scope + +### IN +- Replace the shell-interpolating `execSync` git call in `getRecentGitHistory` (`src/hooks/handlers/audit-context.ts`) with `execFileSync('git', [...argv, '--', filePath])`. +- Replace both shell-interpolating `execSync` git calls in `src/hooks/handlers/freshness.ts` (`getLastCommitInfo` and `checkUncommittedChanges`) the same way, dropping the now-meaningless shell double-quotes around the `--format` value. +- Hostile-filename regression tests that prove neither handler executes an embedded command, exercising each handler's *reachable* path. +- A functional regression proving `freshness` still emits a stale-read warning after the de-shell (guards the `--format` parse trap). +- Remove the top-level `core.bare` probe from `src/genie.ts` so it no longer forks git on the universal invocation path (finding 6). + +### OUT +- Narrowing the repo's `.claude/settings.json` `"*"` PreToolUse matcher — a separate policy change; this wish removes the vulnerability at the source so matcher width stops mattering. +- Findings 2/3/5 (dependency-aware `launch`, enforced completion authority, the phantom DAG) — those are the separate **execution-truth** wish, which carries a product decision. +- Any change to the shipped plugin's hook routing (already safe — it routes only `SendMessage`). +- Broader subprocess-hardening sweeps outside these three sites and the `core.bare` probe. + +## Decisions + +| # | Decision | Rationale | +|---|----------|-----------| +| 1 | Fix via `execFileSync('git', [...argv, '--', filePath])`, not shell-escaping | Removing the shell entirely is the only robust fix; escaping `$`/backticks is a denylist that rots. The `--` pathspec separator also stops a filename beginning with `-` from being parsed as a git flag. | +| 2 | `freshness.ts` `--format` argv element is `--format=%at\|%an\|%s` with **no** surrounding quotes | The `"…"` in the current string are shell quoting that `sh` strips before git sees them. Carried into an argv array they become literal, corrupting the `%at` field so `Number.parseInt` yields `NaN` and freshness silently stops warning. | +| 3 | Regression tests exercise each handler's *reachable* path, not a generic call | `audit-context` has no existence gate (primary vector) — pass the hostile `file_path` directly. `freshness` is `statSync`-gated — the test must create a real on-disk file *literally named* with the payload so `statSync` succeeds and the injectable git call is reached; freshness has **two** such sites (`getLastCommitInfo` for committed files, `checkUncommittedChanges` for uncommitted), so its tests cover both. The static `execSync(` grep gate backstops both regardless of test reachability. | +| 4 | Bundle the `core.bare` probe removal (finding 6) as Group 2 | Panel + performance lane both flag it; both are one-line footgun removals on the same subprocess-hygiene surface, shipping in the same afternoon. Kept a separate group/commit so it reverts independently of the security fix. | +| 5 | Delete the `core.bare` probe rather than relocate it | Its own comment says it "should no longer trigger" (the v4 worktree-corruption path is gone; v5 uses `git clone --shared`), and it can flip a legitimately-bare repo's `core.bare` to false. Relocating a guarded check into `launch` is an acceptable alternative — the Group 2 gate only asserts the probe is gone from `genie.ts`, so it passes either way. | + +## Success Criteria + +- [x] A hostile `file_path` containing `$(…)` executes **no** embedded command through `auditContext` — regression test creates a temp git repo, invokes the handler with `file_path: '$(touch PWNED)'`, and asserts no `PWNED` file exists afterward. +- [x] A hostile on-disk filename containing `$(…)` executes **no** embedded command through `freshness` — regression test creates a fresh-mtime file literally named `$(touch PWNED)`, invokes the handler, and asserts no `PWNED` file exists afterward. +- [x] `freshness` still emits a stale-read warning for a genuinely recent file authored by another agent (proves the `--format` parse survived the de-shell). +- [x] No `execSync(` call remains in `audit-context.ts` or `freshness.ts`; both import and use `execFileSync`. +- [x] The `core.bare` startup probe no longer appears anywhere in `src/genie.ts`. +- [ ] `bun run check` fully green — **all gates PASS for this wish**: typecheck, biome lint, knip dead-code, skills:lint, wishes:lint (this wish conforms), and `bun test` = 729 pass / 1 skip / 0 fail. Box left unchecked because `bun run check` overall still exits 1 for ONE out-of-scope reason: a concurrently-created file, `.genie/wishes/council-workflow/WISH.md`, is missing the same Complexity/Model columns. Not a regression from these changes; owned by another session. +- [x] `bun run build` succeeds and `dist/genie.js --version` works from a non-git directory. + +## Execution Strategy + +### Wave 1 (parallel — both zero-dependency, disjoint files) + +| Group | Agent | Complexity | Model | Description | +|-------|-------|-----------|-------|-------------| +| 1 | engineer | 4 (security fix + reachable-path regression tests) | opus·xhigh | De-shell the three file-path hook git calls (`execFileSync`) + hostile-filename injection tests. The BLOCKED-clearing gate. | +| 2 | engineer | 2 (mechanical deletion + build smoke) | opus·high | Remove the `core.bare` startup probe from `src/genie.ts`. Bundled footgun removal. | + +Both groups touch disjoint files (hook handlers vs. the entry module), so they run fully in parallel with no shared edits. Group 1 is the BLOCKED-clearing gate; Group 2 is the bundled footgun removal. + +--- + +## Execution Groups + +### Group 1: De-shell the file-path hook git calls (+ regression tests) +**Goal:** No PreToolUse hook can execute a command embedded in a `file_path`, and tests permanently own that guarantee. + +**Deliverables:** +1. `src/hooks/handlers/audit-context.ts` — `getRecentGitHistory` uses `execFileSync('git', ['log', '--oneline', '-n', String(MAX_COMMITS), '--', filePath], opts)`; import switches from `execSync` to `execFileSync`. +2. `src/hooks/handlers/freshness.ts` — `getLastCommitInfo` uses `execFileSync('git', ['log', '-1', '--format=%at|%an|%s', '--', filePath], opts)` (no quotes around the format); `checkUncommittedChanges` uses `execFileSync('git', ['status', '--porcelain', '--', filePath], opts)`; import switches to `execFileSync`. +3. `src/hooks/handlers/__tests__/audit-context.test.ts` — **extend the existing file** (this dir uses a `__tests__/` subdir, not colocated tests) with an injection-named test asserting `file_path: '$(touch PWNED)'` creates no `PWNED` file (reuse the existing temp-git-repo fixture). +4. `src/hooks/handlers/__tests__/freshness.test.ts` — **extend the existing file** with injection-named tests covering **both** reachable git sites: (a) a *committed* fresh-mtime file literally named `$(touch PWNED)` — reaches `getLastCommitInfo` (freshness.ts:24); and (b) an *uncommitted* fresh-mtime file literally named `$(touch PWNED2)` with `GENIE_AGENT_NAME` set — reaches `checkUncommittedChanges` (freshness.ts:72); each asserting the sentinel is never created. **Plus** a functional test asserting `freshness` still returns a stale-read warning for a real recent commit by another author (guards the `--format` parse). + +**Acceptance Criteria:** +- [x] `grep -nE '\bexecSync\('` finds nothing in either handler; `grep -q execFileSync` finds it in both. +- [x] Both handler test files (under `__tests__/`) contain injection-named tests and pass under `bun test`; the freshness tests cover both the committed (`getLastCommitInfo`) and uncommitted (`checkUncommittedChanges`) injectable sites. The static `execSync(` grep gate backstops both sites regardless of per-test reachability. +- [x] The freshness functional test proves the `--format` parse still yields a valid timestamp (warning is emitted). +- [x] `bun run typecheck` is clean. + +**Validation:** +```bash +cd "$(git rev-parse --show-toplevel)" +set -euo pipefail + +# 1. No shell-interpolating execSync remains (execFileSync is allowed and expected). +if grep -nE '\bexecSync\(' src/hooks/handlers/audit-context.ts src/hooks/handlers/freshness.ts; then + echo "FAIL: shell-interpolating execSync( still present in a hook handler"; exit 1 +fi + +# 2. Both handlers now call git via execFileSync argv (no shell). +grep -q "execFileSync('git'" src/hooks/handlers/audit-context.ts +grep -q "execFileSync('git'" src/hooks/handlers/freshness.ts + +# 3. Injection regression tests exist (in __tests__/), are named for the threat, and pass. +grep -qi 'inject' src/hooks/handlers/__tests__/audit-context.test.ts +grep -qi 'inject' src/hooks/handlers/__tests__/freshness.test.ts +bun test src/hooks/handlers/__tests__/audit-context.test.ts src/hooks/handlers/__tests__/freshness.test.ts + +# 4. Type gate clean. +bun run typecheck +``` + +**depends-on:** none + +--- + +### Group 2: Remove the `core.bare` startup probe from the universal path +**Goal:** `genie` stops forking git on every `--version`, `--help`, and hook fork, and stops being able to clobber a legitimately-bare repo's `core.bare`. + +**Deliverables:** +1. `src/genie.ts` — delete the top-level `core.bare` guard **including its preceding explanatory comment** (the whole block spans lines ~34-46; note the comment at line 36 contains the literal string `core.bare`, so leaving the comment would trip the Group 2 gate) and its module-scope `require('node:child_process')`. +2. If corruption recovery is still wanted, a guarded equivalent may be relocated into a worktree-creating command (`launch`) only — optional per Decision 5; not required for the gate. + +**Acceptance Criteria:** +- [x] `core.bare` appears nowhere in `src/genie.ts`. +- [x] No module-scope `require('node:child_process')` remains in `src/genie.ts`. +- [x] `bun run build` succeeds and `dist/genie.js --version` runs cleanly from a non-git directory. +- [x] `bun run typecheck` is clean. + +**Validation:** +```bash +cd "$(git rev-parse --show-toplevel)" +set -euo pipefail + +# 1. The core.bare probe is gone from the entry module. +if grep -n 'core.bare' src/genie.ts; then + echo "FAIL: core.bare probe still present in src/genie.ts"; exit 1 +fi + +# 2. No module-scope child_process require remains in the entry module. +if grep -nE "require\('node:child_process'\)" src/genie.ts; then + echo "FAIL: module-scope child_process require still in src/genie.ts"; exit 1 +fi + +# 3. Build + version smoke from a NON-git dir (startup must not hard-depend on git). +bun run build +tmp="$(mktemp -d)" +( cd "$tmp" && bun "$(git rev-parse --show-toplevel)/dist/genie.js" --version ) + +# 4. Type gate clean. +bun run typecheck +``` + +**depends-on:** none From 6e94cbd1a465bc3ade4902ca0063a497e87e6df5 Mon Sep 17 00:00:00 2001 From: "github-actions[bot]" Date: Fri, 10 Jul 2026 04:13:19 +0000 Subject: [PATCH 8/8] chore(version): bump to 5.260710.2 [auto-version] --- .claude-plugin/marketplace.json | 2 +- package.json | 2 +- plugins/genie/.claude-plugin/plugin.json | 2 +- plugins/genie/package.json | 2 +- 4 files changed, 4 insertions(+), 4 deletions(-) diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index ae7b83234..2e82ec2cf 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -10,7 +10,7 @@ "plugins": [ { "name": "genie", - "version": "5.260710.1", + "version": "5.260710.2", "source": "./plugins/genie", "description": "Human-AI partnership for Claude Code. Share a terminal, orchestrate workers, evolve together. Brainstorm ideas, wish them into plans, make with parallel agents, ship as one team. A coding genie that grows with your project." } diff --git a/package.json b/package.json index 68045cb0a..8357e6466 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@automagik/genie", - "version": "5.260710.1", + "version": "5.260710.2", "description": "Collaborative terminal toolkit for human + AI workflows. NOTE: npm distribution discontinued 2026-05-09 — install via `curl -fsSL https://raw.githubusercontent.com/automagik-dev/genie/main/install.sh | bash` (cosign + SLSA verified). See https://automagik.dev/genie/release-process", "type": "module", "bin": { diff --git a/plugins/genie/.claude-plugin/plugin.json b/plugins/genie/.claude-plugin/plugin.json index 7645667c5..4bb82f3b2 100644 --- a/plugins/genie/.claude-plugin/plugin.json +++ b/plugins/genie/.claude-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "genie", - "version": "5.260710.1", + "version": "5.260710.2", "description": "Human-AI partnership for Claude Code. Share a terminal, orchestrate workers, evolve together. Brainstorm ideas, turn them into wishes, execute with /work, validate with /review, and ship as one team.", "author": { "name": "Namastex Labs" diff --git a/plugins/genie/package.json b/plugins/genie/package.json index b1e7219e0..61227f473 100644 --- a/plugins/genie/package.json +++ b/plugins/genie/package.json @@ -1,6 +1,6 @@ { "name": "genie-plugin", - "version": "5.260710.1", + "version": "5.260710.2", "private": true, "description": "Runtime dependencies for genie bundled CLIs", "type": "module",