Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
7223cd9
perf(ci): optimize validation workflow and tooling
Oba-One Aug 16, 2026
fb83541
Merge branch 'develop' into perf/ci-speed-optimization
Oba-One Aug 16, 2026
a017f35
perf(ci): drop the net-negative Bun dependency cache
Oba-One Aug 16, 2026
0086325
fix(claude): keep scoped format from failing on non-Biome paths
Oba-One Aug 16, 2026
2a1fd9b
docs(claude): name the receipt-reuse flag in the validation pipeline
Oba-One Aug 16, 2026
a50b154
perf(claude): run independent package suites concurrently in the loca…
Oba-One Aug 16, 2026
7406ff9
docs(claude): record the concurrency, format, and indexer findings
Oba-One Aug 16, 2026
bbbfe80
docs(claude): pin the indexer cost to processEvent, not parallelism
Oba-One Aug 16, 2026
3846762
docs(claude): quantify the cache verdict and the indexer per-call cost
Oba-One Aug 16, 2026
d03ea28
perf(indexer): batch adjacent mock events in the Hats role tests
Oba-One Aug 16, 2026
3a5c194
fix(claude): use a valid lane status in the validation hub
Oba-One Aug 16, 2026
9f23920
perf(indexer): batch adjacent mock events across four more test files
Oba-One Aug 16, 2026
5f68c3a
docs(claude): record CI-verified results and correct the inflated bas…
Oba-One Aug 16, 2026
e6dba12
docs(claude): record that Envio 3.6.1 removes the per-call test cost
Oba-One Aug 16, 2026
e6c3e90
perf(indexer): upgrade Envio to 3.6.1 and route mock events at indexe…
Oba-One Aug 16, 2026
6a7f66e
docs(claude): record the shipped Envio 3.6.1 upgrade
Oba-One Aug 16, 2026
20926f1
fix(claude): address PR review feedback on the validation branch
Oba-One Aug 17, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/context/admin.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ Loaded when working in `packages/admin/`. Extends CLAUDE.md.
| Command | Purpose |
|---------|---------|
| `bun run test` | Run tests (vitest) |
| `bun build` | Build (includes TypeScript check) |
| `bun run build` | Build (includes TypeScript check) |
| `bun lint` | Lint with oxlint |
| `bun dev` | Start dev server (via PM2 from root) |

Expand Down
2 changes: 1 addition & 1 deletion .claude/context/client.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ Loaded when working in `packages/client/`. Extends CLAUDE.md.
| Command | Purpose |
|---------|---------|
| `bun run test` | Run tests (vitest) |
| `bun build` | Build (includes TypeScript check) |
| `bun run build` | Build (includes TypeScript check) |
| `bun lint` | Lint with oxlint |
| `bun dev` | Start dev server (via PM2 from root) |

Expand Down
2 changes: 1 addition & 1 deletion .claude/context/contracts.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ Loaded when working in `packages/contracts/`. Extends CLAUDE.md.
|---------|---------|
| `bun run test` | Run unit tests (skips E2E) |
| `bun run test:gas` | Tests with gas report |
| `bun build` | Adaptive build (changed Solidity targets with shared-file fallback to `src`) |
| `bun run build` | Adaptive build (changed Solidity targets with shared-file fallback to `src`) |
| `bun build:changed` | Build changed Solidity under `src/test/script` only |
| `bun build:target -- <path...>` | Build explicit Solidity target(s) only |
| `bun build:fast` | Explicit fast mode (`src` only, skips Foundry test/script) |
Expand Down
71 changes: 62 additions & 9 deletions .claude/context/validation-pipeline.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,17 +5,61 @@ instead of restating the pipeline, so a change to the gate (adding a step,
renaming a script) happens in exactly one place. The intent ladder that decides
*which* rung to run lives in `CLAUDE.md § Validation Intent Ladder`.

## Review Readiness Gate (non-mutating)
## Select before executing

The strict evidence gate for plain `/review`. It proves bounded production readiness without
editing tracked files:
Render the repository-owned plan first:

```bash
bun run validation:plan -- --intent <intent>
```

The selector combines intent, changed paths, dependency impact, and criticality. Agents execute the
returned plan instead of inventing a broader command set. If the selector command is unavailable or
fails, fall back to the intent ladder and commands below and report the selector problem. A missing
or failed selector never authorizes omitting a required check or critical override.

Every selected check states:

- **Risk** — the concrete regression, invariant, or acceptance criterion it covers.
- **Expected signal** — the observable pass/fail evidence the command provides.
- **Freshness** — the source inputs, validated paths, validation entrypoint, policy, toolchain, and
environment profile that must remain identical before a passing receipt can be reused.
- **Stop** — which dependent checks stop after a deterministic failure and which explicitly
independent diagnostics may continue.

Receipt reuse is opt-in and off by default. Pass `--reuse-passing-receipts` to
`node scripts/dev/ci-local.js` to skip checks whose exact fingerprint already passed. The store
lives in `.cache/validation`, holds passes only, and any change to the command, policy, toolchain,
validated paths, or environment profile invalidates the fingerprint. A tampered store is rejected
rather than trusted.

Never reuse failures. User cancellation is terminal: stop active validation, schedule nothing else,
and report only evidence already collected. An unavailable browser, RPC, secret, service, or other
capability produces `BLOCKED`, not passing; do not retry the identical check until that capability
changes. Budgets warn and profile but never skip contract, deployment/release, authentication,
JobQueue, Work-provider, mutation-hook, security, ontology, supply-chain, or release gates. Contracts use Bun
wrappers only, never raw Forge.

## Diagnosis and evidence review (non-mutating)

Inspect existing evidence first, then run only the check needed to prove or disprove each material
finding. Keep commands non-mutating and stop dependent work on the first deterministic failure. This
rung supports diagnosis, audit, and ordinary evidence review; it does not certify production
readiness. When the user explicitly asks for production quality, approval, PR/merge readiness, or a
readiness verdict, use the full Production Review Readiness Gate below.

## Production Review Readiness Gate (non-mutating)

The strict evidence gate for an explicit production-readiness review. It proves bounded production
readiness without editing tracked files:

```bash
bun format:check && bun lint && bun run test && VITE_CHAIN_ID=11155111 bun run build
```

Run every stage fresh in the current review. A required failure means `REQUEST_CHANGES`. A required
check that cannot run means `COMMENT_ONLY`; do not downgrade or replace the proof silently.
Run every selected stage fresh unless an exact matching receipt satisfies the freshness contract
above. A required failure means `REQUEST_CHANGES`. A required check that cannot run means
`COMMENT_ONLY`; do not downgrade or replace the proof silently.

Conditional additions when the change touches the relevant surface:

Expand All @@ -35,6 +79,11 @@ Conditional additions when the change touches the relevant surface:
matching Playwright CI project — client: `PLAYWRIGHT_APP=client APP_ENV=test bunx playwright
test --project=client-ci`; admin: `PLAYWRIGHT_APP=admin APP_ENV=test bunx playwright test
--project=admin-ci`
- Agent runtime changes: `bun run build:agent`
- Docs runtime, navigation, or build configuration changes: `bun run build:docs`

The root `bun run build` covers Contracts, Shared, Indexer, Client, and Admin. It does not build
Agent or Docs; the conditional commands above close those scopes.

Visible UI additionally requires rendered proof through the authenticated Brave QA profile. If
that path is unavailable, record browser proof as `BLOCKED` and return `COMMENT_ONLY` unless a
Expand All @@ -46,10 +95,10 @@ clean-room browser-proof commands cannot substitute for authenticated local QA.
The pre-merge/pre-push gate — required before claiming a branch is ready:

```bash
bun format && bun lint && bun run test && bun build
bun format && bun lint && bun run test && bun run build
```

Ship uses the same conditional additions listed in the Review Readiness Gate. Unlike review, ship
Ship uses the same conditional additions listed in the Production Review Readiness Gate. Unlike review, ship
may run mutating format and branch/commit safety steps because the user explicitly requested ship,
PR, commit, merge, or release readiness.

Expand All @@ -67,9 +116,13 @@ node scripts/dev/ci-local.js --quick
Targeted proof for an isolated fix — the package-local test file or command
that proves the touched behavior (see the intent ladder). Common shapes:

- Style only: `bun format && bun lint`
- Style only: `bunx @biomejs/biome format --no-errors-on-unmatched <changed-files...>` and, for changed JavaScript or
TypeScript, `bunx oxlint <changed-source-files...> --deny-warnings`. Both commands are
path-scoped and non-mutating. Do not use workspace-mutating `bun format` or broad `bun lint` for
isolated style-only QA.
- One behavior: `bun run --filter <pkg> test <path/to/file>`
- Baseline capture before a sweep: `bun format && bun lint && bun run test`
- Baseline capture before a sweep: non-mutating `bun run format:check && bun lint`, then the
selector-chosen tests. Use the mutating `bun format` only in explicit fix/Ship intent.
(build intentionally omitted until the sweep lands)

Failing tests are never cached and never skipped around — fix the test, not
Expand Down
4 changes: 2 additions & 2 deletions .claude/skills/clean/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,7 +79,7 @@ Remove unused files/exports/types/deps found by `bunx knip --reporter compact`.

### Agent 4: Circular Dependency Resolution (madge)

Zero out cycles from `npx madge --circular --extensions ts,tsx packages/`. Resolution preference: `import type` → extract shared interface → dependency inversion → merge modules. Rules: respect build order `contracts -> shared -> indexer -> client/admin/agent`; never create upward dependencies; hooks stay in shared; `bun build` must pass after.
Zero out cycles from `npx madge --circular --extensions ts,tsx packages/`. Resolution preference: `import type` → extract shared interface → dependency inversion → merge modules. Rules: respect build order `contracts -> shared -> indexer -> client/admin/agent`; never create upward dependencies; hooks stay in shared; `bun run build` must pass after.
Comment thread
Oba-One marked this conversation as resolved.

### Agent 5: Type Strengthening

Expand Down Expand Up @@ -211,7 +211,7 @@ Use `--no-codex` when:
git diff --check # Whitespace / conflict marker sanity
bun format:check && bun lint # Non-mutating style gate
bun run test # Correctness
bun build # Build integrity
bun run build # Build integrity
madge --circular --extensions ts,tsx packages/ # Only when locally installed or explicitly approved
bunx knip --reporter compact # Checked-in dependency; reduced dead code
```
Expand Down
2 changes: 1 addition & 1 deletion .claude/skills/debug/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -176,7 +176,7 @@ Simple fixes (<10 lines, single file, root cause proven) apply directly. Complex

## Part 3: Verification Before Completion

CLAUDE.md § Verify Before Claiming Success is the contract: evidence in the same turn, no "should work / probably / seems to". Standard proofs: `bun run test` (never `bun test`), `bun build`, `bun lint`, `npx tsc --noEmit` in the touched package.
CLAUDE.md § Verify Before Claiming Success is the contract: evidence in the same turn, no "should work / probably / seems to". Standard proofs: `bun run test` (never `bun test`), `bun run build`, `bun lint`, `npx tsc --noEmit` in the touched package.

---

Expand Down
4 changes: 2 additions & 2 deletions .claude/skills/debug/health-diagnostics.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ Moved from the debug SKILL.md body — load on demand, not on every activation.

```bash
# Compile and check artifacts
cd packages/contracts && bun build
cd packages/contracts && bun run build

# Inspect deployment addresses
cat deployments/11155111-latest.json | jq '.gardenToken'
Expand Down Expand Up @@ -76,7 +76,7 @@ cd packages/shared && npx tsc --noEmit
cd packages/client && npx tsc --noEmit

# Vite build with verbose output
cd packages/client && DEBUG=vite:* bun build
cd packages/client && DEBUG=vite:* bun run build

# Check bundle analysis
cd packages/client && npx vite-bundle-visualizer
Expand Down
48 changes: 41 additions & 7 deletions .claude/skills/review/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,9 +7,15 @@ user-invocable: true

# Review

One command for the standing request: **"review this — ensure no regressions, no remaining gaps, and production quality."** Three passes over one resolved scope, then a verdict. Read-only unless `--fix` is explicitly requested.
One command for change review. Three passes over one resolved scope, then a verdict. Read-only unless
`--fix` is explicitly requested. Evidence/diagnosis review is targeted by default; full production
readiness is a separate, explicit intent.

It answers three questions with fresh evidence: **regression safety** (Pass 1), **requirement closure** (Pass 2), **production readiness** (Pass 3). `APPROVE` is a bounded, evidence-backed readiness verdict for the reviewed scope — not a claim that unrelated repository or production failures are impossible.
It answers three questions with fresh evidence: **regression safety** (Pass 1), **requirement closure**
(Pass 2), and the user's requested **evidence or readiness level** (Pass 3). `APPROVE` is reserved for
an explicit production-quality, approval, PR/merge-readiness, or equivalent request whose full
non-mutating readiness gate passed. It is a bounded verdict for the reviewed scope, not a claim that
unrelated repository or production failures are impossible.

## Scoping

Expand All @@ -24,6 +30,22 @@ Valid package scopes map to `packages/<name>/**` (contracts, indexer, shared, cl
- `--scope cross-package` — verify blast radius in dependency order (contracts → shared → indexer → apps → agent); only cross-boundary findings.
- `--scope design-system` — delegate to [`design/system-alignment-review.md`](../design/system-alignment-review.md), read-only; return its sections directly, don't mix into diff findings. Fires only on explicit invocation, on DESIGN.md-dialect + theme/tokens co-changes, or when a change touches ≥2 visual surfaces at once.

### Review intent

Resolve intent separately from code scope:

- **Evidence review / diagnosis** — ordinary "review this", regression investigation, gap analysis,
audit evidence, or a specific question. Inspect first and run only the non-mutating checks needed
to prove findings. A clean targeted review returns `COMMENT_ONLY`, not a readiness approval.
- **Production readiness** — explicit requests for production quality, approval, PR/merge readiness,
or whether the branch is ready. Run the full Production Review Readiness Gate.

Render the planned checks with `bun run validation:plan -- --intent review`. For explicit production
readiness, use `--intent readiness` so the plan remains non-mutating while criticality can only add
checks. Execute the returned plan. If the selector is unavailable or fails, follow CLAUDE.md's
intent ladder and the shared validation
pipeline directly, report the selector problem, and preserve every hard gate.

### Authoritative requirements

After resolving the code scope, establish the requirement baseline in this order: (1) the user's current request and explicit acceptance criteria; (2) the PR description and any Linear issue linked there with `Fixes`, `Refs`, or `Relates to`, plus the legacy `Closes` and `Linear:` forms while existing PRs transition; (3) any `.plans/` lane referenced by the request, PR, or issue (`brief.md`, `spec.md`, `plan.todo.md`, `status.json`); (4) directly applicable package documentation those sources reference. Never infer issue identity from the branch name. Record which sources were available. If no authoritative requirements can be established, continue with useful findings but set the final verdict to `COMMENT_ONLY` — do not claim that no gaps remain.
Expand Down Expand Up @@ -83,16 +105,26 @@ Then sweep for the repo's recurring gap shapes:

Report gaps in their own section — a gap is not a defect; it's unfinished intent.

## Pass 3 — Production Quality
## Pass 3 — Evidence or Production Quality

Evidence review runs only the selector-chosen, non-mutating checks needed to prove or disprove its
findings. Do not add full tests or builds merely to make the command count look comprehensive. A
clean evidence review is not production certification and returns `COMMENT_ONLY`.

Plain review runs the non-mutating **Review Readiness Gate** defined in [`.claude/context/validation-pipeline.md`](../../context/validation-pipeline.md) (`format:check`, lint, tests, pinned `VITE_CHAIN_ID=11155111` build, plus its scope-conditional additions). Run every required stage fresh in this invocation — never reuse stale evidence. A required stage that fails → `REQUEST_CHANGES`; a required stage that cannot run → `COMMENT_ONLY`, never silently downgraded or substituted.
Explicit production-readiness review runs the non-mutating **Production Review Readiness Gate**
defined in [`.claude/context/validation-pipeline.md`](../../context/validation-pipeline.md)
(`format:check`, lint, tests, pinned `VITE_CHAIN_ID=11155111` root build, plus scope-conditional
additions including Agent or Docs builds). Run every selected stage fresh unless an exact matching
receipt satisfies the shared freshness contract. A required stage that fails → `REQUEST_CHANGES`; a
required stage that cannot run → `COMMENT_ONLY`, never silently downgraded or substituted.

For narrower explicit intents, pick the lightest honest rung per CLAUDE.md § Validation Intent Ladder:

- isolated fix → targeted package-local test/proof
- cross-package or shared-surface impact → Repo Quick Gate
- explicit ship/merge readiness → full Ship Gate + conditional design/vocab/story gates when those surfaces moved

For every selected check, name its risk, expected signal, freshness rule, and stopping condition.
State what ran with real output. Record the tested commit SHA, UTC timestamp, exact command, and
summarized result. Write green, passed, or merge-ready claims only after those commands finish in the
current review and an empty
Expand All @@ -102,8 +134,10 @@ path-scoped `git diff --exit-code <tested>..HEAD -- <validated paths>` proving a
implementation, dependency, configuration, and validation-entrypoint surfaces are unchanged, plus
an empty `git status --porcelain=v1 --untracked-files=all -- <validated paths>` proving no staged,
unstaged, or untracked path changes exist. If a
rung can't run here (env-gated, needs authenticated browser), say
"unverified: X" instead of hedging. Visible-UI claims need rendered proof via the authenticated Brave
rung can't run here (env-gated, or it requires an authenticated browser), mark it `BLOCKED`, name the unavailable
capability, and do not retry until that capability changes. User cancellation is terminal: stop
active validation, schedule no further checks, and report evidence already collected. Visible-UI
claims need rendered proof via the authenticated Brave
QA path or are reported as blocked (CLAUDE.md § Agentic Modern Web Standard). Dated reports under
`.plans/**/reports/` are immutable audit inputs; put corrections or closure evidence in a new report.

Expand Down Expand Up @@ -144,7 +178,7 @@ Lead with findings, keep the list actionable:
3. **Remaining Gaps** — unfinished intent, each with the smallest completing step
4. **Human Call-Outs** — dependencies, auth/permissions, migrations, contract deploys, trust-boundary changes (never auto-fix these)
5. **Verification** — what ran, real results, what remains unverified
6. **Verdict** — `APPROVE` | `REQUEST_CHANGES` | `COMMENT_ONLY`. Rules: any `MISSING` requirement or failed required check → `REQUEST_CHANGES`; any `BLOCKED` requirement/check, or no authoritative requirements available → `COMMENT_ONLY`; `APPROVE` only when every requirement is `SATISFIED`/`OUT_OF_SCOPE` and the readiness gate passed fresh.
6. **Verdict** — `APPROVE` | `REQUEST_CHANGES` | `COMMENT_ONLY`. Rules: any `MISSING` requirement or failed required check → `REQUEST_CHANGES`; any `BLOCKED` requirement/check, no authoritative requirements, or evidence-review-only intent → `COMMENT_ONLY`; `APPROVE` only for explicit production-readiness intent when every requirement is `SATISFIED`/`OUT_OF_SCOPE` and the full readiness gate passed under the freshness contract.

Finding format: `[Title] — severity · type · file:line · why it matters · next step`.

Expand Down
2 changes: 2 additions & 0 deletions .codex/config.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
project_doc_fallback_filenames = ["CLAUDE.md"]
project_doc_max_bytes = 40960

model_verbosity = "low"
model_reasoning_summary = "concise"
[features]
hooks = true

Expand Down
26 changes: 26 additions & 0 deletions .github/actions/setup-js/action.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
name: Setup Green Goods JavaScript
description: Pin Node and Bun, then install the frozen lockfile

runs:
using: composite
steps:
- name: Setup Node.js
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020
with:
node-version: "22.22.1"

- name: Setup Bun
uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6
with:
bun-version: "1.3.14"

# Deliberately uncached. Measured on this repository (PR #719, run
# 31963610888): restoring the ~878 MB Bun download store cost 20.7s
# (4s transfer, 16.6s extraction) to make `bun install` 21.8s -> 6.9s,
# a net loss of ~6s per job. The store also consumed 1.02 GB of an
# already-over-quota 10 GB repository cache, evicting the per-commit
# Foundry build caches that do pay for themselves. Reintroduce only with
# fresh before/after evidence that restore cost is below install savings.
- name: Install dependencies from frozen lockfile
shell: bash
run: bun install --frozen-lockfile
Loading
Loading