Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions .github/workflows/qa-smoke-preview-deploy.yml
Original file line number Diff line number Diff line change
Expand Up @@ -517,7 +517,7 @@ jobs:
run_project missionary_status "missionary" "development-missionary" "${MISSIONARY_URL:-}"

if [[ "${result}" == "FAIL" ]]; then
evidence="GitHub Actions artifact: pr-preview-smoke-playwright-artifacts; report paths: playwright-report/pr-preview-smoke-*; test result paths: test-results/pr-preview-smoke-*"
evidence="Sanitized diagnostics: playwright-report/pr-preview-smoke-*/sanitized/{index.html,results.json}. See the artifact-upload step for availability; raw test outputs are excluded."
fi

{
Expand All @@ -534,13 +534,13 @@ jobs:
with:
name: pr-preview-smoke-playwright-artifacts
path: |
playwright-report/pr-preview-smoke-*
test-results/pr-preview-smoke-*
if-no-files-found: ignore
playwright-report/pr-preview-smoke-*/sanitized/index.html
playwright-report/pr-preview-smoke-*/sanitized/results.json
if-no-files-found: error
Comment thread
cobmojo marked this conversation as resolved.
retention-days: 7

- name: Comment headless smoke QA result
if: steps.gate.outputs.should_run == 'true'
if: always() && steps.gate.outputs.should_run == 'true'
Comment thread
cobmojo marked this conversation as resolved.
env:
ADMIN_STATUS: ${{ steps.smoke.outputs.admin }}
ADMIN_URL: ${{ steps.deploy_admin.outputs.url }}
Expand Down
6 changes: 6 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,12 @@ commit metadata does not. CODEOWNERS routes reviews but does not grant access.
see `docs/guides/development/contributing.md`.
- **Required local PR/push-readiness gate:** `bun run ci:preflight` (exact stages
and focused debugging commands are documented in `docs/ci.md`).
- **Eve build boundary:** web builds and hosted admin previews emit unqualified
Eve artifacts without sandbox provisioning. Production services retain full
prewarming; use `bun run --cwd packages/eve-runtime build:full` in the approved
target environment for full qualification. Follow the
[build runbook](docs/guides/development/build-runbook.md#eve-artifacts-and-qualification)
and keep Release Off until the separate launch requirements are met.
- **Production E2E:** `bun run test:e2e:production-gate` is the bounded
release gate required for `production`; broader `bun run test:e2e` remains useful
for local feature validation.
Expand Down
7 changes: 7 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -170,6 +170,13 @@ behavior; its source-owner amendments remain proposed.

Use the per-app `dev:*` scripts when you only need one surface, or `bun run dev` / `bun run dev:all` when you need several (see root `package.json`).

`bun run build` and the `build:<app>` commands compile Eve dependency artifacts
without provisioning sandboxes. Hosted admin previews use the same unqualified
artifact mode and leave Eve Release Off. Standalone Eve and production service
builds retain full qualification; see the
[build runbook](docs/guides/development/build-runbook.md#eve-artifacts-and-qualification)
before treating any preview as launch evidence.

### AI Agent Guidance System

Agent-oriented docs live under `docs/ai/`:
Expand Down
4 changes: 4 additions & 0 deletions apps/admin/eslint.config.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -5,4 +5,8 @@ import { designSystemConfig } from "@asym/eslint-config/design-system.mjs";
export default [
...scopeWorkspaceConfig(nextjsConfig, import.meta.url),
...designSystemConfig({ workspace: "apps/admin" }),
{
// withEve emits deployment bundles here, not authored app/runtime source.
ignores: [".eve/vercel-services/**", ".vercel/output/**"],
},
];
1 change: 1 addition & 0 deletions apps/admin/next.config.ts
Original file line number Diff line number Diff line change
Expand Up @@ -101,4 +101,5 @@ const sentryConfig = withSentryConfig(

export default withEve(sentryConfig, {
eveRoot: `${WORKSPACE_ROOT}/packages/eve-runtime`,
eveBuildCommand: "bun run build:service",
Comment thread
cobmojo marked this conversation as resolved.
Comment thread
cobmojo marked this conversation as resolved.
});
1 change: 1 addition & 0 deletions bun.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

37 changes: 36 additions & 1 deletion docs/guides/development/build-runbook.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,12 +23,47 @@ This runbook defines the canonical build workflow for the Bun + Turborepo monore

Notes:

- Internal packages define `build` scripts as `tsc --noEmit` (type-check only, no JavaScript emit).
- Most shared packages define `build` as `tsc --noEmit` (type-check only). Eve
emits runtime artifacts; its qualification boundary is described below.
- Example: `bunx turbo run build --filter=@asym/ui` now executes the `@asym/ui` package build task.
- Turbo `build` caching tracks app artifacts (`.next/**`) and package/typecheck artifacts (`dist/**`, `*.tsbuildinfo`).
- Cache output globs in `turbo.json` must never be able to match a package's own `node_modules`. Use package-relative globs (`*.tsbuildinfo`, `dist/*.tsbuildinfo`), not recursive ones (`**/*.tsbuildinfo`), and keep the `"!node_modules/**"` guard in `build.outputs` / `typecheck.outputs`. See [Turbo cache restore replaces workspace symlinks](#turbo-cache-restore-replaces-workspace-symlinks).
- For troubleshooting, use Turbo filters directly: `bunx turbo run build --filter=@asym/<package>`.

## Eve artifacts and qualification

The generic web build planner sets `CORE_EVE_BUILD_MODE=artifacts` for its
children. Eve 0.25.1 then compiles with `--skip-sandbox-prewarm`; this applies to
both dependency and app phases without changing Turbo's dependency order.
Admin uses `bun run build:service` through the supported `withEve` build-command
option. This stable command selects artifact mode only when it executes on a
hosted preview. Production services retain the full SDK build, even when web
dependencies were compiled in artifact mode. The decision happens at service
execution because the SDK preserves previously generated service commands when
the same checkout builds another target.

These are **unqualified Release-Off artifacts**. They can expose the deployed
candidate for inspection, but do not prove a sandbox template exists or permit
autonomous effects. Auth, governance, policy, audit and release admission remain
unchanged. The bundled Vercel runtime refuses a missing sandbox template; preview
health or web smoke success is not sandbox qualification.

- `bun run --cwd packages/eve-runtime build` defaults to the full SDK build when
`CORE_EVE_BUILD_MODE` is unset. Unknown modes fail before the SDK starts.
- `bun run --cwd packages/eve-runtime build:artifacts` explicitly compiles only.
- `bun run --cwd packages/eve-runtime build:full` always selects the full SDK
build, including Vercel sandbox prewarming in a Vercel-targeted environment.

Turbo hashes the mode and does not cache the Eve `build` task. An artifact build
or an earlier prewarm cannot be replayed as new provider qualification. A local
non-Vercel build remains local compilation under the SDK's own behavior.

Before release, run full qualification in the correctly configured, authorized
target environment and collect the exact target-bound evidence required by the
[Eve launch runbook](../operations/eve-launch.md). Missing credentials, unavailable
governance, Release Off, or a denied prewarm remain qualification blockers. Do
not enable release to make CI pass or count artifact compilation as that proof.

## Environment profiles

## 1) Default local profile (CI-equivalent)
Expand Down
9 changes: 9 additions & 0 deletions docs/guides/operations/eve-launch.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,15 @@ Do not copy credential values into the manifest. The panel must show **Runtime
target: Configured** and **Release: Off**. If the governance state changes,
produce a new manifest against the new state version.

Hosted web previews compile Eve artifacts without sandbox prewarming so this
Release-Off inspection can occur. These outputs are unqualified: a healthy web
preview or Eve health endpoint is not proof of sandbox availability or release
readiness. Before activation, full runtime/sandbox qualification and all evidence
below are still required for the exact target. Use the full build described in
the [build runbook](../development/build-runbook.md#eve-artifacts-and-qualification);
if credentials or governance deny prewarming, report that qualification blocker.
Never activate Eve or relax its policy merely to make a build pass.

## 2. Collect launch evidence while release is off

Build one `eve-launch-manifest-v1` JSON document using the schema exported by
Expand Down
58 changes: 47 additions & 11 deletions docs/qa/development-headless-smoke.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,7 +86,15 @@ The Playwright config is `playwright.development-smoke.config.ts`. It:
(`QA_<SURFACE>_BASE_URL`, `VERCEL_<SURFACE>_AUTOMATION_BYPASS_SECRET`)
- sends bypass via headers, not query params
- runs headless Chromium, one worker
- writes report artifacts under `playwright-report/development-smoke/`
- writes HTML and JSON reports under `PLAYWRIGHT_REPORT_DIR`, defaulting to
`playwright-report/development-smoke/`
- writes bounded test evidence under
`PLAYWRIGHT_OUTPUT_DIR`, defaulting to `test-results/`

The preview workflow sets both directories per surface so a later Playwright
invocation does not overwrite an earlier surface's failure evidence. Blank
overrides use the local defaults above. The suite-specific reporter preserves
the original failed exit and test status; it does not relax any assertion.

The helpers in `tests/e2e/development-smoke/helpers.ts` cover:

Expand All @@ -98,19 +106,47 @@ The helpers in `tests/e2e/development-smoke/helpers.ts` cover:

## How to view the report

```bash
bunx playwright show-report playwright-report/development-smoke
```
Open `<report directory>/sanitized/index.html`. This is a bounded diagnostic summary,
not Playwright's interactive trace/report viewer.

## Evidence captured on failure

For any failed test, Playwright keeps under
`playwright-report/development-smoke/`:

- HTML report
- JSON report at `results.json`
- screenshot, trace, and video when retained by Playwright
- non-secret `evidence.json` attachments
For any failed test, Playwright keeps in the configured report and test-output
directories:

- HTML summary
- JSON report at `<report directory>/sanitized/results.json`
- test title, project, status and duration
- `evidence.json` with URL origin/path, page title/heading, and visible password
input count; known QA and bypass values are redacted
- the last 50 document/fetch/XHR response method/status/origin/path entries
observed during authentication, without request headers, cookies or bodies

Known QA and bypass values are redacted before truncation in their raw/trimmed,
URI, URI-component, form-URL-encoded and UTF-8 base64/base64url forms (with or
without padding). Percent escapes may use either hex case. This is a bounded
set of representations, not detection of arbitrary transformations. Response
metadata describes only browser-visible requests; server-side profile reads
may be absent and an access-denied route alone does not establish its cause.

URL queries and fragments are excluded. Raw trace, screenshot, video and the
automatic DOM error prompt are disabled for this credential-bearing suite:
traces retain bypass headers, authentication bodies and API arguments, and DOM
snapshots can retain password input values. The reporter replaces run-owned
test output (including generated error-context files) with the bounded
evidence above and never removes arbitrary attachment sources outside the
resolved run directories. It emits no assertion bodies or API-step text into
the HTML/JSON bundle. Reports and output directories must be separate and
must not be a workspace root or its ancestor. These checks resolve symlinks,
including the nearest existing parent of a new directory, before Playwright
startup. The reporter uses the checked canonical paths and checks them again
before cleanup.

CI uploads only `sanitized/index.html` and `sanitized/results.json`; raw test
output is never part of the upload allowlist. If a reporter or worker fails
before producing a sanitized bundle, the missing-artifact step fails instead
of uploading leftovers. Successful cleanup is not a prerequisite for keeping
raw credentials out of uploaded artifacts.

## Safety Rules

Expand Down
20 changes: 20 additions & 0 deletions docs/qa/pr-preview-smoke.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,6 +115,26 @@ Playwright. The preferred flow is that GitHub Actions deploys previews, runs
Playwright, uploads sanitized failure artifacts, and comments the PASS/FAIL
result for Claude to read.

Each surface invocation also sets `PLAYWRIGHT_REPORT_DIR` to
`playwright-report/pr-preview-smoke-<surface>` and `PLAYWRIGHT_OUTPUT_DIR` to
`test-results/pr-preview-smoke-<surface>`. The development smoke configuration
uses these paths for its bounded HTML/JSON summaries and redacted test evidence,
preserving earlier surfaces' evidence. The upload step allowlists only
`playwright-report/pr-preview-smoke-*/sanitized/index.html` and
`playwright-report/pr-preview-smoke-*/sanitized/results.json`; it never uploads
raw test output and fails if no sanitized file exists. Raw trace/media, DOM
snapshots and API-step/assertion bodies are excluded because they can retain
QA credentials and Vercel bypass values. URL
queries and fragments are removed; test status, title, duration, and safe
origin/path/title/heading remain available. Up to 50 authentication response
method/status/origin/path entries show browser-visible responses without
recording headers, cookies or bodies. Server-side profile reads may be absent;
these entries alone do not establish a profile, membership, tenant or role
failure. See the
[artifact boundary](./development-headless-smoke.md#evidence-captured-on-failure).
An artifact records the failed check; it does not turn a failed smoke test into
a pass.

## Rerun Method

Use one of these:
Expand Down
3 changes: 3 additions & 0 deletions eslint.config.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -181,6 +181,9 @@ const eslintConfig = defineConfig([
ignores: [
".next/**",
"**/.next/**",
// SDK-generated admin service bundles; authored Eve/app code stays linted.
"apps/admin/.eve/vercel-services/**",
"apps/admin/.vercel/output/**",
"out/**",
"**/out/**",
"build/**",
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# Design

The web build planner passes `CORE_EVE_BUILD_MODE=artifacts` to its child
commands while retaining Turbo's dependency graph. The Eve workspace build
dispatcher defaults to full compilation/prewarming and rejects unknown modes.
Explicit `build:artifacts` and `build:full` commands remain available.

Admin's existing `withEve` integration uses one stable `build:service` command.
That dispatcher selects artifact mode only when it executes on a hosted preview.
The SDK preserves existing generated service commands, so choosing the mode only
when generating config could carry a preview skip into a later production build.
Service execution ignores the generic web artifact variable; production and
ordinary standalone commands retain the SDK's full build. A production target
overrides a conflicting preview environment marker. The dispatcher uses the
canonical deployment-environment helpers so case and whitespace cannot hide a
production target. Hosted core-development and legacy staging retain artifact
mode; an explicit local development target retains full mode.

Both named build commands select their explicit mode through the same dispatcher,
overriding an inherited generic build mode. Full mode rejects the SDK's
skip-prewarm flag before starting Eve, while ordinary SDK arguments still pass
through. Explicit selectors cannot be combined with a forwarded service or
another mode selector. The Node 24 dispatcher imports the declared environment workspace
directly; it does not introduce another environment classifier.

Turbo hashes the mode. The Eve build task does not cache provider qualification:
an artifact build or a prior successful prewarm cannot stand in for a fresh
required full build. All auth, governance, sandbox and release code is unchanged.

Admin's generated `.eve/vercel-services` and `.vercel/output` bundles are excluded
from authored-source linting and data-boundary scans. Regression coverage checks
the exact output paths and confirms admin app/config and Eve agent/runtime source
remain checked. Actual scanner CLI fixtures reject raw database imports in
authored admin files, neighboring `.eve`/`.vercel` files, and another app's
same-named directory; retired CRM markers in authored Eve runtime still fail.

Installed Eve 0.25.1 source shows that skipping prewarm still emits app and
workflow functions. Its bundled Vercel runtime refuses a missing template rather
than prewarming it on demand. An isolated SDK fixture verified generated service
output and health without credentials or network access. Actual Core validation
must additionally exercise the package dispatcher and generated service.

Release-Off previews are unqualified artifacts. Later full sandbox/runtime
qualification and target-bound launch evidence remain required. If governance
denies that qualification, the denial remains a blocker; this change supplies no
alternate authorization path. No API, database, migration or runtime policy
change is part of this repair.

## Shared-context validation follow-up

The canonical full gate exposed a pre-existing relation-ID false positive:
the valid UUID `01234567-8910-4111-8123-456789012345` contains a substring matching
the payment-number detector. A generated related claim ID could therefore
reject an otherwise valid disagreement. This is the narrow repair already
preserved from #1862 in #1905's reviewed integration candidate.

After the existing UUID schema validation, sensitive-content scanning excludes
only the top-level `relatedClaimIds` metadata. It still scans claim values,
provenance, evidence and other content, including UUID-shaped payment text.
The original IDs are retained in the claim and still require visible existing
claims for the same field, tenant and root run. Invalid IDs and relationships
remain rejected. This restores the accepted structured context and disagreement
preservation requirements in `eve-subagent-catalog-shared-run-context`; it
does not grant new authority or change any release gate.

The deterministic relation-ID regression and UUID-shaped-content rejection
control travel with this repair. #1862's design packs and the remaining #1905
work are still pending separate integration; this small repair does not
establish either PR's full supersession.
Loading
Loading