Skip to content

sync: update pipeline-control with nearai/main 2026-06-19 - #12

Merged
theredspoon merged 2 commits into
pipeline-controlfrom
sync/pipeline-control-upstream
Jun 20, 2026
Merged

theredspoon merged 2 commits into
pipeline-controlfrom
sync/pipeline-control-upstream

Conversation

@personal-upstream-sync

Copy link
Copy Markdown

Automated sync from nearai/ironclaw:main into pipeline-control.

This PR is prepared on sync/pipeline-control-upstream; the workflow never updates pipeline-control directly.

aiworkbot and others added 2 commits June 19, 2026 09:51
* Bound approval command previews

* Tighten approval payload length guard

* Rebuild WebUI v2 approval bundle

* fix(reborn): reset approval preview expansion

---------

Co-authored-by: aiworkbot <220660587+aiworkbot@users.noreply.github.com>
Co-authored-by: Robert Yan <mstr.raphael@gmail.com>
Co-authored-by: aiworkbot <aiworkbot@users.noreply.github.com>
Co-authored-by: Robert Yan <46699230+think-in-universe@users.noreply.github.com>
@github-actions github-actions Bot added size: XL Changed-line size classification risk: low Risk classification contributor: regular Contributor history classification labels Jun 19, 2026
@theredspoon
theredspoon merged commit fa9662c into pipeline-control Jun 20, 2026
14 checks passed
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…tion (nearai#238)

* feat: add extension registry with metadata catalog, CLI, and onboarding integration

Adds a central registry that catalogs all 14 available extensions (10 tools,
4 channels) with their capabilities, auth requirements, and artifact references.
The onboarding wizard now shows installable channels from the registry and
offers tool installation as a new Step 7.

- registry/ folder with per-extension JSON manifests and bundle definitions
- src/registry/ module: manifest structs, catalog loader, installer
- `ironclaw registry list|info|install|install-defaults` CLI commands
- Setup wizard enhanced: channels from registry, new extensions step (8 steps)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(setup): resolve workspace errors for tool crates and channels-only onboarding

Tool crates in tools-src/ and channels-src/ failed `cargo metadata` during
onboard install because Cargo resolved them as part of the root workspace.
Add `[workspace]` table to each standalone crate and extend the root
`workspace.exclude` list so they build independently.

Channels-only mode (`onboard --channels-only`) failed with "Secrets not
configured" and "No database connection" because it skipped database and
security setup. Add `reconnect_existing_db()` to establish the DB connection
and load saved settings before running channel configuration.

Also improve the tunnel "already configured" display to show full provider
details (domain, mode, command) instead of just the provider name.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(registry): address PR review feedback on installer and catalog

- Use manifest.name (not crate_name) for installed filenames so
  discovery, auth, and CLI commands all agree on the stem (#1)
- Add AlreadyInstalled error variant instead of misleading
  ExtensionNotFound (#2)
- Add DownloadFailed error variant with URL context instead of
  stuffing URLs into PathBuf (#3)
- Validate HTTP status with error_for_status() before reading
  response bytes in artifact downloads (#4)
- Switch build_wasm_component to tokio::process::Command with
  status() so build output streams to the terminal (#6)
- Find WASM artifact by crate_name specifically instead of picking
  the first .wasm file in the release directory (#7)
- Add is_file() guard in catalog loader to skip directories (#8)
- Detect ambiguous bare-name lookups when both tools/<name> and
  channels/<name> exist, with get_strict() returning an error (#9)
- Fix wizard step_extensions to check tool.name for installed
  detection, consistent with the new naming (#11, #12)
- Fix redundant closures and map_or clippy warnings in changed files

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(setup): restore DB connection fields after settings reload

reconnect_postgres() and reconnect_libsql() called Settings::from_db_map()
which overwrote database_url / libsql_path / libsql_url set from env vars.
Also use get_strict() in cmd_info to surface ambiguous bare-name errors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: fix clippy collapsible_if and print_literal warnings

Collapse nested if-let chains and inline string literals in format
macros to satisfy CI clippy lint checks (deny warnings).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(registry): prefer artifacts for install-defaults and improve dir lookup

- InstallDefaults now defaults to downloading pre-built artifacts
  (matching `registry install` behavior), with --build flag for source builds.
- find_registry_dir() walks up 3 ancestor levels from the exe and adds
  a CARGO_MANIFEST_DIR fallback, matching load_registry_catalog() logic.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
)

* feat: add inbound attachment support to WASM channel system

Add attachment record to WIT interface and implement inbound media
parsing across all four channel implementations (Telegram, Slack,
WhatsApp, Discord). Attachments flow from WASM channels through
EmittedMessage to IncomingMessage with validation (size limits,
MIME allowlist, count caps) at the host boundary.

- Add `attachment` record to `emitted-message` in wit/channel.wit
- Add `IncomingAttachment` struct to channel.rs and re-export
- Add host-side validation (20MB total, 10 max, MIME allowlist)
- Telegram: parse photo, document, audio, video, voice, sticker
- Slack: parse file attachments with url_private
- WhatsApp: parse image, audio, video, document with captions
- Discord: backward-compatible empty attachments
- Update FEATURE_PARITY.md section 7
- Add fixture-based tests per channel and host integration tests

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: integrate outbound attachment support and reconcile WIT types (nearai#409)

Reconcile PR nearai#409's outbound attachment work with our inbound attachment
support into a unified design:

WIT type split:
- `inbound-attachment` in channel-host: metadata-only (id, mime_type,
  filename, size_bytes, source_url, storage_key, extracted_text)
- `attachment` in channel: raw bytes (filename, mime_type, data) on
  agent-response for outbound sending

Outbound features (from PR nearai#409):
- `on-broadcast` WIT export for proactive messages without prior inbound
- Telegram: multipart sendPhoto/sendDocument with auto photo→document
  fallback for files >10MB
- wrapper.rs: `call_on_broadcast`, `read_attachments` from disk,
  attachment params threaded through `call_on_respond`
- HTTP tool: `save_to` param for binary downloads to /tmp/ (50MB limit,
  path traversal protection, SSRF-safe redirect following)
- Message tool: allow /tmp/ paths for attachments alongside base_dir
- Credential env var fallback in inject_channel_credentials

Channel updates:
- All 4 channels implement on_broadcast (Telegram full, others stub)
- Telegram: polling_enabled config, adjusted poll timeout
- Inbound attachment types renamed to InboundAttachment in all channels

Tests: 1965 passing (9 new), 0 clippy warnings

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add audio transcription pipeline and extensible WIT attachment design

Add host-side transcription middleware (OpenAI Whisper) that detects audio
attachments with inline data on incoming messages and transcribes them
automatically. Refactor WIT inbound-attachment to use extras-json and a
store-attachment-data host function instead of typed fields, so future
attachment properties (dimensions, codec, etc.) don't require WIT changes
that invalidate all channel plugins.

- Add src/transcription/ module: TranscriptionProvider trait,
  TranscriptionMiddleware, AudioFormat enum, OpenAI Whisper provider
- Add src/config/transcription.rs: TRANSCRIPTION_ENABLED/MODEL/BASE_URL
- Wire middleware into agent message loop via AgentDeps
- WIT: replace data + duration-secs with extras-json + store-attachment-data
- Host: parse extras-json for well-known keys, merge stored binary data
- Telegram: download voice files via store-attachment-data, add duration
  to extras-json, add /file/bot to HTTP allowlist, voice-only placeholder
- Add reqwest multipart feature for Whisper API uploads
- 5 regression tests for transcription middleware

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: wire attachment processing into LLM pipeline with multimodal image support

Attachments on incoming messages are now augmented into user text via XML tags
before entering the turn system, and images with data are passed as multimodal
content parts (base64 data URIs) to LLM providers. This enables audio transcripts,
document text, and image content to reach the LLM without changes to ChatMessage
serialization or provider interfaces.

- Add src/agent/attachments.rs with augment_with_attachments() and 9 unit tests
- Add ContentPart/ImageUrl types to llm::provider with OpenAI-compatible serde
- Carry image_content_parts transiently on Turn (skipped in serialization)
- Update nearai_chat and rig_adapter to serialize multimodal content
- Add 3 e2e tests verifying attachments flow through the full agent loop

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: CI failures — formatting, version bumps, and Telegram voice test

- Fix cargo fmt formatting in attachments.rs, nearai_chat.rs, rig_adapter.rs,
  e2e_attachments.rs
- Bump channel registry versions 0.1.0 → 0.2.0 (discord, slack, telegram,
  whatsapp) to satisfy version-bump CI check
- Fix Telegram test_extract_attachments_voice: add missing required `duration`
  field to voice fixture JSON

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: bump WIT channel version to 0.3.0, fix Telegram voice test, add pre-commit hook

- Bump wit/channel.wit package version 0.2.0 → 0.3.0 (interface changed with
  store-attachment-data)
- Update WIT_CHANNEL_VERSION constant and registry wit_version fields to match
- Fix Telegram test_extract_attachments_voice: gate voice download behind
  #[cfg(target_arch = "wasm32")] so host functions aren't called in native tests,
  update assertions for generated filename and extras_json duration
- Add @0.3.0 linker stubs in wit_compat.rs
- Add .githooks/pre-commit hook that runs scripts/check-version-bumps.sh when
  WIT or extension sources are staged
- Symlink commit-msg regression hook into .githooks/

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: extract voice download from extract_attachments into handle_message

Move download_voice_file + store_attachment_data calls out of
extract_attachments into a separate download_and_store_voice function
called from handle_message. This keeps extract_attachments as a pure
data-mapping function with no host calls, making it fully testable
in native unit tests without #[cfg(target_arch)] gates.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR review comments — security, correctness, and code quality

Security fixes:
- Add path validation to read_attachments (restrict to /tmp/) preventing
  arbitrary file reads from compromised tools
- Escape XML special characters in attachment filenames, MIME types, and
  extracted text to prevent prompt injection via tag spoofing
- Percent-encode file_id in Telegram getFile URL to prevent query injection
- Clone SecretString directly instead of expose_secret().to_string()

Correctness fixes:
- Fix store_attachment_data overwrite accounting: subtract old entry size
  before adding new to prevent inflated totals and false rejections
- Use max(reported, stored_size) for attachment size accounting to prevent
  WASM channels from under-reporting size_bytes to bypass limits
- Add application/octet-stream to MIME allowlist (channels default unknown
  types to this)

Code quality:
- Extract send_response helper in Telegram, deduplicating on_respond and
  on_broadcast
- Rename misleading Discord test to test_parse_slash_command_interaction
- Fix .githooks/commit-msg to use relative symlink (portable across machines)

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add tool_upgrade command + fix TOCTOU in save_to path validation

Add `tool_upgrade` — a new extension management tool that automatically
detects and reinstalls WASM extensions with outdated WIT versions.
Preserves authentication secrets during upgrade. Supports upgrading a
single extension by name or all installed WASM tools/channels at once.

Fix TOCTOU in `validate_save_to_path`: validate the path *before*
creating parent directories, so traversal paths like `/tmp/../../etc/`
cannot cause filesystem mutations outside /tmp before being rejected.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: unify WIT package version to 0.3.0 across tool.wit and all capabilities

tool.wit and channel.wit share the `near:agent` package namespace, so they
must declare the same version. Bumps tool.wit from 0.2.0 to 0.3.0 and
updates all capabilities files and registry entries to match.

Fixes `cargo component build` failure: "package identifier near:agent@0.2.0
does not match previous package name of near:agent@0.3.0"

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: move WIT file comments after package declaration

WIT treats `//` comments before `package` as doc comments. When both
tool.wit and channel.wit had header comments, the parser rejected them
as "doc comments on multiple 'package' items". Move comments after the
package declaration in both files.

Also bumps tool registry versions to 0.2.0 to match the WIT 0.3.0 bump.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: display extension versions in gateway Extensions tab

Add version field to InstalledExtension and RegistryEntry types, pipe
through the web API (ExtensionInfo, RegistryEntryInfo), and render as
a badge in the gateway UI for both installed and available extensions.

For installed WASM extensions, version is read from the capabilities
file with a fallback to the registry entry when the local file has no
version (old installations). Bump all extension Cargo.toml and registry
JSON versions from 0.1.0 to 0.2.0 to keep them in sync.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add document text extraction middleware for PDF, Office, and text files

Extract text from document attachments (PDF, DOCX, PPTX, XLSX, RTF, plain text,
code files) so the LLM can reason about uploaded documents. Uses pdf-extract for
PDFs, zip+XML parsing for Office XML formats, and UTF-8 decode for text files.
Wired into the agent loop after transcription middleware.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: download document files in Telegram channel for text extraction

The DocumentExtractionMiddleware needs file bytes in the attachment `data`
field, but only voice files were being downloaded. Document attachments
(PDFs, DOCX, etc.) had empty `data` and a source_url with a credential
placeholder that only works inside the WASM host's http_request.

Add `download_and_store_documents()` that downloads non-voice, non-image,
non-audio attachments via the existing two-step getFile→download flow and
stores bytes via `store_attachment_data` for host-side extraction.

Also rename `download_voice_file` → `download_telegram_file` since it's
generic for any file_id.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: allow Office MIME types and increase file download limit for Telegram

Two issues preventing document extraction from Telegram:

1. PPTX/DOCX/XLSX MIME types (application/vnd.*) were dropped by the
   WASM host attachment allowlist — add application/vnd., application/msword,
   and application/rtf prefixes.

2. Telegram file downloads over 10 MB failed with "Response body too large" —
   set max_response_bytes to 20 MB in Telegram capabilities.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: report document extraction errors back to user instead of silently skipping

- Bump max_response_bytes to 50 MB for Telegram file downloads
- When document extraction fails (too large, download error, parse error),
  set extracted_text to a user-friendly error message instead of leaving it
  None. This ensures the LLM tells the user what went wrong.
- On Telegram download failure, set extracted_text with the error so the
  user sees feedback even when the file never reaches the extraction middleware.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: store extracted document text in workspace memory for search/recall

After document extraction succeeds, write the extracted text to workspace
memory at `documents/{date}/{filename}`. This enables:
- Full-text and semantic search over past uploaded documents
- Cross-conversation recall ("what did that PDF say?")
- Automatic chunking and embedding via the workspace pipeline

Documents are stored with metadata header (uploader, channel, date, MIME type).
Error messages (extraction failures) are not stored — only successful extractions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: CI failures — formatting, unused assignment warning

- Run cargo fmt on document_extraction and agent_loop modules
- Suppress unused_assignments warning on trace_llm_ref (used only
  behind #[cfg(feature = "libsql")])

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR review comments — security, correctness, and code quality

Security fixes:
- Remove SSRF-prone download() from DocumentExtractionMiddleware (#13)
- Sanitize filenames in workspace path to prevent directory traversal (#11)
- Pre-check file size before reading in WASM wrapper to prevent OOM (#2)
- Percent-encode file_id in Telegram source URLs (#7)

Correctness fixes:
- Clear image_content_parts on turn end to prevent memory leak (#1)
- Find first *successful* transcription instead of first overall (#3)
- Enforce data.len() size limit in document extraction (#10)
- Use UTF-8 safe truncation with char_indices() (#12)

Robustness & code quality:
- Add 120s timeout to OpenAI Whisper HTTP client (#5)
- Trim trailing slash from Whisper base_url (#6)
- Allow ~/.ironclaw/ paths in WASM wrapper (#8)
- Return error from on_broadcast in Slack/Discord/WhatsApp (#9)
- Fix doc comment in HTTP tool (#4)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: formatting — cargo fmt

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address latest PR review — doc comments, error messages, version bumps

- Fix DocumentExtractionMiddleware doc comment (no longer downloads from source_url)
- Fix error message: "no inline data" instead of "no download URL"
- Log error + fallback instead of silent unwrap_or_default on Whisper HTTP client
- Bump all capabilities.json versions from 0.1.0 to 0.2.0 to match Cargo.toml

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: remove unsupported profile: minimal from CI workflows [skip-regression-check]

dtolnay/rust-toolchain@stable does not accept the 'profile' input
(it was a parameter for the deprecated actions-rs/toolchain action).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: merge with latest main — resolve compilation errors and PR review nits

- Add version: None to RegistryEntry/InstalledExtension test constructors
- Fix MessageContent type mismatches in nearai_chat tests (String → MessageContent::Text)
- Fix .contains() calls on MessageContent — use .as_text().unwrap()
- Remove redundant trace_llm_ref = None assignment in test_rig
- Check data size before clone in document extraction to avoid unnecessary allocation

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
* feat: add AWS Bedrock LLM provider via native Converse API

* fix: use JSON parsing for tool result error detection instead of brittle substring matching

* refactor: extract duplicated inference config builder into helper function

* fix: address review feedback — safe casts, input validation, and tests

- Safe u32→i32 cast for max_tokens using try_from with clamp
- Remove brittle string-based error detection fallback for tool results
- Validate BEDROCK_CROSS_REGION against allowed values (us/eu/apac/global)
- Validate message list is non-empty before Converse API call
- Log when using default us-east-1 region
- Update llm_backend doc comment to list all backends
- Add tests for build_inference_config and empty message handling

* fix: persist AWS_PROFILE for Bedrock named profile auth

The wizard collected the profile name but only printed a hint to set
it manually. Now it saves to settings and writes AWS_PROFILE to the
bootstrap .env, consistent with how BEDROCK_REGION and other Bedrock
settings are persisted.

* feat: gate AWS Bedrock behind optional `bedrock` feature flag

The AWS SDK dependencies (aws-config, aws-sdk-bedrockruntime,
aws-smithy-types) require cmake and a C compiler to build aws-lc-sys.
Gate them behind an opt-in `bedrock` feature flag so default builds
are unaffected.

Build with: cargo build --features bedrock
All config, settings, and wizard code stays unconditional (no AWS deps)
so users can configure Bedrock even without the feature compiled — they
get a clear error at startup directing them to rebuild.

* fix: address review feedback and adapt Bedrock provider to registry architecture (takeover nearai#345)

- Resolve merge conflicts with main's registry-based provider system
- Add missing cache_creation_input_tokens/cache_read_input_tokens fields
- Add missing content_parts field in test ChatMessage
- Fix string literal type mismatches in wizard env_vars (.to_string())
- Remove non-functional bearer token auth (AWS_BEARER_TOKEN_BEDROCK) from
  wizard and documentation per reviewer feedback from @zmanian and @serrrfirat
- Remove stale BEDROCK_ACCESS_KEY proxy entry from provider table
- Update Bedrock provider to use is_bedrock string check (LlmBackend enum removed)
- Add bedrock_profile fallback from settings in config resolution

[skip-regression-check]

Co-Authored-By: cgorski <cgorski@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: use main's Cargo.lock as base to preserve dependency versions

Regenerating Cargo.lock from scratch caused transitive dependency version
drift that broke the html_to_markdown fixture test in CI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: bedrock config bugs — spurious warning, alias normalization, profile fallback

- Move is_bedrock check before unknown-backend warning to prevent
  spurious "unknown backend" log for bedrock users
- Normalize backend aliases ("aws", "aws_bedrock") to "bedrock" so
  the provider factory matches correctly
- Add settings.bedrock_profile fallback for AWS_PROFILE, consistent
  with region and cross_region resolution

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address Copilot review feedback — bearer token cleanup, stop_sequences, model dedup

- Remove stale bearer token refs from setup README and CHANGELOG
- Remove dead bedrock_api_key secret injection mapping
- Pass stop_sequences through to Bedrock InferenceConfiguration
- Remove "API key" from wizard menu description (bearer token removed)
- Skip duplicate LLM_MODEL write for bedrock backend in wizard
- Fix cargo fmt formatting

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address review feedback — async new(), remove LiteLLM entry, wizard fixes

- Remove dead LiteLLM-based bedrock entry from providers.json (native
  Converse API intercepts before registry lookup)
- Make BedrockProvider::new() async to avoid block_in_place panic in
  current_thread runtimes; propagate async to create_llm_provider,
  build_provider_chain, and init_llm
- Document CMake build prerequisite in docs/LLM_PROVIDERS.md
- Clear bedrock_profile when user selects "default credentials" in wizard
- Fix selected_model clearing to match established pattern (conditional
  on provider switch, not unconditional)
- Add regression tests for bedrock model preservation and profile clearing

Addresses review feedback from @zmanian on PR nearai#713.
Streaming support tracked in nearai#741.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address remaining review comments — CLAUDE.md backends, wizard UX

- Add `bedrock` to CLAUDE.md inline backend list (#10)
- Skip full setup re-run when keeping existing Bedrock config (#11)
- Clear stale bedrock_profile on empty named-profile input (#12)
- Add regression test for empty profile clearing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Chris Gorski <cgorski@cgorski.org>
Co-authored-by: cgorski <cgorski@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
* feat(engine-v2): mount-backend abstraction for per-project sandbox (Phase 1)

Adds the engine-side `MountBackend` trait + minimal `WorkspaceMounts` registry
and a host-side bridge interceptor that routes sandbox-eligible tool calls
(`file_read`, `file_write`, `list_dir`, `apply_patch`, `shell`) through a
backend when their path argument starts with `/project/`. Default behavior is
unchanged: until `EffectBridgeAdapter::set_workspace_mounts(Some(...))` is
called (Phase 6), the interception path is dormant.

This is the first phase of the per-project sandbox plan
(`docs/plans/2026-04-10-engine-v2-sandbox.md`) and a deliberately small subset
of the unified Workspace VFS proposed in nearai#1894 — just enough
abstraction so the sandbox can be a `MountBackend` rather than a special case
in the bridge. When nearai#1894's full mount table lands, the sandbox backend slots
in unchanged.

Engine crate (`crates/ironclaw_engine/src/workspace/`):
- `mount.rs` — `MountBackend` trait, `MountError` (NotFound / InvalidPath /
  PermissionDenied / Io / Tool / Backend / Unsupported), `DirEntry`,
  `EntryKind`, `ShellOutput`
- `filesystem.rs` — `FilesystemBackend`: passthrough host-fs implementation
  with two-layer path validation (lexical reject of absolute / `..`, then
  symlink-escape canonicalization). `read`/`write`/`list` fully implemented;
  `patch`/`shell` return `Unsupported` so the bridge falls through to the
  host tool until Phase 5
- `registry.rs` — `WorkspaceMounts` per-project registry with lazy
  `ProjectMountFactory`, longest-prefix-match resolution, cached and
  invalidatable

Bridge (`src/bridge/sandbox/`):
- `intercept.rs` — `maybe_intercept` and `SANDBOX_TOOL_NAMES`. Returns
  `Handled(json)` on a successful backend dispatch, `FellThrough` for
  non-sandbox tools, host paths, missing path params, or `Unsupported`
  backend ops
- `effect_adapter.rs` — `workspace_mounts` field + `set_workspace_mounts`
  setter; interception block in `execute_action_internal` right before
  `execute_tool_with_safety`, gated on the optional mount table

Tests (31 new):
- 17 engine workspace unit tests covering trait error mapping, path safety
  (lexical + symlink), longest-prefix routing, and lazy factory caching
- 9 bridge sandbox unit tests including `intercept_actually_dispatches_into_backend`
  (counting backend) which proves the interceptor reaches the backend
- 5 integration tests in `tests/engine_v2_sandbox_integration.rs` driving
  `EffectBridgeAdapter::execute_action()` end-to-end per the
  "Test Through the Caller" rule (`.claude/rules/testing.md`), including
  a host-path-falls-through test that asserts the sandbox tempdir was
  not touched, and a `..`-escape test that verifies no `/etc/passwd`
  content leaks even after safety-layer redaction

Drive-by: feature-gate two pre-existing dead-code helpers in
`crates/ironclaw_skills/src/parser.rs` on `#[cfg(feature = "registry")]` to
match their only call site, fixing a pre-existing clippy warning that blocked
the workspace's `-D warnings` policy when `ironclaw_skills` is built with
`default-features = false` (as the engine crate does).

Verification:
- `cargo fmt --check` clean
- `cargo clippy --all --benches --tests --examples --all-features` zero warnings
- 31 / 31 new tests passing; no existing tests broken

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine-v2): per-project sandbox — Phases 2–7 + live Docker e2e test

Completes the per-project sandbox plan (docs/plans/2026-04-10-engine-v2-sandbox.md
Phases 2–7), building on Phase 1's mount-backend abstraction (nearai#2211).

Phase 2 — Project workspace folder:
- `Project.workspace_path: Option<PathBuf>` field + `with_workspace_path()`
- Host-side `project_workspace_path()`, `ensure_project_workspace_dir()` (creates
  `~/.ironclaw/projects/<id>/` mode 0700, idempotent)
- `FilesystemMountFactory` taking a `ProjectPathResolver` closure (decoupled from
  `Store`); wired into `EffectBridgeAdapter` via `set_workspace_mounts()`

Phase 3 — Standalone daemon binary:
- `src/bin/sandbox_daemon.rs` — NDJSON over stdin/stdout, health/shutdown/execute_tool
- Constructs ReadFileTool/WriteFileTool/ListDirTool/ApplyPatchTool/ShellTool with
  `base_dir=/project` (override via `IRONCLAW_SANDBOX_BASE_DIR`)

Phase 4 — Dockerfile.sandbox:
- Multi-stage build: rust-slim builder (+ python3 for pyo3) compiles sandbox_daemon;
  debian-slim runtime with tini PID 1, common build tools, `/project` mount target

Phase 5 — ProjectSandboxManager + ContainerizedFilesystemBackend:
- protocol.rs: Request/Response/RpcError matching daemon wire format
- transport.rs: `SandboxTransport` trait (seam for testing without Docker)
- containerized_backend.rs: `ContainerizedFilesystemBackend` impls `MountBackend`,
  translates relative→`/project/<rel>`, maps tool-error→MountError
- docker_transport.rs: real bollard exec session, serialized Mutex, lazy reconnect
- lifecycle.rs: deterministic `ironclaw-sandbox-<pid>` naming, ensure_running/stop/remove
- manager.rs: `ProjectSandboxManager` per-project transport cache

Phase 6 — Router gating on ENGINE_V2_SANDBOX:
- `engine_v2_sandbox_enabled()` helper (truthy: 1/true/yes/on)
- Router selects `ContainerizedMountFactory` when enabled + Docker reachable;
  falls back to `FilesystemMountFactory` with warning otherwise

Live e2e bugs caught and fixed:
- Shell without explicit `workdir` defaulted to host (not sandbox); fixed by
  defaulting to `/project/` in `extract_path_param`
- `ContainerizedFilesystemBackend::shell` parsed `stdout`/`stderr` but host
  ShellTool returns merged `output` field; fixed with fallback key lookup
- SANDBOX_TOOL_NAMES only had v2 names (`file_read`/`file_write`) but host
  registry uses v1 names (`read_file`/`write_file`); added both aliases

Tests (62 sandbox-related, all green):
- 27 bridge sandbox unit tests (intercept, workspace_path, factory, protocol,
  lifecycle, containerized_backend with ScriptedTransport mock)
- 7 containerized-backend tests (including 2 regression tests for the shell bugs)
- 5 engine v2 sandbox integration tests (EffectBridgeAdapter end-to-end)
- 5 daemon binary smoke tests (real subprocess + NDJSON I/O)
- 17 engine workspace unit tests
- 1 live Docker e2e test: agent clones nearai/ironclaw into sandbox, renames
  to megaclaw via sed, verifies with grep — 70s, $0.09, recorded trace committed

Verification:
- `cargo fmt --check` clean
- `cargo clippy --all --benches --tests --examples --all-features` zero warnings
- All 62 sandbox tests passing; no existing tests broken

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: replace .expect() with Result in DockerTransport::ensure_session

CI's no-panics checker flagged the .expect("just inserted") in production
code. Replace with .ok_or_else() returning MountError::Backend.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: multi-tenant project paths + unify sandbox env var with v1

Two issues addressed:

1. Project workspace paths now namespace by user_id:
   `~/.ironclaw/projects/<user_id>/<project_id>/` instead of
   `~/.ironclaw/projects/<project_id>/`. Prevents filesystem collisions
   in multi-tenant deployments where two users could theoretically have
   the same project UUID.

2. Sandbox enablement now reads `SANDBOX_ENABLED` (same env var as v1
   sandbox) in addition to `ENGINE_V2_SANDBOX`. Either being truthy
   enables the per-project sandbox. This means a single flag governs
   sandbox behavior regardless of engine version, while the v2-specific
   override remains available for transitional setups.

Tests: 30 bridge sandbox unit tests passing (added multi-tenant path
tests + env var combination tests).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review — TOCTOU race, shell env passthrough, canonicalize guard

Three issues flagged by the code review bot on nearai#2211:

1. TOCTOU race in WorkspaceMounts::resolve (HIGH): Added double-checked
   locking — re-check the cache after acquiring the write lock so two
   threads racing on the same project's first access don't both call
   factory.build(). The second thread finds the insert from the first.

2. Shell intercept ignores env parameter (MEDIUM): The shell arm in
   maybe_intercept was passing HashMap::new() instead of forwarding
   the tool call's env map. Fixed to parse parameters["env"] and pass
   it through to backend.shell().

3. Canonicalization fails when root doesn't exist (MEDIUM): When
   self.root hasn't been created yet (first write to a new project),
   canonicalize_under_root would walk up to a real ancestor and the
   starts_with check against the non-existent root would always fail.
   Now skips canonicalization entirely when root doesn't exist — lexical
   safety is already guaranteed by safe_join.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review round 2 — apply_patch schema, content validation, dir perms, docs

- Fix apply_patch schema mismatch: MountBackend::patch now takes
  (old_string, new_string, replace_all) matching ApplyPatchTool's
  actual contract. Previously sent {patch: diff} which would fail
  with invalid_params in the containerized daemon.
- Validate file_write content param: return error instead of silently
  writing empty string when content is missing.
- Log stderr frames from sandbox daemon at debug! instead of silently
  discarding them in docker_transport StreamReader.
- Tighten permissions on intermediate directories created by
  ensure_project_workspace_dir (projects/, <user_id>/) to 0o700,
  not just the leaf.
- Fix stale module doc in sandbox/mod.rs (referenced "Phase 5 will
  add" but all phases shipped).
- Fix doc path mismatch: workspace path is <user_id>/<project_id>/,
  not <project_id>/ (workspace_path.rs, CLAUDE.md, design plan).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review round 3 — symlink safety, visibility, debug logging

- Close TOCTOU window in canonicalize_under_root: re-canonicalize and
  verify containment when the reassembled path exists on disk
- Fix list_dir_recursive: use symlink_metadata (lstat) so symlinks are
  detected instead of followed; validate directories against root before
  recursive traversal
- Tighten is_mountable_path to /project/, /memory/, /home/ prefixes
  instead of any absolute path (defense-in-depth)
- Narrow sandbox module visibility to pub(crate) and remove unused
  pub use re-exports
- Remove concrete types (FilesystemBackend, DirEntry, EntryKind,
  ShellOutput) from engine crate top-level re-exports; access via
  ironclaw_engine::workspace:: module path
- Add debug! tracing to sandbox intercept routing decisions
- Add read_file/write_file v1 aliases to daemon SUPPORTED_TOOLS health
  response
- Remove developer-local path from sandbox mod.rs doc comment
- Merge staging to fix CI (user_timezone field on ThreadExecutionContext)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review round 4 — safety validation, network isolation, binary writes

- Add pre-intercept safety param validation so sandbox-dispatched calls
  go through the same checks as host-dispatched calls (#1)
- Set network_mode: "none" on sandbox containers to prevent outbound
  network access (#3)
- Reject binary content in containerized write instead of silently
  corrupting via from_utf8_lossy (#5)
- Cap list_dir depth to 10 to prevent unbounded traversal (#8)
- Change container creation log from info! to debug! to avoid breaking
  REPL/TUI output (#10)
- Make is_truthy case-insensitive so SANDBOX_ENABLED=True works (#11)
- Return error instead of unwrap_or_default for missing container ID (#12)
- Propagate set_permissions errors instead of silently ignoring (#13)
- Return error for missing daemon output key instead of defaulting to
  empty object (#14)
- Add env mutex guard in sandbox_live_e2e test (#15)
- Fix rustfmt formatting for let-chain in canonicalize_under_root

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review round 5 — path traversal, error types, tests

Security fixes:
- Sanitize user_id in workspace path to prevent directory traversal via
  malicious user IDs containing `..` or `/`
- Add Component::ParentDir check in ContainerizedFilesystemBackend::container_path
  matching the defense-in-depth approach of FilesystemBackend::safe_join

Correctness:
- Use MountError::Tool instead of MountError::InvalidPath for missing
  tool parameters (content, old_string, new_string) — fixes confusing
  LLM-visible error messages
- Fix clippy sort_by_key suggestion in registry.rs

Cleanup:
- Remove spurious Notify import and dead _notify_link function

New tests:
- ContainerizedFilesystemBackend path traversal rejection (read + write)
- container_path unit tests for safe and unsafe paths
- Adversarial user_id test in workspace_path
- Daemon-side path traversal test in sandbox_daemon_smoke

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review round 6 — param normalization, error types, edge cases

- Normalize sandbox params via prepare_tool_params() before validation,
  matching the host execution path (fixes inconsistent validation)
- Return ToolError::InvalidParameters instead of EngineError::Effect for
  sandbox param validation failures (consistent error surface)
- ensure_dir checks path.is_dir() not path.exists() (rejects files)
- Empty user_id returns "_anonymous" sentinel instead of empty hex string
  that would drop the tenant namespace via PathBuf::join("")
- Restore ENGINE_V2_SANDBOX env var after sandbox live E2E test
- Tighten is_mountable_path to /project/ only (no mounts for /memory/
  or /home/ yet)
- Add v1 tool name aliases (read_file, write_file) to SUPPORTED_TOOLS

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: unify sandbox env var — remove ENGINE_V2_SANDBOX, use SANDBOX_ENABLED only

Single env var controls sandboxing for both engine versions. The
transitional ENGINE_V2_SANDBOX override is removed from code, tests,
docs, and Dockerfile.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: double-checked locking in transport_for, explicit stdin close in smoke test

- ProjectSandboxManager::transport_for no longer holds the mutex across
  the Docker ensure_running await. Uses double-checked locking so
  concurrent projects initialize in parallel.
- sandbox_daemon_smoke: explicitly take() stdin before wait_with_output
  so EOF is sent even without a shutdown request.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review — network mode, error types, race, protocol dedup

- Change sandbox container network_mode from "none" to default bridge
  so git clone / cargo build / pip install work inside the container
- Fix binary content rejection to use MountError::Tool instead of
  MountError::InvalidPath (semantic mismatch)
- Fix list depth: use actual depth value instead of depth.max(1)
- Fix orphan container race in transport_for by holding lock across
  container creation instead of double-checked locking
- Deduplicate protocol types: daemon now imports from shared
  bridge::sandbox::protocol instead of defining its own copies
- Make bridge::sandbox pub (narrow exposure: only protocol and
  workspace_path sub-modules are pub)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: update plan doc — sandbox uses bridge networking, not network_mode=none

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…3920)

* Implement installed WASM hook runtime

Adds crates/ironclaw_hooks/docs/threat-model-wasm.md and follows the reviewed design ack: 1) module bytes are resolved, digest-cached, and compiled in the tool-WASM style while reusing its resource limiter; 2) each invocation gets a fresh wasmtime Store; 3) the ABI is a wasmtime::Linker surface, not wit-bindgen; 4) host-import sink shims enforce call, patch-byte, observer-fact, and decision budgets.

* Harden WASM hook string and metadata budgets

* fix(hooks): validate WASM hook ABI at install time (serrrfirat #3 on PR nearai#3634)

Address serrrfirat MEDIUM finding #3: `WasmHookRuntime::prepare()` compiled
and cached module bytes but did not validate imports or the requested
export. ABI mismatches (unsupported import, missing export, wrong export
signature) were deferred to first live dispatch — and the prior
`wasm_unsupported_host_import_fails_closed` test codified that a
bad-import module would install successfully and only fail closed at
invocation. Malformed untrusted modules should never reach live traffic.

Changes:
- `prepare()` derives the target hook point from `request.kind`, then
  runs `validate_module_abi()`: scratch-instantiate the module against
  the point-specific linker (catches unsupported / wrong-type imports)
  and resolve the typed export `() -> ()` (catches missing export and
  wrong signature). Failures surface as new
  `WasmHookRuntimeError::InvalidImports` or existing
  `WasmHookRuntimeError::InvalidExport`, both of which bubble up as
  `HookError::RegistryConstruction` from the registrar.
- `wasm_point_for_kind(HookManifestKind)` helper centralizes the
  kind → wasm-point mapping; the previous `execute_*` paths can share
  it in a follow-up but kept inline for now to minimize churn.

Tests:
- `wasm_unsupported_host_import_is_rejected_at_install_time`: replaces
  the prior test that codified late-failure behavior; asserts the
  registrar returns `RegistryConstruction` citing the bad import.
- `wasm_missing_export_is_rejected_at_install_time`: new module that
  compiles but lacks the manifest-declared export; same install-time
  rejection.

* fix(hooks): address henrypark133 must-fix #1, #2, #3 on PR nearai#3634

Three items from the 5-15 review:

**#1 (must-fix) Extract ironclaw_wasm_limiter micro-crate**
Replace `#[path = "../../../ironclaw_wasm/src/limiter.rs"]` cross-crate
file import with a proper Cargo edge. The 111-line `WasmResourceLimiter`
moves into a new `crates/ironclaw_wasm_limiter` micro-crate that both
`ironclaw_wasm` and `ironclaw_hooks` depend on. The architecture rule
forbidding `ironclaw_hooks -> ironclaw_wasm` is preserved (the new
crate sits below both consumers and pulls in only `wasmtime` +
`tracing`); `cargo check`, `cargo doc`, and architecture-linting tests
now see the edge, and the file can't be moved out from under one of
the consumers silently.

Mechanical changes:
- new `crates/ironclaw_wasm_limiter/` (Cargo.toml + src/lib.rs with the
  type exposed as `pub` instead of `pub(crate)`)
- workspace `members` entry added
- `crates/ironclaw_wasm/src/limiter.rs` deleted
- `crates/ironclaw_wasm/src/lib.rs`: `mod limiter` removed
- `crates/ironclaw_wasm/src/store.rs`: import switched to
  `ironclaw_wasm_limiter::WasmResourceLimiter`
- `crates/ironclaw_wasm/Cargo.toml`: dep added
- `crates/ironclaw_hooks/Cargo.toml`: dep added
- `crates/ironclaw_hooks/src/wasm/runtime.rs`: `#[path = ...]` block
  removed; import switched to the crate

**#2 + #3 (must-fix) Dead WASM arms in dispatch**
`run_before_capability_hook`, `run_before_prompt_hook`, and
`run_observer_hook` each had an early-return guard that dispatched
WASM hooks with `catch_unwind` + timeout, then ALSO had a matching
WASM arm in the inner `match` that ran without those protections. The
prompt-path arm additionally swallowed `WasmHookFailure` via `|_| ()`,
making the must-fix #2 problem worse on that path specifically.

If a future refactor removed any of the early-return guards, those
inner arms would silently take over and drop panic isolation, deadline
enforcement, AND (for prompts) the failure category. Replaced each
inner arm with `unreachable!()` carrying a comment that explains
why the arm exists and references the early-return guard above it.
A future refactor that removes the guard will now trip the
`unreachable!` at first call instead of silently degrading.

All 154 hooks lib + 29 reborn integration tests still pass.

* fix(hooks): plumb context to WASM hooks + runtime hardening

Critical #1 on PR nearai#3634: WASM hooks previously received no context. The
`execute_*` entry points dropped the `&BeforeCapabilityHookContext` /
`&BeforePromptHookContext` / `&ObserverHookContext` value and invoked
the guest export with `()`, so a WASM gate could never decide based on
the capability name, tenant, provider, or other dispatch-time facts. Add
an `ic:hooks/context@1` host-import module exposing two read-only
calls — `ctx_size() -> i32` and `ctx_read(ptr, len) -> i32` — backed by
a JSON-serialized blob the dispatcher writes per-invocation into the
fresh store. Modules that don't import these continue to link; modules
that do import them get a stable, non-empty payload to read. An
integration test (`wasm_before_capability_hook_reads_context_blob`)
asserts the contract end-to-end: a guest that fails to read a non-empty
blob traps before its `deny` call.

Also rolls up the other reviewer-flagged WASM runtime issues, all of
which touch `wasm/runtime.rs`:

HIGH #2: epoch-tick background thread now holds a shutdown
`AtomicBool` and joins on `Drop`. Previously it looped forever and
leaked an Engine clone on every runtime drop.

MED #4: compiled-module cache is now an `lru::LruCache` bounded by
`MODULE_CACHE_CAPACITY = 128`. Replaces the unbounded `HashMap`.

MED #7: `prepare()` no longer compiles under the cache lock. Fast
path reads from LRU under a brief lock; slow path compiles outside
the lock and re-checks on insert to avoid the TOCTOU window where
two concurrent installs of the same module both compile.

Bug #9: post-call `deadline_exceeded()` re-check on the Ok branch
is gone. wasmtime epoch-interrupt is the authoritative wall-clock
signal; an Ok return is no longer reclassified as a timeout because
the wall ticked over during host-side return.

Bug #10: `add_milestone_metadata` returns a distinct
"metadata value exceeds the u32 byte-length ceiling" error when the
guest-supplied `value.len()` overflows u32, instead of misreporting it
as "exceeded total prompt-patch byte budget".

Existing integration tests for WASM hooks are also re-wired through
`HookRegistrar::with_verified_grants` so the grants-store gate added in
the foundation-01 merge stops failing the pre-existing fixtures.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): run WASM hooks on the blocking pool

HIGH #3 on PR nearai#3634: `tokio::time::timeout` does NOT cancel synchronous
wasmtime execution. The previous code awaited a `catch_unwind(async { h.evaluate(ctx) })`
future whose body completed in one poll, so the timeout could only fire
*around* the WASM call rather than against it; a hook that wedged inside
wasmtime simply pinned the calling tokio task.

Route gate, prompt, and observer WASM dispatch paths through
`tokio::task::spawn_blocking` via a shared `run_wasm_blocking` helper.
The outer `tokio::time::timeout` now governs the JoinHandle, so a stuck
blocking task stops blocking the dispatcher's caller; the wasmtime
epoch interrupt configured in the runtime (10 ms tick) is the
authoritative in-WASM wall-clock cancel signal. JoinError (panic in
the blocking task) maps to `FailureCategory::Panic`, matching the
pre-existing semantics for synchronous panics caught via
`catch_unwind`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(hooks): O(1) hook-id lookup via side index

Finding #8 on PR nearai#3634: `set_priority`, `poison`, `is_poisoned`, and
`contains_hook` all did full-registry scans over every binding at every
point. Each is called per-dispatch (poison-checks on the snapshot loop
in particular), so the cost is `O(registered_hooks)` per
`(installed_hook, registered_hook)` pair.

Maintain a denormalized `HashMap<HookId, (HookPointSpec, usize)>` side
index in lock-step with `by_point` so every per-hook-id operation
becomes a single hash lookup + a direct vec indexed access. The
duplicate-id rejection in `insert` now reads from the side index too,
turning what used to be a flat-map scan into a `HashMap::contains_key`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(hooks): wall-clock timeout, observer memory, limiter rollback, registrar happy path

Round out the test set for the WASM hook execution path:

#11 / #12: gate + observer wall-clock timeout. The pre-fix dispatcher
ran wasmtime synchronously on the executor, so the outer
`tokio::time::timeout` `Err(_elapsed)` arm was effectively unreachable.
Now that WASM execution runs on the blocking pool, the timeout actually
fires; the new tests give the wasm budget headroom (1B fuel, 5s wall)
and the dispatcher a 20 ms timeout, then assert the failure
classification (FailClosed for gate, FailIsolated for observer).

#13: observer memory exhaustion. Mirrors
`wasm_memory_exhaustion_fails_closed_for_gate` against the observer
dispatch path so the FailIsolated branch of the failure matrix has
explicit memory coverage, not just fuel/wall.

#15: `WasmResourceLimiter::memory_grow_failed` rollback. Stages an
approved grow, simulates the OS-level grow failing, and asserts a
subsequent grow of the full ceiling succeeds — the inflated
`memory_used` from the failed attempt must be released.

#16: registrar WASM happy path. Companion to the existing
`install_wasm_body_requires_runtime` negative case: a valid module
installs, the binding is visible via the public registry accessor, and
is not pre-poisoned.

#14 (`add_milestone_metadata` happy path) is intentionally omitted —
the BeforePrompt dispatch path is currently unreachable due to a
pre-existing manifest-vs-registry scope conflict (`OwnCapabilities` is
the only valid `BeforePrompt` scope per manifest validation, but the
registry rejects `OwnCapabilities` at `BeforePrompt` because the point
has no provider context). That contradiction sits outside this PR's
scope; flagging for a follow-up.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(hooks): typed WASM version material, reconcile design doc

LOW #20 on PR nearai#3634: extract the
`{extension_version}+wasm:{module_digest_hex}` concatenation into a
`WasmVersionMaterial` newtype with a single `Display` impl. The
identity material no longer floats free as a stringly-typed argument
inside the registrar.

Reconcile `docs/successors/02-wasm-runtime.md` with the implementation:

- Spell out that wall-clock cancellation depends on the
  `tokio::time::timeout(tokio::task::spawn_blocking(...))` pair, and
  explain why a bare timeout over a synchronous wasmtime call cannot
  actually cancel.
- Define `FailIsolated` and `FailClosed` as `FailureDisposition`
  values, distinct from the older `HookFailureMode::{FailOpen,
  FailClosed}` policy switch that applies to predicates.
- Clarify the generic `evaluate` export contract — name is whatever
  the manifest declares, signature is `(): ()`, context arrives
  through the new `ic:hooks/context@1` host imports — and note the
  intentional divergence from `WitToolRuntime`'s hardcoded interface.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): drop .expect() in WASM module cache capacity

Pre-commit no-panics CI flagged the .expect() on the LruCache capacity.
Move the validity check to a const match, so the NonZeroUsize is fixed at
compile time and the no-panics regex is satisfied.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): use HookLocalId::new after newtype privatization

The newtype-privatization landed in reborn-integration after the
hooks-fu-wasm-runtime branch's WASM scaffolding tests were written;
update the affected test/registrar sites to use HookLocalId::new
instead of the now-private tuple constructor.

* style: cargo fmt after newtype-privatization fixups

* test(hooks): ignore 3 BeforePrompt WASM tests with manifest/registry conflict

These tests were failing on the original branch tip too (verified against
origin/hooks-fu-wasm-runtime @ 571efdf). The Installed-tier BeforePrompt
WASM install path has no valid scope today:
  - OwnCapabilities is rejected by the registry C3 check (finding #2 on
    PR nearai#3573) since BeforePrompt has no per-capability invocation
    context.
  - SameTenant is rejected by manifest validation ("cannot combine
    scope = same_tenant with kind = before_prompt").

The budget-overflow paths these tests exercise are point-agnostic; the
follow-up is to either rewrite the helper to install through
BeforeCapability or add a Global manifest scope. Tracked as a deferred
item on the new PR.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…#3122)

* feat(web): support externally-provided tools in Responses API

Lets callers of `/v1/responses` (and `/api/v1/responses`) declare their
own `function`-typed tools and feed back results via
`function_call_output` items, matching the OpenAI Responses wire shape.

Since IronClaw's engine has no per-request tool surface, integration
happens at the prompt level: the catalog is rendered as
`<external-tools>` in the user message and the agent signals a call by
ending its response with a fenced ```` ```tool_call ```` block. When
that fence is recognised, the reply is split into a leading `Message`
plus a `function_call` `ResponseOutputItem`.

Validation rejects unsupported tool types (`web_search`, `file_search`,
`code_interpreter`) and tools missing `name` with 400, with two new
integration tests covering both paths.

* refactor(responses-api): switch external tools to engine v2 native path

Replace the prompt-level fence protocol from PR nearai#3122 with engine v2
native tool calls: caller-supplied `tools[]` are surfaced as real
LLM-callable actions, the engine pauses with `ResumeKind::External`
when one is invoked, and the bridge router projects the pause to a
new `AppEvent::ExternalToolCall` carrying the OpenAI-shaped
`function_call` wire fields.

The integration is small because v2 already has the right primitives:

- `ResumeKind::External { callback_id }` and
  `GateResolution::ExternalCallback { payload }` already existed for
  OAuth-style callbacks.
- `agent_loop.rs:1480` already routes Responses API messages to
  `handle_with_engine` when `ENGINE_V2=true`, so no v2 migration of
  the endpoint itself is needed.
- `EffectBridgeAdapter::execute_action` is the single chokepoint
  where caller tools can be detected before they reach the dispatch
  pipeline.

Changes:

- New `src/bridge/external_tools.rs` (`ExternalToolCatalog`) — per-thread
  registry of caller-supplied `ActionDef`s, plus the `ext_tool:`
  callback-id helpers used to disambiguate external-tool pauses from
  OAuth/pairing pauses (which also use `ResumeKind::External`).
- `EffectBridgeAdapter` consults the catalog: any name in it is
  short-circuited to a `GatePaused { resume_kind: External {
  callback_id: ext_tool:<call_id> } }` before any registry dispatch,
  and `available_action_inventory` merges the catalog into the
  LLM-visible action surface (internal beats external on collision).
- `Submission::ExternalCallback` gains an optional `payload` field;
  `bridge::handle_external_callback` plumbs it into
  `GateResolution::ExternalCallback { payload }`. Fallback predicate
  `gate_resume_is_external` lets non-auth External pauses (i.e.
  caller-tool resumes) resolve through the same handler.
- New `AppEvent::ExternalToolCall` projected by `notify_pending_gate`
  when a paused gate carries an `ext_tool:` callback id; OAuth/
  pairing flows keep flowing through the existing `GateRequired`
  channel.
- `responses_api.rs` is gutted of the prompt rendering and fence
  parsing (`render_external_tools_preamble`, `extract_trailing_tool_call`,
  `parse_external_tool_call`, `ParsedToolCall`, `external_tool_names`
  accumulator field, and the `TOOL_CALL_FENCE` constants). The handler
  now: rejects `tools[]` when `ENGINE_V2=false`, registers caller
  tools in the catalog under the resolved thread id, detects resume
  requests (`previous_response_id` + `function_call_output` items in
  `input`) and submits them as `Submission::ExternalCallback` with
  the outputs as the resolution payload, and surfaces
  `AppEvent::ExternalToolCall` as a `function_call` `ResponseOutputItem`
  in both streaming (`output_item.added`+`done`) and non-streaming.
- All existing OAuth/pairing `ExternalCallback` constructors updated
  to pass `payload: None` (no behaviour change).
- Fence-protocol unit tests removed; replaced with coverage for the
  new `responses_tools_to_action_defs` converter and the accumulator's
  `ExternalToolCall` arm.

Existing 9 integration tests in `tests/responses_api_path_prefix.rs`
still pass.

Note for reviewers:
- The accumulator-side text response no longer tries to split the
  reply on a fenced `tool_call` block. The wire shape that callers
  receive for caller-tool invocations is purely event-driven now.
- Internal vs external collision is handled silently by the dedup in
  `available_action_inventory` (internal wins). A request-time
  rejection for shadowing names is a follow-up — the current behavior
  is safe (the LLM only sees the internal version) but could surprise
  a caller who expects their tool to run.

* test(responses-api): cover ENGINE_V2-off and resume-without-pending-gate

Two new integration tests for behaviours added by the engine-native
external-tool refactor:

- `external_tools_rejected_when_engine_v2_disabled`: a request with
  caller-supplied `tools[]` while `ENGINE_V2` is off must 400 with a
  message naming the flag, not silently fall through.
- `resume_without_pending_gate_returns_400`: a request with
  `function_call_output` items and a `previous_response_id` that
  doesn't correspond to a live external-tool gate must 400, not start
  a fresh turn against the (unrelated) thread.

Both tests drive the full router (`start_test_server` + bearer auth)
per `.claude/rules/testing.md` "Test Through the Caller".

* test(responses-api): integration tests + drop unsafe env mutation

Three groups of changes:

1. **Drop unsafe env-var mutation in tests.** `responses_api.rs` no
   longer reads `ENGINE_V2` directly: it keys off the presence of the
   live `ExternalToolCatalog` (initialized by `init_engine`) as the
   "engine v2 is up" signal. The path-prefix test that exercises the
   no-engine branch no longer needs `unsafe { std::env::remove_var }`
   — the absence of `init_engine` in `TestGatewayBuilder` is what
   makes the catalog absent, which is what makes the request reject.

2. **Engine-tier integration tests** (`tests/e2e_responses_api_external_tools.rs`).
   Drives engine v2 via the existing `TestRigBuilder` + `TraceLlm`
   replay infrastructure rather than spinning up an HTTP gateway:

   - `catalog_isolates_by_thread_id` — register-under-A doesn't bleed
     into B.
   - `catalog_register_overwrites_not_merges` — Responses API contract
     is "each request restates the full tools[]"; the catalog must
     replace, not merge.
   - `catalog_sweep_evicts_only_stale_entries` — TTL backstop.
   - `catalog_clear_on_terminal_state_explicit` — what we want; the
     test name flags that the production hook is missing.
   - `catalog_handles_concurrent_registrations` — 32 concurrent
     register-then-contains tasks; protects per-thread isolation
     under contention.
   - `callback_id_disambiguates_external_from_oauth` — `ext_tool:`
     vs `pairing:` prefix is the single bit that routes a paused gate
     to `AppEvent::ExternalToolCall` vs `AppEvent::GateRequired`. If
     the prefix invariant breaks, the wrong UI renders.

   Two more tests deliberately `#[ignore]` and document concrete bugs
   the implementation has today:

   - `engine_pauses_when_llm_calls_registered_external_tool` — running
     it surfaces "engine never paused on external tool". Confirms the
     **thread-id mismatch** bug: the catalog is keyed by engine
     `ThreadId`, but the responses_api handler registers under a
     separately-generated UUID before the engine spawns the thread.
   - `round_trip_resume_payload_reaches_llm` — the load-bearing
     end-to-end test. Documents the **resume-payload-not-materialised**
     bug: `bridge::router::resolve_gate` uses `pending.resume_output`
     for `ExternalCallback` resolutions and ignores the payload, so
     caller-supplied tool outputs never reach the LLM's context.
   - `external_collision_with_registry_action_is_rejected` — stub
     asserting the desired validation behaviour for caller tool
     names that shadow registry actions; today silently accepted,
     and the catalog-wins-in-dispatch ordering means the LLM thinks
     it called the internal tool but actually ran caller code.

   The two #[ignore] tests are the deliberate failure documentation:
   running them with `--ignored` panics with messages naming the
   underlying gap. The fixes go on a follow-up commit.

3. **`TestRigBuilder.send_external_callback_with_payload`** — new
   helper that mirrors the OAuth `send_external_callback` but carries
   a JSON payload. Used by the round-trip test; the existing
   payload-less variant kept for OAuth/pairing tests.

Quality gates: `cargo fmt`, `cargo clippy --all --tests --all-features`
clean. 14 of 17 tests pass; the 3 ignored ones are deliberate
documentation of the gaps.

* fix(responses-api): close 4 bugs surfaced by integration tests

The engine-native external-tools path landed in 44135ca had four
real bugs surfaced by the integration tests in a2c13ee. This commit
fixes all four; every previously-ignored test now passes.

**Bug 1 — Thread-id mismatch.** `responses_api.rs` registers tools in
the catalog under a `thread_uuid` it generates from `previous_response_id`
(or freshly), but `ConversationManager::handle_user_message` creates
the engine's *actual* `ThreadId` internally. The catalog entries were
under a UUID the engine never executed in.

Fix: `ExternalToolCatalog::transfer(from, to)` and a hook in
`bridge::handle_with_engine_inner` that calls it after the engine
returns the spawned `ThreadId`. The handler-supplied conversation
scope UUID is rebound onto the actual ThreadId before the LLM call
lands, so `EffectBridgeAdapter::execute_action`'s catalog check
finds the registered tools.

**Bug 2 — Resume payload never materialised.** `bridge::router::resolve_gate`'s
`GateResolution::ExternalCallback` branch only consulted
`pending.resume_output` to construct the resumed `ActionResult`.
`EffectBridgeAdapter::execute_action` sets that to `None` for
caller-tool gates (the output isn't known at gate-fire time), so
the resume fell through to `execute_pending_gate_action`, which
re-ran the original action — re-pausing forever. The caller's tool
output (passed via `GateResolution::ExternalCallback { payload }`)
was dropped on the floor.

Fix: special-case `ext_tool:` callback ids in `resolve_gate`'s
ExternalCallback branch. Extract the matching output from the
payload (Responses API wire shape: `{ outputs: [{ call_id, output }] }`)
via the new `extract_external_tool_output` helper, synthesise an
`ActionResult`-shaped ThreadMessage, and resume the thread directly.
OAuth/pairing flows (which use `pairing:` callback ids and don't
carry an output payload) keep the original `pending.resume_output`
path unchanged.

**Bug 3 — Internal/external collision in dispatch.** The catalog
short-circuit in `EffectBridgeAdapter::execute_action` ran before
the registry, but `available_action_inventory` dedupes the opposite
way (internal beats external in the LLM-visible list). Result: an
LLM call to (say) `shell` would land in caller-side execution even
though the LLM saw the *internal* `shell` description in its action
surface — a confused-deputy where the caller can return any output
and the LLM trusts it as the internal tool's reply.

Fix: reject the collision at request validation in `responses_api.rs`.
After `validate_external_tools(...)`, look up registered tool names
via `state.tool_registry.tool_definitions()` and 400 any caller name
that shadows an internal action.

**Bug 4 — Catalog cleanup never happens.** `sweep_older_than` existed
but nothing scheduled it, and there was no terminal-state hook —
so catalog entries accumulated for every thread that ever ran.

Fix:
- `await_thread_outcome` in the bridge router calls
  `catalog.clear(thread_id)` on every non-`GatePaused` outcome
  (Completed, Stopped, MaxIterations, Failed). `GatePaused` keeps
  the entry so resume requests can still find it.
- A periodic sweep task in `init_engine` runs
  `catalog.sweep_older_than(1 hour)` every 5 minutes as a backstop
  for callers that abandon a paused thread without resuming.

**Tests now passing:**
- `engine_pauses_when_llm_calls_registered_external_tool` — proves Bug 1.
- `round_trip_resume_payload_reaches_llm` — proves Bug 2 (and
  Bug 1 by extension).
- `external_tool_name_shadowing_registered_action_is_rejected` — proves Bug 3.
- `catalog_cleared_on_terminal_completed_outcome` — proves Bug 4.

Plus four catalog-tier unit tests for the new `transfer` method, and
the two pre-existing collision-prevention tests
(`ext_tool:` vs `pairing:` callback id disambiguation).

`TestGatewayBuilder.tool_registry(...)` is a new builder hook for
tests that need to exercise the registry-aware handler paths
(currently only the collision-rejection test, but the seam is there
for future ones).

Quality gates: `cargo fmt`, `cargo clippy --all --benches --tests
--examples --all-features` clean. 17/17 engine v2 tests pass; 12/12
HTTP path-prefix tests pass.

* fix(responses-api): address review feedback from PR nearai#3122

Closes the race-window where caller-supplied external tools could be
invisible to the LLM on a thread's first turn, plus a batch of smaller
review findings.

Race fix (Bug #3 from review):
- Plumb `conversation_scope: Option<Uuid>` through
  `ThreadExecutionContext`, populated from thread metadata.
- `ConversationManager::handle_user_message` accepts an
  `extra_initial_metadata` map; the bridge stamps the parsed scope into
  it so the engine sees it on the in-memory thread the executor task
  reads from (post-spawn `set_thread_metadata` is invisible to that
  task).
- `EffectBridgeAdapter::execute_action` and
  `available_action_inventory` now look up the catalog under both
  `ctx.thread_id` and `ctx.conversation_scope`, so the executor task
  that starts immediately after spawn can find caller tools even
  before the bridge's post-spawn `transfer` rebinds them onto the
  engine `thread_id`. The transfer remains for terminal-state
  cleanup bookkeeping.

Other review fixes:
- Rewrite the stale `## Externally-provided tools` module doc in
  `responses_api.rs` to describe the engine-native flow (the original
  prompt-level fence text was left over from the first commit on the
  branch).
- `validate_external_tools` size check now fails closed on
  serialization error (`unwrap_or(MAX + 1)`) instead of silently
  passing oversized payloads.
- `notify_pending_gate` debug-logs the no-broadcaster path so a
  future SSE-less channel that grows an external-tool surface can
  be diagnosed instead of silently hanging.
- Update `catalog_clear_on_terminal_state_explicit` test docstring
  and rename to `catalog_clear_removes_entry` — Bug 4's cleanup hook
  in `await_thread_outcome` already wires the production cleanup
  (covered separately by `catalog_cleared_on_terminal_completed_outcome`).
- Delete dead `wait_for_first_engine_thread` helper and the
  `_harness_compiles` shim that kept it alive.

New test coverage:
- `bridge::effect_adapter::tests`: three race-window regression tests
  exercising both `available_action_inventory` and `execute_action`
  via the conversation_scope fallback path, plus a unit test on the
  `external_tool_catalog_keys` helper.
- `bridge::router::tests`: four `extract_external_tool_output` tests
  covering match-by-call_id, missing-call_id-returns-null,
  no-outputs-array fallback, and find-after-misses.
- `channels::web::responses_api::tests`: a full streaming round-trip
  test driving `streaming_worker` end-to-end with synthetic
  StreamChunk + ExternalToolCall events, parsing the actual SSE
  byte stream, and asserting the wire-frame ordering an OpenAI
  client would observe.

* test: fix three pre-existing engine test failures

- `executor::structured::call_id_preserved_when_no_lease`: the test
  was asserting on an error string ("no lease") that the preflight
  path stopped emitting when it added the "not callable in this
  execution context" check ahead of the lease lookup. The empty
  `MockEffects` used by the test exposed no actions, so the call
  short-circuited before reaching the lease check the test name
  describes. Register `web_search` in the inventory so the lease-miss
  path the test is named for actually fires.

- `runtime::manager::stop_thread_works`: the test races the
  consecutive-action-error guard added in nearai#2325. The thread loops on
  a deliberately unregistered `test_tool`, and 7 errors land before
  the 10ms sleep + `stop_thread` round-trip can deliver the stop
  signal in fast environments. The test's intent is "calling
  `stop_thread` doesn't deadlock the join", so accept `Failed` as a
  valid terminal outcome in addition to `Stopped`/`Completed`/
  `MaxIterations`. Asserts a more specific message on the join
  result so a true regression (non-terminal outcome) still trips.

- `tests::catalog_cleared_on_terminal_completed_outcome`: was racing
  two ways. (1) The cleanup-poll compared `catalog.len()` to a
  pre-snapshot — racy because `engine_external_tool_catalog` is
  process-global, so concurrent tests could keep the count from ever
  dropping below the pre-snapshot. (2) Multiple engine-touching
  tests in the file race the `OnceLock<Option<EngineState>>` global
  init, so messages sent by test A could be processed by test B's
  engine state.

  Fixes:
  - Add `ExternalToolCatalog::contains_action_anywhere(name)` so the
    cleanup poll can verify a unique marker action is gone regardless
    of what other tests have registered, and switch the test to use
    a per-test `format!("cleanup_marker_{uuid}")` action name.
  - Add a process-static `tokio::sync::Mutex` (`engine_state_lock`)
    that the three engine-touching tests in the file acquire for
    their duration, serializing their use of the global engine state.
    Per-test engine isolation belongs in the bridge itself; this is a
    test-side workaround until that lands.

* fix(responses-api): close gemini bot review findings

Three review fixes from gemini-code-assist on PR nearai#3122 plus a related
fix for fence-style tool calls reported by the user when running
gpt-5.3-codex through `/v1/responses`.

streaming_worker: finalize message item even when resolved text is empty
=========================================================================
The Response branch only emitted `output_item.done` when the resolved
text was non-empty. If `StreamChunks` had already opened a Message
item (`output_item.added` fired, `message_output_index` is `Some`)
and the terminal Response then resolved to an empty string,
`output_item.done` was skipped — leaving the OpenAI client with a
dangling in-progress message in its UI. Now we finalize whenever
either text is available or a message item is in flight, and skip
the redundant delta emit when there's nothing to deliver.

Regression test `streaming_worker_finalizes_item_when_resolved_text_is_empty`
drives the worker with an empty StreamChunk + empty Response and
asserts `added_count == done_count` for output_item events.

validate_external_tools: stream the size check, no allocation
=========================================================================
Replaced the `Vec<Value>` + `String::len()` size measurement with a
streaming `serde_json::Serializer` writing into a counting `io::Write`
sink. Same byte count, no intermediate heap allocation, correctness
contract unchanged (still fails closed if the serialization stream
errors mid-flight).

recover_tool_calls_from_content: recognize markdown-fenced tool calls
=========================================================================
Some OpenAI-compatible models — notably gpt-5.3-codex via the Codex
Responses API — emit caller-supplied tool calls as

    ```tool_call
    {"name": "get_balances", "arguments": {}}
    ```

instead of via the structured `function_call` output channel. Root
cause is engine v2's CodeAct preamble pushing "always respond in a
```repl block" hard enough that the model generalizes the fenced
protocol to a sibling `tool_call` fence for tools it can't dispatch
through Python. Tools ARE passed to the provider correctly
(`OpenAiCodexProvider::build_request_body` sets `body["tools"]` with
strict-OpenAI schema and `tool_choice: "auto"`); the model just opts
out of the structured surface in favor of a fence.

The structural fix is to suppress the CodeAct preamble for Responses
API turns that carry caller-supplied `tools[]`, but that's a larger
piece of work. As a defense-in-depth fix:

- `recover_tool_calls_from_content` now also matches markdown-fenced
  blocks with `tool_call`, `function_call`, or `tool_calls` info
  strings. Opening fence must be at line start to avoid matching
  inline backtick references inside prose. JSON body is parsed and
  validated against the available tool name set; unknown names are
  ignored.
- `clean_response` strips the same fenced blocks via a new
  `strip_markdown_fence_block` helper so any malformed-JSON fence
  the recovery skipped doesn't leak fence syntax to the user.

Six new unit tests in `llm::reasoning::tests` cover the recovery
(JSON, with arguments, function_call alias, unknown-tool ignored,
inline-reference ignored) and the clean_response stripping (clean
case + malformed-JSON case).

* fix(responses-api): emit ExternalToolCall on CodeAct GatePaused outcome

The Responses API was timing out on caller-supplied tool calls when the
LLM emitted them through CodeAct's Python (which is the default mode
for engine v2). Symptoms reported: gateway shows
'Tool X requires external confirmation (gate: external_tool)' as a
generic gate card and the /v1/responses POST returns response.failed.

There are two paths that handle a `GatePaused { External }` from the
engine:

1. `notify_pending_gate` — fires when the engine creates the gate
   mid-execution (Tier 0 structured tool calls). My earlier review-
   feedback fix wired this path to emit `AppEvent::ExternalToolCall`
   so the Responses API handler can surface a `function_call`
   ResponseOutputItem.

2. `await_thread_outcome`'s `ThreadOutcome::GatePaused` arm — fires
   when CodeAct's Python catches `EngineError::GatePaused`, raises a
   `RuntimeError("execution paused by gate 'external_tool'")`, the
   script unwinds, and the thread terminates with a `GatePaused`
   outcome. This arm calls `send_pending_gate_status`, which has an
   empty branch for `ResumeKind::External` (line 514:
   `External { .. } => {}`). No `AppEvent::ExternalToolCall` was
   emitted, so the Responses API handler waited for a never-arriving
   event and the response timed out as failed.

Fix: in the `GatePaused` arm, when `resume_kind` is External with the
`ext_tool:` callback prefix, broadcast `AppEvent::ExternalToolCall`
through the SSE manager and short-circuit past the approval-card
delivery path (which is for human-in-the-loop UX and doesn't apply to
caller-executed tools). Mirrors the `notify_pending_gate` projection.
Annotated with `// projection-exempt: bridge dispatcher, ...` per
`.claude/rules/gateway-events.md` since the broadcast is a projection
of a `ThreadOutcome` (one of the canonical source logs) rather than
an unscheduled side-channel emit.

Also update the `persist_v2_tool_calls_only_called_from_completed_arm`
regression test, which was matching on bare `ThreadOutcome::Completed`
and `ThreadOutcome::GatePaused` text. The Bug 4 catalog-cleanup hook
(commit a4ba764) introduced an early `if !matches!(outcome,
ThreadOutcome::GatePaused { .. })` guard above the match block, so the
first text occurrence of `ThreadOutcome::GatePaused` is now in that
guard rather than the match arm — false-failing the assertion. Anchor
the test on the match-arm destructuring patterns (with first-field
names) instead, so it pins the structural invariant the test name
describes.

* fix(responses-api): filter synthetic engine markers; clearer resume error

Two issues from the user's live test of caller-supplied tools:

1. `__codeact__` was leaking to the response output as a `function_call`
   item with that name. The orchestrator emits ActionFailed events with
   `action_name: "__codeact__"` when a CodeAct script crashes
   (orchestrator.rs:940), and the responses_api accumulator + streaming
   worker were dutifully projecting those into `function_call` items.

   Fix: add `is_synthetic_engine_action(name)` (any `__double_underscore__`
   name) and skip those events in `ResponseAccumulator::process` and
   `streaming_worker` for the ToolStarted/ToolCompleted/ToolResult arms.
   Internal markers no longer surface to the caller.

2. The "function_call_output supplied but no pending external tool
   call" error was opaque — it gave the caller no way to diagnose why
   their resume failed. Most common cause: the LLM ran caller tools
   through CodeAct (Python) instead of structured tool calls, which
   currently doesn't pause the thread (script crashes with a Python
   RuntimeError, no GatePaused outcome, no PendingGate persisted).
   Engine-side fix is in PR nearai#3157 and needs an extension to cover
   `External` resume kinds.

   Improved the error to point at:
   - The diagnostic check (verify prior response.output had a
     function_call item for this call_id)
   - The known limitation (CodeAct path doesn't dispatch caller tools)
   - PR nearai#3157 as the in-progress engine fix

Two new tests:
- `accumulator_filters_synthetic_engine_actions`: drives ToolStarted +
  ToolCompleted with `name: "__codeact__"` and asserts the output
  array stays empty.
- `is_synthetic_engine_action_recognizes_double_underscore`: pins the
  `__double_underscore__` predicate (positive: `__codeact__`,
  `__init__`; negative: regular names, single-underscore, leading- or
  trailing-only doubles).

* fix(responses-api): tighten resume validation, sanitize external payloads, address review findings

Addresses PR nearai#3122 review comments plus the user-directed clean-up:

- responses_api.rs: reject `function_call_output` items with missing/empty
  `call_id` (Copilot 474). Validate the pending gate is actually an
  external-tool gate (ResumeKind::External + `ext_tool:` prefix) and that
  at least one supplied `call_id` matches the pending callback before
  submitting the ExternalCallback (Copilot 1211).
- bridge/router.rs: run the synthesized external-tool payload through
  `SafetyLayer::sanitize_tool_output` before it reaches the LLM —
  external tool payloads originate outside `EffectBridgeAdapter`'s
  pipeline and need the same leak/policy/sanitizer pass internal tool
  outputs get. Move projection-exempt annotations onto the
  `broadcast_for_user` call lines so the gateway-events check passes.
- bridge/effect_adapter.rs: synthesize a `call_ext_<uuid>` call id when
  the executor reaches the external-tool short-circuit without
  `current_call_id` (Copilot effect_adapter.rs:1853), and document
  multi-call batching as an expected limitation at the short-circuit
  site so future readers know the engine pauses on first.
- reasoning.rs: correct the stale fence-recovery comment (Copilot 1738).
- e2e_responses_api_external_tools.rs: remove the "currently expected
  to fail" note; the test now pins the resume materialisation contract
  (Copilot 153).

* fix(responses-api): finalize streaming placeholder before function_call; strip stale PR pointer

Two follow-ups from review:

1. Streaming external-tool dangling `output_item.added`. When a StreamChunk
   arrived before the ExternalToolCall (the placeholder Message was
   already emitted via `output_item.added`), the prose-flush path created
   a *new* Message item at `acc.output.len()` and emitted a fresh
   added+done pair for it — never finalizing the original placeholder.
   OpenAI clients render the unmatched placeholder as "in progress"
   forever. Fix: when `message_output_index` is set, take it, fold the
   accumulated chunks into the existing item at that index, and emit
   `output_item.done` for the same index. Falls through to the original
   no-placeholder behaviour when there is no in-flight Message.

   Updated `streaming_worker_external_tool_call_emits_correct_frame_sequence`
   to match the corrected sequence (one Message added, one Message done,
   then FunctionCall added/done) and added a pairing-invariant
   assertion (`added_count == done_count`) so the regression can't be
   re-locked by an incorrect literal sequence.

2. Wire-visible error message in `responses_api.rs` referenced PR nearai#3157
   for the CodeAct fix. Once the PR merges the pointer is misleading,
   and external callers can't follow the link anyway. Dropped the
   trailing note; the behavioural part of the message stays.

* fix(responses-api): tighten external tool name validation and shadow check

Four follow-ups from review of 87fc8d4:

- responses_api.rs (Copilot #11): the byte-counter comment claimed it
  measured "what the caller actually sent over the wire" but the count
  is canonicalised JSON, not raw request bytes. Rewrote the comment to
  describe the canonicalised-size cap correctly. The `Serializer`
  import is kept — it's needed to bring the trait method into scope
  for the concrete `serde_json::Serializer::serialize_seq` call below
  (an earlier attempt to drop it failed to compile).

- responses_api.rs (Copilot #12): caller-supplied tool names were only
  checked for non-empty + uniqueness. Whitespace, control chars, and
  over-long names could propagate into engine action surfaces, SSE
  payloads, and downstream LLM clients (which all enforce the OpenAI
  Responses spec `^[A-Za-z0-9_-]{1,64}$` and would reject anyway).
  Validate at request time. Added regression tests covering
  whitespace/control/non-ASCII names and the length cap.

- responses_api.rs / bridge/router.rs / bridge/mod.rs (Copilot #13):
  the shadow check only consulted `ToolRegistry::tool_definitions()`
  and missed engine v2 capability actions (`mission_*`, `skill_*`,
  `memory_*`, etc.). A caller registering `mission_create` as an
  external tool would still hit the catalog short-circuit in
  `EffectBridgeAdapter::execute_action`, since the LLM-visible dedup
  in `available_action_inventory` doesn't extend to the execute path.
  Added `engine_capability_action_names()` accessor that pulls the
  full capability-action surface from the bridge's
  `CapabilityRegistry`, and merged it into the collision set the
  responses_api handler checks against.

- runtime/conversation.rs (Copilot #14): documented that
  `extra_initial_metadata` is spawn-only. The `Running` (inject) and
  `Resumable` (resume) branches ignore it; callers needing
  per-request metadata on existing threads must use
  `ThreadManager::set_thread_metadata` instead. The bridge's
  external-tool-catalog `transfer` already handles this for the one
  in-tree caller, but documenting the contract prevents future
  surprise.
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
* Bound approval command previews

* Tighten approval payload length guard

* Rebuild WebUI v2 approval bundle

* fix(reborn): reset approval preview expansion

---------

Co-authored-by: aiworkbot <robert.yan@near.ai>
Co-authored-by: aiworkbot <220660587+aiworkbot@users.noreply.github.com>
Co-authored-by: Robert Yan <mstr.raphael@gmail.com>
Co-authored-by: aiworkbot <aiworkbot@users.noreply.github.com>
Co-authored-by: Robert Yan <46699230+think-in-universe@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
personal-upstream-sync Bot pushed a commit that referenced this pull request Aug 6, 2026
…2 100% gate (nearai#7263)

* docs(target-arch): resolve the await-edge design question by measurement (D-S) and re-walk the WS9 verify row

Appends §12.13 D-S under delegated authority at owner direction, flagged
for post-hoc review by Illia Polosukhin (nearai#6696's author): the await-edge
store is measured to be a pure projection over ProcessDependencyPort
(that half of the shed happened inside nearai#6696 itself), and the resolver
is a genuine loop-tier responsibility journal edges cannot express
(owner recovery, sanitized transcript result materialization, batch-gate
resume-once drain, BlockedDependentRunGate resume policy). §6.7.3 is
amended (scheduler DONE / store DONE / resolver KEEP) instead of the
shed being executed; the 2.9k figure is corrected to 1,459 production +
1,448 cfg(test) lines. The §12.10 bullet, §2 divergence flag, §9 row 49,
§13 validation row, CHECKLIST header/WS4 pointer, README and PLAN all
carry the dated resolution.

WS9 verify row ticked with evidence: one lifecycle authority (the
process journal; TurnRunState/TurnRunRecord are projections via
AgentTurnProcessRuntime, ProcessRecord is a capability-invocation view,
no bare RunRecord exists) and §7 T4 re-walked clause-by-clause against
merged code — matches, including the checkpoint-gated no-auto-retry
mechanism (BeforeModel precedes ModelStage; requeue only when
checkpoint-free under the 3-claim cap).

Docs-only; no code, no tests, no gates touched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(ws12): rows 1-2 — package-set tick (64==64/1/0, gate+selftest+independent rederivation) and the 74-row §9 mapping audit (45 L / 15 L-A / 14 OBD / 0 NOT-LANDED; 3 findings recorded)

Row 1: check-target-tree.py reports 64 workspace members == 64 documented
packages, 1 documented exclusion (tools/ironclaw_silk_decoder), 0 owned
exceptions (EXCEPTIONS table empty — §5 steady state); self-test 17/17;
cargo-metadata name set diffed empty against an independent §5 parse.

Row 2: docs/reborn/target-architecture/ws12-mapping-audit.md is the audit
record — per-row executed-evidence, delete-clauses read against WS8's
execution notes, all 14 open rows cite their owning CHECKLIST/PROPOSAL
row or issue. Findings (recorded, not fixed): F1 prompt_envelope
manifest-description fix has no owner row; F2 WS6:429's 'nearai#5618 residue
deleted' overstates vs the live adopt_migrated_identity + open WS8:523;
F3 stale-docs cluster where the tree is ahead of the prose (trace
re-export drop, TurnRunTransitionPort, processes->resources).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fold(7154): squash-port fix/red-main-7119 onto family-world main — defect train nearai#7146/nearai#7115/nearai#7104/nearai#7103/nearai#7144 (+nearai#7119 CI lane), 34-hunk contribution.rs port into the split modules, planner entrypoint classification, D-R loopback exception on the widened HTTPS credential guard

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(extractors): issue-number + assertion-rationale doc refinement (rescued 844964f from rescue/7154-parked-guard)

Ports only the doc/assertion refinement commit; the guard-parking commit
e8f5a31 on that branch is deliberately NOT taken — superseded by the
D-R loopback ruling.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): record D-R — the loopback credential-guard ruling, wiring choice, and regression pins (PROPOSAL §12.13, 2026-08-05)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review(7154): CodeRabbit round-1 triage — fail-closed tracing-target scan traversal (+node_modules), bounded sidecar output draining (capped capture + discard drain), deadlock regression asserts successful redaction (no seq), XLSX/DOCX empty-classification via extract_document, raise_for_status annotations

Threads already addressed by the fold: latency.rs caller-contract wording
(merged doc scopes the requirement to latency-trace callers), BodyJsonPointer
coverage (the plaintext-refusal test drives all four injection shapes).
Deliberately not taken: un-xfailing the four Slack-catalog projections —
the xfail is a documented tripwire (unexpected-pass goes red) and clearing
them is the nearai#6520 projection-modeling follow-on its comment specs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(assistant): re-point the one field-form tracing target the nearai#7146 gate caught — main's relocated triggered_run_delivery_services carried the drift the PR fixed at its old channel_host address

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* closure fixups: execute the mapping-audit findings — prompt_envelope manifest description (F1), dated ✎ corrections for the nearai#5618 overstatement (F2) and the stale-prose cluster (F3)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(ws12): second-reviewer security spot-audit + extension-journey re-verification (rows 5-6)

Adversarial second-reviewer pass over PROPOSAL §12.1a/b/c and the batch's own
§12.13 D-R loopback carve-out, plus a re-run of the five extension journeys.
Attacks were executed rather than argued: two sabotage files and a 38-shape
hostile-URL probe were planted, run, and reverted.

Verdicts — mint consolidation HOLDS-WITH-RESIDUAL, secrets tightening
HOLDS-WITH-RESIDUAL, host/verifier colocation HOLDS, D-R HOLDS. No HOLE.

Four findings recorded rather than fixed (report-not-repair):
- F1 test_verified/_for_tenant are ungranted mint constructors gated only by
  the `test-support` feature, in no mint-name table, with nothing pinning the
  feature to [dev-dependencies]; the shipped binary is measured feature-free.
- F2 §12.1b's products-layer residue undercounts by one (ironclaw_assistant).
- F3 journey coverage hole: gsuite-with-credential-injection is proven in two
  halves that no committed test joins.
- F4 both recorded census evasions and both fail-open reads are CLOSED on this
  tree, so §11.2.5/§12.1a/CHECKLIST:552/:597 now understate the seal.

Rows 5-6 ticked; only lines 631-632 of CHECKLIST touched so the concurrent
rows 3-4 edit folds cleanly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ratchet(closure): lock the budget gate at the program's end state

Dispatch ceiling 1122 -> 814 (today's observed, nudge taken; WS0 record 827
stays within effective 829). Mass-share ceiling 2398 -> 658 bp (the WS0
baseline floor — the arch-test assert refuses lower, and observed 578 bp sits
inside the nudge window). Absolute LOC re-equalized at 40423: nearai#6831 added 4
governed LOC through the queue's tolerance window; ceiling, observed, and
COMPOSITION_ABSOLUTE_SRC_LOC move together here. Both tightenings
sabotage-verified red (dispatch 9-over at 790; abs 73-over at 40200).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(ws12): gauntlet report — row 3 ticked (full gauntlet green, 0 REAL in scope), row 4 verified-but-open on two pre-existing Postgres-leg test-isolation defects

WS12 rows 3-4 verification on the assembled batch tip 0c6c0cf:

Row 3 (ticked): fmt, clippy default/all-features/--lib --bins, workspace
tests (495 targets, 15,203 passed, 0 failed; the smoke.rs:3132
CPU-saturation flake passed first try), arch suite 285/0, the
integration-feature lane 1,665/0, recorded-fixture QA (61 fixtures clean,
41/0), frontend (typecheck 1,588 files; vitest 1,088/0; build + bundle
budgets), e2e smoke = the CI browser lane under the hermetic wrapper
(50 + 21 + 5 passed), and all 41 scripts/ci self-tests (two mapfile/bash-3.2
casualties green under bash 5, the CI shape).

Row 4 (stays open, dated note added): both-backend parity proven with
legs demonstrably executed for the fabric (57 pg + 81 libsql), triggers
(ADR 0003, REQUIRE_POSTGRES), hooks (ADR 0004, all three backends),
composition, processes journal, extension-registry, host-runtime libSQL
restart, and the backend matrix; fabric-delegated domains enumerated.
Two REAL blockers (one class): the Postgres legs of the event-store and
assistant-ledger contract suites assert against shared-database state and
cannot pass as-written (each failing test passes alone on a virgin
database; files byte-identical to origin/main; no CI lane sets their env
vars). Full evidence: docs/reborn/target-architecture/ws12-gauntlet-report.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(tests): per-test isolated Postgres databases for the two WS12 parity-blocking contract suites

The WS12 gauntlet (ws12-gauntlet-report.md §P6/§P8) measured the Postgres
legs of ironclaw_event_store's durable_event_store_contract and
ironclaw_assistant's durable_ledger_contract as test-isolation-defective:
absolute database-global asserts (event cursors; settled-entry prune
bookkeeping) run against the single external database named by their
IRONCLAW_*_POSTGRES_URL env vars. Every failing test passes alone on a
virgin database - store semantics correct, suites not self-isolating
(PROPOSAL §12.13 D-T).

Fix: each affected test provisions a private database on the configured
server - the fabric contract's IsolatedDatabase pattern
(db_root_filesystem_contract.rs) ported locally into each suite: CREATE
DATABASE per test, store/pool + migrations against it, courtesy
DROP ... WITH (FORCE), and a once-per-binary stale-name sweep. Every
assertion preserved byte-identical; libsql/jsonl twins untouched. In the
ledger suite only the two retention tests move - the other six Postgres
tests keep their proven fingerprint-suffix isolation.

Regression pins are the fixed tests themselves:
- postgres_replay_advances_next_cursor_past_trailing_filtered_records
- postgres_runtime_and_audit_logs_survive_rebuild_with_filtered_cursor_semantics
- postgres_settled_entry_limit_prunes_oldest_when_configured
- postgres_settled_prune_interval_defers_until_interval_when_configured
Green proven on a shared dirty database twice in a row (parallel default
threading) and serially on a virgin database; red-first reproduction
captured before the fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(reborn): record §12.13 D-T (parity-suite isolation ruling) and close CHECKLIST WS12 row 4

D-T (after D-S): the WS12 gauntlet's two REAL findings were one defect
class — absolute database-global asserts against the single shared
env-var Postgres database — in two suites (event store cursor contract,
assistant settled-ledger retention). Ruling executed in commit 864d93e:
per-test isolated databases via the fabric contract's IsolatedDatabase
pattern, assertions preserved; alternatives (baseline-relative asserts,
serial-only, leave-open) recorded with why they lost; regression pin =
the four fixed tests themselves.

CHECKLIST WS12 backend-parity row ticks [x] with a dated addendum: red-first
reproduction, the three green isolation runs (dirty shared DB twice in
parallel; failing pairs serial on virgin), parity now green 10/10.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): three measured corrections surfaced by the guidance program

memory packages are substrates-layer, not products (families/extensions.md);
memory_native declares no extension_contracts dep (PROPOSAL §6.8.4); wasm's
extension_contracts edge is dev-only and the wasm 'never depends on' bullet is
lane-scoped, not family-wide (families/lanes.md).

Three further reported defects were checked and NOT corrected — they were
misreads: the sandbox 'never above the runtime tier' rule holds (substrates sit
below it), and PROPOSAL's safety consumer count already reads 17, matching the
tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): repair the corrupted kernel bullet and correct two family laws

kernel.md: ironclaw_authorization's 'Security & authority role' bullet has been
textually corrupted since nearai#6918 — an approvals sentence was spliced into it
mid-clause, orphaning its continuation line. Reconstructed, with the spliced
sentence restored to the approvals entry where it is true.

lanes.md: 'a lane never depends on a substrate' is false as a family-wide law
(ironclaw_sandbox holds network/safety/secrets normal deps, which its own entry
licenses); the accurate law is the layer ladder, and the narrow claim holds for
ironclaw_wasm alone.

lanes.md + events.md: the 'every crate ships both an AGENTS.md and a CLAUDE.md'
requirement is superseded by docs/reborn/guidance-conventions.md — two files
restating one rule is the drift the guidance program removes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(arch): govern the ProtocolAuthEvidence test seam — WS12 audit F1

Two new gates in reborn_sealed_evidence_mint_ratchet (closed paths #12/#13),
per the audit's remedy spec:
(a) TEST_SEAM_MINT_FNS governs test_verified/test_verified_for_tenant — any
    production-text call site outside ironclaw_host_api is an offender
    (comments/strings stripped, #[cfg(test)] blocks stripped, tests.rs /
    *_tests.rs and cfg-test-only files excluded via the shared census);
(b) test-support may appear in no normal dependency table workspace-wide
    (dependencies / build-dependencies / target.* variants /
    workspace.dependencies), and no [features] key other than test-support
    may forward to it — the laundering shape that would evade (b) by one
    rename. [dev-dependencies] enablement stays legal (cargo-features.md
    bar 4, the sanctioned dev seam).

Measured zero offenders on this tree in both directions before pinning;
sabotage-proven red->green both ways (planted production call named with
file:line-text; [dependencies] enablement named with its table path).
Self-tests drive the same pipelines the gates run (zero-match principle);
the definition-location and partition tests now cover the new table.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(integration): join the gsuite credential-injection journey — WS12 audit F3

WS12 row 5 leg 3 was verified in two halves no committed test joined: gsuite
handler -> staged credential (crate tier) and staged obligation -> wire
(GitHub/Slack only). Scenario 5 already drives gmail.list_messages through
production dispatch on a Google-OAuth-configured group; it now also asserts
the JOIN: the seeded google account's token (itest-google-token) lands on
the recorded outbound gmail.googleapis.com request as
'authorization: Bearer ...', injected at the host egress chokepoint
(apply_credential_injection) per the gmail manifest's declared recipe —
store -> dispatch-time staging -> chokepoint -> wire, through the caller.

Sabotage-proven: disabling the Header injection arm reds exactly this
scenario with 'no network egress request matching url gmail.googleapis.com
has header authorization' while the request itself still reaches the wire
(headers seen: content-type only) — the injection reason, not a setup
error; restore -> green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): correct five measured dependency claims in families/domains.md

conversations does not depend on safety (its BoundaryRule now forbids it);
triggers depends on libsql_runtime + safety and NOT filesystem, so its
'filesystem-routed persistence path alongside SQL' is one path, not two;
memory's live set is host_api alone (prompt_envelope is allowlisted, unused);
auth was short by extension_contracts + product_contracts.

Each verified against the manifest before editing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): record the closed scan evasions (F4) and the secrets-consumer correction (F2)

The sealed-mint census weaknesses PROPOSAL §11.2.5/§12.1a and CHECKLIST recorded
as live and owed to WS10 are all closed on this tree, verified by re-attacking
the seam with both evasions at once; the docs understated the seal. Ratchet is
23 tests. One residual replaces them: the test_verified test-seam constructors,
now pinned by two gates.

§12.1b's 'only products-layer crate with the edge' is false by one —
ironclaw_assistant carries ironclaw_secrets as port-declaration vocabulary with
no expose_secret call. Not a value-reach bypass; joins nearai#7095's inventory.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): correct the app-family layer, config's consumer set, and the webui route count

ironclaw_config declares layer=substrates while living in crates/app/;
its consumers include operator, extension_manager and extension_host, not just
the assembly crate and the binary; webui is 93 contract-locked routes, not 92
(nearai#6780 landed after the last recount).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(ws12): tick row 7 — the fresh-agent placement probe passed on the final tree

All three placements correct with high confidence, each naming the trait, the
tests, and the tempting wrong place it rejected. The probe doubled as a docs
audit and independently hit four defects, three of which the stacked guidance
PR fixes — it succeeded despite them.

WS12 is now 7/7. The restructure is complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): the product→loop_host recount was wrong on the day it was written

Eight importing files across four seams, not seven across three — the fourth
being a skill-activation-observer seam (projection.rs, projection/live_progress.rs)
this bullet never named, which §6.4.7's own same-day note already implied.
Surfaced by the plan-conformance audit.

The recount history is 3→5→6→7→8, wrong at four of five attempts. That retires
the prose count as a method: the sever slice should land an inventory ratchet
before or with the move, not another number.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review(7263): CodeRabbit round-1 triage — 4 code fixes (2 sabotage-proven, 2 red-first) + 6 doc-truth corrections

Code, each verified red-first or by sabotage matrix:
- sealed-mint ratchet: per-name sighting floor for TEST_SEAM_MINT_FNS
  (closed path #12). Proven: renaming test_verified_for_tenant away plus one
  extra legitimate sibling mention passed the old aggregate floor (silent
  disarm) and fails the new per-name floor naming the constructor; suite
  23/23 after revert. (CodeRabbit's claimed baseline ">2 mentions today" is
  wrong — each name has exactly one kept sighting — but the doc/enforcement
  mismatch and at-threshold fragility were real.)
- trace credit: non-finite novelty_score/duplicate_score are treated as
  absent before clamping (clamp preserves NaN, which poisoned online_score
  and credit_points_estimate); NaN cases added to the nearai#7144 regression test,
  red first.
- trace submission: a 2xx whose body stream dies mid-read now maps through
  request_failed (network telemetry kind, true I/O cause) instead of
  collapsing to an empty body that the nearai#7144 strict parse misreported as
  response_invalid/Submission; truncated-body regression test, red first.
- Postgres contract suites (event store + assistant ledger): isolated-DB
  names now carry a creation epoch and the once-per-binary sweep is
  age-gated (1h), closing the cross-process window where a sibling's fresh
  zero-backend database (between CREATE DATABASE and first connection) was
  sweepable; legacy pid-scheme leftovers still collect immediately. Proven
  on live Postgres 16: planted stale name swept, planted fresh name
  survives, 13/13 x2 and 20/20 x2 with zero leftovers.

Docs (target-architecture truth pass):
- PROPOSAL section 9: the WS6 rename sweep (nearai#7152) had rewritten the source
  column of the 12 renamed rows to their post-rename names, turning their
  rename dispositions into no-ops (rows 13/14/28/30/49/51/59/61/64/66/67/70);
  pre-restructure names restored with a dated footnote.
- PROPOSAL:69: removed the superseded 3->5->6->7 recount sentence (the
  corrected 3->5->6->7->8 passage subsumes it).
- PROPOSAL row 34: ToolPermissionOverrideStorePort deletion marked landed
  (2026-08-05 WS8, matching section 6.5.3; zero workspace hits).
- CHECKLIST:631: dated note recording that the WS12 F3 gsuite join landed in
  this batch (scenario_uninstalled_tool_call_denied_until_active.rs asserts
  the seeded google token on the gmail.googleapis.com wire; suite run green).
- CHECKLIST:632: dated note spending F4 (the audit's 19 was correct at its
  SHA; the ratchet file now holds 23 tests, re-counted at lines 552/597).
- ws12-gauntlet-report P6 heading: first of TWO real failures (one class),
  matching P8 and the report's own summary.
- ws12-mapping-audit rows 49/137: dated D-S closure notes (await-edge store
  half = journal projection already; resolver retained loop-tier; no shed
  owed) so the backlog register no longer lists it as in-flight.

Not fixed, with evidence: the span-helper macros gate suggestion
(info_span!(target = ...) is a hard compile error, E0425 — no silent trap),
the webui tracing-subscriber workspace-dep suggestion (no
[workspace.dependencies] entry exists; suggestion would not build; 8
siblings use the identical direct shape), and the mapping-audit
regeneration (the audit is accurate at its pinned SHA; the in-batch F1 fix
is recorded in its dated coordinator note).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review(7263): CodeRabbit round-2 — rejection-body read keeps its cause; 200 {} is not a submission acknowledgement; lanes.md family dep rule matches measured Cargo.tomls

- submission.rs non-2xx path: a failed rejection-body read no longer collapses
  to an empty detail via .unwrap_or_default() (banned by
  .claude/rules/error-handling.md); the read error folds into the
  http_rejection detail so the received status keeps driving the 401/403
  auth-retry and the Credential/HttpRejection telemetry split.
  Regression: submit_preserves_rejection_body_read_failure_cause_with_status.

- TraceSubmissionReceipt.status: serde default removed — it fabricated
  status "submitted" from a proxy's 200 {} (the nearai#7144 synthesis, resurfacing
  through the wire type's defaults), after which the flush caller recorded
  Submitted and deleted the only retryable queued copy. The acknowledgement is
  the server naming what happened to the submission — every workspace fixture
  sends status and callers persist it unconditionally as server_status — so a
  status-less 2xx body now fails the strict receipt parse as response_invalid.
  Regression: submit_rejects_success_response_without_explicit_server_status
  (covers 200 {} and a status-less non-empty object).

- docs(lanes.md): the family Dependency-direction rule no longer claims every
  lane takes the extension-surface vocabulary crate — measured across
  crates/lanes/*/Cargo.toml: mcp + sandbox hold ironclaw_extension_contracts
  under [dependencies], wasm only under [dev-dependencies]; dated ✎
  cross-references the ironclaw_wasm entry's 2026-08-05 correction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review(7263): CodeRabbit round-3 — shared Postgres test provisioner (the "new dep edge" premise measured false), entrypoint self-test armed (sabotage-proven), six doc self-contradictions reconciled

Code:
- ironclaw_filesystem gains a `postgres_isolation` test-support module — the
  single home of the per-test isolated-database scaffolding (once-per-binary
  age-gated stale sweep, epoch-in-name convention, DROP WITH (FORCE) cleanup),
  parameterised by suite/env-var/prefix/unreachable-policy. Zero new
  production edges: event_store already normal-deps filesystem, filesystem
  already owns tokio-postgres, and the dev-dep+feature pattern is the one 17
  crates already use. The event-store and product-workflow-ledger suites
  migrate onto it; both Postgres legs proven live against postgres:16 (12
  tests, zero leftover databases). The fabric original keeps its older
  variant with the differences documented at its IsolatedDatabase.
- ironclaw_event_store drops the duplicate tokio-postgres dev-dep (the normal
  dep already reaches tests).
- test-reborn-docker-entrypoint.sh: the missing-argv check now exits the
  command-substitution subshell instead of incrementing a counter the parent
  never sees — red-proven (a migrate-but-never-exec entrypoint passed with 7
  FAIL lines printed), green after the fix both sabotaged and restored.
- trace_commons submission test additionally pins !auth_rejection() for the
  503 rejection (the structural assert the API affords; the prescribed
  payload asserts are refuted — status is private and source is None by
  design, with the message derived from the structured status in the same
  constructor).

Docs (each reconciled to one canonical statement, measured):
- kernel.md: lease ownership decided from code — authorization stores,
  matches, and expires leases (CapabilityLeaseStore + port + expiry all live
  there); approvals constructs and issues into that store. The round-1
  re-homing of the spliced sentence into approvals was wrong and is corrected
  in the dated repair note.
- app.md: "nothing depends on app" scoped to the three app-layer crates;
  ironclaw_config's consumers restated by dependency kind (normal:
  composition, cli, operator, extension_host; dev-only: extension_manager,
  root integration-tests package).
- lanes.md: the mediated-services sentence now states the family law as
  layer-ladder + injected authority; the no-secrets/network/filesystem-dep
  claim is scoped to ironclaw_wasm, matching the file's own corrections.
- CHECKLIST 429/430: the one open traces clause is named (ScopedFilesystem
  adoption); the stale "other two" count corrected against the F3a strike.
- PROPOSAL:69 + CHECKLIST:72: the project-create route repointed —
  first_party_extension_ports dissolved into loop_host::skill_activation
  (WS8, §9 row 55) — still unattempted.
- PROPOSAL §9 rows 57/62 synced to §6.8.4 (telegram: dependency-set equality
  with Slack's four contract-tier crates) and §6.9.4 (webui -> assistant is a
  charter-permanent edge, §12.11 D-B).
- PLAN top summary records Wave 6's design question as resolved (D-S,
  2026-08-05).
- deploy-reborn-cli-docker.md: the two migration paragraphs unified on the
  entrypoint's actual behavior — only enabled = false beside
  signing_secret_env/bot_token_env is migrated; every other retired-key shape
  fails startup with the migration pointer.
- composition-budget.toml: the stale "2398 bp, a true ratchet" header
  replaced with the WS0-floor truth the baselines test asserts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: move the guidance convention into this PR so its citations resolve

families/lanes.md and families/events.md cite docs/reborn/guidance-conventions.md
when superseding their 'every crate ships both an AGENTS.md and a CLAUDE.md'
requirement, but the file was only on the stacked guidance branch — a forward
reference that dangles if this PR merges alone. The convention is the rule those
notes invoke, so it belongs with them.

Caught by the CodeRabbit round-3 pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): give the hoisted postgres provisioner its safety rationales

The round-3 hoist moved test provisioning into a production src/ path, so
check_no_panics flagged its four panic/expect sites and reddened Code Style via
fast-checks. The gate is right to flag them: it deliberately does NOT exempt
#[cfg(feature = "test-support")] modules, because a cargo feature is not a
privilege boundary in this workspace (PROPOSAL 12.1a proved exactly that) —
so a test-support module still compiles into a build where any sibling enables
the feature.

Suppressed with the gate's documented inline rationale, which must trail the
statement rather than precede it. The panics themselves stay: a configured but
unusable Postgres must fail the suite loudly rather than skip it, which is the
inert-guard rule the isolation fix exists to serve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
personal-upstream-sync Bot pushed a commit that referenced this pull request Aug 6, 2026
…ry crate, and a repo-wide stale sweep (nearai#7264)

* docs(target-arch): resolve the await-edge design question by measurement (D-S) and re-walk the WS9 verify row

Appends §12.13 D-S under delegated authority at owner direction, flagged
for post-hoc review by Illia Polosukhin (nearai#6696's author): the await-edge
store is measured to be a pure projection over ProcessDependencyPort
(that half of the shed happened inside nearai#6696 itself), and the resolver
is a genuine loop-tier responsibility journal edges cannot express
(owner recovery, sanitized transcript result materialization, batch-gate
resume-once drain, BlockedDependentRunGate resume policy). §6.7.3 is
amended (scheduler DONE / store DONE / resolver KEEP) instead of the
shed being executed; the 2.9k figure is corrected to 1,459 production +
1,448 cfg(test) lines. The §12.10 bullet, §2 divergence flag, §9 row 49,
§13 validation row, CHECKLIST header/WS4 pointer, README and PLAN all
carry the dated resolution.

WS9 verify row ticked with evidence: one lifecycle authority (the
process journal; TurnRunState/TurnRunRecord are projections via
AgentTurnProcessRuntime, ProcessRecord is a capability-invocation view,
no bare RunRecord exists) and §7 T4 re-walked clause-by-clause against
merged code — matches, including the checkpoint-gated no-auto-retry
mechanism (BeforeModel precedes ModelStage; requeue only when
checkpoint-free under the 3-claim cap).

Docs-only; no code, no tests, no gates touched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(ws12): rows 1-2 — package-set tick (64==64/1/0, gate+selftest+independent rederivation) and the 74-row §9 mapping audit (45 L / 15 L-A / 14 OBD / 0 NOT-LANDED; 3 findings recorded)

Row 1: check-target-tree.py reports 64 workspace members == 64 documented
packages, 1 documented exclusion (tools/ironclaw_silk_decoder), 0 owned
exceptions (EXCEPTIONS table empty — §5 steady state); self-test 17/17;
cargo-metadata name set diffed empty against an independent §5 parse.

Row 2: docs/reborn/target-architecture/ws12-mapping-audit.md is the audit
record — per-row executed-evidence, delete-clauses read against WS8's
execution notes, all 14 open rows cite their owning CHECKLIST/PROPOSAL
row or issue. Findings (recorded, not fixed): F1 prompt_envelope
manifest-description fix has no owner row; F2 WS6:429's 'nearai#5618 residue
deleted' overstates vs the live adopt_migrated_identity + open WS8:523;
F3 stale-docs cluster where the tree is ahead of the prose (trace
re-export drop, TurnRunTransitionPort, processes->resources).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fold(7154): squash-port fix/red-main-7119 onto family-world main — defect train nearai#7146/nearai#7115/nearai#7104/nearai#7103/nearai#7144 (+nearai#7119 CI lane), 34-hunk contribution.rs port into the split modules, planner entrypoint classification, D-R loopback exception on the widened HTTPS credential guard

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(extractors): issue-number + assertion-rationale doc refinement (rescued 844964f from rescue/7154-parked-guard)

Ports only the doc/assertion refinement commit; the guard-parking commit
e8f5a31 on that branch is deliberately NOT taken — superseded by the
D-R loopback ruling.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): record D-R — the loopback credential-guard ruling, wiring choice, and regression pins (PROPOSAL §12.13, 2026-08-05)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review(7154): CodeRabbit round-1 triage — fail-closed tracing-target scan traversal (+node_modules), bounded sidecar output draining (capped capture + discard drain), deadlock regression asserts successful redaction (no seq), XLSX/DOCX empty-classification via extract_document, raise_for_status annotations

Threads already addressed by the fold: latency.rs caller-contract wording
(merged doc scopes the requirement to latency-trace callers), BodyJsonPointer
coverage (the plaintext-refusal test drives all four injection shapes).
Deliberately not taken: un-xfailing the four Slack-catalog projections —
the xfail is a documented tripwire (unexpected-pass goes red) and clearing
them is the nearai#6520 projection-modeling follow-on its comment specs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(assistant): re-point the one field-form tracing target the nearai#7146 gate caught — main's relocated triggered_run_delivery_services carried the drift the PR fixed at its old channel_host address

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* closure fixups: execute the mapping-audit findings — prompt_envelope manifest description (F1), dated ✎ corrections for the nearai#5618 overstatement (F2) and the stale-prose cluster (F3)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(ws12): second-reviewer security spot-audit + extension-journey re-verification (rows 5-6)

Adversarial second-reviewer pass over PROPOSAL §12.1a/b/c and the batch's own
§12.13 D-R loopback carve-out, plus a re-run of the five extension journeys.
Attacks were executed rather than argued: two sabotage files and a 38-shape
hostile-URL probe were planted, run, and reverted.

Verdicts — mint consolidation HOLDS-WITH-RESIDUAL, secrets tightening
HOLDS-WITH-RESIDUAL, host/verifier colocation HOLDS, D-R HOLDS. No HOLE.

Four findings recorded rather than fixed (report-not-repair):
- F1 test_verified/_for_tenant are ungranted mint constructors gated only by
  the `test-support` feature, in no mint-name table, with nothing pinning the
  feature to [dev-dependencies]; the shipped binary is measured feature-free.
- F2 §12.1b's products-layer residue undercounts by one (ironclaw_assistant).
- F3 journey coverage hole: gsuite-with-credential-injection is proven in two
  halves that no committed test joins.
- F4 both recorded census evasions and both fail-open reads are CLOSED on this
  tree, so §11.2.5/§12.1a/CHECKLIST:552/:597 now understate the seal.

Rows 5-6 ticked; only lines 631-632 of CHECKLIST touched so the concurrent
rows 3-4 edit folds cleanly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ratchet(closure): lock the budget gate at the program's end state

Dispatch ceiling 1122 -> 814 (today's observed, nudge taken; WS0 record 827
stays within effective 829). Mass-share ceiling 2398 -> 658 bp (the WS0
baseline floor — the arch-test assert refuses lower, and observed 578 bp sits
inside the nudge window). Absolute LOC re-equalized at 40423: nearai#6831 added 4
governed LOC through the queue's tolerance window; ceiling, observed, and
COMPOSITION_ABSOLUTE_SRC_LOC move together here. Both tightenings
sabotage-verified red (dispatch 9-over at 790; abs 73-over at 40200).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(ws12): gauntlet report — row 3 ticked (full gauntlet green, 0 REAL in scope), row 4 verified-but-open on two pre-existing Postgres-leg test-isolation defects

WS12 rows 3-4 verification on the assembled batch tip 0c6c0cf:

Row 3 (ticked): fmt, clippy default/all-features/--lib --bins, workspace
tests (495 targets, 15,203 passed, 0 failed; the smoke.rs:3132
CPU-saturation flake passed first try), arch suite 285/0, the
integration-feature lane 1,665/0, recorded-fixture QA (61 fixtures clean,
41/0), frontend (typecheck 1,588 files; vitest 1,088/0; build + bundle
budgets), e2e smoke = the CI browser lane under the hermetic wrapper
(50 + 21 + 5 passed), and all 41 scripts/ci self-tests (two mapfile/bash-3.2
casualties green under bash 5, the CI shape).

Row 4 (stays open, dated note added): both-backend parity proven with
legs demonstrably executed for the fabric (57 pg + 81 libsql), triggers
(ADR 0003, REQUIRE_POSTGRES), hooks (ADR 0004, all three backends),
composition, processes journal, extension-registry, host-runtime libSQL
restart, and the backend matrix; fabric-delegated domains enumerated.
Two REAL blockers (one class): the Postgres legs of the event-store and
assistant-ledger contract suites assert against shared-database state and
cannot pass as-written (each failing test passes alone on a virgin
database; files byte-identical to origin/main; no CI lane sets their env
vars). Full evidence: docs/reborn/target-architecture/ws12-gauntlet-report.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(guidance): set the crate/family guidance convention

The base commit for the family-guidance program: one canonical home per fact,
measured-not-aspirational claims, boundaries stated as exclusions, and the note
that guidance files can be gate-pinned. Every family/crate document written on
top of this branch follows this shape.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(tests): per-test isolated Postgres databases for the two WS12 parity-blocking contract suites

The WS12 gauntlet (ws12-gauntlet-report.md §P6/§P8) measured the Postgres
legs of ironclaw_event_store's durable_event_store_contract and
ironclaw_assistant's durable_ledger_contract as test-isolation-defective:
absolute database-global asserts (event cursors; settled-entry prune
bookkeeping) run against the single external database named by their
IRONCLAW_*_POSTGRES_URL env vars. Every failing test passes alone on a
virgin database - store semantics correct, suites not self-isolating
(PROPOSAL §12.13 D-T).

Fix: each affected test provisions a private database on the configured
server - the fabric contract's IsolatedDatabase pattern
(db_root_filesystem_contract.rs) ported locally into each suite: CREATE
DATABASE per test, store/pool + migrations against it, courtesy
DROP ... WITH (FORCE), and a once-per-binary stale-name sweep. Every
assertion preserved byte-identical; libsql/jsonl twins untouched. In the
ledger suite only the two retention tests move - the other six Postgres
tests keep their proven fingerprint-suffix isolation.

Regression pins are the fixed tests themselves:
- postgres_replay_advances_next_cursor_past_trailing_filtered_records
- postgres_runtime_and_audit_logs_survive_rebuild_with_filtered_cursor_semantics
- postgres_settled_entry_limit_prunes_oldest_when_configured
- postgres_settled_prune_interval_defers_until_interval_when_configured
Green proven on a shared dirty database twice in a row (parallel default
threading) and serially on a virgin database; red-first reproduction
captured before the fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(reborn): record §12.13 D-T (parity-suite isolation ruling) and close CHECKLIST WS12 row 4

D-T (after D-S): the WS12 gauntlet's two REAL findings were one defect
class — absolute database-global asserts against the single shared
env-var Postgres database — in two suites (event store cursor contract,
assistant settled-ledger retention). Ruling executed in commit 864d93e:
per-test isolated databases via the fabric contract's IsolatedDatabase
pattern, assertions preserved; alternatives (baseline-relative asserts,
serial-only, leave-open) recorded with why they lost; regression pin =
the four fixed tests themselves.

CHECKLIST WS12 backend-parity row ticks [x] with a dated addendum: red-first
reproduction, the three green isolation runs (dirty shared DB twice in
parallel; failing pairs serial on virgin), parity now green 10/10.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(extensions): family guidance layer — AGENTS.md rewrite to the guidance-conventions shape, READMEs for all 4 family crates and 14 packages, duplicate-guidance consolidation

The family AGENTS.md now teaches the unified extension model (extension =
the only product object; channel/tool/auth are manifest surfaces; runtime
is loading, never taxonomy; ExtensionId vs VendorId; retired vocabulary
pinned by reborn_retired_taxonomy.rs), carries the self-containment and
package-to-crate rules from families/extensions.md, the four-responsibility
lookup, the measured package catalog, the exclusion list, and the armed
gates by test name.

Every crate and package gains a README.md (ironclaw_extension_host had no
guidance of any kind). ironclaw_extension_registry and memory-native each
had both an AGENTS.md and a CLAUDE.md saying overlapping things: AGENTS.md
is now canonical, CLAUDE.md a pointer, and memory-native's stale v1
references (src/workspace, src/db/libsql) are dropped in the merge. The
slack/telegram agent maps get package framing and a contracts-tier pointer
in place of the stale ironclaw_assistant one. Every path literal verified
to resolve on disk; all figures (tool counts, dep sets, consumers, layer
declarations) measured from the tree at 8d13454.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(guidance): substrates + lanes family guidance per guidance-conventions.md

Family AGENTS.md rewritten to the spec shape for crates/substrates/ and
crates/lanes/: boundary, crate table, exclusion lists (mechanism-not-authority
for substrates; kernel-decides-lane-executes for lanes), armed gates by test
name, and measured deviations stated as deviations (sandbox's three substrate
deps, script.rs direct spawn). The lanes wit/-is-load-bearing note is kept.

A README.md for every crate in both families (10 new), measured against
cargo metadata 2026-08-05: public surface, workspace edges, consumer counts,
and enforced invariants each citing their gate. ironclaw_libsql_runtime and
ironclaw_wasm_limiter previously had no guidance of any kind; their READMEs
carry the sole-pool-home rule (ADDITIONAL_DRIVER_ALLOWLISTS: deadpool =
{filesystem, libsql_runtime}) and the outbound-only limiter gate
(wasm_sandbox_core_module_stays_domain_free_v1_parity_kernel; no BoundaryRule
names the limiter).

Duplicate guidance consolidated per rule 1: for the six crates holding both
AGENTS.md and CLAUDE.md (filesystem, network, secrets, mcp, sandbox, wasm),
CLAUDE.md stays canonical (module spec for filesystem; gate-pinned wording for
mcp and wasm) and AGENTS.md becomes a short pointer. No gate-pinned file was
edited. Stale reference removed: safety AGENTS.md pointed at
src/NETWORK_SECURITY.md, which exists nowhere in the tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(crates-map): rewrite the three top-level maps family-first after the restructure

crates/AGENTS.md (264 -> 175 lines): routing map only — the ten families,
the read order (family AGENTS.md -> crate README.md -> working rules/module
spec -> docs/reborn/contracts/), the enforced seven-layer matrix with the
family/layer divergences, measured workspace facts (64 packages, 1 documented
exclusion, 0 owned exceptions per scripts/ci/check-target-tree.py), and a
verified command block. The 40-row per-crate map is gone: family AGENTS.md
files own crate routing per docs/reborn/guidance-conventions.md.

crates/README.md (141 -> 119 lines): human map — mental model in family
vocabulary, the ten families with measured crate counts, the 14 extension
packages (4 crates + 10 data-only), and the two workspace members outside
crates/.

crates/Architecture.md (1019 -> 1059 lines): audited against the live tree;
every named symbol/path re-verified 2026-08-05. Corrected: retired
ProductAdapter vocabulary (zero residue in code), the stale pre-rename
dependency ladder that still cited the deleted gateway/TUI crates, run-state
store mentions, lane-table crate anchors (sandbox/extension_support),
declared-in vs minted-by owners in the core data model, and the subagent
deny-filter status note (re-verified). Marked the pre-restructure
'partial or evolving' list as unmeasured rather than asserting it.

Also documents that scripts/check-boundaries.sh fails on a clean tree
(check-5 grep false positives) and greps the deleted v1 src/ in 4 of 6
checks — boundary enforcement for crates/ is the architecture suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(crates-map): package directories carry their own README.md (coordinator sync with extensions-family agent)

Every extensions package dir — the 10 data-only ones included — now ships a
README.md, so both maps extend the read order to package level. The sibling
branch also confirmed what this map already derived per-crate: packages/ is
not uniformly products-layer (memory-native and mem0 declare substrates).
The other two coordinator corrections targeted rows of the old per-crate
map, which this rewrite deleted wholesale; nothing here cites
memory-native's CLAUDE.md or claims ironclaw_extension_host lacks guidance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(guidance): contracts + events family guidance layer per docs/reborn/guidance-conventions.md

- crates/contracts/AGENTS.md and crates/events/AGENTS.md rewritten to the
  family shape: exclusion lists with destinations, armed gates by test name,
  layer-matrix rows, crossing guide, measured header counts.
- README.md added for all 10 crates (ironclaw_prompt_envelope previously had
  no guidance of any kind — the CHECKLIST WS11 gap).
- One canonical guidance file per crate, other file a pointer:
  A+C merges for ironclaw_host_api, ironclaw_event_log,
  ironclaw_event_projections, ironclaw_event_streams; CLAUDE-only content
  moved to AGENTS.md for ironclaw_loop_contracts,
  ironclaw_extension_contracts, ironclaw_product_contracts (none of these are
  root module-spec crates, so AGENTS.md is the working-rules home).
- Stale guidance fixed against the live tree:
  * loop_contracts dep list contradicted the enforced allowlist (manifest is
    host_api + extension_contracts; common/prompt_envelope are permitted,
    unused).
  * event_log still documented the deleted jsonl parse/replay helpers.
  * event_projections still claimed EventStreamManager,
    DurableMemoryAuditSink, MemoryAuditProjectionMetadata, and
    PendingGateProjection — all deleted per PROPOSAL 6.3.3.
  * product_contracts still carried the pre-D-E open vendor decision under
    the nonexistent module name llm_config, and a Deferred section
    contradicting its own operator_llm/operator_service rows.
  * extension_contracts module table was missing the WS3 runtime module
    while counting 18.
  * common's llm_costs note carried the ModelCostTable seam claim refuted by
    PROPOSAL 12.11 D-F; now cites the pricer-port ruling and the vendor
    census residue.
- Deleted crates/events/ironclaw_event_projections/PENDING_GATE_PROJECTION.md:
  every claim in it referenced deleted symbols or the removed v1 src/ tree,
  and its only inbound reference was the crate's own CLAUDE.md.

Verified: all consumer counts reproduce via the printed grep commands; 147
path literals across the 28 touched files resolve on disk; no architecture
test reads any of these files by name; conflict-marker scan clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): three measured corrections surfaced by the guidance program

memory packages are substrates-layer, not products (families/extensions.md);
memory_native declares no extension_contracts dep (PROPOSAL §6.8.4); wasm's
extension_contracts edge is dev-only and the wasm 'never depends on' bullet is
lane-scoped, not family-wide (families/lanes.md).

Three further reported defects were checked and NOT corrected — they were
misreads: the sandbox 'never above the runtime tier' rule holds (substrates sit
below it), and PROPOSAL's safety consumer count already reads 17, matching the
tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(kernel): family guidance layer — perimeter AGENTS.md, nine crate READMEs, AGENTS/CLAUDE consolidation

Family-guidance program, kernel family (guidance-conventions.md shape):

- crates/kernel/AGENTS.md rewritten to the family shape: the nine-stage
  effect pipeline with stage ownership, the sealed-mint table (witness /
  trust ceiling / approval lease / verified-inbound evidence, each with its
  mint site and its seal mechanism), the per-stage fail-closed table with
  file:line or test citations, the sharp exclusion list, and the armed
  gates by test name (authorized-seal ratchet, sealed-evidence mint
  ratchet, BoundaryRules, same-layer edge inventory at 21 kernel edges,
  empty LAYER_MATRIX_EXCEPTIONS register, driver boundary, process storage
  scan, origin-gate matrix ratchet).
- A README.md for each of the nine crates, per the crate shape: measured
  workspace deps and consumer counts (cargo metadata), public surface with
  verified citations, enforced invariants naming their gates.
  ironclaw_processes states the single-lifecycle-authority direction of
  truth (journal = store; TurnRunState/ProcessRecord/await-edge =
  projections; PROPOSAL §12.13 D-S); ironclaw_host_runtime documents the
  D-R literal-loopback carve-out and names its two regression tests.
- Duplicate guidance reconciled in all nine crates: AGENTS.md is canonical
  (guardrails absorbed), CLAUDE.md reduced to a pointer; ironclaw_trust's
  CONTRACT.md untouched as the co-located cross-crate contract.
- Stale references fixed inside owned paths: the deleted capability-profile
  conformance module (evaluate_profile_conformance — zero hits
  workspace-wide) removed from ironclaw_capabilities guidance; trust's
  'staging branch' / 'PR3' phrasing updated; capabilities' 'later
  obligation slices' updated to the landed host_runtime obligations split;
  cross-crate path mentions fully qualified. Every path literal in all 29
  kernel .md files verified to resolve on disk; every named symbol swept
  against crate sources.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(domains): family guidance layer — AGENTS.md boundary doc, 12 crate READMEs, duplicate-guidance consolidation, stale-path fixes

Family guidance for crates/domains/ per docs/reborn/guidance-conventions.md:

- crates/domains/AGENTS.md rewritten to the family shape: charter table with
  go-here-when routing, the exclusion list, every armed gate named by test
  (BoundaryRules + identity/memory allowlists, the 5-entry in-family edge
  inventory, the naming gates, trusted-trigger ownership, the memory-provider
  residue ledger, persistence-driver boundary, the two module-charter gates).
- A measured README.md for each of the 12 crates: charter, use-when /
  don't-use-when routing, public surface, measured normal deps + named
  consumers, enforced invariants with their gates, exact test commands.
  ironclaw_attachments and ironclaw_identity had no guidance of any kind;
  identity's README points at CONTRACT.md (the module spec), llm's at its
  CLAUDE.md module spec.
- Duplicate guidance consolidated to one canonical file + pointer per crate:
  threads/conversations/memory/outbound rules now live in AGENTS.md (CLAUDE.md
  is a pointer); auth/llm keep CLAUDE.md canonical because their
  tests/module_charter.rs gates read it (AGENTS.md is the pointer). One
  misstatement fixed in the conversations merge: transcript content belongs to
  ironclaw_threads' SessionThreadService, not InboundConversationService.
- Staleness fixed inside the family: identity CONTRACT.md two-edge allowlist
  claim reconciled with D-Q's three entries; trace_commons CLAUDE.md gains the
  capture module row and strikes its two discharged Known Gaps (recording/paths
  shims deleted, rename done); llm CLAUDE.md reasoning.rs caller corrected to
  crates/loop/ironclaw_loop_host; triggers lib.rs 'feature-gated' repo doc
  comments corrected; pre-family path literals in comments repointed
  (kernel/approvals+processes, loop/hooks, app/architecture_tests,
  domains/auth) and the deleted-v1-engine references in skills marked
  historical.

Verified: cargo test -p ironclaw_llm --no-fail-fast (922 passed, exit 0 —
CLAUDE.md is gate-pinned); cargo check --all-targets on all six crates with
source edits; every cited path literal resolves on disk; conflict-marker scan
clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): repair the corrupted kernel bullet and correct two family laws

kernel.md: ironclaw_authorization's 'Security & authority role' bullet has been
textually corrupted since nearai#6918 — an approvals sentence was spliced into it
mid-clause, orphaning its continuation line. Reconstructed, with the spliced
sentence restored to the approvals entry where it is true.

lanes.md: 'a lane never depends on a substrate' is false as a family-wide law
(ironclaw_sandbox holds network/safety/secrets normal deps, which its own entry
licenses); the accurate law is the layer ladder, and the narrow claim holds for
ironclaw_wasm alone.

lanes.md + events.md: the 'every crate ships both an AGENTS.md and a CLAUDE.md'
requirement is superseded by docs/reborn/guidance-conventions.md — two files
restating one rule is the drift the guidance program removes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(arch): govern the ProtocolAuthEvidence test seam — WS12 audit F1

Two new gates in reborn_sealed_evidence_mint_ratchet (closed paths #12/#13),
per the audit's remedy spec:
(a) TEST_SEAM_MINT_FNS governs test_verified/test_verified_for_tenant — any
    production-text call site outside ironclaw_host_api is an offender
    (comments/strings stripped, #[cfg(test)] blocks stripped, tests.rs /
    *_tests.rs and cfg-test-only files excluded via the shared census);
(b) test-support may appear in no normal dependency table workspace-wide
    (dependencies / build-dependencies / target.* variants /
    workspace.dependencies), and no [features] key other than test-support
    may forward to it — the laundering shape that would evade (b) by one
    rename. [dev-dependencies] enablement stays legal (cargo-features.md
    bar 4, the sanctioned dev seam).

Measured zero offenders on this tree in both directions before pinning;
sabotage-proven red->green both ways (planted production call named with
file:line-text; [dependencies] enablement named with its table path).
Self-tests drive the same pipelines the gates run (zero-match principle);
the definition-location and partition tests now cover the new table.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(integration): join the gsuite credential-injection journey — WS12 audit F3

WS12 row 5 leg 3 was verified in two halves no committed test joined: gsuite
handler -> staged credential (crate tier) and staged obligation -> wire
(GitHub/Slack only). Scenario 5 already drives gmail.list_messages through
production dispatch on a Google-OAuth-configured group; it now also asserts
the JOIN: the seeded google account's token (itest-google-token) lands on
the recorded outbound gmail.googleapis.com request as
'authorization: Bearer ...', injected at the host egress chokepoint
(apply_credential_injection) per the gmail manifest's declared recipe —
store -> dispatch-time staging -> chokepoint -> wire, through the caller.

Sabotage-proven: disabling the Header injection arm reds exactly this
scenario with 'no network egress request matching url gmail.googleapis.com
has header authorization' while the request itself still reaches the wire
(headers seen: content-type only) — the injection reason, not a setup
error; restore -> green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): correct five measured dependency claims in families/domains.md

conversations does not depend on safety (its BoundaryRule now forbids it);
triggers depends on libsql_runtime + safety and NOT filesystem, so its
'filesystem-routed persistence path alongside SQL' is one path, not two;
memory's live set is host_api alone (prompt_envelope is allowlisted, unused);
auth was short by extension_contracts + product_contracts.

Each verified against the manifest before editing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(guidance): family AGENTS.md + crate READMEs + guidance consolidation for loop/product/app

Family-guidance program, families 8-10 (the top of the stack), per
docs/reborn/guidance-conventions.md:

- Rewrite crates/{loop,product,app}/AGENTS.md from routing stubs to the
  spec's family shape: exclusion lists, armed gates by test name, layer
  rows, crossing guides. Loop carries the trust story + the declared
  Loop*Port decorator chain; product carries the frozen-surface rule,
  the transports-consume-contracts rule (with the D-B frozen-constant
  qualification), the evidence-mint prohibition, and the two vendor
  exceptions; app carries the wires-owners-never-becomes-one charter,
  the binary-names-packages rule, config's zero-dep guarantee, and the
  composition mass ratchet (loc 40423 / Arc<dyn> 814).
- Add a README.md to all 13 crates (12 new; webui's rewritten to the
  spec shape) with measured public surface, deps, and consumer counts.
- Consolidate duplicate AGENTS.md/CLAUDE.md per spec rule 1: AGENTS.md
  is canonical and CLAUDE.md a pointer for agent_loop, loop_host,
  turn_runner, hooks, host_ingress, openai_compat, operator, and
  architecture_tests; CLAUDE.md stays canonical (module spec /
  gate-pinned) for webui, composition, and assistant, with
  composition's AGENTS.md reduced to the pointer.
- Fix stale references in owned paths: hooks' dependency diagram and
  AgentLoopDriver home (ironclaw_loop_contracts, not ironclaw_turns),
  loop_host/agent_loop port-home claims, turn_runner's pre-nearai#6696
  scheduler description, webui's ProductSurface path
  (product_contracts, not host_api), route count (93, measured), and
  webui's allowed-dependency list (7 of 10 were listed), the D-S
  await-edge ruling reflected in turn_runner guidance, composition's
  llm_admin residue (nearai_login_serve left for operator).

Verified: cargo test -p ironclaw_architecture_tests --no-fail-fast
(39 binaries, 0 failures — covers the CLI AGENTS.md phrase pin and the
composition guidance-markdown scan), scripts/ci/check-target-tree.py,
path-literal resolution over all 37 changed files, conflict-marker scan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): record the closed scan evasions (F4) and the secrets-consumer correction (F2)

The sealed-mint census weaknesses PROPOSAL §11.2.5/§12.1a and CHECKLIST recorded
as live and owed to WS10 are all closed on this tree, verified by re-attacking
the seam with both evasions at once; the docs understated the seal. Ratchet is
23 tests. One residual replaces them: the test_verified test-seam constructors,
now pinned by two gates.

§12.1b's 'only products-layer crate with the edge' is false by one —
ironclaw_assistant carries ironclaw_secrets as port-declaration vocabulary with
no expose_secret call. Not a value-reach bypass; joins nearai#7095's inventory.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): correct the app-family layer, config's consumer set, and the webui route count

ironclaw_config declares layer=substrates while living in crates/app/;
its consumers include operator, extension_manager and extension_host, not just
the assembly crate and the binary; webui is 93 contract-locked routes, not 92
(nearai#6780 landed after the last recount).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(stale-sweep): fix agent guidance outside crates/ for the family restructure

Audit-and-fix pass over every stale document outside crates/ (PR 2 of the
family-guidance program). Live guidance verified against the tree; records
kept with dated notes instead of rewrites.

Guidance fixes (verified against HEAD before writing):
- .claude/commands/trace.md: MCP tool prefix codebase-memory -> codebase-memory-mcp
  (allowed-tools never matched the real server), ProductSurface home ->
  ironclaw_product_contracts, capabilities host.rs -> host/ module split,
  scripts lane -> script-sandbox; deleted the redundant v1-anchors section.
- .claude/commands/add-sse-event.md: deleted the banner-quarantined v1 scaffold
  steps (every path deleted with the monolith); now an honest redirect to the
  Reborn projection/SSE path. Frontmatter no longer advertises a working scaffold.
- .claude/commands/deslop-reborn.md: three dead crates/*/Cargo.toml globs (family
  layout added a level), ls crates/ -> family-aware listing, v1-only consumer
  logic retired, per-crate --features integration phrasing.
- .claude/rules/type-placement.md: crates/*/src globs matched nothing; recipes
  re-pointed and numbers re-measured 2026-08 (3,495 structs/enums, 385 traits,
  fan-in host_api 53 / common 20 / turns 12).
- .claude/rules/skills.md: paths trigger pointed at a nonexistent
  bundled_skills.rs (rule never fired); SKILLS_REGEX_ACTIVATION_ENABLED /
  SKILLS_MAX_TOKENS env vars are read by nothing -> documented the real
  config-file setting and DEFAULT_MAX_SKILL_CONTEXT_TOKENS.
- .claude/rules/testing.md, ironclaw-reborn-testing skill, CONTRIBUTING.md,
  .github/pull_request_template.md, testing-playbook, deslop: the workspace-root
  `integration` feature is empty with zero consumers - all "cargo test
  --features integration" guidance re-pointed to crate-level suites.
- .claude/skills/reborn-extension-surfaces: four pre-colocation assets/ paths,
  CapabilitySurfaceKind home, conformance-suite move to
  ironclaw_extension_contracts, ingestion test move to the registry crate,
  gate-banned migration exemplar replaced with the live behavioral pin, [mcp]
  instead-of claim softened (nearai-mcp pins a static [[tools]]).
- .claude/skills/ironclaw-reborn-orientation: turn_runner labels, prompt-crate
  list re-derived (turns/first_party_extension_ports out; host_api,
  loop_contracts, assistant in), consumer-grep glob fixed.
- .claude/skills/reborn-feature + docs/reborn/how-to-port-channel-to-reborn.md:
  ProductSurface/ProductView/descriptors/caller types live in
  ironclaw_product_contracts; recipes re-pointed.
- CLAUDE.md: dead root --features integration line replaced; project tree
  redrawn with the ten families; trait homes corrected; ProviderId -> VendorId;
  CapabilitySurfaceKind + ChannelAdapter homes; [channel.config] ->
  [channel.connection]/[admin_configuration]; v1 Job State Machine section
  deleted (no such machine in Reborn); prompt-crates recipe fixed; MCP server
  name; LLM backend list re-derived from LlmBackendKind.
- docs/extensions/building-a-tool.md: product-adapter crates row -> channel
  surface model; package registration -> PACKAGES collector in
  ironclaw_extension_support (available_extensions.rs is being dissolved);
  hosted-MCP policy home -> ironclaw_extension_host/src/mcp.rs; dead v1 bullets
  dropped.
- docs/internal/mutation-audit.md: runnable command blocks re-pointed (family
  paths; ironclaw_dispatcher example replaced - crate deleted in WS0).
- docs/reborn/harness/e2e.md: dispatcher row -> the capabilities dispatch
  contract suites. docs/reborn/contracts/host-api.md: three ironclaw_dispatcher
  mentions -> capabilities dispatch module. standard-operations.md: renamed
  crate + arch-test package name.
- scripts: mutation-audit.sh usage header, check-hermetic-env.sh env_helpers
  pointer, check-generic-without-concrete.sh mirror pointer,
  telegram_smoke/README regression step (target deleted with v1 in nearai#6375).
- .env.example: dead SKILLS_REGEX_ACTIVATION_ENABLED entry -> config-file doc.
- docs/qa/telegram-coverage-map.md: nine not-automated reasons re-worded to the
  crate-level integration tier.

Records (dated notes, no rewrites): ADR 0003/0004 path notes (evidence pinned
to their measured SHA), FEATURE_PARITY state-migration paragraph marked
historical with a git-show recovery pointer, engine-v2 parity record's
"coexist on main" claim corrected with a historical note, subagent-spawn
legacy scope re-tensed.

Pre-family path reproduction count: 73 -> 70 files; every remaining file is a
dated record (docs/plans, docs/superpowers, ADRs, audits, CHANGELOG history,
historical-marked train docs) or a deliberate past-tense mention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(ws12): tick row 7 — the fresh-agent placement probe passed on the final tree

All three placements correct with high confidence, each naming the trait, the
tests, and the tempting wrong place it rejected. The probe doubled as a docs
audit and independently hit four defects, three of which the stacked guidance
PR fixes — it succeeded despite them.

WS12 is now 7/7. The restructure is complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(target-arch): the product→loop_host recount was wrong on the day it was written

Eight importing files across four seams, not seven across three — the fourth
being a skill-activation-observer seam (projection.rs, projection/live_progress.rs)
this bullet never named, which §6.4.7's own same-day note already implied.
Surfaced by the plan-conformance audit.

The recount history is 3→5→6→7→8, wrong at four of five attempts. That retires
the prose count as a method: the sever slice should land an inventory ratchet
before or with the move, not another number.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review(7263): CodeRabbit round-1 triage — 4 code fixes (2 sabotage-proven, 2 red-first) + 6 doc-truth corrections

Code, each verified red-first or by sabotage matrix:
- sealed-mint ratchet: per-name sighting floor for TEST_SEAM_MINT_FNS
  (closed path #12). Proven: renaming test_verified_for_tenant away plus one
  extra legitimate sibling mention passed the old aggregate floor (silent
  disarm) and fails the new per-name floor naming the constructor; suite
  23/23 after revert. (CodeRabbit's claimed baseline ">2 mentions today" is
  wrong — each name has exactly one kept sighting — but the doc/enforcement
  mismatch and at-threshold fragility were real.)
- trace credit: non-finite novelty_score/duplicate_score are treated as
  absent before clamping (clamp preserves NaN, which poisoned online_score
  and credit_points_estimate); NaN cases added to the nearai#7144 regression test,
  red first.
- trace submission: a 2xx whose body stream dies mid-read now maps through
  request_failed (network telemetry kind, true I/O cause) instead of
  collapsing to an empty body that the nearai#7144 strict parse misreported as
  response_invalid/Submission; truncated-body regression test, red first.
- Postgres contract suites (event store + assistant ledger): isolated-DB
  names now carry a creation epoch and the once-per-binary sweep is
  age-gated (1h), closing the cross-process window where a sibling's fresh
  zero-backend database (between CREATE DATABASE and first connection) was
  sweepable; legacy pid-scheme leftovers still collect immediately. Proven
  on live Postgres 16: planted stale name swept, planted fresh name
  survives, 13/13 x2 and 20/20 x2 with zero leftovers.

Docs (target-architecture truth pass):
- PROPOSAL section 9: the WS6 rename sweep (nearai#7152) had rewritten the source
  column of the 12 renamed rows to their post-rename names, turning their
  rename dispositions into no-ops (rows 13/14/28/30/49/51/59/61/64/66/67/70);
  pre-restructure names restored with a dated footnote.
- PROPOSAL:69: removed the superseded 3->5->6->7 recount sentence (the
  corrected 3->5->6->7->8 passage subsumes it).
- PROPOSAL row 34: ToolPermissionOverrideStorePort deletion marked landed
  (2026-08-05 WS8, matching section 6.5.3; zero workspace hits).
- CHECKLIST:631: dated note recording that the WS12 F3 gsuite join landed in
  this batch (scenario_uninstalled_tool_call_denied_until_active.rs asserts
  the seeded google token on the gmail.googleapis.com wire; suite run green).
- CHECKLIST:632: dated note spending F4 (the audit's 19 was correct at its
  SHA; the ratchet file now holds 23 tests, re-counted at lines 552/597).
- ws12-gauntlet-report P6 heading: first of TWO real failures (one class),
  matching P8 and the report's own summary.
- ws12-mapping-audit rows 49/137: dated D-S closure notes (await-edge store
  half = journal projection already; resolver retained loop-tier; no shed
  owed) so the backlog register no longer lists it as in-flight.

Not fixed, with evidence: the span-helper macros gate suggestion
(info_span!(target = ...) is a hard compile error, E0425 — no silent trap),
the webui tracing-subscriber workspace-dep suggestion (no
[workspace.dependencies] entry exists; suggestion would not build; 8
siblings use the identical direct shape), and the mapping-audit
regeneration (the audit is accurate at its pinned SHA; the in-batch F1 fix
is recorded in its dated coordinator note).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review(7263): CodeRabbit round-2 — rejection-body read keeps its cause; 200 {} is not a submission acknowledgement; lanes.md family dep rule matches measured Cargo.tomls

- submission.rs non-2xx path: a failed rejection-body read no longer collapses
  to an empty detail via .unwrap_or_default() (banned by
  .claude/rules/error-handling.md); the read error folds into the
  http_rejection detail so the received status keeps driving the 401/403
  auth-retry and the Credential/HttpRejection telemetry split.
  Regression: submit_preserves_rejection_body_read_failure_cause_with_status.

- TraceSubmissionReceipt.status: serde default removed — it fabricated
  status "submitted" from a proxy's 200 {} (the nearai#7144 synthesis, resurfacing
  through the wire type's defaults), after which the flush caller recorded
  Submitted and deleted the only retryable queued copy. The acknowledgement is
  the server naming what happened to the submission — every workspace fixture
  sends status and callers persist it unconditionally as server_status — so a
  status-less 2xx body now fails the strict receipt parse as response_invalid.
  Regression: submit_rejects_success_response_without_explicit_server_status
  (covers 200 {} and a status-less non-empty object).

- docs(lanes.md): the family Dependency-direction rule no longer claims every
  lane takes the extension-surface vocabulary crate — measured across
  crates/lanes/*/Cargo.toml: mcp + sandbox hold ironclaw_extension_contracts
  under [dependencies], wasm only under [dev-dependencies]; dated ✎
  cross-references the ironclaw_wasm entry's 2026-08-05 correction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review(7263): CodeRabbit round-3 — shared Postgres test provisioner (the "new dep edge" premise measured false), entrypoint self-test armed (sabotage-proven), six doc self-contradictions reconciled

Code:
- ironclaw_filesystem gains a `postgres_isolation` test-support module — the
  single home of the per-test isolated-database scaffolding (once-per-binary
  age-gated stale sweep, epoch-in-name convention, DROP WITH (FORCE) cleanup),
  parameterised by suite/env-var/prefix/unreachable-policy. Zero new
  production edges: event_store already normal-deps filesystem, filesystem
  already owns tokio-postgres, and the dev-dep+feature pattern is the one 17
  crates already use. The event-store and product-workflow-ledger suites
  migrate onto it; both Postgres legs proven live against postgres:16 (12
  tests, zero leftover databases). The fabric original keeps its older
  variant with the differences documented at its IsolatedDatabase.
- ironclaw_event_store drops the duplicate tokio-postgres dev-dep (the normal
  dep already reaches tests).
- test-reborn-docker-entrypoint.sh: the missing-argv check now exits the
  command-substitution subshell instead of incrementing a counter the parent
  never sees — red-proven (a migrate-but-never-exec entrypoint passed with 7
  FAIL lines printed), green after the fix both sabotaged and restored.
- trace_commons submission test additionally pins !auth_rejection() for the
  503 rejection (the structural assert the API affords; the prescribed
  payload asserts are refuted — status is private and source is None by
  design, with the message derived from the structured status in the same
  constructor).

Docs (each reconciled to one canonical statement, measured):
- kernel.md: lease ownership decided from code — authorization stores,
  matches, and expires leases (CapabilityLeaseStore + port + expiry all live
  there); approvals constructs and issues into that store. The round-1
  re-homing of the spliced sentence into approvals was wrong and is corrected
  in the dated repair note.
- app.md: "nothing depends on app" scoped to the three app-layer crates;
  ironclaw_config's consumers restated by dependency kind (normal:
  composition, cli, operator, extension_host; dev-only: extension_manager,
  root integration-tests package).
- lanes.md: the mediated-services sentence now states the family law as
  layer-ladder + injected authority; the no-secrets/network/filesystem-dep
  claim is scoped to ironclaw_wasm, matching the file's own corrections.
- CHECKLIST 429/430: the one open traces clause is named (ScopedFilesystem
  adoption); the stale "other two" count corrected against the F3a strike.
- PROPOSAL:69 + CHECKLIST:72: the project-create route repointed —
  first_party_extension_ports dissolved into loop_host::skill_activation
  (WS8, §9 row 55) — still unattempted.
- PROPOSAL §9 rows 57/62 synced to §6.8.4 (telegram: dependency-set equality
  with Slack's four contract-tier crates) and §6.9.4 (webui -> assistant is a
  charter-permanent edge, §12.11 D-B).
- PLAN top summary records Wave 6's design question as resolved (D-S,
  2026-08-05).
- deploy-reborn-cli-docker.md: the two migration paragraphs unified on the
  entrypoint's actual behavior — only enabled = false beside
  signing_secret_env/bot_token_env is migrated; every other retired-key shape
  fails startup with the migration pointer.
- composition-budget.toml: the stale "2398 bp, a true ratchet" header
  replaced with the WS0-floor truth the baselines test asserts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: move the guidance convention into this PR so its citations resolve

families/lanes.md and families/events.md cite docs/reborn/guidance-conventions.md
when superseding their 'every crate ships both an AGENTS.md and a CLAUDE.md'
requirement, but the file was only on the stacked guidance branch — a forward
reference that dangles if this PR merges alone. The convention is the rule those
notes invoke, so it belongs with them.

Caught by the CodeRabbit round-3 pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): give the hoisted postgres provisioner its safety rationales

The round-3 hoist moved test provisioning into a production src/ path, so
check_no_panics flagged its four panic/expect sites and reddened Code Style via
fast-checks. The gate is right to flag them: it deliberately does NOT exempt
#[cfg(feature = "test-support")] modules, because a cargo feature is not a
privilege boundary in this workspace (PROPOSAL 12.1a proved exactly that) —
so a test-support module still compiles into a build where any sibling enables
the feature.

Suppressed with the gate's documented inline rationale, which must trail the
statement rather than precede it. The panics themselves stay: a configured but
unusable Postgres must fail the suite loudly rather than skip it, which is the
inert-guard rule the isolation fix exists to serve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): classify the three planner-unknown paths this PR touches

The Reborn PR test planner fails closed on any unclassified path and
raises on the FIRST failure in sorted order, so CI only ever showed
.github/pull_request_template.md. Classifying that unmasked two more
paths in this PR's own diff: scripts/mutation-audit.sh and
scripts/telegram_smoke/README.md. All three are classified; the
fail-closed arm is untouched:

* .github/pull_request_template.md -> IGNORED_PREFIXES, beside its
  exact sibling .github/ISSUE_TEMPLATE/ (both GitHub UI templates;
  classify-test-scope.sh already pairs them in its docs-only arm).
* scripts/mutation-audit.sh -> PR_STATIC_CONTROL_PATHS, beside its
  self-test scripts/test-mutation-audit.sh; both run only in
  nightly-deep-ci.yml's mutation-frontier job.
* scripts/telegram_smoke/ -> QA_HARNESS_PREFIXES; a live, by-hand
  release smoke harness referenced by no workflow, same class as
  scripts/reborn_qa_matrix/.

Each entry is pinned red-first in test_reborn_pr_test_plan.py (entry
commented out, new assertion fails with the exact production error,
entry restored, green): a new PR-template test with paired
accept-AND-select-nothing assertions plus unknown-.github/-sibling
refusal probes, and the two existing class tests extended. Planner
self-test: 65 tests OK. The planner CLI over this PR's full 209-path
diff now exits 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: regular Contributor history classification risk: low Risk classification size: XL Changed-line size classification

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants