Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
44 commits
Select commit Hold shift + click to select a range
980d60e
[codex] Stabilize auth readiness and gate flows (#2050)
ilblackdragon Apr 9, 2026
af9b59a
feat: unified tool dispatch + schema-validated workspace (#2049)
ilblackdragon Apr 9, 2026
b819d70
feat(web): add scroll-to-bottom arrow in gateway chat (#2202)
ilblackdragon Apr 9, 2026
9399fcc
fix(auth) first-pass Gmail OAuth auth prompt in chat (#2038)
serrrfirat Apr 9, 2026
e0bdd74
feat(admin): admin tool policy to disable tools for users (#2154)
serrrfirat Apr 9, 2026
aaeb904
feat(tui): ship TUI in default binary (#2195)
serrrfirat Apr 9, 2026
26e5e4c
feat(docker): pre-bundle WASM extensions in staging image (#2210)
henrypark133 Apr 9, 2026
6a8e581
fix(wasm): upgrade Wasmtime to 43.0.1 and restore CI (#2224)
henrypark133 Apr 10, 2026
580165c
feat(railway): build staging target with pre-bundled WASM extensions …
henrypark133 Apr 10, 2026
2d8e6e6
fix(ci): resolve 3 staging test failures (#2207)
henrypark133 Apr 10, 2026
e0e0fcd
fix(agent): stop intercepting bare yes/no/always as approval when not…
henrypark133 Apr 10, 2026
b9b239e
fix(gateway): suppress duplicate text response during auth flow and u…
henrypark133 Apr 10, 2026
efdb738
fix(docker): consume CACHE_BUST arg so BuildKit invalidates cache
henrypark133 Apr 10, 2026
bd7f2b7
Create QA Bug Report issue template (#2228)
joe-rlo Apr 10, 2026
4147c6d
feat(gateway): extract gateway frontend into ironclaw_gateway crate w…
ilblackdragon Apr 10, 2026
55cdbf2
fix(docs): explain in more details `activation` block & installation …
denbite Apr 10, 2026
e5b82fc
fix(bridge): sanitize auth_url on engine v2 path (#2206) (#2215)
ilblackdragon Apr 10, 2026
494636d
docs: add amazon tutorial (#2261)
matiasbenary Apr 10, 2026
a8e6533
fix(oauth): use localhost for redirect URI when bound to 0.0.0.0 (#2247)
ilblackdragon Apr 10, 2026
8dfedfa
feat: add native Composio tool for third-party app integrations (#920)
vutran1710 Apr 10, 2026
cd8f3f2
feat(skills): commitments system — active intake for personal AI assi…
ilblackdragon Apr 10, 2026
56eb0ad
ci: trigger ironclaw-dind image build (#2190)
think-in-universe Apr 10, 2026
f37a26f
fix(engine): mission cron scheduling + timezone propagation (#1944) (…
ilblackdragon Apr 10, 2026
152e8b0
feat: add extensible deployment profiles (IRONCLAW_PROFILE) (#2203)
serrrfirat Apr 10, 2026
2cc5546
feat(tools): production-grade coding tools, file history, and skills …
ilblackdragon Apr 10, 2026
1b8e1cc
fix(docker): copy profiles/ into build stages (#2289)
henrypark133 Apr 10, 2026
b4502cf
fix(ci): resolve 4 staging test failures (#2273)
henrypark133 Apr 10, 2026
33580d4
fix(v2): tool naming, auth gates, schema flatten, WASM traps, workspa…
ilblackdragon Apr 10, 2026
2f2fe30
fix(test): case-insensitive hint matching in TraceLlm step_matches (#…
henrypark133 Apr 10, 2026
4896996
Merge pull request #2293 from nearai/staging-promote/2f2fe306-2426380…
henrypark133 Apr 10, 2026
7cc172d
Merge pull request #2290 from nearai/staging-promote/b4502cf9-2425811…
henrypark133 Apr 10, 2026
fe37205
Merge pull request #2278 from nearai/staging-promote/2cc55460-2425498…
henrypark133 Apr 10, 2026
0f700af
Merge pull request #2272 from nearai/staging-promote/f37a26f7-2425259…
henrypark133 Apr 10, 2026
db449ee
Merge pull request #2266 from nearai/staging-promote/56eb0adf-2425000…
henrypark133 Apr 10, 2026
f9639af
Merge pull request #2264 from nearai/staging-promote/a8e6533a-2424764…
henrypark133 Apr 10, 2026
66c7de1
Merge pull request #2256 from nearai/staging-promote/e5b82fc6-2424031…
henrypark133 Apr 10, 2026
1538cf8
Merge pull request #2253 from nearai/staging-promote/55cdbf2b-2423827…
henrypark133 Apr 10, 2026
ac1f301
Merge pull request #2251 from nearai/staging-promote/4147c6d5-2423615…
henrypark133 Apr 10, 2026
8c55fee
Merge pull request #2243 from nearai/staging-promote/efdb738a-2422991…
henrypark133 Apr 10, 2026
c5a9237
Merge pull request #2226 from nearai/staging-promote/b9b239ee-2422188…
henrypark133 Apr 10, 2026
30b2d17
Merge pull request #2218 from nearai/staging-promote/26e5e4cf-2421365…
henrypark133 Apr 10, 2026
c8d824d
Merge pull request #2213 from nearai/staging-promote/e0bdd74f-2420621…
henrypark133 Apr 10, 2026
e462b5e
Merge pull request #2208 from nearai/staging-promote/9399fccc-2420383…
henrypark133 Apr 10, 2026
de20175
Merge pull request #2205 from nearai/staging-promote/b819d704-2420134…
henrypark133 Apr 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
13 changes: 8 additions & 5 deletions .claude/rules/skills.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,16 +31,19 @@ activation:
tags:
- "devops"
max_context_tokens: 2000
metadata:
openclaw:
requires:
bins: [docker, kubectl]
env: [KUBECONFIG]
requires:
bins: [docker, kubectl]
env: [KUBECONFIG]
---

# Skill instructions here...
```

Only the top-level `requires:` block is supported. The legacy nested shape
`metadata.openclaw.requires` is unsupported and ignored by the current parser,
so older external skills must be migrated instead of relying on silent
compatibility.

## Selection Pipeline

1. **Gating** -- Check binary/env/config requirements; skip skills whose prerequisites are missing
Expand Down
41 changes: 41 additions & 0 deletions .claude/rules/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,3 +23,44 @@ Run `bash scripts/check-boundaries.sh` to verify test tier gating.
- Use `tempfile` crate for test directories, never hardcode `/tmp/`
- Regression test with every bug fix (enforced by commit-msg hook)
- Integration tests (`--test workspace_integration`) require PostgreSQL; skipped if DB is unreachable

## Test Through the Caller, Not Just the Helper

**When a helper gates a side-effecting flow, the test must go through the caller — not just the helper in isolation.**

A whole class of bugs in this repo has the same shape: a wrapper function silently loses one of its inputs, and the unit test for the helper passes because it never crosses the layer where the input gets dropped.

Real examples (do not let these recur):

| Bug | Helper | What got lost | How a caller-level test would have caught it |
|-----|--------|--------------|------------------------------------------------|
| nearai/ironclaw#1948 | `McpServerConfig::has_custom_auth_header()` | Helper existed but `requires_auth()` never consulted it, so MCP triggered OAuth/DCR even with a user-set `Authorization` header | A test driving `mcp::factory::create_client_from_config()` with a header-bearing config and asserting zero OAuth-state side effects |
| nearai/ironclaw#1921 | `derive_activation_status(ext, has_owner_binding)` | Wrapper hardcodes the underlying classifier's `has_paired` axis to `false`, even though `classify_wasm_channel_activation` takes both bools | A test driving `extensions_list_handler` against a DB with a real `channel_identities` row and asserting `Active`, not `Pairing` |
| nearai/ironclaw#1502 | `window.open` mock `(url) => { window._lastOpenedUrl = url }` | Mock captured only the URL, silently swallowing `target` and `windowFeatures`; a regression to same-tab open would not fail | A mock capturing all three args plus an assert that `target === '_blank'` |

### When the rule applies

You must add a caller-level test (not just a helper-level unit test) when **all** of the following are true:

1. The helper is a **predicate, classifier, or transform** whose return value gates a side effect (HTTP call, DB write, UI mutation, OAuth flow, secret read, tool execution, sandbox launch, etc.).
2. There is **at least one wrapper or call site** between the helper and the side effect.
3. The helper has **more than one input** *or* its caller computes any of the inputs from the surrounding context.

If all three are true, a unit test on the helper alone is **not sufficient regression coverage**. You must additionally either:

- Add a test that drives the call site (`*_handler`, `factory::create_*`, `manager::*`), **or**
- Inline the helper into its single caller so there is no wrapper to silently drop an input.

### Where the test belongs

Most of these gaps are above unit-test scope and below e2e scope. Default to the **integration tier** (`cargo test --features integration`):

- `tests/<module>_integration.rs` for Rust integration tests against the public handler/factory surface
- `tests/multi_tenant_integration.rs` when the lost axis is per-user state
- `tests/e2e/scenarios/test_*.py` when the lost axis is browser-visible

Unit tests in `mod tests {}` are still fine for the helper itself, but they do not satisfy this rule.

### Mock hygiene corollary

When you mock a browser/runtime API in a test, the mock's signature must match the production call site's signature, and assertions should cover **every argument** the production code passes. A `(url) => {}` stub for a `window.open(url, target, features)` call site is a silent argument-loss bug waiting to happen.
101 changes: 101 additions & 0 deletions .claude/rules/tools.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@
paths:
- "src/tools/**"
- "tools-src/**"
- "src/channels/**"
- "src/cli/**"
---
# Tool Architecture

Expand Down Expand Up @@ -37,3 +39,102 @@ impl Tool for MyTool {
fn requires_sanitization(&self) -> bool { true } // External data
}
```

## Everything Goes Through Tools

**All actions originating from any non-agent caller — gateway handlers, CLI
commands, routine engine, WASM channels, future channel extensions — MUST
go through `ToolDispatcher::dispatch()`, never directly through the
database, workspace, or domain managers.**

This is the core design principle behind #2049. The reasons are concrete:

1. **Audit trail.** Every dispatched call creates an `ActionRecord` linked
to a system job, so UI-initiated mutations are visible in job history
alongside agent-initiated ones. Direct DB calls bypass this entirely.
2. **Safety pipeline parity.** The dispatcher runs the same pipeline as
`Worker::execute_tool`: parameter normalization, schema validation,
`sensitive_params()` redaction, per-tool timeout, output sanitization.
Direct calls skip all of it and risk leaking secrets into logs or
persisting unsafe content.
3. **Channel-agnostic.** Channels are interchangeable extensions (gateway,
CLI, telegram, WASM, future custom channels). Routing through a single
dispatch function means new channels inherit the full pipeline for free.
4. **Agent parity.** The agent can do anything channels can do (and vice
versa), because both call the same tools. No more "the UI can install
extensions but the agent can only list them" gaps.

### Required pattern

```rust
// In any gateway handler, CLI command, or routine engine callback:
use crate::tools::dispatch::{DispatchSource, ToolDispatcher};

let dispatcher: &ToolDispatcher = state
.tool_dispatcher
.as_ref()
.ok_or((StatusCode::SERVICE_UNAVAILABLE, "dispatcher unavailable"))?;

let output = dispatcher
.dispatch(
"memory_write",
serde_json::json!({ "target": path, "content": content }),
&user.user_id,
DispatchSource::Channel("gateway".into()),
)
.await
.map_err(|e| (StatusCode::INTERNAL_SERVER_ERROR, e.to_string()))?;
```

### Forbidden pattern

```rust
// DO NOT do this in a gateway handler, CLI command, or routine callback:
let store = state.store.as_ref().ok_or(...)?;
store.set_setting(&user.user_id, &key, &value).await?; // BYPASSES dispatch

let workspace = resolve_workspace(&state, &user).await?;
workspace.write(path, content).await?; // BYPASSES dispatch + safety pipeline

let ext_mgr = state.extension_manager.as_ref().ok_or(...)?;
ext_mgr.install(name, url, kind, &user.user_id).await?; // BYPASSES audit trail
```

### When direct access IS allowed

The dispatch principle applies to **non-agent callers** acting on behalf of
a user. These are exempt:

| Layer | Why exempt |
|---|---|
| `Worker::execute_tool()` (agent loop) | Has its own atomic sequence-numbered audit trail; the dispatcher would conflict |
| `EffectBridgeAdapter::execute_action()` (v2 engine) | Same — its own audit via `ThreadEvent` event sourcing |
| The tool implementations themselves | Tools are the leaves; they need direct `Workspace`, `Database`, etc. handles to do their work |
| Background jobs (scheduler, hygiene, mission runner) inside the engine | These ARE the engine; they emit their own events |
| Pure read endpoints that need to JOIN/aggregate from multiple sources | A single tool call cannot express "list all jobs across users with filters X, Y, Z" — these are queries, not actions, and the audit value is low |

### Annotating intentional exceptions

If a handler legitimately needs direct access (rare — usually only for
read aggregation), suppress the pre-commit check with a trailing comment
on the offending line:

```rust
let rows = state.store.list_agent_jobs().await?; // dispatch-exempt: read-only aggregation
```

The pre-commit hook (`scripts/pre-commit-safety.sh`) flags any newly
added line in `src/channels/web/handlers/*.rs` or `src/cli/*.rs` that
touches `state.{store,workspace,workspace_pool,extension_manager,
skill_registry,session_manager}.*` without a trailing
`// dispatch-exempt: <reason>` comment on the same line. The check only
looks at added lines (`+` lines in the diff), so existing untouched code
doesn't trip it during incremental migration.

### Migration status

As of #2049, `ToolDispatcher` is wired into `GatewayState` but per-handler
migration is incomplete. New handlers MUST use the dispatcher. Existing
handlers should be migrated incrementally; each handler family
(settings, memory, extensions, skills, routines, jobs, threads) is its
own follow-up PR.
1 change: 0 additions & 1 deletion .dockerignore
Original file line number Diff line number Diff line change
Expand Up @@ -5,4 +5,3 @@ target/
*.md
!CLAUDE.md
node_modules/
tools-src/
2 changes: 1 addition & 1 deletion .githooks/pre-commit
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ if $NEEDS_CHECK; then
fi

# i18n parity: when any language pack changes, all languages must stay in sync.
if echo "$STAGED" | grep -qE '^src/channels/web/static/i18n/.*\.js$'; then
if echo "$STAGED" | grep -qE '^crates/ironclaw_gateway/static/i18n/.*\.js$'; then
echo "pre-commit: checking i18n parity..."
if ! ./scripts/check-i18n-parity.sh; then
echo ""
Expand Down
94 changes: 94 additions & 0 deletions .github/ISSUE_TEMPLATE/qa-bug.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,94 @@
name: QA Bug Report
description: Bug found during QA testing on staging or hosted environments
title: "[QA] "
labels: ["qa-bug"]
body:
- type: dropdown
id: environment
attributes:
label: Environment
description: Where was this bug found?
options:
- hosted-staging (crab shack)
- hosted-production
- local (cloned ironclaw)
- railway-staging
validations:
required: true

- type: input
id: version
attributes:
label: Version / Commit Hash
description: Paste the commit hash from staging at time of discovery (run `git rev-parse HEAD` or find it on the Railway deploy)
placeholder: "e.g. abcdef1"
validations:
required: true

- type: input
id: qa-date
attributes:
label: QA Test Date
description: Date you discovered this (YYYY-MM-DD)
placeholder: "e.g. 2026-04-12"
validations:
required: true

- type: input
id: feature-area
attributes:
label: Feature Area
description: What part of the app? (e.g. Google Suite extension, Telegram pairing, auth flow)
placeholder: "e.g. Extensions → Google Suite install"
validations:
required: true

- type: textarea
id: steps
attributes:
label: Steps to Reproduce
description: Exact steps — numbered, specific, no summaries
placeholder: |
1. Open extensions tab
2. Click "Install Google Suite"
3. Fill in credentials and click Save
4. ...
validations:
required: true

- type: textarea
id: expected
attributes:
label: Expected Behavior
description: What should happen?
validations:
required: true

- type: textarea
id: actual
attributes:
label: Actual Behavior
description: What actually happened? Include the exact error message/text.
placeholder: "Error: 'Failed to authenticate with Google: invalid_grant' shown in red toast"
validations:
required: true

- type: textarea
id: logs
attributes:
label: Logs / Screenshots
description: Paste relevant logs, error output, or attach screenshots. Drag files here.
validations:
required: false

- type: checkboxes
id: checklist
attributes:
label: Pre-submit checklist
options:
- label: Title is specific (not "fix Google Suite" but "Google Suite install throws invalid_grant on OAuth step")
required: true
- label: Commit hash is filled in
required: true
- label: Steps are numbered and reproducible
required: true
38 changes: 38 additions & 0 deletions .github/workflows/docker.yml
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,13 @@ jobs:
echo "tags=${TAGS}" >> "$GITHUB_OUTPUT"
echo "worker_tags=${WORKER_TAGS}" >> "$GITHUB_OUTPUT"

# Staging builds get pre-bundled WASM extensions
if [[ "${EVENT_NAME}" == "schedule" || "${INPUT_TAG}" == "staging" ]]; then
echo "target=runtime-staging" >> "$GITHUB_OUTPUT"
else
echo "target=runtime" >> "$GITHUB_OUTPUT"
fi

- name: Set up Docker Buildx
uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3

Expand All @@ -100,6 +107,7 @@ jobs:
context: .
push: true
tags: ${{ steps.tags.outputs.tags }}
target: ${{ steps.tags.outputs.target }}
platforms: linux/amd64
cache-from: type=gha
cache-to: type=gha,mode=max
Expand All @@ -115,6 +123,36 @@ jobs:
cache-from: type=gha,scope=worker
cache-to: type=gha,mode=max,scope=worker

- name: Create releases-manager app token
id: app-token
continue-on-error: true
uses: actions/create-github-app-token@fee1f7d63c2ff003460e3d139729b119787bc349 # v2
with:
app-id: ${{ secrets.GH_RELEASES_MANAGER_APP_ID }}
private-key: ${{ secrets.GH_RELEASES_MANAGER_APP_PRIVATE_KEY }}
owner: nearai
repositories: ironclaw-dind

- name: Trigger ironclaw-dind Build & Push
if: steps.app-token.outcome == 'success'
continue-on-error: true
env:
GH_TOKEN: ${{ steps.app-token.outputs.token }}
EVENT_NAME: ${{ github.event_name }}
INPUT_TAG: ${{ inputs.tag }}
VERSION: ${{ steps.version.outputs.version }}
run: |
if [[ "${EVENT_NAME}" == "workflow_call" && -n "${VERSION}" ]]; then
gh api repos/nearai/ironclaw-dind/dispatches \
--method POST \
-f event_type="ironclaw_image_published" \
-f client_payload[version]="${VERSION}"
elif [[ "${EVENT_NAME}" == "schedule" ]] || [[ "${INPUT_TAG}" == "staging" ]]; then
gh api repos/nearai/ironclaw-dind/dispatches \
--method POST \
-f event_type="ironclaw_image_published"
fi

- name: Summary
run: |
{
Expand Down
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ target/
# Python
__pycache__/
*.pyc
/tests/e2e/.venv/

# Benchmark results (local runs, not committed)
bench-results/
Expand Down
1 change: 1 addition & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,6 +77,7 @@ Start with these deeper docs as needed:
- If you change implementation status for any feature tracked in `FEATURE_PARITY.md`, update that file in the same branch.
- Do not open a PR that changes feature behavior without checking `FEATURE_PARITY.md` for needed status updates (`❌`, `🚧`, `✅`, notes, and priorities).
- Add the narrowest tests that validate the change: unit tests for local logic, integration tests for runtime/DB/routing behavior, and E2E or trace coverage for gateway, approvals, extensions, or other user-visible flows.
- **Test through the caller, not just the helper.** When a predicate/classifier/transform helper gates a side effect (HTTP, DB write, OAuth flow, UI mutation, tool execution) and has any wrapper or computed input between it and that side effect, a unit test on the helper alone is not sufficient regression coverage. Add a test that drives the actual call site (`*_handler`, `factory::create_*`, `manager::*`) at the integration tier or higher. Mocks of multi-arg runtime APIs must capture every argument the production caller passes. See `.claude/rules/testing.md` for the full rule and bug examples.

## Risk and Change Discipline

Expand Down
Loading
Loading