Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/live-canary.yml
Original file line number Diff line number Diff line change
Expand Up @@ -538,6 +538,8 @@ jobs:
REBORN_WEBUI_V2_LIVE_QA_LLM_API_KEY_ENV: NEARAI_API_KEY
REBORN_WEBUI_V2_LIVE_QA_LLM_PROVIDER_ID: nearai
REBORN_WEBUI_V2_LIVE_QA_LLM_MODEL: ${{ vars.REBORN_WEBUI_V2_LIVE_QA_LLM_MODEL || 'deepseek-ai/DeepSeek-V4-Flash' }}
REBORN_WEBUI_V2_LIVE_QA_LLM_JUDGE: ${{ vars.REBORN_WEBUI_V2_LIVE_QA_LLM_JUDGE || '1' }}
REBORN_WEBUI_V2_LIVE_QA_LLM_JUDGE_MODEL: ${{ vars.REBORN_WEBUI_V2_LIVE_QA_LLM_JUDGE_MODEL || vars.REBORN_WEBUI_V2_LIVE_QA_LLM_MODEL || 'deepseek-ai/DeepSeek-V4-Flash' }}
IRONCLAW_REBORN_GOOGLE_CLIENT_ID: ${{ secrets.IRONCLAW_REBORN_GOOGLE_CLIENT_ID || secrets.GOOGLE_CLIENT_ID || secrets.GOOGLE_OAUTH_CLIENT_ID }}
IRONCLAW_REBORN_GOOGLE_OAUTH_REDIRECT_URI: ${{ vars.IRONCLAW_REBORN_GOOGLE_OAUTH_REDIRECT_URI || vars.GOOGLE_OAUTH_REDIRECT_URI }}
IRONCLAW_REBORN_GOOGLE_HOSTED_DOMAIN_HINT: ${{ vars.IRONCLAW_REBORN_GOOGLE_HOSTED_DOMAIN_HINT || vars.GOOGLE_ALLOWED_HD }}
Expand Down
76 changes: 76 additions & 0 deletions .github/workflows/openwiki-update.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
name: OpenWiki Update

# Regenerates the narrative code wiki (openwiki/) and opens a PR for HUMAN review.
#
# Decisions (see docs/plans/2026-07-02 discovery rollout):
# - Provider: Anthropic via ANTHROPIC_API_KEY (already a repo secret).
# - Publish: open a pull request; a human reviews and merges. NO auto-merge —
# SOC 2 change management (CC8.1 / separation of duties) requires every change
# to main to be human-reviewed and approved. The bot only authors the proposal.
#
# Notes:
# - The PR is opened via the GH_RELEASES_MANAGER GitHub App (not GITHUB_TOKEN) so
# the required-check workflows actually trigger on it — a GITHUB_TOKEN-opened PR
# does not trigger other workflows, so a human reviewer would see no checks and
# the ruleset would block the merge.
# - OPENWIKI_MODEL_ID must match OpenWiki's Anthropic model resolution; if the
# run errors on the model, adjust the string (e.g. an "anthropic/"/"anthropic:"
# prefix, or claude-haiku-4-5-20251001 for a cheaper pass).

on:
workflow_dispatch:
schedule:
- cron: "0 8 * * 1" # Mondays 08:00 UTC (weekly)

permissions:
contents: write
pull-requests: write

jobs:
update:
runs-on: ubuntu-latest
steps:
- name: Mint GitHub App token
id: app
uses: actions/create-github-app-token@v1
with:
app-id: ${{ secrets.GH_RELEASES_MANAGER_APP_ID }}
private-key: ${{ secrets.GH_RELEASES_MANAGER_APP_PRIVATE_KEY }}

- name: Check out repository
uses: actions/checkout@v4
with:
token: ${{ steps.app.outputs.token }}
persist-credentials: true

- name: Set up Node.js
uses: actions/setup-node@v4
with:
node-version: "22"

- name: Install OpenWiki
run: npm install --global openwiki

- name: Regenerate wiki
run: openwiki --update --print
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
OPENWIKI_MODEL_ID: claude-sonnet-4-6

- name: Open review PR
id: cpr
uses: peter-evans/create-pull-request@v7
with:
token: ${{ steps.app.outputs.token }}
add-paths: openwiki
branch: openwiki/update
delete-branch: true
commit-message: "docs: update OpenWiki wiki"
title: "docs: update OpenWiki wiki"
labels: documentation
body: |
Automated OpenWiki narrative-docs refresh (`openwiki/`).

Review and merge as usual — this PR is NOT auto-merged; a human
approves it per change-management policy. Precise structural discovery
lives in the codebase-memory graph; this is the prose "what/why" layer.
5 changes: 5 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,11 @@ tests/fixtures/llm_traces/live/*.log
.anvil/
.codegraph/

# codebase-memory knowledge graph — build artifact, not source.
# Rebuilt from code via the codebase-memory MCP; per-environment, never committed.
# See CLAUDE.md -> "Code Discovery".
.codebase-memory/

# Downstream sync exception: nearai/main currently tracks these
# source/control assets even though broad local-ignore rules also
# match them. Keep the explicit exceptions in generated sync PRs so
Expand Down
8 changes: 8 additions & 0 deletions .mcp.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"mcpServers": {
"codebase-memory-mcp": {
"command": "codebase-memory-mcp",
"args": []
}
}
}
15 changes: 15 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,18 @@ Start with these deeper docs as needed:
- `src/NETWORK_SECURITY.md`
- `tests/e2e/CLAUDE.md`

## Code Discovery — Query the Knowledge Graph First

This repo is indexed into a **codebase knowledge graph** via the `codebase-memory` MCP server (~464k nodes / ~2.5M edges over `src/` + `crates/`). For any *where-is / who-calls / how-does-data-flow / what-does-this-touch* question, **query the graph before `grep`** — text search can't see cross-crate call chains, and a feature here crosses many crates (`product_workflow → composition → webui_v2 → runtime → frontend`).

- **Location:** `.codebase-memory/graph.db.zst` — a git-ignored build artifact (rebuilt from code, one per environment, never committed).
- **Freshness:** run `bash scripts/codebase-graph.sh status` (compares indexed commit vs `HEAD`). If missing → `index_repository(repo_path=".")`; if stale → `detect_changes(since="<indexed-commit>")` or re-index. The graph is point-in-time — verify its claims against live code before acting.
- **Recipes:** define → `search_graph(name_pattern=…)` + `get_code_snippet(…)`; callers → `trace_path(mode="calls")`; value flow → `trace_path(mode="data_flow")`; cross-crate flow → `trace_path(mode="cross_service")`; area shape → `get_architecture(…)`; Cypher → `query_graph(…)`.

Use `grep`/read for text, config, and non-code files, and after the graph for code structure — not before.

**Narrative docs (auto-generated, read-only):** prose subsystem docs live in `openwiki/`, regenerated by `.github/workflows/openwiki-update.yml`. `Read` them for *what/how* orientation; use the graph for precise *where/who*. Don't hand-edit `openwiki/`.

## Architecture Mental Model

- Channels normalize external input into `IncomingMessage`; `ChannelManager` merges all active channel streams.
Expand All @@ -25,6 +37,9 @@ Start with these deeper docs as needed:

## Where to Work

**Build new features Reborn-side, in `crates/` — not the v1 `src/` monolith.** A Reborn feature crosses `product_workflow → composition → webui_v2 → runtime/serve → frontend`; the binary entry point is `crates/ironclaw_reborn_cli` (`reborn_cli`), **not** `src/main.rs`. Start from the `reborn-feature` skill, which maps the layers. `src/` is v1, being retired under "Clean up old architecture" — touch it only to maintain existing v1 behavior, never to add new features.

Existing subsystem locations (mostly v1 `src/`; maintain, don't extend):
- Agent/runtime behavior: `src/agent/`
- Web gateway/API/SSE/WebSocket: `src/channels/web/`
- Persistence and DB abstractions: `src/db/`
Expand Down
29 changes: 29 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,35 @@

**IronClaw** is a secure personal AI assistant — user-first security, self-expanding tools, defense in depth, multi-channel access with proactive background execution.

## Code Discovery — Query the Knowledge Graph First

This repo is indexed into a **codebase knowledge graph** (the `codebase-memory` MCP server): ~464k nodes / ~2.5M edges spanning all of `src/` and `crates/`. For any *where-is / who-calls / how-does-data-flow / what-does-this-touch* question, **query the graph before reaching for `Grep`** — text search cannot see cross-crate call chains, and this codebase's real cost is cross-crate (a feature crosses `product_workflow → composition → webui_v2 → runtime → frontend`).

**Where it lives:** `.codebase-memory/graph.db.zst` — a **git-ignored build artifact, not source**. One per environment, rebuilt from code. Never commit it.

**Freshness (check at the start of a discovery task):** run `bash scripts/codebase-graph.sh status` — it compares the graph's indexed commit against `HEAD`. Then:
- **Missing** → `index_repository(repo_path=".")` once to build it.
- **Stale** → `detect_changes(since="<indexed-commit>")` for the changed symbols + blast radius, or re-run `index_repository` to fully refresh.
- The graph is a point-in-time index — verify anything it asserts against live code before acting.

**Discovery recipes (use these instead of `Grep` for code structure):**
- Where a symbol is defined → `search_graph(name_pattern=…)`, then `get_code_snippet(qualified_name=…)`
- Who calls X / what X calls → `trace_path(function_name=…, mode="calls")`
- How a value flows across layers → `trace_path(mode="data_flow")`
- Cross-crate / cross-service path (the reborn 5-layer feature flow) → `trace_path(mode="cross_service")`
- Structure of an area → `get_architecture(…)`; graph-augmented text search → `search_code(pattern=…)`
- Arbitrary structural queries → `query_graph(<Cypher>)`

`Grep`/`Glob`/`Read` remain correct for text, config, and non-code files — and for reading a file the graph pointed you to. For *code structure*, the graph comes first.

**Narrative orientation (what/why, not where):** prose docs for each subsystem live in `openwiki/` — an auto-generated wiki kept fresh by `.github/workflows/openwiki-update.yml`. For *"what does this subsystem do / how does this flow work"* questions, `Read` the relevant `openwiki/` page; use the graph for precise structure. Do not hand-edit `openwiki/` — it is regenerated. The two layers are complementary: `openwiki/` = prose map, the graph = exact index.

## Where to Build — Reborn-First

**New feature work targets the Reborn stack in `crates/`, not the v1 `src/` monolith.** A Reborn feature crosses `product_workflow → composition → webui_v2 → runtime/serve → frontend`; the binary entry point is `crates/ironclaw_reborn_cli` (`reborn_cli`), **not** `src/main.rs`. Start from the `reborn-feature` skill — it maps those layers so you wire a feature in one pass instead of layer-by-layer.

`src/` is the **v1 monolith**, being retired under the roadmap's "Clean up old architecture." Maintain existing v1 behavior there when a bug requires it, but **do not build new features into `src/`** — they belong Reborn-side. The detailed `src/` layout in "Project Structure" below documents v1 for maintenance, not as the default place to add code.

## Build & Test

```bash
Expand Down
59 changes: 59 additions & 0 deletions scripts/codebase-graph.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
#!/usr/bin/env bash
#
# codebase-graph.sh — freshness status for the codebase-memory knowledge graph.
#
# The graph (.codebase-memory/graph.db.zst) is a git-ignored build artifact indexed
# by the codebase-memory MCP server. This script only INSPECTS freshness; the actual
# (re)indexing is done through the MCP tools invoked by an agent:
# - build: index_repository(repo_path=".")
# - delta/impact: detect_changes(since="<indexed-commit>")
#
# Usage: bash scripts/codebase-graph.sh [status]
#
set -euo pipefail

repo_root="$(git rev-parse --show-toplevel)"
artifact="$repo_root/.codebase-memory/artifact.json"

read_field() { python3 -c "import json,sys;print(json.load(open('$artifact')).get('$1','?'))"; }

status() {
if [ ! -f "$artifact" ]; then
echo "graph: MISSING (no .codebase-memory/artifact.json)"
echo "action: build it once — call index_repository(repo_path=\".\") via the codebase-memory MCP"
return 2
fi

local indexed head nodes edges when
indexed="$(read_field commit)"
nodes="$(read_field nodes)"
edges="$(read_field edges)"
when="$(read_field indexed_at)"
head="$(git rev-parse HEAD)"

echo "graph: indexed @ ${indexed:0:9} (${nodes} nodes / ${edges} edges, ${when})"
echo "HEAD: ${head:0:9}"

if [ "$indexed" = "$head" ]; then
echo "status: FRESH (matches HEAD)"
return 0
fi

if git merge-base --is-ancestor "$indexed" HEAD 2>/dev/null; then
local n
n="$(git rev-list --count "$indexed"..HEAD 2>/dev/null || echo '?')"
echo "status: STALE — ${n} commit(s) behind HEAD"
echo "action: delta — detect_changes(since=\"$indexed\") (changed symbols + blast radius)"
echo " or full refresh — index_repository(repo_path=\".\")"
return 1
fi

echo "status: DIVERGED — indexed commit is not in current history (rebase/force-push?)"
echo "action: full re-index — index_repository(repo_path=\".\")"
return 1
}

case "${1:-status}" in
status) status ;;
*) echo "usage: $0 [status]" >&2; exit 64 ;;
esac
15 changes: 15 additions & 0 deletions scripts/dev-setup.sh
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,21 @@ else
echo " Skipped: not a git repository"
fi

echo ""
# Codebase knowledge graph (codebase-memory MCP) — powers agent code discovery.
# Single static binary, no deps, no API keys, 100% local. See CLAUDE.md -> "Code Discovery".
if command -v codebase-memory-mcp &>/dev/null; then
echo "[graph] codebase-memory-mcp found: $(command -v codebase-memory-mcp)"
else
echo "[graph] Installing codebase-memory-mcp (agent code-discovery graph)..."
if curl -fsSL https://raw.githubusercontent.com/DeusData/codebase-memory-mcp/main/install.sh | bash; then
echo " installed. The repo's .mcp.json wires it into Claude Code automatically."
else
echo " WARN: install failed — agents will fall back to grep."
echo " Install manually: https://github.com/DeusData/codebase-memory-mcp"
fi
fi

echo ""
echo "=== Setup complete ==="
echo ""
Expand Down
22 changes: 21 additions & 1 deletion scripts/reborn_webui_v2_live_qa/run_live_qa.py
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,11 @@
_root_filesystem_json,
_root_filesystem_secret_by_handle,
)
from scripts.reborn_webui_v2_live_qa.semantic_judge import ( # noqa: E402
_compact_json,
_judge_assistant_reply_completion,
_semantic_judge_passed,
)
from scripts.reborn_webui_v2_live_qa.slack_helpers import ( # noqa: E402
_append_slack_channel_route,
_append_slack_channel_route_if_configured,
Expand Down Expand Up @@ -856,6 +861,7 @@ async def action(page: object) -> None:
marker=marker,
required_text=required_text,
timeout=timeout,
semantic_goal=prompt,
)
if forbidden_text:
text = str(observed["text_excerpt"]).lower()
Expand Down Expand Up @@ -967,6 +973,7 @@ async def action(page: object) -> None:
marker=marker,
required_text=required_text,
timeout=timeout,
semantic_goal=prompt,
)

try:
Expand Down Expand Up @@ -996,6 +1003,7 @@ async def _wait_for_assistant_reply(
marker: str | None,
required_text: list[str],
timeout: float,
semantic_goal: str | None = None,
) -> str:
deadline = time.monotonic() + timeout
assistant = page.locator("[data-testid='msg-assistant']").last # type: ignore[attr-defined]
Expand Down Expand Up @@ -1030,10 +1038,22 @@ async def _wait_for_assistant_reply(
main_text = await page.locator("main").inner_text(timeout=1000) # type: ignore[attr-defined]
except Exception:
pass
semantic_judge: dict[str, object] | None = None
if last_text and (not marker or marker in last_text):
semantic_judge = await _judge_assistant_reply_completion(
marker=marker,
required_text=required_text,
assistant_text=last_text,
main_text=main_text,
semantic_goal=semantic_goal,
)
if _semantic_judge_passed(semantic_judge):
return last_text[-2000:]
raise AssertionError(
"assistant reply did not contain required text before timeout. "
f"marker={marker!r} required_text={required_text!r} "
f"last_assistant={last_text[-500:]!r} main_excerpt={main_text[-1000:]!r}"
f"last_assistant={last_text[-500:]!r} main_excerpt={main_text[-1000:]!r} "
f"semantic_judge={_compact_json(semantic_judge)}"
)


Expand Down
Loading
Loading