Skip to content

sync: update Matrix pilot with nearai/main 2026-07-01 - #41

Merged
github-actions[bot] merged 5 commits into
native-matrix-channel-pilotfrom
sync/native-matrix-pilot
Jul 1, 2026
Merged

github-actions[bot] merged 5 commits into
native-matrix-channel-pilotfrom
sync/native-matrix-pilot

Conversation

@personal-upstream-sync

Copy link
Copy Markdown

Automated sync from nearai/ironclaw. The upstream-main branch is an exact fast-forward mirror of nearai/main; this PR imports it into native-matrix-channel-pilot for CI with upstream source winning unrelated-history conflicts. When checks pass, the Matrix pilot branch is fast-forwarded so upstream ancestry is preserved.

italic-jinxin and others added 5 commits July 1, 2026 14:29
…input" (nearai#5338)

* fix(reborn): surface specific failure summaries for loop-exit categories (nearai#5289)

Terminal failures that reach the WebUI projection via the normal
loop-exit path carry a category from `LoopFailureKind::as_str()`
(e.g. `capability_protocol_error`). `reborn_failure_summary_for_category`
only mapped the driver-error and scheduler categories plus three loop
kinds, so the rest — `capability_protocol_error`, `model_error`,
`invalid_model_output`, `checkpoint_*`, `transcript_write_failed`,
`driver_bug`, `policy_denied`, `compaction_unavailable`, and
`driver_protocol_violation` — degraded to the generic
"The run failed before producing a reply." The LLM failure explainer,
fed only that generic fallback, then paraphrased it into the vague
"driver protocol error" the user saw, masking the real tool failure.

Map each loop-exit category to a specific, honest, user-facing summary.
This also improves the explainer's input, since the fallback it receives
now describes the actual failure stage.

Regression coverage:
- unit: `reborn_failure_summary_describes_capability_protocol_error` and
  `reborn_failure_summary_maps_loop_failure_categories_specifically`
  (no loop-exit category degrades to the generic fallback).
- caller-level: `webui_event_stream_projects_capability_protocol_error_summary`
  drives the category through the projection with no explainer wired and
  asserts the specific summary reaches the run-status item.

* fix(reborn): surface capability failure detail in per-tool UI preview (nearai#5289)

A failed capability (e.g. the `json` builtin returning `invalid_input`)
showed only the bare error-kind string in the WebUI Activity panel's
per-tool Error tab. The rich `CapabilityFailureDetail::InvalidInput`
field issues reached the model transcript but never the display-preview
path the UI renders: failures don't call `write_capability_result`, so
no preview record was staged and the projection fell back to
`failed_capability_display_preview`, which only renders the kind.

Stage a failure display-preview record (approach B), so the existing
rich projection path surfaces the detail:

- ironclaw_loop_support: `LoopCapabilityResultWriter` gains a default
  no-op `stage_capability_failure_preview`. `runtime_outcome_to_loop`
  renders a bounded, host-authored summary from the InvalidInput issues
  (`capability_failure_display_summary`) and stages it on the Failed
  arm. Only schema-derived fields (path/code/expected) are rendered;
  `received` (raw tool input) is deliberately omitted.
- ironclaw_reborn_composition: implement the writer hook on
  `LocalDevCapabilityIo` and `ProductLiveCapabilityIo`; add
  `CapabilityDisplayPreviewStore::record_failure_preview`, which mirrors
  the success path (title/input pulled from the staged input) and stores
  the rendered summary as the output, no result ref.
- frontend (history-messages.js): `toolCardFromPreview` now prefers the
  backend `output_summary`/`output_preview` over the bare error kind for
  the Error tab. Bundle rebuilt (static/dist/app.js).

Regression coverage:
- unit: `capability_failure_display_summary_renders_invalid_input_issues`
  (+ asserts `received` never leaks) and `_is_none_for_non_invalid_input`.
- caller-level: `capability_display_preview_uses_staged_failure_summary_over_bare_kind`
  drives the projection chokepoint and asserts the detailed summary
  reaches the preview view instead of "tool failed: <kind>".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): also surface failure safe_summary when no structured issues (nearai#5289)

The previous change only rendered a failure display preview for
`InvalidInput` failures carrying structured field issues. Builtin tools
like `json` report invalid_input with a descriptive message (e.g.
"invalid JSON: expected value at line 1 column 1") but no structured
issues, so `capability_failure_display_summary` returned `None`, nothing
was staged, and the per-tool preview fell back to the bare error kind.

Extend the helper: when there are no structured issues, surface the
failure's host-authored `safe_summary` (already sanitized) unless it is a
generic placeholder ("capability invocation failed" /
"capability authorization denied") that adds nothing over the kind.

Tests: replaced the non-invalid-input case with
`capability_failure_display_summary_uses_safe_summary_without_issues`
(asserts the json-style message is surfaced) and
`_is_none_for_generic_placeholder`.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address PR review on failure-preview staging (nearai#5289)

Three review findings on the per-tool failure-preview change:

- Reuse `ironclaw_host_api::truncate_capability_display_text` for UTF-8
  boundary truncation in `capability_failure_display_summary` instead of
  a hand-rolled helper (gemini-code-assist).
- Fix a TOCTOU between `record_failure_preview` and `prune_run`:
  acquire the pending and completed locks together and hold them across
  the remove-from-pending + insert-into-completed pair so a concurrent
  prune cannot interleave and leak an unprunable completed record. Lock
  order (pending before completed) matches every other site, so holding
  both cannot deadlock (gemini-code-assist).
- Persist the failure preview to the durable timeline, not just the
  in-memory store: `stage_capability_failure_preview` is now async, and
  `LocalDevCapabilityIo` appends a durable display-preview message with
  status `Failed`, mirroring the success path so the detail survives
  refresh/replay. `try_append_durable_display_preview` takes a status
  parameter. `ProductLiveCapabilityIo` stays in-memory, matching its own
  success path which does not persist previews durably (coderabbitai).

Test: `capability_io_writes_failure_display_preview_to_durable_history`
asserts the durable timeline carries a Failed-status preview with the
rendered summary.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): carry tool failure detail to the live per-tool UI card (nearai#5289)

The live, in-progress per-tool activity card showed only the bare error
kind (e.g. "invalid_input") on failure. The display-preview store fix
covered the history/timeline path, but the live card is driven by a
different projection path: the `CapabilityFailed` loop milestone ->
`ThreadLiveProjectionItem::CapabilityActivity` -> `CapabilityActivityView`,
which carried `error_kind` only — the sanitized failure message never
reached it.

Plumb the host-authored sanitized failure summary additively through that
live path so the card shows the real reason (e.g. "invalid JSON: ..."):

- ironclaw_turns: add `safe_summary: Option<String>` to
  `LoopHostMilestoneKind::CapabilityFailed` and
  `LoopProgressEvent::CapabilityActivityFailed` (additive, serde-default).
- ironclaw_loop_support: `runtime_terminal_milestone` populates it from
  the `RuntimeCapabilityFailure` message on the model-visible failure arm;
  host/infra and gate-denied paths emit `None`.
- ironclaw_event_streams: add `error_detail` to the live
  `CapabilityActivity` projection item.
- ironclaw_product_adapters: add `error_detail` to `CapabilityActivityView`
  (+ input, wire ser/de) with the same bounded/sanitized boundary
  validation as the other display fields.
- ironclaw_reborn_composition: live progress sanitizes the milestone
  summary and maps it onto the view; the history/runtime-payload path
  keeps carrying detail via the separate CapabilityDisplayPreview.
- frontend (history-messages.js): `toolCardFromActivity` prefers
  `error_detail` over the bare kind. Bundle rebuilt.

The durable runtime event log still records only the failure kind (no
backend-detail persistence), per ironclaw_turns guardrails.

Regression: `webui_event_stream_projects_live_tool_failure` now drives a
`CapabilityFailed` milestone with a `safe_summary` and asserts the live
`CapabilityActivityView.error_detail` carries it end-to-end.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): render dispatch failure kinds as plain language, not category tokens (nearai#5289)

When a capability dispatch fails without a host-authored safe_summary,
the failure message fell back to the error's Display, which exposed the
stable redacted category token (e.g. "dispatch failed: InputEncode") in
the per-tool UI Error tab.

Add `human_summary()` to `RuntimeDispatchErrorKind` and
`DispatchFailureKind` (fixed host-authored sentences, no raw content) and
use it as the fallback in `sanitized_failure_message`'s dispatch arm. The
stable `as_str()` token stays the contract for routing/metrics/audit; only
the user-facing message changes. So "InputEncode" now reads "the tool
input could not be encoded".

Tests: `dispatch_failure_kind_human_summary_is_plain_language_not_category_token`;
updated the two production.rs tests that pinned the old token wording
(they now assert the human summary and still verify no raw backend string
or secret leaks).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): close prune_run race in record_result_with_preview (nearai#5289)

Code review flagged a race between the display-preview store's
remove-from-pending + insert-into-completed pair and `prune_run`: the two
locks were acquired independently, so a concurrent prune could interleave
and leak a completed preview record that is never pruned.

The failure path (`record_failure_preview`) was already fixed to hold both
locks; apply the same fix to the pre-existing success path
(`record_result_with_preview`) so the whole pattern is consistent. Both
sites and `prune_run` acquire pending-before-completed, so holding both
cannot deadlock.

(The other two review findings — non-durable failure staging, and a
hand-rolled char-boundary truncation helper — were already resolved in
earlier commits on this branch.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: update webui v2 capability activity fixtures

* test: expect human dispatch failure summary

* test: stabilize reborn composer active-run smoke

* fix(webui): stop bare error-kind activity frame clobbering failure detail (nearai#5289)

A failed capability surfaces its detail on two frames that race in the
live stream: the runtime-payload `capability_activity` deliberately
carries no `error_detail` (its `toolError` is the bare kind, e.g.
"invalid_input"), while the `capability_display_preview` carries the
real reason. The tool-activity merge used `incoming.toolError ||
current.toolError`, so a late bare-kind activity frame overwrote the
detailed reason already set by the preview — the live card showed
"invalid_input" while a page refresh (replay) showed the real message.

`mergedToolError` now refuses to downgrade: if the incoming error is
only the bare kind and the current one is a different (detailed)
message, the detail is kept. First-frame and genuine upgrades are
unaffected.

Test: tool-activity-state.test.mjs drives the real preview-then-activity
race and asserts the detail survives.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: stream capability failure detail in webui events

* test: update event stream capability activity fixture

* fix: redact filename-bearing failure summaries

* fix: preserve redacted capability failure summaries

* fix: keep invalid tool failure summaries nonfatal

* fix: harden capability failure summary redaction

* fix: require filesystem context for filename summaries

* fix: tighten workspace failure summary matching

* fix: preserve invalid input failure summary

* fix: keep input encode summaries replay-safe

* fix: preserve safe invalid input summaries

* fix: harden capability failure summaries

* fix: redact sensitive capability issue fields

* fix: document projected capability error details

* fix: catch sensitive issue marker variants

* docs: clarify projected error detail boundary

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…nearai#5430)

* ci(reborn): add cargo-llvm-cov integration-tier coverage job (T0-COV)

Adds a per-PR coverage signal for the Reborn backend. A new `Reborn
Coverage` workflow runs cargo-llvm-cov over the in-process
integration-tier test binaries (tests/reborn_integration_*.rs +
tests/reborn_group_*/) and surfaces a Reborn-crate-scoped line-coverage
percentage plus a per-crate hole list in the PR job summary.

This makes the roadmap's "100% int-tier coverage" goal measurable
(docs/reborn/reborn-backend-coverage-roadmap.md, T0-COV) and yields the
real hole list. It is informational only and does not gate PRs.

- .github/workflows/reborn-coverage.yml: the job. Combined
  `cargo llvm-cov --workspace --test <int-tier>` invocation so the report
  is scoped to every linked member crate (the standalone `report`
  subcommand cannot take --workspace and would scope to the root package
  only). Default features, matching run-reborn-root-partition.sh; no live
  Postgres needed. Reuses coverage.yml's pinned action SHAs / install
  action. Surfaces via $GITHUB_STEP_SUMMARY per repo norm.
- scripts/ci/reborn-coverage-int-tier-tests.sh: dynamically discovers the
  int-tier `--test` targets (auto-expands as new suites land).
- scripts/ci/reborn-coverage-summary.sh: filters the llvm-cov JSON export
  to the Reborn crate families and renders the % + per-crate table.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(reborn): sticky PR comment for int-tier coverage (T0-COV)

Surface the Reborn coverage %/hole-list as an upserted sticky PR comment so
it lives in the conversation, not just the Actions job summary. Visibility
only — never gates: the step is `continue-on-error: true` and PR-only, so a
fork PR's read-only GITHUB_TOKEN (comment API 403s) can't red the check, and
nothing in the job exits non-zero on a coverage number.

- reborn-coverage.yml: grant `pull-requests: write`; add a "Post sticky
  coverage comment" step after the existing $GITHUB_STEP_SUMMARY render
  (kept as-is), before the artifact upload.
- reborn-coverage-comment.sh (new): upsert one marker'd comment via pure
  `gh api` (no new action). Reuses reborn-coverage-summary.sh for both the
  body and the breadth holes (--zero-crates) — no duplicated aggregation jq.
  Prepends a 0-coverage breadth callout (informational; "target: 0" is the
  roadmap goal, not a check). Lookup uses `--paginate` so a busy PR's aged
  sticky isn't missed (which would post duplicates); marker passed via env
  to jq (no string interpolation); ids captured before `head` to avoid a
  pipefail SIGPIPE.
- reborn-coverage-summary.sh: add `--zero-crates` mode (single owner of the
  crate filter + aggregation); fold the per-crate table into <details> so
  the comment/summary stays compact (headline % + callout always visible);
  reword the informational note to cover the callout too.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(reborn): gate coverage rust-cache saves to main (T0-COV)

The reborn-coverage job had no `save-if`, so it wrote a multi-GB instrumented
rust-cache on every PR push. GitHub scopes caches per-branch, so those PR-branch
saves are invisible to other PRs and only churn the ~10 GB repo cache LRU —
evicting the one main-seeded cache that PRs actually restore from.

Gate saves to `push` on `main` (the workflow already triggers there), matching
the convention in test.yml/code_style.yml/replay-gate.yml. PRs now restore the
warm main-seeded instrumented cache and skip the wasteful save. Instrumented
caches stay separate from the non-instrumented reborn-tests caches by design
(coverage RUSTFLAGS change every crate's fingerprint — no cross-reuse possible).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(reborn): address PR review — multi-dataset jq, token guard, job timeout (T0-COV)

Fixes from PR nearai#5430 review (valid findings only):

- summary.sh: iterate all llvm-cov `data[]` datasets (`.data[]?.files[]?`)
  instead of `.data[0]` — the export format permits multiple datasets and
  index-0-only would undercount. The trailing `?`s keep the empty/missing-data
  path crash-free (unchanged behavior; `{"data":[]}` / `{}` still fall through
  to the no-data message).
- comment.sh: fast-fail if GH_TOKEN is unset, matching the existing
  GITHUB_REPOSITORY/PR_NUMBER guards (clearer failure for local runs; gh needs
  it for the API calls).
- reborn-coverage.yml: add `timeout-minutes: 180` as a runaway-hang backstop
  (well above any real build; halves GitHub's 360-minute default). Corrects the
  prior comment that claimed coverage.yml runs its instrumented job without a
  timeout — its matrix job caps at 60.

Rejected as non-issues: jq null-index "crash" (jq indexes null as null, not an
error — verified); group_by requiring pre-sorted input (jq's group_by sorts
internally); needing `issues: write` (the sticky comment already posts on a
same-repo PR with pull-requests: write). `clean --workspace` preserving the
cache is intentional (it scopes to workspace members, keeping instrumented
deps).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): regression tests for coverage CI helpers (T0-COV)

Add scripts/ci/test-reborn-coverage.sh (mirrors the test-classify-test-scope.sh
precedent) covering the three new coverage helpers — closes the "no committed
edge-case test" review findings on PR nearai#5430:

- summary.sh report mode: mixed Reborn/non-Reborn filtering, no-match no-data
  message, empty/absent `data` no-crash, multi-dataset `.data[]` aggregation,
  zero-covered crate sorted to top.
- summary.sh --zero-crates: exact zero-covered name list, all-covered/empty →
  empty output.
- comment.sh upsert (fake `gh` on PATH): no-sticky → POST with marker body,
  existing sticky at a non-first list position → PATCH (not a duplicate POST),
  zero-covered callout prepended.
- int-tier discovery: empty tree → exit 1, single file/dir suites, mixed
  suites sorted+deduped.

43 cases, self-contained (mktemp + trap cleanup), shellcheck-clean. Not wired
into a workflow — matches the unwired test-classify-test-scope.sh precedent;
run manually as a local regression signal.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): cover comment.sh env-var guards (T0-COV)

Add C4-C6 to test-reborn-coverage.sh: run reborn-coverage-comment.sh with each
of GH_TOKEN / PR_NUMBER / GITHUB_REPOSITORY unset and assert a non-zero exit
plus the matching "must be set" message on stderr. Pins the fast-fail guards
(the GH_TOKEN one added in 7b617b7) that C1-C3 didn't exercise. Uses the
existing `capture env -u …` pattern (49 cases, shellcheck-clean).

Addresses the CodeRabbit review finding on PR nearai#5430.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): make assert_line_before non-fatal on missing needle (T0-COV)

`assert_line_before`'s `grep -n | head | cut` pipelines run under
`set -euo pipefail`; a missing needle made grep exit 1, propagated by pipefail
to the assignment, aborting the whole suite before the function's own empty
checks could report a normal FAIL — contradicting the "run every case" harness
design. Append `|| true` so a missing needle yields an empty line number and
falls through to the FAIL path.

Addresses the Copilot review finding on PR nearai#5430.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(reborn): PR-number concurrency key + drop dead roadmap citation (T0-COV)

Two PR review fixes:

- reborn-coverage.yml: key the pull_request concurrency group on
  `github.event.pull_request.number` instead of `github.head_ref`. head_ref is
  just the source branch name, so two fork PRs sharing a common branch name
  (main/fix/…) collided on the same group and, with cancel-in-progress, one
  push could cancel the other PR's coverage run. number is null off-PR, so it
  falls back to github.ref for push/dispatch (pattern from pr-label-classify.yml).
- Drop the citation of docs/reborn/reborn-backend-coverage-roadmap.md from the
  workflow header and the int-tier discovery script — that file exists on
  neither this branch nor main, so it was a dangling source-of-truth pointer.
  The int-tier scope is already stated inline (reborn_integration_*.rs +
  reborn_group_*/), so nothing is lost.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): allowlist-boundary + missing-file coverage cases (T0-COV)

Add three regression cases to test-reborn-coverage.sh (PR nearai#5430 review):

- A6: crate-filter boundary — fixture with exact single-crate matches
  (ironclaw_architecture, ironclaw_slack_v2_adapter), a family-prefix crate
  (ironclaw_reborn_config), and a lookalike (ironclaw_architecture_extra).
  Asserts the exact + prefix crates appear, the lookalike is excluded, and its
  999 lines are dropped from the aggregate (16/30, not 16/1029) — pins the
  prefix-vs-exact regex semantics against a future edit that widens the match.
- A7: summary.sh on a nonexistent JSON path -> non-zero exit + "coverage JSON
  not found".
- C7: comment.sh on a nonexistent JSON path -> non-zero exit + "coverage JSON
  not found" + NO gh mutation (guard fires before any gh call).

60 cases, shellcheck-clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): malformed-JSON aborts comment before any mutation (T0-COV)

Add case C8 to test-reborn-coverage.sh: an existing-but-malformed coverage
JSON. Distinct from C7 (missing file) — it passes comment.sh's `[ -f ]` guard
and must instead fail at the reborn-coverage-summary.sh render (jq parse error
under `set -e`), before any POST/PATCH. Asserts non-zero exit and no fake-gh
mutation, pinning the render-before-mutate ordering. Exit asserted non-zero
(not an exact code) since jq's parse-error exit varies across versions.

62 cases, shellcheck-clean. Addresses the PR nearai#5430 review finding.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(webui-v2): avoid duplicate logs header during chat runs

* fix(webui-v2): hide chat run logs shortcut

* test(webui-v2): remove redundant chat logs assertion

* fix(webui-v2): move chat logs link into message actions

* fix(webui-v2): float chat logs shortcut

* fix(webui-v2): use terminal icon for chat logs shortcut

* fix(webui-v2): reserve chat logs space with spacer

* style(webui-v2): soften floating logs shortcut

* style(webui-v2): increase chat logs shortcut contrast

* fix(webui-v2): reuse scoped logs path builder

* test(e2e): handle busy reborn sends before settling

* fix(webui-v2): keep chat logs route construction in chat

* test(e2e): retry transient submit errors while settling
@github-actions github-actions Bot added scope: ci scope: docs size: XL Changed-line size classification risk: medium Risk classification contributor: regular Contributor history classification labels Jul 1, 2026
@github-actions
github-actions Bot merged commit eb861d4 into native-matrix-channel-pilot Jul 1, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: regular Contributor history classification risk: medium Risk classification scope: ci scope: docs size: XL Changed-line size classification

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants