sync: update Matrix pilot with nearai/main 2026-07-01 - #41
Merged
github-actions[bot] merged 5 commits intoJul 1, 2026
Merged
Conversation
…input" (nearai#5338) * fix(reborn): surface specific failure summaries for loop-exit categories (nearai#5289) Terminal failures that reach the WebUI projection via the normal loop-exit path carry a category from `LoopFailureKind::as_str()` (e.g. `capability_protocol_error`). `reborn_failure_summary_for_category` only mapped the driver-error and scheduler categories plus three loop kinds, so the rest — `capability_protocol_error`, `model_error`, `invalid_model_output`, `checkpoint_*`, `transcript_write_failed`, `driver_bug`, `policy_denied`, `compaction_unavailable`, and `driver_protocol_violation` — degraded to the generic "The run failed before producing a reply." The LLM failure explainer, fed only that generic fallback, then paraphrased it into the vague "driver protocol error" the user saw, masking the real tool failure. Map each loop-exit category to a specific, honest, user-facing summary. This also improves the explainer's input, since the fallback it receives now describes the actual failure stage. Regression coverage: - unit: `reborn_failure_summary_describes_capability_protocol_error` and `reborn_failure_summary_maps_loop_failure_categories_specifically` (no loop-exit category degrades to the generic fallback). - caller-level: `webui_event_stream_projects_capability_protocol_error_summary` drives the category through the projection with no explainer wired and asserts the specific summary reaches the run-status item. * fix(reborn): surface capability failure detail in per-tool UI preview (nearai#5289) A failed capability (e.g. the `json` builtin returning `invalid_input`) showed only the bare error-kind string in the WebUI Activity panel's per-tool Error tab. The rich `CapabilityFailureDetail::InvalidInput` field issues reached the model transcript but never the display-preview path the UI renders: failures don't call `write_capability_result`, so no preview record was staged and the projection fell back to `failed_capability_display_preview`, which only renders the kind. Stage a failure display-preview record (approach B), so the existing rich projection path surfaces the detail: - ironclaw_loop_support: `LoopCapabilityResultWriter` gains a default no-op `stage_capability_failure_preview`. `runtime_outcome_to_loop` renders a bounded, host-authored summary from the InvalidInput issues (`capability_failure_display_summary`) and stages it on the Failed arm. Only schema-derived fields (path/code/expected) are rendered; `received` (raw tool input) is deliberately omitted. - ironclaw_reborn_composition: implement the writer hook on `LocalDevCapabilityIo` and `ProductLiveCapabilityIo`; add `CapabilityDisplayPreviewStore::record_failure_preview`, which mirrors the success path (title/input pulled from the staged input) and stores the rendered summary as the output, no result ref. - frontend (history-messages.js): `toolCardFromPreview` now prefers the backend `output_summary`/`output_preview` over the bare error kind for the Error tab. Bundle rebuilt (static/dist/app.js). Regression coverage: - unit: `capability_failure_display_summary_renders_invalid_input_issues` (+ asserts `received` never leaks) and `_is_none_for_non_invalid_input`. - caller-level: `capability_display_preview_uses_staged_failure_summary_over_bare_kind` drives the projection chokepoint and asserts the detailed summary reaches the preview view instead of "tool failed: <kind>". Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn): also surface failure safe_summary when no structured issues (nearai#5289) The previous change only rendered a failure display preview for `InvalidInput` failures carrying structured field issues. Builtin tools like `json` report invalid_input with a descriptive message (e.g. "invalid JSON: expected value at line 1 column 1") but no structured issues, so `capability_failure_display_summary` returned `None`, nothing was staged, and the per-tool preview fell back to the bare error kind. Extend the helper: when there are no structured issues, surface the failure's host-authored `safe_summary` (already sanitized) unless it is a generic placeholder ("capability invocation failed" / "capability authorization denied") that adds nothing over the kind. Tests: replaced the non-invalid-input case with `capability_failure_display_summary_uses_safe_summary_without_issues` (asserts the json-style message is surfaced) and `_is_none_for_generic_placeholder`. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn): address PR review on failure-preview staging (nearai#5289) Three review findings on the per-tool failure-preview change: - Reuse `ironclaw_host_api::truncate_capability_display_text` for UTF-8 boundary truncation in `capability_failure_display_summary` instead of a hand-rolled helper (gemini-code-assist). - Fix a TOCTOU between `record_failure_preview` and `prune_run`: acquire the pending and completed locks together and hold them across the remove-from-pending + insert-into-completed pair so a concurrent prune cannot interleave and leak an unprunable completed record. Lock order (pending before completed) matches every other site, so holding both cannot deadlock (gemini-code-assist). - Persist the failure preview to the durable timeline, not just the in-memory store: `stage_capability_failure_preview` is now async, and `LocalDevCapabilityIo` appends a durable display-preview message with status `Failed`, mirroring the success path so the detail survives refresh/replay. `try_append_durable_display_preview` takes a status parameter. `ProductLiveCapabilityIo` stays in-memory, matching its own success path which does not persist previews durably (coderabbitai). Test: `capability_io_writes_failure_display_preview_to_durable_history` asserts the durable timeline carries a Failed-status preview with the rendered summary. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn): carry tool failure detail to the live per-tool UI card (nearai#5289) The live, in-progress per-tool activity card showed only the bare error kind (e.g. "invalid_input") on failure. The display-preview store fix covered the history/timeline path, but the live card is driven by a different projection path: the `CapabilityFailed` loop milestone -> `ThreadLiveProjectionItem::CapabilityActivity` -> `CapabilityActivityView`, which carried `error_kind` only — the sanitized failure message never reached it. Plumb the host-authored sanitized failure summary additively through that live path so the card shows the real reason (e.g. "invalid JSON: ..."): - ironclaw_turns: add `safe_summary: Option<String>` to `LoopHostMilestoneKind::CapabilityFailed` and `LoopProgressEvent::CapabilityActivityFailed` (additive, serde-default). - ironclaw_loop_support: `runtime_terminal_milestone` populates it from the `RuntimeCapabilityFailure` message on the model-visible failure arm; host/infra and gate-denied paths emit `None`. - ironclaw_event_streams: add `error_detail` to the live `CapabilityActivity` projection item. - ironclaw_product_adapters: add `error_detail` to `CapabilityActivityView` (+ input, wire ser/de) with the same bounded/sanitized boundary validation as the other display fields. - ironclaw_reborn_composition: live progress sanitizes the milestone summary and maps it onto the view; the history/runtime-payload path keeps carrying detail via the separate CapabilityDisplayPreview. - frontend (history-messages.js): `toolCardFromActivity` prefers `error_detail` over the bare kind. Bundle rebuilt. The durable runtime event log still records only the failure kind (no backend-detail persistence), per ironclaw_turns guardrails. Regression: `webui_event_stream_projects_live_tool_failure` now drives a `CapabilityFailed` milestone with a `safe_summary` and asserts the live `CapabilityActivityView.error_detail` carries it end-to-end. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn): render dispatch failure kinds as plain language, not category tokens (nearai#5289) When a capability dispatch fails without a host-authored safe_summary, the failure message fell back to the error's Display, which exposed the stable redacted category token (e.g. "dispatch failed: InputEncode") in the per-tool UI Error tab. Add `human_summary()` to `RuntimeDispatchErrorKind` and `DispatchFailureKind` (fixed host-authored sentences, no raw content) and use it as the fallback in `sanitized_failure_message`'s dispatch arm. The stable `as_str()` token stays the contract for routing/metrics/audit; only the user-facing message changes. So "InputEncode" now reads "the tool input could not be encoded". Tests: `dispatch_failure_kind_human_summary_is_plain_language_not_category_token`; updated the two production.rs tests that pinned the old token wording (they now assert the human summary and still verify no raw backend string or secret leaks). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn): close prune_run race in record_result_with_preview (nearai#5289) Code review flagged a race between the display-preview store's remove-from-pending + insert-into-completed pair and `prune_run`: the two locks were acquired independently, so a concurrent prune could interleave and leak a completed preview record that is never pruned. The failure path (`record_failure_preview`) was already fixed to hold both locks; apply the same fix to the pre-existing success path (`record_result_with_preview`) so the whole pattern is consistent. Both sites and `prune_run` acquire pending-before-completed, so holding both cannot deadlock. (The other two review findings — non-durable failure staging, and a hand-rolled char-boundary truncation helper — were already resolved in earlier commits on this branch.) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: update webui v2 capability activity fixtures * test: expect human dispatch failure summary * test: stabilize reborn composer active-run smoke * fix(webui): stop bare error-kind activity frame clobbering failure detail (nearai#5289) A failed capability surfaces its detail on two frames that race in the live stream: the runtime-payload `capability_activity` deliberately carries no `error_detail` (its `toolError` is the bare kind, e.g. "invalid_input"), while the `capability_display_preview` carries the real reason. The tool-activity merge used `incoming.toolError || current.toolError`, so a late bare-kind activity frame overwrote the detailed reason already set by the preview — the live card showed "invalid_input" while a page refresh (replay) showed the real message. `mergedToolError` now refuses to downgrade: if the incoming error is only the bare kind and the current one is a different (detailed) message, the detail is kept. First-frame and genuine upgrades are unaffected. Test: tool-activity-state.test.mjs drives the real preview-then-activity race and asserts the detail survives. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: stream capability failure detail in webui events * test: update event stream capability activity fixture * fix: redact filename-bearing failure summaries * fix: preserve redacted capability failure summaries * fix: keep invalid tool failure summaries nonfatal * fix: harden capability failure summary redaction * fix: require filesystem context for filename summaries * fix: tighten workspace failure summary matching * fix: preserve invalid input failure summary * fix: keep input encode summaries replay-safe * fix: preserve safe invalid input summaries * fix: harden capability failure summaries * fix: redact sensitive capability issue fields * fix: document projected capability error details * fix: catch sensitive issue marker variants * docs: clarify projected error detail boundary --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…nearai#5430) * ci(reborn): add cargo-llvm-cov integration-tier coverage job (T0-COV) Adds a per-PR coverage signal for the Reborn backend. A new `Reborn Coverage` workflow runs cargo-llvm-cov over the in-process integration-tier test binaries (tests/reborn_integration_*.rs + tests/reborn_group_*/) and surfaces a Reborn-crate-scoped line-coverage percentage plus a per-crate hole list in the PR job summary. This makes the roadmap's "100% int-tier coverage" goal measurable (docs/reborn/reborn-backend-coverage-roadmap.md, T0-COV) and yields the real hole list. It is informational only and does not gate PRs. - .github/workflows/reborn-coverage.yml: the job. Combined `cargo llvm-cov --workspace --test <int-tier>` invocation so the report is scoped to every linked member crate (the standalone `report` subcommand cannot take --workspace and would scope to the root package only). Default features, matching run-reborn-root-partition.sh; no live Postgres needed. Reuses coverage.yml's pinned action SHAs / install action. Surfaces via $GITHUB_STEP_SUMMARY per repo norm. - scripts/ci/reborn-coverage-int-tier-tests.sh: dynamically discovers the int-tier `--test` targets (auto-expands as new suites land). - scripts/ci/reborn-coverage-summary.sh: filters the llvm-cov JSON export to the Reborn crate families and renders the % + per-crate table. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(reborn): sticky PR comment for int-tier coverage (T0-COV) Surface the Reborn coverage %/hole-list as an upserted sticky PR comment so it lives in the conversation, not just the Actions job summary. Visibility only — never gates: the step is `continue-on-error: true` and PR-only, so a fork PR's read-only GITHUB_TOKEN (comment API 403s) can't red the check, and nothing in the job exits non-zero on a coverage number. - reborn-coverage.yml: grant `pull-requests: write`; add a "Post sticky coverage comment" step after the existing $GITHUB_STEP_SUMMARY render (kept as-is), before the artifact upload. - reborn-coverage-comment.sh (new): upsert one marker'd comment via pure `gh api` (no new action). Reuses reborn-coverage-summary.sh for both the body and the breadth holes (--zero-crates) — no duplicated aggregation jq. Prepends a 0-coverage breadth callout (informational; "target: 0" is the roadmap goal, not a check). Lookup uses `--paginate` so a busy PR's aged sticky isn't missed (which would post duplicates); marker passed via env to jq (no string interpolation); ids captured before `head` to avoid a pipefail SIGPIPE. - reborn-coverage-summary.sh: add `--zero-crates` mode (single owner of the crate filter + aggregation); fold the per-crate table into <details> so the comment/summary stays compact (headline % + callout always visible); reword the informational note to cover the callout too. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(reborn): gate coverage rust-cache saves to main (T0-COV) The reborn-coverage job had no `save-if`, so it wrote a multi-GB instrumented rust-cache on every PR push. GitHub scopes caches per-branch, so those PR-branch saves are invisible to other PRs and only churn the ~10 GB repo cache LRU — evicting the one main-seeded cache that PRs actually restore from. Gate saves to `push` on `main` (the workflow already triggers there), matching the convention in test.yml/code_style.yml/replay-gate.yml. PRs now restore the warm main-seeded instrumented cache and skip the wasteful save. Instrumented caches stay separate from the non-instrumented reborn-tests caches by design (coverage RUSTFLAGS change every crate's fingerprint — no cross-reuse possible). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(reborn): address PR review — multi-dataset jq, token guard, job timeout (T0-COV) Fixes from PR nearai#5430 review (valid findings only): - summary.sh: iterate all llvm-cov `data[]` datasets (`.data[]?.files[]?`) instead of `.data[0]` — the export format permits multiple datasets and index-0-only would undercount. The trailing `?`s keep the empty/missing-data path crash-free (unchanged behavior; `{"data":[]}` / `{}` still fall through to the no-data message). - comment.sh: fast-fail if GH_TOKEN is unset, matching the existing GITHUB_REPOSITORY/PR_NUMBER guards (clearer failure for local runs; gh needs it for the API calls). - reborn-coverage.yml: add `timeout-minutes: 180` as a runaway-hang backstop (well above any real build; halves GitHub's 360-minute default). Corrects the prior comment that claimed coverage.yml runs its instrumented job without a timeout — its matrix job caps at 60. Rejected as non-issues: jq null-index "crash" (jq indexes null as null, not an error — verified); group_by requiring pre-sorted input (jq's group_by sorts internally); needing `issues: write` (the sticky comment already posts on a same-repo PR with pull-requests: write). `clean --workspace` preserving the cache is intentional (it scopes to workspace members, keeping instrumented deps). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(reborn): regression tests for coverage CI helpers (T0-COV) Add scripts/ci/test-reborn-coverage.sh (mirrors the test-classify-test-scope.sh precedent) covering the three new coverage helpers — closes the "no committed edge-case test" review findings on PR nearai#5430: - summary.sh report mode: mixed Reborn/non-Reborn filtering, no-match no-data message, empty/absent `data` no-crash, multi-dataset `.data[]` aggregation, zero-covered crate sorted to top. - summary.sh --zero-crates: exact zero-covered name list, all-covered/empty → empty output. - comment.sh upsert (fake `gh` on PATH): no-sticky → POST with marker body, existing sticky at a non-first list position → PATCH (not a duplicate POST), zero-covered callout prepended. - int-tier discovery: empty tree → exit 1, single file/dir suites, mixed suites sorted+deduped. 43 cases, self-contained (mktemp + trap cleanup), shellcheck-clean. Not wired into a workflow — matches the unwired test-classify-test-scope.sh precedent; run manually as a local regression signal. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(reborn): cover comment.sh env-var guards (T0-COV) Add C4-C6 to test-reborn-coverage.sh: run reborn-coverage-comment.sh with each of GH_TOKEN / PR_NUMBER / GITHUB_REPOSITORY unset and assert a non-zero exit plus the matching "must be set" message on stderr. Pins the fast-fail guards (the GH_TOKEN one added in 7b617b7) that C1-C3 didn't exercise. Uses the existing `capture env -u …` pattern (49 cases, shellcheck-clean). Addresses the CodeRabbit review finding on PR nearai#5430. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(reborn): make assert_line_before non-fatal on missing needle (T0-COV) `assert_line_before`'s `grep -n | head | cut` pipelines run under `set -euo pipefail`; a missing needle made grep exit 1, propagated by pipefail to the assignment, aborting the whole suite before the function's own empty checks could report a normal FAIL — contradicting the "run every case" harness design. Append `|| true` so a missing needle yields an empty line number and falls through to the FAIL path. Addresses the Copilot review finding on PR nearai#5430. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(reborn): PR-number concurrency key + drop dead roadmap citation (T0-COV) Two PR review fixes: - reborn-coverage.yml: key the pull_request concurrency group on `github.event.pull_request.number` instead of `github.head_ref`. head_ref is just the source branch name, so two fork PRs sharing a common branch name (main/fix/…) collided on the same group and, with cancel-in-progress, one push could cancel the other PR's coverage run. number is null off-PR, so it falls back to github.ref for push/dispatch (pattern from pr-label-classify.yml). - Drop the citation of docs/reborn/reborn-backend-coverage-roadmap.md from the workflow header and the int-tier discovery script — that file exists on neither this branch nor main, so it was a dangling source-of-truth pointer. The int-tier scope is already stated inline (reborn_integration_*.rs + reborn_group_*/), so nothing is lost. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(reborn): allowlist-boundary + missing-file coverage cases (T0-COV) Add three regression cases to test-reborn-coverage.sh (PR nearai#5430 review): - A6: crate-filter boundary — fixture with exact single-crate matches (ironclaw_architecture, ironclaw_slack_v2_adapter), a family-prefix crate (ironclaw_reborn_config), and a lookalike (ironclaw_architecture_extra). Asserts the exact + prefix crates appear, the lookalike is excluded, and its 999 lines are dropped from the aggregate (16/30, not 16/1029) — pins the prefix-vs-exact regex semantics against a future edit that widens the match. - A7: summary.sh on a nonexistent JSON path -> non-zero exit + "coverage JSON not found". - C7: comment.sh on a nonexistent JSON path -> non-zero exit + "coverage JSON not found" + NO gh mutation (guard fires before any gh call). 60 cases, shellcheck-clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(reborn): malformed-JSON aborts comment before any mutation (T0-COV) Add case C8 to test-reborn-coverage.sh: an existing-but-malformed coverage JSON. Distinct from C7 (missing file) — it passes comment.sh's `[ -f ]` guard and must instead fail at the reborn-coverage-summary.sh render (jq parse error under `set -e`), before any POST/PATCH. Asserts non-zero exit and no fake-gh mutation, pinning the render-before-mutate ordering. Exit asserted non-zero (not an exact code) since jq's parse-error exit varies across versions. 62 cases, shellcheck-clean. Addresses the PR nearai#5430 review finding. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(webui-v2): avoid duplicate logs header during chat runs * fix(webui-v2): hide chat run logs shortcut * test(webui-v2): remove redundant chat logs assertion * fix(webui-v2): move chat logs link into message actions * fix(webui-v2): float chat logs shortcut * fix(webui-v2): use terminal icon for chat logs shortcut * fix(webui-v2): reserve chat logs space with spacer * style(webui-v2): soften floating logs shortcut * style(webui-v2): increase chat logs shortcut contrast * fix(webui-v2): reuse scoped logs path builder * test(e2e): handle busy reborn sends before settling * fix(webui-v2): keep chat logs route construction in chat * test(e2e): retry transient submit errors while settling
github-actions
Bot
merged commit Jul 1, 2026
eb861d4
into
native-matrix-channel-pilot
17 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Automated sync from nearai/ironclaw. The
upstream-mainbranch is an exact fast-forward mirror ofnearai/main; this PR imports it intonative-matrix-channel-pilotfor CI with upstream source winning unrelated-history conflicts. When checks pass, the Matrix pilot branch is fast-forwarded so upstream ancestry is preserved.