Skip to content

release: prepare 1.1.1-rc.1 - #7427

Closed
serrrfirat wants to merge 17 commits into
release/1.1.0-rc.1from
release/1.1.1
Closed

serrrfirat wants to merge 17 commits into
release/1.1.0-rc.1from
release/1.1.1

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • Backport urgent IronHub/custom MCP, WebUI, retrieval, runtime-credential, Slack, and Telegram fixes onto the 1.1 release line.
  • Default the retained rc1 Slack/Telegram migration to skip legacy channel state safely; operators may explicitly opt into verified import.
  • Document durable container storage and exact state handling from stable 1.0.0 and 1.1.0.
  • Version the candidate as 1.1.1-rc.1 and gate publishing on artifact upgrade canaries from both stable predecessors.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Related #7268, #7217, #7286, #7267, #7135, #7288, #7307, #7329, #7361, #7363, #7300

Validation

  • cargo fmt --all -- --check
  • cargo clippy -p ironclaw --all-targets --all-features -- -D warnings
  • cargo check -p ironclaw --locked
  • Relevant tests pass: release workflow smoke contracts, release-upgrade canary sabotage tests, cut-release tests, release migration unit tests, and migration-barrier integration test
  • cargo test --features integration: Not applicable; no new database behavior in release-preparation commit, and the backported fixes retain their focused regression coverage
  • Manual testing: seven-platform compile preflight and immutable artifact canaries will run on the pushed candidate SHA
  • Independent release-blocker review completed before requesting review

Test Strategy

User behavior: The candidate preserves channel delivery/pairing, IronHub prompt installation, custom MCP setup, stable WebUI streaming, durable retrieval, and container artifacts across redeploys when configured on durable storage.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: release workflow structural contract requires stable 1.0.0 and 1.1.0 predecessor matrix legs.
  • Reborn integration: release-pair migration barrier test.
  • Recorded fixture: Not applicable; no fixture protocol changed.
  • Browser E2E: Existing focused backport coverage; no new browser behavior in release metadata.
  • Backend or runtime: release migration unit suite and locked CLI compile.
  • Live canary: Tag publisher must pass upgrade/restart/rollback/re-upgrade using exact artifacts from both stable predecessors before publication.

What the tests prove: Publishing remains fail-closed until both supported upgrade paths pass against the exact checksummed candidate artifact, and default channel-state skip evidence does not fabricate migrated counts.

Commands run:

  • cargo fmt --all -- --check
  • cargo test -p ironclaw --test smoke release_ci_ -- --nocapture
  • python3 tests/test_release_upgrade_canary.py
  • uv run --python 3.12 python scripts/ci/test_cut_ironclaw_release.py
  • python3 scripts/ci/ws12_workflow_contracts.py
  • cargo check -p ironclaw --locked
  • cargo test -p ironclaw_release_migration
  • cargo test -p ironclaw_reborn_composition --test release_pair_migration_barrier
  • cargo clippy -p ironclaw --all-targets --all-features -- -D warnings

Security Impact

No trust boundary is weakened. Release artifact jobs remain read-only; only the host publishing job receives contents write after both upgrade matrix legs succeed. Previous release downloads and candidate archives are checksum-verified.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: unchanged.
  • Untrusted content enters prompts only through an envelope/escaping primitive: IronHub prompt asset backport preserves signed asset validation.
  • Hashes declare purpose; release archives use published SHA-256 checksums.
  • New/changed status, exit, policy, runtime, or error variants: no new variants in release-preparation commit.
  • Security/durability serde(default) fields fail closed or have migration tests.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic: unchanged by release preparation.
  • Driver/operator-visible errors have stable class semantics: unchanged.
  • Sandbox/native/host names accurately describe trust boundary: unchanged.

Database Impact

No new schema migration in the release-preparation commit. From 1.0.0, startup applies existing additive v33/v34 migrations and bounded record migration; from 1.1.0, no offline transform is required. Both paths are artifact-canary gated.

Blast Radius

Release metadata, release publishing CI, startup migration policy, IronHub/custom MCP, WebUI streaming, libSQL retrieval, WASM staged credentials, and Slack/Telegram channel behavior.

Rollback Plan

Do not promote the RC. Stop writers, preserve/restore the pre-upgrade database and volume snapshot, and run the exact predecessor binary. The record migration retains old authorities; 1.0.0 workspace snapshots remain unchanged for rollback.

Review Follow-Through

Reviewer judgment is requested on the supported predecessor matrix, upgrade notes, and default Slack/Telegram reconfiguration policy. Stable promotion is intentionally out of scope until RC artifact QA completes.


Review track: C (runtime/persistence/release CI)

henrypark133 and others added 15 commits August 10, 2026 10:04
* fix(ironhub): install signed prompt assets

* fix(ironhub): reject colliding prompt assets
)

`useChatEvents` kept three refs of its own -- `settledRunsRef`,
`latestRunIdRef`, `promptRunIdRef` -- while the thread-switch reset lives
in `useChat`. Nothing cleared them. `latestRunId` is only cleared on a
TERMINAL run status, so a run that got stuck pinned it for the life of
the mounted chat page and then followed the user into every thread they
opened afterwards, where it was consumed as:

  1. the run-id fallback for capability frames that omit `turn_run_id`,
     attributing another thread's tool cards to the stuck run, and
  2. the seed for the stale-terminal check, so the newly opened thread's
     own terminal status was silently discarded -- `onRunSettled` never
     fired, the durable timeline was never refetched, and leftover live
     assistant text was never replaced by the finalized reply.

Move the three slots into `lib/run-tracking-state.ts` behind one ref that
`useChat` owns and resets in the `[threadId]` effect it already has,
mirroring `createToolActivityState` / `resetToolActivityState`.
`useChatEvents` now holds no state at all, so "what does a thread switch
mean" is answered in exactly one place.

The test harness could not express this bug: its `useRef` stub minted a
fresh object per call, so every thread switch looked like a clean slate.
It now uses call-order ref slots and a deps-aware `useEffect`, matching
React. All 54 pre-existing tests pass unchanged against the faithful
stub, which is what says the old one was lenient rather than
load-bearing.

Tests: 9 cases in useChatEvents.test.ts (7 reproductions across untagged
capability_activity / capability_display_preview / projection items,
consecutive switches, completed and failed terminal statuses, and
leftover streaming assistant text; 2 controls pinning same-thread run-id
inference and the normal settle path), plus a caller-level test driving
the real `useChat` across a thread switch -- the handler suite stubs
`useChatEvents`, so only that one proves `useChat` performs the reset.
All 7 fail before this change; the caller-level test was verified red by
removing the reset line.
* fix(loop): preserve result_read continuation reference

Return the original pageable result reference from result_read while retaining inline-only chunk persistence and evidence. Key transcript dedup by provider call so multiple pages can safely share one durable source reference, with regression coverage for replay and a two-page continuation.

* test(composition): align result read continuation assertions

* fix(loop): pin result updates to provider calls

* fix(ci): update subagent metadata fixture

* test(composition): keep result lookup helper test-only

* test(reborn): restore await-edge fixture compatibility
)

* fix(filesystem): treat FTS filters as plain text

* fix(filesystem): address review — PG stop list, term-match tests, DRY (#7288)

Addresses multi-agent review on #7288:

- Plain-FTS stop words now mirror PostgreSQL's fixed english stop list
  (shared/english.stop) verbatim, so in-memory/libSQL required terms match
  plainto_tsquery('english', ...) exactly; documents the remaining
  stemming divergence (FTS5 matches literal terms).
- Exhaustive table-driven test pins every stop word (case-insensitive) plus
  required non-stop words (please/tell/would/could).
- libsql FTS contract test adds a partial-match negative document so the
  FTS5 implicit-AND join is distinguishable from an OR join, and proves the
  negative doc is searchable by its own terms.
- In-memory reference matcher now tokenizes stored text (whole-token
  matching, mirroring FTS5 unicode61) instead of substring containment,
  fixing contractions divergence; FTS queries are tokenized once per query
  instead of once per scanned record.
- core_builtin harness: shared-filesystem variant reuses the recording
  harness assembly tail instead of re-copying it.

* test(integration): prove memory recall is scope-isolated on the libSQL path (#7288)

The proactive-recall scenario only checked that a never-written marker was
absent, which says nothing about scope isolation. Seed a second user's
MEMORY.md — word-for-word the canonical document apart from the marker —
into the same libSQL composite, then assert the canonical user's explicit
memory_search and proactive prompt both still return plum-42 and never the
other user's marker, while that marker stays retrievable in its own scope.

The seed goes through the native provider rather than a second actor's
thread: this group pins capability dispatch to one fixed user, so a second
actor's memory write would land in the canonical scope anyway.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…) (#7329)

Third-party WASM guests (ironhub tools such as attio) gate on the
secret-exists host import before issuing any request, but production
wired the sandbox with the deny-all default, so the probe always
returned false: attio aborted pre-network with "API key not
configured" and the host classified the plain-string guest error as
operation_failed, never auth_required.

Introduce StagedWasmHostSecrets, a per-invocation WasmHostSecrets
implementation over the staged secret injection store: exists(name) is
true exactly when authorization leased and staged non-empty credential
material for (scope, capability_id, handle), read non-destructively so
the HTTP egress still receives the material. Wire it into
WasmRuntimeAdapter::host_for_scope on every host variant and plumb the
shared store through the builder.

Credential staging now rejects empty resolved material as
AuthRequired (obligation handler and host-driven staging), so a
configured-but-blank key surfaces the typed re-auth signal instead of
an opaque guest failure. No prose heuristics: the structured
{"kind":"auth_required"} guest contract remains the fallback.

Adds unit tests for the probe semantics and WASM contract tests with a
secret-exists probe component (staged -> true, absent -> false, empty
material -> AuthRequired staging error).
…signal, builtin description trust, docs (#7361)

* fix(extensions): host-bundled description trust + already-connected install confirmation

Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread
e79a994f, run 251aec0b on ironclaw-qa-testing-libsql):

1. Host-bundled capability descriptions were description_trust=Untrusted,
   so the loop-tier prompt-text denylist strict-scanned compiled-in text
   and silently omitted builtin.extension_register_hosted_mcp from every
   model prompt's capability surface ("browser authorization-code flow"
   matched the "authorization" credential pattern). HostBundled is the
   only source eligible for effective FirstParty/System trust, so its
   repo-authored descriptions now cross the verified-catalog boundary
   like signature/digest-verified registry installs. Untrusted provenance
   (InstalledLocal, UserRegistered, unknown) keeps the strict scan.

2. When install-driven activation passed the credential gate because the
   caller's declared requirements were all satisfied, the response never
   said so — the model got only conditional guidance ("If WebChat shows
   an account connection panel...") and deflected an explicit "connect
   account" request to the web interface even though the account was
   already connected. The install response now appends an explicit
   already-connected confirmation exactly when declared requirements
   were verified present for the calling user.

Regression tests: manager surface test pins VerifiedCatalog trust for all
model-visible lifecycle capabilities through the real host runtime;
instruction-bundle tests pin retain/omit behavior for auth-vocabulary
descriptions by trust; install-path tests pin the confirmation on the
seeded-credential path and its absence for credential-free extensions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channels): chat can drive the personal half of channel connect

The onboarding and channels pages claimed "asking the agent to connect a
channel doesn't work" and that the agent "may tell you it can't help".
That describes only the operator half (registering app/bot credentials).
The per-user half has shipped since early July: extension_install runs the
same activation credential gate as the Channels card, raises the in-chat
OAuth connection panel when the account is unconnected, and (as of the
sibling fix) confirms when it is already connected.

The self-knowledge protocol makes these pages the model's authority on
IronClaw's own capabilities, so the stale claim scripted the exact
refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an
account that was already connected. Correct both pages to distinguish
the operator step from the chat-drivable personal connect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channels): align slack and telegram setup notes with the connect contract

The slack page's operator-step note and the telegram troubleshooting
accordion still carried the blanket "asking the agent to connect will not
work" claim the overview/onboarding correction removed — same drift,
different phrasing (review catch on #7361, plus one more instance found
by a broader sweep). Both now state the two-step contract: the operator
half stays in the web interface; after it, chat drives the personal half
(install/activate -> in-chat connection or pairing panel, or an
already-connected confirmation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(golden): recapture surface digests over the description-trust change

The queue run failed golden_payload because the branch predated current
main and its own surface.rs trust fix changes the surface digest. The
recaptured snapshots differ ONLY in the surface sha256 lines (verified
char-by-char) — no prompt text or capability-list changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Users habitually type /pair from the earlier pairing flow. Keep every
suggested wording on /start (the vendor deep-link convention) and accept
/pair <CODE> as a declared inbound-code-prefix alias so those users pair
instead of looping through the connect nudge.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 10, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • staging
  • reborn-integration

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: de5c2974-6faa-4be6-92b0-ad2b680e4a2f

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added scope: ci CI/CD workflows scope: docs Documentation scope: dependencies Dependency updates size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Aug 10, 2026
@ironloopai

ironloopai Bot commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

🧭 IronLoop Run · Review

This comment updates in place as the Run moves through its stages.

⬛ Final result · Stopped

🟨 Queued → 🟦 Working → ⬛ Stopped

Automatic trigger · attempt 1 of 3 · stopped after 7m 34s

IronLoop stopped because the pull request target branch or head changed while this Run was active.

Run details

Run: a76408ba-fa4d-43ef-8c8f-cd0d2baa8d13
Base: release/1.1.0-rc.1 at f4c361f
Head: release/1.1.1 at 913e5eb
Created: 2026-08-10 09:55 UTC
Updated: 2026-08-10 10:03 UTC

This branch was previously deployed

1 inactive deployment
ironclaw-rebornneari / rc-1-testing — 6fb60b1d Deployed Aug 10, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: ci CI/CD workflows scope: dependencies Dependency updates scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants