Skip to content

refactor(reborn): complete host-side gate reconstitution — resume input_ref, local-dev gate persistence, GateRef reconcile (§5.3 flip prep / Stage 0) - #6278

Merged
ilblackdragon merged 4 commits into
refactor/reborn-resolution-model-visiblefrom
refactor/reborn-host-reconstitution-stage0
Jul 19, 2026
Merged

ilblackdragon merged 4 commits into
refactor/reborn-resolution-model-visiblefrom
refactor/reborn-host-reconstitution-stage0

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

Stage 0 of the capability-result flip (docs/reborn/2026-07-17-architecture-simplification-dto-dyn-local.md, §5.2.9/§5.3): a no-flip host-reconstitution slice closing two verified correctness hazards ahead of the atomic flip. The loop-facing trait is UNCHANGED and CapabilityOutcome is untouched; only host-side plumbing moves. Base tip verified at/after PR-B (a0c497b2d ancestor); 2a-i ReplayPayloadStore present.

Fix 1 — resume-time input_ref reconstitution (hazard 3)

On a gate/auth resume, the effective input_ref used for the idempotency key and validation is now derived from the host-private ReplayPayload persisted at the fresh gate raise (loaded by the resume's invocation id) — not the advisory loop-supplied approval_resume.input_ref. A new resume_replay_payload helper centralizes the derivation so the invoke_capability seam and invoke_capability_dispatch compute the same byte-stable key.

The idempotency key is now derived lazily inside persist_gate_record_for_outcome (after dispatch, only once a gate record exists) so a missing/stale resume payload can never pre-empt dispatch's own resume identity/activity validation — a malformed resume still surfaces InvalidInvocation, a genuinely-missing payload still fails closed (the 2a-i missing-payload test stays green).

  • Wiring: invoke_capability → invoke_capability_dispatch(request.clone()) → persist_gate_record_for_outcome(&request, &outcome); dispatch loads the payload once (after resume_mode validation) and reuses it for both the store-derived input_ref and {input, estimate}.
  • Test (failed first): approval_resume_derives_input_ref_and_key_from_store_not_loop_supplied — a resume with a WRONG loop-supplied input_ref still reconstitutes the original input from the store, and a second resume differing ONLY in that advisory ref replays the cached outcome instead of re-dispatching. Verified red before the fix (resume_request_count == 2, re-dispatch), green after. Updated invoke_capability_checks_registered_activity_on_approval_resume_input_ref to seed the payload it now reconstitutes from.

Fix 2 — gate-record persistence for the local-dev synthetic approval producer (hazard 2)

OutboundDeliveryTargetSetHandler::request_approval raises its approval gate outside the loop-host persist seam (the synthetic-capability wrapper bypasses persist_gate_record_for_outcome), so it now persists a GateRecord::Approval itself at the raise via the wired GateRecordStore, keyed by the canonical GateRef::for_approval_request, in the run's resource-owner scope. Without it the §5.2.9 host-side gate rendering would have no record to read after the flip.

  • Test (failed first): the existing local-dev approval journey now asserts the GateRecord persisted at the raise loads and round-trips. Red before the fix (load returned None).

Fix 3 — reconcile the two approval GateRef encodings

Canonical encoding chosen: GateRef::for_approval_request(id) = GateRef::from_uuid(id.as_uuid()) — a new ironclaw_host_api helper — as the GateRecordStore key. This is the ironclaw_host_api::GateRef (a uuid newtype, the record key), a different type from the loop-facing gate:approval-{id} routing ref (ironclaw_turns::GateRef). ironclaw_capabilities::host now uses the helper at both authorize sites.

Because the canonical is the uuid record-key GateRef and not the routing string, this touches no persisted wire format and does not break the is_approval_gate_ref (gate:approval- prefix) predicate.

  • Test: a host-persisted approval gate ref resolves through the read model — the routing gate:approval-{id} ref recovers the id via approval_request_id_from_gate_ref, the canonical key re-derives, and the persisted GateRecord is found (asserted in the local-dev approval test); plus a host_api round-trip unit test.

Deviation note: the brief suggested canonicalizing on from_uuid — that is exactly what this does for the record-key GateRef. The RISKY direction (rewriting the gate:approval-{id} routing string to a bare uuid) was not taken: it would break is_approval_gate_ref prefix routing used across production routing/projection/channel-delivery, plus persisted adapter command strings ("approve gate:approval-X") and delivered-gate route records — the STOP-and-report condition in the brief.

Verification (all green)

  • cargo test -p ironclaw_loop_host — 378 lib (incl. new + updated resume/seam/2a-i)
  • cargo test -p ironclaw_capabilities -p ironclaw_product_workflow -p ironclaw_run_state -p ironclaw_host_api --all-features
  • cargo test -p ironclaw_reborn_composition --lib --all-features — 1629 (env-clean; the only 6 failures are pre-existing llm_admin nearai env-var tests polluted by the machine's NEARAI_API_KEY/LLM_BACKEND, green with those unset, and touch none of this diff)
  • Integration: reborn_integration_outbound_target (16), reborn_integration_auth_gate (18), reborn_integration_reopen_resume_through_gate (13), reborn_group_approvals (15)
  • cargo test -p ironclaw_architecture green
  • cargo clippy --all-targets --all-features -- -D warnings clean on all 6 touched crates
  • scripts/pre-commit-safety.sh clean (exit 0)

Trait unchanged; CapabilityOutcome untouched; invoke_capability not flipped; executor Resolution consumers not touched.

🤖 Generated with Claude Code

…ut_ref, local-dev gate persistence, GateRef reconcile (§5.3 flip prep / Stage 0)

Stage 0 no-flip host-reconstitution slice: closes two verified correctness
hazards ahead of the atomic capability-result flip. The loop-facing trait and
`CapabilityOutcome` are UNCHANGED; only host-side plumbing moves.

Fix 1 — resume-time input_ref reconstitution (hazard 3). On a gate/auth resume
the effective input_ref used for the idempotency key + validation is now derived
from the host-private `ReplayPayload` persisted at the fresh gate raise (loaded
by the resume's invocation id), not the advisory loop-supplied
`approval_resume.input_ref`. A new `resume_replay_payload` helper centralizes the
derivation so `invoke_capability` (seam) and `invoke_capability_dispatch` compute
the SAME byte-stable key. The key is now derived lazily inside
`persist_gate_record_for_outcome` (after dispatch, only for a gate outcome) so a
missing/stale resume payload never pre-empts dispatch's own resume
identity/activity validation — a malformed resume still surfaces
`InvalidInvocation`, and a genuinely-missing payload still fails closed
(2a-i missing-payload test stays green). Test proves a resume derives the same
input_ref/key from the store even when the loop-supplied ref differs (a second
resume differing only in that advisory ref replays the cached outcome instead of
re-dispatching); failed first (re-dispatched, count 2) before the fix.

Fix 2 — gate-record persistence for the local-dev synthetic approval producer
(hazard 2). `OutboundDeliveryTargetSetHandler::request_approval` raises its gate
OUTSIDE the loop-host persist seam, so it now persists a `GateRecord::Approval`
itself at the raise (via the wired `GateRecordStore`), keyed by the canonical
`GateRef::for_approval_request`, so host-side gate rendering (§5.2.9) has a record
to read. Test proves the record persists and round-trips; failed first (load
returned None) before the fix.

Fix 3 — reconcile the two approval GateRef encodings. Canonicalized on the typed
`GateRef::for_approval_request(id) = GateRef::from_uuid(id.as_uuid())` — a new
`ironclaw_host_api` helper — as the GateRecordStore key. This is the host_api
uuid GateRef (record key), a DIFFERENT type from the loop-facing
`gate:approval-{id}` routing ref (`ironclaw_turns::GateRef`), so reconciling here
touches NO persisted wire format and does not break the `is_approval_gate_ref`
prefix-routing predicate. `ironclaw_capabilities::host` now uses the helper at
both authorize sites. Test proves a host-persisted approval gate ref resolves
through the read model: the routing ref recovers the approval id
(`approval_request_id_from_gate_ref`), the canonical key re-derives, and the
persisted GateRecord is found.

Note: the task suggested canonicalizing on `from_uuid`; that is exactly what this
does for the record-key GateRef. The RISKY direction (rewriting the
`gate:approval-{id}` routing string to a bare uuid) was NOT taken — it would
break the `is_approval_gate_ref` prefix predicate used in production routing,
projection, and channel delivery, plus persisted adapter command strings and
delivered-gate route records.

Tests: loop_host 378 lib (incl. new + updated resume/seam/2a-i); host_api helper
round-trip; composition 1629 (env-clean); capabilities/product_workflow/run_state
green; architecture green; integration outbound_target/auth_gate/
reopen_resume_through_gate/group_approvals green; clippy -D warnings clean on all
touched crates; pre-commit-safety clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@ironloopai

ironloopai Bot commented Jul 19, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: c250e47c7e0c605f7d145f3fe5b2467bb59125cf
Result: One or more review results were superseded by a newer PR head.
Next: Run @ironloopai review on the latest PR head.
Updated: 2026-07-19T09:30:21.173Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Superseded N/A N/A 2026-07-19T09:14:27.846Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Superseded by a newer PR head. New head: c8595fd. Previous verdict: Changes requested.
Recent activity
Time Reviewer State Detail
2026-07-19T07:55:16.361Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head f1eaf2f.
2026-07-19T07:55:16.361Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-19T07:55:17.092Z ironloop/common-reviewer (reviewer) Started Reviewer worker started.
2026-07-19T07:55:19.442Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (head_ref) at f1eaf2f.
2026-07-19T08:01:20.321Z ironloop/common-reviewer (reviewer) Result captured Changes requested; 1 blocking finding.
2026-07-19T08:01:20.321Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
2026-07-19T09:14:27.846Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (c8595fd).
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@coderabbitai

coderabbitai Bot commented Jul 19, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • staging
  • reborn-integration

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: b9aa301e-8509-4ec4-84bd-b98d195b1d54

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 19, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a canonical GateRef::for_approval_request helper to derive gate references from approval request IDs. It updates the host-side resume reconstitution logic to derive the effective input_ref and idempotency key from the host-persisted replay payload rather than relying on the advisory loop-supplied resume.input_ref, ensuring byte-stable idempotency keys. Additionally, the local-dev synthetic approval producer now persists a durable GateRecord at the gate raise, and corresponding unit tests have been added. There are no review comments, so I have no feedback to provide.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

❌ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
❌ Changes requested 1 0 1 f1eaf2fd4b6a

Head: f1eaf2fd4b6aa1282ea5107c46b286ed5b776788
Next: Fix the blocking findings, push the PR branch, then re-run this reviewer.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

Normal runtime approval gates still persist under a random GateRef, so the new canonical approval-record key cannot retrieve them after the flip.

Findings

Blocking: 1 / Notes: 0

Blocking findings

1. ❌ [MEDIUM] Use the canonical key when persisting normal approval gates

Location: crates/ironclaw_capabilities/src/host.rs:733
This derived ref is discarded by CapabilityHost::invoke_json, which returns only AuthorizationRequiresApproval; DefaultHostRuntime then rebuilds RuntimeApprovalGate from the approval request ID. When the loop-host persists that ordinary ApprovalRequired outcome, capability_outcome_to_resolution still mints GateRef::new(), so the GateRecord is stored under a random key. The approval read model instead derives GateRef::for_approval_request from gate:approval-{id}, and therefore cannot load records for normal runtime capability approvals. The new local-dev synthetic test covers only its separately keyed producer. Persist normal approval records under GateRef::for_approval_request (or make the reader follow a durable loop-to-host binding) and add an end-to-end normal-runtime approval lookup test.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.

Ok(AuthorizeFold::Blocked {
result: AuthorizeResult::Blocked(Blocked::Approval(GateWaypoint::new(
GateRef::from_uuid(approval_request_id.as_uuid()),
GateRef::for_approval_request(approval_request_id),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This ref does not become the normal runtime gate-record key: invoke_json discards AuthorizeResult, and loop-host persistence maps CapabilityOutcome::ApprovalRequired to a freshly minted GateRef::new(). The read model derives GateRef::for_approval_request from gate:approval-{id}, so it will not find ordinary runtime approval records. Please wire the canonical key through that persistence path (or durably follow the binding) and test the normal runtime case.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Accurate observation, and I verified it — but it describes the deferred atomic-flip work, not a defect in this additive slice. Two facts settle the disposition:

  1. No live consumer. I grepped every GateRef::for_approval_request / GateRecordStore::load call site: the only read-model that re-derives for_approval_request is the local-dev synthetic path (local_dev/outbound_delivery.rs + its test), which this PR fully wires and tests. The normal runtime approval-resume path still flows end-to-end through CapabilityOutcome + the gate:approval-{id} loop string ref via approval_interaction/gate_ref.rs — nothing reads the host_api GateRecord the loop-host seam persists for runtime gates. So the fresh-GateRef runtime record is unconsumed prep, not an unfindable-live-record bug.

  2. Correct wiring needs the flip, not a mapping hack. On a fresh raise the loop-host persist seam (persist_gate_record_for_outcome) has only the gate:approval-{id} string — the typed ApprovalRequestId is available solely via approval_resume, which is None on first raise. Deriving the canonical key there would mean string-parsing the approval-id convention into the pure ironclaw_turns→host_api mapping (a composition/product-layer convention leaking into a boundary mapper). The correct home is the flip's authorize→Resolution path, where the typed id is in hand — which is exactly why for_approval_request was introduced at authorize now (forward-looking, value-identical to the prior from_uuid), and why the runtime persist wasn't flipped (CapabilityOutcome untouched, executor Resolution consumers not touched — the PR's stated scope).

So: the write-side canonical key belongs with the flip that adds the read-side consumer; doing only the write half now (via a string-parse in the mapper) would be the misleading state. The local-dev path — which DOES have a consumer — is wired and tested here. Tracking the runtime canonical-key + consumer as the flip slice.

ilblackdragon and others added 3 commits July 19, 2026 09:14
…-visible' into HEAD

# Conflicts:
#	crates/ironclaw_loop_host/src/capability_port.rs
Merge of refactor/reborn-resolution-model-visible collapsed a
GateRef::for_approval_request call across two lines; rustfmt --check
flagged it. Formatting-only, no behavior change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ume test helpers

The Pending{Auth,Approval}Resume structs migrated to the ReplayPayloadStore
model — raw runtime input/estimate no longer ride the checkpoint (they are
persisted host-side at the fresh gate raise and reconstituted on resume,
keyed by invocation id). The structs dropped `replay`/`input`/`estimate` in
favor of `provider_replay`/`input_ref`, but three test helpers in
planned_driver.rs still set the removed fields.

This never broke base/#6273/#6275 CI because the clippy matrix is
change-scoped and those PRs did not pull ironclaw_runner into the compiled
set; this stacked change does, surfacing the latent E0560. Remove the nine
write-only stale field settings (3× replay, 3× input, 3× estimate) and the
now-unused ResourceEstimate import. No behavior change: the fields did not
exist on the structs, so no test read them back. Full runner suite passes
(344+ tests) and clippy --all-features is clean.

[skip-regression-check] compile-only fix in test helpers; the regression
coverage is the now-compiling runner test module under the all-features lane.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@ilblackdragon

Copy link
Copy Markdown
Member Author

✅ Ready for merge — host-side gate reconstitution (§5.3 flip prep / Stage 0). CI green 19/19, MERGEABLE/CLEAN.

Rebased onto its base branch (refactor/reborn-resolution-model-visible) — the branch was CONFLICTING; I merged the base in and reconciled the resume-code semantic conflict:

  1. capability_port.rs {input, estimate} reconstitution — kept this PR's design (consume the once-loaded resume_payload, which also derives effective_input_ref) over the base's lazy replay_payload_for_resume re-load, using the base's model-visible-updated None-branch body (carries the failure detail). No double-load, no orphaned local.
  2. local_dev/outbound_delivery.rs — auto-merged cleanly; this PR's gate_record_store wiring coexists with the base's changes.
  3. Formatting — one auto-merged test line rustfmt wanted collapsed.
  4. Latent ironclaw_runner E0560 (×9) — the merge pulls ironclaw_runner into CI's change-scoped clippy set for the first time, surfacing three test helpers that still constructed the replay/input/estimate fields dropped in the Pending{Auth,Approval}Resume → ReplayPayloadStore migration (raw input/estimate no longer ride the checkpoint). Verified pre-existing (reproduces on the base tip too), removed the 9 write-only stale settings + the now-unused ResourceEstimate import. All 344+ runner tests pass.

Verification: loop_host + composition + runner tests pass; the exact CI command cargo clippy --all --tests --examples --all-features -- -D warnings is clean workspace-wide locally (exit 0); all three feature lanes (default / libsql-only / all-features) clean; cargo fmt --all --check clean.

Review comment: the IronLoop finding (canonical GateRef key on the runtime persist path) is already dispositioned by the author as deferred flip work — no live consumer re-derives for_approval_request on the runtime path; the local-dev synthetic path that does is wired and tested here. Nothing outstanding.

Note for the stack collapse: this sits on refactor/reborn-resolution-model-visible; once that branch lands, retarget/rebase onto main (should be clean, since this now contains the base).

@ilblackdragon
ilblackdragon merged commit 9ec82d8 into refactor/reborn-resolution-model-visible Jul 19, 2026
23 checks passed
@ilblackdragon
ilblackdragon deleted the refactor/reborn-host-reconstitution-stage0 branch July 19, 2026 22:14
ilblackdragon added a commit that referenced this pull request Jul 19, 2026
… + denial content (§5.3 flip prep / PR-B) (#6273)

* refactor(reborn): Resolution carries model-visible failure diagnostic + denial content (§5.3 flip prep / PR-B)

Additive vocabulary slice (§3/§5.2.9/§5.3): make `host_api::Resolution`
carry the model-visible result content the executor's
`handle_capability_outcome` reads off `CapabilityOutcome` today, so a later
slice (PR-C) can flip `invoke_capability` to `Resolution` without the loop
reading host storage (its charter forbids that).

host_api additions (redacted vocabulary only — charter fit):
- `ModelInputIssue`: redacted mirror of the loop's `CapabilityInputIssue` —
  `DispatchInputIssueCode` (existing host_api enum) + bounded redacted
  `SafeSummary` path/expected/received/schema_path.
- `ModelFailureDiagnostic { InvalidInput { issues }, Diagnostic { text } }`:
  redacted mirror of `CapabilityFailureDetail`; rides
  `ToolVerdict::RecoverableFailure.diagnostic` (additive Option).
- `Denial { deny, reason_kind, summary }`: `Resolution::Denied(DenyRef)` →
  `Resolution::Denied(Denial)`, carrying the model-visible `DenyReason` +
  redacted `SafeSummary` on the channel (a projection of the sibling
  DenyRecord).

Mapping (`resolution_mapping.rs`): `failed_outcome` now redacts the loop
`detail` into the diagnostic (per-field SafeSummary; path-shaped free text
degrades to the placeholder); the Denied arm populates the Denial channel
from the same reason/summary the DenyRecord holds.

Tests (failed first, then green): structured `InvalidInput` issues and
free-text `Diagnostic` round-trip through the mapping; Denied carries
reason_kind + redacted summary; redaction proof — path- and secret-shaped
content is redacted in the `Resolution` output, never carried raw. GateRecord/
DenyRecord persistence round-trips (ironclaw_run_state) stay green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(host_api): enforce the 16-issue diagnostic cap in the type, not one producer

ModelFailureDiagnostic::InvalidInput now carries a bounded
ModelInputIssues newtype: validated at construction, revalidated on the
wire (try_from), with a truncating() producer constructor for the
mapping. A 17-item payload is rejected at every entry point — pinned by
model_input_issues_cap_is_enforced_at_construction_and_on_the_wire.

Reported-by: ironloopai (PR #6273 review)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style: cargo fmt (inherited from base pre-fmt fork)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(reborn): complete host-side gate reconstitution — resume input_ref, local-dev gate persistence, GateRef reconcile (§5.3 flip prep / Stage 0) (#6278)

* refactor(reborn): complete host-side gate reconstitution — resume input_ref, local-dev gate persistence, GateRef reconcile (§5.3 flip prep / Stage 0)

Stage 0 no-flip host-reconstitution slice: closes two verified correctness
hazards ahead of the atomic capability-result flip. The loop-facing trait and
`CapabilityOutcome` are UNCHANGED; only host-side plumbing moves.

Fix 1 — resume-time input_ref reconstitution (hazard 3). On a gate/auth resume
the effective input_ref used for the idempotency key + validation is now derived
from the host-private `ReplayPayload` persisted at the fresh gate raise (loaded
by the resume's invocation id), not the advisory loop-supplied
`approval_resume.input_ref`. A new `resume_replay_payload` helper centralizes the
derivation so `invoke_capability` (seam) and `invoke_capability_dispatch` compute
the SAME byte-stable key. The key is now derived lazily inside
`persist_gate_record_for_outcome` (after dispatch, only for a gate outcome) so a
missing/stale resume payload never pre-empts dispatch's own resume
identity/activity validation — a malformed resume still surfaces
`InvalidInvocation`, and a genuinely-missing payload still fails closed
(2a-i missing-payload test stays green). Test proves a resume derives the same
input_ref/key from the store even when the loop-supplied ref differs (a second
resume differing only in that advisory ref replays the cached outcome instead of
re-dispatching); failed first (re-dispatched, count 2) before the fix.

Fix 2 — gate-record persistence for the local-dev synthetic approval producer
(hazard 2). `OutboundDeliveryTargetSetHandler::request_approval` raises its gate
OUTSIDE the loop-host persist seam, so it now persists a `GateRecord::Approval`
itself at the raise (via the wired `GateRecordStore`), keyed by the canonical
`GateRef::for_approval_request`, so host-side gate rendering (§5.2.9) has a record
to read. Test proves the record persists and round-trips; failed first (load
returned None) before the fix.

Fix 3 — reconcile the two approval GateRef encodings. Canonicalized on the typed
`GateRef::for_approval_request(id) = GateRef::from_uuid(id.as_uuid())` — a new
`ironclaw_host_api` helper — as the GateRecordStore key. This is the host_api
uuid GateRef (record key), a DIFFERENT type from the loop-facing
`gate:approval-{id}` routing ref (`ironclaw_turns::GateRef`), so reconciling here
touches NO persisted wire format and does not break the `is_approval_gate_ref`
prefix-routing predicate. `ironclaw_capabilities::host` now uses the helper at
both authorize sites. Test proves a host-persisted approval gate ref resolves
through the read model: the routing ref recovers the approval id
(`approval_request_id_from_gate_ref`), the canonical key re-derives, and the
persisted GateRecord is found.

Note: the task suggested canonicalizing on `from_uuid`; that is exactly what this
does for the record-key GateRef. The RISKY direction (rewriting the
`gate:approval-{id}` routing string to a bare uuid) was NOT taken — it would
break the `is_approval_gate_ref` prefix predicate used in production routing,
projection, and channel delivery, plus persisted adapter command strings and
delivered-gate route records.

Tests: loop_host 378 lib (incl. new + updated resume/seam/2a-i); host_api helper
round-trip; composition 1629 (env-clean); capabilities/product_workflow/run_state
green; architecture green; integration outbound_target/auth_gate/
reopen_resume_through_gate/group_approvals green; clippy -D warnings clean on all
touched crates; pre-commit-safety clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): rustfmt the auto-merged local-dev gate-key test line

Merge of refactor/reborn-resolution-model-visible collapsed a
GateRef::for_approval_request call across two lines; rustfmt --check
flagged it. Formatting-only, no behavior change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(reborn): drop stale replay/input/estimate from planned_driver resume test helpers

The Pending{Auth,Approval}Resume structs migrated to the ReplayPayloadStore
model — raw runtime input/estimate no longer ride the checkpoint (they are
persisted host-side at the fresh gate raise and reconstituted on resume,
keyed by invocation id). The structs dropped `replay`/`input`/`estimate` in
favor of `provider_replay`/`input_ref`, but three test helpers in
planned_driver.rs still set the removed fields.

This never broke base/#6273/#6275 CI because the clippy matrix is
change-scoped and those PRs did not pull ironclaw_runner into the compiled
set; this stacked change does, surfacing the latent E0560. Remove the nine
write-only stale field settings (3× replay, 3× input, 3× estimate) and the
now-unused ResourceEstimate import. No behavior change: the fields did not
exist on the structs, so no test read them back. Full runner suite passes
(344+ tests) and clippy --all-features is clean.

[skip-regression-check] compile-only fix in test helpers; the regression
coverage is the now-compiling runner test module under the all-features lane.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): ResolutionBatch + parks() loop-suspension predicate + scripted-outcome fixture (§5.3 flip prep / Stage 1) (#6275)

* feat(reborn): ResolutionBatch + parks() loop-suspension predicate + scripted-outcome fixture (§5.3 flip prep / Stage 1)

Additive vocabulary + test-fixture slice ahead of the capability-result
flip. Nothing consumes the new items yet, so the tree stays green.

Closes a verified correctness hazard: `Resolution::is_suspension()` is
`Suspended(_)` ONLY, but the batch loops the flip rewrites must also stop
on the re-entrant gate variants (`Resolution::Blocked` — Approval/Auth/
Resource), which map to `Blocked`, not `Suspended`. A naive
`is_suspension()` batch guard would silently let a gate fall through as
if the call had completed (the #6137 bug class). This provides the
correct predicate now so the flip has one canonical, tested site to call.

- `ironclaw_host_api::Resolution::parks()`: the loop-semantics suspension
  predicate — true for every `Blocked` variant AND every `Suspended`
  variant; false for `Done` (incl. non-suspending `ChildSpawned`) and
  terminal `Denied`. `parks()` ⊋ `is_suspension()`: a `Blocked::Approval`
  parks but is not a suspension. `is_suspension()` keeps its narrow
  meaning unchanged.
- `ironclaw_host_api::ResolutionBatch { resolutions: Vec<Resolution>,
  stopped_on_suspension: bool }`: the loop-facing batch result the flip
  adopts, mirroring `CapabilityBatchOutcome`'s shape/semantics over
  `Resolution`. Wire-stable serde.
- `ironclaw_agent_loop::test_support::{resolution_from_scripted_outcome,
  resolution_batch_from_scripted}` (behind the existing `test-support`
  feature): convert existing `CapabilityOutcome` fixtures to
  `Resolution`/`ResolutionBatch` via the production mapping, for the
  flip's ~150 test sites.

Tests (written first, watched fail):
- host_api: acceptance table extended with a `parks` column across all 10
  CapabilityOutcome→Resolution rows; a dedicated exhaustive `match`
  (compile-fails if a new `Resolution` variant is added) proving
  `parks()` ⊋ `is_suspension()`; a `ResolutionBatch` round-trip.
  Verified red: a naive `is_suspension()`-based `parks()` fails
  ("parks() disagrees for blocked").
- agent_loop test-support: round-trips scripted Completed/Failed/
  ApprovalRequired → the right Resolution channel, and a batch preserving
  order + stop flag.

No trait flip, no consumer change, `CapabilityOutcome` untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(host_api): Resolution::parks uses an exhaustive match, not matches!

A new Resolution variant must be a compile error in parks() (§11.9
no-wildcard) so its parking behavior is decided deliberately rather than
silently defaulting to false — the exact class of silent fall-through
parks() exists to prevent (#6137).

Reported-by: gemini-code-assist (PR #6275 review)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runner): drop removed Pending{Approval,Auth}Resume replay fields (inherited base debt)

The replay-field removal lands in a lower stack slice; until this branch's
base picks it up, planned_driver.rs's cfg(test) fixtures still set the
removed fields and fail the workspace clippy lane. Harmless once the base
merges.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ilblackdragon added a commit that referenced this pull request Jul 20, 2026
…loadStore; wire real stores (§5.3 Stage 2a-i) (#6271)

* refactor(reborn): resume replay payload moves host-side via ReplayPayloadStore; wire real stores (§5.3 Stage 2a-i)

Moves the raw gate/auth resume replay payload out of the untrusted loop and
into the host-private `ReplayPayloadStore`, and wires the real durable stores
in production composition. The loop-facing result stays `CapabilityOutcome`
(no `Resolution` flip — that is Stage 2a-ii).

Security fix: the loop checkpoint no longer stores raw tool args, retiring the
charter violation in `crates/ironclaw_agent_loop/CLAUDE.md` ("State stores
refs... Do not store raw prompts, raw model output, tool args ... in state").

Write (two gate-raise paths): at a FRESH runtime gate raise
(`HostRuntimeLoopCapabilityPort::invoke_capability_dispatch`, gated on
`is_fresh_dispatch` before `finish_runtime_outcome`) and at the synthetic
outbound-delivery approval-gate raise (`outbound_delivery.rs::request_approval`),
the host `save`s a `ReplayPayload{input, estimate, prior_approval:None,
input_ref, correlation_id}` keyed by `InvocationId` (write-once; a benign
same-invocation duplicate is tolerated).

Read on resume: `runtime_outcome_to_loop` no longer embeds input/estimate;
`invocation_replay_input` is replaced by `ReplayPayloadStore::load` keyed by the
`InvocationId` recovered from the resume token, in both the runtime seam
(`replay_payload_for_resume`) and the synthetic decorator/handler. Fail closed
on a miss: an absent payload (including a wrong-scope read) is a sanitized
terminal `Unavailable`, never a silent empty-input dispatch. The fingerprinted
approval lease claim ordering is preserved.

Dropped loop fields (security): `input`/`estimate` from `CapabilityApprovalResume`
and `PendingApprovalResume`; the whole `CapabilityAuthResumeReplay` type and the
`replay` field on `CapabilityAuthResume`/`PendingAuthResume`. `approval_request_id`/
`correlation_id`/`input_ref`/`prior_approval` stay on the wire (they move in 2a-ii).

Composition (real stores): `capability_wiring` builds `FilesystemGateRecordStore`
(closes the #6245 production gap where it defaulted to `NoopGateRecordStore`) and
`FilesystemReplayPayloadStore` over the shared scoped filesystem and threads both
through the refreshing capability port. Adds the `/replay-payloads` per-user mount
alias. `loop_host` gains a dependency-inversion dep on `ironclaw_capabilities` for
the port trait.

Tests (all failed first): fail-closed on missing payload (loop_host seam);
extended the #6245 approval-resume seam test to assert the store persisted the
payload and the resume redispatches the correct original input; the
`ReplayPayloadStore` contract (round-trip / write-once / cross-tenant /
within-tenant) rides the base commit. Full-infra resume edges pass through the
production-shaped harness: group_approvals (runtime approval resume),
group_journeys (approval→auth convergence), outbound_target (synthetic approval
resume).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(turns): move the removed-replay design note out of prior_approval's rustdoc

Reported-by: gemini-code-assist (PR #6271 review)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runner): drop removed replay fields from planned-driver test fixtures

The Pending{Approval,Auth}Resume replay-field removal missed six
cfg(test) construction sites in planned_driver.rs (outside the PR's
verified crates; caught by the workspace all-features clippy lane).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style: cargo fmt

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(reborn): Resolution carries model-visible failure diagnostic + denial content (§5.3 flip prep / PR-B) (#6273)

* refactor(reborn): Resolution carries model-visible failure diagnostic + denial content (§5.3 flip prep / PR-B)

Additive vocabulary slice (§3/§5.2.9/§5.3): make `host_api::Resolution`
carry the model-visible result content the executor's
`handle_capability_outcome` reads off `CapabilityOutcome` today, so a later
slice (PR-C) can flip `invoke_capability` to `Resolution` without the loop
reading host storage (its charter forbids that).

host_api additions (redacted vocabulary only — charter fit):
- `ModelInputIssue`: redacted mirror of the loop's `CapabilityInputIssue` —
  `DispatchInputIssueCode` (existing host_api enum) + bounded redacted
  `SafeSummary` path/expected/received/schema_path.
- `ModelFailureDiagnostic { InvalidInput { issues }, Diagnostic { text } }`:
  redacted mirror of `CapabilityFailureDetail`; rides
  `ToolVerdict::RecoverableFailure.diagnostic` (additive Option).
- `Denial { deny, reason_kind, summary }`: `Resolution::Denied(DenyRef)` →
  `Resolution::Denied(Denial)`, carrying the model-visible `DenyReason` +
  redacted `SafeSummary` on the channel (a projection of the sibling
  DenyRecord).

Mapping (`resolution_mapping.rs`): `failed_outcome` now redacts the loop
`detail` into the diagnostic (per-field SafeSummary; path-shaped free text
degrades to the placeholder); the Denied arm populates the Denial channel
from the same reason/summary the DenyRecord holds.

Tests (failed first, then green): structured `InvalidInput` issues and
free-text `Diagnostic` round-trip through the mapping; Denied carries
reason_kind + redacted summary; redaction proof — path- and secret-shaped
content is redacted in the `Resolution` output, never carried raw. GateRecord/
DenyRecord persistence round-trips (ironclaw_run_state) stay green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(host_api): enforce the 16-issue diagnostic cap in the type, not one producer

ModelFailureDiagnostic::InvalidInput now carries a bounded
ModelInputIssues newtype: validated at construction, revalidated on the
wire (try_from), with a truncating() producer constructor for the
mapping. A 17-item payload is rejected at every entry point — pinned by
model_input_issues_cap_is_enforced_at_construction_and_on_the_wire.

Reported-by: ironloopai (PR #6273 review)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style: cargo fmt (inherited from base pre-fmt fork)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(reborn): complete host-side gate reconstitution — resume input_ref, local-dev gate persistence, GateRef reconcile (§5.3 flip prep / Stage 0) (#6278)

* refactor(reborn): complete host-side gate reconstitution — resume input_ref, local-dev gate persistence, GateRef reconcile (§5.3 flip prep / Stage 0)

Stage 0 no-flip host-reconstitution slice: closes two verified correctness
hazards ahead of the atomic capability-result flip. The loop-facing trait and
`CapabilityOutcome` are UNCHANGED; only host-side plumbing moves.

Fix 1 — resume-time input_ref reconstitution (hazard 3). On a gate/auth resume
the effective input_ref used for the idempotency key + validation is now derived
from the host-private `ReplayPayload` persisted at the fresh gate raise (loaded
by the resume's invocation id), not the advisory loop-supplied
`approval_resume.input_ref`. A new `resume_replay_payload` helper centralizes the
derivation so `invoke_capability` (seam) and `invoke_capability_dispatch` compute
the SAME byte-stable key. The key is now derived lazily inside
`persist_gate_record_for_outcome` (after dispatch, only for a gate outcome) so a
missing/stale resume payload never pre-empts dispatch's own resume
identity/activity validation — a malformed resume still surfaces
`InvalidInvocation`, and a genuinely-missing payload still fails closed
(2a-i missing-payload test stays green). Test proves a resume derives the same
input_ref/key from the store even when the loop-supplied ref differs (a second
resume differing only in that advisory ref replays the cached outcome instead of
re-dispatching); failed first (re-dispatched, count 2) before the fix.

Fix 2 — gate-record persistence for the local-dev synthetic approval producer
(hazard 2). `OutboundDeliveryTargetSetHandler::request_approval` raises its gate
OUTSIDE the loop-host persist seam, so it now persists a `GateRecord::Approval`
itself at the raise (via the wired `GateRecordStore`), keyed by the canonical
`GateRef::for_approval_request`, so host-side gate rendering (§5.2.9) has a record
to read. Test proves the record persists and round-trips; failed first (load
returned None) before the fix.

Fix 3 — reconcile the two approval GateRef encodings. Canonicalized on the typed
`GateRef::for_approval_request(id) = GateRef::from_uuid(id.as_uuid())` — a new
`ironclaw_host_api` helper — as the GateRecordStore key. This is the host_api
uuid GateRef (record key), a DIFFERENT type from the loop-facing
`gate:approval-{id}` routing ref (`ironclaw_turns::GateRef`), so reconciling here
touches NO persisted wire format and does not break the `is_approval_gate_ref`
prefix-routing predicate. `ironclaw_capabilities::host` now uses the helper at
both authorize sites. Test proves a host-persisted approval gate ref resolves
through the read model: the routing ref recovers the approval id
(`approval_request_id_from_gate_ref`), the canonical key re-derives, and the
persisted GateRecord is found.

Note: the task suggested canonicalizing on `from_uuid`; that is exactly what this
does for the record-key GateRef. The RISKY direction (rewriting the
`gate:approval-{id}` routing string to a bare uuid) was NOT taken — it would
break the `is_approval_gate_ref` prefix predicate used in production routing,
projection, and channel delivery, plus persisted adapter command strings and
delivered-gate route records.

Tests: loop_host 378 lib (incl. new + updated resume/seam/2a-i); host_api helper
round-trip; composition 1629 (env-clean); capabilities/product_workflow/run_state
green; architecture green; integration outbound_target/auth_gate/
reopen_resume_through_gate/group_approvals green; clippy -D warnings clean on all
touched crates; pre-commit-safety clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): rustfmt the auto-merged local-dev gate-key test line

Merge of refactor/reborn-resolution-model-visible collapsed a
GateRef::for_approval_request call across two lines; rustfmt --check
flagged it. Formatting-only, no behavior change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(reborn): drop stale replay/input/estimate from planned_driver resume test helpers

The Pending{Auth,Approval}Resume structs migrated to the ReplayPayloadStore
model — raw runtime input/estimate no longer ride the checkpoint (they are
persisted host-side at the fresh gate raise and reconstituted on resume,
keyed by invocation id). The structs dropped `replay`/`input`/`estimate` in
favor of `provider_replay`/`input_ref`, but three test helpers in
planned_driver.rs still set the removed fields.

This never broke base/#6273/#6275 CI because the clippy matrix is
change-scoped and those PRs did not pull ironclaw_runner into the compiled
set; this stacked change does, surfacing the latent E0560. Remove the nine
write-only stale field settings (3× replay, 3× input, 3× estimate) and the
now-unused ResourceEstimate import. No behavior change: the fields did not
exist on the structs, so no test read them back. Full runner suite passes
(344+ tests) and clippy --all-features is clean.

[skip-regression-check] compile-only fix in test helpers; the regression
coverage is the now-compiling runner test module under the all-features lane.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): ResolutionBatch + parks() loop-suspension predicate + scripted-outcome fixture (§5.3 flip prep / Stage 1) (#6275)

* feat(reborn): ResolutionBatch + parks() loop-suspension predicate + scripted-outcome fixture (§5.3 flip prep / Stage 1)

Additive vocabulary + test-fixture slice ahead of the capability-result
flip. Nothing consumes the new items yet, so the tree stays green.

Closes a verified correctness hazard: `Resolution::is_suspension()` is
`Suspended(_)` ONLY, but the batch loops the flip rewrites must also stop
on the re-entrant gate variants (`Resolution::Blocked` — Approval/Auth/
Resource), which map to `Blocked`, not `Suspended`. A naive
`is_suspension()` batch guard would silently let a gate fall through as
if the call had completed (the #6137 bug class). This provides the
correct predicate now so the flip has one canonical, tested site to call.

- `ironclaw_host_api::Resolution::parks()`: the loop-semantics suspension
  predicate — true for every `Blocked` variant AND every `Suspended`
  variant; false for `Done` (incl. non-suspending `ChildSpawned`) and
  terminal `Denied`. `parks()` ⊋ `is_suspension()`: a `Blocked::Approval`
  parks but is not a suspension. `is_suspension()` keeps its narrow
  meaning unchanged.
- `ironclaw_host_api::ResolutionBatch { resolutions: Vec<Resolution>,
  stopped_on_suspension: bool }`: the loop-facing batch result the flip
  adopts, mirroring `CapabilityBatchOutcome`'s shape/semantics over
  `Resolution`. Wire-stable serde.
- `ironclaw_agent_loop::test_support::{resolution_from_scripted_outcome,
  resolution_batch_from_scripted}` (behind the existing `test-support`
  feature): convert existing `CapabilityOutcome` fixtures to
  `Resolution`/`ResolutionBatch` via the production mapping, for the
  flip's ~150 test sites.

Tests (written first, watched fail):
- host_api: acceptance table extended with a `parks` column across all 10
  CapabilityOutcome→Resolution rows; a dedicated exhaustive `match`
  (compile-fails if a new `Resolution` variant is added) proving
  `parks()` ⊋ `is_suspension()`; a `ResolutionBatch` round-trip.
  Verified red: a naive `is_suspension()`-based `parks()` fails
  ("parks() disagrees for blocked").
- agent_loop test-support: round-trips scripted Completed/Failed/
  ApprovalRequired → the right Resolution channel, and a batch preserving
  order + stop flag.

No trait flip, no consumer change, `CapabilityOutcome` untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(host_api): Resolution::parks uses an exhaustive match, not matches!

A new Resolution variant must be a compile error in parks() (§11.9
no-wildcard) so its parking behavior is decided deliberately rather than
silently defaulting to false — the exact class of silent fall-through
parks() exists to prevent (#6137).

Reported-by: gemini-code-assist (PR #6275 review)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runner): drop removed Pending{Approval,Auth}Resume replay fields (inherited base debt)

The replay-field removal lands in a lower stack slice; until this branch's
base picks it up, planned_driver.rs's cfg(test) fixtures still set the
removed fields and fail the workspace clippy lane. Harmless once the base
merges.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): commit the dispatch reservation only after the replay-payload write (#6271 IronLoop)

`guard.commit()` ran before the fallible `persist_replay_payload_for_fresh_gate`
store write. On a transient store error the `?` returned with the dispatch
reservation still `InFlight` and the committed guard skipping its cleanup —
stranding retries/duplicates of the same key waiting forever. Commit after the
write succeeds, so a store error leaves the guard uncommitted and its drop
clears the reservation and wakes waiters.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(reborn): record the drain-then-deploy rollout for pre-#6271 gates (§5.3 Stage 2a-i)

IronLoop finding 2 flagged that gates raised/checkpointed before this change
carried `{input, estimate}` on the loop resume, whereas the new code reads them
from the host-side `ReplayPayloadStore`; a resume of a pre-deployment gate finds
no store entry and fails closed as `Unavailable`.

Author decision (PR #6271 thread): drain-then-deploy — no backward-compat
bridge. In-flight gates from before the cutover are drained before rollout, and
none exist pre-release, so a store miss in production is a genuine fault, never a
migration gap. Record that rationale at the fail-closed miss site so a future
reader debugging an `Unavailable` resume sees why there is no fallback.

Comment-only; verified `cargo check -p ironclaw_loop_host --all-features` clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant