Skip to content

feat(host_api): result-record vocabulary (GateRecord/DenyRecord + OutcomeRefs preview/child_run) (#6168) - #6237

Merged
ilblackdragon merged 3 commits into
mainfrom
feat/host-api-result-record-vocab
Jul 18, 2026
Merged

ilblackdragon merged 3 commits into
mainfrom
feat/host-api-result-record-vocab

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

What

The §5.2.9 "render from record" contract for the capability-result collapse — the model-visible content that Resolution's opaque refs point at, so the thin new channels don't drop data the loop consumes (the G1–G5 gaps the analysis surfaced).

  • gate_record.rs (new): GateRecord { Approval | Auth{credential_requirements} | Resource | DependentRun{result,byte_len} | ExternalTool } keyed by GateRef (carries resume/credential payloads + dependent-run staged result); DenyRecord { reason: DenyReason, summary } keyed by DenyRef.
  • resolution.rs: OutcomeRefs gains preview: Option<SafeSummary> + child_run: Option<RunId> (set only for the ChildSpawned verdict).

Additive + unused (like the C.1–C.7 vocabulary), pending the result-channel migration. Off main — lands independently of the #6229 review gate.

Security

Records are model-visible → they carry redacted SafeSummary/DenyReason; the loop renders FROM them and never reconstructs credential requirements from model-visible data (tool-evidence.md/safety-and-sandbox.md).

Checks

cargo test -p ironclaw_host_api (64+57) · clippy -D warnings clean · ironclaw_capabilities/ironclaw_turns build clean (the OutcomeRefs field-add) · ironclaw_architecture green.

🤖 Generated with Claude Code

…tcomeRefs preview/child_run (#6168)

The §5.2.9 "render from record" contract for the capability-result collapse: the
model-visible content that `Resolution`'s opaque refs (GateRef/DenyRef/ResultRef)
point at, so the thin result channels don't lose data the loop consumes.

- gate_record.rs (new):
  - `GateRecord { Approval | Auth{credential_requirements} | Resource |
    DependentRun{result,byte_len} | ExternalTool }` keyed by `GateRef` — carries
    the resume/credential payloads (G3) + dependent-run staged result (G2) that
    ride inline on today's `CapabilityOutcome`. `summary()`/`kind()` accessors.
  - `DenyRecord { reason: DenyReason, summary }` keyed by `DenyRef` — the
    model-visible denial content behind `Resolution::Denied`.
- resolution.rs: `OutcomeRefs` gains `preview: Option<SafeSummary>` (was
  model_observation) and `child_run: Option<RunId>` (was SpawnedChildRun.child_run_id,
  set only for the ChildSpawned verdict).

Additive + unused (like C.1–C.7) pending the result-channel migration slice. The
records are model-visible → they carry redacted SafeSummary/DenyReason; the loop
renders FROM them and never reconstructs credential requirements from
model-visible data (tool-evidence.md / safety-and-sandbox.md). Off main.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@ironloopai

ironloopai Bot commented Jul 18, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: cd6d0fdfd9d90cd6174193014fcbb148ebda9990
Result: One or more review results were superseded by a newer PR head.
Next: Run @ironloopai review on the latest PR head.
Updated: 2026-07-18T08:22:57.752Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Superseded N/A N/A 2026-07-18T08:21:17.339Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Superseded by a newer PR head. New head: 97fd82d. Previous verdict: Review declined.
Recent activity
Time Reviewer State Detail
2026-07-18T07:43:25.403Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head 555da53.
2026-07-18T07:43:25.403Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-18T07:43:26.348Z ironloop/common-reviewer (reviewer) Started Reviewer worker started.
2026-07-18T07:43:28.979Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at 6a49e07.
2026-07-18T07:44:22.445Z ironloop/common-reviewer (reviewer) Result captured Skipped; 0 blocking findings.
2026-07-18T07:44:22.445Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
2026-07-18T08:21:17.339Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (97fd82d).
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6237 July 18, 2026 07:43 Destroyed
@github-actions github-actions Bot added size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 18, 2026
@coderabbitai

coderabbitai Bot commented Jul 18, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@ilblackdragon, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 4 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7542fc7b-9266-40a1-96ea-a0a90b9735a3

📥 Commits

Reviewing files that changed from the base of the PR and between 97fd82d and cd6d0fd.

📒 Files selected for processing (1)
  • crates/ironclaw_host_api/src/gate_record.rs
📝 Walkthrough

Walkthrough

Adds public gate and denial records with redacted summaries, stable serde discriminants, and round-trip tests. Changes ChildSpawned to carry a child run ID and adds optional OutcomeRefs previews omitted from serialized output when absent.

Changes

Host API outcome contracts

Layer / File(s) Summary
Gate and denial record contracts
crates/ironclaw_host_api/src/gate_record.rs, crates/ironclaw_host_api/src/lib.rs
Defines DenyRecord and tagged GateRecord variants, exposes summary and kind accessors, re-exports the module, and tests serde preservation and snake_case tags.
Child-run verdict contract
crates/ironclaw_host_api/src/resolution.rs
Changes ChildSpawned to carry RunId, adds the child_run() accessor, and updates serialization and resolution tests.
Outcome preview metadata
crates/ironclaw_host_api/src/resolution.rs
Adds optional SafeSummary preview metadata to OutcomeRefs, omits absent values on the wire, and updates constructors and round-trip tests.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description covers the change and security intent, but it omits most required template sections, including Summary, Linked Issue, and Validation. Add the missing template sections: Summary bullets, Change Type, Linked Issue, Validation details, Security Impact, DB/Blast Radius/Rollback, and the trust-boundary checklist.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed Conventional-commit title matches the host_api record-vocabulary changes and is specific enough.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⏭️ IronLoop Review Declined: reviewer

Review at a glance

Disposition Head
⏭️ Review declined 555da533504c

Head: 555da533504ca71b29c3a32330c23dea6c3a2e4e
Reason: The required base (0ae4197) and head (555da53) are neither ancestor of the other (merge base ba31ee1). Their comparison spans 54 files and 11,635 changed lines, including removal of unrelated CLI onboarding, WebUI, composition, LLM, secrets, and Docker behavior. A reliable complete review is not possible within the configured focused-review scope.
Next: Rebase or merge the PR head onto base 0ae4197 (preserving the intended host_api change), then request review of the resulting focused comparison.

Run details

Status: Current
Trustworthy review produced: no

Summary

Skipped: the supplied base/head comparison is an oversized, mixed divergence rather than the focused host_api change described by the PR.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces the gate_record module containing DenyRecord and GateRecord to support the "render from record" result contract. Additionally, it updates OutcomeRefs in resolution.rs to include optional preview and child_run fields. The review feedback suggests adding validation to enforce that child_run is only populated when the verdict is ToolVerdict::ChildSpawned to prevent inconsistent states.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +190 to +193
/// The child run this outcome spawned (was `SpawnedChildRun.child_run_id`).
/// Set only for the `ToolVerdict::ChildSpawned` outcome; `None` otherwise.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub child_run: Option<RunId>,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The documentation states that child_run is set only for the ToolVerdict::ChildSpawned outcome and should be None otherwise. However, there is currently no validation enforcing this invariant at construction or deserialization boundaries. An inconsistent state (such as a ToolVerdict::ChildSpawned with child_run: None, or a ToolVerdict::Success with child_run: Some(...)) could be deserialized or constructed, which could lead to unexpected behavior or logic errors in the execution loop.

Consider implementing a validation step (e.g., via a custom Deserialize implementation or a validation helper) on Outcome to strictly enforce this invariant, similar to how Obligations validates its invariants.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right that the invariant was unenforced — fixed one step stronger than a validation hook: the RunId now lives ON the variant (ToolVerdict::ChildSpawned { child_run: RunId }) and OutcomeRefs.child_run is gone. Both illegal states are now unrepresentable rather than validated: a Success/RecoverableFailure has no field to carry a child ref, and a child_spawned wire tag without the run fails deserialization structurally (pinned by child_spawned_verdict_carries_the_run_on_the_variant, including the {"child_spawned": {}} rejection case). A ToolVerdict::child_run() accessor gives uniform Option<RunId> access where the old field shape was wanted. Nothing outside host_api consumed the field yet (additive vocabulary), so the wire-shape change is free.

@railway-app

railway-app Bot commented Jul 18, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6237 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 18, 2026 at 8:33 am

…ned — invariant unrepresentable, not validated

OutcomeRefs.child_run was an Option whose doc promised 'set only for
the ChildSpawned verdict' with nothing enforcing it: a Success outcome
carrying a child ref (or a ChildSpawned without one) could be built or
deserialized. Moving the RunId onto the variant makes both illegal
states impossible to express — a child_spawned wire tag without the run
fails deserialization structurally (pinned by test), and no validation
hook is needed.

Reported-by: gemini-code-assist (PR #6237 review)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6237 July 18, 2026 08:21 Destroyed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6237 July 18, 2026 08:22 Destroyed
@github-actions

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.49% (309494 / 362020 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 362020 lines now vs 320188 at floor capture (+41832 lines, +13.06%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.49% — 309494 / 362020 lines

Per-crate breakdown (65 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_runtime_policy 31.75% 80 / 252
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_observability 61.54% 16 / 26
ironclaw_authorization 62.46% 604 / 967
ironclaw_mcp 64.89% 595 / 917
ironclaw_triggers 65.44% 2142 / 3273
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_filesystem 67.69% 3932 / 5809
ironclaw_channel_host 68.65% 219 / 319
ironclaw_memory 69.2% 773 / 1117
ironclaw_reborn_migration 71.64% 1551 / 2165
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.36% 1685 / 2266
ironclaw_reborn_cli 74.57% 8880 / 11909
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.67% 958 / 1283
ironclaw_extractors 74.72% 538 / 720
ironclaw_projects 76.48% 400 / 523
ironclaw_llm 78.36% 20306 / 25915
ironclaw_product_context 78.57% 11 / 14
ironclaw_telegram_extension 80.18% 4842 / 6039
ironclaw_wasm_product_adapters 80.36% 1448 / 1802
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_first_party_extensions 81.06% 5965 / 7359
ironclaw_memory_native 81.22% 3205 / 3946
ironclaw_events 81.43% 1539 / 1890
ironclaw_secrets 82.78% 2827 / 3415
ironclaw_network 82.98% 673 / 811
ironclaw_reborn_identity 83.59% 433 / 518
ironclaw_processes 83.76% 939 / 1121
ironclaw_run_state 83.96% 424 / 505
ironclaw_reborn_config 84.02% 1830 / 2178
ironclaw_wasm 84.44% 1069 / 1266
ironclaw_auth 84.81% 3233 / 3812
ironclaw_product_workflow 84.91% 11031 / 12992
ironclaw_turns 85.04% 13722 / 16136
ironclaw_channel_delivery 85.79% 1383 / 1612
ironclaw_common 86.13% 1714 / 1990
ironclaw_threads 86.93% 4708 / 5416
ironclaw_slack_v2_adapter 87.3% 1491 / 1708
ironclaw_host_api 87.47% 3371 / 3854
ironclaw_skills 87.6% 4471 / 5104
ironclaw_hooks 87.78% 9921 / 11302
ironclaw_reborn_composition 87.97% 70201 / 79803
ironclaw_product_adapter_registry 88.06% 531 / 603
ironclaw_product_adapters 88.1% 3384 / 3841
ironclaw_reborn_traces 88.2% 11946 / 13544
ironclaw_host_runtime 88.76% 18005 / 20284
ironclaw_webui 88.9% 7652 / 8607
ironclaw_extensions 89.38% 2971 / 3324
ironclaw_reborn_openai_compat 89.5% 3778 / 4221
ironclaw_runner 89.65% 17365 / 19370
ironclaw_telegram_v2_adapter 89.7% 2717 / 3029
ironclaw_approvals 90.18% 1598 / 1772
ironclaw_conversations 90.39% 3123 / 3455
ironclaw_event_streams 90.82% 1009 / 1111
ironclaw_resources 91.65% 4476 / 4884
ironclaw_loop_host 92.25% 15051 / 16316
ironclaw_attachments 93.06% 630 / 677
ironclaw_agent_loop 94.88% 9184 / 9680
ironclaw_safety 95.04% 3677 / 3869
ironclaw_outbound 95.52% 3451 / 3613
ironclaw_first_party_extension_ports 95.62% 3672 / 3840

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@ilblackdragon

Copy link
Copy Markdown
Member Author

✅ Ready for merge

Reviewed, review comment addressed, CI fully green — 59 pass / 0 fail.

  • Lands the §5.2.9 "render from record" vocabulary: GateRecord (Approval / Auth+credential-requirements / Resource / DependentRun+staged-result / ExternalTool) and DenyRecord keyed by their opaque refs, plus OutcomeRefs.preview. Redaction invariants hold — every record carries a SafeSummary, the auth gate's credential requirements stay host-owned (rendered FROM the record, never reconstructed from model-visible data), and wire round-trips are pinned per variant.
  • Gemini's child_run⇔ChildSpawned invariant finding — fixed one step stronger than the suggested validation (97fd82d): the RunId moved onto the variant (ToolVerdict::ChildSpawned { child_run }) and OutcomeRefs.child_run is deleted. Both illegal states are now unrepresentable rather than validated; a {"child_spawned": {}} wire payload fails deserialization structurally (pinned by test). Nothing consumed the field yet, so the wire-shape change was free.

🤖 Generated with Claude Code

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6237 — cd6d0fdf Deployed Jul 18, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant