Skip to content

fix(host-runtime): classify HTTP error responses as failures - #7342

Merged
serrrfirat merged 13 commits into
mainfrom
codex/fix-http-error-semantics
Aug 10, 2026
Merged

serrrfirat merged 13 commits into
mainfrom
codex/fix-http-error-semantics

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Aug 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • classify builtin.http and builtin.http.save HTTP 4xx/5xx responses as recoverable OperationFailed capability outcomes
  • preserve the bounded, sanitized HTTP response as model-visible diagnostic context so the model can inspect status, body, and recovery hints
  • keep informational, successful, and redirect responses as inspectable successful results
  • add caller-level regression coverage and wire it into the documented architecture-runtime gate

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

None. Reproduced from an exported production thread where an HTTP 403 was reported to the model as a successful tool result.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings
  • cargo build
  • Relevant tests pass: all 50 builtin_http_* tests and the complete architecture-runtime gate
  • cargo test -p <owning-crate> --features integration if database-backed or runtime-integration behavior changed (the root integration feature is empty — the flag is per-crate)
  • Manual testing: Not applicable; deterministic host-runtime tests cover the response classification
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review

Additional scoped validation:

  • cargo clippy -p ironclaw_host_runtime --all-targets --all-features -- -D warnings
  • git diff --check
  • bash -n scripts/reborn-e2e-rust.sh

Test Strategy

User behavior:

Given an authorized builtin.http call whose server returns HTTP 403, when the response reaches the capability boundary, then the model receives a recoverable failed tool outcome with the sanitized status, body, and authentication hint instead of a successful tool result. HTTP redirects remain inspectable results.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: added caller-path host-runtime tests for HTTP 403 failure classification and HTTP 302 compatibility
  • Reborn integration: Not applicable; the behavior is deterministic at the public HostRuntime::invoke_capability seam and does not require model orchestration
  • Recorded fixture: Not applicable; tool selection and argument shape are unchanged
  • Browser E2E: Not applicable; no browser-visible behavior changed
  • Backend or runtime: ran scripts/reborn-e2e-rust.sh architecture-runtime
  • Live canary: Not applicable; no live provider is required to verify status classification

What the tests prove:

  • HTTP 403 produces FailureKind::OperationFailed
  • the sanitized response body, status, and authentication recovery hint remain available as diagnostic context
  • redirects remain successful inspectable results
  • existing HTTP limits, redaction, credential mediation, saved-body behavior, and capability dispatch contracts remain green

Commands run:

  • cargo fmt --all -- --check
  • cargo test -p ironclaw_host_runtime --test first_party_builtin_tools builtin_http_
  • cargo clippy -p ironclaw_host_runtime --all-targets --all-features -- -D warnings
  • scripts/reborn-e2e-rust.sh architecture-runtime
  • git diff --check
  • bash -n scripts/reborn-e2e-rust.sh

Security Impact

Changes model-visible network tool execution semantics but does not change permissions, allowlists, credential injection, redaction, or transport policy. HTTP error bodies continue through the existing bounded, sanitized diagnostic channel and are re-scrubbed and injection-fenced before model observation.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: no new public or trust-bearing types; the existing OperationFailed and diagnostic contracts are reused
  • Untrusted content enters prompts only through an envelope/escaping primitive: HTTP output remains sanitized and crosses the existing untrusted diagnostic scrub/fence path
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check: Not applicable; no hashes changed
  • New/changed status, exit, policy, runtime, or error variants: downstream match sites audited. Command/output: no new variant; scripts/reborn-e2e-rust.sh architecture-runtime passed using the existing OperationFailed mapping
  • Security/durability serde(default) fields fail closed or have migration tests: Not applicable; no serialized fields changed
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic: no new collection; diagnostic output retains existing response and model-observation bounds
  • Driver/operator-visible errors have stable class semantics (Transient, Permanent, Misconfigured, PolicyDenied or equivalent): HTTP 4xx/5xx consistently use existing OperationFailed
  • Sandbox/native/host names accurately describe trust boundary: Not applicable; no trust-boundary naming changed

Database Impact

None.

Blast Radius

builtin.http and builtin.http.save response classification only. Clients that intentionally inspect 4xx/5xx responses now receive a recoverable failed capability outcome rather than success, but retain the sanitized response as diagnostic context. 1xx, 2xx, and 3xx behavior is unchanged.

Rollback Plan

Revert this commit to restore the previous behavior where every received HTTP response was returned as a successful capability result. No data migration or persisted schema rollback is required.

Review Follow-Through

Reviewer judgment requested on the deliberate compatibility change for expected 4xx statuses such as 404 or 409. The model retains the full bounded diagnostic and can continue, but the verdict is now correctly non-successful.


Review track: C (runtime error semantics)

@railway-app

railway-app Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-7342 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 10, 2026 at 10:46 am

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 13:11 Destroyed
@github-actions github-actions Bot added the scope: docs Documentation label Aug 7, 2026
@serrrfirat serrrfirat added size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules labels Aug 7, 2026
@github-actions github-actions Bot added the contributor: core 20+ merged PRs label Aug 7, 2026
@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes

    • HTTP 4xx and 5xx responses now report failed operations instead of appearing successful.
    • Failure diagnostics are sanitized, valid JSON, size-limited, and retain relevant status details.
    • Redirects remain visible without being automatically followed.
    • Save-mode requests preserve response metadata when requests fail.
    • Failure outcomes include wall-clock usage; other status responses remain successful.
  • Documentation

    • Updated HTTP behavior, retry guidance, rate-limit handling, and authentication failure guidance.
  • Tests

    • Added coverage for status boundaries, redirects, saved responses, and bounded diagnostics.

Walkthrough

First-party HTTP tools classify 4xx and 5xx responses as OperationFailed outcomes. Diagnostics remain valid, sanitized, and bounded. Save mode preserves saved-body metadata. Redirects remain successful and are not followed.

Changes

HTTP status classification

Layer / File(s) Summary
Status classification and bounded diagnostics
crates/kernel/ironclaw_host_runtime/src/first_party_tools/http.rs, crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs
Dispatch records wall-clock time, applies diagnostic limits, and passes shaped responses to classify_status. Error responses produce bounded JSON diagnostics with status, hints, truncation metadata, and resource usage.
Runtime regression coverage
crates/kernel/ironclaw_host_runtime/tests/first_party_builtin_tools.rs, crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs, scripts/reborn-e2e-rust.sh
Tests cover error statuses, boundaries, redirects, saved-body metadata, valid JSON, binary bodies, control characters, fallback output, usage accounting, and diagnostic limits.
Contract and integration validation
docs/reborn/contracts/host-runtime.md, tests/integration/http_matcher.rs, tests/integration/CLAUDE.md, tests/integration/support/http_matcher.rs
The contract, integration tests, and test documentation describe failed HTTP outcomes, redirect handling, save-mode metadata, bounded diagnostics, and retry behavior.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant HTTPDispatch
  participant classify_status
  participant FirstPartyCapabilityError
  HTTPDispatch->>classify_status: shaped response, status, and wall-clock time
  classify_status->>FirstPartyCapabilityError: 4xx/5xx bounded diagnostic
  classify_status-->>HTTPDispatch: success or OperationFailed result
Loading

Possibly related PRs

  • nearai/ironclaw#7330: Implements overlapping HTTP status classification, bounded diagnostics, save-mode behavior, and related tests and documentation.

Suggested reviewers: benkurrek

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title follows Conventional Commits style and accurately describes the HTTP error classification change.
Description check ✅ Passed The description covers all required sections, explains the runtime behavior, documents validation, and identifies security, database, blast-radius, and rollback impacts.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ironloopai

ironloopai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

🧭 IronLoop Run · Review

This comment updates in place as the Run moves through its stages.

🟩 Final result · Completed

🟨 Queued → 🟦 Working → 🟦 Posting results → 🟩 Completed

Automatic trigger · attempt 1 of 3 · completed in 7m 36s

IronLoop completed the review and posted it to GitHub.

🔗 Result

Open submitted review →

Run details

Run: bf109816-7a5f-498d-b8b6-220736c01264
Base: main at ce2d6f8
Head: codex/fix-http-error-semantics at 78b95ad
Created: 2026-08-07 13:12 UTC
Updated: 2026-08-07 13:20 UTC

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs`:
- Around line 139-170: Update the trimming loop around the inline body branches
so body_text or body_base64 is selected only when its current content can still
shrink; once it reaches zero length, continue to the headers branch instead of
repeatedly selecting the empty body. Preserve the existing truncation
bookkeeping, then add a caller-level regression test covering an oversized
response with a body that trims to empty and headers exceeding
MODEL_DIAGNOSTIC_MAX_BYTES, verifying headers and other diagnostic fields are
retained rather than returning fallback_diagnostic.

In `@docs/reborn/contracts/host-runtime.md`:
- Around line 110-118: Update the save-mode body contract in the surrounding
documentation to state that the original sanitized body is retained only when
the initial call uses builtin.http.save, and only up to the existing save-mode
limit. Clarify that the failure diagnostic replaces only body content with
saved_body metadata; preserve status, auth_hint, and truncation metadata in the
diagnostic, and remove the implication that re-issuing the request is the normal
inspection mechanism.
- Around line 119-121: Update the retry guidance in the ToolRecoveryObservation
contract paragraph to document that 429/503 failures may populate retry_after_ms
with provider-delay metadata. Preserve the existing backoff requirement and
explicitly retain that retry_after_ms being None does not permit an immediate
retry.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4f9d1fe-450a-4bd2-84ee-3a4381565230

📥 Commits

Reviewing files that changed from the base of the PR and between ce2d6f8 and 78b95ad.

📒 Files selected for processing (5)
  • crates/kernel/ironclaw_host_runtime/src/first_party_tools/http.rs
  • crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs
  • crates/kernel/ironclaw_host_runtime/tests/first_party_builtin_tools.rs
  • docs/reborn/contracts/host-runtime.md
  • scripts/reborn-e2e-rust.sh

Comment thread crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs Outdated
Comment thread docs/reborn/contracts/host-runtime.md Outdated
Comment thread docs/reborn/contracts/host-runtime.md Outdated

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 IronLoop review

Found 1 medium-severity model-visible diagnostic truncation issue.

Findings: 🟠 Medium 1

🟠 Medium · Reserve space for the downstream safety envelope

Inline on crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs:120. See the inline comment for details.

Validation

  • ✅ Focused HTTP integration tests — `cargo test -p ironclaw_host_runtime --test first_party_builtin_tools builtin_http_` passed: 50 tests.
  • ✅ HTTP output unit tests — `cargo test -p ironclaw_host_runtime --lib first_party_tools::http_output::tests` passed: 5 tests.
  • ✅ Static checks — `cargo fmt --all -- --check`, `git diff --check`, and `bash -n scripts/reborn-e2e-rust.sh` passed.
Review details
  • Run: bf109816-7a5f-498d-b8b6-220736c01264
  • Workflow: Review
  • Attempts: 1

return fallback_diagnostic(status);
};
let mut output = output.clone();
if serialized_output_len(&output) <= MODEL_DIAGNOSTIC_MAX_BYTES {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 IronLoop review · Inline finding

🟠 Medium · Reserve space for the downstream safety envelope

The 4 KiB cap is applied before the diagnostic enters the loop’s model-visible scrubber. If an error body contains an injection pattern, that scrubber wraps the entire JSON diagnostic in a several-hundred-byte external-content notice, and the resolution boundary then truncates it back to 4 KiB. For a large 4xx/5xx body this cuts the wrapped JSON before its tail fields (`status` and `truncation` are serialized near the end), so the model does not receive the complete valid JSON/context this change promises. Cover an injection-bearing, near-limit HTTP error through the loop resolution boundary and trim with enough headroom (or preserve the structured fields after fencing).

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Railway preview QA — PASS

Tested head SHA: 78b95addf44a22319bf03d370a2fa9959359f3ca
Railway state: success (bound to head; asset assets/app-BkDfgLHI.js)
Preview URL: https://ironclaw-ironclaw-pr-7342.up.railway.app
Route: /chat (Gateway V2 console, token login, browser-driven via local CDP harness)
PR claim: builtin.http / builtin.http.save classify HTTP 4xx/5xx as recoverable OperationFailed capability outcomes with the bounded sanitized response (status, body, recovery hint) as model-visible diagnostic; 1xx/2xx/3xx remain successful inspectable results.

Given/When/Then matrix

# Classification Intended contract Actual contract exercised Observed result Status
1 Required builtin.http GET returning HTTP 403 surfaces to the model as a failed tool outcome (OperationFailed) carrying sanitized status, body, and authentication hint builtin.http GET https://postman-echo.com/status/403 (httpbin.org/httpstat.us are transport-blocked from the preview; postman-echo.com/status/403 is a deterministic 403 endpoint reachable from Railway) status: "error", summary: "Capability failed with operation_failed.", failure_kind: "operation_failed", detail.status: 403, detail.body_text: "{\"status\":403}", detail.auth_hint: "This request was rejected for authentication/authorization...", recovery: {recovery_hint: "respect_failure_constraint", same_call_retry: "allowed"} PASS
2 Required HTTP 3xx redirect responses remain successful inspectable results (classifier draws the boundary at 400) builtin.http GET https://postman-echo.com/status/302 (deterministic 302, host does not follow redirects) status: "success", detail.status: 302, detail.body_text: "{\"status\":302}" — redirect returned inline as an inspectable successful result PASS
3 Required builtin.http.save 4xx response also classified as failed outcome while saved-body behavior is preserved builtin.http.save GET https://postman-echo.com/status/403 status: "error", failure_kind: "operation_failed", summary "Capability failed with operation_failed.", recovery respect_failure_constraint; body still saved — 14 bytes {"status":403} written to /workspace/echo-403-response.txt PASS
4 Supplemental Preview deployment health + egress builtin.http GET https://example.com; gateway state map HTTP 200, full inspectable successful result (status: 200, headers, body preview); gateway model + 12 extensions live PASS

Status derivation

  • Required cases passed: 3 (403 classification, 302 compatibility, http.save 403 + saved body)
  • Required cases failed: 0
  • Required cases blocked / not executed: 0
  • Overall: PASS — every required case executed against the intended contract and matched the PR's acceptance claim.

The deployed build demonstrably contains the PR change: old code reported every received HTTP response as a successful result; the preview reports received 403 responses as operation_failed with the new diagnostic fields (detail.status, detail.body_text, detail.auth_hint, respect_failure_constraint recovery) — exactly the PR's regression claim. Tool identity confirmed by the unique builtin.http output contract shapes (status/body_text/headers detail with result_reference preview, operation_failed failure kind, saved-body side effect); the chat UI's Activity surface shows one tool call per turn but does not render the tool name.

Environment notes (not PR behavior)

  • httpbin.org and httpstat.us are unreachable at the transport layer from the Railway preview (failure_kind: "network", no HTTP response received). Initial 403/302 attempts against httpbin.org therefore never exercised the classification path and were discarded; all required cases were re-run against postman-echo.com/status/*, which is reachable and returns the exact status.
  • example.com (Cloudflare) reachable; the preview's egress itself is healthy.

Cleanup status

  • Test data: one 14-byte file /workspace/echo-403-response.txt created inside the preview's ephemeral per-deployment workspace (empty workspace, no shared state; deployment-scoped, cleaned with the environment).
  • Local CDP harness profile (browser session storage) removed after evidence capture; no credentials in this comment.

Skipped cases

  • Streaming/upload/responsive/auth recipes skipped: not relevant to a backend response-classification change; login was exercised once to reach /chat.
  • Regression case "4xx reported as success" (pre-fix behavior) intentionally not re-verified — the preview runs the fixed head only.

…cover header/base64 branches

- share one budget-trim engine (fit_output_to_budget) between the success
  path and the failure diagnostic; the diagnostic trim now converges in a
  strict-progress loop instead of an 8-iteration cap, and the headers
  branch stays reachable when the inline body is empty or absent, so
  header-heavy 4xx/5xx diagnostics keep status/auth_hint/truncation
  instead of collapsing to the fallback verdict
- scrub DEL/C1 control bytes (U+007F..U+009F) from the serialized
  diagnostic: serde_json does not escape them, and ModelDiagnostic
  validation would otherwise replace the whole diagnostic with the fixed
  fallback sentence at the resolution boundary
- move the shaped output by value instead of cloning it per error
  response
- add unit coverage for the empty-body/header trim, base64 alignment on
  the failure path, and control-char scrubbing; wire the redirect
  regression test into the architecture-runtime gate
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 17:09 Destroyed
…ck, pin envelope size

- shape 4xx/5xx responses at the 4 KiB diagnostic budget instead of the
  success inline limit so the discarded success-budget trim pass (and its
  serializations) no longer runs on the failure path
- attach dispatch wall_clock_ms to the OperationFailed usage, matching
  the sibling first-party dispatches' failure-path accounting
- pin the truncation envelope below its reserved budget with a unit test
- correct the fit_output_to_budget doc to describe the incremental
  re-measure loop rather than a fixed three-serialization bound
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 17:39 Destroyed
…lback shape

- single is_error_status predicate shared by the dispatch shape-limit
  selection and classify_status, so the 400..=599 boundary cannot drift
- pin failure usage accounting: classify_status unit test asserts egress
  bytes + wall_clock_ms on the OperationFailed outcome; the 403
  integration test asserts egress bytes reach the governor for failed
  calls
- pin the fallback diagnostic payload shape for non-object output and
  note why the serde-failure branch is unreachable by construction
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 18:13 Destroyed
…ed-outcome contract

- reborn_integration_http_matcher asserted the pre-change contract that a
  scripted HTTP 500 surfaces as a successful tool result; the new
  classification makes it a recoverable OperationFailed outcome, so the
  test now asserts ToolErrorClass::Failed with the operation_failed kind
  (run still completes; docstring updated to the contract doc)
- pin the post-fit fallback safety valve with an oversized untrimmable
  key test; correct the governor-accounting comment; dedupe the 400
  boundary rationale onto is_error_status
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 18:40 Destroyed
…cument fence interaction

- tests/integration/CLAUDE.md .with_status entry now states 4xx/5xx
  classify as a Failed tool outcome (operation_failed) with sanitized
  diagnostic context; other statuses remain Completed results
- host-runtime contract doc records the loop-host injection-fence
  interaction: verdict semantics never depend on the fenced diagnostic
  surviving the observation bound (OperationFailed + safe summary always
  reach the model)
- document the fit_output_to_budget convergence bound (<= 3 passes)
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 18:59 Destroyed
serde_json serializes every Value string without revalidating UTF-8
(probed: even an unsafe lone-surrogate string serializes Ok), so the
serde-failure arm is a pure defensive guard for future Value shapes, not
a reachable lone-surrogate path. Correct the doc comment and the
fallback test note to state the empirical fact.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 19:24 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs (2)

617-662: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Assert decoded byte metadata for base64 trimming.

This test checks four-character alignment but not truncation.bytes_returned. That field is byte-based. Assert that it equals the decoded length of trimmed_base64; otherwise an encoded-character count can pass this test.

Proposed assertion
         assert_eq!(
             trimmed_base64.len() % 4,
             0,
             "trimmed base64 must stay multiple-of-4 aligned"
         );
+        let decoded_bytes = BASE64_STANDARD
+            .decode(trimmed_base64)
+            .expect("trimmed base64 must decode")
+            .len();
+        assert_eq!(
+            parsed["truncation"]["bytes_returned"],
+            json!(decoded_bytes),
+            "truncation bytes must count decoded body bytes"
+        );
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs`
around lines 617 - 662, Update
failure_diagnostic_trims_base64_body_with_large_headers to assert that
truncation.bytes_returned equals the decoded byte length of trimmed_base64,
using the existing base64 decoding support. Keep the current four-character
alignment and JSON validity assertions unchanged.

699-719: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Use a non-zero request size in the usage regression.

response_with_status(403) sets request_bytes to 0. The assertion can pass if failure handling always reports zero egress. Set a non-zero request size and assert that usage.network_egress_bytes preserves it.

Proposed test adjustment
-        let shaped = shape_response(response_with_status(403), 48 * 1024);
+        let mut response = response_with_status(403);
+        response.request_bytes = 17;
+        let shaped = shape_response(response, 48 * 1024);
...
-            usage.network_egress_bytes, 0,
-            "empty GET request carries no egress bytes"
+            usage.network_egress_bytes, 17,
+            "failure must preserve request egress bytes"
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs`
around lines 699 - 719, Update
classify_status_failure_carries_egress_and_wall_clock_usage to use a non-zero
request size when creating the 403 response, then assert
usage.network_egress_bytes equals that request size while retaining the
wall-clock assertion and successful 2xx check.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs`:
- Around line 685-690: Rename
failure_diagnostic_falls_back_on_unserializable_output to reflect that it tests
bounded_failure_diagnostic’s non-object fallback for serde_json::Value::Null,
not an unserializable output or serializer-error path.

---

Outside diff comments:
In `@crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs`:
- Around line 617-662: Update
failure_diagnostic_trims_base64_body_with_large_headers to assert that
truncation.bytes_returned equals the decoded byte length of trimmed_base64,
using the existing base64 decoding support. Keep the current four-character
alignment and JSON validity assertions unchanged.
- Around line 699-719: Update
classify_status_failure_carries_egress_and_wall_clock_usage to use a non-zero
request size when creating the 403 response, then assert
usage.network_egress_bytes equals that request size while retaining the
wall-clock assertion and successful 2xx check.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 1d95ac69-89fe-4b9b-90fa-23f6c4b53b02

📥 Commits

Reviewing files that changed from the base of the PR and between 2a98ada and 1a53e3d.

📒 Files selected for processing (1)
  • crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs

Comment thread crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs Outdated
…ostic envelope

The re-inserted truncation envelope carried only the diagnostic-budget
trim state, so a 4xx/5xx response whose shape stage had already marked
headers or body as truncated (e.g. more than 32 headers) could end up
with headers_truncated:true beside an envelope claiming headers:false.
OR the surviving keys into the envelope and pin with a regression test;
boundary doc now states the full complement (outside 400..=599 stays
inspectable, including out-of-spec 600+).
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 19:46 Destroyed
@github-actions github-actions Bot added size: XL 500+ changed lines and removed size: L 200-499 changed lines labels Aug 7, 2026
…t, fix trim rationale

- tests/integration/support/http_matcher.rs module doc still claimed
  .with_status non-2xx stays Completed; now documents 4xx/5xx as a
  model-visible Failed operation_failed outcome
- bounded_failure_diagnostic doc: head-keeping truncation cuts only the
  last-sorted keys (status, truncation envelope); auth_hint sorts first
  and survives
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 7, 2026 19:57 Destroyed

@serrrfirat serrrfirat left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Reclassify builtin.http and builtin.http.save HTTP 4xx/5xx responses from successful results into recoverable OperationFailed outcomes carrying sanitized diagnostic context.

Shape: normal primary mode, no modifiers (5-file PR reviewed against exact local diff; 8-file surface by final round after doc/test sync).

Coverage: complete — local-git diff source, 8 files (2 production, 3 tests, 2 docs, 1 CI), no packetization, 0 failed reviewers, 0 limitations.

Stats: 0 findings at head efa76f7ef0 (18 findings across 9 review-fix rounds, all fixed and re-verified). Reviewers run: security, bugs, performance, tests, conventions, local-patterns, maintainability, approach. Reviewers failed: none. Body-only: 0.

Fix-loop history (all verified fixed at head)

  • Round 1 (8 findings): headers-trim branch unreachable with empty body_text → shared fit_output_to_budget with independent branches; DEL/C1 control chars surviving serde escaping → scrub_model_diagnostic_controls on both diagnostic paths; up-to-8x re-serialization → convergent strict-progress loop; per-error Map clone → move by value; untested base64/headers trim branches → 2 new unit tests; duplicated safety comment → removed; duplicate trim engine → single shared pass; missing gate wiring → redirect test added to architecture-runtime gate.
  • Round 2 (4): envelope budget unpinned → truncation_envelope_fits_its_reserved_budget; discarded success-budget trim on error path → shape 4xx/5xx at MODEL_DIAGNOSTIC_MAX_BYTES; missing wall_clock on failure usage → set_wall_clock_ms in classify_status; stale doc bound → rewritten.
  • Round 3 (3): failure usage never asserted → unit pin + governor egress assertion; fallback payload unpinned → test; duplicated 400..=599 predicate → single is_error_status.
  • Round 4 (1 HIGH + 3): stale reborn_integration_http_matcher 5xx-as-success contract test (broke on this head) → migrated to Failed/operation_failed contract; misleading governor comment → corrected; post-fit fallback branch → test; duplicated boundary doc → deduped.
  • Round 5 (3): stale tests/integration/CLAUDE.md .with_status doc → synced; injection-fence overflow interaction → documented in contract (verdict independent of fenced diagnostic); loop pass-count concern → convergence bound documented (≤3 passes).
  • Round 6 (1): serde-failure arm "untested" → probed empirically (serde_json serializes every Value string without revalidating UTF-8; arm is a pure defensive guard, no in-process test can exercise it without UB) → doc corrected.
  • Round 7 (2): envelope dropped shape-stage truncation flags → OR surviving keys, pinned by 33-header regression test; boundary doc omitted 600+ tail → complement stated.
  • Round 8 (2): stale support-module doc → synced; trim-rationale doc overstated cut list → corrected (auth_hint sorts first and survives).
  • Round 9: zero findings from all 8 reviewers.

Validation

  • cargo test -p ironclaw_host_runtime --lib first_party_tools::http_output — 13 passed
  • cargo test -p ironclaw_host_runtime --test first_party_builtin_tools builtin_http_ — 50 passed
  • cargo test -p ironclaw_integration_tests --test reborn_integration_http_matcher — 22 passed (incl. migrated 5xx test)
  • cargo clippy -p ironclaw_host_runtime --all-targets --all-features -- -D warnings — clean
  • cargo fmt --all -- --check — clean

Remaining risks (documented, accepted)

  • Injection-fenced error diagnostics may be tail-truncated at the loop-host observation bound; verdict + safe summary are independent (documented in docs/reborn/contracts/host-runtime.md).
  • serialize_diagnostic serde-failure arm is an untestable defensive guard (serde_json cannot fail on Value serialization; probed).
  • Save-mode failure writes the sanitized body before the verdict; retry may duplicate the write (documented).

@think-in-universe

Copy link
Copy Markdown
Collaborator

@ironloopai review

@ironloopai

ironloopai Bot commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

🧭 IronLoop Run · Review

This comment updates in place as the Run moves through its stages.

🟩 Final result · Completed

🟨 Queued → 🟦 Working → 🟦 Posting results → 🟩 Completed

Manual command by think-in-universe · attempt 1 of 3 · completed in 20m 1s

IronLoop completed the review and posted it to GitHub.

🔗 Result

Open submitted review →

Run details

Run: ce15715d-1554-431d-9649-ec5ca84b389f
Base: main at 9faa6cd
Head: codex/fix-http-error-semantics at efa76f7
Created: 2026-08-10 09:41 UTC
Updated: 2026-08-10 10:01 UTC

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 IronLoop review

Found one medium and one low correctness issue in failed-response diagnostic handling.

Findings: 🟠 Medium 1 · 🟡 Low 1

🟠 Medium · Preserve the original egress cap when shaping failed responses

Inline on crates/kernel/ironclaw_host_runtime/src/first_party_tools/http.rs:262. See the inline comment for details.

🟡 Low · Retain saved-body evidence when the diagnostic falls back

Inline on crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs:142. See the inline comment for details.

Validation

  • ✅ Diff hygiene — No whitespace errors were reported for the pull request diff.
  • ⚪ Focused test execution — Not run. Not run; static tracing of the changed classifier through the partial-response egress and body-store paths was sufficient to establish the findings.
Review details
  • Run: ce15715d-1554-431d-9649-ec5ca84b389f
  • Workflow: Review
  • Attempts: 1

Comment on lines +257 to +262
let shape_limit = if is_error_status(status) {
u64::try_from(MODEL_DIAGNOSTIC_MAX_BYTES).unwrap_or(response_body_limit)
} else {
response_body_limit
};
let shaped = shape_response(response, shape_limit);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 IronLoop review · Inline finding

🟠 Medium · Preserve the original egress cap when shaping failed responses

This replaces the caller’s cap with 4 KiB before `shape_response`, which uses that argument to determine whether the egress already truncated the body. The model-visible egress returns a partial response when a body crosses the configured cap: a 403 with `response_body_limit: 1` and a longer body has one returned byte and `response_bytes = 2`, but is shaped against 4096 and loses `body_truncated` and the truncation envelope. The model consequently sees a partial error body as complete. Keep the original egress cap for truncation accounting and apply the diagnostic display budget separately.

Comment on lines +140 to +142
let trim = fit_output_to_budget(&mut output, final_budget);
if serialized_output_len(&output) > final_budget {
return fallback_diagnostic(status);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 IronLoop review · Inline finding

🟡 Low · Retain saved-body evidence when the diagnostic falls back

`save_to` accepts inputs up to the general 1 MiB tool-input limit and `ScopedPath` has no path-length bound. A long accepted save target can make `saved_body.path` push a 4xx/5xx diagnostic over 4 KiB; the trim routine only reduces bodies and headers, so this branch replaces the entire detail with the fallback even though `builtin.http.save` already wrote the response. The failed outcome then omits both the destination and `bytes_written`, undermining the new guidance to inspect the saved body before retrying. Bound the save path or retain compact saved-body metadata in the fallback.

…p, fence headroom, saved-body fallback (#7342)

- shape failed responses at the caller's response_body_limit again so
  egress-truncation accounting (body_was_truncated_by_egress) stays
  correct; the diagnostic display budget is applied separately. Pinned by
  builtin_http_error_diagnostic_preserves_egress_truncation_flag.
- reserve MODEL_DIAGNOSTIC_FENCE_HEADROOM_BYTES so a diagnostic wrapped
  in the loop-host external-content fence still fits the observation
  budget; pinned by failure_diagnostic_stays_within_budget_when_fenced.
- retain compact saved_body evidence (bounded path prefix +
  bytes_written) in the fallback verdict instead of dropping the save
  destination; pinned by failure_diagnostic_fallback_retains_saved_body_evidence.
- rename failure_diagnostic_falls_back_on_unserializable_output to the
  non-object contract it actually exercises (Value::Null serializes).
- host-runtime contract: correct the save-mode body-retention wording and
  document retry_after_ms semantics (None does not permit immediate retry).
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7342 August 10, 2026 10:37 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Addressed review feedback at head dc2f65177d:

Comment Reviewer Resolution
http_output.rs — continue header trimming after body reaches zero CodeRabbit Already fixed in fit_output_to_budget (!body_text.is_empty() / !body_base64.is_empty() guards + independent headers pass) — no change needed
host-runtime.md — correct the save-mode body contract CodeRabbit Fixed — doc now states the original body from a network-only failed call is not retained; only the initial builtin.http.save call retains it (save-mode limit applies); status/auth_hint/truncation metadata remain in the diagnostic
host-runtime.md — document retry metadata on 429/503 CodeRabbit Fixed — doc now covers retry_after_ms: populated when the host holds provider delay metadata; builtin.http does not parse Retry-After so it stays None; None does not permit immediate retry
http_output.rs — rename test to match non-object fallback CodeRabbit Fixed — failure_diagnostic_falls_back_on_unserializable_output → failure_diagnostic_non_object_output_falls_back_to_verdict (Value::Null is serializable; the test exercises the non-object guard)
http_output.rs:132 — reserve space for the safety envelope IronLoop Fixed — new MODEL_DIAGNOSTIC_FENCE_HEADROOM_BYTES (768) subtracted from the trim budget; regression test wraps the diagnostic with the real wrap_external_content fence and asserts it stays ≤ 4096
http.rs:262 — preserve the original egress cap when shaping failed responses IronLoop Fixed — shape_response again receives the caller's response_body_limit, so a body the egress truncated at the caller's cap keeps body_truncated + truncation envelope; diagnostic display budget applied separately. Regression test: builtin_http_error_diagnostic_preserves_egress_truncation_flag
http_output.rs:142 — retain saved-body evidence when the diagnostic falls back IronLoop Fixed — fallback verdict now carries compact saved_body (path bounded to 1024 chars + bytes_written); regression test with an 8 KiB save path

Validation at head: 15 http_output unit + 51 builtin_http_* integration + 22 http_matcher integration tests pass; clippy -D warnings and fmt clean.

Note: the head moved after the prior approval, so re-review is welcome (bots re-run automatically on push).

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs`:
- Around line 140-145: Update bounded_failure_diagnostic so its early return
only applies when serialized_output_len(&output) is at most
MODEL_DIAGNOSTIC_MAX_BYTES minus MODEL_DIAGNOSTIC_FENCE_HEADROOM_BYTES; larger
diagnostics must continue through fit_output_to_budget. Add a caller-level
regression at the nearest observation seam that triggers fencing for a
diagnostic in this range and verifies the final observation is bounded and
complete.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: d544b7c1-b2f3-4f77-9383-d9072bc86069

📥 Commits

Reviewing files that changed from the base of the PR and between efa76f7 and dc2f651.

📒 Files selected for processing (4)
  • crates/kernel/ironclaw_host_runtime/src/first_party_tools/http.rs
  • crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs
  • crates/kernel/ironclaw_host_runtime/tests/first_party_builtin_tools.rs
  • docs/reborn/contracts/host-runtime.md

Comment on lines +140 to +145
fn bounded_failure_diagnostic(output: Value, status: u16) -> String {
let Value::Object(mut output) = output else {
return fallback_diagnostic(status, None);
};
if serialized_output_len(&output) <= MODEL_DIAGNOSTIC_MAX_BYTES {
return scrub_model_diagnostic_controls(serialize_diagnostic(&output, status));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Reserve fence headroom before the early return.

Line 144 returns a 3–4 KiB diagnostic without fitting it to the fenced budget. If the diagnostic triggers the external-content fence, the fence pushes it past MODEL_DIAGNOSTIC_MAX_BYTES. The observation boundary can then truncate the diagnostic.

Only return early when the serialized output fits below MODEL_DIAGNOSTIC_MAX_BYTES - MODEL_DIAGNOSTIC_FENCE_HEADROOM_BYTES. Otherwise, use fit_output_to_budget. Add a caller-level regression that fences a diagnostic in this range and verifies the final observation remains bounded and complete.

Proposed fix
 fn bounded_failure_diagnostic(output: Value, status: u16) -> String {
     let Value::Object(mut output) = output else {
         return fallback_diagnostic(status, None);
     };
-    if serialized_output_len(&output) <= MODEL_DIAGNOSTIC_MAX_BYTES {
+    let unfenced_budget =
+        MODEL_DIAGNOSTIC_MAX_BYTES.saturating_sub(MODEL_DIAGNOSTIC_FENCE_HEADROOM_BYTES);
+    if serialized_output_len(&output) <= unfenced_budget {
         return scrub_model_diagnostic_controls(serialize_diagnostic(&output, status));
     }

As per coding guidelines, “For new or changed production-wired behavior, add a caller-level test at the nearest meaningful seam.” As per path instructions, “Test through the caller” applies when a helper gates a side effect.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/kernel/ironclaw_host_runtime/src/first_party_tools/http_output.rs`
around lines 140 - 145, Update bounded_failure_diagnostic so its early return
only applies when serialized_output_len(&output) is at most
MODEL_DIAGNOSTIC_MAX_BYTES minus MODEL_DIAGNOSTIC_FENCE_HEADROOM_BYTES; larger
diagnostics must continue through fit_output_to_budget. Add a caller-level
regression at the nearest observation seam that triggers fencing for a
diagnostic in this range and verifies the final observation is bounded and
complete.

Sources: Coding guidelines, Path instructions

@serrrfirat
serrrfirat added this pull request to the merge queue Aug 10, 2026
Merged via the queue into main with commit c517786 Aug 10, 2026
47 checks passed
@serrrfirat
serrrfirat deleted the codex/fix-http-error-semantics branch August 10, 2026 12:16
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…7342)

* fix(host-runtime): classify HTTP error responses as failures

* fix(host-runtime): address review — bound HTTP error diagnostics and add status regression tests (nearai#7330)

* docs(host-runtime): correct saved-body retrieval sequence for failed HTTP calls (nearai#7330)

* fix(host-runtime): converge failure-diagnostic trim, scrub controls, cover header/base64 branches

- share one budget-trim engine (fit_output_to_budget) between the success
  path and the failure diagnostic; the diagnostic trim now converges in a
  strict-progress loop instead of an 8-iteration cap, and the headers
  branch stays reachable when the inline body is empty or absent, so
  header-heavy 4xx/5xx diagnostics keep status/auth_hint/truncation
  instead of collapsing to the fallback verdict
- scrub DEL/C1 control bytes (U+007F..U+009F) from the serialized
  diagnostic: serde_json does not escape them, and ModelDiagnostic
  validation would otherwise replace the whole diagnostic with the fixed
  fallback sentence at the resolution boundary
- move the shaped output by value instead of cloning it per error
  response
- add unit coverage for the empty-body/header trim, base64 alignment on
  the failure path, and control-char scrubbing; wire the redirect
  regression test into the architecture-runtime gate

* fix(host-runtime): shape errors at diagnostic budget, attach wall clock, pin envelope size

- shape 4xx/5xx responses at the 4 KiB diagnostic budget instead of the
  success inline limit so the discarded success-budget trim pass (and its
  serializations) no longer runs on the failure path
- attach dispatch wall_clock_ms to the OperationFailed usage, matching
  the sibling first-party dispatches' failure-path accounting
- pin the truncation envelope below its reserved budget with a unit test
- correct the fit_output_to_budget doc to describe the incremental
  re-measure loop rather than a fixed three-serialization bound

* fix(host-runtime): pin error-status predicate, failure usage, and fallback shape

- single is_error_status predicate shared by the dispatch shape-limit
  selection and classify_status, so the 400..=599 boundary cannot drift
- pin failure usage accounting: classify_status unit test asserts egress
  bytes + wall_clock_ms on the OperationFailed outcome; the 403
  integration test asserts egress bytes reach the governor for failed
  calls
- pin the fallback diagnostic payload shape for non-object output and
  note why the serde-failure branch is unreachable by construction

* fix(host-runtime): migrate stale 5xx-success integration test to failed-outcome contract

- reborn_integration_http_matcher asserted the pre-change contract that a
  scripted HTTP 500 surfaces as a successful tool result; the new
  classification makes it a recoverable OperationFailed outcome, so the
  test now asserts ToolErrorClass::Failed with the operation_failed kind
  (run still completes; docstring updated to the contract doc)
- pin the post-fit fallback safety valve with an oversized untrimmable
  key test; correct the governor-accounting comment; dedupe the 400
  boundary rationale onto is_error_status

* docs(host-runtime): sync matcher guide to failed-outcome contract, document fence interaction

- tests/integration/CLAUDE.md .with_status entry now states 4xx/5xx
  classify as a Failed tool outcome (operation_failed) with sanitized
  diagnostic context; other statuses remain Completed results
- host-runtime contract doc records the loop-host injection-fence
  interaction: verdict semantics never depend on the fenced diagnostic
  surviving the observation bound (OperationFailed + safe summary always
  reach the model)
- document the fit_output_to_budget convergence bound (<= 3 passes)

* docs(host-runtime): correct serialize_diagnostic guard rationale

serde_json serializes every Value string without revalidating UTF-8
(probed: even an unsafe lone-surrogate string serializes Ok), so the
serde-failure arm is a pure defensive guard for future Value shapes, not
a reachable lone-surrogate path. Correct the doc comment and the
fallback test note to state the empirical fact.

* fix(host-runtime): keep shape-stage truncation flags in failure-diagnostic envelope

The re-inserted truncation envelope carried only the diagnostic-budget
trim state, so a 4xx/5xx response whose shape stage had already marked
headers or body as truncated (e.g. more than 32 headers) could end up
with headers_truncated:true beside an envelope claiming headers:false.
OR the surviving keys into the envelope and pin with a regression test;
boundary doc now states the full complement (outside 400..=599 stays
inspectable, including out-of-spec 600+).

* docs(host-runtime): sync support-module doc to failed-outcome contract, fix trim rationale

- tests/integration/support/http_matcher.rs module doc still claimed
  .with_status non-2xx stays Completed; now documents 4xx/5xx as a
  model-visible Failed operation_failed outcome
- bounded_failure_diagnostic doc: head-keeping truncation cuts only the
  last-sorted keys (status, truncation envelope); auth_hint sorts first
  and survives

* fix(host-runtime): address coderabbitai/ironloopai review — egress cap, fence headroom, saved-body fallback (nearai#7342)

- shape failed responses at the caller's response_body_limit again so
  egress-truncation accounting (body_was_truncated_by_egress) stays
  correct; the diagnostic display budget is applied separately. Pinned by
  builtin_http_error_diagnostic_preserves_egress_truncation_flag.
- reserve MODEL_DIAGNOSTIC_FENCE_HEADROOM_BYTES so a diagnostic wrapped
  in the loop-host external-content fence still fits the observation
  budget; pinned by failure_diagnostic_stays_within_budget_when_fenced.
- retain compact saved_body evidence (bounded path prefix +
  bytes_written) in the fallback verdict instead of dropping the save
  destination; pinned by failure_diagnostic_fallback_retains_saved_body_evidence.
- rename failure_diagnostic_falls_back_on_unserializable_output to the
  non-object contract it actually exercises (Value::Null serializes).
- host-runtime contract: correct the save-mode body-retention wording and
  document retry_after_ms semantics (None does not permit immediate retry).

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7342 — dc2f6517 Deployed Aug 10, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants