fix(safety): redact model-bound secrets without rejecting turns - #7509
Conversation
…6364638 # Conflicts: # crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rs
# Conflicts: # crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rs
|
🚅 Deployed to the ironclaw-pr-7509 environment in ironclaw-ci-preview
|
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe change removes prompt-content secret denylists, preserves structural validation, and adds deterministic redaction at memory admission, model-context projection, and provider dispatch. ChangesPrompt security boundaries
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant PromptBuilder
participant MemoryContext
participant ModelGateway
participant Provider
PromptBuilder->>MemoryContext: materialize structurally valid content
MemoryContext->>ModelGateway: pass content after memory redaction
ModelGateway->>Provider: dispatch recursively redacted request
🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Temporarily closing and reopening to retrigger the Railway preview after the superseded fork PR cleanup cancelled the first deployment. |
🧭 IronLoop Run · ReviewThis comment updates in place as the Run moves through its stages. ⬛ Final result · Stopped
Automatic trigger · stopped after <1s IronLoop stopped because the pull request was closed while this Run was active. |
🧭 IronLoop Run · ReviewThis comment updates in place as the Run moves through its stages. 🟩 Final result · Completed
Automatic trigger · attempt 1 of 3 · completed in 18m 53s IronLoop completed the review and posted it to GitHub. 🔗 Result |
There was a problem hiding this comment.
Actionable comments posted: 8
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
crates/kernel/ironclaw_host_runtime/src/memory_context.rs (1)
294-305: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win
.ok()drops the validation error with no diagnostic.
from_untrusted_memoryreturnsAgentLoopHostError, and.ok()discards it. The caller at line 148 then skips the snippet silently. Every sibling drop path in this file logs first:sanitize_context_snippetat line 229 andadmit_laneat line 180 both emittracing::debug!. A memory snippet that fails structural validation now disappears with no signal.The repo rule is explicit: fail loud, or name the fallback. Log the rejection at
debug!(notinfo!) and add the// silent-ok:marker.🔍 Proposed fix
- LoopContextSnippet::from_untrusted_memory(snippet_ref, snippet.text).ok() + // silent-ok: a structurally invalid memory snippet degrades to "not admitted" + // rather than failing the turn; the drop is recorded below. + LoopContextSnippet::from_untrusted_memory(snippet_ref, snippet.text) + .inspect_err(|error| { + tracing::debug!( + error_kind = ?error.kind, + error_safe_summary = %error.safe_summary, + "dropping memory snippet rejected by structural prompt validation" + ); + }) + .ok()As per coding guidelines: "Do not use
.unwrap_or_default()onResult,.ok()?... justified fallbacks must include an inline// silent-ok: <reason>comment naming the operation."🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/kernel/ironclaw_host_runtime/src/memory_context.rs` around lines 294 - 305, Update to_loop_context_snippet so a failed LoopContextSnippet::from_untrusted_memory validation is logged with tracing::debug! before returning None, including the rejection error and relevant snippet context. Add the required inline // silent-ok: marker documenting the intentional fallback, while preserving successful conversions.Sources: Coding guidelines, Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rs`:
- Around line 730-735: Update the ratchet entry for ironclaw_loop_contracts to
describe the actual production growth: validation now checks structural limits
and control characters only, with Basic-auth samples limited to tests. Set the
ceiling to 13,495 to remain within the 400-line tolerance of the 13,095
production count, and append “Count read from this test's own failure message.”
In `@crates/contracts/ironclaw_loop_contracts/src/prompt_text.rs`:
- Around line 155-157: Update the documentation comment above
validate_prompt_text_with_diagnostics to remove the obsolete claim that a trust
gate relaxes credential-shaped value checks. Describe structural
control-character rejection across all surfaces and direct readers to the
provider-bound redaction boundary for credential handling.
In `@crates/kernel/ironclaw_turns/tests/agent_loop_host_contract.rs`:
- Around line 1542-1559: Align the trust-level terminology in both affected test
names and fixtures: either rename the tests to indicate Installed skills, or
change their skill_instruction_request fixtures to SkillTrustLevel::Untrusted if
that variant exists. Ensure the names accurately describe the trust level
actually exercised.
In `@crates/loop/ironclaw_loop_host/src/model_gateway.rs`:
- Around line 1653-1847: Move the self-contained redaction helpers, including
redact_completion_request, redact_tool_completion_request, redact_json_object,
and related functions, into a new model_gateway/redaction.rs submodule. Import
only the required ironclaw_llm types, serde_json collections, and
redact_model_input_text, then update model_gateway to reference the module
without changing dispatch behavior. Place or relocate focused unit tests for the
JSON-schema key remapping there.
- Around line 1653-1662: Measure the redaction cost in redact_completion_request
and redact_tool_completion_request using the existing CONTEXT_SHADOW_TARGET
shadow-measurement path, then eliminate repeated scans of unchanged replayed
content by caching redaction results keyed by immutable content_ref and/or
redacting tool schemas once at the capability-surface seam. Preserve redaction
behavior while ensuring each unchanged message and tool definition is processed
only once across provider dispatches.
- Around line 1678-1697: Update redact_chat_message to process
ContentPart::ImageUrl image_url.url before provider dispatch, using URL-aware
redaction that removes query credentials while preserving data: payloads; do not
rely on redact_string alone. Include the redacted URL changes in the returned
count and add a regression test for the recording provider covering
credential-bearing image URLs.
In `@crates/loop/ironclaw_loop_host/tests/llm_gateway.rs`:
- Around line 225-286: Extend
gateway_redacts_tool_descriptions_and_schema_strings_before_dispatch to exercise
prompt-cache signatures: submit two otherwise identical requests whose only
difference is a secret removed by redaction, then assert their recorded
system_prompt_cache_signature values are equal. Use the existing
provider/request recording symbols and preserve the current dispatch redaction
assertions.
In `@crates/substrates/ironclaw_safety/src/model_input_redaction.rs`:
- Line 82: Decouple the redaction filter in the model-input redaction flow from
the detector-internal "high_entropy_hex" literal by using a typed LeakDetector
accessor or enum-based pattern classification. Preserve the behavior that
high-entropy hexadecimal findings, including SHA-256 fingerprints, are excluded
from redaction; add detector-crate coverage that locks this classification or
name contract.
---
Outside diff comments:
In `@crates/kernel/ironclaw_host_runtime/src/memory_context.rs`:
- Around line 294-305: Update to_loop_context_snippet so a failed
LoopContextSnippet::from_untrusted_memory validation is logged with
tracing::debug! before returning None, including the rejection error and
relevant snippet context. Add the required inline // silent-ok: marker
documenting the intentional fallback, while preserving successful conversions.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 2f8e73f6-346b-4a0e-90be-a7db34ad3a9e
📒 Files selected for processing (16)
crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rscrates/app/ironclaw_architecture_tests/tests/reborn_extension_specificity.rscrates/contracts/ironclaw_loop_contracts/src/host/context.rscrates/contracts/ironclaw_loop_contracts/src/instruction_bundle.rscrates/contracts/ironclaw_loop_contracts/src/prompt_text.rscrates/contracts/ironclaw_loop_contracts/src/runtime_context/tests.rscrates/kernel/ironclaw_host_runtime/src/memory_context.rscrates/kernel/ironclaw_host_runtime/tests/memory_prompt_context.rscrates/kernel/ironclaw_turns/tests/agent_loop_host_contract.rscrates/loop/ironclaw_loop_host/src/lib.rscrates/loop/ironclaw_loop_host/src/model_gateway.rscrates/loop/ironclaw_loop_host/tests/llm_gateway.rscrates/loop/ironclaw_loop_host/tests/thread_loop_host_contract.rscrates/substrates/ironclaw_safety/README.mdcrates/substrates/ironclaw_safety/src/lib.rscrates/substrates/ironclaw_safety/src/model_input_redaction.rs
💤 Files with no reviewable changes (1)
- crates/app/ironclaw_architecture_tests/tests/reborn_extension_specificity.rs
Railway preview QA — FAIL
Given / When / Then matrix
Status derivation
Exact regression resultThe PR fixes whole-turn rejection: credential-shaped attachments remain usable and their benign context reaches the assistant. However, the stronger security claim is not satisfied for this live attachment path. With neutral prompts, the assistant received and echoed the complete synthetic password value in two independent tasks. The second disclosure persisted after refresh. Excluded prior evidenceThe earlier instruction “do not reveal any credential value” is excluded from the automatic-redaction acceptance result because it tested model instruction-following rather than the boundary guarantee. Remaining riskThe observed behavior indicates that at least one attachment/document-read path can deliver a credential-shaped value to the model despite the final model-bound scan. The exact missing boundary requires code/log diagnosis; this browser run does not infer which internal hop bypassed redaction. CleanupCleanup complete: both synthetic tasks were deleted, both local fixtures were removed, and the browser test tab was finalized. |
There was a problem hiding this comment.
🔍 IronLoop review
Found two high-severity security defects in the provider-bound redaction flow.
Findings: 🔴 High 2
🔴 High · Redact values paired with sensitive JSON keys
Inline on crates/loop/ironclaw_loop_host/src/model_gateway.rs:1722. See the inline comment for details.
🔴 High · Keep raw host paths out of provider prompts
Inline on crates/contracts/ironclaw_loop_contracts/src/prompt_text.rs:192. See the inline comment for details.
Validation
- ✅ Focused safety redaction tests — 6 targeted model-input redaction tests passed.
- ✅ Focused memory-boundary test — The targeted changed memory prompt-boundary test passed (1 test).
- ⚪ Broad validation — Not run. Not run; static call-path analysis and focused checks were sufficient for this review.
Review details
- Run:
7e19a04a-f33a-4dde-9a4c-15fb9e04ee08 - Workflow: Review
- Attempts: 1
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
crates/loop/ironclaw_loop_host/tests/llm_gateway.rs (1)
177-226: 🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy liftAdd an attachment-to-provider regression test.
gateway_redacts_every_message_role...setsimage_parts: Vec::new()and bypasses attachment ingress. The existing attachment test stops atRecordingGatewayand asserts raw image bytes, not the provider request. This violates theAGENTS.md“Test through the caller” invariant. Drive the attachment path throughLlmProviderModelGatewayand assert[REDACTED_SECRET]without the original value before claiming provider-bound redaction.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/loop/ironclaw_loop_host/tests/llm_gateway.rs` around lines 177 - 226, The current regression test only verifies text redaction and does not exercise attachment ingress through the provider gateway. Extend the test around gateway.stream_model and the model_request setup to include an attachment containing a secret, route it through LlmProviderModelGateway to the provider, and assert the provider-bound request contains [REDACTED_SECRET] while excluding the original secret value.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@crates/substrates/ironclaw_safety/src/model_input_redaction.rs`:
- Around line 15-18: The credential-redaction regex in the model-input redaction
logic currently stops at whitespace and delimiters inside quoted values. Update
the relevant pattern and matching logic to support complete single-, double-,
and backtick-quoted values, including escaped quotes, for both structured
credentials and Authorization values; preserve unquoted matching behavior. Add
regression tests covering spaces, commas, semicolons, and escaped quotes, and
verify the full credential is redacted according to the documented safety
invariant.
---
Outside diff comments:
In `@crates/loop/ironclaw_loop_host/tests/llm_gateway.rs`:
- Around line 177-226: The current regression test only verifies text redaction
and does not exercise attachment ingress through the provider gateway. Extend
the test around gateway.stream_model and the model_request setup to include an
attachment containing a secret, route it through LlmProviderModelGateway to the
provider, and assert the provider-bound request contains [REDACTED_SECRET] while
excluding the original secret value.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: c2d9e54d-7f2c-4925-90ec-966c6b339349
📒 Files selected for processing (2)
crates/loop/ironclaw_loop_host/tests/llm_gateway.rscrates/substrates/ironclaw_safety/src/model_input_redaction.rs
…very-prompt-denylist # Conflicts: # crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rs
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
crates/loop/ironclaw_loop_host/src/lib.rs (1)
2990-2997: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick winConsume the complete tool-result group.
load_task_pinned_context_windowremoves only oneToolResultReferenceafter an assistant. Multiple distinct results can follow one assistant. With[Assistant, Ref1, Ref2],Ref2remains model-visible andrecent_window_truncationstops atRef1. Remove the full consecutive group and extendtask_pin_evicts_complete_tool_exchange_at_a_compactable_boundarywith two references. This follows the repository’s caller-level test-through-the-caller invariant.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/loop/ironclaw_loop_host/src/lib.rs` around lines 2990 - 2997, Update load_task_pinned_context_window to remove the entire consecutive ToolResultReference group following a displaced Assistant, not just the first reference, so later references are not model-visible and truncation reaches the group boundary. Extend task_pin_evicts_complete_tool_exchange_at_a_compactable_boundary with a case containing two references and verify both are consumed through the caller-level behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@crates/loop/ironclaw_loop_host/src/lib.rs`:
- Around line 2990-2997: Update load_task_pinned_context_window to remove the
entire consecutive ToolResultReference group following a displaced Assistant,
not just the first reference, so later references are not model-visible and
truncation reaches the group boundary. Extend
task_pin_evicts_complete_tool_exchange_at_a_compactable_boundary with a case
containing two references and verify both are consumed through the caller-level
behavior.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 375d06c0-5948-4d98-b525-7faf4b22f7cb
📒 Files selected for processing (4)
crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rscrates/contracts/ironclaw_loop_contracts/src/host/context.rscrates/loop/ironclaw_loop_host/src/lib.rscrates/loop/ironclaw_loop_host/tests/thread_loop_host_contract.rs
…very-prompt-denylist # Conflicts: # crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rs
…very-prompt-denylist # Conflicts: # crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rs
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/integration/extension_visibility.rs`:
- Around line 95-102: The system-prompt assertion in the extension visibility
test is not specific to the local description and can pass using the verified
description. Update the local description fixture and its corresponding
assertion around assert_system_prompt_contains to include a unique local-only
marker, then assert that marker in the local system-prompt projection while
preserving the existing verified-source checks.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 06c1729e-a4e6-4f59-b777-bb8df51462f4
📒 Files selected for processing (3)
tests/CLAUDE.mdtests/integration/extension_visibility.rstests/integration/support/harness/profiles/extension.rs
| .assert_model_tool_description_contains( | ||
| "verifiedprompt__invoke", | ||
| AUTH_VOCABULARY_DESCRIPTION, | ||
| ) | ||
| .await | ||
| .expect("verified catalog description reaches the model intact, including Bearer"); | ||
| harness | ||
| .assert_system_prompt_contains(PROMPT_DENIAL_DESCRIPTION) | ||
| .assert_system_prompt_contains(AUTH_VOCABULARY_DESCRIPTION) |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Make the local system-prompt assertion source-specific.
assert_system_prompt_contains(AUTH_VOCABULARY_DESCRIPTION) at Line 102 can be satisfied by the verified description. The local assertion at Lines 118-120 checks only localprompt.unsafe. The test can therefore pass if the local description is still omitted from the system prompt. Add a local-only marker and assert it in the local system-prompt projection.
Also applies to: 114-120
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/integration/extension_visibility.rs` around lines 95 - 102, The
system-prompt assertion in the extension visibility test is not specific to the
local description and can pass using the verified description. Update the local
description fixture and its corresponding assertion around
assert_system_prompt_contains to include a unique local-only marker, then assert
that marker in the local system-prompt projection while preserving the existing
verified-source checks.
…wn-ratchet Second refresh fold of the day (#7509 + #7365). Conflicts were again the composition bookkeeping pair; resolved by keeping both dated chains and taking the merged tree's own measurement: - composition-budget.toml + reborn_restructure_baselines.rs: main's #7365 re-ratcheted DOWN 41810 -> 41533 (memory-save guidance evicted to the memory-native package); this branch adds no composition code, and the merged tree measures 41533 LOC / 831 Arc<dyn> exactly, so main's banked eviction stands and the arc_dyn re-equalization from the previous fold carries through. Verified with the gate's --print. - Main's three SIZE_CEILINGS bumps (extension_contracts 7947, host_api 19086, loop_contracts 13524) auto-merged; the ceiling gate passes on the merged tree with the windows intact. Verified green: reborn_dependency_boundaries 42/42 (incl. the window-edge fixture), reborn_restructure_baselines, composition-budget gate, clippy clean, fmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ates re-measured on the merged tree Conflicts were the three measurement files only. Ceilings re-measured, not summed (composition 41780 LOC / 838 governed Arc<dyn> sites; contracts-tier ceilings verified passing under the union records). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ai#7509) * fix(loop): allow security prose in recovered context * test(turns): align prompt safety contract coverage * fix(loop): address review feedback on prompt recovery (nearai#7434) * fix(loop): reject filler-separated credentials (nearai#7434) * fix(safety): redact model-bound secrets without rejecting turns * fix(safety): preserve non-secret sha256 fingerprints * test(safety): align channel context with gateway redaction * fix(gateway): address coderabbit review — preserve redacted JSON shape (nearai#7434) * fix(safety): redact quoted structured credentials * fix(safety): scan encoded tool result content * fix(safety): redact structured credential values * fix(safety): redact character-dump credentials * fix(safety): close provider-bound redaction gaps (nearai#7509) * fix(safety): close structured redaction review gaps (nearai#7509) * fix(safety): redact nested schema and URL fragment secrets (nearai#7509) * test(integration): align prompt trust expectation (nearai#7509)
Summary
[REDACTED_SECRET].Change Type
Linked Issue
None.
Validation
cargo fmt --all -- --checkcargo clippy --all --benches --tests --examples --all-features -- -D warningscargo test -p ironclaw_safety --no-fail-fast(308 passed)cargo test -p ironclaw_loop_contracts --no-fail-fastcargo test -p ironclaw_turns --test agent_loop_host_contract --no-fail-fast(87 passed)cargo test -p ironclaw_host_runtime --test memory_prompt_context --no-fail-fast(18 passed)cargo test -p ironclaw_loop_host --test llm_gateway --no-fail-fast(92 passed)cargo test -p ironclaw_loop_host --lib provider_bound_redaction_covers_stop_sequences --no-fail-fastcargo test -p ironclaw_architecture_tests --no-fail-fastRUST_MIN_STACK=16777216 cargo test -p ironclaw_integration_tests --test reborn_integration_golden_payload --no-fail-fast(21 passed)cargo test -p <owning-crate> --features integration: Not applicable; no database-backed behavior changed.Test Strategy
User behavior: Given recovered text containing security vocabulary, host paths, credential-shaped values, or prompt-injection markers, when IronClaw reconstructs and dispatches a model request, benign context remains usable, detected credential values are placeholders, injection-bearing memory snippets are isolated, and the turn continues.
Risk areas:
Tests added or updated:
[REDACTED_SECRET]and the canary was absent.What the tests prove: detected secrets do not reach provider-visible prompt content, while false positives no longer reject a whole turn or remove unrelated context. Prompt-injection containment remains independent of secret handling.
Security Impact
The model-input boundary now scans immediately before provider dispatch, after all prompt assembly and before prompt-cache hashing. It covers plain/tool/streaming/repair paths; message content across roles; text content parts; plain reasoning and summaries; tool-call argument string keys/values and parse errors; tool descriptions and JSON schema strings/keys; and stop sequences. It logs only aggregate redaction counts, never detected values.
Opaque protocol material that requires exact replay—encrypted/redacted reasoning payloads, signatures, tool-call IDs/names, route/model identifiers, and provider metadata—is not treated as prompt text. Managed credentials remain host-side and continue to use mediated injection rather than prompt construction.
This guarantees redaction for formats recognized by the existing leak detector plus labeled weak values such as
password: letmein; it is not a cryptographic proof that arbitrary unlabeled prose cannot contain a secret. The strongest invariant remains: secrets known to IronClaw must stay in the secret store and never be assembled into model input. The final scan is defense in depth for untrusted text and accidental leakage.Warn-only high-entropy hex findings remain visible unless they are credential-labeled, so ordinary SHA-256 fingerprints do not mutate prompts or prompt-cache identity.
Prompt injection is intentionally not conflated with secret detection. Existing trust envelopes and injection checks still isolate offending untrusted snippets. SkillSpector-style skill trust scanning and NeMo Guardrails-style broader input/output policy remain follow-up layers rather than dependencies of this recovery hotfix.
Reborn Trust-Boundary Checklist
serde(default)fields: none changed.Database Impact
None. Raw stored LLM/thread/memory data is retained; redaction applies to the transient model-facing view.
Blast Radius
Prompt construction contracts, production memory admission, and all loop-host provider request shapes. A regression could leak a recognized secret, corrupt opaque replay data, or reintroduce thread-wide rejection. Caller-level provider captures, memory integration tests, structural contract tests, all-features clippy, and architecture ratchets cover those directions.
Compatibility and Rollback
No wire or storage migration. Removing the unused
base64dependency fromironclaw_loop_contractsshrinks the contract layer. Roll back by reverting this PR; stored data requires no repair. During rollback, the prior credential denylist behavior—and its false-positive thread failures—would return.Follow-up
Review track: C (security/runtime/DB/CI)
Prior exact-head Railway QA (269c020)
The first live attempt exposed a second path: after direct and JSON-encoded results were redacted, the agent retried through
shellwithod -c. Its character-separated output (p a s s w o r d) bypassed label matching even though the model could reconstruct the value. The final fix decodes only offset-prefixed character dumps for detection and replaces the encoded field when the reconstructed text contains a credential assignment. A benignSecretary of the Treasurydump remains unchanged.Validation on
269c0206f8bf26ac233cd62a4e92cab9e19bba5b:cargo test -p ironclaw_safety --no-fail-fast(308 passed)llm_gatewayironclaw_safetyand thellm_gatewaytest target with-D warningsreborn_dependency_boundaries(41 passed)[REDACTED_SECRET], canary absentNo real credential was used. The synthetic QA task and local fixture were deleted after verification.
Final exact-head Railway QA (362966a)
362966a2417eaa2babb7d28cbabaa9da7595d7a4.Read the attached JSON and report every field and its value.passwordrendered as[REDACTED_SECRET]; the exact synthetic canary and distinctive tail were absent.Final review regressions on this head cover propagation of sensitive schema context through nested composition/applicator keywords, credential redaction in URL fragments while benign state remains visible, and the corrected memory-path expectation. CodeRabbit reported no new actionable comments.