Skip to content

refactor(triggers,conversations): scan trusted trigger prompts at the mint, not inside one submitter (WS6) - #7136

Closed
BenKurrek wants to merge 1 commit into
mainfrom
ws6/conversations-trigger-safety
Closed

BenKurrek wants to merge 1 commit into
mainfrom
ws6/conversations-trigger-safety

Conversation

@BenKurrek

Copy link
Copy Markdown
Collaborator

What

PROPOSAL §6.4.2 asks for the trusted-trigger prompt safety scan to move "behind the triggers/kernel seam it guards (module move)". This does that, and reports two things the row got wrong.

One of the four items on CHECKLIST WS6.

It was not a module, and its location was a fail-open

The scan was three lines inside ConversationTrustedTriggerSubmitter::submit_trusted_trigger_fire (crates/ironclaw_conversations/src/inbound.rs), holding its own Arc<dyn InjectionScanner> built from Sanitizer::new(). src/trusted_trigger.rs — the module whose name suggests it — is TurnError classification and did not move.

That placement is the defect, not just the address: a guard inside one implementation of a port is lost the moment a second implementation exists. TrustedTriggerFireSubmitter is a public trait with a dyn wiring point (TriggerPollerWorkerDeps::trusted_submitter); nothing in the tree forced a new impl to re-run the scan, and the crate-tier suite never covered it (the only coverage was composition's worker-driven test).

Where it went, and why there

The seam is TrustedTriggerFireSubmitter, whose only input is the sealed TrustedTriggerSubmitRequest — which §6.4.3 already makes ironclaw_triggers the sole minter of. So the scan moved to the mint:

pub(crate) fn new(…) -> Result<Self, TriggerError> {
    crate::prompt_safety::validate_trusted_trigger_fire_prompt(&fire.prompt)?;
    Ok(Self { … })
}

"This prompt passed the trusted-prompt scan" is now an invariant of the type, not a step a submitter performs. new_for_test delegates to new, so the test-support seal bypasses visibility only, never the scan.

Severity policy did not move: ironclaw_safety::validate_trusted_trigger_prompt still owns high/critical-reject, medium-and-below-audit-only. The new crates/ironclaw_triggers/src/prompt_safety.rs owns only the scanner instance (a LazyLock<Sanitizer> — Sanitizer::new compiles an Aho-Corasick automaton plus a regex set) and the mapping to TriggerError::InvalidMaterialization.

Rejected alternatives

Alternative Why not
Scan in due_fire.rs just before the mint Same effect today, but leaves a second mint site free to skip it — and new is pub(crate) precisely so this crate can add one.
Scan before materialization Strictly stronger (a non-scanning materializer could not record an unsafe prompt into a thread), but it changes composition-owned ordering and makes composition's own scan dead. That is a composition-tier call, not this row's. Exposure is unchanged from before this move: composition scans first in production, so no shipped path records an unsafe prompt.
Invert into a TrustedTriggerPromptScanner port wired by composition A new trait with one real implementation, whose no-op wiring would be the same fail-open in a new place.

Behaviour: identical per fire, wider per port

Same rejection point, same TriggerError::InvalidMaterialization, same poller disposition — InvalidMaterialization classifies Permanent in both classify_failure and classify_submit_failure, which differ only on NotFound. Composition's pre-materialization scan (trigger_poller_trusted_submit.rs:172) is untouched, so the two-scan defence in depth survives with the second scan relocated. What changes is coverage: the second scan now applies to every TrustedTriggerFireSubmitter, not to whichever one composition happens to wire.

⚠ §6.4.2's dep list is wrong by one

§6.4.2 says the move is a module move and lists safety among conversations' target deps. Both cannot hold: the scan was that crate's only use of ironclaw_safety (inbound.rs:4, four imported symbols, one call site). The move drops the dependency. Correct target list: extension_contracts, filesystem, host_api, triggers + turn vocabulary via host_api. Amended in place with the sentence it replaces quoted.

Enforcement

  • ironclaw_triggers' boundary rule stops forbidding ironclaw_safety, with the reason inline: safety is a same-layer, I/O-free substrates leaf (no ironclaw_* deps at all), so this is a peer edge, not a reach upward, and a mint that does not validate what it seals is not a mint.
  • A new BoundaryRule for ironclaw_conversations — the crate had none — forbids ironclaw_safety, so the fail-open cannot be re-introduced by re-adding the dependency. It also pins ironclaw_threads (§6.4.2's "Never: … transcript content") and the usual retired/kernel-and-above set. Enforced against cargo metadata, not source text.
  • LAYER_MATRIX_EXCEPTIONS unchanged at 10 (counted in Python between const LAYER_MATRIX_EXCEPTIONS and its ];, so neither the struct definition nor the four out-of-array fixtures are miscounted).

Sabotage-proved at the caller tier

tick_rejects_injection_prompt_before_any_trusted_submitter_is_reached drives the real TriggerPollerWorker::tick_once with a materializer that does not scan (RecordingMaterializer — the shape of any materializer that skips or loses the check) and a submitter configured to accept, then asserts the submitter is never reached.

Deleting the one scan line from new:

Test Result under sabotage
tick_rejects_injection_prompt_before_any_trusted_submitter_is_reached (triggers) FAILED — "an injection prompt must fail the fire permanently, got Some(Submitted { run_id: … })"
unsafe_trigger_prompt_is_rejected_before_turn_submission (composition, pre-existing) FAILED — the production-wired guarantee tracks the new location

Both restored green (476 passed; 0 failed across the three touched crates; composition's test ok). A companion test pins that a medium-severity-only prompt still submits, so the mint cannot drift into a blanket filter — ordinary automation prompts trip act as.

Test accounting (unfiltered --list, quiescent tree)

Crate Before After Delta
ironclaw_conversations 97 97 name-identical, zero diff
ironclaw_triggers 169 173 +prompt_safety::tests::{high_severity_prompt_maps_to_invalid_materialization, medium_severity_warning_is_audit_only}, +worker::tests::{tick_rejects_injection_prompt_before_any_trusted_submitter_is_reached, tick_submits_a_prompt_whose_only_injection_warning_is_audit_only}
ironclaw_architecture 206 206 rules are data inside existing tests

No test edited for content. Function-roster diff against origin/main for every touched file shows exactly one removal (trigger_prompt_safety_rejection, which moved) and the two added tests — nothing silently deleted.

Also verified: cargo check --tests -p ironclaw_reborn_integration_tests (exit 0) for the two new_for_test call sites in tests/integration/support/triggered_submit.rs, and cargo clippy --all-targets --all-features -D warnings on the three touched crates.

🤖 Generated with Claude Code

… mint (WS6)

PROPOSAL §6.4.2 asked for the trusted-trigger prompt safety scan to move
"behind the triggers/kernel seam it guards". It was not a module: it was
three lines inside `ConversationTrustedTriggerSubmitter::submit_trusted_trigger_fire`
— one of the two implementations of `ironclaw_triggers::TrustedTriggerFireSubmitter`
— holding its own `Arc<dyn InjectionScanner>` from `Sanitizer::new()`.

That placement is a fail-open: a guard that lives inside one implementation
of a port is lost the moment a second implementation exists, and nothing in
the tree forced a new submitter to re-run it.

The seam is `TrustedTriggerFireSubmitter`, whose only input is the sealed
`TrustedTriggerSubmitRequest`, which `ironclaw_triggers` is the sole minter
of. So the scan moved to the mint: `TrustedTriggerSubmitRequest::new` is now
fallible and calls the new `ironclaw_triggers::prompt_safety` first, making
"this prompt passed the trusted-prompt scan" an invariant of the type rather
than a step some submitter performs. `new_for_test` delegates to `new`, so
the test-support seal bypasses visibility only, never the scan.

Behaviour at the fire level is unchanged — same rejection point, same
`TriggerError::InvalidMaterialization`, same permanent disposition — and
composition's pre-materialization scan is untouched, so defence in depth
survives with the second scan relocated and now covering every submitter.

`ironclaw_conversations` drops `ironclaw_safety` entirely (the scan was its
only use). Enforcement: triggers' boundary rule stops forbidding
`ironclaw_safety` (a same-layer, I/O-free `substrates` leaf — a peer edge,
not a reach upward), and a NEW `BoundaryRule` for `ironclaw_conversations`
forbids it, plus `ironclaw_threads` (§6.4.2's "Never: transcript content"),
a crate that was unruled until now.

Regression coverage at the caller tier, not on the helper:
`tick_rejects_injection_prompt_before_any_trusted_submitter_is_reached`
drives the real `TriggerPollerWorker::tick_once` with a materializer that
does NOT scan and a submitter configured to accept, and asserts the
submitter is never reached. A companion pins that a medium-severity-only
prompt still submits, so the mint cannot drift into a blanket filter.

Tests: conversations 97 -> 97 (name-identical), triggers 169 -> 173
(+2 worker, +2 prompt_safety unit), architecture 206 -> 206.
LAYER_MATRIX_EXCEPTIONS unchanged at 10.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@railway-app

railway-app Bot commented Aug 4, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-7136 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 4, 2026 at 11:29 am

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7136 August 4, 2026 11:20 Destroyed
@github-actions github-actions Bot added scope: docs Documentation scope: dependencies Dependency updates size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Aug 4, 2026
@coderabbitai

coderabbitai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 069926e2-2fbb-4bf3-b141-ae19565b0a79

📥 Commits

Reviewing files that changed from the base of the PR and between 1e2a294 and 847729d.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock, !**/Cargo.lock
📒 Files selected for processing (15)
  • crates/ironclaw_architecture/tests/reborn_dependency_boundaries.rs
  • crates/ironclaw_conversations/AGENTS.md
  • crates/ironclaw_conversations/CLAUDE.md
  • crates/ironclaw_conversations/Cargo.toml
  • crates/ironclaw_conversations/src/inbound.rs
  • crates/ironclaw_triggers/AGENTS.md
  • crates/ironclaw_triggers/Cargo.toml
  • crates/ironclaw_triggers/src/lib.rs
  • crates/ironclaw_triggers/src/prompt_safety.rs
  • crates/ironclaw_triggers/src/worker/due_fire.rs
  • crates/ironclaw_triggers/src/worker/ports.rs
  • crates/ironclaw_triggers/src/worker/tests.rs
  • docs/reborn/target-architecture/CHECKLIST.md
  • docs/reborn/target-architecture/PROPOSAL.md
  • tests/integration/support/triggered_submit.rs
💤 Files with no reviewable changes (1)
  • crates/ironclaw_conversations/Cargo.toml

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Trusted-trigger prompts are now safety-checked before submission.
    • High-severity injection attempts are blocked and recorded as permanent failures.
    • Medium-severity warnings remain eligible for submission while being audited.
  • Bug Fixes

    • Invalid prompts no longer reach trusted submitters.
    • Failed trigger requests now advance schedules appropriately instead of being retried indefinitely.
  • Documentation

    • Updated architecture and contributor guidance to clarify prompt-safety ownership and validation behavior.

Walkthrough

Trusted-trigger prompt scanning moved from ironclaw_conversations to ironclaw_triggers. TrustedTriggerSubmitRequest::new now validates prompts and returns a Result. The worker handles rejected requests before submission.

Changes

Trusted trigger safety ownership

Layer / File(s) Summary
Safety ownership and dependency boundaries
crates/ironclaw_architecture/tests/..., crates/ironclaw_triggers/..., crates/ironclaw_conversations/*.md, docs/reborn/target-architecture/*
ironclaw_triggers now owns trusted-trigger prompt scanning and may depend on ironclaw_safety. Dependency rules and architecture guidance prohibit the safety dependency and scanning logic in ironclaw_conversations.
Fallible request minting and submitter cleanup
crates/ironclaw_triggers/src/worker/ports.rs, crates/ironclaw_conversations/src/inbound.rs, tests/integration/support/triggered_submit.rs
Trusted request constructors validate prompts and return Result. Conversation submission removes duplicate scanning. Integration callers propagate construction errors.
Worker failure handling and regression coverage
crates/ironclaw_triggers/src/worker/due_fire.rs, crates/ironclaw_triggers/src/worker/tests.rs
The worker creates requests before submission. Rejected prompts produce permanent materialization failures and advance the schedule. Tests cover rejection and audit-only acceptance.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant TriggerWorker
  participant TrustedTriggerSubmitRequest
  participant prompt_safety
  participant TrustedTriggerSubmitter
  TriggerWorker->>TrustedTriggerSubmitRequest: Create request
  TrustedTriggerSubmitRequest->>prompt_safety: Validate prompt
  alt Rejected
    prompt_safety-->>TrustedTriggerSubmitRequest: InvalidMaterialization
    TrustedTriggerSubmitRequest-->>TriggerWorker: Error
    TriggerWorker->>TriggerWorker: Persist permanent failure
  else Accepted
    prompt_safety-->>TrustedTriggerSubmitRequest: Validated prompt
    TrustedTriggerSubmitRequest-->>TriggerWorker: Sealed request
    TriggerWorker->>TrustedTriggerSubmitter: Submit request
  end
Loading

Possibly related PRs

Suggested reviewers: serrrfirat

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description explains the implementation and testing well but omits most required template sections, including change type, linked issue, security impact, rollback, and review follow-through. Reformat the description to include every required template section, using "None" or "N/A" with reasons where applicable.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title uses Conventional Commits style and clearly describes moving trusted-trigger prompt scanning to request minting.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ironloopai

ironloopai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #7136

🟢 Completed · Review submitted

Submitted review →

Reviewed the complete trusted base-to-head comparison. No concrete correctness, security, architecture, maintainability, or test-coverage defects were found. The prompt scan is enforced at the sealed request mint, constructor failures retain the prior permanent-failure disposition, all mint call sites handle the new fallible API, and dependency/architecture documentation changes are consistent.

Automatic · PR opened · attempt 1 of 3 · completed in 1m 12s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 1e2a294
  • Head: ws6/conversations-trigger-safety at 847729d
  • Created: Aug 4, 2026, 11:25 AM UTC
  • Updated: Aug 4, 2026, 11:26 AM UTC
  • Run: 9af4f366-13ca-4b84-b2ce-cc47bc635815
  • Latest attempt: 1 · Completed · 6064d2f4-a29d-4d50-8bba-65bb465af9c3

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #7136

✅ No actionable findings

Reviewed the complete trusted base-to-head comparison. No concrete correctness, security, architecture, maintainability, or test-coverage defects were found. The prompt scan is enforced at the sealed request mint, constructor failures retain the prior permanent-failure disposition, all mint call sites handle the new fallible API, and dependency/architecture documentation changes are consistent.

Validation and technical details
  • Inspected all 16 changed files across refs/ironloop/base (1e2a294) and refs/ironloop/head (847729d).
  • Traced every TrustedTriggerSubmitRequest production/test constructor call and the TrustedTriggerFireSubmitter path.
  • Verified InvalidMaterialization from mint validation flows through classify_submit_failure and permanent fire persistence as before.
  • Verified the production composition materializer retains its independent pre-materialization safety scan.
  • Checked dependency-boundary changes, manifests, Cargo.lock, local guidance, and architecture documentation for consistency.
  • git diff --check completed successfully.
  • Targeted cargo tests could not be executed because cargo is unavailable in the review environment (/bin/bash: cargo: command not found).
  • Base: main
  • Head: ws6/conversations-trigger-safety at 847729d
  • Run: 9af4f366-13ca-4b84-b2ce-cc47bc635815

@BenKurrek

Copy link
Copy Markdown
Collaborator Author

Superseded: this branch's content merged to main inside the #7152 rename PR (verified by content, not ancestry — every added line present under the rename map, prompt_safety.rs byte-identical, the mint call-site live, the dep moved, and the caller-tier regression present; verification recorded in the #7152 refresh report). Closing as merged-via-#7152, not abandoned.

@BenKurrek BenKurrek closed this Aug 5, 2026

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7136 — 847729d2 Deployed Aug 4, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: dependencies Dependency updates scope: docs Documentation size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant