Skip to content

fix(reborn): hide routine implementation details - #6038

Closed
italic-jinxin wants to merge 15 commits into
mainfrom
issue-5707-hide-routine-internals
Closed

italic-jinxin wants to merge 15 commits into
mainfrom
issue-5707-hide-routine-internals

Conversation

@italic-jinxin

@italic-jinxin italic-jinxin commented Jul 13, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Adds fail-closed, user-facing display previews for routine create, list, pause, resume, and remove operations.
  • Updates the routine skill and trigger tool guidance so final replies use plain-language schedules and omit internal ids, raw cron, stored prompts, capability names, and host metadata.
  • Adds caller-level durable-history coverage and recorded Reborn QA contracts/replays for safe routine confirmations.
image

Linked Issue

Closes #5707

Validation

  • cargo test -p ironclaw_architecture
  • cargo test -p ironclaw_host_runtime --test tool_surface_contract
  • cargo test -p ironclaw_reborn_composition --lib
  • cargo test --test reborn_qa_recorded_behavior contract_routine_ -- --nocapture
  • cargo test --test reborn_qa_recorded_behavior replay_routine_ -- --nocapture
  • scripts/ci/check-reborn-qa-fixtures.sh
  • Targeted clippy with -D warnings
  • cargo fmt --all -- --check

Security Impact

Reduces accidental disclosure of internal routine identifiers, raw schedules, stored prompts, capability names, and host metadata through activity previews and assistant replies.

Database Impact

None. No schema or persistence-contract changes.

Blast Radius

Limited to Reborn routine-management display previews, model-facing routine guidance, and recorded QA expectations.

Rollback Plan

Revert the two commits to restore the previous generic capability previews and routine response guidance.


Review track: B

@italic-jinxin italic-jinxin added size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Jul 13, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 13, 2026 11:19 Destroyed
@coderabbitai

coderabbitai Bot commented Jul 13, 2026 •

Copy link
Copy Markdown

Review Change Stack

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: d72a3541-0139-4fd8-a0c2-77390757f4b6

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Routine trigger prompts are embedded from Markdown, routine previews use bounded allowlisted summaries, and capability-owned final replies flow through the loop before transcript finalization. Tests cover redaction, persistence, malformed fields, lifecycle handling, and integration behavior.

Changes

Routine presentation and final-reply boundary

Layer / File(s) Summary
Routine prompt contracts and capability wiring
crates/ironclaw_host_runtime/src/first_party_tools/prompts/*, crates/ironclaw_host_runtime/src/first_party_tools/trigger_management.rs, skills/routine-advisor/SKILL.md, crates/ironclaw_host_runtime/tests/tool_surface_contract.rs
Trigger verbs use embedded prompts and descriptions that define caller scope, delivery routing, plain-language schedules, and internal-field redaction.
Capability-owned routine previews
crates/ironclaw_host_runtime/src/first_party_tools/trigger_presentation.rs, crates/ironclaw_reborn_composition/src/projection/display_preview.rs
Routine inputs and outputs receive bounded titles, summaries, schedule/state labels, list truncation, and capability-aware matching.
Final-reply presentation propagation
crates/ironclaw_host_api/src/dispatch.rs, crates/ironclaw_turns/src/run_profile/model_observation.rs, crates/ironclaw_loop_host/src/capability_port.rs, crates/ironclaw_agent_loop/src/{state.rs,executor/*}
A typed safe reply is attached to completed capability results, queued in execution state, applied during finalization, and excluded from model-visible serialization.
Preview, contract, and integration validation
crates/ironclaw_reborn_composition/src/projection/tests/*, crates/ironclaw_reborn_composition/src/runtime/local_dev/tests/tests/display_preview.rs, crates/ironclaw_loop_host/**/tests/*, tests/integration/group_triggers/*, tests/reborn_qa_recorded_behavior.rs
Tests validate allowlisting, durable redaction, lifecycle result handling, malformed-field fallback, final-reply wording, and observation compatibility.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant TriggerManagement
  participant RoutinePresentation
  participant CapabilityWriteResult
  participant AgentLoop
  participant FinalTranscript
  TriggerManagement->>RoutinePresentation: project routine result
  RoutinePresentation-->>TriggerManagement: bounded preview and safe reply
  TriggerManagement->>CapabilityWriteResult: attach safe reply
  CapabilityWriteResult->>AgentLoop: forward result reference
  AgentLoop->>FinalTranscript: apply safe reply before finalization
Loading

Possibly related issues

Possibly related PRs

  • nearai/ironclaw#4780: Updates trigger descriptions and outbound delivery-target prompt wiring.
  • nearai/ironclaw#5131: Adds the pause/resume capabilities exercised by these prompt and presentation changes.
  • nearai/ironclaw#6059: Modifies the same ResultReference observation contract extended with final_reply_presentation.

Suggested reviewers: serrrfirat, henrypark133, ilblackdragon

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed Conventional Commits format is used and the summary matches the routine-privacy fix.
Description check ✅ Passed Mostly complete: it includes summary, linked issue, validation, security, impact, rollback, and review track sections.
Linked Issues check ✅ Passed The routine presentation and prompt updates hide internal IDs, cron, prompts, and metadata as #5707 requires.
Out of Scope Changes check ✅ Passed The added presentation plumbing, docs, and tests all support the same routine redaction goal; no unrelated change stands out.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size: L 200-499 changed lines size: XL 500+ changed lines scope: docs Documentation risk: low Changes to docs, tests, or low-risk modules and removed size: XL 500+ changed lines size: L 200-499 changed lines risk: medium Business logic, config, or moderate-risk modules labels Jul 13, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 13, 2026 11:19 Destroyed
@github-actions github-actions Bot added size: L 200-499 changed lines and removed size: XL 500+ changed lines labels Jul 13, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request transitions user-facing terminology from 'trigger' to 'routine' and ensures that internal implementation details, such as raw cron expressions and internal IDs, are redacted from user-facing replies and previews. It updates capability descriptions, display preview formatting, test assertions, and LLM trace fixtures to enforce this boundary. Feedback was provided regarding a potential discrepancy in routine_list_preview_lines where the reported count of routines could mismatch the actual listed routines if any are skipped due to missing names; a suggestion was made to filter the triggers first.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@italic-jinxin

Copy link
Copy Markdown
Contributor Author

@claude review

@claude

This comment was marked as resolved.

@railway-app

railway-app Bot commented Jul 13, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6038 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 14, 2026 at 11:27 am

@github-actions

github-actions Bot commented Jul 13, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.57% (302718 / 353765 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 353765 lines now vs 320188 at floor capture (+33577 lines, +10.49%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.57% — 302718 / 353765 lines

Per-crate breakdown (63 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_runtime_policy 31.75% 80 / 252
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_run_state 53.07% 225 / 424
ironclaw_authorization 53.89% 464 / 861
ironclaw_triggers 59.89% 1792 / 2992
ironclaw_observability 61.54% 16 / 26
ironclaw_webui_v2 62.93% 2679 / 4257
ironclaw_mcp 63.03% 578 / 917
ironclaw_reborn_cli 66.18% 4488 / 6781
ironclaw_filesystem 67.1% 3833 / 5712
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_memory 69.2% 773 / 1117
ironclaw_reborn_migration 71.57% 1551 / 2167
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.39% 1685 / 2265
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.67% 958 / 1283
ironclaw_extractors 74.72% 538 / 720
ironclaw_llm 78.36% 20328 / 25941
ironclaw_product_context 78.57% 11 / 14
ironclaw_first_party_extensions 78.81% 5576 / 7075
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_wasm_product_adapters 80.71% 1448 / 1794
ironclaw_memory_native 81.22% 3205 / 3946
ironclaw_secrets 82.7% 2791 / 3375
ironclaw_events 82.86% 1765 / 2130
ironclaw_reborn_identity 83.59% 433 / 518
ironclaw_wasm 83.97% 1011 / 1204
ironclaw_auth 83.99% 3147 / 3747
ironclaw_reborn_config 84.06% 1814 / 2158
ironclaw_processes 84.44% 993 / 1176
ironclaw_common 84.85% 1490 / 1756
ironclaw_turns 84.99% 13669 / 16084
ironclaw_host_api 85.13% 2663 / 3128
ironclaw_product_workflow 85.57% 10845 / 12674
ironclaw_projects 85.92% 659 / 767
ironclaw_network 86.12% 670 / 778
ironclaw_threads 86.7% 4594 / 5299
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_product_adapters 87.18% 3265 / 3745
ironclaw_skills 87.58% 4470 / 5104
ironclaw_hooks 87.78% 9921 / 11302
ironclaw_product_adapter_registry 88.06% 531 / 603
ironclaw_reborn_traces 88.19% 11946 / 13546
ironclaw_host_runtime 88.59% 17395 / 19635
ironclaw_reborn_composition 89.31% 80436 / 90066
ironclaw_extensions 89.38% 2971 / 3324
ironclaw_approvals 89.41% 1587 / 1775
ironclaw_runner 89.42% 16917 / 18919
ironclaw_reborn_openai_compat 89.55% 3798 / 4241
ironclaw_conversations 90.33% 3121 / 3455
ironclaw_event_streams 90.82% 1009 / 1111
ironclaw_loop_host 92.52% 14811 / 16008
ironclaw_resources 92.83% 4736 / 5102
ironclaw_attachments 93.06% 630 / 677
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.62% 2511 / 2682
ironclaw_agent_loop 94.83% 9148 / 9647
ironclaw_safety 95.04% 3677 / 3869
ironclaw_first_party_extension_ports 95.24% 3343 / 3510
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 13, 2026 11:59 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 13, 2026 12:25 Destroyed
@italic-jinxin italic-jinxin added the human-verified Manually tested and verified label Jul 13, 2026
@italic-jinxin
italic-jinxin marked this pull request as ready for review July 13, 2026 12:50
@italic-jinxin

Copy link
Copy Markdown
Contributor Author

@claude review

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 13, 2026 13:22 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 13, 2026 13:48 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 13, 2026 13:54 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_reborn_composition/src/projection/display_preview.rs (1)

917-1210: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Keep routine list counts consistent with rendered entries.

routine_list_preview_lines derives the count, capacity, and overflow marker from triggers.len(), but skips records without a non-empty name. A malformed record can therefore produce "2 routines found" while rendering one row; ten malformed records before a valid one can also hide the valid routine behind the limit.

Filter to displayable named routines first, then derive the count and limit from that collection (or render an explicit safe placeholder). Add a caller-level regression test for malformed list entries.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_reborn_composition/src/projection/display_preview.rs` around
lines 917 - 1210, Update routine_list_preview_lines to filter out triggers
lacking a non-empty name before computing the displayed count, capacity,
overflow marker, and ROUTINE_LIST_PREVIEW_LIMIT slice, so counts and rendered
rows remain consistent and valid routines are not hidden by malformed entries.
Add a caller-level regression test covering malformed list records, including
entries before a valid routine.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_reborn_composition/src/projection/display_preview.rs`:
- Around line 917-1210: Update routine_list_preview_lines to filter out triggers
lacking a non-empty name before computing the displayed count, capacity,
overflow marker, and ROUTINE_LIST_PREVIEW_LIMIT slice, so counts and rendered
rows remain consistent and valid routines are not hidden by malformed entries.
Add a caller-level regression test covering malformed list records, including
entries before a valid routine.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: f0fc9ffd-e9ea-4560-bcc6-e22b68b96419

📥 Commits

Reviewing files that changed from the base of the PR and between 2437783 and 3852d22.

📒 Files selected for processing (1)
  • crates/ironclaw_reborn_composition/src/projection/display_preview.rs

…ne-internals

# Conflicts:
#	crates/ironclaw_host_runtime/src/first_party_tools/trigger_management.rs
#	crates/ironclaw_host_runtime/tests/tool_surface_contract.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 14, 2026 05:45 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 14, 2026 06:12 Destroyed

@hanakannzashi hanakannzashi left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few blocking correctness issues remain:

  1. P2 - routine_capability() at display_preview.rs:967-975 classifies by namespace-stripped suffix. A valid third-party capability such as acme.trigger_create is therefore presented as a built-in Routine, and its real/custom preview is replaced. Please exact-match the canonical first-party IDs and supported provider aliases, with a negative third-party test.

  2. P2 - final-reply redaction does not cover every routine path. Only create/list descriptions gained the presentation constraint, while pause/resume/remove remain generic and trigger_output() still gives the model IDs, raw schedules, and host metadata. The routine skill is not guaranteed to activate for direct pause/delete/disable wording. Please provide the safe presentation contract for all five verbs.

  3. The recorded QA assertions replay manually edited canned final text, so they do not demonstrate that the new prompts produce safe replies. Please re-record under the new prompt and add caller-level coverage that attempts to return an ID/raw cron and verifies the product boundary.

  4. P2 - trigger_list.md asks the model to summarize tasks, but trigger_output() provides no task or safe task summary. The model must guess from the name. Either expose a bounded user-facing task summary or remove that requirement.

  5. P3 - routine_list_preview_lines() counts and truncates the raw array before skipping nameless entries. This can report 12 routines while rendering fewer, and malformed leading entries can hide valid later entries. Filter displayable entries before count/limit calculations.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 14, 2026 11:17 Destroyed
@github-actions github-actions Bot added size: XL 500+ changed lines and removed size: L 200-499 changed lines labels Jul 14, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6038 July 14, 2026 11:27 Destroyed
@github-actions github-actions Bot added size: L 200-499 changed lines and removed size: XL 500+ changed lines labels Jul 14, 2026
@italic-jinxin
italic-jinxin marked this pull request as draft July 14, 2026 11:34

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_host_api/src/dispatch.rs`:
- Around line 59-81: Update CapabilityFinalReplyPresentation to enforce
validation during deserialization by removing direct Deserialize derivation and
using serde’s try_from = "String" with its fallible new constructor. Rename
safe_reply() to the required as_str() accessor, then update downstream callers
such as assistant_reply.rs to use as_str() while preserving trimming,
truncation, and empty-value rejection.

In
`@crates/ironclaw_loop_host/src/capability_port/tests/runtime_lifecycle_tests.rs`:
- Around line 174-188: In the assertion for the extracted final-reply
presentation, replace the safe_reply() call with as_str() to match the newtype
validation API introduced in dispatch.rs, while preserving the expected "Edited
1 file" value.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: f7b44e47-2675-462c-bd10-68859eb095ab

📥 Commits

Reviewing files that changed from the base of the PR and between 6352ac0 and 242d2e7.

📒 Files selected for processing (37)
  • crates/ironclaw_agent_loop/src/executor.rs
  • crates/ironclaw_agent_loop/src/executor/assistant_reply.rs
  • crates/ironclaw_agent_loop/src/executor/capability_helpers.rs
  • crates/ironclaw_agent_loop/src/executor/loop_exit.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/state.rs
  • crates/ironclaw_architecture/tests/routine_presentation_boundary.rs
  • crates/ironclaw_first_party_extensions/src/coding/diff_preview.rs
  • crates/ironclaw_host_api/src/dispatch.rs
  • crates/ironclaw_host_runtime/src/first_party_tools/mod.rs
  • crates/ironclaw_host_runtime/src/first_party_tools/prompts/trigger_list.md
  • crates/ironclaw_host_runtime/src/first_party_tools/prompts/trigger_pause.md
  • crates/ironclaw_host_runtime/src/first_party_tools/prompts/trigger_remove.md
  • crates/ironclaw_host_runtime/src/first_party_tools/prompts/trigger_resume.md
  • crates/ironclaw_host_runtime/src/first_party_tools/trigger_management.rs
  • crates/ironclaw_host_runtime/src/first_party_tools/trigger_presentation.rs
  • crates/ironclaw_host_runtime/src/lib.rs
  • crates/ironclaw_host_runtime/src/obligations.rs
  • crates/ironclaw_loop_host/src/capability_port.rs
  • crates/ironclaw_loop_host/src/capability_port/tests/runtime_lifecycle_tests.rs
  • crates/ironclaw_loop_host/src/subagent_spawn_port/tests.rs
  • crates/ironclaw_loop_host/tests/thread_loop_host_contract.rs
  • crates/ironclaw_reborn_composition/src/extension_host/extension_lifecycle_capabilities.rs
  • crates/ironclaw_reborn_composition/src/projection/display_preview.rs
  • crates/ironclaw_reborn_composition/src/projection/tests/display_preview.rs
  • crates/ironclaw_reborn_composition/src/projection/tests/routine_display_preview.rs
  • crates/ironclaw_reborn_composition/src/root/product_live_adapters.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/result_read.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/tests/tests/display_preview.rs
  • crates/ironclaw_turns/src/run_profile/model_observation.rs
  • crates/ironclaw_turns/tests/turn_coordinator_contract.rs
  • docs/reborn/contracts/triggers.md
  • tests/integration/group_triggers/main.rs
  • tests/integration/group_triggers/scenario_final_reply_boundary.rs
  • tests/integration/support/golden.rs
  • tests/reborn_qa_recorded_behavior.rs

Comment on lines 59 to 81
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct CapabilityFinalReplyPresentation {
safe_reply: String,
}

impl CapabilityFinalReplyPresentation {
pub fn new(safe_reply: impl Into<String>) -> Option<Self> {
let safe_reply = truncate_capability_display_text(
safe_reply.into().trim(),
CAPABILITY_FINAL_REPLY_MAX_BYTES,
)
.text;
if safe_reply.is_empty() {
return None;
}

Some(Self { safe_reply })
}

pub fn safe_reply(&self) -> &str {
&self.safe_reply
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Enforce newtype validation on deserialization.

CapabilityFinalReplyPresentation derives Deserialize directly, allowing storage or network payloads to bypass the new() constructor and deserialize empty or oversized strings. As per coding guidelines, validated newtypes must use #[serde(try_from = "String")], a fallible new, and explicit as_str() methods.

Note: You will also need to update downstream callers (e.g., assistant_reply.rs) to use .as_str() instead of .safe_reply().

🛡️ Proposed fix to secure the deserialization boundary
-#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
-pub struct CapabilityFinalReplyPresentation {
-    safe_reply: String,
-}
+#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
+#[serde(try_from = "String", into = "String")]
+pub struct CapabilityFinalReplyPresentation(String);
 
 impl CapabilityFinalReplyPresentation {
     pub fn new(safe_reply: impl Into<String>) -> Option<Self> {
         let safe_reply = truncate_capability_display_text(
             safe_reply.into().trim(),
             CAPABILITY_FINAL_REPLY_MAX_BYTES,
         )
         .text;
         if safe_reply.is_empty() {
             return None;
         }
 
-        Some(Self { safe_reply })
+        Some(Self(safe_reply))
     }
 
-    pub fn safe_reply(&self) -> &str {
-        &self.safe_reply
+    pub fn as_str(&self) -> &str {
+        &self.0
     }
 }
+
+impl TryFrom<String> for CapabilityFinalReplyPresentation {
+    type Error = &'static str;
+
+    fn try_from(value: String) -> Result<Self, Self::Error> {
+        Self::new(value).ok_or("Final reply presentation cannot be empty")
+    }
+}
+
+impl From<CapabilityFinalReplyPresentation> for String {
+    fn from(val: CapabilityFinalReplyPresentation) -> Self {
+        val.0
+    }
+}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_host_api/src/dispatch.rs` around lines 59 - 81, Update
CapabilityFinalReplyPresentation to enforce validation during deserialization by
removing direct Deserialize derivation and using serde’s try_from = "String"
with its fallible new constructor. Rename safe_reply() to the required as_str()
accessor, then update downstream callers such as assistant_reply.rs to use
as_str() while preserving trimming, truncation, and empty-value rejection.

Source: Coding guidelines

Comment on lines +174 to +188
let CapabilityOutcome::Completed(completed) = outcome else {
panic!("expected completed outcome");
};
let presentation = match completed
.model_observation
.as_ref()
.map(|observation| &observation.detail)
{
Some(ironclaw_turns::run_profile::ToolObservationDetail::ResultReference {
final_reply_presentation: Some(presentation),
..
}) => presentation,
detail => panic!("expected final-reply presentation, got {detail:?}"),
};
assert_eq!(presentation.safe_reply(), "Edited 1 file");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Update method call to align with newtype validation guidelines.

If you apply the recommended newtype fix in crates/ironclaw_host_api/src/dispatch.rs, update this assertion to use .as_str() instead of .safe_reply().

♻️ Proposed refactor
     };
-    assert_eq!(presentation.safe_reply(), "Edited 1 file");
+    assert_eq!(presentation.as_str(), "Edited 1 file");
     let previews = result_writer.display_previews();
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
let CapabilityOutcome::Completed(completed) = outcome else {
panic!("expected completed outcome");
};
let presentation = match completed
.model_observation
.as_ref()
.map(|observation| &observation.detail)
{
Some(ironclaw_turns::run_profile::ToolObservationDetail::ResultReference {
final_reply_presentation: Some(presentation),
..
}) => presentation,
detail => panic!("expected final-reply presentation, got {detail:?}"),
};
assert_eq!(presentation.safe_reply(), "Edited 1 file");
let CapabilityOutcome::Completed(completed) = outcome else {
panic!("expected completed outcome");
};
let presentation = match completed
.model_observation
.as_ref()
.map(|observation| &observation.detail)
{
Some(ironclaw_turns::run_profile::ToolObservationDetail::ResultReference {
final_reply_presentation: Some(presentation),
..
}) => presentation,
detail => panic!("expected final-reply presentation, got {detail:?}"),
};
assert_eq!(presentation.as_str(), "Edited 1 file");
let previews = result_writer.display_previews();
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/ironclaw_loop_host/src/capability_port/tests/runtime_lifecycle_tests.rs`
around lines 174 - 188, In the assertion for the extracted final-reply
presentation, replace the safe_reply() call with as_str() to match the newtype
validation API introduced in dispatch.rs, while preserving the expected "Edited
1 file" value.

@hanakannzashi hanakannzashi left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed the latest head. The current tree is identical to the previously reviewed 0ea00e862 tree, so the functional findings in my earlier review remain current: third-party suffix misclassification, incomplete mutation-path redaction, canned rather than boundary-level final-reply coverage, missing task data for list summaries, and malformed list count/limit handling. Not approving this head yet.

@hanakannzashi hanakannzashi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved after follow-up review. The remaining discussion items are non-blocking for this change.

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6038 — 7c3aa0ba Deployed Jul 14, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs human-verified Manually tested and verified risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Routine creation response exposes internal implementation details

2 participants