Skip to content

Fix compaction failures after tool results - #5895

Merged
henrypark133 merged 12 commits into
mainfrom
issue-5838-compaction-error
Jul 10, 2026
Merged

henrypark133 merged 12 commits into
mainfrom
issue-5838-compaction-error

Conversation

@henrypark133

Copy link
Copy Markdown
Collaborator

Closes #5838.

Summary

  • Treat non-cancellation compaction errors/timeouts as recoverable prompt-step skips instead of terminal CompactionUnavailable run failures.
  • Emit CompactionFailed, clear the forced-compaction bit, remember the deferred watermark, and continue with the existing prompt path.
  • Preserve cancellation behavior: explicit compaction cancellation still exits as cancelled, and cancellation after a failed compaction still wins before the next model call.

Bug evidence

Before the production fix, the new executor regression failed after a successful large tool result because forced compaction returned SecurityRejected and the executor exited with CompactionUnavailable / compaction_security_rejected instead of letting the post-tool model turn finish.

Tests

  • cargo fmt
  • cargo test -p ironclaw_agent_loop compaction -- --nocapture
  • cargo test --test reborn_integration_http_matcher multi_tool_turn_survives_failed_forced_compaction_after_results -- --nocapture
  • cargo clippy -p ironclaw_agent_loop --all-targets -- -D warnings
  • cargo clippy --test reborn_integration_http_matcher -- -D warnings
  • git diff --check

Note: cargo prints the existing workspace warning unused config key net.retries.

Copilot AI review requested due to automatic review settings July 9, 2026 17:01
@ironloopai

ironloopai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: d30ef6228c4b1895b541f14c5584db4297d0b88d
Result: No reviewer jobs are scheduled yet.
Next: Run @ironloopai review to start reviewers.
Updated: 2026-07-10T22:20:36.581Z

Current reviewers:

Reviewer State Verdict Findings Last update
none Queued N/A No reviewer jobs scheduled yet. N/A
Reviewer summaries
Reviewer Detail
none No reviewer jobs scheduled yet.
Recent activity
Time Reviewer State Detail
N/A N/A Waiting No progress events recorded yet.
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
  • @ironloopai status
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5895 July 9, 2026 17:01 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules labels Jul 9, 2026
@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes
    • Phase 1 compaction failures are now non-terminal: the loop emits a compaction-failed progress event and continues with the prepared prompt flow (including inference failures and security rejections); cancellation remains terminal.
    • Compaction-failure continuation is deferred and avoids intermediate checkpoint writes while still persisting the final checkpoint with the deferred watermark.
    • Forced compaction failures after tool-result overflow no longer break multi-tool turn completion.
  • Tests
    • Updated unit/integration/failure-matrix expectations for continued execution and deferred-watermark assertions.
    • Added an integration regression test for forced compaction failure survival.
  • Documentation
    • Refreshed the compaction error-handling spec to match the new deferred continuation and timeout semantics.

Walkthrough

Compaction errors and inference failures now defer compaction and continue through the existing prompt path. Tests, integration milestone assertions, failure-category coverage, and the frozen compaction specification reflect this non-terminal behavior.

Changes

Deferred compaction continuation

Layer / File(s) Summary
Compaction failure continuation
crates/ironclaw_agent_loop/src/executor/prompt.rs
Failures emit CompactionFailed, record deferred state, clear forced compaction, and return Skipped; outer timeout racing and terminal failure categorization are removed.
Executor continuation regression tests
crates/ironclaw_agent_loop/src/executor/tests.rs, crates/ironclaw_agent_loop/src/executor/tests/failure_matrix.rs
Tests expect prepared or completed continuation, deferred watermarks, absent checkpoints, and completed divergence outcomes.
Completed loop assertions
crates/ironclaw_agent_loop/tests/executor_happy_paths.rs
Compaction failure coverage expects successful completion and a CompactionFailed event.
Integration milestone observability
tests/integration/support/{builder,group,assertions}.rs
The harness retains milestone state and adds baseline-scoped milestone and history assertions.
Forced compaction rejection integration test
tests/integration/http_matcher.rs
Multi-tool coverage verifies ordered tool calls, failure milestones, unsafe-summary exclusion, and the final reply.
Failure-summary category alignment
crates/ironclaw_reborn_composition/src/projection/tests/failure_explanation.rs
Granular compaction categories are removed from the current agent-loop parity set.
Behavioral contract update
docs/reborn/2026-05-26-context-compaction.md
The specification documents non-terminal failures, watermark deferral, and inner inference deadline enforcement.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant ToolExecution
  participant PromptCompactionStep
  participant LoopState
  participant ModelTurn

  ToolExecution->>PromptCompactionStep: forced compaction after tool results
  PromptCompactionStep->>LoopState: emit failure and record deferred watermark
  PromptCompactionStep->>LoopState: clear forced compaction
  PromptCompactionStep->>ModelTurn: continue with existing prompt
  ModelTurn-->>ToolExecution: finalize assistant reply
Loading

Possibly related issues

  • 4474 — Covers forced compaction after capability-result overflow and continuation behavior.
  • 4313 — Concerns CompactionFailed milestone reason payloads, which are asserted by this change.
🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description covers summary, bug evidence, and tests, but omits major required template sections and checklist items. Add the missing template sections, especially Change Type, Validation checklist, Security Impact, Rollback Plan, and Review Follow-Through.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title is directly about the compaction-after-tool-results fix and matches the main change.
Linked Issues check ✅ Passed The implementation matches #5838 by making post-tool compaction failures recoverable and preserving successful tool results.
Out of Scope Changes check ✅ Passed The docs and test harness changes stay aligned with the compaction-failure fix; no unrelated scope is evident.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the contributor: core 20+ merged PRs label Jul 9, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request modifies the agent loop executor to make prompt compaction failures recoverable. Instead of terminating the execution with a failure exit, the loop now continues on the normal prompt path using the existing prompt candidate. Corresponding unit and integration tests have been updated or added to verify this fallback behavior and to track compaction milestones. The review feedback suggests improving the test assertions in tests/integration/http_matcher.rs by replacing a generic .is_err() check with a specific typed error assertion using .expect_err() to prevent false positives from unrelated infrastructure issues.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread tests/integration/http_matcher.rs Outdated
Comment on lines +137 to +145
assert!(
h.assert_conversation_history_role_contains(
MessageKind::Summary,
"ignore previous instructions"
)
.await
.is_err(),
"unsafe compaction summary must not be persisted"
);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using a generic .is_err() assertion here can lead to false positives if the harness encounters an unrelated database or infrastructure error. Additionally, avoid using string-matching on error messages (such as err.to_string().contains(...)) to classify failures. Instead, assert against the specific expected error variant using typed checks (e.g., matches!), and prefer using .expect_err() with a descriptive message instead of .unwrap_err() to ensure failures are explicitly reported with clear context.

    let err = h
        .assert_conversation_history_role_contains(
            MessageKind::Summary,
            "ignore previous instructions",
        )
        .await
        .expect_err("expected conversation history assertion to fail");
    assert!(
        matches!(err, ConversationError::NotFound),
        "expected NotFound error, got: {err:?}"
    );
References
  1. Avoid using generic is_err() assertions in tests when verifying that an operation fails. Instead, assert against the specific expected error message or variant to prevent infrastructure or harness-level failures from causing false positives.
  2. Avoid using string-matching on error messages (e.g., err.to_string().contains(...)) to classify or allow-list failures in tests. Prefer typed checks or introducing typed seams in the test harness to prevent unrelated failures from being silently retried or swallowed.
  3. In Rust tests, prefer using .expect() with a descriptive message instead of .unwrap() or fallbacks like unwrap_or_else() to ensure that failures are explicitly reported with clear context.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
✅ Approved 0 0 0 868957fb1c81

Head: 868957fb1c81ad787e91a6f2f4633e2e0833f294
Next: No reviewer action needed.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

No concrete actionable issues found in the compaction-failure recovery change or its test/support updates.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.
  4. Use @ironloopai status to check queued/running/completed/failed/superseded state while reviewers run.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_agent_loop/src/executor/prompt.rs (1)

505-530: 📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

Duplicate deferred-compaction state mutation between Deferred branch and compaction_failed_continue.

The LoopCompactionOutcome::Deferred arm (lines 512-530) and the new compaction_failed_continue helper (lines 657-669) both: clear force_compact_on_next_iteration, build an identical DeferredCompactionWatermark { through_seq: drop_through_seq, prompt_fingerprint: state.compaction_prompt.fingerprint() }, run cancel_if_requested_after_pending_input_ack, and return Skipped(state). Only the progress-event emission differs. Any future change to the watermark/cancellation sequencing now has two call sites to keep in sync.

♻️ Proposed refactor: extract the shared tail into one helper
+async fn defer_compaction(
+    ctx: StageContext<'_>,
+    mut state: LoopExecutionState,
+    pending_input_ack: &mut PendingInputAck,
+    drop_through_seq: u64,
+) -> Result<PromptCompactionOutcome, AgentLoopExecutorError> {
+    state.compaction_state.force_compact_on_next_iteration = false;
+    state.compaction_state.last_deferred = Some(DeferredCompactionWatermark {
+        through_seq: drop_through_seq,
+        prompt_fingerprint: state.compaction_prompt.fingerprint(),
+    });
+    state = match CheckpointStage
+        .cancel_if_requested_after_pending_input_ack(ctx, state, pending_input_ack)
+        .await?
+    {
+        CancelCheck::Continue(state) => *state,
+        CancelCheck::Exit(exit) => return Ok(PromptCompactionOutcome::Exited(exit)),
+    };
+    Ok(PromptCompactionOutcome::Skipped(state))
+}

Then both the Deferred arm and compaction_failed_continue call defer_compaction(...) after doing their own (differing) progress-event emission.

As per coding guidelines, "Keep functions focused and extract helpers when logic is reused."

Also applies to: 639-670

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_agent_loop/src/executor/prompt.rs` around lines 505 - 530,
The Deferred compaction handling is duplicated between the
LoopCompactionOutcome::Deferred branch in PromptCompactionOutcome processing and
the compaction_failed_continue helper. Extract the shared tail into a single
helper, such as defer_compaction, that clears force_compact_on_next_iteration,
sets last_deferred with the DeferredCompactionWatermark built from
drop_through_seq and state.compaction_prompt.fingerprint(), performs
cancel_if_requested_after_pending_input_ack, and returns Skipped(state). Keep
only the progress-event emission differences in the existing call sites and
route both paths through the shared helper.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_agent_loop/src/executor/prompt.rs`:
- Around line 505-530: The Deferred compaction handling is duplicated between
the LoopCompactionOutcome::Deferred branch in PromptCompactionOutcome processing
and the compaction_failed_continue helper. Extract the shared tail into a single
helper, such as defer_compaction, that clears force_compact_on_next_iteration,
sets last_deferred with the DeferredCompactionWatermark built from
drop_through_seq and state.compaction_prompt.fingerprint(), performs
cancel_if_requested_after_pending_input_ack, and returns Skipped(state). Keep
only the progress-event emission differences in the existing call sites and
route both paths through the shared helper.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 13263780-d8f5-4553-b836-6e95f1196692

📥 Commits

Reviewing files that changed from the base of the PR and between 8e05a37 and 868957f.

📒 Files selected for processing (7)
  • crates/ironclaw_agent_loop/src/executor/prompt.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/executor/tests/failure_matrix.rs
  • crates/ironclaw_agent_loop/tests/executor_happy_paths.rs
  • tests/integration/http_matcher.rs
  • tests/integration/support/builder.rs
  • tests/integration/support/group.rs

@railway-app

railway-app Bot commented Jul 9, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5895 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 10, 2026 at 10:29 pm

henrypark133 and others added 2 commits July 10, 2026 10:05
…ty guard

PR #5895 (868957f) removed compaction_failure_category from the agent
loop, changing non-cancellation compaction failures from a terminal exit
to a deferred continuation (issue #5838, design doc section 5). The
composition-core safe-summary parity guard still expected the agent loop
to mint the 7 granular compaction_* categories, which are no longer
produced. crate::failure_summary keeps display support for them so
historical failure records still render.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Deduplicate the identical tail of the Deferred outcome arm and
compaction_failed_continue into a crate-private defer_compaction
helper; each caller emits its own progress event then delegates.
Behavior-preserving: mutate-then-cancel-check ordering unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 17:12
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5895 July 10, 2026 17:13 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions github-actions Bot added size: L 200-499 changed lines and removed size: M 50-199 changed lines labels Jul 10, 2026
@github-actions

github-actions Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.25% (292535 / 343140 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 343140 lines now vs 320188 at floor capture (+22952 lines, +7.17%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.25% — 292535 / 343140 lines

Per-crate breakdown (63 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_runtime_policy 31.75% 80 / 252
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_run_state 52.36% 222 / 424
ironclaw_authorization 53.66% 462 / 861
ironclaw_triggers 59.89% 1792 / 2992
ironclaw_observability 61.54% 16 / 26
ironclaw_webui_v2 62.76% 2659 / 4237
ironclaw_reborn_cli 62.83% 3842 / 6115
ironclaw_mcp 63.03% 578 / 917
ironclaw_reborn_migration 67.09% 1215 / 1811
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_filesystem 67.2% 3833 / 5704
ironclaw_memory 69.2% 773 / 1117
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.39% 1685 / 2265
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.67% 958 / 1283
ironclaw_extractors 74.72% 538 / 720
ironclaw_first_party_extensions 77.66% 5400 / 6953
ironclaw_llm 78.31% 20258 / 25870
ironclaw_product_context 78.57% 11 / 14
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_wasm_product_adapters 80.71% 1448 / 1794
ironclaw_reborn_openai_compat 81.16% 978 / 1205
ironclaw_memory_native 81.22% 3205 / 3946
ironclaw_secrets 82.7% 2791 / 3375
ironclaw_wasm 82.72% 996 / 1204
ironclaw_events 82.86% 1765 / 2130
ironclaw_auth 83.87% 3078 / 3670
ironclaw_reborn_config 84.33% 1814 / 2151
ironclaw_processes 84.44% 993 / 1176
ironclaw_turns 84.66% 13447 / 15884
ironclaw_common 84.85% 1490 / 1756
ironclaw_host_api 85.17% 2664 / 3128
ironclaw_product_workflow 85.56% 10836 / 12665
ironclaw_projects 85.92% 659 / 767
ironclaw_threads 86.04% 4234 / 4921
ironclaw_network 86.12% 670 / 778
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_product_adapters 86.98% 3207 / 3687
ironclaw_reborn_identity 87.03% 557 / 640
ironclaw_skills 87.58% 4470 / 5104
ironclaw_hooks 87.75% 9917 / 11302
ironclaw_product_adapter_registry 88.06% 531 / 603
ironclaw_reborn_traces 88.19% 11946 / 13546
ironclaw_host_runtime 88.41% 17188 / 19442
ironclaw_reborn_composition 89.01% 75922 / 85292
ironclaw_extensions 89.03% 2864 / 3217
ironclaw_approvals 89.24% 1584 / 1775
ironclaw_runner 89.28% 16693 / 18697
ironclaw_conversations 90.33% 3120 / 3454
ironclaw_event_streams 90.82% 1009 / 1111
ironclaw_loop_support 92.51% 14752 / 15947
ironclaw_resources 92.83% 4736 / 5102
ironclaw_attachments 93.06% 630 / 677
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.62% 2511 / 2682
ironclaw_agent_loop 94.66% 8771 / 9266
ironclaw_safety 94.88% 3671 / 3869
ironclaw_first_party_extension_ports 95.24% 3343 / 3510
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Make failed forced compaction recoverable after tool results while preserving cancellation semantics so the post-tool model turn can complete.
Stats: 6 findings (from 9 raw, 6 after dedup) across 4 files. Reviewers run: security, bugs, performance, tests, conventions, local-patterns, maintainability, approach. Reviewers failed: none. Body-only: 0.

Security

  1. Medium Timeout recovery leaves compaction inference running beside model call (crates/ironclaw_agent_loop/src/executor/prompt.rs:530-542, confidence 94) — anchor: crates/ironclaw_agent_loop/src/executor/prompt.rs:534
    When the outer compaction timeout wins, the pinned future is dropped and the new recovery path immediately continues. In production, GuardedSystemInferencePort has already spawned a worker whose JoinHandle is then detached, so its model request and post-accounting can remain active for up to another deadline while the normal model turn starts. Repeated timeout recoveries can accumulate concurrent model work and exhaust gateway or connection-pool capacity.
    Fix: Make the compaction inference worker abort-safe and await its termination before returning the recoverable timeout outcome.
  2. High SecurityRejected compaction failures bypass the safety gate (crates/ironclaw_agent_loop/src/executor/prompt.rs:519-528, confidence 95) — anchor: crates/ironclaw_agent_loop/src/executor/prompt.rs:519
    SecurityRejected is emitted when transcript input contains injection or secret-leak markers, but this branch routes it through defer_compaction and sends the already-built prompt to the next model turn. Newly returned malicious tool content can therefore reach the model unsanitized, enabling prompt-injection tool calls or secret disclosure.
    Fix: Distinguish input-security rejection from sanitized-summary rejection; block or sanitize the former and only recover from safe operational/output failures.

Tests

  1. Medium Cancellation test does not verify deferred state persistence (crates/ironclaw_agent_loop/src/executor/prompt.rs:653-665, confidence 95) — anchor: crates/ironclaw_agent_loop/src/executor/prompt.rs:659
    defer_compaction mutates the force flag and deferred watermark before the cancellation boundary writes the Final checkpoint. The cancellation test only checks the cancelled exit and prompt count, so it would pass if those mutations were lost from the staged checkpoint.
    Fix: Add executor coverage for compaction-failure cancellation that verifies the Final checkpoint persists the cleared force flag and expected last_deferred watermark.
  2. Medium Scope the summary-history assertion to the tested turn (tests/integration/http_matcher.rs:137-143, confidence 98) — anchor: tests/integration/http_matcher.rs:143
    This test submits three seed turns before calling the full-history assert_conversation_history_role_contains. The integration-test contract says full-history assertions are safe only for single-turn harnesses; multi-turn tests must use a baseline captured with history_len() and a *_since assertion. Earlier summaries can therefore contaminate this regression check.
    Fix: Capture history_len() before the fetch turn and use a baseline-aware role-filtered history assertion, adding that helper if necessary.

Local Patterns

  1. Low Remove reference to unavailable design document (crates/ironclaw_reborn_composition/src/projection/tests/failure_explanation.rs:351-353, confidence 100) — anchor: crates/ironclaw_reborn_composition/src/projection/tests/failure_explanation.rs:351
    The explanatory comment links to .issue-work/5838-compaction-robustness-design.md, but that file is not tracked or present in the repository, leaving readers with a dead navigation path for the test's rationale.
    Fix: Replace the path with a concise invariant-based explanation or link to a committed design/contract document.

Maintainability

  1. Low Keep milestone matching inside the assertion layer (tests/integration/support/builder.rs:634-640, confidence 75) — anchor: tests/integration/support/builder.rs:634
    The new public loop_milestones accessor exposes raw LoopHostMilestone records so the integration test can reimplement milestone matching directly. This leaks the sink representation into callers and separates this check from the existing assertion API.
    Fix: Make the raw accessor private to the support module and add a focused assertion helper in tests/integration/support/assertions.rs, analogous to assert_turn_event_recorded, so tests depend on behavior rather than the sink's record shape.

Comment thread crates/ironclaw_agent_loop/src/executor/prompt.rs Outdated
Comment thread crates/ironclaw_agent_loop/src/executor/prompt.rs
Comment thread crates/ironclaw_agent_loop/src/executor/prompt.rs
Comment thread tests/integration/http_matcher.rs
Comment thread crates/ironclaw_reborn_composition/src/projection/tests/failure_explanation.rs Outdated
Comment thread tests/integration/support/builder.rs Outdated
henrypark133 and others added 4 commits July 10, 2026 12:53
…tion checkpoint

Strengthen compaction_failure_cancellation_skips_explanation_and_returns_cancelled
to decode the Final checkpoint payload and assert force_compact_on_next_iteration
is cleared and the deferred watermark is persisted, not just the in-memory
Cancelled exit reason.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…access

Full-history conversation asserts are unsafe outside single-turn harnesses
(CLAUDE.md); add role-scoped *_since baseline variants and use them after
3 seed turns. Restrict loop_milestones() to the support module and add a
named assert_compaction_failed helper instead of pattern-matching raw
LoopHostMilestoneKind at the test call site.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…hed inference workers

await_compaction_with_cancellation raced an outer tokio::time::sleep
against the compaction future using the same deadline_ms already
enforced by ModelGatewayBackedSystemInferencePort's inner timeout. When
the outer race won first, the future was dropped, detaching the
GuardedSystemInferencePort worker it had spawned. The inner timeout
already surfaces as InferenceFailed -> compaction_failed_continue with
identical observable behavior, so the outer race added risk with no
benefit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… note

Review finding on #5895: the comment referenced .issue-work/, a local
untracked file other readers cannot open.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 20:07
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5895 July 10, 2026 20:07 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Make non-cancellation compaction failures recoverable so post-tool model turns continue while preserving cancellation semantics.

Stats: 2 new findings (from 7 raw, 4 after overlap dedup, 2 after same-line and existing-thread suppression) across 2 files. Reviewers run: security, bugs, performance, tests, conventions, local-patterns, maintainability, approach. Reviewers failed: none. Body-only: 0.

The security lane reproduced an existing resolved current-head thread and was not reposted.

Performance

  1. Medium Removing outer timeout leaves compaction DB work unbounded (crates/ironclaw_agent_loop/src/executor/prompt.rs:494-498, confidence 92) — anchor: crates/ironclaw_agent_loop/src/executor/prompt.rs:568

    The inner timeout only wraps model inference, but compact_loop_context also performs transcript reads, validation, scanning, and summary persistence. A stalled database or persistence call can now block the agent loop indefinitely; the removed outer timeout previously bounded the entire compaction operation.

Tests

  1. Medium Scope compaction milestone assertion to the tested turn (tests/integration/http_matcher.rs:133-135, confidence 95) — anchor: tests/integration/CLAUDE.md:202

    This multi-turn test asserts against assert_compaction_failed, which scans milestones from harness construction and therefore includes the three seed turns. The assertion can pass because of a prior compaction failure rather than the forced compaction in the turn under test. Capture a milestone baseline immediately before fetch items and orders and assert only that slice.

Comment thread crates/ironclaw_agent_loop/src/executor/prompt.rs
Comment thread tests/integration/http_matcher.rs Outdated
henrypark133 and others added 2 commits July 10, 2026 13:45
assert_compaction_failed scanned all milestones since harness construction, so a multi-turn test (3 seed turns then the turn under test) could match a stale milestone from an earlier turn. Add milestone_len()/assert_compaction_failed_since mirroring the history_len()/*_since pattern; the old full-scan variant had one caller and is replaced.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Amend two Design-lock rows: compaction errors are now deferred-continue
rather than terminal CompactionUnavailable (#5838), and the wall-clock
deadline is scoped to the inference call inside SystemInferencePort —
store phases follow the loop's uniform store-I/O semantics, matching
step 7's existing language.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 20:47
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5895 July 10, 2026 20:48 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions github-actions Bot added the scope: docs Documentation label Jul 10, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/reborn/2026-05-26-context-compaction.md`:
- Line 797: Update the failure state machine section around the descriptions of
InvalidCutPoint, InputTooLarge, InjectionDetected, LeakDetected,
InferenceFailed, and PersistenceFailed to remove the obsolete terminal
LoopFailureKind::CompactionUnavailable behavior. Document that these failures
emit CompactionFailed, record a deferred watermark, and continue the candidate
prompt in the same iteration; retain only Cancelled as terminal, consistent with
the executor path.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 5fe69532-ebef-45a0-acf7-e12556de7dbb

📥 Commits

Reviewing files that changed from the base of the PR and between 1458846 and 7487f96.

📒 Files selected for processing (4)
  • docs/reborn/2026-05-26-context-compaction.md
  • tests/integration/http_matcher.rs
  • tests/integration/support/assertions.rs
  • tests/integration/support/builder.rs

Comment thread docs/reborn/2026-05-26-context-compaction.md
Non-terminal compaction failures were silent: no durable event and no
log at emission. Adds a warn! with reason_kind + task_id (closed-vocab,
safe) at the compaction_failed_continue funnel so operators can see the
failure; durable RuntimeEvent surfacing is deferred to Slice C.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 21:59
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5895 July 10, 2026 21:59 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_agent_loop/src/executor/prompt.rs`:
- Around line 617-621: Change the recovered compaction failure log in the
executor path from tracing::warn! to tracing::debug!, preserving the task_id,
reason_kind, and existing message; the failure is already reported through
LoopProgressEvent::CompactionFailed.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 682f0b06-bd35-408d-9456-007873b6b137

📥 Commits

Reviewing files that changed from the base of the PR and between 7487f96 and ea25cb3.

📒 Files selected for processing (1)
  • crates/ironclaw_agent_loop/src/executor/prompt.rs

Comment thread crates/ironclaw_agent_loop/src/executor/prompt.rs Outdated
Review feedback on #5895: keep background loop diagnostics at debug
per the repo logging convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 22:17
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5895 July 10, 2026 22:17 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Review follow-up on #5895: the failure state machine (§10) and three
procedural walkthroughs still described the pre-#5838 terminal
CompactionUnavailable behavior. All compaction-path references now
match the deferred-continue contract; the Phase-2 GoalRefresh section
is untouched (different stage, unchanged behavior).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 22:20
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5895 July 10, 2026 22:20 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/reborn/2026-05-26-context-compaction.md`:
- Around line 985-997: The failed-compaction recovery path must check
cancellation before resuming inference. In the non-terminal error handling
branch for CompactionFailed, after clearing force_compact_on_next_iteration and
before continuing the candidate prompt, add an explicit cancellation checkpoint
that exits terminally if cancellation is requested.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 4a7ec146-d4e6-4467-a15e-9a1ab0b01e4d

📥 Commits

Reviewing files that changed from the base of the PR and between 0a2e98e and d30ef62.

📒 Files selected for processing (1)
  • docs/reborn/2026-05-26-context-compaction.md

Comment thread docs/reborn/2026-05-26-context-compaction.md

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5895 — d30ef622 Deployed Jul 10, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Run fails with context compaction error despite successful tool execution

2 participants