Skip to content

test(channels): complete delivery variant evidence - #6885

Merged
serrrfirat merged 7 commits into
mainfrom
codex/ws8-delivery-evidence
Jul 30, 2026
Merged

serrrfirat merged 7 commits into
mainfrom
codex/ws8-delivery-evidence

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Complete WS8's proven residual delivery evidence without changing production behavior: Telegram forum-topic whole-path coverage now proves message_thread_id survives normalization through coordinated provider delivery.
  • Make Slack DM, shared-channel, and threaded-channel destinations mechanically citable from the journey inventory, including opaque conversation ID, optional thread anchor, important content, and exactly-once count.
  • Extend the existing Telegram pairing journey with exact unthreaded chat delivery plus authenticated status read-back before and after durable unpair/re-pair.
  • Derive required variants from production outbound manifests and adapter thread-anchor support; no closed vendor destination enum is introduced.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Related #6524. Webhook ingress remains covered by #6828. Reconciled against open #6831, #6776, and #6364; none implements this caller-seam delivery evidence.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — both touched targets are clean locally with -D warnings; GitHub's fresh Clippy (all-features) gate also passed.
  • cargo build — not applicable: test-only change; both touched test targets compiled under Clippy and focused test execution.
  • Relevant tests pass: journey inventory (44), Slack thread, Telegram topic, Telegram pairing/status/unpair, and scheduled-trigger Slack DM/channel.
  • cargo test --features integration if database-backed or integration behavior changed — test-only coverage change; focused libSQL integration cases pass. PostgreSQL provisioning was attempted and failed before scenario execution because no Docker daemon is reachable.
  • Manual testing: not applicable; hermetic recording adapters assert provider wire evidence.
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review — pr-shepherd readiness pass completed; the full final diff was reviewed across correctness, security, architecture, test quality, and repository conventions.

Test Strategy

User behavior: Replies from Slack threads and Telegram forum topics return to the exact originating channel/thread/topic once; Telegram unthreaded chat pairing status and unpair are durably observable; scheduled Slack DM/channel delivery stays exact.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: tests/e2e/scenarios/test_journey_coverage.py mechanically validates production-supported outbound surfaces, opaque conversation IDs, optional thread anchors, exact count, and reachable named caller assertions.
  • Reborn integration: Slack threaded reply, Telegram topic reply, and Telegram chat pairing/status/unpair/re-pair journeys in reborn_integration_extension_delivery.
  • Recorded fixture: Not applicable: no model/provider selection behavior changed; network recording adapters are asserted directly.
  • Browser E2E: Not applicable: WebUI delivery behavior did not change; the inventory retains the existing exact durable WebUI response proof.
  • Backend or runtime: libSQL focused cases pass. PostgreSQL case was attempted but cannot provision because this host has no reachable Docker/Colima daemon; this is a local environment limitation, not a skip in the test. GitHub's integration matrix is the authoritative PostgreSQL-capable readiness gate.
  • Live canary: Not applicable: provider requests are hermetically captured and credentials remain host-injected.

What the tests prove:

  • Telegram topic: durable assistant state, normalized external actor and paired subject account, exact chat -1008675309, topic 77, important content, provider-issued cleanup message ID, and exactly one final reply.
  • Slack thread: durable assistant state, normalized actor/configured subject account, exact channel C777, thread 1710000200.000050, important content, injected credential, and exactly one final reply.
  • Slack DM/channel: existing restart proof remains citable as exact one-message delivery to D-TRIGGER-DEFAULT and C-TRIGGER-OVERRIDE.
  • Telegram chat/pairing: exact unthreaded delivery to 515151, authenticated status connected=true, durable unpair, status connected=false, and fresh thread after re-pair.
  • Completeness: production outbound manifests currently expose Slack and Telegram; optional thread_anchor support mechanically requires both threaded and unthreaded evidence.

Commands run:

  • bash scripts/codebase-graph.sh status → graph missing; immediately used targeted live-code/contracts per guidance.
  • tests/e2e/.venv/bin/pytest -q tests/e2e/scenarios/test_journey_coverage.py → 44 passed.
  • cargo test -q -p ironclaw_reborn_integration_tests --test reborn_integration_extension_delivery 'slack_final_reply_flows_through_the_real_delivery_coordinator::case_1_libsql' -- --exact --nocapture → 1 passed.
  • cargo test -q -p ironclaw_reborn_integration_tests --test reborn_integration_extension_delivery 'telegram_update_becomes_a_turn_and_a_coordinated_reply::case_1_libsql' -- --exact --nocapture → 1 passed.
  • cargo test -q -p ironclaw_reborn_integration_tests --test reborn_integration_extension_delivery 'unbound_telegram_actor_pairs_via_web_minted_code_then_turns_attribute_to_the_paired_user::case_1_libsql' -- --exact --nocapture → 1 passed.
  • cargo test -q -p ironclaw_reborn_composition --test trigger_poller_e2e scheduled_trigger_results_reach_exact_slack_targets_once_across_restart -- --exact --nocapture → 1 passed.
  • cargo clippy -p ironclaw_reborn_integration_tests --test reborn_integration_extension_delivery -- -D warnings → clean.
  • cargo clippy -p ironclaw_reborn_composition --test trigger_poller_e2e -- -D warnings → clean.
  • cargo fmt --all -- --check and git diff --check → clean.
  • git merge origin/main → conflicts in journey_types.py and test_journey_coverage.py resolved by preserving both WS12 live/browser evidence and WS8 exact delivery-address evidence; merge commit 09fb55723.
  • git merge origin/main at 6535edc96 → WS9 state-machine and IronHub changes merged without additional file conflicts; merge commit 7aeb1aa27.
  • tests/e2e/.venv/bin/pytest -q tests/e2e/scenarios/test_journey_coverage.py tests/e2e/scenarios/test_product_surface_coverage.py → 51 passed on the merged head.
  • tests/e2e/.venv/bin/pytest -q tests/e2e/scenarios/test_state_machine_coverage.py → 12 passed on the merged head.
  • tests/e2e/.venv/bin/pytest -q scripts/ci/test_ws12_suite_shards.py scripts/ci/test_ws12_workflow_contracts.py scripts/ci/test_reborn_changed_coverage.py → 19 passed, 36 subtests passed on the merged head.
  • Maintainer-quality review of every changed file and the complete origin/main...HEAD diff → eight valid automated review findings fixed; no remaining known correctness, security, architecture, or test-quality findings.
  • GitHub CI on prior merge head 09fb55723 → 56 passed, 0 failed, 0 pending, 6 policy-selected skips. Fresh CI on current merge head 7aeb1aa27 is in progress.
  • PostgreSQL matrix attempt reached libSQL success, then failed during PostgreSQL provisioning with StorageMode::Postgres requires a reachable Docker daemon / container connection failure.

Railway follow-through:

  • The private Railway preview initially reported only Deployment failed on two consecutive heads, including a clean current-main rebase. Its third deployment on the final head succeeded, so no readiness blocker remains. If the failure recurs, the linked logs require an authenticated Railway account with membership in project 76f27460-1806-4eed-b168-e9b150113141 (service 3d19f4f3-6060-4143-8faf-ef546ed6fb97, environment fdfae749-b48a-43b0-a8b1-89b2f9aab672).

Sabotage baseline:

  • Removing Telegram topic inventory evidence failed with threaded delivery is implemented but lacks exact evidence.
  • Removing Telegram message_thread_id propagation failed the whole-path topic assertion.
  • Removing Slack thread_ts propagation failed exact threaded delivery evidence.
  • Duplicating a Telegram final reply failed the exactly-once count.
  • Omitting a named Slack caller assertion failed mechanical citability.
  • Replacing the structured inventory destination with C-FAKE-STRUCTURED-PROBE failed because it was not bound to expected_conversation_id.
  • Changing the cited Rust helper's expected_count from 1 to 2 failed because it no longer matched the inventory.

Security Impact

None. Test-only changes preserve authenticated pairing routes, host-side credential injection, typed actor/subject identity assertions, and provider mediation.

Reborn Trust-Boundary Checklist

N/A: no production policy/evidence/trust-bearing types, prompt ingress, hashes, variants, serde fields, queues, errors, or runtime boundary names changed. Tests explicitly assert the existing authenticated pairing caller, normalized actor/account, credential injection, durable state, and mediated provider evidence.

Database Impact

None. No migrations or schema changes. libSQL integration evidence passes; PostgreSQL execution is blocked by unavailable Docker on this host and remains visible above.

Blast Radius

Test inventory metadata and the Slack/Telegram delivery integration fixtures. The fail-loud inventory parser intentionally permits a cited test or its direct _impl delegate; a source-layout rename without updated evidence will fail the contract test.

Rollback Plan

Revert commits 0d469e935, a52dd3ce1, d34f29c75, 89427d61c, and 88724e031; production runtime behavior and persistent schemas are unchanged.

Review Follow-Through

Checkbox-to-test/PR map:

Reviewer judgment requested on the intentionally narrow mechanical rule: delivery assertions must be called by the cited Rust test or its direct _impl delegate.

Automated review follow-through:

  • ironloopai: “Delivery metadata is not tied to the cited assertion” — fixed in 89427d61c. Each named Rust helper now binds expected_conversation_id, expected_thread_anchor, and expected_count; the inventory verifies those exact bindings gate the provider evidence and mutation count.
  • CodeRabbit: “CREDENTIAL_INJECTION looks unsupported by the cited pairing test” — fixed in d34f29c75; the pairing row now claims only evidence asserted by its cited journey.
  • CodeRabbit: “Substring grep for thread_anchor will over-trigger the threaded-evidence requirement” — fixed in d34f29c75; requirements now derive from channel.presentation.supports_threads in production manifests.
  • CodeRabbit: “Literal-preserving extraction allows comments to spoof evidence” — fixed in a52dd3ce1; comments are masked even when string literals are preserved, with a regression test.
  • CodeRabbit: “Unthreaded delivery accepts a present wrong-typed thread anchor” — fixed in a52dd3ce1; Telegram chat, Slack DM, and Slack channel evidence now asserts direct absence of the provider thread/topic field.
  • CodeRabbit: “Threading requirement fails open when supports_threads is absent” — fixed in 0d469e935; missing, misspelled, and non-boolean declarations each fail a sabotage regression.
  • CodeRabbit: “Manifest id is indexed unguarded” — fixed in 0d469e935; a malformed-manifest regression requires a non-empty string ID and path-specific diagnostic.
  • CodeRabbit/Ruff: “Use contextlib.suppress for delegate-body lookup” — fixed in 0d469e935; behavior remains covered by the complete inventory suite.

Review track: A (tests/chore)

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@coderabbitai

coderabbitai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes

    • Improved Slack delivery verification by distinguishing DM vs per-trigger override evidence, tightening message matching, and enforcing exact delivery counts.
    • Improved Telegram delivery verification by switching to structured thread/topic evidence for message and cleanup validation, and updating pairing/unpair/repair state checks.
    • Strengthened result-delivery assertions to confirm evidence is recorded exactly once for the expected destination.
  • Tests

    • Expanded end-to-end journey coverage for Slack and Telegram with typed delivery-address evidence and exact mutation/count assertions.
    • Strengthened journey coverage completeness gates for threaded vs unthreaded delivery variants driven by capability manifests.

Walkthrough

Slack and Telegram E2E coverage now models exact delivery addresses, validates structured provider payloads and pairing state, and enforces source-citable evidence for outbound channel surfaces. Scheduled Slack tests separately verify default and override targets exactly once.

Changes

External delivery evidence

Layer / File(s) Summary
Evidence contracts and journey declarations
tests/e2e/journey_types.py, tests/e2e/journey_cases.py
Adds typed delivery-address evidence and assigns exact counts, thread identifiers, and assertion hooks to Slack and Telegram journeys.
Provider delivery flow assertions
tests/integration/extension_delivery.rs
Validates threaded Slack, Telegram topic, and unthreaded Telegram payloads as JSON, propagates vendor actor identity, and checks pairing status before and after unpairing.
Whole-path evidence coverage gate
tests/e2e/scenarios/test_journey_coverage.py
Checks that delivery evidence is citable from Rust helpers and covers each outbound channel surface and supported threading mode.
Scheduled Slack target evidence
crates/ironclaw_reborn_composition/tests/trigger_poller_e2e.rs
Verifies exactly one unthreaded message reaches the default DM and per-trigger override channel.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant SlackWebhook
  participant ExtensionDelivery
  participant SlackProvider
  SlackWebhook->>ExtensionDelivery: threaded app_mention event
  ExtensionDelivery->>SlackProvider: postMessage with channel and thread_ts
  SlackProvider-->>ExtensionDelivery: structured delivery payload
Loading
sequenceDiagram
  participant TelegramWebhook
  participant ExtensionDelivery
  participant TelegramProvider
  participant PairingRouter
  TelegramWebhook->>ExtensionDelivery: forum-topic or chat message
  ExtensionDelivery->>TelegramProvider: sendMessage with chat and topic identifiers
  TelegramProvider-->>ExtensionDelivery: structured delivery payload
  PairingRouter->>PairingRouter: report pairing status before and after unpair
Loading

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The Conventional Commits-style title matches the PR’s main test-only delivery evidence change.
Description check ✅ Passed The description closely follows the template and fills the required sections with specific validation, test strategy, impact, and rollback details.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app

railway-app Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6885 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 30, 2026 at 11:48 am

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6885 July 29, 2026 22:44 Destroyed
@github-actions github-actions Bot added size: XS < 10 changed lines (excluding docs) risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 29, 2026
@ironloopai

ironloopai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6885

🟢 Completed · Review submitted

1 actionable findings →

Reviewed the complete trusted comparison. The added delivery journeys exercise the intended Slack and Telegram paths, but the new mechanical evidence inventory does not verify that its declared addresses match the cited Rust assertions.

Automatic · PR opened · attempt 1 of 3 · completed in 1m 58s

Run details
  • Repository: nearai/ironclaw
  • Base: main at bed3f68
  • Head: codex/ws8-delivery-evidence at 23eb40b
  • Created: Jul 29, 2026, 10:49 PM UTC
  • Updated: Jul 29, 2026, 10:51 PM UTC
  • Run: 19c8f12a-6811-4085-897d-694c84a4f8b5
  • Latest attempt: 1 · Completed · 856ab932-961a-4714-81ae-feee4a6e5dd8

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #6885

💬 1 finding

Reviewed the complete trusted comparison. The added delivery journeys exercise the intended Slack and Telegram paths, but the new mechanical evidence inventory does not verify that its declared addresses match the cited Rust assertions.

Findings

  1. 🟡 Low · Delivery metadata is not tied to the cited assertion — tests/e2e/scenarios/test_journey_coverage.py:366-383
    Details are attached to the relevant diff.
Validation and technical details
  • Inspected the full trusted base/head PR delta via the merge-base comparison: 5 files, 438 insertions, 49 deletions.
  • Reviewed all changed files and surrounding journey inventory, Rust delivery helpers, pairing flow, provider request capture, and composition test guidance.
  • git diff --check refs/ironloop/base...refs/ironloop/head passed.
  • Focused tests could not be executed in this checkout: the bundled pytest environment is absent, system Python has no pytest module, and cargo is unavailable.
  • Base: main
  • Head: codex/ws8-delivery-evidence at 23eb40b
  • Run: 19c8f12a-6811-4085-897d-694c84a4f8b5

Comment thread tests/e2e/scenarios/test_journey_coverage.py
@github-actions

github-actions Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.5% (316868 / 370600 lines)
  floor:    85.54% (tolerance 0.5pp -> effective floor 85.04%)
  denominator: 370600 lines now vs 368757 at floor capture (+1843 lines, +0.5%) — not a material change

RATCHET PASS: ironclaw_runner
  observed: 86.87% (14674 / 16892 lines)
  floor:    86.87% (tolerance 0.5pp -> effective floor 86.37%)
  floor_covered_lines: 14669 (tolerance 20 lines -> effective floor 14649)
  denominator: 16892 lines now vs 16887 at floor capture (+5 lines, +0.03%) — not a material change

RATCHET PASS: ironclaw_processes
  observed: 87.98% (5833 / 6630 lines)
  floor:    88.07% (tolerance 0.5pp -> effective floor 87.57%)
  floor_covered_lines: 5839 (tolerance 20 lines -> effective floor 5819)
  denominator: 6630 lines now vs 6630 at floor capture (+0 lines, +0%) — not a material change

RATCHET PASS: ironclaw_turns
  observed: 86.24% (9522 / 11041 lines)
  floor:    86.21% (tolerance 0.5pp -> effective floor 85.71%)
  floor_covered_lines: 9453 (tolerance 20 lines -> effective floor 9433)
  denominator: 11041 lines now vs 10965 at floor capture (+76 lines, +0.69%) — not a material change

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.5% — 316868 / 370600 lines

Per-crate breakdown (60 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_process_sandbox 34.38% 120 / 349
ironclaw_host_ingress 42.5% 17 / 40
ironclaw_event_projections 43.51% 684 / 1572
ironclaw_libsql_runtime 59.91% 127 / 212
ironclaw_observability 61.54% 16 / 26
ironclaw_authorization 63.02% 610 / 968
ironclaw_memory 64.41% 959 / 1489
ironclaw_telegram_v2_adapter 72.07% 756 / 1049
ironclaw_trust 73.21% 664 / 907
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_extractors 74.72% 538 / 720
ironclaw_capabilities 74.95% 2882 / 3845
ironclaw_projects 76.48% 400 / 523
ironclaw_mcp 76.6% 779 / 1017
ironclaw_filesystem 77.24% 5966 / 7724
ironclaw_reborn_cli 77.89% 10795 / 13860
ironclaw_wasm 79.72% 735 / 922
ironclaw_llm 80.82% 23319 / 28854
ironclaw_memory_native 81.02% 3299 / 4072
ironclaw_auth 81.88% 6679 / 8157
ironclaw_first_party_extensions 82.38% 6682 / 8111
ironclaw_host_api 83.68% 9728 / 11625
ironclaw_events 83.8% 1536 / 1833
ironclaw_reborn_identity 83.8% 450 / 537
ironclaw_operator 84.41% 5561 / 6588
ironclaw_secrets 84.56% 2798 / 3309
ironclaw_skills 84.74% 4494 / 5303
ironclaw_network 85.09% 959 / 1127
ironclaw_reborn_config 85.23% 2101 / 2465
ironclaw_extension_host 85.24% 21126 / 24783
ironclaw_telegram_extension 85.47% 1300 / 1521
ironclaw_approvals 86.08% 1818 / 2112
ironclaw_triggers 86.11% 2803 / 3255
ironclaw_reborn_composition 86.17% 22072 / 25613
ironclaw_turns 86.24% 9522 / 11041
ironclaw_webui 86.37% 11171 / 12934
ironclaw_hooks 86.73% 9950 / 11473
ironclaw_runner 86.87% 14674 / 16892
ironclaw_common 87.04% 1773 / 2037
ironclaw_reborn_event_store 87.19% 1327 / 1522
ironclaw_extensions 87.45% 4927 / 5634
ironclaw_product 87.49% 21354 / 24406
ironclaw_processes 87.98% 5833 / 6630
ironclaw_reborn_openai_compat 88% 3724 / 4232
ironclaw_reborn_traces 88.13% 11987 / 13601
ironclaw_threads 88.57% 4943 / 5581
ironclaw_host_runtime 89.1% 20638 / 23163
ironclaw_slack_extension 89.3% 2028 / 2271
ironclaw_conversations 90.03% 3171 / 3522
ironclaw_resources 90.84% 4474 / 4925
ironclaw_loop_host 90.93% 17742 / 19512
ironclaw_event_streams 91.24% 1063 / 1165
ironclaw_attachments 93.06% 630 / 677
ironclaw_outbound 93.91% 4101 / 4367
ironclaw_agent_loop 94.42% 10424 / 11040
ironclaw_safety 95.22% 3941 / 4139
ironclaw_first_party_extension_ports 95.71% 3837 / 4009
ironclaw_runtime_policy 96.56% 814 / 843

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@serrrfirat
serrrfirat force-pushed the codex/ws8-delivery-evidence branch from 23eb40b to 9f81e64 Compare July 30, 2026 08:56
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6885 July 30, 2026 08:56 Destroyed
@serrrfirat
serrrfirat marked this pull request as ready for review July 30, 2026 09:14
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@ironloopai

ironloopai Bot commented Jul 30, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6885

⚫ Cancelled · Target changed

The target changed before this Run could finish.

Automatic · PR opened · attempt 0 of 3 · cancelled after <1s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 5923789
  • Head: codex/ws8-delivery-evidence at 9f81e64
  • Created: Jul 30, 2026, 9:19 AM UTC
  • Updated: Jul 30, 2026, 9:19 AM UTC
  • Run: 9f1ef000-7dc4-4c61-89a5-61a9014c5a56

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/e2e/journey_cases.py`:
- Around line 347-352: Remove ObservableAssertion.CREDENTIAL_INJECTION from the
assertions for the row citing
unbound_telegram_actor_pairs_via_web_minted_code_then_turns_attribute_to_the_paired_user_impl,
unless that integration test is updated to explicitly verify host-side token
substitution; keep the row’s assertions limited to behaviors covered by the
cited test.

In `@tests/e2e/scenarios/test_journey_coverage.py`:
- Around line 686-695: Update the threaded-evidence check in the journey
coverage test to inspect each adapter’s declared manifest or trait capability
rather than searching channel.rs text for “thread_anchor”. Resolve adapter
metadata through the existing surface/manifest discovery mechanism so valid
channels are not rejected due to hardcoded crate/file naming, and require
address.thread_anchor evidence only when that declared capability indicates
threaded delivery.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 45732a8b-5b71-4f55-9d31-eb824e9cf6bd

📥 Commits

Reviewing files that changed from the base of the PR and between 5923789 and 9f81e64.

📒 Files selected for processing (5)
  • crates/ironclaw_reborn_composition/tests/trigger_poller_e2e.rs
  • tests/e2e/journey_cases.py
  • tests/e2e/journey_types.py
  • tests/e2e/scenarios/test_journey_coverage.py
  • tests/integration/extension_delivery.rs

Comment thread tests/e2e/journey_cases.py
Comment thread tests/e2e/scenarios/test_journey_coverage.py Outdated
@serrrfirat
serrrfirat force-pushed the codex/ws8-delivery-evidence branch from 9f81e64 to b762361 Compare July 30, 2026 09:29
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6885 July 30, 2026 09:29 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/e2e/scenarios/test_journey_coverage.py`:
- Around line 399-447: Update _rust_function_body extraction in the assertion
validation path to mask Rust comments while preserving string literals before
applying the binding and gating regex checks. Ensure commented-out declarations
or comparisons cannot satisfy checks for expected_conversation_id,
expected_thread_anchor, or expected_count, while literals remain available for
value validation.

In `@tests/integration/extension_delivery.rs`:
- Around line 691-710: Update assert_telegram_chat_delivery_evidence so the
unthreaded predicate distinguishes a missing message_thread_id from a present
non-integer value. Replace the current get(...).and_then(Value::as_i64)
comparison with a direct absence check while retaining the
expected_thread_anchor declaration and comparison required by the coverage gate.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 5e7ef519-e66c-4fd6-b489-92714afa307f

📥 Commits

Reviewing files that changed from the base of the PR and between 9f81e64 and b762361.

📒 Files selected for processing (5)
  • crates/ironclaw_reborn_composition/tests/trigger_poller_e2e.rs
  • tests/e2e/journey_cases.py
  • tests/e2e/journey_types.py
  • tests/e2e/scenarios/test_journey_coverage.py
  • tests/integration/extension_delivery.rs

Comment thread tests/e2e/scenarios/test_journey_coverage.py
Comment thread tests/integration/extension_delivery.rs
@serrrfirat
serrrfirat force-pushed the codex/ws8-delivery-evidence branch from b762361 to 9944ea6 Compare July 30, 2026 09:54
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6885 July 30, 2026 09:54 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/e2e/scenarios/test_journey_coverage.py`:
- Around line 709-717: Update the outbound capability assertion in the test loop
over _production_channel_capabilities so supports_threads must be explicitly
declared for every surface, rather than using
capabilities.get("supports_threads") and failing open when absent or misspelled.
Preserve the threaded evidence check for declared true values, and add or update
regression coverage that executes this enforcing test command and verifies
missing declarations fail.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: f1dac8d5-5951-4091-96b7-8d86e13eaa3e

📥 Commits

Reviewing files that changed from the base of the PR and between b762361 and 9944ea6.

📒 Files selected for processing (5)
  • crates/ironclaw_reborn_composition/tests/trigger_poller_e2e.rs
  • tests/e2e/journey_cases.py
  • tests/e2e/journey_types.py
  • tests/e2e/scenarios/test_journey_coverage.py
  • tests/integration/extension_delivery.rs

Comment thread tests/e2e/scenarios/test_journey_coverage.py Outdated
@serrrfirat
serrrfirat force-pushed the codex/ws8-delivery-evidence branch from 9944ea6 to a52dd3c Compare July 30, 2026 10:06
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6885 July 30, 2026 10:06 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
tests/e2e/scenarios/test_journey_coverage.py (1)

179-242: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Ruff flags this scanner over the branch/statement budget.

PLR0912 (19 > 12) and PLR0915 (56 > 50) both point at Line 179 after the preserve_strings additions. Splitting the raw-string and quoted-string scanning into a small position-returning helper brings it back under budget without changing the masking contract.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/e2e/scenarios/test_journey_coverage.py` around lines 179 - 242,
Refactor _rust_code_without_comments_or_strings to extract raw-string and
quoted-string scanning into a small helper that returns the next position and,
when masking is required, updates the result while preserving newlines. Replace
the corresponding string branches with the helper call, retaining the existing
preserve_strings behavior and masking contract while reducing PLR0912 and
PLR0915 counts.

Source: Linters/SAST tools

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/e2e/journey_cases.py`:
- Around line 412-424: The _production_channel_capabilities function indexes
manifest["id"] without validating that the manifest declares an id, causing an
opaque KeyError. Add a named assertion or equivalent explicit validation before
indexing manifest["id"], with a diagnostic message identifying the missing
top-level id and manifest_path, while preserving the existing capability mapping
behavior.

In `@tests/e2e/scenarios/test_journey_coverage.py`:
- Around line 464-470: Replace the try/except AssertionError/pass around
_rust_function_body in the delegate loop with
contextlib.suppress(AssertionError), and add the contextlib import. Preserve the
existing behavior of appending reachable delegate bodies while ignoring lookup
failures.

---

Outside diff comments:
In `@tests/e2e/scenarios/test_journey_coverage.py`:
- Around line 179-242: Refactor _rust_code_without_comments_or_strings to
extract raw-string and quoted-string scanning into a small helper that returns
the next position and, when masking is required, updates the result while
preserving newlines. Replace the corresponding string branches with the helper
call, retaining the existing preserve_strings behavior and masking contract
while reducing PLR0912 and PLR0915 counts.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 99c55db7-d26d-4fc1-847a-32a7acba026c

📥 Commits

Reviewing files that changed from the base of the PR and between 9944ea6 and a52dd3c.

📒 Files selected for processing (5)
  • crates/ironclaw_reborn_composition/tests/trigger_poller_e2e.rs
  • tests/e2e/journey_cases.py
  • tests/e2e/journey_types.py
  • tests/e2e/scenarios/test_journey_coverage.py
  • tests/integration/extension_delivery.rs

Comment thread tests/e2e/journey_cases.py Outdated
Comment thread tests/e2e/scenarios/test_journey_coverage.py Outdated
@serrrfirat
serrrfirat marked this pull request as draft July 30, 2026 10:14
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6885 July 30, 2026 10:20 Destroyed
@serrrfirat
serrrfirat marked this pull request as ready for review July 30, 2026 10:42
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@ironloopai

ironloopai Bot commented Jul 30, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6885

🔴 Failed

GitHub request failed

IronLoop could not complete a required GitHub request.

Automatic · PR opened · attempt 1 of 3 · failed after 1m 55s

Failure details
  • Repository: nearai/ironclaw
  • Base: main at 6eb1931
  • Head: codex/ws8-delivery-evidence at 0d469e9
  • Created: Jul 30, 2026, 10:47 AM UTC
  • Updated: Jul 30, 2026, 10:49 AM UTC
  • Run: a973afa1-6673-4b69-aef6-6d14fcc4edb4
  • Latest attempt: 1 · Completed · e184a0c1-5403-49a0-9215-8b9cdec0db55
  • Failed during: GitHub writeback
  • Retryable: No
  • Failure: 50276f27-920a-4690-b373-04528829dc00

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
tests/e2e/scenarios/test_journey_coverage.py (1)

410-411: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Add failure messages to these two asserts.

Every other assert in _assert_delivery_address_is_citable carries an f"{case.case_id}: ..." diagnostic; these two bare asserts will surface as an opaque AssertionError with no indication of which required ObservableAssertion is missing or for which case.

🩺 Add diagnostics
-    assert ObservableAssertion.EXACT_DESTINATION in case.assertions
-    assert ObservableAssertion.EXACT_MUTATION_COUNT in case.assertions
+    assert ObservableAssertion.EXACT_DESTINATION in case.assertions, (
+        f"{case.case_id}: delivery address evidence requires EXACT_DESTINATION"
+    )
+    assert ObservableAssertion.EXACT_MUTATION_COUNT in case.assertions, (
+        f"{case.case_id}: delivery address evidence requires EXACT_MUTATION_COUNT"
+    )
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/e2e/scenarios/test_journey_coverage.py` around lines 410 - 411, The two
assertions in _assert_delivery_address_is_citable must include case-specific
diagnostic failure messages. Add f"{case.case_id}: ..." messages identifying the
missing EXACT_DESTINATION and EXACT_MUTATION_COUNT ObservableAssertion
respectively, matching the diagnostics used by the surrounding assertions.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@tests/e2e/scenarios/test_journey_coverage.py`:
- Around line 410-411: The two assertions in _assert_delivery_address_is_citable
must include case-specific diagnostic failure messages. Add f"{case.case_id}:
..." messages identifying the missing EXACT_DESTINATION and EXACT_MUTATION_COUNT
ObservableAssertion respectively, matching the diagnostics used by the
surrounding assertions.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 1009d890-759b-41ad-bea2-d6cefa973baa

📥 Commits

Reviewing files that changed from the base of the PR and between a52dd3c and 0d469e9.

📒 Files selected for processing (2)
  • tests/e2e/journey_cases.py
  • tests/e2e/scenarios/test_journey_coverage.py

…idence

# Conflicts:
#	tests/e2e/journey_types.py
#	tests/e2e/scenarios/test_journey_coverage.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
tests/e2e/scenarios/test_journey_coverage.py (1)

180-239: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Do not treat preserved string contents as executable Rust evidence.

With preserve_literals=True, _assert_rust_assignment and the subsequent regexes scan raw string payloads as well as code. A raw/error string containing let expected_conversation_id = ..., the expected anchor comparison, and matching.count() can therefore satisfy the gate without executable delivery assertions. Use token/AST-aware matching, or keep code masked and extract literal RHS values separately; add a regression containing code-looking strings.

As per path instructions, delivery coverage must be backed by executable, cited journey evidence.

Also applies to: 394-460, 479-492

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/e2e/scenarios/test_journey_coverage.py` around lines 180 - 239, Update
_assert_rust_assignment and the subsequent delivery-coverage regex checks so
preserved string contents cannot satisfy executable Rust evidence; keep code
masked for token matching and extract literal RHS values separately when needed.
Apply the same protection to the related checks and add a regression fixture
containing code-looking raw/error strings, while requiring delivery coverage to
come from executable, cited journey assertions.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@tests/e2e/scenarios/test_journey_coverage.py`:
- Around line 180-239: Update _assert_rust_assignment and the subsequent
delivery-coverage regex checks so preserved string contents cannot satisfy
executable Rust evidence; keep code masked for token matching and extract
literal RHS values separately when needed. Apply the same protection to the
related checks and add a regression fixture containing code-looking raw/error
strings, while requiring delivery coverage to come from executable, cited
journey assertions.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 90ba0400-6968-41a8-92dd-e3a3f72b53ef

📥 Commits

Reviewing files that changed from the base of the PR and between 0d469e9 and 09fb557.

📒 Files selected for processing (3)
  • tests/e2e/journey_cases.py
  • tests/e2e/journey_types.py
  • tests/e2e/scenarios/test_journey_coverage.py

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6885 July 30, 2026 11:05 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6885 July 30, 2026 11:38 Destroyed
@serrrfirat
serrrfirat merged commit 458c9dd into main Jul 30, 2026
62 checks passed
@serrrfirat
serrrfirat deleted the codex/ws8-delivery-evidence branch July 30, 2026 12:10
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
* test(channels): complete delivery variant evidence

* test(channels): bind inventory to provider assertions

* test(channels): align inventory with declared evidence

* test(channels): harden exact evidence parsing

* test(channels): fail closed on manifest metadata

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6885 — 7aeb1aa2 Deployed Jul 30, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XS < 10 changed lines (excluding docs)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant