Skip to content

fix(github): reduce PR-prioritization tool churn - #6953

Closed
serrrfirat wants to merge 14 commits into
mainfrom
codex/github-tool-efficiency-regression
Closed

serrrfirat wants to merge 14 commits into
mainfrom
codex/github-tool-efficiency-regression

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 31, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Compact GitHub pull-request list and issue/PR search results before they enter model context, while preserving pagination and the fields needed for prioritization.
  • Route self-scoped GitHub searches through the relationship the user actually named, using @me directly without an unnecessary identity lookup.
  • Preserve the original durable result_read continuation reference when inline preview text is suppressed.
  • Promote the captured 65-call run into a minimal one-call authored-PR regression and add a reusable run-regression promotion skill.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Related #6524; Related #5838

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — not run; the focused GitHub WASM crate clippy check passed.
  • cargo build — covered by the focused test builds below; a standalone full-workspace build was not run.
  • Relevant tests pass: GitHub WASM unit suite (58), first-party extension suite (152), host-runtime GitHub caller tests (2), recorded behavior contract, continuation unit contract, and composition caller-path contract.
  • cargo test --features integration if database-backed or integration behavior changed — not applicable; no database behavior changed and the package-qualified recorded integration test passed.
  • Manual testing: imported and measured the supplied run artifact as RED evidence (65 calls: 17 GitHub and 34 result_read), then verified the minimized fixture enforces a one-call authored route without overfitting other self relationships.
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review — PR remains draft pending normal CI/review.

Test Strategy

User behavior:

A user explicitly asking for open nearai/ironclaw PRs they authored receives a prioritized answer after one author: "@me" search. Large provider responses remain bounded, and suppressed preview text cannot destroy the durable continuation authority needed for subsequent reads.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: GitHub WASM response compaction tests; continuation metadata contract.
  • Reborn integration: package-qualified recorded model behavior contract.
  • Recorded fixture: github_open_pr_priority.json, minimized to an explicit authored-by-me request and one self-scoped search rather than copying its ambiguous 65-call trajectory.
  • Browser E2E: Not applicable: no browser or frontend behavior changed.
  • Backend or runtime: host-runtime GitHub WASM contracts and composition caller-path continuation contract.
  • Live canary: Not applicable: deterministic provider fixtures cover this change without live credentials or network variability.

What the tests prove:

  • Realistic 100-item GitHub list/search payloads are compacted below the model-context budget while preserving items, identity, ranking fields, and pagination.
  • An explicit authored-by-me request uses one search_issues_pull_requests call with author: "@me", while the prompts preserve distinct authored, assigned, involved, and review-requested relationships.
  • A credential-like rejected preview remains hidden while the original result reference and continuation metadata survive through the production composition wrapper.

Commands run:

python3 scripts/test-import-reborn-run-artifact.py
bash scripts/ci/check-reborn-qa-fixtures.sh
python3 /Users/firatsertgoz/.codex/skills/.system/skill-creator/scripts/quick_validate.py .claude/skills/promote-run-regression
cargo fmt --all -- --check
git diff --check
cargo test --manifest-path crates/ironclaw_first_party_extensions/assets/github/wasm-src/Cargo.toml
cargo clippy --manifest-path crates/ironclaw_first_party_extensions/assets/github/wasm-src/Cargo.toml --all-targets -- -D warnings
cargo test -p ironclaw_first_party_extensions
cargo test -p ironclaw_host_runtime --test github_wasm_runtime_contract host_runtime_services_compact_github
CARGO_PROFILE_TEST_DEBUG=0 cargo test -p ironclaw_reborn_integration_tests --test reborn_qa_recorded_behavior contract_github_open_pr_priority_uses_scoped_search -- --exact
CARGO_PROFILE_TEST_DEBUG=0 cargo test -p ironclaw_reborn_composition standalone_runtime_result_read_suppressed_preview_keeps_durable_continuation_ref -- --nocapture

Security Impact

No permission, secret, network-policy, filesystem, sandbox, or approval behavior changes. GitHub credentials remain host-mediated. Provider output still crosses the existing model-visible observation/redaction boundary; rejected credential-like preview text remains suppressed.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: no new public types; continuation authority remains constructed by the existing host result-reference path.
  • Untrusted content enters prompts only through an envelope/escaping primitive: unchanged model-observation validation and credential redaction remain in force.
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check: the promotion records the source artifact SHA-256 only as provenance.
  • New/changed status, exit, policy, runtime, or error variants: none added. Audit command: git diff origin/main...HEAD -- '*error*' '*status*'.
  • Security/durability serde(default) fields fail closed or have migration tests: not applicable; no persisted schema/default was added.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic: compact GitHub collections are capped at 100 items.
  • Driver/operator-visible errors have stable class semantics: no driver/operator error class changed; malformed provider responses remain explicit errors.
  • Sandbox/native/host names accurately describe trust boundary: not applicable; no runtime lane or trust-boundary name changed.

Database Impact

None.

Blast Radius

The first-party GitHub extension’s list/search response shape is intentionally narrower. Consumers needing omitted provider fields must use the existing detail tools. Reborn result observations can now carry continuation metadata without inline preview text; the referenced bytes and persistence semantics are unchanged.

Rollback Plan

Revert the branch commits. This restores the previous GitHub WASM asset/manifest and the prior observation-collapse behavior without a data migration. If only provider compatibility regresses, revert the compaction commit independently while retaining the continuation and regression-test commits.

Review Follow-Through

Please focus reviewer judgment on compact response field sufficiency and the metadata-only continuation observation. The PR is draft so normal CI and maintainer review can run before readiness.


Review track: C (runtime observation behavior and first-party provider execution)

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@railway-app

railway-app Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6953 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 4, 2026 at 9:49 am

@coderabbitai

coderabbitai Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • GitHub searches now support @me for author, assignee, reviewer-requested, and involvement filters.
    • “My pull requests” requests distinguish authorship from other relationships.
    • Sensitive large-result previews can be suppressed while preserving continuation details.
  • Bug Fixes

    • Prevented unnecessary account lookups during self-scoped searches.
    • Improved handling of rejected previews for large, sensitive results.
  • Documentation

    • Updated GitHub search and regression-testing guidance.
    • Added procedures for creating deterministic regression tests from run artifacts.
  • Tests

    • Added coverage for sensitive result previews and scoped GitHub searches.

Walkthrough

The PR adds regression-promotion guidance, suppresses unsafe large-result previews while preserving continuation metadata, updates GitHub self-scoped search contracts, and extends PR test-plan routing for guidance and extension assets.

Changes

Regression promotion workflow

Layer / File(s) Summary
Artifact promotion procedure
.claude/skills/promote-run-regression/SKILL.md, AGENTS.md
Defines provenance, scrubbing, deterministic test-seam selection, assertions, RED/GREEN validation, verification commands, and promotion evidence.

Safe result preview continuation

Layer / File(s) Summary
Suppressed-preview runtime coverage
crates/ironclaw_loop_host/src/result_read.rs, crates/ironclaw_reborn_composition/src/runtime/tests/core.rs
Unsafe result_read previews are omitted while the durable reference, byte count, and continuation offset remain available. Tests cover sensitive payloads and ordinary document content.

GitHub self-scoped search contracts

Layer / File(s) Summary
Recorded QA search contract
tests/reborn_qa_recorded_behavior.rs, docs/internal/testing-playbook.md
Updates package-qualified commands and adds a contract for one scoped open pull-request search authored by the authenticated user. The contract forbids identity lookup, legacy listing, and result reads.
Relationship-aware search prompts
crates/extensions/packages/github/prompts/github/*.md, crates/extensions/packages/github/wasm-src/src/lib.rs
Maps relationship-based requests to author, assignee, involves, and review-request qualifiers. Tests verify that @me produces one correctly formed request.

PR test-plan path routing

Layer / File(s) Summary
Guidance and extension asset classification
scripts/ci/reborn_pr_test_plan.py, scripts/ci/test_reborn_pr_test_plan.py, .gitignore, scripts/ci/wasm-src-digests.toml
The planner ignores .claude/ and repository hygiene paths, and routes extension package assets to ironclaw_extension_support. Tests cover representative asset types. WASM build artifacts are ignored, and the GitHub source digest is updated.

Estimated code review effort: 3 (Moderate) | ~30 minutes

Possibly related PRs

Suggested reviewers: think-in-universe

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title follows Conventional Commits style and accurately summarizes the GitHub search optimization changes.
Description check ✅ Passed The description follows the repository template and documents scope, validation, test strategy, security, rollback, and review details.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the scope: docs Documentation label Jul 31, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 July 31, 2026 09:14 Destroyed
@github-actions github-actions Bot added size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 31, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 July 31, 2026 09:14 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 July 31, 2026 09:14 Destroyed
@serrrfirat
serrrfirat marked this pull request as ready for review July 31, 2026 09:21
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@ironloopai

ironloopai Bot commented Jul 31, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6953

🔴 Failed

Execution result is invalid

The structured result could not be verified.

Automatic · PR opened · attempt 1 of 3 · failed after 3m 6s

Failure details
  • Repository: nearai/ironclaw
  • Base: main at 945d926
  • Head: codex/github-tool-efficiency-regression at 8f0c90b
  • Created: Jul 31, 2026, 9:26 AM UTC
  • Updated: Jul 31, 2026, 9:29 AM UTC
  • Run: 5d77d9f4-441a-4beb-919d-d41f7b2447c9
  • Latest attempt: 1 · Completed · deda740c-f22e-44c8-8f00-36a3fb93c521
  • Failed during: Verification
  • Retryable: No
  • Failure: 9f4a83d5-ca46-4cd2-9c07-10e362383468

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.claude/skills/promote-run-regression/SKILL.md:
- Around line 116-117: Update the final fixture check around the rg command to
avoid printing matching JSON content; use rg -l or a structured jq validation
that reports only the affected filename and failing key while still checking
"_review", PENDING_REPLAY_COMMIT, and PR_NUMBER.

In `@crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/lib.rs`:
- Around line 1655-1748: Add a small table-driven caller-level test alongside
search_issues_pull_requests_compacts_items_and_preserves_envelope_and_errors
that invokes execute_inner for malformed providers: a non-object search body, a
search object missing its items array, and a list_pull_requests response
containing a non-object item. Set each response through
test_support::set_response and assert execute_inner returns
github_api_invalid_response, while preserving the existing happy-path and
provider-error coverage.

In `@crates/ironclaw_host_runtime/tests/github_wasm_runtime_contract.rs`:
- Around line 351-357: Extend the list pull requests test around the existing
requests[0] assertions to verify the request has an empty body, the expected
applied network policy, and the injected authorization header, matching the
search-path contract assertions. Reuse github_wasm_services_for_test! in the
search test setup, retaining local policy and slot_handle bindings so its
existing assertions remain valid.

In `@tests/reborn_qa_recorded_behavior.rs`:
- Around line 427-443: Update the test around the github.get_authenticated_user
and github.search_issues_pull_requests assertions to capture the
authenticated-user result and assert that its login is the value used for the
search author. Replace the standalone literal author check with a comparison
against the returned login while preserving the existing owner, repo, state,
type, and call-sequence assertions.
- Around line 448-452: Strengthen the assertion in the priority-list workflow
around final_text_reply so it verifies the caller-visible ranked-list outcome,
not merely the presence of “priority.” Assert stable expected entries from the
fixture and the expected list structure, preserving the existing final reply
extraction and failure context.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9ec9df7b-3a2c-4d47-a382-d1704d084191

📥 Commits

Reviewing files that changed from the base of the PR and between 16b26f3 and 8f0c90b.

⛔ Files ignored due to path filters (2)
  • crates/ironclaw_first_party_extensions/assets/github/wasm/github_tool.wasm is excluded by !**/*.wasm, !**/*.wasm
  • tests/fixtures/llm_traces/reborn_qa/github_open_pr_priority.json is excluded by !tests/fixtures/**
📒 Files selected for processing (18)
  • .claude/skills/promote-run-regression/SKILL.md
  • AGENTS.md
  • crates/ironclaw_agent_loop/src/executor/capabilities.rs
  • crates/ironclaw_first_party_extensions/assets/github/manifest.toml
  • crates/ironclaw_first_party_extensions/assets/github/prompts/github/get_authenticated_user.md
  • crates/ironclaw_first_party_extensions/assets/github/prompts/github/list_pull_requests.md
  • crates/ironclaw_first_party_extensions/assets/github/prompts/github/search_issues.md
  • crates/ironclaw_first_party_extensions/assets/github/prompts/github/search_issues_pull_requests.md
  • crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/api/pulls.rs
  • crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/api/search.rs
  • crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/lib.rs
  • crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/response.rs
  • crates/ironclaw_host_api/src/resolution.rs
  • crates/ironclaw_host_runtime/tests/github_wasm_runtime_contract.rs
  • crates/ironclaw_reborn_composition/src/runtime/tests/core.rs
  • crates/ironclaw_turns/src/run_profile/resolution.rs
  • docs/internal/testing-playbook.md
  • tests/reborn_qa_recorded_behavior.rs

Comment on lines +116 to +117
rg -n '"_review"|PENDING_REPLAY_COMMIT|PR_NUMBER' \
tests/fixtures/llm_traces/reborn_qa --glob '*.json'

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Do not print raw JSON lines during the final fixture check.

rg -n prints the complete matching JSON line. A minified fixture can expose the entire artifact in terminal logs. Use rg -l or a structured jq check that reports only the file and failing key.

As per coding guidelines, never commit secrets or PII.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.claude/skills/promote-run-regression/SKILL.md around lines 116 - 117,
Update the final fixture check around the rg command to avoid printing matching
JSON content; use rg -l or a structured jq validation that reports only the
affected filename and failing key while still checking "_review",
PENDING_REPLAY_COMMIT, and PR_NUMBER.

Source: Coding guidelines

Comment on lines +1655 to +1748
#[test]
fn search_issues_pull_requests_compacts_items_and_preserves_envelope_and_errors() {
let large_body = "x".repeat(32 * 1024);
let provider_items = (0..100)
.map(|index| {
json!({
"id": 20_000 + index,
"node_id": format!("I_{index}"),
"number": 6_000 + index,
"title": format!("Search result {index}"),
"body": large_body,
"state": "open",
"state_reason": null,
"locked": false,
"draft": index % 2 == 0,
"html_url": format!("https://github.com/nearai/ironclaw/pull/{}", 6_000 + index),
"repository_url": "https://api.github.com/repos/nearai/ironclaw",
"user": {"login": format!("author-{index}"), "avatar_url": "https://example.test/avatar"},
"labels": [{"name": "bug", "description": large_body}],
"assignees": [{"login": "owner", "avatar_url": "https://example.test/avatar"}],
"milestone": {"title": "next", "description": large_body},
"comments": 17,
"created_at": "2026-07-01T00:00:00Z",
"updated_at": "2026-07-31T00:00:00Z",
"closed_at": null,
"author_association": "MEMBER",
"pull_request": {
"url": format!("https://api.github.com/repos/nearai/ironclaw/pulls/{}", 6_000 + index),
"html_url": format!("https://github.com/nearai/ironclaw/pull/{}", 6_000 + index),
"diff_url": "https://example.test/large.diff",
"patch_url": "https://example.test/large.patch"
},
"reactions": {"total_count": 999, "url": "https://api.github.com/large"},
"performed_via_github_app": {"description": large_body},
"score": 1.0
})
})
.collect::<Vec<_>>();
let provider_output = json!({
"total_count": 12_345,
"incomplete_results": true,
"items": provider_items
})
.to_string();
assert!(provider_output.len() > 6 * 1024 * 1024);
test_support::set_response(Ok(provider_output));

let output = execute_inner(
r#"{"repo":"nearai/ironclaw","type":"pr","page":4,"limit":100}"#,
Some(r#"{"capability_id":"github.search_issues_pull_requests"}"#),
)
.expect("github.search_issues_pull_requests should compact the provider response");

assert!(
output.len() < 100 * 1024,
"100 compact search results should stay model-useful, got {} bytes",
output.len()
);
let parsed: serde_json::Value =
serde_json::from_str(&output).expect("compact search output should be JSON");
assert_eq!(parsed["total_count"], 12_345);
assert_eq!(parsed["incomplete_results"], true);
assert_eq!(parsed["items"].as_array().map(Vec::len), Some(100));
assert_eq!(parsed["items"][0]["user"], json!({"login": "author-0"}));
assert_eq!(parsed["items"][0]["labels"], json!([{"name": "bug"}]));
assert_eq!(
parsed["items"][0]["pull_request"],
json!({
"url": "https://api.github.com/repos/nearai/ironclaw/pulls/6000",
"html_url": "https://github.com/nearai/ironclaw/pull/6000"
})
);
assert!(
parsed["items"][0].get("body").is_none()
&& parsed["items"][0].get("reactions").is_none()
&& parsed["items"][0].get("performed_via_github_app").is_none(),
"large detail must remain available through github.get_issue or github.get_pull_request"
);
assert_eq!(
test_support::requests()[0].path,
"/search/issues?q=repo%3Anearai%2Fironclaw%20is%3Apr&per_page=100&page=4"
);

test_support::set_response(Err("github_api_error_status_503".to_string()));
assert_eq!(
execute_inner(
r#"{"repo":"nearai/ironclaw","type":"pr"}"#,
Some(r#"{"capability_id":"github.search_issues_pull_requests"}"#),
)
.unwrap_err(),
"github_api_error_status_503",
"provider errors must pass through the compacting response path"
);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add caller-level coverage for the github_api_invalid_response branches.

The happy path and the provider-error passthrough are both covered. The rejection branches in response.rs are not. validate_page_size and the shape checks (compact_issue_search on a non-object body, a missing items array, or a non-object item) all collapse to one model-visible code, github_api_invalid_response. Add a small table-driven case through execute_inner so a future change to the compactor cannot silently turn a valid provider page into that error.

💚 Suggested additional test
#[test]
fn search_and_list_reject_malformed_provider_shapes() {
    for (capability, params, provider) in [
        (
            "github.search_issues_pull_requests",
            r#"{"repo":"nearai/ironclaw","type":"pr"}"#,
            json!([]).to_string(),
        ),
        (
            "github.search_issues_pull_requests",
            r#"{"repo":"nearai/ironclaw","type":"pr"}"#,
            json!({"total_count": 1}).to_string(),
        ),
        (
            "github.list_pull_requests",
            r#"{"owner":"nearai","repo":"ironclaw"}"#,
            json!(["not-an-object"]).to_string(),
        ),
    ] {
        test_support::set_response(Ok(provider));
        assert_eq!(
            execute_inner(
                params,
                Some(&format!(r#"{{"capability_id":"{capability}"}}"#))
            )
            .unwrap_err(),
            "github_api_invalid_response",
            "{capability} must reject malformed provider shapes"
        );
    }
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/lib.rs`
around lines 1655 - 1748, Add a small table-driven caller-level test alongside
search_issues_pull_requests_compacts_items_and_preserves_envelope_and_errors
that invokes execute_inner for malformed providers: a non-object search body, a
search object missing its items array, and a list_pull_requests response
containing a non-object item. Set each response through
test_support::set_response and assert execute_inner returns
github_api_invalid_response, while preserving the existing happy-path and
provider-error coverage.

Comment on lines +351 to +357
let requests = network.requests();
assert_eq!(requests.len(), 1);
assert_eq!(
requests[0].url,
"https://api.github.com/repos/nearai/ironclaw/pulls?state=open&per_page=30&page=3"
);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Assert credential injection and network policy on the list path too.

github.list_pull_requests declares effects = ["network", "use_secret"]. The search test at lines 249-262 asserts the method, the empty body, the applied policy, and the injected authorization header. This test asserts only the request count and the URL. A regression that dropped the InjectCredentialAccountOnce obligation or the ApplyNetworkPolicy obligation for the list capability would still pass.

🔒️ Proposed fix
     let requests = network.requests();
     assert_eq!(requests.len(), 1);
+    assert_eq!(requests[0].method, NetworkMethod::Get);
+    assert_eq!(requests[0].body, Vec::<u8>::new());
+    assert_eq!(requests[0].policy, github_policy());
+    assert_eq!(
+        requests[0]
+            .headers
+            .iter()
+            .find(|(name, _)| name == "authorization"),
+        Some(&(
+            "authorization".to_string(),
+            "Bearer ghp_fake_fixture_token".to_string(),
+        ))
+    );
     assert_eq!(
         requests[0].url,
         "https://api.github.com/repos/nearai/ironclaw/pulls?state=open&per_page=30&page=3"
     );

Separately: the search test builds HostRuntimeServices inline at lines 181-208 while this test uses the new github_wasm_services_for_test! macro. Reusing the macro in the search test would remove that duplication, though the search test needs local policy and slot_handle bindings for its assertions.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
let requests = network.requests();
assert_eq!(requests.len(), 1);
assert_eq!(
requests[0].url,
"https://api.github.com/repos/nearai/ironclaw/pulls?state=open&per_page=30&page=3"
);
}
let requests = network.requests();
assert_eq!(requests.len(), 1);
assert_eq!(requests[0].method, NetworkMethod::Get);
assert_eq!(requests[0].body, Vec::<u8>::new());
assert_eq!(requests[0].policy, github_policy());
assert_eq!(
requests[0]
.headers
.iter()
.find(|(name, _)| name == "authorization"),
Some(&(
"authorization".to_string(),
"Bearer ghp_fake_fixture_token".to_string(),
))
);
assert_eq!(
requests[0].url,
"https://api.github.com/repos/nearai/ironclaw/pulls?state=open&per_page=30&page=3"
);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_host_runtime/tests/github_wasm_runtime_contract.rs` around
lines 351 - 357, Extend the list pull requests test around the existing
requests[0] assertions to verify the request has an empty body, the expected
applied network policy, and the injected authorization header, matching the
search-path contract assertions. Reuse github_wasm_services_for_test! in the
search test setup, retaining local policy and slot_handle bindings so its
existing assertions remain valid.

Comment thread tests/reborn_qa_recorded_behavior.rs Outdated
Comment on lines +427 to +443
assert_tool_sequence(
&trace,
&[
"github.get_authenticated_user",
"github.search_issues_pull_requests",
],
);
assert_tool_called_with(
&trace,
"github.search_issues_pull_requests",
&[
r#""owner":"nearai""#,
r#""repo":"ironclaw""#,
r#""author":"fixture-user""#,
r#""state":"open""#,
r#""type":"pr""#,
],

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Assert that the search author comes from the authenticated-user result.

The test checks call order and a literal "author":"fixture-user", but it does not compare that value with the result of github.get_authenticated_user. A regression can ignore the lookup and still pass this test. Assert that the returned login equals the search author.

Based on the GitHub search contract, the authenticated login must populate author.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/reborn_qa_recorded_behavior.rs` around lines 427 - 443, Update the test
around the github.get_authenticated_user and github.search_issues_pull_requests
assertions to capture the authenticated-user result and assert that its login is
the value used for the search author. Replace the standalone literal author
check with a comparison against the returned login while preserving the existing
owner, repo, state, type, and call-sequence assertions.

Comment on lines +448 to +452
let reply = final_text_reply(&trace).expect("priority-list phrase should finalize a reply");
assert!(
reply.to_ascii_lowercase().contains("priority"),
"priority-list workflow should end with a priority list"
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Require a positive priority-list assertion.

reply.to_ascii_lowercase().contains("priority") also accepts a refusal or unrelated text. Assert a stable positive outcome, such as expected ranked entries and list structure from the fixture.

Based on .claude/skills/promote-run-regression/SKILL.md, the regression must assert the caller-visible final outcome.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/reborn_qa_recorded_behavior.rs` around lines 448 - 452, Strengthen the
assertion in the priority-list workflow around final_text_reply so it verifies
the caller-visible ranked-list outcome, not merely the presence of “priority.”
Assert stable expected entries from the fixture and the expected list structure,
preserving the existing final reply extraction and failure context.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 July 31, 2026 09:33 Destroyed
@github-actions github-actions Bot added size: XL 500+ changed lines and removed size: L 200-499 changed lines labels Jul 31, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/lib.rs (1)

569-677: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Make the PR cover list_pull_requests with the real WASM HTTP limit.

list_pull_requests calls github_request(...).and_then(compact_pull_request_list), but github_request maps host body-size failures to github_api_body_limit. This test uses test_support::set_response, which bypasses the production WASM egress path and can let compaction run when host transport would reject a large body. Add host-runtime/real-transport coverage for the host response-limit ordering so CLAUDE.md’s “Test through the caller” invariant covers this egress gate.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/lib.rs`
around lines 569 - 677, Extend coverage around list_pull_requests and
github_request using the host-runtime real HTTP transport rather than
test_support::set_response. Configure a response exceeding the WASM host body
limit, invoke list_pull_requests through its normal caller path, and assert
github_api_body_limit is returned before compact_pull_request_list can run.
Preserve a separate successful compact-response assertion if needed, ensuring
the test validates the production egress gate and caller ordering.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@crates/ironclaw_first_party_extensions/assets/github/prompts/github/search_issues_pull_requests.md`:
- Line 9: Update the search prompt description around compact_search_item to
stop promising a standalone type marker and instead describe the optional
pull_request object as the discriminator between pull requests and issues. Keep
the documented fields aligned with the schema implemented by compact_search_item
and its tests.

---

Outside diff comments:
In `@crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/lib.rs`:
- Around line 569-677: Extend coverage around list_pull_requests and
github_request using the host-runtime real HTTP transport rather than
test_support::set_response. Configure a response exceeding the WASM host body
limit, invoke list_pull_requests through its normal caller path, and assert
github_api_body_limit is returned before compact_pull_request_list can run.
Preserve a separate successful compact-response assertion if needed, ensuring
the test validates the production egress gate and caller ordering.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: db3c0b82-8853-4fca-bcce-0d0083e6efe8

📥 Commits

Reviewing files that changed from the base of the PR and between 8f0c90b and 2285aeb.

⛔ Files ignored due to path filters (1)
  • tests/fixtures/llm_traces/reborn_qa/github_open_pr_priority.json is excluded by !tests/fixtures/**
📒 Files selected for processing (6)
  • crates/ironclaw_first_party_extensions/assets/github/prompts/github/get_authenticated_user.md
  • crates/ironclaw_first_party_extensions/assets/github/prompts/github/list_pull_requests.md
  • crates/ironclaw_first_party_extensions/assets/github/prompts/github/search_issues.md
  • crates/ironclaw_first_party_extensions/assets/github/prompts/github/search_issues_pull_requests.md
  • crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/lib.rs
  • tests/reborn_qa_recorded_behavior.rs


Prefer the structured `owner`, `repo`, `author`, `assignee`, `involves`, `state`, and `type` fields over duplicating those qualifiers in `query`.

The result keeps GitHub's `total_count`, `incomplete_results`, and `items` search envelope while returning compact item summaries with the repository URL, number, title, type marker, state/draft status, URL, author, labels, assignees, milestone, comment count, timestamps, and score. Use `page` and `limit` to continue through results. For bodies or other full detail, call `github.get_pull_request` for pull request items or `github.get_issue` for issue items.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Align the prompt with the compact search schema.

Line 9 promises a standalone “type marker”, but response.rs::compact_search_item does not copy a type field or add another marker. It preserves the optional pull_request object instead. Either add an explicit type field and test it, or describe pull_request presence as the discriminator.

As per path instructions, documentation promising guarantees must match code and tests.

Suggested prompt fix
-... number, title, type marker, state/draft status, URL, author, ...
+... number, title, state/draft status, URL, author, ...
+For pull requests, use the returned `pull_request` metadata to identify the result type.
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
The result keeps GitHub's `total_count`, `incomplete_results`, and `items` search envelope while returning compact item summaries with the repository URL, number, title, type marker, state/draft status, URL, author, labels, assignees, milestone, comment count, timestamps, and score. Use `page` and `limit` to continue through results. For bodies or other full detail, call `github.get_pull_request` for pull request items or `github.get_issue` for issue items.
The result keeps GitHub's `total_count`, `incomplete_results`, and `items` search envelope while returning compact item summaries with the repository URL, number, title, state/draft status, URL, author, labels, assignees, milestone, comment count, timestamps, and score. Use `page` and `limit` to continue through results. For pull requests, use the returned `pull_request` metadata to identify the result type. For bodies or other full detail, call `github.get_pull_request` for pull request items or `github.get_issue` for issue items.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/ironclaw_first_party_extensions/assets/github/prompts/github/search_issues_pull_requests.md`
at line 9, Update the search prompt description around compact_search_item to
stop promising a standalone type marker and instead describe the optional
pull_request object as the discriminator between pull requests and issues. Keep
the documented fields aligned with the schema implemented by compact_search_item
and its tests.

Source: Path instructions

@serrrfirat serrrfirat left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Reduce PR-prioritization tool churn by compacting GitHub list/search outputs, routing self-scoped searches via @me, and preserving durable continuation refs when preview text is suppressed.

Shape: primary mode normal; no modifiers. XL PR (992+/54−, 20 files) packetized into 4 buckets (production/tests/config/docs). Local-git diff against main.

Coverage: ⚠️ Partial. Diff source: local-git. 4/8 reviewers returned clean strict receipts (security, performance, conventions, maintainability). 2 returned narrative reviews without the strict receipt — findings reconstructed from reviewer prose and verified against the worktree (local-patterns: 4 anchored; approach: 0). 1 reviewer failed (bugs stalled at 112 turns without emitting a receipt). Packetization complete (4/4 packets, 20 files, 0 oversized). Approval is forbidden under partial coverage.

Stats: 6 findings (from 6 raw, 6 after dedup) across 2 files. Reviewers run: security, bugs, performance, tests, conventions, local-patterns, maintainability, approach. Reviewers failed: bugs (stalled). Receipt-only-but-noted: local-patterns, approach. Body-only: 0 (all 6 resolve to valid diff lines).

Reviewers failed / partial:

  • bugs — stalled at 112 turns reading the test packet, no receipt emitted. Its lens (production error-class behavior) was partially recovered: the github_api_invalid_response error-class finding below was verified by the orchestrator against lib.rs:guest_error_kind and is flagged local-patterns + cross-noted. A dedicated bugs re-review is recommended.
  • local-patterns, approach — completed review but emitted the delegate agent's acceptance-report scaffold instead of the strict {coverage, findings} receipt. local-patterns findings were recovered from its narrative and verified; approach reported 0 findings in its narrative.

Tests

  1. Medium compact_pull_request_list error paths untested (crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/response.rs:5-13, confidence 80) — anchor: crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/response.rs:5
    New compact_pull_request_list returns github_api_invalid_response via ?/ok_or_else for malformed JSON, non-array top-level response, item_count > MAX_PAGE_ITEMS (validate_page_size), item not an object, and wrong-type nested fields (user/labels/head/base not object/array). Only the happy 100-item path and the upstream provider Err pass-through are exercised; none of the response.rs-authored invalid_response_error branches have a test. GitHub pagination drift or proxy-injected malformed payloads could hit these paths silently.
    Fix: tests::list_pull_requests_compaction_rejects_malformed_provider_response covering malformed JSON, non-array body, 101-item oversized page, item-not-object, and head/labels/user wrong-type fields returning github_api_invalid_response

  2. Medium compact_issue_search envelope error paths untested (crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/response.rs:16-28, confidence 80) — anchor: crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/response.rs:16
    New compact_issue_search returns github_api_invalid_response when the response is not a JSON object, lacks the items array, items is not an array, item_count > 100, or search items have wrong-type nested fields (user/labels/milestone/pull_request shape). The compact path is only exercised on the happy 100-item fixture; the response.rs-authored invalid branches lack tests.
    Fix: tests::search_issues_compaction_rejects_malformed_provider_response covering non-object body, missing items, 101-item oversized page, and wrong-type nested fields returning github_api_invalid_response

Local Patterns

  1. Medium New github_api_invalid_response code erased to generic operation_failed (crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/response.rs:178-179, confidence 78) — anchor: crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/lib.rs:99
    New error code github_api_invalid_response invented in response.rs:178 diverges from sibling github_api_error_status_{N} convention in request.rs:42,45. It is not enumerated in guest_error_kind (lib.rs:90-99), so it falls through the _ => "operation_failed" default arm. A malformed/oversized/wrong-type GitHub response — a distinct provider-output contract violation — is reported to the host/driver under the generic operation_failed class, erasing its stable class semantics and making it indistinguishable from unrelated host failures. AGENTS.md trust-boundary checklist requires driver/operator-visible errors have stable class semantics.
    Fix: Add an explicit arm in guest_error_kind mapping github_api_invalid_response to a stable class such as input (malformed provider output / contract violation), and keep the name under the github_api_* sibling prefix.

  2. Low Stale TRUNCATED-preview wording in continuation doc comment (crates/ironclaw_turns/src/run_profile/resolution.rs:544-546, confidence 72) — anchor: crates/ironclaw_turns/src/run_profile/resolution.rs:544
    Doc comment says ResultPreviewMeta carries TRUNCATED-preview continuation info, but the behavior change preserves continuation metadata when a preview is rejected (not truncated). Sibling host_api/resolution.rs:449-454 was reworded to say 'remains present when an unsafe preview is suppressed'; this turns copy was not reconciled. A reader mis-scopes ResultPreviewMeta as truncated-only.
    Fix: Reconcile the turns/resolution.rs:544-546 doc sentence with the host_api rewording — replace 'TRUNCATED-preview continuation info' with 'continuation metadata, preserved even when a preview is suppressed'.

  3. Nit response.rs module lacks doc comment (crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/response.rs:1-1, confidence 65) — anchor: crates/ironclaw_first_party_extensions/assets/github/wasm-src/src/response.rs:1
    New response.rs module performs non-obvious compaction/projection of GitHub PR/search results but has no module-level doc comment. Sibling validation.rs:7-8 documents exported items. A short //! module doc would help maintainers understand the compaction contract and the 100-item cap intent.
    Fix: Add a //! module doc comment at response.rs:1 describing compaction, field projection, and the MAX_PAGE_ITEMS cap.

  4. Nit Test helper rename refs_preview->outcome_refs drifts from sibling naming (crates/ironclaw_turns/src/run_profile/resolution.rs:1597-1597, confidence 55) — anchor: crates/ironclaw_turns/src/run_profile/resolution.rs:1599
    Test helper renamed from refs_preview to outcome_refs after a scope change. Sibling completion-test naming elsewhere uses refs_preview-shaped verbs. Harmless local rename but a reader grepping for refs_preview may miss the renamed helper.
    Fix: Optional: keep a brief alias or align sibling completion-test naming to outcome_refs.

Security / Bugs / Performance / Conventions / Maintainability / Approach

All returned 0 findings (security/performance/conventions/maintainability/approach). Bugs lens not fully covered — see partial-coverage note. Note: the github_api_invalid_response error-class erasure flagged under Local Patterns is also a bugs-class concern (production error semantics) and was the most likely hit a completed bugs review would have surfaced.


Note on coverage: This review ran on the delegate builtin subagent which injected an acceptance-report scaffold that caused several reviewers to emit narrative instead of the strict receipt contract. Findings marked as reconstructed were verified by the orchestrator against the live worktree at head_sha. Recommend re-running with a stricter reviewer agent for the bugs lens.

serde_json::to_string(value).map_err(|_| invalid_response_error())
}

fn invalid_response_error() -> String {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — New github_api_invalid_response code erased to generic operation_failed.

New error code github_api_invalid_response invented in response.rs:178 diverges from sibling github_api_error_status_{N} convention in request.rs:42,45. It is not enumerated in guest_error_kind (lib.rs:90-99), so it falls through the _ => "operation_failed" default arm. A malformed/oversized/wrong-type GitHub response — a distinct provider-output contract violation — is reported to the host/driver under the generic operation_failed class, erasing its stable class semantics and making it indistinguishable from unrelated host failures. AGENTS.md trust-boundary checklist requires driver/operator-visible errors have stable class semantics.

Fix: Add an explicit arm in guest_error_kind mapping github_api_invalid_response to a stable class such as input (malformed provider output / contract violation), and keep the name under the github_api_* sibling prefix.

"github_api_invalid_response" => "input",

/// referenced result ref, full byte size, next offset, and JSON-array element
/// count, so the model reads the full result. Detail kinds other than
/// `ResultReference` have no inline content.
/// count, so the model reads the full result. This metadata is preserved even

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low — Stale TRUNCATED-preview wording in continuation doc comment.

Doc comment says ResultPreviewMeta carries TRUNCATED-preview continuation info, but the behavior change preserves continuation metadata when a preview is rejected (not truncated). Sibling host_api/resolution.rs:449-454 was reworded to say 'remains present when an unsafe preview is suppressed'; this turns copy was not reconciled. A reader mis-scopes ResultPreviewMeta as truncated-only.

Fix: Reconcile the turns/resolution.rs:544-546 doc sentence with the host_api rewording — replace 'TRUNCATED-preview continuation info' with 'continuation metadata, preserved even when a preview is suppressed'.

/// the model reads the full result. This metadata is preserved even when

@@ -0,0 +1,180 @@
use serde_json::{Map, Value};

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit — response.rs module lacks doc comment.

New response.rs module performs non-obvious compaction/projection of GitHub PR/search results but has no module-level doc comment. Sibling validation.rs:7-8 documents exported items. A short //! module doc would help maintainers understand the compaction contract and the 100-item cap intent.

Fix: Add a //! module doc comment at response.rs:1 describing compaction, field projection, and the MAX_PAGE_ITEMS cap.

//! Compacts GitHub PR-list and issue/PR-search provider responses before they

other => panic!("expected GateRecord::Resource, got {other:?}"),
}
}

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit — Test helper rename refs_preview->outcome_refs drifts from sibling naming.

Test helper renamed from refs_preview to outcome_refs after a scope change. Sibling completion-test naming elsewhere uses refs_preview-shaped verbs. Harmless local rename but a reader grepping for refs_preview may miss the renamed helper.

Fix: Optional: keep a brief alias or align sibling completion-test naming to outcome_refs.

// outcome_refs (was refs_preview; scope expanded to preserve continuation meta)


const MAX_PAGE_ITEMS: usize = 100;

pub(crate) fn compact_pull_request_list(response: String) -> Result<String, String> {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — compact_pull_request_list error paths untested.

New compact_pull_request_list returns github_api_invalid_response via ?/ok_or_else for malformed JSON, non-array top-level response, item_count > MAX_PAGE_ITEMS (validate_page_size), item not an object, and wrong-type nested fields (user/labels/head/base not object/array). Only the happy 100-item path and the upstream provider Err pass-through are exercised; none of the response.rs-authored invalid_response_error branches have a test. GitHub pagination drift or proxy-injected malformed payloads could hit these paths silently.

Fix: tests::list_pull_requests_compaction_rejects_malformed_provider_response covering malformed JSON, non-array body, 101-item oversized page, item-not-object, and head/labels/user wrong-type fields returning github_api_invalid_response

// set_response(Ok("not json".into())); assert!(matches!(execute_inner(...), Err(e) if e.contains("github_api_invalid_response")))

serialize(&compact)
}

pub(crate) fn compact_issue_search(response: String) -> Result<String, String> {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — compact_issue_search envelope error paths untested.

New compact_issue_search returns github_api_invalid_response when the response is not a JSON object, lacks the items array, items is not an array, item_count > 100, or search items have wrong-type nested fields (user/labels/milestone/pull_request shape). The compact path is only exercised on the happy 100-item fixture; the response.rs-authored invalid branches lack tests.

Fix: tests::search_issues_compaction_rejects_malformed_provider_response covering non-object body, missing items, 101-item oversized page, and wrong-type nested fields returning github_api_invalid_response

// set_response(Ok("[]".into())); assert!(matches!(execute_inner(...), Err(e) if e.contains("github_api_invalid_response")))

Resolve conflicts from the extensions colocation (WS2) refactor:
- github wasm-src lib.rs: keep self-qualifier search test on new path
- response.rs: identical both sides, accept at new path
- list_pull_requests.md / search_issues_pull_requests.md: keep PR's
  @me guidance at new extensions path
- github_tool.wasm: drop old ironclaw_first_party_extensions path,
  keep new extensions path (main deleted old, PR modified bytes)
- resolution.rs: keep main's suppressed_array meta test

Local validation:
- cargo test -p ironclaw_loop_contracts --lib: 132 passed
- github wasm-src cargo test: 58 passed
- cargo test -p ironclaw_host_runtime --test github_wasm_runtime_contract
  --features test-support --all-targets: 48 passed
- clippy clean on ironclaw_loop_contracts and github wasm-src
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 August 3, 2026 14:12 Destroyed
Three CI gates regressed when origin/main was merged into this PR:

1. Reborn PR test planner rejected .claude/skills/promote-run-regression/
   SKILL.md as an unclassified path. .claude/ holds agent guidance, not
   compiled test surface; add it to IGNORED_PREFIXES.

2. The planner also rejected crates/extensions/packages/<pkg>/{prompts,
   schemas, manifest.toml, wasm-src/, wasm/} paths as unmapped crate
   paths. WS2 colocated first-party extension packages as siblings of
   the support crate; they are not workspace members, so the
   workspace-package lookup cannot map them. Route them to
   ironclaw_extension_support (the freshness-gate anchor; buckets into
   wasm-sandbox) so prompt/schema/guest-source changes run the right
   bucket instead of failing fast.

3. check-wasm-artifact-freshness failed: the PR rebuilt the github
   wasm from updated wasm-src but did not re-record the source digest.
   Rebuild + re-record so the committed artifact matches the recorded
   digest. Also drop the accidentally-tracked guest Cargo.lock (guests
   resolve fresh; the freshness gate excludes Cargo.lock as non-source)
   and gitignore all guest wasm-src lockfiles + target dirs.

Tests:
- python3.11 -m unittest scripts/ci/test_reborn_pr_test_plan.py: 34 OK
  (added coverage for .claude/ ignore + extension-package routing)
- python3.11 scripts/ci/check-wasm-artifact-freshness.py: OK, 6 packages
- scripts/ci/test-check-wasm-artifact-freshness.sh: 19 passed, 0 failed
- cargo test -p ironclaw_loop_contracts --lib: 132 passed
- github wasm-src cargo test: 58 passed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 August 3, 2026 16:10 Destroyed
@github-actions github-actions Bot added size: M 50-199 changed lines and removed size: XL 500+ changed lines labels Aug 3, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/ci/test_reborn_pr_test_plan.py`:
- Around line 61-67: Update the canonical_packages selection in the helper
method calling planner.build_plan so only a None value falls back to
self.canonical; preserve an explicitly provided empty list by using an explicit
None check instead of truthiness.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 3845f5b0-ce68-4627-bcaa-943fee44231a

📥 Commits

Reviewing files that changed from the base of the PR and between bd755ad and 5a89460.

📒 Files selected for processing (4)
  • .gitignore
  • scripts/ci/reborn_pr_test_plan.py
  • scripts/ci/test_reborn_pr_test_plan.py
  • scripts/ci/wasm-src-digests.toml

Comment on lines +61 to +67
canonical_packages: list[str] | None = None,
) -> dict:
return planner.build_plan(
event=event,
changed_paths=paths,
metadata=metadata(),
canonical_packages=self.canonical,
canonical_packages=canonical_packages or self.canonical,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Preserve an explicit empty canonical package list.

Line [67] uses truthiness, so canonical_packages=[] silently selects self.canonical. This prevents the helper from testing an intentionally empty canonical set and can hide the canonical-set failure path. Use an explicit None check.

Proposed fix
-            canonical_packages=canonical_packages or self.canonical,
+            canonical_packages=(
+                self.canonical
+                if canonical_packages is None
+                else canonical_packages
+            ),
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
canonical_packages: list[str] | None = None,
) -> dict:
return planner.build_plan(
event=event,
changed_paths=paths,
metadata=metadata(),
canonical_packages=self.canonical,
canonical_packages=canonical_packages or self.canonical,
canonical_packages: list[str] | None = None,
) -> dict:
return planner.build_plan(
event=event,
changed_paths=paths,
metadata=metadata(),
canonical_packages=(
self.canonical
if canonical_packages is None
else canonical_packages
),
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/ci/test_reborn_pr_test_plan.py` around lines 61 - 67, Update the
canonical_packages selection in the helper method calling planner.build_plan so
only a None value falls back to self.canonical; preserve an explicitly provided
empty list by using an explicit None check instead of truthiness.

.gitignore is a repository hygiene file: it changes which untracked files a
checkout sees, not which workspace crates or tests run. The planner's
unclassified-path guard rejected it, failing 'Detect Reborn test scope' on
this PR (which adds a gitignore rule for guest wasm-src Cargo.lock/target).

Add .gitignore, .dockerignore, .gitattributes to IGNORED_ROOT_PATHS with a
reason line, and cover them in the contract tests.

python3.11 -m unittest scripts/ci/test_reborn_pr_test_plan.py: 35 OK
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 August 3, 2026 16:15 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
scripts/ci/reborn_pr_test_plan.py (1)

408-413: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Add ironclaw_extension_support to the production canonical package set.

scripts/ci/reborn_pr_test_plan.py adds ironclaw_extension_support for changed crates/extensions/packages/ assets, but scripts/ci/discover-reborn-package-crates.sh only collects Cargo workspace crates from cargo metadata. That keeps changed_packages outside canonical_set in CI’s workflow path, which then raises changed packages are outside the canonical Reborn package set. Add ironclaw_extension_support to the canonical crate set, plus a CI test exercise without the test’s manually supplied canonical_packages.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/ci/reborn_pr_test_plan.py` around lines 408 - 413, Add
ironclaw_extension_support to the canonical package/crate set used by
scripts/ci/discover-reborn-package-crates.sh so it matches the production
package added by the reborn_pr_test_plan.py extension-package path. Update the
relevant CI test to exercise canonical discovery without manually supplying
canonical_packages, ensuring the changed package is included and no
outside-canonical-set error occurs.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@scripts/ci/reborn_pr_test_plan.py`:
- Around line 408-413: Add ironclaw_extension_support to the canonical
package/crate set used by scripts/ci/discover-reborn-package-crates.sh so it
matches the production package added by the reborn_pr_test_plan.py
extension-package path. Update the relevant CI test to exercise canonical
discovery without manually supplying canonical_packages, ensuring the changed
package is included and no outside-canonical-set error occurs.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c6e22ff9-a017-43f0-bd46-8f416dc6ebcd

📥 Commits

Reviewing files that changed from the base of the PR and between 5a89460 and 598f20c.

📒 Files selected for processing (2)
  • scripts/ci/reborn_pr_test_plan.py
  • scripts/ci/test_reborn_pr_test_plan.py

builtin.result_read built its model-visible observation with
preview: Some(content) unconditionally and relied on downstream
ToolResultReferenceEnvelope::new_best_effort_model_observation repair
(strip_unsafe_result_reference_preview) to drop a credential-bearing
preview later. CI showed a path where the model replay still contained
"preview", so the suppression leaked: the standalone runtime test
standalone_runtime_result_read_suppressed_preview_keeps_durable_continuation_ref
failed in the affected-2 bucket (gateway asserted
!contains("\"preview\""), then the panicked gateway turned the
turn into scheduler_executor_panic).

Construct the preview through ModelResultPreview::new(content).ok() at
the source — the same credential/content gate the initial first-look
preview uses (resolution::result_preview_parts). A chunk that fails the
gate drops only the inline preview; result_ref, total_bytes, next_offset
continuation metadata survives so the model can page past the
suppressed chunk. No longer dependent on best-effort downstream repair.

Tests:
- result_read_observation_drops_preview_for_credential_chunk_but_keeps_continuation
- result_read_observation_keeps_preview_for_ordinary_document_chunk
- cargo test -p ironclaw_loop_host --lib result_read: 3 passed
- cargo test -p ironclaw_reborn_composition --features test-support --lib     standalone_runtime_result_read_suppressed_preview_keeps_durable_continuation_ref: ok
- clippy -p ironclaw_loop_host clean
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 August 3, 2026 16:55 Destroyed
…es.md

Code review (local-patterns) found the result-envelope/pagination
paragraph appeared twice consecutively after this branch's edit — an
appended copy left the original in place. Sibling
search_issues_pull_requests.md states the equivalent sentence once.
Keep the single-occurrence pattern so future intent edits aren't masked.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 August 4, 2026 09:35 Destroyed
Resolve conflicts in scripts/ci/reborn_pr_test_plan.py and
test_reborn_pr_test_plan.py. Main landed its own .claude/ classification
plus IGNORED_GUIDANCE_PATHS, QA_HARNESS_PREFIXES, INTEGRATION_SUPPORT_OWNERS,
and an expanded PR_STATIC_CONTROL_PATHS. Re-apply this branch's two
additions on top of main's planner:
- EXTENSION_PACKAGES_PREFIX routing (crates/extensions/packages/<pkg>/
  prompts/schemas/manifest/wasm-src/wasm -> ironclaw_extension_support; not
  workspace members, so the workspace-package lookup cannot map them)
- IGNORED_ROOT_PATHS (.gitignore/.dockerignore/.gitattributes hygiene files
  own no Rust/E2E test surface)
Re-apply the paired contract tests (repository hygiene ignored; extension
package asset routes to support owner). Main's own .claude/ tests cover
the third path this branch added earlier.

python3.11 -m unittest scripts/ci/test_reborn_pr_test_plan.py: 45 OK
python3.11 scripts/ci/check-wasm-artifact-freshness.py: OK, 6 packages
PR plan builds clean against the merged tree.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6953 August 4, 2026 09:42 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_reborn_composition/src/runtime/tests/core.rs`:
- Around line 1035-1071: Strengthen the suppressed replay assertion in the
suppress_result_read branch of the test harness by verifying that
tool_result.content does not contain the sensitive chunk marker "secret " before
parsing its JSON metadata. Keep the existing preview-field and continuation
metadata assertions unchanged, and ensure the test exercises the real caller
path already covered by this branch.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 781ce11b-7efb-453d-a745-d1872df3ce5e

📥 Commits

Reviewing files that changed from the base of the PR and between abc47fb and 7988b07.

📒 Files selected for processing (1)
  • crates/ironclaw_reborn_composition/src/runtime/tests/core.rs

Comment on lines +1035 to +1071
if self.suppress_result_read_preview {
assert!(
!tool_result.content.contains("\"preview\""),
"the rejected chunk preview must remain suppressed"
);
let observation: serde_json::Value =
serde_json::from_str(&tool_result.content).expect("result_read observation");
let detail = &observation["model_observation"]["detail"];
let source_result_ref = self
.source_result_ref
.lock()
.expect("source result ref lock poisoned")
.clone()
.expect("source result ref captured");
assert_eq!(
detail["result_ref"].as_str(),
Some(source_result_ref.as_str()),
"suppressed preview replay must retain the durable source result reference"
);
assert_ne!(
detail["result_ref"], observation["result_ref"],
"the inline result_read invocation ref must never become continuation authority"
);
assert!(
detail["total_bytes"]
.as_u64()
.is_some_and(|total_bytes| total_bytes > 2048),
"suppressed preview replay must retain total bytes: {}",
tool_result.content
);
assert_eq!(
detail["next_offset"].as_u64(),
Some(2048),
"suppressed preview replay must retain the next continuation offset"
);
return Ok(HostManagedModelResponse::assistant_reply("tool ok"));
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Assert that the sensitive payload is absent from the replay.

Lines 1035-1039 only assert that JSON has no "preview" field. A regression could place the rejected secret ... chunk in another model-visible field and still pass. Assert that tool_result.content does not contain "secret " before parsing the continuation metadata.

As per coding guidelines, “Every new feature and bug fix must begin with a test that demonstrates the intended behavior.” As per path instructions, “Test through the caller” requires the real caller path to prove the side-effect gate.

Proposed test assertion
             if self.suppress_result_read_preview {
                 assert!(
                     !tool_result.content.contains("\"preview\""),
                     "the rejected chunk preview must remain suppressed"
                 );
+                assert!(
+                    !tool_result.content.contains("secret "),
+                    "the model-visible replay must not expose the rejected sensitive chunk"
+                );
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if self.suppress_result_read_preview {
assert!(
!tool_result.content.contains("\"preview\""),
"the rejected chunk preview must remain suppressed"
);
let observation: serde_json::Value =
serde_json::from_str(&tool_result.content).expect("result_read observation");
let detail = &observation["model_observation"]["detail"];
let source_result_ref = self
.source_result_ref
.lock()
.expect("source result ref lock poisoned")
.clone()
.expect("source result ref captured");
assert_eq!(
detail["result_ref"].as_str(),
Some(source_result_ref.as_str()),
"suppressed preview replay must retain the durable source result reference"
);
assert_ne!(
detail["result_ref"], observation["result_ref"],
"the inline result_read invocation ref must never become continuation authority"
);
assert!(
detail["total_bytes"]
.as_u64()
.is_some_and(|total_bytes| total_bytes > 2048),
"suppressed preview replay must retain total bytes: {}",
tool_result.content
);
assert_eq!(
detail["next_offset"].as_u64(),
Some(2048),
"suppressed preview replay must retain the next continuation offset"
);
return Ok(HostManagedModelResponse::assistant_reply("tool ok"));
}
if self.suppress_result_read_preview {
assert!(
!tool_result.content.contains("\"preview\""),
"the rejected chunk preview must remain suppressed"
);
assert!(
!tool_result.content.contains("secret "),
"the model-visible replay must not expose the rejected sensitive chunk"
);
let observation: serde_json::Value =
serde_json::from_str(&tool_result.content).expect("result_read observation");
let detail = &observation["model_observation"]["detail"];
let source_result_ref = self
.source_result_ref
.lock()
.expect("source result ref lock poisoned")
.clone()
.expect("source result ref captured");
assert_eq!(
detail["result_ref"].as_str(),
Some(source_result_ref.as_str()),
"suppressed preview replay must retain the durable source result reference"
);
assert_ne!(
detail["result_ref"], observation["result_ref"],
"the inline result_read invocation ref must never become continuation authority"
);
assert!(
detail["total_bytes"]
.as_u64()
.is_some_and(|total_bytes| total_bytes > 2048),
"suppressed preview replay must retain total bytes: {}",
tool_result.content
);
assert_eq!(
detail["next_offset"].as_u64(),
Some(2048),
"suppressed preview replay must retain the next continuation offset"
);
return Ok(HostManagedModelResponse::assistant_reply("tool ok"));
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_reborn_composition/src/runtime/tests/core.rs` around lines
1035 - 1071, Strengthen the suppressed replay assertion in the
suppress_result_read branch of the test harness by verifying that
tool_result.content does not contain the sensitive chunk marker "secret " before
parsing its JSON metadata. Keep the existing preview-field and continuation
metadata assertions unchanged, and ensure the test exercises the real caller
path already covered by this branch.

Sources: Coding guidelines, Path instructions

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Closing for now. Railway testing showed that the live path still resolves the authenticated user and fans out into per-PR detail calls, so the PR’s one-call behavior is not achieved. The current regression fixture validates a scripted trajectory rather than live model selection.

This optimization is nice to have, and a proper fix would require broader changes to the model-visible GitHub contract and realistic canary coverage. We can revisit it separately if tool-call churn becomes a priority.

@serrrfirat serrrfirat closed this Aug 4, 2026

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6953 — 7988b071 Deployed Aug 4, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants