Skip to content

test(stress): add API capacity admin-user harness - #5855

Merged
serrrfirat merged 3 commits into
mainfrom
codex/api-capacity-harness-only
Jul 9, 2026
Merged

serrrfirat merged 3 commits into
mainfrom
codex/api-capacity-harness-only

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

This is the harness-only base PR for the API capacity / chat latency work. It intentionally avoids the runtime performance changes so the measurement surface can be reviewed and merged independently.

Changes:

  • Extends ironclaw_stress --scenario api-user-capacity to provision many real API users through WebUI admin CRUD.
  • Adds admin provisioner fanout, per-user bearer tokens, read-hammer traffic (list_threads, timeline, session), and mock-LLM request timing stats.
  • Fixes the multi-operation wait semantics so each accepted send waits for that user's Nth finalized assistant message instead of accepting any old assistant message in the thread.
  • Documents the admin-user capacity shape in the stress README.

Benchmark Context

This PR is the measurement foundation. The current local baseline captured on the stacked runtime branch was:

  • c100/u100, 1 op/user, no read hammer: full-flow p95 1.85s, throughput 53.0 ops/sec.
  • c100/u100, 1 op/user, 2 read QPS/user: full-flow p95 1.97-1.99s, throughput 49.6-50.8 ops/sec.
  • read hammer: list_threads p95 771-887ms, timeline p95 835-876ms.
  • mock LLM: 100/100 requests, p95 ~2.5-3ms, max in-flight 5-6.

Those runtime gains are not part of this PR; they will be stacked on top.

Timeline So Far

  • Started from hosted-single-tenant Postgres/libSQL storage latency work where blob-shaped turn/thread/resource state made DB time dominate.
  • Moved hot stores toward row-shaped RootFilesystem-backed data and group-commit patterns.
  • Replaced direct Postgres-only store bypasses with unified filesystem-backed paths where appropriate.
  • Added production-shaped API capacity testing so we can test 100 users / 100 concurrency instead of unrealistic single-user-only pressure.
  • Found that with mock LLM the chat path is still dominated by pre-model host/context setup; the next PR starts isolating those runtime wins.

Verification

  • CARGO_TARGET_DIR=/Volumes/NVME/ironclaw-target-api-admin CARGO_INCREMENTAL=0 cargo test -p ironclaw_stress finalized_assistant -- --nocapture
  • CARGO_TARGET_DIR=/Volumes/NVME/ironclaw-target-api-admin CARGO_INCREMENTAL=0 cargo build -p ironclaw_stress

@ironloopai

ironloopai Bot commented Jul 8, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: 8b442f8b261653b4550571b0a3f234bd86a8ed4b
Result: One or more review results were superseded by a newer PR head.
Next: Run @ironloopai review on the latest PR head.
Updated: 2026-07-09T09:23:35.315Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Superseded N/A N/A 2026-07-09T09:23:35.292Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Superseded by a newer PR head. New head: 8b442f8. Previous verdict: Changes requested.
Recent activity
Time Reviewer State Detail
2026-07-09T09:17:19.366Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-09T09:17:19.378Z ironloop/common-reviewer (reviewer) Queued Added to the local review work handoff.
2026-07-09T09:17:20.468Z ironloop/common-reviewer (reviewer) Started Reviewer worker started attempt 1.
2026-07-09T09:17:23.280Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at 0c8c7f2.
2026-07-09T09:22:42.266Z ironloop/common-reviewer (reviewer) Superseded Old-head reviewer is still running after newer head 8b442f8 replaced it. Codex is reviewing; process live; elapsed 5m 20s; timeout in 14m 40s; last heartbeat 2026-07-09T09:22:42.266Z. Codex emitted stderr output at 2026-07-09T09:22:36.286Z.
2026-07-09T09:23:13.414Z ironloop/common-reviewer (reviewer) Result captured Changes requested; 1 blocking finding.
2026-07-09T09:23:13.414Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
2026-07-09T09:23:35.292Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (8b442f8).
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
  • @ironloopai status
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@coderabbitai

coderabbitai Bot commented Jul 8, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 711acdaa-a3ff-4ff0-b8af-6706b3f56c9e

📥 Commits

Reviewing files that changed from the base of the PR and between 1f697cc and 8b442f8.

📒 Files selected for processing (3)
  • tools/ironclaw_stress/src/api_capacity.rs
  • tools/ironclaw_stress/src/main.rs
  • tools/ironclaw_stress/src/tests.rs

📝 Walkthrough

Summary by CodeRabbit

  • New Features
    • Added api-user-capacity stress scenario for end-to-end API send/read pressure testing.
    • Added CLI support for provisioning API users via admin bearer token, plus configurable background user load (users, concurrency, operations, start delay) and mock background latency.
  • Bug Fixes
    • Improved run completion detection by waiting for a target count of finalized assistant timeline messages.
    • Enhanced request/concurrency and latency reporting, including bearer-aware request execution and admin token replacement.
  • Chores
    • Updated stress scenario documentation and example commands to reflect the new options and admin provisioning flow.

Walkthrough

Adds api-user-capacity stress-scenario support with admin-backed user provisioning, background API load controls, mock-LLM request metrics, and finalized-assistant-count waiting. Updates CLI validation, docs, and tests to match the new execution path.

Changes

ironclaw_stress api-user-capacity scenario

Layer / File(s) Summary
CLI flags, validation, and docs
tools/ironclaw_stress/Cargo.toml, tools/ironclaw_stress/src/main.rs, tools/ironclaw_stress/src/tests.rs, tools/ironclaw_stress/README.md
Adds admin provisioning flags, background load flags, mock background latency, scenario validation, startup labeling, and matching test/docs updates.
Mock LLM metrics and responses
tools/ironclaw_stress/src/api_capacity.rs
Reworks mock LLM state and summary tracking for request counts, concurrency, timing, and OpenAI-compatible completion payloads.
Window orchestration and finalized-count waits
tools/ironclaw_stress/src/api_capacity.rs
Splits writer/read/background execution, threads finalized-assistant expectations through virtual-user flows, and replaces assistant existence checks with count-based waiting.
Admin user provisioning path
tools/ironclaw_stress/src/api_capacity.rs
Adds typed admin create-user requests, bearer-aware request plumbing, concurrent user provisioning, and identity substitution before scenario startup.
Estimated code review effort: 4 (Complex) ~60 minutes

Possibly related PRs

  • nearai/ironclaw#5313: Introduced the ironclaw_stress harness and CLI structure extended here.
  • nearai/ironclaw#5781: Adds the api_capacity scenario plumbing that this PR extends with provisioning, timing, and finalized-count changes.

Suggested reviewers: italic-jinxin

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The summary and verification are present, but the template is missing several required sections, including Change Type and Linked Issue. Add the missing template sections, especially Change Type, a Linked Issue for this new feature, Security Impact, Blast Radius, Rollback Plan, and Review Follow-Through.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title is Conventional Commits-style and accurately summarizes the stress harness change.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5855 July 8, 2026 21:47 Destroyed
@github-actions github-actions Bot added scope: docs Documentation scope: dependencies Dependency updates size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 8, 2026
@serrrfirat
serrrfirat marked this pull request as ready for review July 8, 2026 21:49

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces the api-user-capacity scenario to the ironclaw_stress tool, enabling end-to-end WebUI API load testing with support for provisioning test users via an admin API. It also enhances the mock LLM server to track and report detailed request metrics such as latencies, concurrency, and request spreads. Feedback on these changes highlights a critical issue in the wait_for_assistant polling loop, which currently lacks any delay between iterations, resulting in a tight busy-waiting loop that could overwhelm both server and client resources.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines 661 to 679
let deadline = Instant::now() + Duration::from_millis(args.api_terminal_timeout_ms);
loop {
let timeline = harness.timeline(user).await;
let value = timeline.value.clone();
api_samples.push(timeline.sample);
match value {
Ok(value) => {
if timeline_has_finalized_assistant(&value) {
return None;
let finalized_count = timeline_finalized_assistant_count(&value);
if finalized_count >= target_finalized_count {
return Ok(finalized_count);
}
}
Err(failure) => return Some(failure),
Err(failure) => return Err(failure),
}
if Instant::now() >= deadline {
return Some(FailureCause::new(
return Err(FailureCause::new(
"api_full_flow_timeout",
"timeline",
format!(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The polling loop in wait_for_assistant continuously calls harness.timeline(user).await without any delay/sleep between iterations. This creates a tight busy-waiting loop that can exhaust connection pools, consume excessive CPU on the client, and hammer the server with unnecessary requests under stress testing.

Consider adding a small delay (e.g., 100ms or 200ms) between polling attempts using tokio::time::sleep as a pragmatic tradeoff for waiting in non-critical paths.

    let deadline = Instant::now() + Duration::from_millis(args.api_terminal_timeout_ms);
    loop {
        let timeline = harness.timeline(user).await;
        let value = timeline.value.clone();
        api_samples.push(timeline.sample);
        match value {
            Ok(value) => {
                let finalized_count = timeline_finalized_assistant_count(&value);
                if finalized_count >= target_finalized_count {
                    return Ok(finalized_count);
                }
            }
            Err(failure) => return Err(failure),
        }
        if Instant::now() >= deadline {
            return Err(FailureCause::new(
                "api_full_flow_timeout",
                "timeline",
                format!(
                    "timed out waiting for assistant message count {target_finalized_count}",
                ),
            ));
        }
        tokio::time::sleep(Duration::from_millis(200)).await;
    }
References
  1. A fixed-duration sleep (e.g., tokio::time::sleep) can be a pragmatic tradeoff for waiting on resource initialization in non-critical paths, especially if a fallback mechanism (like logging) exists and a more robust solution is out of scope.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

❌ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
❌ Changes requested 2 0 2 b634ecbdb93a

Head: b634ecbdb93a789fef162ee5338ad62a140506de
Next: Fix the blocking findings, push the PR branch, then re-run this reviewer.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

Found blocking issues in the new api-user-capacity assistant wait accounting. Static review only; Rust tests could not be run because cargo is unavailable in this environment.

Findings

Blocking: 2 / Notes: 0

Blocking findings

1. ❌ [MEDIUM] Cumulative assistant count cannot work with paged timelines

Location: tools/ironclaw_stress/src/api_capacity.rs:668-669
The new wait logic compares the cumulative per-user target against the number of finalized assistant messages in a single timeline page. The timeline endpoint is bounded by --api-page-size (default 50, server max 200), while expected_finalized_assistant_count grows with every operation in the thread. With default --operations 200, the count returned here tops out at the current page's assistants, so later operations time out even though their assistant replies finalized. Track the submitted message/run, use pagination, or maintain the baseline from a sufficiently complete source instead of comparing a cumulative count to one page.

2. ❌ [MEDIUM] Failed waits advance the expected assistant count

Location: tools/ironclaw_stress/src/api_capacity.rs:611-612
On any wait failure this returns target_finalized_count, so the next operation assumes the missing assistant reply exists. If the timeout or error was caused by a real model/server failure and no assistant is ever finalized, all following operations wait for an unreachable count and get reported as failures too. Keep the previous observed count on failure, or return the last observed finalized count from wait_for_assistant.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.
  4. Use @ironloopai status to check queued/running/completed/stale/stalled state while reviewers run.

Ok(value) => {
if timeline_has_finalized_assistant(&value) {
return None;
let finalized_count = timeline_finalized_assistant_count(&value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This compares a cumulative target to only the current timeline page. Since timeline responses are bounded by --api-page-size (default 50, server max 200), longer runs eventually time out despite successful assistant replies once older messages fall off the page. The wait needs to track the submitted message/run, page through history, or use a baseline from a complete count.

stages: None,
},
api_samples,
target_finalized_count,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Returning target_finalized_count on a wait failure advances state as if this assistant reply finalized. If the failure is real and no reply is ever added, every later operation waits for a count that cannot be reached and gets misreported as failed. Preserve the previous count or return the last observed count from the waiter.

@github-actions

github-actions Bot commented Jul 8, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.12% (283863 / 333481 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 333481 lines now vs 320188 at floor capture (+13293 lines, +4.15%) — not a material change

⚠️ 3 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts, ironclaw_skill_learning

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.12% — 283863 / 333481 lines

Per-crate breakdown (65 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 347
ironclaw_skill_learning 0% 0 / 61
ironclaw_wasm_sandbox_core 7.37% 7 / 95
ironclaw_runtime_policy 33.2% 80 / 241
ironclaw_event_projections 43.34% 673 / 1553
ironclaw_run_state 52.36% 222 / 424
ironclaw_authorization 53.54% 461 / 861
ironclaw_triggers 59.79% 1740 / 2910
ironclaw_observability 61.54% 16 / 26
ironclaw_webui_v2 62.62% 2632 / 4203
ironclaw_reborn_cli 62.84% 3816 / 6073
ironclaw_mcp 63.15% 581 / 920
ironclaw_reborn_migration 67.01% 1172 / 1749
ironclaw_memory 67.12% 747 / 1113
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_filesystem 67.44% 3815 / 5657
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.08% 1658 / 2238
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.61% 958 / 1284
ironclaw_extractors 74.72% 538 / 720
ironclaw_first_party_extensions 77.62% 5411 / 6971
ironclaw_llm 78.07% 19998 / 25614
ironclaw_product_context 78.57% 11 / 14
ironclaw_wasm_product_adapters 80.58% 1510 / 1874
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_reborn_openai_compat 80.95% 956 / 1181
ironclaw_memory_native 81.86% 3226 / 3941
ironclaw_wasm 82.54% 950 / 1151
ironclaw_secrets 82.7% 2791 / 3375
ironclaw_events 83.47% 1762 / 2111
ironclaw_processes 84.06% 965 / 1148
ironclaw_turns 84.26% 13127 / 15579
ironclaw_host_api 85.17% 2549 / 2993
ironclaw_product_workflow 85.34% 10748 / 12594
ironclaw_projects 85.92% 659 / 767
ironclaw_network 86.12% 670 / 778
ironclaw_common 86.46% 1514 / 1751
ironclaw_threads 86.62% 4132 / 4770
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_auth 86.94% 2995 / 3445
ironclaw_reborn_config 86.98% 1730 / 1989
ironclaw_product_adapters 86.98% 3207 / 3687
ironclaw_reborn_identity 87.03% 557 / 640
ironclaw_skills 87.35% 4336 / 4964
ironclaw_hooks 87.84% 9916 / 11289
ironclaw_product_adapter_registry 87.96% 526 / 598
ironclaw_reborn_traces 88.21% 11931 / 13526
ironclaw_extensions 88.26% 2631 / 2981
ironclaw_reborn_composition 88.77% 70143 / 79019
ironclaw_host_runtime 88.94% 17522 / 19700
ironclaw_reborn 89.14% 15796 / 17721
ironclaw_conversations 90% 2924 / 3249
ironclaw_approvals 90.51% 1507 / 1665
ironclaw_event_streams 91.48% 1009 / 1103
ironclaw_loop_support 92.44% 14731 / 15936
ironclaw_attachments 93.06% 630 / 677
ironclaw_resources 93.09% 4637 / 4981
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.87% 2452 / 2612
ironclaw_agent_loop 94.58% 8776 / 9279
ironclaw_safety 94.8% 3668 / 3869
ironclaw_first_party_extension_ports 95% 3094 / 3257
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (4 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_oauth v1-only: consumed only by root ironclaw (src/auth/oauth.rs); no crates/* dependents. Crate's own doc comment confirms v1-only. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tools/ironclaw_stress/src/api_capacity.rs`:
- Around line 222-236: The mutex handling in MockLlmState is silently swallowing
PoisonError in summary, record_request_start, and record_request_latency, which
can hide broken metrics. Update the MockLlmState methods so lock failures are
not defaulted or ignored: either propagate the mutex error from summary or log
it explicitly before returning, and make the
record_request_start/record_request_latency helpers emit a diagnostic when
locking fails instead of dropping the sample. Use the existing MockLlmState
method names to keep the change localized.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9941add1-7979-4e41-8cb0-92f04a69b886

📥 Commits

Reviewing files that changed from the base of the PR and between 88f8d17 and b634ecb.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock, !**/Cargo.lock
📒 Files selected for processing (5)
  • tools/ironclaw_stress/Cargo.toml
  • tools/ironclaw_stress/README.md
  • tools/ironclaw_stress/src/api_capacity.rs
  • tools/ironclaw_stress/src/main.rs
  • tools/ironclaw_stress/src/tests.rs

Comment on lines +222 to +236
let request_latencies = self
.request_latencies
.lock()
.map(|latencies| latencies.clone())
.unwrap_or_default();
let mut request_latency_us = request_latencies
.iter()
.map(Duration::as_micros)
.collect::<Vec<_>>();
request_latency_us.sort_unstable();
let request_start_offsets = self
.request_start_offsets
.lock()
.map(|offsets| offsets.clone())
.unwrap_or_default();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Silent-failure on mutex poisoning in MockLlmState.

summary() uses .unwrap_or_default() on Mutex::lock() results (lines 226, 236), and record_request_start/record_request_latency use if let Ok(...) (lines 262, 268). All three silently swallow PoisonError, producing empty metrics or dropped samples with no diagnostic. Per the "fail loud" invariant, errors should propagate or at minimum be logged — a poisoned mutex means the summary silently reports zero latencies/offsets.

🔧 Proposed fix: log on poison instead of silent default
     let request_latencies = self
         .request_latencies
         .lock()
-        .map(|latencies| latencies.clone())
-        .unwrap_or_default();
+        .map(|latencies| latencies.clone())
+        .unwrap_or_else(|e| {
+            eprintln!("mock llm: request_latencies mutex poisoned: {e}");
+            Vec::new()
+        });
     // ...
     let request_start_offsets = self
         .request_start_offsets
         .lock()
-        .map(|offsets| offsets.clone())
-        .unwrap_or_default();
+        .map(|offsets| offsets.clone())
+        .unwrap_or_else(|e| {
+            eprintln!("mock llm: request_start_offsets mutex poisoned: {e}");
+             Vec::new()
+         });

And for the record helpers:

 fn record_request_start(&self, offset: Duration) {
-    if let Ok(mut offsets) = self.request_start_offsets.lock() {
+    match self.request_start_offsets.lock() {
+        Ok(mut offsets) => offsets.push(offset),
+        Err(e) => eprintln!("mock llm: request_start_offsets mutex poisoned: {e}"),
     }
 }

 fn record_request_latency(&self, latency: Duration) {
-    if let Ok(mut latencies) = self.request_latencies.lock() {
+    match self.request_latencies.lock() {
+        Ok(mut latencies) => latencies.push(latency),
+        Err(e) => eprintln!("mock llm: request_latencies mutex poisoned: {e}"),
     }
 }

Also applies to: 261-271

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tools/ironclaw_stress/src/api_capacity.rs` around lines 222 - 236, The mutex
handling in MockLlmState is silently swallowing PoisonError in summary,
record_request_start, and record_request_latency, which can hide broken metrics.
Update the MockLlmState methods so lock failures are not defaulted or ignored:
either propagate the mutex error from summary or log it explicitly before
returning, and make the record_request_start/record_request_latency helpers emit
a diagnostic when locking fails instead of dropping the sample. Use the existing
MockLlmState method names to keep the change localized.

Source: Path instructions

@railway-app

railway-app Bot commented Jul 8, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5855 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 9, 2026 at 9:31 am

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5855 July 9, 2026 09:02 Destroyed
@serrrfirat
serrrfirat marked this pull request as ready for review July 9, 2026 09:17

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

❌ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
❌ Changes requested 1 0 1 1f697cc8fdeb

Head: 1f697cc8fdeb4ba4e8ddd925b08079d15174304e
Next: Fix the blocking findings, push the PR branch, then re-run this reviewer.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

Found a blocking correctness issue in the API capacity harness: the new assistant-count wait logic can time out under the default workload once timeline pagination hides older messages.

Findings

Blocking: 1 / Notes: 0

Blocking findings

1. ❌ [MEDIUM] Do not compare absolute assistant count to a single timeline page

Location: tools/ironclaw_stress/src/api_capacity.rs:696-698
target_finalized_count is accumulated over the whole thread, but harness.timeline() only requests the latest api_page_size messages. With the default page size of 50, each turn adds a user and assistant message, so after about 25 turns the current page can never contain enough finalized assistant messages to satisfy target_finalized_count; the default --operations 200 API capacity run will start timing out even when assistants are being produced. Page through history, raise the requested limit based on the target, or wait on a per-operation identifier/run instead of a whole-thread count from one page.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.
  4. Use @ironloopai status to check queued/running/completed/failed/superseded state while reviewers run.

if timeline_has_finalized_assistant(&value) {
return None;
let finalized_count = timeline_finalized_assistant_count(&value);
if finalized_count >= target_finalized_count {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This compares an absolute whole-thread assistant count to only the current timeline page. With the default api_page_size of 50, a user doing more than about 25 turns can never see finalized_count >= target_finalized_count, so the default --operations 200 run times out despite successful assistants. Please page through history, adjust the requested limit for the target, or wait on a per-operation/run identifier.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5855 July 9, 2026 09:23 Destroyed
@serrrfirat
serrrfirat merged commit 3adc58f into main Jul 9, 2026
64 checks passed
@serrrfirat
serrrfirat deleted the codex/api-capacity-harness-only branch July 9, 2026 14:50

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5855 — 8b442f8b Deployed Jul 9, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: dependencies Dependency updates scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant