Skip to content

feat(stress): scripted tool-call workload with durable write read-back (#7360) - #7382

Merged
serrrfirat merged 5 commits into
nearai:mainfrom
serrrfirat:review-issue-7360
Aug 8, 2026
Merged

serrrfirat merged 5 commits into
nearai:mainfrom
serrrfirat:review-issue-7360

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Aug 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Phase 1 of Expand stress coverage across built-in and durable write paths #7360: scripted tool-call workloads for the API stress scenario — the mock LLM sidecar now drives deterministic builtin/memory tool sequences, and the driver verifies durable write read-back through the production path.
  • New --api-scripted-tool mode on api-user-capacity with read-back verdicts (confirmed / contended / leak / missing / undisclosed), per-document-size buckets (4 KiB–1 MiB), submit-to-tool-visible / submit-to-finalize stage timings, and --api-hot-writers same-user contention.
  • Nightly leg added to the hosted-single-tenant Postgres stress job (server stays up; leg rebinds the mock sidecar on the same port).
  • 40 new unit tests (scripted state machine, verdict computation incl. leak precedence, disclosure fallback, wire-shape rig deserialization, timeline helpers, summary buckets, flag validation).

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Part of #7360 — Phase 1 foundation slice: scripted-tool harness + persistent-memory roundtrip coverage. Reindex/churn workloads, the libSQL leg, and Phases 2–3 remain tracked in the issue.

Validation

  • cargo fmt -p ironclaw_stress -- --check
  • cargo clippy -p ironclaw_stress --all-targets --all-features -- -D warnings
  • cargo build -p ironclaw_stress --release
  • Relevant tests pass: cargo test -p ironclaw_stress (119 passed)
  • cargo test -p <owning-crate> --features integration — Not applicable: no crate-level integration behavior changed; E2E path covered by the new nightly stress job (postgres-api-capacity scripted leg), which requires a live server + Postgres and runs on schedule
  • Manual testing: full local E2E — real hosted-single-tenant server (dist build) + local Postgres 16 + pgvector, scripted leg driven through the real capability host: memory_roundtrip 9/9 clean (6 confirmed + 3 contended under one hot writer, 0 leaks), memory_grow 8/8 confirmed, per-size buckets and stage latencies present. E2E caught a hot-writer client_action_id collision (409 duplicate) that is fixed and covered by the rerun.
  • review-pr / pr-shepherd — not run (not available in this session)

Test Strategy

User behavior: Not applicable — developer tooling (ironclaw_stress) and CI; no product UX change.

Risk areas:

  • Model behavior — mock LLM only, scripted/deterministic
  • Browser
  • Side effect — durable writes (memory docs, workspace files) are the workload under test; verified via read-back verdicts
  • Persistence — memory + workspace write paths exercised through the production capability host; both read-back and cross-user isolation verified per op
  • Security or permissions — tools exercised through normal origin gates; auto-approve enabled via the production settings API, same as a real user
  • External provider
  • Cross-component behavior — HTTP → turn → capability host → scoped filesystem → event/projection path (API-driven by design)

Tests added or updated:

  • Unit or contract: 31 tests — marker parsing/roundtrip, per-op step sequencing, tool-name resolution (encoded/dotted/bare), verdict computation (confirmed/contended/leak/missing, leak precedence), disclosure fallback (placeholder → undisclosed), no-redrive after completion, content padding determinism, timeline tool/verdict/placeholder helpers, per-size summary buckets, --api-* flag validation
  • Reborn integration: Not applicable — the scripted leg is the integration surface; it runs in the nightly Postgres stress job
  • Recorded fixture: Not applicable
  • Browser E2E: Not applicable
  • Backend or runtime: new nightly postgres-api-capacity scripted leg (memory_roundtrip, 4 sizes, 2 hot writers, failure-rate + p95 thresholds; loose first-run ceilings pending baseline artifact)
  • Live canary: Not applicable

What the tests prove: the scripted state machine emits the right tool call per step from the conversation alone; read-back verdicts correctly distinguish confirmed/contended/leak/missing/undisclosed; operations are not redriven after completion; the driver attributes per-op outcomes to document size and stage; invalid flag combinations fail fast.

Commands run: cargo fmt -p ironclaw_stress -- --check, cargo clippy -p ironclaw_stress --all-targets --all-features -- -D warnings, cargo test -p ironclaw_stress, cargo build -p ironclaw_stress --release

Compatibility / rollback notes

  • New flags are opt-in; existing api-user-capacity runs unchanged (sidecar falls through to text when no marker is present).
  • Summary JSON gains optional scripted sections (skipped when absent) — --compare-json consumers unaffected.
  • CI: new steps only inside the existing nightly-only Postgres job.
  • Follow-ups (in issue Expand stress coverage across built-in and durable write paths #7360): libSQL leg of the capability-write matrix, memory_grow/memory_mixed/write_file_roundtrip nightly legs, Phase 2 (workspace/spill/trigger/config/extension families), Phase 3 inventory ratchet, per-case threshold tightening after the first baseline.

nearai#7360)

Phase 1 of issue nearai#7360: teach the stress harness to drive real builtin
and memory tool calls through the production capability path and verify
their durable side effects.

The api-user-capacity mock LLM sidecar learns a deterministic scripted
state machine: the driver embeds an `ironclaw-stress-tool` marker in the
user message, the sidecar emits the scripted tool call for a tool
advertised in the request, the server executes it through the real
capability host, and the driver verifies the read-back verdict in the
final assistant message. Verdicts: confirmed / contended (same-user CAS
race, counted) / leak (cross-user isolation, hard failure) / missing
(write lost, hard failure) / undisclosed (tool never advertised).

Scripts: write_file_roundtrip (write_file + read_file of a unique
workspace path), memory_roundtrip / memory_grow / memory_mixed
(ironclaw.memory.write/replace-append + read of the shared
stress/shared.md target — every run doubles as a same-relative-path
isolation check). --api-scripted-doc-sizes cycles 4 KiB..1 MiB
documents with per-size buckets and submit-to-tool-visible /
submit-to-finalize stage latencies; --api-hot-writers spawns concurrent
same-user writers for hot-document CAS contention. Gated tools are
exercised through the per-user Tools auto-approve setting enabled during
setup via the production settings API.

Wired as a nightly leg in the hosted-single-tenant Postgres job (the
existing server stays up; the leg rebinds the mock sidecar on the same
port). Unit coverage: marker parsing, per-op step sequencing, tool-name
resolution (encoded/dotted/bare), verdict computation incl. leak
precedence, disclosure fallback, timeline helpers, per-size summary
buckets, and flag validation.
@github-actions github-actions Bot added scope: ci CI/CD workflows scope: docs Documentation size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules labels Aug 7, 2026
@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 09f2bfc2-4998-4fec-9bbb-bd39056f4a48

📥 Commits

Reviewing files that changed from the base of the PR and between 1e8ceb8 and 760190c.

📒 Files selected for processing (1)
  • tools/ironclaw_stress/src/tests.rs

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added scripted tool-workload testing for API capacity scenarios.
    • Supports file and shared-memory round trips with configurable document sizes.
    • Added concurrent hot-writer testing for contention and data-isolation checks.
    • Reports operation outcomes, tool results, failure categories, stage latency, and verdicts.
    • Stress runs now generate JSONL data, summaries, and report artifacts with failure-rate and p95 latency checks.
  • Documentation

    • Added setup guidance, configuration requirements, verdict definitions, and examples.
  • Tests

    • Expanded coverage for workload validation, sequencing, sizing, result handling, and configuration behavior.

Walkthrough

The stress runner adds deterministic scripted tool workloads for file and durable-memory operations. It adds mock-LLM tool-call handling, hot-writer contention, verdict and latency aggregation, CLI validation, documentation, and a hosted Postgres CI workload.

Changes

Scripted API stress workload

Layer / File(s) Summary
Deterministic scripted workload engine
tools/ironclaw_stress/src/scripted.rs
Defines scripted operation sequences, markers, bounded arguments, conversation decisions, tool-call handling, verdict classification, and unit tests.
CLI configuration and validation
tools/ironclaw_stress/src/main.rs, tools/ironclaw_stress/src/tests.rs
Adds scripted-tool, document-size, and hot-writer options. Validates scenario, tool, mock LLM, polling, size, and workload constraints.
Capacity runner execution and reporting
tools/ironclaw_stress/src/api_capacity.rs
Runs scripted foreground, background, and hot-writer flows. Enables tool auto-approval, handles scripted mock responses, polls timelines, and aggregates samples by document size and verdict.
Postgres workload and operating documentation
.github/workflows/ironclaw-stress.yml, tools/ironclaw_stress/README.md
Adds a hosted Postgres scripted stress invocation, artifact uploads, and usage documentation for sizing, contention, verdicts, and required flags.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant StressRunner
  participant IronclawAPI
  participant MockLLM
  participant DurableMemory
  StressRunner->>IronclawAPI: submit scripted marker
  IronclawAPI->>MockLLM: request completion with advertised tools
  MockLLM-->>IronclawAPI: return tool call
  IronclawAPI->>DurableMemory: execute memory_roundtrip
  DurableMemory-->>IronclawAPI: return tool result
  IronclawAPI-->>StressRunner: expose timeline verdict
Loading

Possibly related issues

Possibly related PRs

  • nearai/ironclaw#7062 — The scripted memory_roundtrip workload exercises caller-scoped memory and workspace behavior implemented by this PR.

Suggested reviewers: ilblackdragon

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description is detailed but omits several required template sections, including Security Impact, Database Impact, Blast Radius, Reborn Trust-Boundary Checklist, and Review Follow-Through. Add the missing required sections with explicit impact, trust-boundary, database, blast-radius, rollback, review follow-through, and review-track details.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title follows Conventional Commits style and accurately summarizes the scripted tool-call stress workload.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the contributor: core 20+ merged PRs label Aug 7, 2026
@ironloopai

ironloopai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

🧭 IronLoop Run · Review

This comment updates in place as the Run moves through its stages.

⬛ Final result · Stopped

🟨 Queued → 🟦 Working → ⬛ Stopped

Automatic trigger · attempt 1 of 3 · stopped after 7m 42s

IronLoop stopped because the pull request target branch or head changed while this Run was active.

Run details

Run: 1e2d00aa-b9a4-4291-926b-ef1c2352223c
Base: main at 254483d
Head: review-issue-7360 at e586c67
Created: 2026-08-07 22:12 UTC
Updated: 2026-08-07 22:20 UTC

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 18

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/ironclaw-stress.yml:
- Around line 392-422: Update the workflow’s server lifecycle around the
capacity-step server startup and the “Run hosted-single-tenant Postgres scripted
tool writes” step so the server remains running when the scripted leg targets
port 18080. Move the server startup into the same run block as the scripted
invocation, or remove the per-step EXIT trap and add an explicit teardown step
guarded by if: always(); preserve cleanup while ensuring the scripted operations
execute against a live server.

In `@tools/ironclaw_stress/src/api_capacity.rs`:
- Around line 964-1007: Remove the unused fifth tuple element from the match in
the foreground flow, including the Some(())/None values and the let _ =
scripted_op discard. Apply the same cleanup to the corresponding match in
run_background_user, preserving the existing scripted_samples push and other
returned values.
- Around line 1320-1324: Update the verdict handling around my_final_content to
call scripted::parse_result_verdict only once, store the parsed result, assign
verdict from it, and reuse that result when constructing failure instead of
reparsing the same content.
- Around line 623-626: Move the tool_results accumulation into the existing
first bucket lookup block around the size-bucket processing, reusing the bucket
borrowed there instead of calling get_mut again. Add sample.tools_executed as
u64 to that bucket alongside the other metrics, and remove the later lookup and
conditional block.
- Around line 728-753: Move hot-writer spawning from writer_tasks into a
dedicated JoinSet, while preserving the existing run_hot_writer arguments and
concurrency behavior. Keep the regular writer drain/refill loop limited to
regular virtual-user tasks, then drain the hot-writer JoinSet after that loop
completes so hot-writer completion never triggers run_virtual_user refills.
- Around line 2687-2795: Add rig-shape deserialization assertions in the
existing response-shape test near mock_completion_response: validate
mock_tool_call_response as
rig::providers::openai::completion::CompletionResponse, and validate the
header/body/done tuple from mock_streaming_tool_call_chunks using the
corresponding rig streaming chunk type or parser. Assert deserialization
succeeds and preserves the tool-call data needed by the client.
- Around line 2611-2620: The completion request path currently parses messages
twice by calling scripted::latest_op and scripted::decide separately. Update
scripted::decide to return the decision together with the latest operation, or
otherwise expose both results from one conversation parse, then use that
combined result in the shown flow while preserving the existing no-operation
behavior.
- Around line 333-337: Update the scripted_counters locking in the summary path
to recover poisoned mutexes with the mutex error’s into_inner() value, matching
scripted_decision_for, instead of falling back to a default
ScriptedMockCounters. Preserve cloning the recovered counters before releasing
the lock.
- Around line 1433-1446: Update timeline_finalized_verdict_content to match
verdict_prefix only when it is followed by a space, rather than using an
unrestricted contains check. Preserve the existing finalized assistant-message
filtering and return the full matching content.
- Around line 610-638: Sort each stage’s latency values before passing them to
latency_summary in the stage_latencies aggregation loop. Update the values
handling around latency_summary(&values) so percentile calculations and min/max
use ascending order, while preserving the existing stage_latency_us insertion
behavior.
- Around line 1037-1040: Namespace scripted marker identities by user cohort so
foreground and background users cannot collide. Update the background identity
construction near ScriptedTaskIdentity to give marker_user or op_prefix a
distinct background namespace, and apply the matching namespace change in the
foreground path around its identity setup; preserve the existing per-user
operation numbering while ensuring every cross-cohort identity is distinct.
- Around line 1296-1302: Replace the paginated absolute count from
timeline_tool_result_count in the tools_seen update with operation-scoped
evidence: inspect tool-result messages returned by the timeline and count only
those matching the current operation’s result/token context. Ensure
tools_executed is derived from this operation-specific count rather than
subtracting prev_tool_result_count from a page-limited cumulative value.

In `@tools/ironclaw_stress/src/main.rs`:
- Around line 233-234: The script key should have one typed source of truth
instead of duplicated raw-string mappings. In tools/ironclaw_stress/src/main.rs
lines 233-234, derive clap::ValueEnum for ScriptKey, change the CLI field to
Option<ScriptKey>, remove the hand-written validation in validate_args, and
update the three ScriptKey::parse re-derivations in api_capacity.rs to use the
typed value directly. In tools/ironclaw_stress/src/scripted.rs lines 86-113,
remove parse and known_keys, or derive them from a single ALL slice only if
parse_marker still requires standalone wire-format parsing; preserve exhaustive
as_str and steps handling.
- Around line 1152-1159: Raise the minimum accepted scripted document size in
the argument validation loop to a shared exported floor constant defined beside
scripted::MAX_SCRIPTED_DOC_SIZE_BYTES, choosing a value that survives the /4
memory_grow split and token prefix. Update the validation error text and
scripted_api_rejects_out_of_range_doc_sizes to use the new lower-bound wording.

In `@tools/ironclaw_stress/src/scripted.rs`:
- Around line 539-551: Update resolve_wire_name to remove the bare-name
candidate from resolution, matching only the encoded and dotted capability
identifiers. Preserve the existing candidate lookup order and return behavior so
namespaced capabilities cannot bind to an unrelated extension tool.
- Around line 611-616: Update parse_result_verdict to locate RESULT_PREFIX plus
the operation identity with substring matching, using find before parsing the
trailing verdict instead of requiring strip_prefix at the start. Keep its
parsing behavior unchanged after the marker, so it remains consistent with
conversation_has_result and still returns None for missing or unparsable
verdicts.
- Around line 406-436: Update the verdict calculation in the scripted decision
flow around compute_verdict so it receives only the tool result corresponding to
the read step, excluding MemoryWrite/WriteFile results identified by the
operation’s step table. Preserve the existing step sequencing and verdict
behavior, and add coverage for a write-result echo containing the operation
token followed by a read result without that token, asserting the Missing
verdict.

In `@tools/ironclaw_stress/src/tests.rs`:
- Around line 200-208: Extend scripted_api_rejects_out_of_range_doc_sizes to
also validate the upper boundary by setting api_scripted_doc_sizes to
MAX_SCRIPTED_DOC_SIZE_BYTES + 1 and asserting the same range-validation error.
Add an empty-vector case and assert it returns the expected validation error for
missing scripted document sizes, covering both untested branches in
validate_args.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: dfbe9d7a-d6f7-45c1-bd08-62a19eb37d33

📥 Commits

Reviewing files that changed from the base of the PR and between 254483d and e586c67.

📒 Files selected for processing (6)
  • .github/workflows/ironclaw-stress.yml
  • tools/ironclaw_stress/README.md
  • tools/ironclaw_stress/src/api_capacity.rs
  • tools/ironclaw_stress/src/main.rs
  • tools/ironclaw_stress/src/scripted.rs
  • tools/ironclaw_stress/src/tests.rs

Comment thread .github/workflows/ironclaw-stress.yml Outdated
Comment thread tools/ironclaw_stress/src/api_capacity.rs Outdated
Comment thread tools/ironclaw_stress/src/api_capacity.rs
Comment thread tools/ironclaw_stress/src/api_capacity.rs Outdated
Comment thread tools/ironclaw_stress/src/api_capacity.rs Outdated
Comment thread tools/ironclaw_stress/src/main.rs
Comment thread tools/ironclaw_stress/src/scripted.rs Outdated
Comment thread tools/ironclaw_stress/src/scripted.rs
Comment thread tools/ironclaw_stress/src/scripted.rs
Comment thread tools/ironclaw_stress/src/tests.rs
…ver lifecycle

Review fixes for the nearai#7360 Phase 1 scripted workload:

- Hot writers now run on distinct threads of the first user instead of
  sharing one thread, so concurrent operations exercise real per-user
  memory-document CAS contention rather than per-thread turn
  serialization. setup_users creates and records one extra thread per
  hot writer for user 0; run_hot_writer picks its own thread.
- Scripted document sizes are floored at 4 KiB (the token-dominated
  region below is meaningless and the issue's workloads start there);
  enforced in marker parsing and --api-scripted-doc-sizes validation.
- The CI scripted leg runs inside the server's run block so the trap
  does not kill the server before it starts; artifacts upload together.
- Wire-shape tests: mock_tool_call_response deserializes as the rig
  OpenAI CompletionResponse (stringified arguments, finish_reason
  tool_calls) and streaming tool-call chunks carry indexed delta
  tool_calls.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (5)
tools/ironclaw_stress/src/tests.rs (1)

151-227: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Drive one scripted validation case through the CLI caller.

The added cases call validate_args directly. They do not cover clap parsing, default resolution, or the startup path that gates workload execution. Add one command-level test for a scripted flag combination.

As per path instructions, test through the caller when a helper gates a side effect.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tools/ironclaw_stress/src/tests.rs` around lines 151 - 227, Add a
command-level test that constructs the scripted API flag combination through the
CLI parser, then runs the resulting command through the startup path that gates
workload execution instead of calling validate_args directly. Cover a valid or
invalid scripted configuration and assert the observable caller outcome,
including clap parsing/default resolution and validation before the workload
starts.

Source: Path instructions

tools/ironclaw_stress/src/scripted.rs (3)

321-330: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Reject unbounded and incompatible marker identities.

parse_marker accepts any non-whitespace user and op. readback_tokens recognizes only ASCII letters, digits, -, _, and . after the marker. For example, u0/x__1 is accepted here, but its token is truncated at /, so the scripted call cannot produce a matching verdict. A long identity can also expand readback_token() and tool arguments before dispatch.

Validate both fields against one bounded grammar before constructing ScriptedOp. This also prevents scripted_content from returning a token larger than the configured document size.

As per path instructions, new ingress must validate and bound the original payload before prompt construction or dispatch.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tools/ironclaw_stress/src/scripted.rs` around lines 321 - 330, Tighten
parse_marker’s identity validation before constructing ScriptedOp: require both
user and op to use only the readback-compatible ASCII letters, digits, '-', '_',
and '.', and enforce the configured length bound for each field or the complete
identity. Reject invalid or oversized identities before prompt construction or
dispatch, preserving the existing None-return behavior.

Source: Path instructions


585-587: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Preserve exact configured sizes across split writes.

fraction_size floors each fraction independently. For size_bytes = 4097, MemoryGrow writes 1024 + 3072 = 4096, and MemoryMixed writes 2048 + 2048 = 4096, while the operation remains labeled 4097. The reported size bucket no longer describes the durable document.

Compute chunks from cumulative fraction boundaries so the remainder is assigned once, or reject non-divisible sizes. Add a regression test for 4097.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tools/ironclaw_stress/src/scripted.rs` around lines 585 - 587, Update
fraction_size and the split-write callers to preserve the configured total size
by deriving each chunk from cumulative fraction boundaries, assigning any
remainder exactly once; alternatively reject sizes not divisible by the
denominator. Add a regression test covering size 4097 and verify MemoryGrow and
MemoryMixed persist exactly that size.

575-582: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Add caller-level memory-read coverage.

document-read.v1 requires path, so the wire shape here is valid. The owning invariant still requires test through the caller: add the missing caller-level read-back test for ironclaw.memory.read, not just generated JSON and memory-write assertions.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tools/ironclaw_stress/src/scripted.rs` around lines 575 - 582, Add a
caller-level read-back test for ironclaw.memory.read, covering the full path
from the caller through the StepKind::MemoryRead wire request and validating the
returned shared-memory content. Keep the existing generated-JSON and
memory-write assertions unchanged.

Source: Path instructions

tools/ironclaw_stress/src/main.rs (1)

1163-1165: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Reject hot writers for write_file_roundtrip.

The validation checks only that a script exists. run_hot_writer in tools/ironclaw_stress/src/api_capacity.rs Lines 916-943 changes the operation identity, and build_arguments in tools/ironclaw_stress/src/scripted.rs Lines 559-582 includes that identity in stress/<identity>.txt. Each hot writer therefore uses a different file.

This combination cannot measure the shared-memory CAS contention advertised by --api-hot-writers. Restrict the option to memory scripts, or change the option contract.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tools/ironclaw_stress/src/main.rs` around lines 1163 - 1165, The validation
around api_hot_writers must reject write_file_roundtrip scripts, not merely
require an API script. Update the argument checks in the main validation flow to
allow --api-hot-writers only for memory scripts, preserving the existing error
behavior for missing scripted tools and preventing distinct file identities from
being used for hot-writer contention tests.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@tools/ironclaw_stress/src/main.rs`:
- Around line 1163-1165: The validation around api_hot_writers must reject
write_file_roundtrip scripts, not merely require an API script. Update the
argument checks in the main validation flow to allow --api-hot-writers only for
memory scripts, preserving the existing error behavior for missing scripted
tools and preventing distinct file identities from being used for hot-writer
contention tests.

In `@tools/ironclaw_stress/src/scripted.rs`:
- Around line 321-330: Tighten parse_marker’s identity validation before
constructing ScriptedOp: require both user and op to use only the
readback-compatible ASCII letters, digits, '-', '_', and '.', and enforce the
configured length bound for each field or the complete identity. Reject invalid
or oversized identities before prompt construction or dispatch, preserving the
existing None-return behavior.
- Around line 585-587: Update fraction_size and the split-write callers to
preserve the configured total size by deriving each chunk from cumulative
fraction boundaries, assigning any remainder exactly once; alternatively reject
sizes not divisible by the denominator. Add a regression test covering size 4097
and verify MemoryGrow and MemoryMixed persist exactly that size.
- Around line 575-582: Add a caller-level read-back test for
ironclaw.memory.read, covering the full path from the caller through the
StepKind::MemoryRead wire request and validating the returned shared-memory
content. Keep the existing generated-JSON and memory-write assertions unchanged.

In `@tools/ironclaw_stress/src/tests.rs`:
- Around line 151-227: Add a command-level test that constructs the scripted API
flag combination through the CLI parser, then runs the resulting command through
the startup path that gates workload execution instead of calling validate_args
directly. Cover a valid or invalid scripted configuration and assert the
observable caller outcome, including clap parsing/default resolution and
validation before the workload starts.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 089ab14c-68d9-49f9-8da9-faf1af4c1c7f

📥 Commits

Reviewing files that changed from the base of the PR and between e586c67 and 990846e.

📒 Files selected for processing (6)
  • .github/workflows/ironclaw-stress.yml
  • tools/ironclaw_stress/README.md
  • tools/ironclaw_stress/src/api_capacity.rs
  • tools/ironclaw_stress/src/main.rs
  • tools/ironclaw_stress/src/scripted.rs
  • tools/ironclaw_stress/src/tests.rs

…iter

A hot writer and the first user's regular writer shared the same user
label and operation index, so their client_action_id values were
identical and the server rejected the second submit with a 409
duplicate conflict. Include the scripted op prefix (h{k}-) in the
operation ref so concurrent writers always submit distinct action ids.

Found by a full local E2E run of the scripted leg against a real
hosted-single-tenant server: after the fix, memory_roundtrip with one
hot writer runs 9/9 clean (6 confirmed + 3 contended, 0 leaks) and
memory_grow runs 8/8 confirmed.
… tool counts, typed script key (nearai#7382)

- compute_verdict: verdict comes from read steps only (write echoes can no
  longer mask missing/contended)
- timeline tool evidence: count tool results by sequence above the op's
  baseline instead of subtracting page-limited absolute counts
- timeline verdict match: delimit prefix by trailing space so op 1 cannot
  terminate on op 10's message; parse_result_verdict aligns on substring
- background users namespace markers as b{index} so cross-cohort leaks
  cannot read back as their own token
- hot writers drain in a dedicated JoinSet (no run_virtual_user refills)
- fraction chunks derive from cumulative boundaries so split writes persist
  exactly the configured size (regression test at 4097)
- ScriptKey derives clap::ValueEnum: CLI, marker wire format, and parsing
  share one string mapping; --api-hot-writers rejects write_file_roundtrip
- parse_marker bounds identity grammar; poisoned mutex recovery; sorted
  stage latencies; single conversation parse per completion request;
  CLI-level scripted validation test and doc-size bound coverage

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tools/ironclaw_stress/src/tests.rs`:
- Around line 251-275: Extend the validation test around validate_args to cover
the --api-wait-for-assistant gate: construct or update arguments with
api_wait_for_assistant enabled and mock_llm_bind unset, call validate_args, and
assert it rejects the configuration with the expected missing-sidecar error.
Keep the existing scripted-tool assertion unchanged.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c381a6c1-0432-42cc-a7c0-1a358b12159d

📥 Commits

Reviewing files that changed from the base of the PR and between 0bd9f6e and 1e8ceb8.

📒 Files selected for processing (4)
  • tools/ironclaw_stress/src/api_capacity.rs
  • tools/ironclaw_stress/src/main.rs
  • tools/ironclaw_stress/src/scripted.rs
  • tools/ironclaw_stress/src/tests.rs

Comment thread tools/ironclaw_stress/src/tests.rs
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Addressed CodeRabbit review — all 18 inline + 5 outside-diff comments:

Correctness

  • Verdict input restricted to read steps: a write result echoing the operation's own token can no longer mask missing/contended (regression test added)
  • Timeline tool evidence is operation-scoped: counts tool results by sequence above the op's baseline instead of subtracting page-limited absolute counts, so later ops no longer false-fail scripted_tools_not_executed
  • Verdict prefix matching delimited by trailing space (u0__1 no longer matches u0__10); parse_result_verdict aligns on substring like conversation_has_result
  • Background users namespace markers as b{index} so a cross-cohort isolation violation cannot read back as its own token
  • Hot writers drain in a dedicated JoinSet — their completion never triggers run_virtual_user refills, keeping --concurrency a hard bound
  • Split writes derive chunks from cumulative fraction boundaries so MemoryGrow/MemoryMixed persist exactly the configured size (regression test at 4097)
  • parse_marker validates identity grammar (readback-compatible charset + 64-byte cap) before constructing the op
  • --api-hot-writers now rejects write_file_roundtrip (per-operation paths cannot contend)

Maintainability

  • ScriptKey derives clap::ValueEnum: CLI flag, marker wire format, and parsing share one string mapping; typed Option<ScriptKey> field replaces three ScriptKey::parse re-derivations
  • decide_with_op returns decision + op from a single conversation parse (no double parse on the mock hot path); latest_op removed
  • Dead scripted_op tuple slot removed from both user loops; tool_results folded into the first bucket; stage latencies sorted before latency_summary; poisoned-mutex recovery via into_inner; bare-name fallback dropped from resolve_wire_name (extension tool shadowing)

Tests

  • CLI-level test drives the scripted flag combination through clap parsing + validation gate (unknown key rejected by clap, missing sidecar and no-polling gates covered)
  • Doc-size validation covers both bounds and the empty-list guard

Skipped (already covered at review time)

  • Caller-level memory-read test: memory_roundtrip_read_step_after_one_tool_result + memory_roundtrip_final_verdict_confirmed already drive the ironclaw.memory.read wire request and read-back through decide

CI is green on the head commit (Code Style, Reborn E2E, libsql bottleneck suite, Hooks parity, Regression enforcement all pass).

@serrrfirat
serrrfirat added this pull request to the merge queue Aug 8, 2026
Merged via the queue into nearai:main with commit 30ae2d5 Aug 8, 2026
41 checks passed
@serrrfirat
serrrfirat deleted the review-issue-7360 branch August 8, 2026 17:08
serrrfirat added a commit that referenced this pull request Aug 9, 2026
- Merge origin/main (#7377 run-acts-as-invoker, #7323, #7382, #6938,
  #7280, #7393, #7389, #7364, #7228, #7371, #7399).
- main's #7377 landed a narrower terminal arm (generic failure notice for
  TurnStatus::Failed only); keep the #6896 arm, which covers Failed and
  RecoveryRequired with sanitized per-category summaries plus Cancelled
  and the timeout grace path, and adapt to the Option<String>
  notice_discriminator main introduced.
- Re-seed the composition budget to the merged-tree measurement
  (40811 -> 40861, the run-failure settlement observer lands +50 governed
  LOC); the arch-test record moves with the manifest.
personal-upstream-sync Bot pushed a commit to theredspoon/ironclaw that referenced this pull request Aug 10, 2026
…rai#6896) (nearai#7131)

* fix(run_delivery): deliver triggered run failures to the creator (nearai#6896)

Scheduled/triggered runs that ended in Failed, Cancelled, or
RecoveryRequired produced no user-visible notification: the triggered
delivery driver minted notifications only for Completed /
BlockedApproval / BlockedAuth and recorded every other terminal status
as Skipped. A run that timed out before reaching an actionable state
only logged a warn and recorded Failed, leaving the creator in silence.

Delivery:
- triggered_notification_for_state now mints a FinalReplyReady
  notification for Failed and RecoveryRequired using the existing
  per-category failure summaries (reborn_failure_summary_for_category)
  over state.failure.category(), with a generic fallback when no
  category is present.
- Cancelled mints the same notification, preferring a failure-category
  summary when one is present and falling back to a fixed cancellation
  notice otherwise.
- The RunWaitTimedOut branch with no prior blocked marker now delivers
  the timeout notice as a terminal reply instead of recording Failed.
- The wildcard arm is replaced with explicit non-actionable statuses
  (Queued, Running, CancelRequested, BlockedResource,
  BlockedDependentRun, BlockedExternalTool) so a future status fails to
  compile rather than silently skipping.

Observer:
- TriggerFireSettlementObserver gains on_failed_fire_settled as a
  default no-op method, plus a TriggerFailedFireSettlement event
  carrying tenant/trigger/fire-slot/run-id/history-status. Noop and
  existing implementors keep compiling.
- The active-cleanup sweep fires on_failed_fire_settled when
  clear_active_fire succeeds with TriggerRunHistoryStatus::Error, so
  post-accept failures are observable for automation health. Ok,
  Running, and already-cleared fires do not fire the hook.

Tests:
- run_delivery_contract: Failed+model_error, Failed without category,
  Cancelled, and timeout-before-actionable all assert a Delivered
  outcome with the expected notice text and footer.
- worker tests: a terminal-Error active fire fires exactly one
  on_failed_fire_settled; a terminal-Ok active fire fires none.

The larger retry/redrive budget for failed post-accept fires
(retry_disposition has zero production callers) is intentionally left
for a follow-up; it is out of scope for this surgical delivery fix.

* style: cargo fmt the nearai#6896 delivery fix

* fix(triggers): address terminal delivery review feedback

* fix(assistant): drop unused UserId import after merge

* fix(run_delivery): address multi-agent review findings

- Extract shared terminal-notice helpers (final_reply_notice,
  outcome_for_delivery_failure, deliver_terminal_notice) so the
  timeout, OAuth-backstop, and generic failure arms share one notice
  shape and outcome taxonomy instead of a third hand-rolled copy.
- Add a bounded race-grace window after the wait backstop: a run that
  crosses into a terminal state during the final wait (cancellation in
  flight, failure landing after the last poll) now delivers the correct
  terminal notice instead of the timeout copy.
- Cancelled runs always deliver the fixed cancellation notice; the
  failure-category branch was unreachable in production and would have
  mislabeled a host/operator cancel as a failure.
- Update the stale invariant doc, the five-output surface contract
  count, and the exhaustiveness-only comment on the non-actionable arm.
- Document the cheap/non-blocking contract on
  TriggerFireSettlementObserver (the worker awaits it inline in the
  poller sweep) and note it at the active-cleanup call site.
- Add contract coverage for the timeout arm's delivery-failure outcome
  (Failed) and a regression test proving the race-grace path delivers
  the cancellation notice; the cancelled-with-category test now asserts
  the cancellation notice wins.

* fix(run_delivery): address review comments and restore CI gates

Review fixes (CodeRabbit on 01e887f/f8af109):
- Grace loop fails loud: log the bound TurnError on state-poll failure and
  the RunDeliveryError on terminal-notice build failure before falling back
  to the timeout copy, with silent-ok markers on both intentional fallbacks.
- Hoist TriggeredReplyTargetAuthority, CodecChannelTargetResolver, and
  TriggeredNotificationContext to one construction before the watcher loop;
  the race-grace arm, timeout arm, and loop body now share it.
- Collapse the duplicated failure-summary expression into one closure and
  name TurnStatus::Failed explicitly so future statuses are compiler-visible.
- Drop the stale "Only three states" count from the surface-contract doc.
- Test fixture: encode the late-terminal flip as one Option<(usize,
  ScriptedRunState)> field instead of two correlated Options with an expect.
- Terminal-crossing test: document why flip_after=30 deterministically
  outruns the wait poll budget and assert the grace loop issues no
  cancellation (cancel_calls == 0).

CI:
- composition-budget: re-seed loc_ceiling 40432 -> 40593 (measured on the
  merged tree; the nearai#7131 settlement observer adds +161 governed LOC of
  wiring) and move the arch-test record with it.
- trigger_poller: use the colon-form tracing target required by nearai#7146.

* ci: re-trigger pull_request workflows for c2460ed

* fix(composition): capture the settlement health warn in the observer test

The traced_test default filter is {crate}=trace, which drops events whose
metadata target is `ironclaw::reborn::…`. The observer warning is emitted
with the colon-form target (required by nearai#7146 — the equals form recorded a
field and never matched RUST_LOG target filters), so the test saw an empty
buffer. Enable tracing-test's no-env-filter feature, the same pattern the
capabilities/host-runtime/mcp/loop crates use for cross-target assertions.

Re-seed the composition budget to the merged-tree measurement (40747 ->
40867): nearai#7131's observer wiring lands on top of post-measurement mainline
inflow; measured with the gate, set to current. The arch-test record moves
with the manifest.

* fix(run_delivery): merge main and adapt to notice_discriminator String

- Merge origin/main (nearai#7377 run-acts-as-invoker, nearai#7323, nearai#7382, nearai#6938,
  nearai#7280, nearai#7393, nearai#7389, nearai#7364, nearai#7228, nearai#7371, nearai#7399).
- main's nearai#7377 landed a narrower terminal arm (generic failure notice for
  TurnStatus::Failed only); keep the nearai#6896 arm, which covers Failed and
  RecoveryRequired with sanitized per-category summaries plus Cancelled
  and the timeout grace path, and adapt to the Option<String>
  notice_discriminator main introduced.
- Re-seed the composition budget to the merged-tree measurement
  (40811 -> 40861, the run-failure settlement observer lands +50 governed
  LOC); the arch-test record moves with the manifest.
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
nearai#7360) (nearai#7382)

* feat(stress): scripted tool-call workload with durable write read-back (nearai#7360)

Phase 1 of issue nearai#7360: teach the stress harness to drive real builtin
and memory tool calls through the production capability path and verify
their durable side effects.

The api-user-capacity mock LLM sidecar learns a deterministic scripted
state machine: the driver embeds an `ironclaw-stress-tool` marker in the
user message, the sidecar emits the scripted tool call for a tool
advertised in the request, the server executes it through the real
capability host, and the driver verifies the read-back verdict in the
final assistant message. Verdicts: confirmed / contended (same-user CAS
race, counted) / leak (cross-user isolation, hard failure) / missing
(write lost, hard failure) / undisclosed (tool never advertised).

Scripts: write_file_roundtrip (write_file + read_file of a unique
workspace path), memory_roundtrip / memory_grow / memory_mixed
(ironclaw.memory.write/replace-append + read of the shared
stress/shared.md target — every run doubles as a same-relative-path
isolation check). --api-scripted-doc-sizes cycles 4 KiB..1 MiB
documents with per-size buckets and submit-to-tool-visible /
submit-to-finalize stage latencies; --api-hot-writers spawns concurrent
same-user writers for hot-document CAS contention. Gated tools are
exercised through the per-user Tools auto-approve setting enabled during
setup via the production settings API.

Wired as a nightly leg in the hosted-single-tenant Postgres job (the
existing server stays up; the leg rebinds the mock sidecar on the same
port). Unit coverage: marker parsing, per-op step sequencing, tool-name
resolution (encoded/dotted/bare), verdict computation incl. leak
precedence, disclosure fallback, timeline helpers, per-size summary
buckets, and flag validation.

* fix(stress): hot writers on distinct user threads, size floor, CI server lifecycle

Review fixes for the nearai#7360 Phase 1 scripted workload:

- Hot writers now run on distinct threads of the first user instead of
  sharing one thread, so concurrent operations exercise real per-user
  memory-document CAS contention rather than per-thread turn
  serialization. setup_users creates and records one extra thread per
  hot writer for user 0; run_hot_writer picks its own thread.
- Scripted document sizes are floored at 4 KiB (the token-dominated
  region below is meaningless and the issue's workloads start there);
  enforced in marker parsing and --api-scripted-doc-sizes validation.
- The CI scripted leg runs inside the server's run block so the trap
  does not kill the server before it starts; artifacts upload together.
- Wire-shape tests: mock_tool_call_response deserializes as the rig
  OpenAI CompletionResponse (stringified arguments, finish_reason
  tool_calls) and streaming tool-call chunks carry indexed delta
  tool_calls.

* fix(stress): hot-writer client action ids collide with the primary writer

A hot writer and the first user's regular writer shared the same user
label and operation index, so their client_action_id values were
identical and the server rejected the second submit with a 409
duplicate conflict. Include the scripted op prefix (h{k}-) in the
operation ref so concurrent writers always submit distinct action ids.

Found by a full local E2E run of the scripted leg against a real
hosted-single-tenant server: after the fix, memory_roundtrip with one
hot writer runs 9/9 clean (6 confirmed + 3 contended, 0 leaks) and
memory_grow runs 8/8 confirmed.

* fix(stress): address coderabbit review — verdict integrity, op-scoped tool counts, typed script key (nearai#7382)

- compute_verdict: verdict comes from read steps only (write echoes can no
  longer mask missing/contended)
- timeline tool evidence: count tool results by sequence above the op's
  baseline instead of subtracting page-limited absolute counts
- timeline verdict match: delimit prefix by trailing space so op 1 cannot
  terminate on op 10's message; parse_result_verdict aligns on substring
- background users namespace markers as b{index} so cross-cohort leaks
  cannot read back as their own token
- hot writers drain in a dedicated JoinSet (no run_virtual_user refills)
- fraction chunks derive from cumulative boundaries so split writes persist
  exactly the configured size (regression test at 4097)
- ScriptKey derives clap::ValueEnum: CLI, marker wire format, and parsing
  share one string mapping; --api-hot-writers rejects write_file_roundtrip
- parse_marker bounds identity grammar; poisoned mutex recovery; sorted
  stage latencies; single conversation parse per completion request;
  CLI-level scripted validation test and doc-size bound coverage

* test(stress): cover --api-wait-for-assistant gate in CLI-level scripted test (nearai#7382)
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
nearai#7360) (nearai#7382)

* feat(stress): scripted tool-call workload with durable write read-back (nearai#7360)

Phase 1 of issue nearai#7360: teach the stress harness to drive real builtin
and memory tool calls through the production capability path and verify
their durable side effects.

The api-user-capacity mock LLM sidecar learns a deterministic scripted
state machine: the driver embeds an `ironclaw-stress-tool` marker in the
user message, the sidecar emits the scripted tool call for a tool
advertised in the request, the server executes it through the real
capability host, and the driver verifies the read-back verdict in the
final assistant message. Verdicts: confirmed / contended (same-user CAS
race, counted) / leak (cross-user isolation, hard failure) / missing
(write lost, hard failure) / undisclosed (tool never advertised).

Scripts: write_file_roundtrip (write_file + read_file of a unique
workspace path), memory_roundtrip / memory_grow / memory_mixed
(ironclaw.memory.write/replace-append + read of the shared
stress/shared.md target — every run doubles as a same-relative-path
isolation check). --api-scripted-doc-sizes cycles 4 KiB..1 MiB
documents with per-size buckets and submit-to-tool-visible /
submit-to-finalize stage latencies; --api-hot-writers spawns concurrent
same-user writers for hot-document CAS contention. Gated tools are
exercised through the per-user Tools auto-approve setting enabled during
setup via the production settings API.

Wired as a nightly leg in the hosted-single-tenant Postgres job (the
existing server stays up; the leg rebinds the mock sidecar on the same
port). Unit coverage: marker parsing, per-op step sequencing, tool-name
resolution (encoded/dotted/bare), verdict computation incl. leak
precedence, disclosure fallback, timeline helpers, per-size summary
buckets, and flag validation.

* fix(stress): hot writers on distinct user threads, size floor, CI server lifecycle

Review fixes for the nearai#7360 Phase 1 scripted workload:

- Hot writers now run on distinct threads of the first user instead of
  sharing one thread, so concurrent operations exercise real per-user
  memory-document CAS contention rather than per-thread turn
  serialization. setup_users creates and records one extra thread per
  hot writer for user 0; run_hot_writer picks its own thread.
- Scripted document sizes are floored at 4 KiB (the token-dominated
  region below is meaningless and the issue's workloads start there);
  enforced in marker parsing and --api-scripted-doc-sizes validation.
- The CI scripted leg runs inside the server's run block so the trap
  does not kill the server before it starts; artifacts upload together.
- Wire-shape tests: mock_tool_call_response deserializes as the rig
  OpenAI CompletionResponse (stringified arguments, finish_reason
  tool_calls) and streaming tool-call chunks carry indexed delta
  tool_calls.

* fix(stress): hot-writer client action ids collide with the primary writer

A hot writer and the first user's regular writer shared the same user
label and operation index, so their client_action_id values were
identical and the server rejected the second submit with a 409
duplicate conflict. Include the scripted op prefix (h{k}-) in the
operation ref so concurrent writers always submit distinct action ids.

Found by a full local E2E run of the scripted leg against a real
hosted-single-tenant server: after the fix, memory_roundtrip with one
hot writer runs 9/9 clean (6 confirmed + 3 contended, 0 leaks) and
memory_grow runs 8/8 confirmed.

* fix(stress): address coderabbit review — verdict integrity, op-scoped tool counts, typed script key (nearai#7382)

- compute_verdict: verdict comes from read steps only (write echoes can no
  longer mask missing/contended)
- timeline tool evidence: count tool results by sequence above the op's
  baseline instead of subtracting page-limited absolute counts
- timeline verdict match: delimit prefix by trailing space so op 1 cannot
  terminate on op 10's message; parse_result_verdict aligns on substring
- background users namespace markers as b{index} so cross-cohort leaks
  cannot read back as their own token
- hot writers drain in a dedicated JoinSet (no run_virtual_user refills)
- fraction chunks derive from cumulative boundaries so split writes persist
  exactly the configured size (regression test at 4097)
- ScriptKey derives clap::ValueEnum: CLI, marker wire format, and parsing
  share one string mapping; --api-hot-writers rejects write_file_roundtrip
- parse_marker bounds identity grammar; poisoned mutex recovery; sorted
  stage latencies; single conversation parse per completion request;
  CLI-level scripted validation test and doc-size bound coverage

* test(stress): cover --api-wait-for-assistant gate in CLI-level scripted test (nearai#7382)
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
nearai#7360) (nearai#7382)

* feat(stress): scripted tool-call workload with durable write read-back (nearai#7360)

Phase 1 of issue nearai#7360: teach the stress harness to drive real builtin
and memory tool calls through the production capability path and verify
their durable side effects.

The api-user-capacity mock LLM sidecar learns a deterministic scripted
state machine: the driver embeds an `ironclaw-stress-tool` marker in the
user message, the sidecar emits the scripted tool call for a tool
advertised in the request, the server executes it through the real
capability host, and the driver verifies the read-back verdict in the
final assistant message. Verdicts: confirmed / contended (same-user CAS
race, counted) / leak (cross-user isolation, hard failure) / missing
(write lost, hard failure) / undisclosed (tool never advertised).

Scripts: write_file_roundtrip (write_file + read_file of a unique
workspace path), memory_roundtrip / memory_grow / memory_mixed
(ironclaw.memory.write/replace-append + read of the shared
stress/shared.md target — every run doubles as a same-relative-path
isolation check). --api-scripted-doc-sizes cycles 4 KiB..1 MiB
documents with per-size buckets and submit-to-tool-visible /
submit-to-finalize stage latencies; --api-hot-writers spawns concurrent
same-user writers for hot-document CAS contention. Gated tools are
exercised through the per-user Tools auto-approve setting enabled during
setup via the production settings API.

Wired as a nightly leg in the hosted-single-tenant Postgres job (the
existing server stays up; the leg rebinds the mock sidecar on the same
port). Unit coverage: marker parsing, per-op step sequencing, tool-name
resolution (encoded/dotted/bare), verdict computation incl. leak
precedence, disclosure fallback, timeline helpers, per-size summary
buckets, and flag validation.

* fix(stress): hot writers on distinct user threads, size floor, CI server lifecycle

Review fixes for the nearai#7360 Phase 1 scripted workload:

- Hot writers now run on distinct threads of the first user instead of
  sharing one thread, so concurrent operations exercise real per-user
  memory-document CAS contention rather than per-thread turn
  serialization. setup_users creates and records one extra thread per
  hot writer for user 0; run_hot_writer picks its own thread.
- Scripted document sizes are floored at 4 KiB (the token-dominated
  region below is meaningless and the issue's workloads start there);
  enforced in marker parsing and --api-scripted-doc-sizes validation.
- The CI scripted leg runs inside the server's run block so the trap
  does not kill the server before it starts; artifacts upload together.
- Wire-shape tests: mock_tool_call_response deserializes as the rig
  OpenAI CompletionResponse (stringified arguments, finish_reason
  tool_calls) and streaming tool-call chunks carry indexed delta
  tool_calls.

* fix(stress): hot-writer client action ids collide with the primary writer

A hot writer and the first user's regular writer shared the same user
label and operation index, so their client_action_id values were
identical and the server rejected the second submit with a 409
duplicate conflict. Include the scripted op prefix (h{k}-) in the
operation ref so concurrent writers always submit distinct action ids.

Found by a full local E2E run of the scripted leg against a real
hosted-single-tenant server: after the fix, memory_roundtrip with one
hot writer runs 9/9 clean (6 confirmed + 3 contended, 0 leaks) and
memory_grow runs 8/8 confirmed.

* fix(stress): address coderabbit review — verdict integrity, op-scoped tool counts, typed script key (nearai#7382)

- compute_verdict: verdict comes from read steps only (write echoes can no
  longer mask missing/contended)
- timeline tool evidence: count tool results by sequence above the op's
  baseline instead of subtracting page-limited absolute counts
- timeline verdict match: delimit prefix by trailing space so op 1 cannot
  terminate on op 10's message; parse_result_verdict aligns on substring
- background users namespace markers as b{index} so cross-cohort leaks
  cannot read back as their own token
- hot writers drain in a dedicated JoinSet (no run_virtual_user refills)
- fraction chunks derive from cumulative boundaries so split writes persist
  exactly the configured size (regression test at 4097)
- ScriptKey derives clap::ValueEnum: CLI, marker wire format, and parsing
  share one string mapping; --api-hot-writers rejects write_file_roundtrip
- parse_marker bounds identity grammar; poisoned mutex recovery; sorted
  stage latencies; single conversation parse per completion request;
  CLI-level scripted validation test and doc-size bound coverage

* test(stress): cover --api-wait-for-assistant gate in CLI-level scripted test (nearai#7382)
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…rai#6896) (nearai#7131)

* fix(run_delivery): deliver triggered run failures to the creator (nearai#6896)

Scheduled/triggered runs that ended in Failed, Cancelled, or
RecoveryRequired produced no user-visible notification: the triggered
delivery driver minted notifications only for Completed /
BlockedApproval / BlockedAuth and recorded every other terminal status
as Skipped. A run that timed out before reaching an actionable state
only logged a warn and recorded Failed, leaving the creator in silence.

Delivery:
- triggered_notification_for_state now mints a FinalReplyReady
  notification for Failed and RecoveryRequired using the existing
  per-category failure summaries (reborn_failure_summary_for_category)
  over state.failure.category(), with a generic fallback when no
  category is present.
- Cancelled mints the same notification, preferring a failure-category
  summary when one is present and falling back to a fixed cancellation
  notice otherwise.
- The RunWaitTimedOut branch with no prior blocked marker now delivers
  the timeout notice as a terminal reply instead of recording Failed.
- The wildcard arm is replaced with explicit non-actionable statuses
  (Queued, Running, CancelRequested, BlockedResource,
  BlockedDependentRun, BlockedExternalTool) so a future status fails to
  compile rather than silently skipping.

Observer:
- TriggerFireSettlementObserver gains on_failed_fire_settled as a
  default no-op method, plus a TriggerFailedFireSettlement event
  carrying tenant/trigger/fire-slot/run-id/history-status. Noop and
  existing implementors keep compiling.
- The active-cleanup sweep fires on_failed_fire_settled when
  clear_active_fire succeeds with TriggerRunHistoryStatus::Error, so
  post-accept failures are observable for automation health. Ok,
  Running, and already-cleared fires do not fire the hook.

Tests:
- run_delivery_contract: Failed+model_error, Failed without category,
  Cancelled, and timeout-before-actionable all assert a Delivered
  outcome with the expected notice text and footer.
- worker tests: a terminal-Error active fire fires exactly one
  on_failed_fire_settled; a terminal-Ok active fire fires none.

The larger retry/redrive budget for failed post-accept fires
(retry_disposition has zero production callers) is intentionally left
for a follow-up; it is out of scope for this surgical delivery fix.

* style: cargo fmt the nearai#6896 delivery fix

* fix(triggers): address terminal delivery review feedback

* fix(assistant): drop unused UserId import after merge

* fix(run_delivery): address multi-agent review findings

- Extract shared terminal-notice helpers (final_reply_notice,
  outcome_for_delivery_failure, deliver_terminal_notice) so the
  timeout, OAuth-backstop, and generic failure arms share one notice
  shape and outcome taxonomy instead of a third hand-rolled copy.
- Add a bounded race-grace window after the wait backstop: a run that
  crosses into a terminal state during the final wait (cancellation in
  flight, failure landing after the last poll) now delivers the correct
  terminal notice instead of the timeout copy.
- Cancelled runs always deliver the fixed cancellation notice; the
  failure-category branch was unreachable in production and would have
  mislabeled a host/operator cancel as a failure.
- Update the stale invariant doc, the five-output surface contract
  count, and the exhaustiveness-only comment on the non-actionable arm.
- Document the cheap/non-blocking contract on
  TriggerFireSettlementObserver (the worker awaits it inline in the
  poller sweep) and note it at the active-cleanup call site.
- Add contract coverage for the timeout arm's delivery-failure outcome
  (Failed) and a regression test proving the race-grace path delivers
  the cancellation notice; the cancelled-with-category test now asserts
  the cancellation notice wins.

* fix(run_delivery): address review comments and restore CI gates

Review fixes (CodeRabbit on 01e887f/f8af109):
- Grace loop fails loud: log the bound TurnError on state-poll failure and
  the RunDeliveryError on terminal-notice build failure before falling back
  to the timeout copy, with silent-ok markers on both intentional fallbacks.
- Hoist TriggeredReplyTargetAuthority, CodecChannelTargetResolver, and
  TriggeredNotificationContext to one construction before the watcher loop;
  the race-grace arm, timeout arm, and loop body now share it.
- Collapse the duplicated failure-summary expression into one closure and
  name TurnStatus::Failed explicitly so future statuses are compiler-visible.
- Drop the stale "Only three states" count from the surface-contract doc.
- Test fixture: encode the late-terminal flip as one Option<(usize,
  ScriptedRunState)> field instead of two correlated Options with an expect.
- Terminal-crossing test: document why flip_after=30 deterministically
  outruns the wait poll budget and assert the grace loop issues no
  cancellation (cancel_calls == 0).

CI:
- composition-budget: re-seed loc_ceiling 40432 -> 40593 (measured on the
  merged tree; the nearai#7131 settlement observer adds +161 governed LOC of
  wiring) and move the arch-test record with it.
- trigger_poller: use the colon-form tracing target required by nearai#7146.

* ci: re-trigger pull_request workflows for c2460ed

* fix(composition): capture the settlement health warn in the observer test

The traced_test default filter is {crate}=trace, which drops events whose
metadata target is `ironclaw::reborn::…`. The observer warning is emitted
with the colon-form target (required by nearai#7146 — the equals form recorded a
field and never matched RUST_LOG target filters), so the test saw an empty
buffer. Enable tracing-test's no-env-filter feature, the same pattern the
capabilities/host-runtime/mcp/loop crates use for cross-target assertions.

Re-seed the composition budget to the merged-tree measurement (40747 ->
40867): nearai#7131's observer wiring lands on top of post-measurement mainline
inflow; measured with the gate, set to current. The arch-test record moves
with the manifest.

* fix(run_delivery): merge main and adapt to notice_discriminator String

- Merge origin/main (nearai#7377 run-acts-as-invoker, nearai#7323, nearai#7382, nearai#6938,
  nearai#7280, nearai#7393, nearai#7389, nearai#7364, nearai#7228, nearai#7371, nearai#7399).
- main's nearai#7377 landed a narrower terminal arm (generic failure notice for
  TurnStatus::Failed only); keep the nearai#6896 arm, which covers Failed and
  RecoveryRequired with sanitized per-category summaries plus Cancelled
  and the timeout grace path, and adapt to the Option<String>
  notice_discriminator main introduced.
- Re-seed the composition budget to the merged-tree measurement
  (40811 -> 40861, the run-failure settlement observer lands +50 governed
  LOC); the arch-test record moves with the manifest.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: ci CI/CD workflows scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants