feat(stress): scripted tool-call workload with durable write read-back (#7360) - #7382
Conversation
nearai#7360) Phase 1 of issue nearai#7360: teach the stress harness to drive real builtin and memory tool calls through the production capability path and verify their durable side effects. The api-user-capacity mock LLM sidecar learns a deterministic scripted state machine: the driver embeds an `ironclaw-stress-tool` marker in the user message, the sidecar emits the scripted tool call for a tool advertised in the request, the server executes it through the real capability host, and the driver verifies the read-back verdict in the final assistant message. Verdicts: confirmed / contended (same-user CAS race, counted) / leak (cross-user isolation, hard failure) / missing (write lost, hard failure) / undisclosed (tool never advertised). Scripts: write_file_roundtrip (write_file + read_file of a unique workspace path), memory_roundtrip / memory_grow / memory_mixed (ironclaw.memory.write/replace-append + read of the shared stress/shared.md target — every run doubles as a same-relative-path isolation check). --api-scripted-doc-sizes cycles 4 KiB..1 MiB documents with per-size buckets and submit-to-tool-visible / submit-to-finalize stage latencies; --api-hot-writers spawns concurrent same-user writers for hot-document CAS contention. Gated tools are exercised through the per-user Tools auto-approve setting enabled during setup via the production settings API. Wired as a nightly leg in the hosted-single-tenant Postgres job (the existing server stays up; the leg rebinds the mock sidecar on the same port). Unit coverage: marker parsing, per-op step sequencing, tool-name resolution (encoded/dotted/bare), verdict computation incl. leak precedence, disclosure fallback, timeline helpers, per-size summary buckets, and flag validation.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe stress runner adds deterministic scripted tool workloads for file and durable-memory operations. It adds mock-LLM tool-call handling, hot-writer contention, verdict and latency aggregation, CLI validation, documentation, and a hosted Postgres CI workload. ChangesScripted API stress workload
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant StressRunner
participant IronclawAPI
participant MockLLM
participant DurableMemory
StressRunner->>IronclawAPI: submit scripted marker
IronclawAPI->>MockLLM: request completion with advertised tools
MockLLM-->>IronclawAPI: return tool call
IronclawAPI->>DurableMemory: execute memory_roundtrip
DurableMemory-->>IronclawAPI: return tool result
IronclawAPI-->>StressRunner: expose timeline verdict
Possibly related issues
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 3 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
🧭 IronLoop Run · ReviewThis comment updates in place as the Run moves through its stages. ⬛ Final result · Stopped
Automatic trigger · attempt 1 of 3 · stopped after 7m 42s IronLoop stopped because the pull request target branch or head changed while this Run was active. |
There was a problem hiding this comment.
Actionable comments posted: 18
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/ironclaw-stress.yml:
- Around line 392-422: Update the workflow’s server lifecycle around the
capacity-step server startup and the “Run hosted-single-tenant Postgres scripted
tool writes” step so the server remains running when the scripted leg targets
port 18080. Move the server startup into the same run block as the scripted
invocation, or remove the per-step EXIT trap and add an explicit teardown step
guarded by if: always(); preserve cleanup while ensuring the scripted operations
execute against a live server.
In `@tools/ironclaw_stress/src/api_capacity.rs`:
- Around line 964-1007: Remove the unused fifth tuple element from the match in
the foreground flow, including the Some(())/None values and the let _ =
scripted_op discard. Apply the same cleanup to the corresponding match in
run_background_user, preserving the existing scripted_samples push and other
returned values.
- Around line 1320-1324: Update the verdict handling around my_final_content to
call scripted::parse_result_verdict only once, store the parsed result, assign
verdict from it, and reuse that result when constructing failure instead of
reparsing the same content.
- Around line 623-626: Move the tool_results accumulation into the existing
first bucket lookup block around the size-bucket processing, reusing the bucket
borrowed there instead of calling get_mut again. Add sample.tools_executed as
u64 to that bucket alongside the other metrics, and remove the later lookup and
conditional block.
- Around line 728-753: Move hot-writer spawning from writer_tasks into a
dedicated JoinSet, while preserving the existing run_hot_writer arguments and
concurrency behavior. Keep the regular writer drain/refill loop limited to
regular virtual-user tasks, then drain the hot-writer JoinSet after that loop
completes so hot-writer completion never triggers run_virtual_user refills.
- Around line 2687-2795: Add rig-shape deserialization assertions in the
existing response-shape test near mock_completion_response: validate
mock_tool_call_response as
rig::providers::openai::completion::CompletionResponse, and validate the
header/body/done tuple from mock_streaming_tool_call_chunks using the
corresponding rig streaming chunk type or parser. Assert deserialization
succeeds and preserves the tool-call data needed by the client.
- Around line 2611-2620: The completion request path currently parses messages
twice by calling scripted::latest_op and scripted::decide separately. Update
scripted::decide to return the decision together with the latest operation, or
otherwise expose both results from one conversation parse, then use that
combined result in the shown flow while preserving the existing no-operation
behavior.
- Around line 333-337: Update the scripted_counters locking in the summary path
to recover poisoned mutexes with the mutex error’s into_inner() value, matching
scripted_decision_for, instead of falling back to a default
ScriptedMockCounters. Preserve cloning the recovered counters before releasing
the lock.
- Around line 1433-1446: Update timeline_finalized_verdict_content to match
verdict_prefix only when it is followed by a space, rather than using an
unrestricted contains check. Preserve the existing finalized assistant-message
filtering and return the full matching content.
- Around line 610-638: Sort each stage’s latency values before passing them to
latency_summary in the stage_latencies aggregation loop. Update the values
handling around latency_summary(&values) so percentile calculations and min/max
use ascending order, while preserving the existing stage_latency_us insertion
behavior.
- Around line 1037-1040: Namespace scripted marker identities by user cohort so
foreground and background users cannot collide. Update the background identity
construction near ScriptedTaskIdentity to give marker_user or op_prefix a
distinct background namespace, and apply the matching namespace change in the
foreground path around its identity setup; preserve the existing per-user
operation numbering while ensuring every cross-cohort identity is distinct.
- Around line 1296-1302: Replace the paginated absolute count from
timeline_tool_result_count in the tools_seen update with operation-scoped
evidence: inspect tool-result messages returned by the timeline and count only
those matching the current operation’s result/token context. Ensure
tools_executed is derived from this operation-specific count rather than
subtracting prev_tool_result_count from a page-limited cumulative value.
In `@tools/ironclaw_stress/src/main.rs`:
- Around line 233-234: The script key should have one typed source of truth
instead of duplicated raw-string mappings. In tools/ironclaw_stress/src/main.rs
lines 233-234, derive clap::ValueEnum for ScriptKey, change the CLI field to
Option<ScriptKey>, remove the hand-written validation in validate_args, and
update the three ScriptKey::parse re-derivations in api_capacity.rs to use the
typed value directly. In tools/ironclaw_stress/src/scripted.rs lines 86-113,
remove parse and known_keys, or derive them from a single ALL slice only if
parse_marker still requires standalone wire-format parsing; preserve exhaustive
as_str and steps handling.
- Around line 1152-1159: Raise the minimum accepted scripted document size in
the argument validation loop to a shared exported floor constant defined beside
scripted::MAX_SCRIPTED_DOC_SIZE_BYTES, choosing a value that survives the /4
memory_grow split and token prefix. Update the validation error text and
scripted_api_rejects_out_of_range_doc_sizes to use the new lower-bound wording.
In `@tools/ironclaw_stress/src/scripted.rs`:
- Around line 539-551: Update resolve_wire_name to remove the bare-name
candidate from resolution, matching only the encoded and dotted capability
identifiers. Preserve the existing candidate lookup order and return behavior so
namespaced capabilities cannot bind to an unrelated extension tool.
- Around line 611-616: Update parse_result_verdict to locate RESULT_PREFIX plus
the operation identity with substring matching, using find before parsing the
trailing verdict instead of requiring strip_prefix at the start. Keep its
parsing behavior unchanged after the marker, so it remains consistent with
conversation_has_result and still returns None for missing or unparsable
verdicts.
- Around line 406-436: Update the verdict calculation in the scripted decision
flow around compute_verdict so it receives only the tool result corresponding to
the read step, excluding MemoryWrite/WriteFile results identified by the
operation’s step table. Preserve the existing step sequencing and verdict
behavior, and add coverage for a write-result echo containing the operation
token followed by a read result without that token, asserting the Missing
verdict.
In `@tools/ironclaw_stress/src/tests.rs`:
- Around line 200-208: Extend scripted_api_rejects_out_of_range_doc_sizes to
also validate the upper boundary by setting api_scripted_doc_sizes to
MAX_SCRIPTED_DOC_SIZE_BYTES + 1 and asserting the same range-validation error.
Add an empty-vector case and assert it returns the expected validation error for
missing scripted document sizes, covering both untested branches in
validate_args.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: dfbe9d7a-d6f7-45c1-bd08-62a19eb37d33
📒 Files selected for processing (6)
.github/workflows/ironclaw-stress.ymltools/ironclaw_stress/README.mdtools/ironclaw_stress/src/api_capacity.rstools/ironclaw_stress/src/main.rstools/ironclaw_stress/src/scripted.rstools/ironclaw_stress/src/tests.rs
…ver lifecycle Review fixes for the nearai#7360 Phase 1 scripted workload: - Hot writers now run on distinct threads of the first user instead of sharing one thread, so concurrent operations exercise real per-user memory-document CAS contention rather than per-thread turn serialization. setup_users creates and records one extra thread per hot writer for user 0; run_hot_writer picks its own thread. - Scripted document sizes are floored at 4 KiB (the token-dominated region below is meaningless and the issue's workloads start there); enforced in marker parsing and --api-scripted-doc-sizes validation. - The CI scripted leg runs inside the server's run block so the trap does not kill the server before it starts; artifacts upload together. - Wire-shape tests: mock_tool_call_response deserializes as the rig OpenAI CompletionResponse (stringified arguments, finish_reason tool_calls) and streaming tool-call chunks carry indexed delta tool_calls.
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (5)
tools/ironclaw_stress/src/tests.rs (1)
151-227: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winDrive one scripted validation case through the CLI caller.
The added cases call
validate_argsdirectly. They do not cover clap parsing, default resolution, or the startup path that gates workload execution. Add one command-level test for a scripted flag combination.As per path instructions, test through the caller when a helper gates a side effect.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tools/ironclaw_stress/src/tests.rs` around lines 151 - 227, Add a command-level test that constructs the scripted API flag combination through the CLI parser, then runs the resulting command through the startup path that gates workload execution instead of calling validate_args directly. Cover a valid or invalid scripted configuration and assert the observable caller outcome, including clap parsing/default resolution and validation before the workload starts.Source: Path instructions
tools/ironclaw_stress/src/scripted.rs (3)
321-330: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick winReject unbounded and incompatible marker identities.
parse_markeraccepts any non-whitespaceuserandop.readback_tokensrecognizes only ASCII letters, digits,-,_, and.after the marker. For example,u0/x__1is accepted here, but its token is truncated at/, so the scripted call cannot produce a matching verdict. A long identity can also expandreadback_token()and tool arguments before dispatch.Validate both fields against one bounded grammar before constructing
ScriptedOp. This also preventsscripted_contentfrom returning a token larger than the configured document size.As per path instructions, new ingress must validate and bound the original payload before prompt construction or dispatch.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tools/ironclaw_stress/src/scripted.rs` around lines 321 - 330, Tighten parse_marker’s identity validation before constructing ScriptedOp: require both user and op to use only the readback-compatible ASCII letters, digits, '-', '_', and '.', and enforce the configured length bound for each field or the complete identity. Reject invalid or oversized identities before prompt construction or dispatch, preserving the existing None-return behavior.Source: Path instructions
585-587: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick winPreserve exact configured sizes across split writes.
fraction_sizefloors each fraction independently. Forsize_bytes = 4097,MemoryGrowwrites1024 + 3072 = 4096, andMemoryMixedwrites2048 + 2048 = 4096, while the operation remains labeled4097. The reported size bucket no longer describes the durable document.Compute chunks from cumulative fraction boundaries so the remainder is assigned once, or reject non-divisible sizes. Add a regression test for
4097.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tools/ironclaw_stress/src/scripted.rs` around lines 585 - 587, Update fraction_size and the split-write callers to preserve the configured total size by deriving each chunk from cumulative fraction boundaries, assigning any remainder exactly once; alternatively reject sizes not divisible by the denominator. Add a regression test covering size 4097 and verify MemoryGrow and MemoryMixed persist exactly that size.
575-582: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick winAdd caller-level memory-read coverage.
document-read.v1requirespath, so the wire shape here is valid. The owning invariant still requires test through the caller: add the missing caller-level read-back test forironclaw.memory.read, not just generated JSON and memory-write assertions.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tools/ironclaw_stress/src/scripted.rs` around lines 575 - 582, Add a caller-level read-back test for ironclaw.memory.read, covering the full path from the caller through the StepKind::MemoryRead wire request and validating the returned shared-memory content. Keep the existing generated-JSON and memory-write assertions unchanged.Source: Path instructions
tools/ironclaw_stress/src/main.rs (1)
1163-1165: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick winReject hot writers for
write_file_roundtrip.The validation checks only that a script exists.
run_hot_writerintools/ironclaw_stress/src/api_capacity.rsLines 916-943 changes the operation identity, andbuild_argumentsintools/ironclaw_stress/src/scripted.rsLines 559-582 includes that identity instress/<identity>.txt. Each hot writer therefore uses a different file.This combination cannot measure the shared-memory CAS contention advertised by
--api-hot-writers. Restrict the option to memory scripts, or change the option contract.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tools/ironclaw_stress/src/main.rs` around lines 1163 - 1165, The validation around api_hot_writers must reject write_file_roundtrip scripts, not merely require an API script. Update the argument checks in the main validation flow to allow --api-hot-writers only for memory scripts, preserving the existing error behavior for missing scripted tools and preventing distinct file identities from being used for hot-writer contention tests.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@tools/ironclaw_stress/src/main.rs`:
- Around line 1163-1165: The validation around api_hot_writers must reject
write_file_roundtrip scripts, not merely require an API script. Update the
argument checks in the main validation flow to allow --api-hot-writers only for
memory scripts, preserving the existing error behavior for missing scripted
tools and preventing distinct file identities from being used for hot-writer
contention tests.
In `@tools/ironclaw_stress/src/scripted.rs`:
- Around line 321-330: Tighten parse_marker’s identity validation before
constructing ScriptedOp: require both user and op to use only the
readback-compatible ASCII letters, digits, '-', '_', and '.', and enforce the
configured length bound for each field or the complete identity. Reject invalid
or oversized identities before prompt construction or dispatch, preserving the
existing None-return behavior.
- Around line 585-587: Update fraction_size and the split-write callers to
preserve the configured total size by deriving each chunk from cumulative
fraction boundaries, assigning any remainder exactly once; alternatively reject
sizes not divisible by the denominator. Add a regression test covering size 4097
and verify MemoryGrow and MemoryMixed persist exactly that size.
- Around line 575-582: Add a caller-level read-back test for
ironclaw.memory.read, covering the full path from the caller through the
StepKind::MemoryRead wire request and validating the returned shared-memory
content. Keep the existing generated-JSON and memory-write assertions unchanged.
In `@tools/ironclaw_stress/src/tests.rs`:
- Around line 151-227: Add a command-level test that constructs the scripted API
flag combination through the CLI parser, then runs the resulting command through
the startup path that gates workload execution instead of calling validate_args
directly. Cover a valid or invalid scripted configuration and assert the
observable caller outcome, including clap parsing/default resolution and
validation before the workload starts.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 089ab14c-68d9-49f9-8da9-faf1af4c1c7f
📒 Files selected for processing (6)
.github/workflows/ironclaw-stress.ymltools/ironclaw_stress/README.mdtools/ironclaw_stress/src/api_capacity.rstools/ironclaw_stress/src/main.rstools/ironclaw_stress/src/scripted.rstools/ironclaw_stress/src/tests.rs
…iter
A hot writer and the first user's regular writer shared the same user
label and operation index, so their client_action_id values were
identical and the server rejected the second submit with a 409
duplicate conflict. Include the scripted op prefix (h{k}-) in the
operation ref so concurrent writers always submit distinct action ids.
Found by a full local E2E run of the scripted leg against a real
hosted-single-tenant server: after the fix, memory_roundtrip with one
hot writer runs 9/9 clean (6 confirmed + 3 contended, 0 leaks) and
memory_grow runs 8/8 confirmed.
… tool counts, typed script key (nearai#7382) - compute_verdict: verdict comes from read steps only (write echoes can no longer mask missing/contended) - timeline tool evidence: count tool results by sequence above the op's baseline instead of subtracting page-limited absolute counts - timeline verdict match: delimit prefix by trailing space so op 1 cannot terminate on op 10's message; parse_result_verdict aligns on substring - background users namespace markers as b{index} so cross-cohort leaks cannot read back as their own token - hot writers drain in a dedicated JoinSet (no run_virtual_user refills) - fraction chunks derive from cumulative boundaries so split writes persist exactly the configured size (regression test at 4097) - ScriptKey derives clap::ValueEnum: CLI, marker wire format, and parsing share one string mapping; --api-hot-writers rejects write_file_roundtrip - parse_marker bounds identity grammar; poisoned mutex recovery; sorted stage latencies; single conversation parse per completion request; CLI-level scripted validation test and doc-size bound coverage
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tools/ironclaw_stress/src/tests.rs`:
- Around line 251-275: Extend the validation test around validate_args to cover
the --api-wait-for-assistant gate: construct or update arguments with
api_wait_for_assistant enabled and mock_llm_bind unset, call validate_args, and
assert it rejects the configuration with the expected missing-sidecar error.
Keep the existing scripted-tool assertion unchanged.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: c381a6c1-0432-42cc-a7c0-1a358b12159d
📒 Files selected for processing (4)
tools/ironclaw_stress/src/api_capacity.rstools/ironclaw_stress/src/main.rstools/ironclaw_stress/src/scripted.rstools/ironclaw_stress/src/tests.rs
|
Addressed CodeRabbit review — all 18 inline + 5 outside-diff comments: Correctness
Maintainability
Tests
Skipped (already covered at review time)
CI is green on the head commit (Code Style, Reborn E2E, libsql bottleneck suite, Hooks parity, Regression enforcement all pass). |
- Merge origin/main (#7377 run-acts-as-invoker, #7323, #7382, #6938, #7280, #7393, #7389, #7364, #7228, #7371, #7399). - main's #7377 landed a narrower terminal arm (generic failure notice for TurnStatus::Failed only); keep the #6896 arm, which covers Failed and RecoveryRequired with sanitized per-category summaries plus Cancelled and the timeout grace path, and adapt to the Option<String> notice_discriminator main introduced. - Re-seed the composition budget to the merged-tree measurement (40811 -> 40861, the run-failure settlement observer lands +50 governed LOC); the arch-test record moves with the manifest.
…rai#6896) (nearai#7131) * fix(run_delivery): deliver triggered run failures to the creator (nearai#6896) Scheduled/triggered runs that ended in Failed, Cancelled, or RecoveryRequired produced no user-visible notification: the triggered delivery driver minted notifications only for Completed / BlockedApproval / BlockedAuth and recorded every other terminal status as Skipped. A run that timed out before reaching an actionable state only logged a warn and recorded Failed, leaving the creator in silence. Delivery: - triggered_notification_for_state now mints a FinalReplyReady notification for Failed and RecoveryRequired using the existing per-category failure summaries (reborn_failure_summary_for_category) over state.failure.category(), with a generic fallback when no category is present. - Cancelled mints the same notification, preferring a failure-category summary when one is present and falling back to a fixed cancellation notice otherwise. - The RunWaitTimedOut branch with no prior blocked marker now delivers the timeout notice as a terminal reply instead of recording Failed. - The wildcard arm is replaced with explicit non-actionable statuses (Queued, Running, CancelRequested, BlockedResource, BlockedDependentRun, BlockedExternalTool) so a future status fails to compile rather than silently skipping. Observer: - TriggerFireSettlementObserver gains on_failed_fire_settled as a default no-op method, plus a TriggerFailedFireSettlement event carrying tenant/trigger/fire-slot/run-id/history-status. Noop and existing implementors keep compiling. - The active-cleanup sweep fires on_failed_fire_settled when clear_active_fire succeeds with TriggerRunHistoryStatus::Error, so post-accept failures are observable for automation health. Ok, Running, and already-cleared fires do not fire the hook. Tests: - run_delivery_contract: Failed+model_error, Failed without category, Cancelled, and timeout-before-actionable all assert a Delivered outcome with the expected notice text and footer. - worker tests: a terminal-Error active fire fires exactly one on_failed_fire_settled; a terminal-Ok active fire fires none. The larger retry/redrive budget for failed post-accept fires (retry_disposition has zero production callers) is intentionally left for a follow-up; it is out of scope for this surgical delivery fix. * style: cargo fmt the nearai#6896 delivery fix * fix(triggers): address terminal delivery review feedback * fix(assistant): drop unused UserId import after merge * fix(run_delivery): address multi-agent review findings - Extract shared terminal-notice helpers (final_reply_notice, outcome_for_delivery_failure, deliver_terminal_notice) so the timeout, OAuth-backstop, and generic failure arms share one notice shape and outcome taxonomy instead of a third hand-rolled copy. - Add a bounded race-grace window after the wait backstop: a run that crosses into a terminal state during the final wait (cancellation in flight, failure landing after the last poll) now delivers the correct terminal notice instead of the timeout copy. - Cancelled runs always deliver the fixed cancellation notice; the failure-category branch was unreachable in production and would have mislabeled a host/operator cancel as a failure. - Update the stale invariant doc, the five-output surface contract count, and the exhaustiveness-only comment on the non-actionable arm. - Document the cheap/non-blocking contract on TriggerFireSettlementObserver (the worker awaits it inline in the poller sweep) and note it at the active-cleanup call site. - Add contract coverage for the timeout arm's delivery-failure outcome (Failed) and a regression test proving the race-grace path delivers the cancellation notice; the cancelled-with-category test now asserts the cancellation notice wins. * fix(run_delivery): address review comments and restore CI gates Review fixes (CodeRabbit on 01e887f/f8af109): - Grace loop fails loud: log the bound TurnError on state-poll failure and the RunDeliveryError on terminal-notice build failure before falling back to the timeout copy, with silent-ok markers on both intentional fallbacks. - Hoist TriggeredReplyTargetAuthority, CodecChannelTargetResolver, and TriggeredNotificationContext to one construction before the watcher loop; the race-grace arm, timeout arm, and loop body now share it. - Collapse the duplicated failure-summary expression into one closure and name TurnStatus::Failed explicitly so future statuses are compiler-visible. - Drop the stale "Only three states" count from the surface-contract doc. - Test fixture: encode the late-terminal flip as one Option<(usize, ScriptedRunState)> field instead of two correlated Options with an expect. - Terminal-crossing test: document why flip_after=30 deterministically outruns the wait poll budget and assert the grace loop issues no cancellation (cancel_calls == 0). CI: - composition-budget: re-seed loc_ceiling 40432 -> 40593 (measured on the merged tree; the nearai#7131 settlement observer adds +161 governed LOC of wiring) and move the arch-test record with it. - trigger_poller: use the colon-form tracing target required by nearai#7146. * ci: re-trigger pull_request workflows for c2460ed * fix(composition): capture the settlement health warn in the observer test The traced_test default filter is {crate}=trace, which drops events whose metadata target is `ironclaw::reborn::…`. The observer warning is emitted with the colon-form target (required by nearai#7146 — the equals form recorded a field and never matched RUST_LOG target filters), so the test saw an empty buffer. Enable tracing-test's no-env-filter feature, the same pattern the capabilities/host-runtime/mcp/loop crates use for cross-target assertions. Re-seed the composition budget to the merged-tree measurement (40747 -> 40867): nearai#7131's observer wiring lands on top of post-measurement mainline inflow; measured with the gate, set to current. The arch-test record moves with the manifest. * fix(run_delivery): merge main and adapt to notice_discriminator String - Merge origin/main (nearai#7377 run-acts-as-invoker, nearai#7323, nearai#7382, nearai#6938, nearai#7280, nearai#7393, nearai#7389, nearai#7364, nearai#7228, nearai#7371, nearai#7399). - main's nearai#7377 landed a narrower terminal arm (generic failure notice for TurnStatus::Failed only); keep the nearai#6896 arm, which covers Failed and RecoveryRequired with sanitized per-category summaries plus Cancelled and the timeout grace path, and adapt to the Option<String> notice_discriminator main introduced. - Re-seed the composition budget to the merged-tree measurement (40811 -> 40861, the run-failure settlement observer lands +50 governed LOC); the arch-test record moves with the manifest.
nearai#7360) (nearai#7382) * feat(stress): scripted tool-call workload with durable write read-back (nearai#7360) Phase 1 of issue nearai#7360: teach the stress harness to drive real builtin and memory tool calls through the production capability path and verify their durable side effects. The api-user-capacity mock LLM sidecar learns a deterministic scripted state machine: the driver embeds an `ironclaw-stress-tool` marker in the user message, the sidecar emits the scripted tool call for a tool advertised in the request, the server executes it through the real capability host, and the driver verifies the read-back verdict in the final assistant message. Verdicts: confirmed / contended (same-user CAS race, counted) / leak (cross-user isolation, hard failure) / missing (write lost, hard failure) / undisclosed (tool never advertised). Scripts: write_file_roundtrip (write_file + read_file of a unique workspace path), memory_roundtrip / memory_grow / memory_mixed (ironclaw.memory.write/replace-append + read of the shared stress/shared.md target — every run doubles as a same-relative-path isolation check). --api-scripted-doc-sizes cycles 4 KiB..1 MiB documents with per-size buckets and submit-to-tool-visible / submit-to-finalize stage latencies; --api-hot-writers spawns concurrent same-user writers for hot-document CAS contention. Gated tools are exercised through the per-user Tools auto-approve setting enabled during setup via the production settings API. Wired as a nightly leg in the hosted-single-tenant Postgres job (the existing server stays up; the leg rebinds the mock sidecar on the same port). Unit coverage: marker parsing, per-op step sequencing, tool-name resolution (encoded/dotted/bare), verdict computation incl. leak precedence, disclosure fallback, timeline helpers, per-size summary buckets, and flag validation. * fix(stress): hot writers on distinct user threads, size floor, CI server lifecycle Review fixes for the nearai#7360 Phase 1 scripted workload: - Hot writers now run on distinct threads of the first user instead of sharing one thread, so concurrent operations exercise real per-user memory-document CAS contention rather than per-thread turn serialization. setup_users creates and records one extra thread per hot writer for user 0; run_hot_writer picks its own thread. - Scripted document sizes are floored at 4 KiB (the token-dominated region below is meaningless and the issue's workloads start there); enforced in marker parsing and --api-scripted-doc-sizes validation. - The CI scripted leg runs inside the server's run block so the trap does not kill the server before it starts; artifacts upload together. - Wire-shape tests: mock_tool_call_response deserializes as the rig OpenAI CompletionResponse (stringified arguments, finish_reason tool_calls) and streaming tool-call chunks carry indexed delta tool_calls. * fix(stress): hot-writer client action ids collide with the primary writer A hot writer and the first user's regular writer shared the same user label and operation index, so their client_action_id values were identical and the server rejected the second submit with a 409 duplicate conflict. Include the scripted op prefix (h{k}-) in the operation ref so concurrent writers always submit distinct action ids. Found by a full local E2E run of the scripted leg against a real hosted-single-tenant server: after the fix, memory_roundtrip with one hot writer runs 9/9 clean (6 confirmed + 3 contended, 0 leaks) and memory_grow runs 8/8 confirmed. * fix(stress): address coderabbit review — verdict integrity, op-scoped tool counts, typed script key (nearai#7382) - compute_verdict: verdict comes from read steps only (write echoes can no longer mask missing/contended) - timeline tool evidence: count tool results by sequence above the op's baseline instead of subtracting page-limited absolute counts - timeline verdict match: delimit prefix by trailing space so op 1 cannot terminate on op 10's message; parse_result_verdict aligns on substring - background users namespace markers as b{index} so cross-cohort leaks cannot read back as their own token - hot writers drain in a dedicated JoinSet (no run_virtual_user refills) - fraction chunks derive from cumulative boundaries so split writes persist exactly the configured size (regression test at 4097) - ScriptKey derives clap::ValueEnum: CLI, marker wire format, and parsing share one string mapping; --api-hot-writers rejects write_file_roundtrip - parse_marker bounds identity grammar; poisoned mutex recovery; sorted stage latencies; single conversation parse per completion request; CLI-level scripted validation test and doc-size bound coverage * test(stress): cover --api-wait-for-assistant gate in CLI-level scripted test (nearai#7382)
nearai#7360) (nearai#7382) * feat(stress): scripted tool-call workload with durable write read-back (nearai#7360) Phase 1 of issue nearai#7360: teach the stress harness to drive real builtin and memory tool calls through the production capability path and verify their durable side effects. The api-user-capacity mock LLM sidecar learns a deterministic scripted state machine: the driver embeds an `ironclaw-stress-tool` marker in the user message, the sidecar emits the scripted tool call for a tool advertised in the request, the server executes it through the real capability host, and the driver verifies the read-back verdict in the final assistant message. Verdicts: confirmed / contended (same-user CAS race, counted) / leak (cross-user isolation, hard failure) / missing (write lost, hard failure) / undisclosed (tool never advertised). Scripts: write_file_roundtrip (write_file + read_file of a unique workspace path), memory_roundtrip / memory_grow / memory_mixed (ironclaw.memory.write/replace-append + read of the shared stress/shared.md target — every run doubles as a same-relative-path isolation check). --api-scripted-doc-sizes cycles 4 KiB..1 MiB documents with per-size buckets and submit-to-tool-visible / submit-to-finalize stage latencies; --api-hot-writers spawns concurrent same-user writers for hot-document CAS contention. Gated tools are exercised through the per-user Tools auto-approve setting enabled during setup via the production settings API. Wired as a nightly leg in the hosted-single-tenant Postgres job (the existing server stays up; the leg rebinds the mock sidecar on the same port). Unit coverage: marker parsing, per-op step sequencing, tool-name resolution (encoded/dotted/bare), verdict computation incl. leak precedence, disclosure fallback, timeline helpers, per-size summary buckets, and flag validation. * fix(stress): hot writers on distinct user threads, size floor, CI server lifecycle Review fixes for the nearai#7360 Phase 1 scripted workload: - Hot writers now run on distinct threads of the first user instead of sharing one thread, so concurrent operations exercise real per-user memory-document CAS contention rather than per-thread turn serialization. setup_users creates and records one extra thread per hot writer for user 0; run_hot_writer picks its own thread. - Scripted document sizes are floored at 4 KiB (the token-dominated region below is meaningless and the issue's workloads start there); enforced in marker parsing and --api-scripted-doc-sizes validation. - The CI scripted leg runs inside the server's run block so the trap does not kill the server before it starts; artifacts upload together. - Wire-shape tests: mock_tool_call_response deserializes as the rig OpenAI CompletionResponse (stringified arguments, finish_reason tool_calls) and streaming tool-call chunks carry indexed delta tool_calls. * fix(stress): hot-writer client action ids collide with the primary writer A hot writer and the first user's regular writer shared the same user label and operation index, so their client_action_id values were identical and the server rejected the second submit with a 409 duplicate conflict. Include the scripted op prefix (h{k}-) in the operation ref so concurrent writers always submit distinct action ids. Found by a full local E2E run of the scripted leg against a real hosted-single-tenant server: after the fix, memory_roundtrip with one hot writer runs 9/9 clean (6 confirmed + 3 contended, 0 leaks) and memory_grow runs 8/8 confirmed. * fix(stress): address coderabbit review — verdict integrity, op-scoped tool counts, typed script key (nearai#7382) - compute_verdict: verdict comes from read steps only (write echoes can no longer mask missing/contended) - timeline tool evidence: count tool results by sequence above the op's baseline instead of subtracting page-limited absolute counts - timeline verdict match: delimit prefix by trailing space so op 1 cannot terminate on op 10's message; parse_result_verdict aligns on substring - background users namespace markers as b{index} so cross-cohort leaks cannot read back as their own token - hot writers drain in a dedicated JoinSet (no run_virtual_user refills) - fraction chunks derive from cumulative boundaries so split writes persist exactly the configured size (regression test at 4097) - ScriptKey derives clap::ValueEnum: CLI, marker wire format, and parsing share one string mapping; --api-hot-writers rejects write_file_roundtrip - parse_marker bounds identity grammar; poisoned mutex recovery; sorted stage latencies; single conversation parse per completion request; CLI-level scripted validation test and doc-size bound coverage * test(stress): cover --api-wait-for-assistant gate in CLI-level scripted test (nearai#7382)
nearai#7360) (nearai#7382) * feat(stress): scripted tool-call workload with durable write read-back (nearai#7360) Phase 1 of issue nearai#7360: teach the stress harness to drive real builtin and memory tool calls through the production capability path and verify their durable side effects. The api-user-capacity mock LLM sidecar learns a deterministic scripted state machine: the driver embeds an `ironclaw-stress-tool` marker in the user message, the sidecar emits the scripted tool call for a tool advertised in the request, the server executes it through the real capability host, and the driver verifies the read-back verdict in the final assistant message. Verdicts: confirmed / contended (same-user CAS race, counted) / leak (cross-user isolation, hard failure) / missing (write lost, hard failure) / undisclosed (tool never advertised). Scripts: write_file_roundtrip (write_file + read_file of a unique workspace path), memory_roundtrip / memory_grow / memory_mixed (ironclaw.memory.write/replace-append + read of the shared stress/shared.md target — every run doubles as a same-relative-path isolation check). --api-scripted-doc-sizes cycles 4 KiB..1 MiB documents with per-size buckets and submit-to-tool-visible / submit-to-finalize stage latencies; --api-hot-writers spawns concurrent same-user writers for hot-document CAS contention. Gated tools are exercised through the per-user Tools auto-approve setting enabled during setup via the production settings API. Wired as a nightly leg in the hosted-single-tenant Postgres job (the existing server stays up; the leg rebinds the mock sidecar on the same port). Unit coverage: marker parsing, per-op step sequencing, tool-name resolution (encoded/dotted/bare), verdict computation incl. leak precedence, disclosure fallback, timeline helpers, per-size summary buckets, and flag validation. * fix(stress): hot writers on distinct user threads, size floor, CI server lifecycle Review fixes for the nearai#7360 Phase 1 scripted workload: - Hot writers now run on distinct threads of the first user instead of sharing one thread, so concurrent operations exercise real per-user memory-document CAS contention rather than per-thread turn serialization. setup_users creates and records one extra thread per hot writer for user 0; run_hot_writer picks its own thread. - Scripted document sizes are floored at 4 KiB (the token-dominated region below is meaningless and the issue's workloads start there); enforced in marker parsing and --api-scripted-doc-sizes validation. - The CI scripted leg runs inside the server's run block so the trap does not kill the server before it starts; artifacts upload together. - Wire-shape tests: mock_tool_call_response deserializes as the rig OpenAI CompletionResponse (stringified arguments, finish_reason tool_calls) and streaming tool-call chunks carry indexed delta tool_calls. * fix(stress): hot-writer client action ids collide with the primary writer A hot writer and the first user's regular writer shared the same user label and operation index, so their client_action_id values were identical and the server rejected the second submit with a 409 duplicate conflict. Include the scripted op prefix (h{k}-) in the operation ref so concurrent writers always submit distinct action ids. Found by a full local E2E run of the scripted leg against a real hosted-single-tenant server: after the fix, memory_roundtrip with one hot writer runs 9/9 clean (6 confirmed + 3 contended, 0 leaks) and memory_grow runs 8/8 confirmed. * fix(stress): address coderabbit review — verdict integrity, op-scoped tool counts, typed script key (nearai#7382) - compute_verdict: verdict comes from read steps only (write echoes can no longer mask missing/contended) - timeline tool evidence: count tool results by sequence above the op's baseline instead of subtracting page-limited absolute counts - timeline verdict match: delimit prefix by trailing space so op 1 cannot terminate on op 10's message; parse_result_verdict aligns on substring - background users namespace markers as b{index} so cross-cohort leaks cannot read back as their own token - hot writers drain in a dedicated JoinSet (no run_virtual_user refills) - fraction chunks derive from cumulative boundaries so split writes persist exactly the configured size (regression test at 4097) - ScriptKey derives clap::ValueEnum: CLI, marker wire format, and parsing share one string mapping; --api-hot-writers rejects write_file_roundtrip - parse_marker bounds identity grammar; poisoned mutex recovery; sorted stage latencies; single conversation parse per completion request; CLI-level scripted validation test and doc-size bound coverage * test(stress): cover --api-wait-for-assistant gate in CLI-level scripted test (nearai#7382)
…rai#6896) (nearai#7131) * fix(run_delivery): deliver triggered run failures to the creator (nearai#6896) Scheduled/triggered runs that ended in Failed, Cancelled, or RecoveryRequired produced no user-visible notification: the triggered delivery driver minted notifications only for Completed / BlockedApproval / BlockedAuth and recorded every other terminal status as Skipped. A run that timed out before reaching an actionable state only logged a warn and recorded Failed, leaving the creator in silence. Delivery: - triggered_notification_for_state now mints a FinalReplyReady notification for Failed and RecoveryRequired using the existing per-category failure summaries (reborn_failure_summary_for_category) over state.failure.category(), with a generic fallback when no category is present. - Cancelled mints the same notification, preferring a failure-category summary when one is present and falling back to a fixed cancellation notice otherwise. - The RunWaitTimedOut branch with no prior blocked marker now delivers the timeout notice as a terminal reply instead of recording Failed. - The wildcard arm is replaced with explicit non-actionable statuses (Queued, Running, CancelRequested, BlockedResource, BlockedDependentRun, BlockedExternalTool) so a future status fails to compile rather than silently skipping. Observer: - TriggerFireSettlementObserver gains on_failed_fire_settled as a default no-op method, plus a TriggerFailedFireSettlement event carrying tenant/trigger/fire-slot/run-id/history-status. Noop and existing implementors keep compiling. - The active-cleanup sweep fires on_failed_fire_settled when clear_active_fire succeeds with TriggerRunHistoryStatus::Error, so post-accept failures are observable for automation health. Ok, Running, and already-cleared fires do not fire the hook. Tests: - run_delivery_contract: Failed+model_error, Failed without category, Cancelled, and timeout-before-actionable all assert a Delivered outcome with the expected notice text and footer. - worker tests: a terminal-Error active fire fires exactly one on_failed_fire_settled; a terminal-Ok active fire fires none. The larger retry/redrive budget for failed post-accept fires (retry_disposition has zero production callers) is intentionally left for a follow-up; it is out of scope for this surgical delivery fix. * style: cargo fmt the nearai#6896 delivery fix * fix(triggers): address terminal delivery review feedback * fix(assistant): drop unused UserId import after merge * fix(run_delivery): address multi-agent review findings - Extract shared terminal-notice helpers (final_reply_notice, outcome_for_delivery_failure, deliver_terminal_notice) so the timeout, OAuth-backstop, and generic failure arms share one notice shape and outcome taxonomy instead of a third hand-rolled copy. - Add a bounded race-grace window after the wait backstop: a run that crosses into a terminal state during the final wait (cancellation in flight, failure landing after the last poll) now delivers the correct terminal notice instead of the timeout copy. - Cancelled runs always deliver the fixed cancellation notice; the failure-category branch was unreachable in production and would have mislabeled a host/operator cancel as a failure. - Update the stale invariant doc, the five-output surface contract count, and the exhaustiveness-only comment on the non-actionable arm. - Document the cheap/non-blocking contract on TriggerFireSettlementObserver (the worker awaits it inline in the poller sweep) and note it at the active-cleanup call site. - Add contract coverage for the timeout arm's delivery-failure outcome (Failed) and a regression test proving the race-grace path delivers the cancellation notice; the cancelled-with-category test now asserts the cancellation notice wins. * fix(run_delivery): address review comments and restore CI gates Review fixes (CodeRabbit on 01e887f/f8af109): - Grace loop fails loud: log the bound TurnError on state-poll failure and the RunDeliveryError on terminal-notice build failure before falling back to the timeout copy, with silent-ok markers on both intentional fallbacks. - Hoist TriggeredReplyTargetAuthority, CodecChannelTargetResolver, and TriggeredNotificationContext to one construction before the watcher loop; the race-grace arm, timeout arm, and loop body now share it. - Collapse the duplicated failure-summary expression into one closure and name TurnStatus::Failed explicitly so future statuses are compiler-visible. - Drop the stale "Only three states" count from the surface-contract doc. - Test fixture: encode the late-terminal flip as one Option<(usize, ScriptedRunState)> field instead of two correlated Options with an expect. - Terminal-crossing test: document why flip_after=30 deterministically outruns the wait poll budget and assert the grace loop issues no cancellation (cancel_calls == 0). CI: - composition-budget: re-seed loc_ceiling 40432 -> 40593 (measured on the merged tree; the nearai#7131 settlement observer adds +161 governed LOC of wiring) and move the arch-test record with it. - trigger_poller: use the colon-form tracing target required by nearai#7146. * ci: re-trigger pull_request workflows for c2460ed * fix(composition): capture the settlement health warn in the observer test The traced_test default filter is {crate}=trace, which drops events whose metadata target is `ironclaw::reborn::…`. The observer warning is emitted with the colon-form target (required by nearai#7146 — the equals form recorded a field and never matched RUST_LOG target filters), so the test saw an empty buffer. Enable tracing-test's no-env-filter feature, the same pattern the capabilities/host-runtime/mcp/loop crates use for cross-target assertions. Re-seed the composition budget to the merged-tree measurement (40747 -> 40867): nearai#7131's observer wiring lands on top of post-measurement mainline inflow; measured with the gate, set to current. The arch-test record moves with the manifest. * fix(run_delivery): merge main and adapt to notice_discriminator String - Merge origin/main (nearai#7377 run-acts-as-invoker, nearai#7323, nearai#7382, nearai#6938, nearai#7280, nearai#7393, nearai#7389, nearai#7364, nearai#7228, nearai#7371, nearai#7399). - main's nearai#7377 landed a narrower terminal arm (generic failure notice for TurnStatus::Failed only); keep the nearai#6896 arm, which covers Failed and RecoveryRequired with sanitized per-category summaries plus Cancelled and the timeout grace path, and adapt to the Option<String> notice_discriminator main introduced. - Re-seed the composition budget to the merged-tree measurement (40811 -> 40861, the run-failure settlement observer lands +50 governed LOC); the arch-test record moves with the manifest.
Summary
--api-scripted-toolmode onapi-user-capacitywith read-back verdicts (confirmed/contended/leak/missing/undisclosed), per-document-size buckets (4 KiB–1 MiB), submit-to-tool-visible / submit-to-finalize stage timings, and--api-hot-writerssame-user contention.Change Type
Linked Issue
Part of #7360 — Phase 1 foundation slice: scripted-tool harness + persistent-memory roundtrip coverage. Reindex/churn workloads, the libSQL leg, and Phases 2–3 remain tracked in the issue.
Validation
cargo fmt -p ironclaw_stress -- --checkcargo clippy -p ironclaw_stress --all-targets --all-features -- -D warningscargo build -p ironclaw_stress --releasecargo test -p ironclaw_stress(119 passed)cargo test -p <owning-crate> --features integration— Not applicable: no crate-level integration behavior changed; E2E path covered by the new nightly stress job (postgres-api-capacity scripted leg), which requires a live server + Postgres and runs on schedulememory_roundtrip9/9 clean (6 confirmed + 3 contended under one hot writer, 0 leaks),memory_grow8/8 confirmed, per-size buckets and stage latencies present. E2E caught a hot-writer client_action_id collision (409 duplicate) that is fixed and covered by the rerun.review-pr/pr-shepherd— not run (not available in this session)Test Strategy
User behavior: Not applicable — developer tooling (ironclaw_stress) and CI; no product UX change.
Risk areas:
Tests added or updated:
--api-*flag validationpostgres-api-capacityscripted leg (memory_roundtrip, 4 sizes, 2 hot writers, failure-rate + p95 thresholds; loose first-run ceilings pending baseline artifact)What the tests prove: the scripted state machine emits the right tool call per step from the conversation alone; read-back verdicts correctly distinguish confirmed/contended/leak/missing/undisclosed; operations are not redriven after completion; the driver attributes per-op outcomes to document size and stage; invalid flag combinations fail fast.
Commands run:
cargo fmt -p ironclaw_stress -- --check,cargo clippy -p ironclaw_stress --all-targets --all-features -- -D warnings,cargo test -p ironclaw_stress,cargo build -p ironclaw_stress --releaseCompatibility / rollback notes
api-user-capacityruns unchanged (sidecar falls through to text when no marker is present).scriptedsections (skipped when absent) —--compare-jsonconsumers unaffected.memory_grow/memory_mixed/write_file_roundtripnightly legs, Phase 2 (workspace/spill/trigger/config/extension families), Phase 3 inventory ratchet, per-case threshold tightening after the first baseline.