feat(engine): add mission_get action for retrieving mission results - #2549
Conversation
The LLM had no tool to retrieve mission outcomes — when users asked "what is the result of the research", the agent fell back to calling list_jobs because no mission results tool existed. Missions spawn threads (not jobs), so list_jobs always returned empty results. Add mission_get action that loads mission details + recent thread outputs so the LLM can answer mission result queries directly. Also map routine_history (v1 alias) to mission_get for compatibility. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request introduces the mission_get action, enabling the retrieval of detailed status, approach history, and recent thread results for missions and routines. The implementation includes a new store() accessor in MissionManager, the mission_get handler in EffectBridgeAdapter, and a mapping for the legacy routine_history action. Review feedback highlights a security vulnerability where the mission_get handler lacks an ownership check, potentially allowing unauthorized access to mission data. Furthermore, it is suggested to truncate the approach_history in the response to prevent exceeding LLM context window limits.
| }); | ||
| match id { | ||
| Ok(id) => match mgr.get_mission(id).await { | ||
| Ok(Some(mission)) => { |
There was a problem hiding this comment.
The mission_get handler lacks an ownership check, allowing any user to retrieve details of any mission by its UUID. This is a security regression compared to other mission actions like mission_fire or mission_pause. An explicit check against context.user_id should be added.
References
- Tools that interact with user-owned resources must verify that the authenticated user ID in the tool context matches the resource owner's ID before performing any read or write operations to prevent unauthorized cross-user access.
| "goal": mission.goal, | ||
| "status": format!("{:?}", mission.status), | ||
| "current_focus": mission.current_focus, | ||
| "approach_history": mission.approach_history, |
There was a problem hiding this comment.
approach_history can grow significantly over time as it stores the full response of every mission run. Returning the entire history in mission_get could eventually exceed the LLM's context window or the safety layer's output limits. Consider returning only the most recent entries (e.g., the last 10).
| "approach_history": mission.approach_history, | |
| "approach_history": mission.approach_history.iter().rev().take(10).rev().cloned().collect::<Vec<_>>(), |
References
- Always truncate tool output for previews or status updates to a reasonable maximum length to prevent excessive memory/bandwidth usage and reduce the risk of leaking sensitive information.
…2549) Add user_id ownership check to mission_get handler to prevent cross-user IDOR (mirrors fire/pause/resume guards). Cap approach_history to last 10 entries to prevent context window overflow. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
Addressed review feedback:
Both fixes in commit 021d5c6. |
henrypark133
left a comment
There was a problem hiding this comment.
Review: ownership fix and response shaping look sound
I re-checked the current head in src/bridge/effect_adapter.rs:344-382. The earlier ownership concern is addressed: mission_get now rejects missions whose mission.user_id does not match the caller, and the response also caps approach_history to the most recent 10 entries instead of returning the full unbounded list.
Prior feedback status: the previously raised ownership issue appears addressed in the current head.
Residual risk: I did not wait for the full local compile to finish for a targeted alias test in this pass, so approval here is based on code inspection rather than completed local test execution.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…ceptance Closes two gaps blocking v2 engine becoming the default: **Phase 4 — token + cost accounting** - Delete orphaned `crates/ironclaw_engine/src/executor/compaction.rs` (176 lines). The Python orchestrator (`default.py::compact_if_needed`) has owned compaction policy since #1557; the Rust module had no callers anywhere in the workspace. - Wire `cost_usd` in `LlmBridgeAdapter` by calling `LlmProvider::calculate_cost()` at both the no-tools and with-tools response paths. The engine's `Thread::total_cost_usd` accumulator and `max_budget_usd` gates were already plumbed — only the adapter was hardcoding 0.0. - Persist `total_cost_usd` through `ThreadArchiveSummary` round-trip in `store_adapter.rs`. Previously, rehydrating an archived thread silently dropped the cost to 0.0. `#[serde(default)]` keeps existing archive files deserializing cleanly. **Phase 6 — mission lifecycle acceptance** Three new integration tests in `bridge/effect_adapter.rs` driving `execute_action()` end-to-end (per `.claude/rules/testing.md` "Test Through the Caller"): - `mission_full_lifecycle_via_execute_action` — create → list → complete → list, asserting the `Completed` status surfaces through `mission_list` after `mission_complete`. - `mission_fire_returns_thread_id_for_manual_cadence_via_execute_action` — fresh manual mission fires successfully and returns a UUID thread_id rather than `not_fired`. - `mission_list_returns_all_user_missions_via_execute_action` — all three created missions appear in `mission_list` output. **Regression tests for cost wiring** Three new tests in `bridge/llm_adapter.rs`: - `complete_no_tools_populates_cost_usd_through_adapter` - `complete_with_tools_populates_cost_usd_through_adapter` - `complete_routes_subcalls_through_cheap_provider_for_cost` — pins that `depth > 0` is priced with the cheap provider, not the primary. Coordinated with in-flight work: skipped paths owned by #2504 (auth E2E), #2631 (paused-lease resume), #2570 (mission re-fire), #2549 (mission_get), #2452 (tool_calls persistence), #2621 (replay snapshot). Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (409 passed), `cargo test -p ironclaw --lib` (5079 passed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…ceptance Closes two gaps blocking v2 engine becoming the default: **Phase 4 — token + cost accounting** - Delete orphaned `crates/ironclaw_engine/src/executor/compaction.rs` (176 lines). The Python orchestrator (`default.py::compact_if_needed`) has owned compaction policy since #1557; the Rust module had no callers anywhere in the workspace. - Wire `cost_usd` in `LlmBridgeAdapter` by calling `LlmProvider::calculate_cost()` at both the no-tools and with-tools response paths. The engine's `Thread::total_cost_usd` accumulator and `max_budget_usd` gates were already plumbed — only the adapter was hardcoding 0.0. - Persist `total_cost_usd` through `ThreadArchiveSummary` round-trip in `store_adapter.rs`. Previously, rehydrating an archived thread silently dropped the cost to 0.0. `#[serde(default)]` keeps existing archive files deserializing cleanly. **Phase 6 — mission lifecycle acceptance** Three new integration tests in `bridge/effect_adapter.rs` driving `execute_action()` end-to-end (per `.claude/rules/testing.md` "Test Through the Caller"): - `mission_full_lifecycle_via_execute_action` — create → list → complete → list, asserting the `Completed` status surfaces through `mission_list` after `mission_complete`. - `mission_fire_returns_thread_id_for_manual_cadence_via_execute_action` — fresh manual mission fires successfully and returns a UUID thread_id rather than `not_fired`. - `mission_list_returns_all_user_missions_via_execute_action` — all three created missions appear in `mission_list` output. **Regression tests for cost wiring** Three new tests in `bridge/llm_adapter.rs`: - `complete_no_tools_populates_cost_usd_through_adapter` - `complete_with_tools_populates_cost_usd_through_adapter` - `complete_routes_subcalls_through_cheap_provider_for_cost` — pins that `depth > 0` is priced with the cheap provider, not the primary. Coordinated with in-flight work: skipped paths owned by #2504 (auth E2E), #2631 (paused-lease resume), #2570 (mission re-fire), #2549 (mission_get), #2452 (tool_calls persistence), #2621 (replay snapshot). Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (409 passed), `cargo test -p ironclaw --lib` (5079 passed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…ceptance (#2660) * feat(engine-v2): Phase 4 cost tracking + Phase 6 mission lifecycle acceptance Closes two gaps blocking v2 engine becoming the default: **Phase 4 — token + cost accounting** - Delete orphaned `crates/ironclaw_engine/src/executor/compaction.rs` (176 lines). The Python orchestrator (`default.py::compact_if_needed`) has owned compaction policy since #1557; the Rust module had no callers anywhere in the workspace. - Wire `cost_usd` in `LlmBridgeAdapter` by calling `LlmProvider::calculate_cost()` at both the no-tools and with-tools response paths. The engine's `Thread::total_cost_usd` accumulator and `max_budget_usd` gates were already plumbed — only the adapter was hardcoding 0.0. - Persist `total_cost_usd` through `ThreadArchiveSummary` round-trip in `store_adapter.rs`. Previously, rehydrating an archived thread silently dropped the cost to 0.0. `#[serde(default)]` keeps existing archive files deserializing cleanly. **Phase 6 — mission lifecycle acceptance** Three new integration tests in `bridge/effect_adapter.rs` driving `execute_action()` end-to-end (per `.claude/rules/testing.md` "Test Through the Caller"): - `mission_full_lifecycle_via_execute_action` — create → list → complete → list, asserting the `Completed` status surfaces through `mission_list` after `mission_complete`. - `mission_fire_returns_thread_id_for_manual_cadence_via_execute_action` — fresh manual mission fires successfully and returns a UUID thread_id rather than `not_fired`. - `mission_list_returns_all_user_missions_via_execute_action` — all three created missions appear in `mission_list` output. **Regression tests for cost wiring** Three new tests in `bridge/llm_adapter.rs`: - `complete_no_tools_populates_cost_usd_through_adapter` - `complete_with_tools_populates_cost_usd_through_adapter` - `complete_routes_subcalls_through_cheap_provider_for_cost` — pins that `depth > 0` is priced with the cheap provider, not the primary. Coordinated with in-flight work: skipped paths owned by #2504 (auth E2E), #2631 (paused-lease resume), #2570 (mission re-fire), #2549 (mission_get), #2452 (tool_calls persistence), #2621 (replay snapshot). Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (409 passed), `cargo test -p ironclaw --lib` (5079 passed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(engine-v2): surface engine capability actions to LLM via available_actions Fixes the gap called out in the PR body: `EffectBridgeAdapter::available_actions` was only enumerating v1 `ToolRegistry` tools + latent OAuth actions, so engine-native capabilities like `missions` never appeared in the LLM's tools list even when a thread held an active lease for them. The LLM was therefore unable to call `mission_create` / `mission_list` / etc. via structured tool calls; the only ways to drive missions were CodeAct Python calls (which relied on the same `known_actions` set and hit the same gap) or `/routine` slash commands falling through to v1. Wire `CapabilityRegistry` into the adapter and iterate active leases to surface every leased, engine-registered capability action. Respects lease grant scope — a lease granting only `mission_list` does not leak `mission_create`. Skips the `"tools"` capability since that lease is already reconciled from the v1 path. Router wires the shared `Arc<CapabilityRegistry>` to both the adapter and `ThreadManager` at setup. Three new regression tests: - `available_actions_surfaces_leased_mission_capability` - `available_actions_respects_partial_lease_grant` - `available_actions_omits_capability_without_lease` Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (409 passed), `cargo test -p ironclaw --lib` (5082 passed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(engine-v2): close review gaps — archive round-trip, v1/engine merge, defensive filters Addresses gaps raised in PR #2660 review: - **Consolidate `use ironclaw_engine::{...}`** into a single grouped import in `effect_adapter.rs` (was split across two statements). - **Apply `is_v1_only_tool` / `is_v1_auth_tool` filters** to the engine capability path in `available_actions`. Defensive guardrail: a future engine capability that registers an action under a v1-denylisted name (`create_job`, `tool_auth`, ...) must not bypass the v2-isolation filters by virtue of coming through a different capability registry. - **`ThreadArchiveSummary` serialization round-trip tests** in `store_adapter.rs`: - `archive_summary_preserves_total_cost_usd_through_round_trip` — pins the regression the PR fixed (cost silently zeroed on rehydration). - `archive_summary_handles_legacy_json_without_total_cost_usd_field` — pins `#[serde(default)]` back-compat for archive files written before this PR. - **`available_actions` combined advertising tests** in `effect_adapter.rs`: - `available_actions_merges_v1_tools_with_engine_capabilities` — v1 tool + mission capability both surface on one call. - `available_actions_filters_v1_denylisted_names_from_engine_capabilities` — pins the new defensive filter. - **`cost_usd_from` subscription-billed-provider test** in `llm_adapter.rs`: - `complete_with_subscription_billed_provider_yields_zero_cost` — zero `cost_per_token` round-trips to exactly `0.0`, no NaN/Inf. Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (409 passed), `cargo test -p ironclaw --lib` (5087 passed, +5 new). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(engine-v2): price cache tokens correctly in LlmBridgeAdapter Addresses PR #2660 review (gemini-code-assist + Copilot, L23/L115/L189): `cost_usd_from` only priced `input_tokens + output_tokens`, ignoring `cache_read_input_tokens` and `cache_creation_input_tokens`. For providers with prompt caching (Anthropic, OpenAI), this undercounted input cost and silently neutered the `max_budget_usd` gate. Extend the helper to mirror the canonical formula in `src/agent/cost_guard.rs::CostGuard::record_llm_call`: uncached_input = input_tokens - (cache_read + cache_write) cache_read_cost = input_rate * cache_read / cache_read_discount() cache_write_cost = input_rate * cache_write * cache_write_multiplier() cost = input_rate * uncached_input + cache_read_cost + cache_write_cost + output_rate * output_tokens All `LlmProvider` implementations already supply `cache_read_discount()` (default 1, Anthropic 10, OpenAI 2) and `cache_write_multiplier()` (default 1, Anthropic 1.25 for 5m / 2.0 for 1h) through the decorator chain, so no trait surgery is required. Regression test: `complete_prices_cache_tokens_with_discount_and_multiplier` uses Anthropic Sonnet 5m-TTL rates, exercises a 10k-input / 2k-read / 1k-write / 500-output response, and pins the correct total ($0.03285) against the old naive $0.0375 that would have undercounted ~14%. Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (435 passed), `cargo test -p ironclaw --lib` (5136 passed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- `EngineState` on staging now has `extension_manager` and `project_root` fields (added in #2549 and the attachments flow). Update the test-only `install_engine_state_with_threads` helper to populate them. - Rewrite the `let Some(...) else { return None }` in `engine_history_entry_to_message` as a `?` — clippy::question_mark is denied on staging's all-features CI. - Add missing `AuthenticatedUser` + `Query` imports to the server.rs test module for the `history_request` helper. - `cargo fmt` on the 500-error test. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…2532) * feat(gateway): expose engine v2 threads in chat history and sidebar Engine v2 threads weren't appearing in the gateway sidebar and deep-linking to one by id (`#/chat/<engine-thread-id>`) returned an empty history because the v1 `assistant` flow dual-writes into the single assistant conversation id, not the engine thread id. Three coordinated fixes: - `chat_history_handler`: extend the ownership check with an engine v2 lookup so an engine thread id is recognized, then fall back to loading messages via `bridge::get_engine_thread` when the v1 conversation table has nothing. - `chat_threads_handler`: merge engine threads from `bridge::list_engine_threads` into the sidebar, label them with their goal, and re-sort by `updated_at`. Bump the v1 conversation cap from 50 to 500 so older threads stop silently aging off the sidebar. - Gateway frontend (`app.js`): when restoring from `#/chat/<id>` on load, switch even if the id is not in the loaded sidebar list — the history endpoint resolves it via the DB. Log a warning instead of silently dropping the URL. Cherry-picked from 8df22ab (feat/skills-engine-fixes), adapted to staging where chat_history_handler and chat_threads_handler still live in both server.rs and handlers/chat.rs. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): address v2 history review comments * test(gateway): caller-level coverage for v2 thread ownership + history Adds tests exercising chat_history_handler through the three ownership branches the PR introduces, and two HTTP-driven e2e scenarios against an ENGINE_V2=true server. Fills the caller-level coverage gap flagged in review. - Rust caller-level tests on chat_history_handler: - v2-owned: engine thread owned by user returns synthesized history - cross-user: alice can't read bob's engine thread (404) - session-only: in-memory session-owned thread returns 200 without DB - Playwright e2e under ENGINE_V2=true: - engine-only thread appears in sidebar with channel=engine - deep-link by engine thread id returns synthesized turns Exposes a minimal bridge::test_support module (ThreadTestStore, install_engine_state_with_threads, clear_engine_state, shared test lock) so cross-module tests can seed ENGINE_STATE without the weight of the bridge's own full-featured TestStore. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): align with staging EngineState fields and clippy rules - `EngineState` on staging now has `extension_manager` and `project_root` fields (added in #2549 and the attachments flow). Update the test-only `install_engine_state_with_threads` helper to populate them. - Rewrite the `let Some(...) else { return None }` in `engine_history_entry_to_message` as a `?` — clippy::question_mark is denied on staging's all-features CI. - Add missing `AuthenticatedUser` + `Query` imports to the server.rs test module for the `history_request` helper. - `cargo fmt` on the 500-error test. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
…ceptance (nearai#2660) * feat(engine-v2): Phase 4 cost tracking + Phase 6 mission lifecycle acceptance Closes two gaps blocking v2 engine becoming the default: **Phase 4 — token + cost accounting** - Delete orphaned `crates/ironclaw_engine/src/executor/compaction.rs` (176 lines). The Python orchestrator (`default.py::compact_if_needed`) has owned compaction policy since nearai#1557; the Rust module had no callers anywhere in the workspace. - Wire `cost_usd` in `LlmBridgeAdapter` by calling `LlmProvider::calculate_cost()` at both the no-tools and with-tools response paths. The engine's `Thread::total_cost_usd` accumulator and `max_budget_usd` gates were already plumbed — only the adapter was hardcoding 0.0. - Persist `total_cost_usd` through `ThreadArchiveSummary` round-trip in `store_adapter.rs`. Previously, rehydrating an archived thread silently dropped the cost to 0.0. `#[serde(default)]` keeps existing archive files deserializing cleanly. **Phase 6 — mission lifecycle acceptance** Three new integration tests in `bridge/effect_adapter.rs` driving `execute_action()` end-to-end (per `.claude/rules/testing.md` "Test Through the Caller"): - `mission_full_lifecycle_via_execute_action` — create → list → complete → list, asserting the `Completed` status surfaces through `mission_list` after `mission_complete`. - `mission_fire_returns_thread_id_for_manual_cadence_via_execute_action` — fresh manual mission fires successfully and returns a UUID thread_id rather than `not_fired`. - `mission_list_returns_all_user_missions_via_execute_action` — all three created missions appear in `mission_list` output. **Regression tests for cost wiring** Three new tests in `bridge/llm_adapter.rs`: - `complete_no_tools_populates_cost_usd_through_adapter` - `complete_with_tools_populates_cost_usd_through_adapter` - `complete_routes_subcalls_through_cheap_provider_for_cost` — pins that `depth > 0` is priced with the cheap provider, not the primary. Coordinated with in-flight work: skipped paths owned by nearai#2504 (auth E2E), nearai#2631 (paused-lease resume), nearai#2570 (mission re-fire), nearai#2549 (mission_get), nearai#2452 (tool_calls persistence), nearai#2621 (replay snapshot). Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (409 passed), `cargo test -p ironclaw --lib` (5079 passed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(engine-v2): surface engine capability actions to LLM via available_actions Fixes the gap called out in the PR body: `EffectBridgeAdapter::available_actions` was only enumerating v1 `ToolRegistry` tools + latent OAuth actions, so engine-native capabilities like `missions` never appeared in the LLM's tools list even when a thread held an active lease for them. The LLM was therefore unable to call `mission_create` / `mission_list` / etc. via structured tool calls; the only ways to drive missions were CodeAct Python calls (which relied on the same `known_actions` set and hit the same gap) or `/routine` slash commands falling through to v1. Wire `CapabilityRegistry` into the adapter and iterate active leases to surface every leased, engine-registered capability action. Respects lease grant scope — a lease granting only `mission_list` does not leak `mission_create`. Skips the `"tools"` capability since that lease is already reconciled from the v1 path. Router wires the shared `Arc<CapabilityRegistry>` to both the adapter and `ThreadManager` at setup. Three new regression tests: - `available_actions_surfaces_leased_mission_capability` - `available_actions_respects_partial_lease_grant` - `available_actions_omits_capability_without_lease` Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (409 passed), `cargo test -p ironclaw --lib` (5082 passed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(engine-v2): close review gaps — archive round-trip, v1/engine merge, defensive filters Addresses gaps raised in PR nearai#2660 review: - **Consolidate `use ironclaw_engine::{...}`** into a single grouped import in `effect_adapter.rs` (was split across two statements). - **Apply `is_v1_only_tool` / `is_v1_auth_tool` filters** to the engine capability path in `available_actions`. Defensive guardrail: a future engine capability that registers an action under a v1-denylisted name (`create_job`, `tool_auth`, ...) must not bypass the v2-isolation filters by virtue of coming through a different capability registry. - **`ThreadArchiveSummary` serialization round-trip tests** in `store_adapter.rs`: - `archive_summary_preserves_total_cost_usd_through_round_trip` — pins the regression the PR fixed (cost silently zeroed on rehydration). - `archive_summary_handles_legacy_json_without_total_cost_usd_field` — pins `#[serde(default)]` back-compat for archive files written before this PR. - **`available_actions` combined advertising tests** in `effect_adapter.rs`: - `available_actions_merges_v1_tools_with_engine_capabilities` — v1 tool + mission capability both surface on one call. - `available_actions_filters_v1_denylisted_names_from_engine_capabilities` — pins the new defensive filter. - **`cost_usd_from` subscription-billed-provider test** in `llm_adapter.rs`: - `complete_with_subscription_billed_provider_yields_zero_cost` — zero `cost_per_token` round-trips to exactly `0.0`, no NaN/Inf. Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (409 passed), `cargo test -p ironclaw --lib` (5087 passed, +5 new). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(engine-v2): price cache tokens correctly in LlmBridgeAdapter Addresses PR nearai#2660 review (gemini-code-assist + Copilot, L23/L115/L189): `cost_usd_from` only priced `input_tokens + output_tokens`, ignoring `cache_read_input_tokens` and `cache_creation_input_tokens`. For providers with prompt caching (Anthropic, OpenAI), this undercounted input cost and silently neutered the `max_budget_usd` gate. Extend the helper to mirror the canonical formula in `src/agent/cost_guard.rs::CostGuard::record_llm_call`: uncached_input = input_tokens - (cache_read + cache_write) cache_read_cost = input_rate * cache_read / cache_read_discount() cache_write_cost = input_rate * cache_write * cache_write_multiplier() cost = input_rate * uncached_input + cache_read_cost + cache_write_cost + output_rate * output_tokens All `LlmProvider` implementations already supply `cache_read_discount()` (default 1, Anthropic 10, OpenAI 2) and `cache_write_multiplier()` (default 1, Anthropic 1.25 for 5m / 2.0 for 1h) through the decorator chain, so no trait surgery is required. Regression test: `complete_prices_cache_tokens_with_discount_and_multiplier` uses Anthropic Sonnet 5m-TTL rates, exercises a 10k-input / 2k-read / 1k-write / 500-output response, and pins the correct total ($0.03285) against the old naive $0.0375 that would have undercounted ~14%. Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features` (0 warnings), `cargo test -p ironclaw_engine` (435 passed), `cargo test -p ironclaw --lib` (5136 passed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…earai#2549) * feat(engine): add mission_get action for retrieving mission results The LLM had no tool to retrieve mission outcomes — when users asked "what is the result of the research", the agent fell back to calling list_jobs because no mission results tool existed. Missions spawn threads (not jobs), so list_jobs always returned empty results. Add mission_get action that loads mission details + recent thread outputs so the LLM can answer mission result queries directly. Also map routine_history (v1 alias) to mission_get for compatibility. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): address review — ownership check + approach_history cap (nearai#2549) Add user_id ownership check to mission_get handler to prevent cross-user IDOR (mirrors fire/pause/resume guards). Cap approach_history to last 10 entries to prevent context window overflow. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: remove redundant .into_iter() to satisfy clippy useless_conversion Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…earai#2532) * feat(gateway): expose engine v2 threads in chat history and sidebar Engine v2 threads weren't appearing in the gateway sidebar and deep-linking to one by id (`#/chat/<engine-thread-id>`) returned an empty history because the v1 `assistant` flow dual-writes into the single assistant conversation id, not the engine thread id. Three coordinated fixes: - `chat_history_handler`: extend the ownership check with an engine v2 lookup so an engine thread id is recognized, then fall back to loading messages via `bridge::get_engine_thread` when the v1 conversation table has nothing. - `chat_threads_handler`: merge engine threads from `bridge::list_engine_threads` into the sidebar, label them with their goal, and re-sort by `updated_at`. Bump the v1 conversation cap from 50 to 500 so older threads stop silently aging off the sidebar. - Gateway frontend (`app.js`): when restoring from `#/chat/<id>` on load, switch even if the id is not in the loaded sidebar list — the history endpoint resolves it via the DB. Log a warning instead of silently dropping the URL. Cherry-picked from 8df22ab (feat/skills-engine-fixes), adapted to staging where chat_history_handler and chat_threads_handler still live in both server.rs and handlers/chat.rs. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): address v2 history review comments * test(gateway): caller-level coverage for v2 thread ownership + history Adds tests exercising chat_history_handler through the three ownership branches the PR introduces, and two HTTP-driven e2e scenarios against an ENGINE_V2=true server. Fills the caller-level coverage gap flagged in review. - Rust caller-level tests on chat_history_handler: - v2-owned: engine thread owned by user returns synthesized history - cross-user: alice can't read bob's engine thread (404) - session-only: in-memory session-owned thread returns 200 without DB - Playwright e2e under ENGINE_V2=true: - engine-only thread appears in sidebar with channel=engine - deep-link by engine thread id returns synthesized turns Exposes a minimal bridge::test_support module (ThreadTestStore, install_engine_state_with_threads, clear_engine_state, shared test lock) so cross-module tests can seed ENGINE_STATE without the weight of the bridge's own full-featured TestStore. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): align with staging EngineState fields and clippy rules - `EngineState` on staging now has `extension_manager` and `project_root` fields (added in nearai#2549 and the attachments flow). Update the test-only `install_engine_state_with_threads` helper to populate them. - Rewrite the `let Some(...) else { return None }` in `engine_history_entry_to_message` as a `?` — clippy::question_mark is denied on staging's all-features CI. - Add missing `AuthenticatedUser` + `Query` imports to the server.rs test module for the `history_request` helper. - `cargo fmt` on the 500-error test. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
Summary
mission_getaction so the LLM can retrieve mission results, status, approach history, and recent thread outputsroutine_history→mission_getso the v1 alias also works in engine v2list_jobsbecause no mission results tool existed — missions spawn threads (not jobs), solist_jobsalways returned emptyChanges
src/bridge/router.rsmission_getActionDef in capability registrysrc/bridge/effect_adapter.rsroutine_history→mission_get; regression testcrates/ironclaw_engine/src/runtime/mission.rspub fn store()accessor on MissionManagerTest plan
cargo fmt --check— cleancargo clippy --all --all-features— zero warningscargo test -- routine_history_maps— alias test passesmission_getinstead oflist_jobs🤖 Generated with Claude Code