engine-v2: add capability projection and two-surface prompt baseline - #2826
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces a capability background status system, allowing the engine to surface information about non-callable or indirectly-callable integrations in the system prompt. Key changes include the addition of CapabilitySummary types, a new available_capabilities method in the EffectExecutor trait, and the introduction of ActionProjector and CapabilityProjector to handle the surfacing logic. Feedback focuses on performance improvements, specifically recommending the use of write! for efficient string building, parallelizing asynchronous channel route lookups, and prefetching tool provider data to avoid sequential awaits and potential lock contention within loops.
There was a problem hiding this comment.
Pull request overview
Adds engine-v2 capability “background” projection (separate from callable tool schemas) and wires that capability background into the initial two-surface system prompt, while tightening available_actions() to only include currently callable actions.
Changes:
- Introduces
CapabilitySummary/CapabilitySummaryKindand addsavailable_capabilities(...)to theEffectExecutorboundary. - Adds bridge-side
CapabilityProjector+ActionProjectorto derive background capability summaries and filter callable actions using runtime truth. - Plumbs
ThreadExecutionContextthrough action/capability discovery and renders capability background into the CodeAct system prompt baseline.
Reviewed changes
Copilot reviewed 24 out of 24 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
crates/ironclaw_engine/src/types/capability.rs |
Adds capability background summary types + serde tests. |
crates/ironclaw_engine/src/traits/effect.rs |
Extends EffectExecutor with available_capabilities and adds ThreadExecutionContext to available_actions. |
src/bridge/effect_adapter.rs |
Reworks action listing via ActionProjector and adds capability listing via CapabilityProjector. |
src/bridge/action_projector.rs |
Filters callable actions (incl. extension-backed) using runtime extension inventory + surface policy. |
src/bridge/capability_projector.rs |
Projects runtime truth (extensions/latent/channel routing) into compact capability summaries. |
src/bridge/tool_surface.rs |
Updates surface assignment so Error entries land in capability background. |
crates/ironclaw_engine/src/executor/prompt.rs |
Renders capability background into the CodeAct system prompt. |
crates/ironclaw_engine/src/executor/thread_context.rs |
Centralizes thread→ThreadExecutionContext extraction. |
crates/ironclaw_engine/src/executor/* + crates/ironclaw_engine/src/runtime/lease_refresh.rs |
Threads execution context through action discovery and prompt build. |
tests/*.rs and various mod tests blocks |
Updates mocks for new trait methods/signatures and adds coverage for capability background + filtering behavior. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2b48fe37ea
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| let Some(extension_statuses) = extension_statuses.as_ref() else { | ||
| continue; | ||
| }; | ||
| let Some(extension) = |
There was a problem hiding this comment.
Medium Severity · Confidence: Likely
Latent actions silently dropped from LLM tool list — auth-on-first-call pattern may be broken
Previously, available_actions included latent extension actions (tools from inactive/uninstalled providers like gmail_send), which allowed the LLM to call them and trigger an auth flow on first use. This new projector filters them out — latent actions now only appear in available_capabilities (background context the LLM can read but not invoke).
The old test was renamed from available_actions_include_latent_inactive_provider_actions to available_actions_omit_latent_inactive_provider_actions, confirming this is intentional. However, if the latent-action-triggers-auth-gate pattern is still expected to work for providers like Gmail, this breaks that UX flow — the LLM can see Gmail exists but cannot attempt a call that would trigger the auth gate.
Suggestion: Verify downstream that the auth gate flow still works for latent providers through the capabilities-only surface. If the LLM needs to be able to invoke latent tools to trigger auth, this needs a different approach (e.g., keeping latent actions in available_actions with a flag, or having the engine auto-trigger auth when the LLM references a capability).
| let tool_defs = tools.tool_definitions().await; | ||
| let extension_statuses = if let Some(auth_manager) = auth_manager { | ||
| match auth_manager | ||
| .list_capability_extensions(&context.user_id) |
There was a problem hiding this comment.
Medium Severity · Confidence: Likely
Double call to list_capability_extensions per LLM step
Both ActionProjector::project (line 35 here) and CapabilityProjector::project (capability_projector.rs:40) independently call auth_manager.list_capability_extensions(&context.user_id), which hits ExtensionManager each time. In loop_engine.rs:240-262, both projectors are called in sequence for the initial system prompt, duplicating the extension list fetch.
While this only affects the system prompt path (not every step), it's an easy optimization to share the result.
Suggestion: Either cache the extension list at the EffectBridgeAdapter level for the duration of a single projection round, or have the caller fetch once and pass the list into both projectors.
…Auth in actions - Fetch list_capability_extensions once in EffectBridgeAdapter and pass to both ActionProjector and CapabilityProjector via prefetched_extensions - Keep NeedsAuth provider tools in available_actions so the LLM can trigger auth gates by attempting to call them - Add unit tests for NeedsAuth preservation and latent tool omission at the ActionProjector level where extension maps can be controlled Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
| let status = capability_status_for_extension(extension, false); | ||
| // NeedsAuth tools remain in available_actions so the LLM can | ||
| // trigger the auth-on-first-call gate by attempting to call them. | ||
| if status == CapabilityStatus::NeedsAuth { |
There was a problem hiding this comment.
Both findings addressed in 3e6359d:
Fix 1 — Double fetch eliminated: Added fetch_extension_map / fetch_extension_list to EffectBridgeAdapter with a short-lived per-user cache. Both available_actions and available_capabilities now go through this shared fetch — the second call within the same projection round hits the cache. ActionProjector::project and CapabilityProjector::project accept an optional prefetched_extensions parameter so the adapter can pass the pre-fetched map.
Fix 2 — NeedsAuth preserved in actions: ActionProjector now bypasses the assign_surface filter for tools whose provider extension has NeedsAuth status (lines 93-96), keeping them callable so the LLM can trigger the auth gate on first use. Unit test needs_auth_provider_tools_remain_in_available_actions in action_projector.rs verifies this with a manually-constructed extension map. The original EffectBridgeAdapter-level test was removed because fake WASM files produce installed=false (→ AvailableNotInstalled), not NeedsAuth.
serrrfirat
left a comment
There was a problem hiding this comment.
Paranoid architect re-review -- approved.
Both Medium findings from the initial review have been correctly addressed in 3e6359d:
- Double list_capability_extensions fetch eliminated via fetch_extension_list with short-lived cache in EffectBridgeAdapter
- NeedsAuth provider tools preserved in available_actions for auth-on-first-call gate triggering
Remaining findings are all Low (channel N+1 lookup, no default trait impl for test ergonomics, behavioral change when auth_manager is None). None are blocking.
Well-structured two-surface separation, clean DRY extraction of thread_execution_context, thorough test coverage across projectors and prompt builder.
…ction Resolve conflicts from base branch's approval_gated removal and ReadyScoped/Error surface policy fix: - tool_surface.rs: take base branch version (ReadyScoped -> capabilities_only, Error -> neither via early return) - effect_adapter.rs: keep projector imports, drop deleted bridge::auth_manager import (now in auth::extension) - mod.rs: keep projector module declarations, remove deleted auth_manager module - action_projector.rs / capability_projector.rs: remove all approval_gated fields from SurfacePolicyInput construction sites, update auth_manager import path to auth::extension - capability_projector test: update assertion for Error status items which are now excluded from capabilities surface Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Add canonical engine capability status enum * Add bridge tool surface assignment policy * fix(engine): tighten scoped surface assignment * fix(bridge): remove premature approval_gated field, surface ReadyScoped in capabilities - Remove approval_gated from SurfacePolicyInput (YAGNI until policy uses it) - Change ReadyScoped fallback from neither() to capabilities_only() so scoped subjects remain visible in background context - Update tests to match new behavior Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(ci): unblock section 2 policy PR * engine-v2: add capability projection and two-surface prompt baseline (#2826) * Add capability projection and two-surface prompt baseline * Reduce step-context args for clippy-clean two-surface stack * fix(engine): address two-surface review follow-ups * fix(engine): normalize alias-aware capability projection * fix(bridge): share extension fetch between projectors, preserve NeedsAuth in actions - Fetch list_capability_extensions once in EffectBridgeAdapter and pass to both ActionProjector and CapabilityProjector via prefetched_extensions - Keep NeedsAuth provider tools in available_actions so the LLM can trigger auth gates by attempting to call them - Add unit tests for NeedsAuth preservation and latent tool omission at the ActionProjector level where extension maps can be controlled Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: serrrfirat <f@nuff.tech> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: serrrfirat <f@nuff.tech> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Add canonical engine capability status enum * Add bridge tool surface assignment policy * fix(engine): tighten scoped surface assignment * fix(bridge): remove premature approval_gated field, surface ReadyScoped in capabilities - Remove approval_gated from SurfacePolicyInput (YAGNI until policy uses it) - Change ReadyScoped fallback from neither() to capabilities_only() so scoped subjects remain visible in background context - Update tests to match new behavior Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(ci): unblock section 2 policy PR * engine-v2: add capability projection and two-surface prompt baseline (nearai#2826) * Add capability projection and two-surface prompt baseline * Reduce step-context args for clippy-clean two-surface stack * fix(engine): address two-surface review follow-ups * fix(engine): normalize alias-aware capability projection * fix(bridge): share extension fetch between projectors, preserve NeedsAuth in actions - Fetch list_capability_extensions once in EffectBridgeAdapter and pass to both ActionProjector and CapabilityProjector via prefetched_extensions - Keep NeedsAuth provider tools in available_actions so the LLM can trigger auth gates by attempting to call them - Add unit tests for NeedsAuth preservation and latent tool omission at the ActionProjector level where extension maps can be controlled Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: serrrfirat <f@nuff.tech> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: serrrfirat <f@nuff.tech> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Summary
#2767task plan: Epic: Separate engine v2 capability background from callable tool schemas #2767 (comment)Finished
available_actions()cleanupCapabilitySummary/CapabilitySummaryKindplus requiredavailable_capabilities(...)on the engine boundarythread_execution_context(...)so engine callers reuse one thread-to-context extraction pathCapabilityProjectorfor provider/channel capability background from runtime truthActionProjectorfor callable-onlyavailable_actions()filteringavailable_capabilities()fallbackavailable_actions()Left
CapabilitySummaryKind::Runtimeprojection yetVerification
cargo fmtcargo test -p ironclaw_engine --libcargo test provider_extension_lookup_accepts_legacy_hyphen_alias --libcargo test available_actions_omit_registered_tool_when_provider_is_not_installed --libcargo test -p ironclaw_engine prompt_with_capabilities_includes_background_statuses --libcargo test -p ironclaw_engine capability_summary_allows_minimal_payload_and_omits_none_fields --libcargo test -p ironclaw_engine system_prompt_includes_capability_background --lib