feat(engine): add code execution failure categorization instrumentation - #2483
Conversation
Add structured error classification to the v2 engine's CodeAct execution path so we can measure whether REPL failures come from Monty VM limitations, LLM logic errors, tool dispatch issues, or resource limits. - Add CodeExecutionFailure enum (8 categories: SyntaxError, RuntimeError, NameLookup, VmPanic, ResourceLimit, ToolError, GatePause, OsDenied) - Add CodeExecutionFailed event kind to EventKind for event sourcing - Tag every error return path in scripting.rs with the correct category - Emit CodeExecutionFailed events from the orchestrator on code errors - Enhance trace analyzer to use structured events (with fallback to message-level pattern matching for pre-instrumentation threads) - Expand fallback error patterns from 4 to 10 Python exception types - Add 13 regression tests covering classification and trace detection This enables aggregate queries like "what % of code failures are Monty VM panics vs LLM generating bad Python" to inform runtime decisions. https://claude.ai/code/session_018jFKVTjv1pkzwJobw43HwP
There was a problem hiding this comment.
Code Review
This pull request introduces structured instrumentation for code execution failures by adding a CodeExecutionFailure enum and emitting CodeExecutionFailed events during the execution process. The analyze_trace logic is updated to utilize these events for more precise issue reporting, while maintaining a fallback for older message-level patterns. Review feedback recommends replacing DefaultHasher with a stable hashing algorithm like XxHash64 to ensure consistency across compiler versions for persisted events, and suggests using to_ascii_lowercase() for error classification to follow project conventions.
| pub fn code_hash(code: &str) -> String { | ||
| use std::collections::hash_map::DefaultHasher; | ||
| use std::hash::{Hash, Hasher}; | ||
| let mut hasher = DefaultHasher::new(); | ||
| code.hash(&mut hasher); | ||
| format!("{:016x}", hasher.finish()) | ||
| } |
There was a problem hiding this comment.
DefaultHasher is not guaranteed to produce stable hashes across different compiler versions. Since these hashes are used for correlating events that may be persisted, this could lead to inconsistencies over time. It's recommended to use a hashing algorithm with a specified, stable output, such as XxHash64 from the twox-hash crate.
| fn classify_runtime_error(error_msg: &str) -> crate::types::step::CodeExecutionFailure { | ||
| use crate::types::step::CodeExecutionFailure; | ||
|
|
||
| let lower = error_msg.to_lowercase(); |
There was a problem hiding this comment.
Per project conventions, to_ascii_lowercase() should be used for case-insensitive comparisons on ASCII text. This avoids locale-dependent behavior and is more efficient. The error messages from Monty are expected to be ASCII.
| let lower = error_msg.to_lowercase(); | |
| let lower = error_msg.to_ascii_lowercase(); |
References
- For case-insensitive comparisons, use
to_ascii_lowercase()as it is the project convention, especially when the text is known to be ASCII.
| if let Some(ref category) = result.failure_category { | ||
| let error_text: String = result.stdout.chars().take(500).collect(); | ||
| let instrumentation_event = ThreadEvent::new( | ||
| thread.id, |
There was a problem hiding this comment.
Medium Severity — error_text takes first 500 chars of stdout instead of last
let error_text: String = result.stdout.chars().take(500).collect();When code has print output before the error (e.g., a loop printing 100 items then hitting a NameError), the first 500 chars capture the print statements, not the error traceback. The project's own compact_output_metadata() in scripting.rs:88-96 takes the last N chars for exactly this reason.
Suggestion: Extract from the tail instead:
let chars: Vec<char> = result.stdout.chars().collect();
let start = chars.len().saturating_sub(500);
let error_text: String = chars[start..].iter().collect();Note: the existing ActionFailed path at line 806 (result.stdout.chars().take(500)) has the same issue — worth fixing both with a shared tail-extract helper.
| @@ -641,6 +656,9 @@ pub async fn execute_code_with_skills( | |||
| recursive_tokens, | |||
| final_answer: None, | |||
| had_error, | |||
There was a problem hiding this comment.
Medium Severity — GatePaused sets failure_category but the event is never emitted
At this return site, the code sets failure_category: Some(GatePause) but had_error is false (verified: every had_error = true in this function precedes an early return, so it's still false when GatePaused is reached).
In orchestrator.rs, the CodeExecutionFailed event emission is nested inside if result.had_error { ... }, so the GatePause category is silently discarded — it's dead data that will never appear in aggregate analysis.
Suggestion: Either:
- Don't set
failure_categoryhere (a gate pause is a suspension, not a failure), or - Emit
CodeExecutionFailedinorchestrator.rsbased onfailure_category.is_some()independently ofhad_error, so gate pauses are tracked in the instrumentation
ReviewUseful instrumentation, low blast radius. Three things worth fixing before merge: (1) error-text truncation captures the wrong end, (2) Correctness
Trace fallback
Persistence / forward-compat
Conventions
Test coverage
Suggestions (file:line)
|
…ntation correctness (#2483) - Replace `had_error: bool` + `failure_category: Option<_>` with single `failure: Option<CodeExecutionFailure>` field, making invalid states unrepresentable - Remove `GatePause` variant (gate pauses are suspensions, not failures) - Convert all 9 catch_unwind Err(_) paths to emit VmPanic instead of propagating EngineError, so panics get proper instrumentation events - Fix error_text truncation: take last 500 chars (where tracebacks are), not first 500 chars (where print output is) - Thread real `Instant::now()` timing through `duration_ms` instead of hardcoded 0 - Tighten `classify_runtime_error`: "syntax" → "syntaxerror", remove loose `"os" && "denied"` substring match, reorder checks - Replace `DefaultHasher` with FNV-1a for stable cross-version hashing - Add `#[serde(rename_all = "snake_case")]` so Serialize matches Display - Add `#[serde(other)] Unknown` to EventKind for forward-compat - Fix misleading CodeExecutionFailed docstring - Add orchestrator caller test for CodeExecutionFailed event emission - Use `to_ascii_lowercase()` per project convention Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Addressed review feedback (88d3f63)All feedback from @ilblackdragon, @gemini-code-assist[bot], and self-review has been addressed: Correctness
Stability
Tests
Style
All 395 tests pass, |
|
@ilblackdragon — addressed all your feedback in 88d3f63. Here's a walkthrough of the design decisions: 1.
|
…e tautological assert - Replace `!result.failure.is_some()` with `result.failure.is_none()` in scripting tests - Remove tautological `duration_ms >= 0` assertion on unsigned type in orchestrator test - Add missing V24 migration checksum to checksums.lock Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ReviewOverviewIntroduces Strengths
Correctness concerns1. Behavior change on VM panic ( 2. PR description mismatch — 3. Unrelated migration churn. 4. Fallback branch is all-or-nothing. Per the comment in Nits / quality
Security / perfNo new I/O or allocations in a hot path. SummarySolid, well-tested instrumentation. Three things to resolve before merge:
|
…ighten fuel match, clippy (#2483) - Extract `tail_chars(s, n)` helper to deduplicate last-N-chars logic in `handle_execute_code_step` (used by both ActionFailed and CodeExecutionFailed event emission) - Tighten `classify_runtime_error` fuel match from `contains("fuel")` to `contains("out of fuel") || contains("fuel exhausted")` to avoid miscategorizing runtime errors that mention the word "fuel" - Fix test message typo: "should set had_error" → "should set failure" - Fix clippy warnings: `!x.is_some()` → `x.is_none()`, remove tautological `u64 >= 0` assertion Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
Addressed review feedback from #2483 (comment): Code fixes (d5baf70):
PR description updates:
|
Rebase artifact — no V24 migration exists in this PR. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Rewrite the code-review skill from a 6-bullet checklist into a paranoid-architect workflow that handles both local diffs and GitHub PRs end-to-end: - Two input shapes: local `git diff` or `owner/repo N` / `github.com/.../pull/N` URLs. - Step 1 wraps GitHub fetches in `async def` + `FINAL(await ...)` to avoid the closure-capture quirk that kept tripping LLMs (see the paired codeact preamble update); reads metadata, diff, and files via three sequential awaits instead of `asyncio.gather`. - Step 2 reads each changed file in full (raw media type, no base64 module needed) so reviews account for surrounding context. - Step 3 runs the change through six lenses: correctness, edge cases, security (with a real adversarial checklist), test coverage, docs, architecture. - Step 4 renders findings as a severity table and asks which to post. - Step 5 posts line-level comments via the PR comments endpoint with the captured head SHA, falling back to issue comments for multi-file findings. Bumps `requires.skills` to include `github` so the activation pulls in the GitHub API recipes via the chain-loader. Adds a live e2e test (`e2e_live_code_review.rs`) plus a recorded trace fixture (PR #2483) so the workflow is replayable without hitting GitHub.
henrypark133
left a comment
There was a problem hiding this comment.
Review: Code execution failure categorization
No verified findings in the current diff. The change threads structured code-execution failure categories from the scripting runtime into the orchestrator caller and trace analyzer, and it keeps the fallback message-based detection for older threads that predate the new event.
Code reviewFound 5 issues:
Consider iterating once: collect into a Vec, or use byte-based slicing if ASCII.
Notes:
|
Rewrite the code-review skill from a 6-bullet checklist into a paranoid-architect workflow that handles both local diffs and GitHub PRs end-to-end: - Two input shapes: local `git diff` or `owner/repo N` / `github.com/.../pull/N` URLs. - Step 1 wraps GitHub fetches in `async def` + `FINAL(await ...)` to avoid the closure-capture quirk that kept tripping LLMs (see the paired codeact preamble update); reads metadata, diff, and files via three sequential awaits instead of `asyncio.gather`. - Step 2 reads each changed file in full (raw media type, no base64 module needed) so reviews account for surrounding context. - Step 3 runs the change through six lenses: correctness, edge cases, security (with a real adversarial checklist), test coverage, docs, architecture. - Step 4 renders findings as a severity table and asks which to post. - Step 5 posts line-level comments via the PR comments endpoint with the captured head SHA, falling back to issue comments for multi-file findings. Bumps `requires.skills` to include `github` so the activation pulls in the GitHub API recipes via the chain-loader. Adds a live e2e test (`e2e_live_code_review.rs`) plus a recorded trace fixture (PR #2483) so the workflow is replayable without hitting GitHub.
…ates (#2528) * feat(skills): paranoid-architect code-review skill v2 Rewrite the code-review skill from a 6-bullet checklist into a paranoid-architect workflow that handles both local diffs and GitHub PRs end-to-end: - Two input shapes: local `git diff` or `owner/repo N` / `github.com/.../pull/N` URLs. - Step 1 wraps GitHub fetches in `async def` + `FINAL(await ...)` to avoid the closure-capture quirk that kept tripping LLMs (see the paired codeact preamble update); reads metadata, diff, and files via three sequential awaits instead of `asyncio.gather`. - Step 2 reads each changed file in full (raw media type, no base64 module needed) so reviews account for surrounding context. - Step 3 runs the change through six lenses: correctness, edge cases, security (with a real adversarial checklist), test coverage, docs, architecture. - Step 4 renders findings as a severity table and asks which to post. - Step 5 posts line-level comments via the PR comments endpoint with the captured head SHA, falling back to issue comments for multi-file findings. Bumps `requires.skills` to include `github` so the activation pulls in the GitHub API recipes via the chain-loader. Adds a live e2e test (`e2e_live_code_review.rs`) plus a recorded trace fixture (PR #2483) so the workflow is replayable without hitting GitHub. * docs(github): clarify search endpoints, response envelope, @me queries LLMs kept inventing a `search_issues` action and looping over `/repos/{owner}/{repo}/pulls` for "my PRs" queries. Clarify the GitHub tool surface in three places: - `tools-src/github/src/lib.rs` and `registry/tools/github.json`: enumerate the three real search actions and call out that `search_issues_pull_requests` covers both. Add the canonical `is:pr author:@me sort:updated-desc` recipe for cross-repo "my PRs". - `skills/github/SKILL.md`: add an "Authenticated User & Cross-Repo Queries" section with copy-paste recipes for `@me`, the search endpoints with proper URL encoding, and the response-envelope contract (`body` is parsed JSON for application/json, raw `str` for diff endpoints — never call `json.loads()` on it, never write `.get("body", body)` as a fallback). * fix: resolve CI failures — clippy useless_conversion + missing test harness methods - Remove `.into_iter()` on `details` in catalog.rs (clippy::useless_conversion) - Add `with_skills_dir` to `LiveTestHarnessBuilder` for e2e_live_code_review test - Add `active_skill_names` to `TestRig` extracting from SkillActivated status events Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(skills): address zmanian + gemini review — URL encoding, multi-line comments, description trimming (#2528) - URL-encode file paths in GitHub API content URLs - Add start_line/start_side to multi-line comment example - Add 'locally' keyword override for mode detection - Trim overly long schema descriptions - Remove duplicated /search/issues note from Common Mistakes - Fetch PR title from trace fixture instead of hard-coding Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(test): propagate skills_dir into TestRig config (#2528) LiveTestHarnessBuilder::with_skills_dir() stored a PathBuf but only used it as an is_some() flag — the actual SkillRegistry always pointed at an empty temp directory. Now the stored path flows through TestRigBuilder::with_skills_dir() into config.skills.local_dir and the SkillRegistry constructor. Also generalizes the hardcoded nearai/ironclaw repo name in the github skill's response-handling example to {owner}/{repo}. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…on (nearai#2483) * feat(engine): add code execution failure categorization instrumentation Add structured error classification to the v2 engine's CodeAct execution path so we can measure whether REPL failures come from Monty VM limitations, LLM logic errors, tool dispatch issues, or resource limits. - Add CodeExecutionFailure enum (8 categories: SyntaxError, RuntimeError, NameLookup, VmPanic, ResourceLimit, ToolError, GatePause, OsDenied) - Add CodeExecutionFailed event kind to EventKind for event sourcing - Tag every error return path in scripting.rs with the correct category - Emit CodeExecutionFailed events from the orchestrator on code errors - Enhance trace analyzer to use structured events (with fallback to message-level pattern matching for pre-instrumentation threads) - Expand fallback error patterns from 4 to 10 Python exception types - Add 13 regression tests covering classification and trace detection This enables aggregate queries like "what % of code failures are Monty VM panics vs LLM generating bad Python" to inform runtime decisions. https://claude.ai/code/session_018jFKVTjv1pkzwJobw43HwP * fix(engine): address ilblackdragon + gemini review — failure instrumentation correctness (nearai#2483) - Replace `had_error: bool` + `failure_category: Option<_>` with single `failure: Option<CodeExecutionFailure>` field, making invalid states unrepresentable - Remove `GatePause` variant (gate pauses are suspensions, not failures) - Convert all 9 catch_unwind Err(_) paths to emit VmPanic instead of propagating EngineError, so panics get proper instrumentation events - Fix error_text truncation: take last 500 chars (where tracebacks are), not first 500 chars (where print output is) - Thread real `Instant::now()` timing through `duration_ms` instead of hardcoded 0 - Tighten `classify_runtime_error`: "syntax" → "syntaxerror", remove loose `"os" && "denied"` substring match, reorder checks - Replace `DefaultHasher` with FNV-1a for stable cross-version hashing - Add `#[serde(rename_all = "snake_case")]` so Serialize matches Display - Add `#[serde(other)] Unknown` to EventKind for forward-compat - Fix misleading CodeExecutionFailed docstring - Add orchestrator caller test for CodeExecutionFailed event emission - Use `to_ascii_lowercase()` per project convention Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: resolve clippy warnings — simplify boolean expressions and remove tautological assert - Replace `!result.failure.is_some()` with `result.failure.is_none()` in scripting tests - Remove tautological `duration_ms >= 0` assertion on unsigned type in orchestrator test - Add missing V24 migration checksum to checksums.lock Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): address ilblackdragon review nits — tail_chars helper, tighten fuel match, clippy (nearai#2483) - Extract `tail_chars(s, n)` helper to deduplicate last-N-chars logic in `handle_execute_code_step` (used by both ActionFailed and CodeExecutionFailed event emission) - Tighten `classify_runtime_error` fuel match from `contains("fuel")` to `contains("out of fuel") || contains("fuel exhausted")` to avoid miscategorizing runtime errors that mention the word "fuel" - Fix test message typo: "should set had_error" → "should set failure" - Fix clippy warnings: `!x.is_some()` → `x.is_none()`, remove tautological `u64 >= 0` assertion Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): drop stray V24 checksums.lock entry (nearai#2483) Rebase artifact — no V24 migration exists in this PR. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com>
…ates (nearai#2528) * feat(skills): paranoid-architect code-review skill v2 Rewrite the code-review skill from a 6-bullet checklist into a paranoid-architect workflow that handles both local diffs and GitHub PRs end-to-end: - Two input shapes: local `git diff` or `owner/repo N` / `github.com/.../pull/N` URLs. - Step 1 wraps GitHub fetches in `async def` + `FINAL(await ...)` to avoid the closure-capture quirk that kept tripping LLMs (see the paired codeact preamble update); reads metadata, diff, and files via three sequential awaits instead of `asyncio.gather`. - Step 2 reads each changed file in full (raw media type, no base64 module needed) so reviews account for surrounding context. - Step 3 runs the change through six lenses: correctness, edge cases, security (with a real adversarial checklist), test coverage, docs, architecture. - Step 4 renders findings as a severity table and asks which to post. - Step 5 posts line-level comments via the PR comments endpoint with the captured head SHA, falling back to issue comments for multi-file findings. Bumps `requires.skills` to include `github` so the activation pulls in the GitHub API recipes via the chain-loader. Adds a live e2e test (`e2e_live_code_review.rs`) plus a recorded trace fixture (PR nearai#2483) so the workflow is replayable without hitting GitHub. * docs(github): clarify search endpoints, response envelope, @me queries LLMs kept inventing a `search_issues` action and looping over `/repos/{owner}/{repo}/pulls` for "my PRs" queries. Clarify the GitHub tool surface in three places: - `tools-src/github/src/lib.rs` and `registry/tools/github.json`: enumerate the three real search actions and call out that `search_issues_pull_requests` covers both. Add the canonical `is:pr author:@me sort:updated-desc` recipe for cross-repo "my PRs". - `skills/github/SKILL.md`: add an "Authenticated User & Cross-Repo Queries" section with copy-paste recipes for `@me`, the search endpoints with proper URL encoding, and the response-envelope contract (`body` is parsed JSON for application/json, raw `str` for diff endpoints — never call `json.loads()` on it, never write `.get("body", body)` as a fallback). * fix: resolve CI failures — clippy useless_conversion + missing test harness methods - Remove `.into_iter()` on `details` in catalog.rs (clippy::useless_conversion) - Add `with_skills_dir` to `LiveTestHarnessBuilder` for e2e_live_code_review test - Add `active_skill_names` to `TestRig` extracting from SkillActivated status events Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(skills): address zmanian + gemini review — URL encoding, multi-line comments, description trimming (nearai#2528) - URL-encode file paths in GitHub API content URLs - Add start_line/start_side to multi-line comment example - Add 'locally' keyword override for mode detection - Trim overly long schema descriptions - Remove duplicated /search/issues note from Common Mistakes - Fetch PR title from trace fixture instead of hard-coding Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(test): propagate skills_dir into TestRig config (nearai#2528) LiveTestHarnessBuilder::with_skills_dir() stored a PathBuf but only used it as an is_some() flag — the actual SkillRegistry always pointed at an empty temp directory. Now the stored path flows through TestRigBuilder::with_skills_dir() into config.skills.local_dir and the SkillRegistry constructor. Also generalizes the hardcoded nearai/ironclaw repo name in the github skill's response-handling example to {owner}/{repo}. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Summary
CodeExecutionFailureenum (7 variants: SyntaxError, RuntimeError, NameLookup, VmPanic, ResourceLimit, ToolError, OsDenied) to classify why code execution failsfailurethroughCodeExecutionResult(replacinghad_error: bool) and emits structuredCodeExecutionFailedevents for aggregate analysis of failure modes (Monty limitation vs LLM logic error vs tool dispatch failure)Caller audit (Err → Ok(failure=VmPanic) shift)
execute_code/execute_code_with_skillspreviously returnedErr(EngineError::Effect)on VM panic; now returnsOk(CodeExecutionResult { failure: Some(VmPanic) }). The only production call site isorchestrator.rs:handle_execute_code_stepwhich correctly dispatches onresult.failure. Test-only callers inscripting.rswere also updated. No other callers exist.Known limitation: mixed-era fallback
The trace analyzer's backward-compatible fallback (message-scraping for pre-instrumentation threads) is all-or-nothing: if a thread has any
CodeExecutionFailedevent, message-scraping is skipped entirely. Errors from pre-instrumentation steps in a mixed-era thread will go unreported. This is acceptable during the transition period since all new threads will have full instrumentation.Test plan
cargo test -p ironclaw_engine --lib— 395 passcargo clippy --all --all-features— zero engine warnings🤖 Generated with Claude Code