Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .changeset/error-recovery-skill.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"@bradygaster/squad-cli": minor
"@bradygaster/squad-sdk": minor
---
feat: add error-recovery skill for standard agent failure recovery patterns
69 changes: 69 additions & 0 deletions docs/proposals/error-recovery.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# Proposal: Error Recovery Skill

**Issue:** bradygaster/squad#623
**Author:** tamirdresher
**Date:** 2026-03-26
**Status:** Proposal

---

## Problem Statement

When a Squad agent fails (model timeout, tool error, invalid output, context overflow), there is
no standardized recovery pattern. Individual coordinators implement ad-hoc retry logic or simply
fail the task. This causes inconsistent user experience and missed opportunities for graceful
degradation across the squad.

---

## Proposed Approach

A skill library providing five recovery patterns an agent can apply when it encounters a failure:

| Pattern | When to use |
|---------|-------------|
| **retry** | Transient errors (rate limit, timeout) — wait and retry same approach |
| **fallback** | Primary approach consistently fails — switch to alternative method |
| **diagnose** | Unclear failure cause — gather diagnostics before deciding |
| **escalate** | Blocked beyond agent capability — surface to coordinator/human |
| **degrade** | Full functionality unavailable — deliver partial result with caveat |

The skill provides a selection guide table mapping error symptoms to the appropriate pattern,
plus prompt templates for each pattern that agents can use in their reasoning.

Copilot AI Mar 28, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This proposal says the skill includes “prompt templates for each pattern,” but the new SKILL.md content only contains narrative guidance/examples (no reusable prompt templates). Either add the prompt-template sections to the skill files or adjust this proposal text so it matches what’s actually being shipped.

Suggested change
plus prompt templates for each pattern that agents can use in their reasoning.
plus narrative guidance and example prompts for each pattern that agents can adapt in their reasoning.

Copilot uses AI. Check for mistakes.

---

## Fit with Existing Architecture

- **Complements** existing gent-conduct skill (which covers behavior) — this skill covers failure states

Copilot AI Mar 28, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There’s a non-printable/control character before gent-conduct in this line, which will render incorrectly in Markdown and makes the skill reference harder to read/search. Replace it with plain text agent-conduct.

Suggested change
- **Complements** existing gent-conduct skill (which covers behavior) — this skill covers failure states
- **Complements** existing agent-conduct skill (which covers behavior) — this skill covers failure states

Copilot uses AI. Check for mistakes.
- **No code changes** — template-only, agents apply this via their prompt reasoning
- **Coordinator-agnostic** — works with any existing coordinator style
- **Consistent with** the 3-cycle protocol from iterative-retrieval skill

---

## What Changes

- New skill: packages/squad-cli/templates/skills/error-recovery/SKILL.md
- New skill: packages/squad-sdk/templates/skills/error-recovery/SKILL.md
- New changeset: .changeset/error-recovery-skill.md

## What Stays the Same

- No existing skills modified
- No CLI or SDK runtime code changed

---

## Risks and Mitigations

| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|------------|
| Agents over-apply retry (masking root causes) | Low | Medium | Skill explicitly limits retry to 3 attempts max |
| Conflicts with future built-in error handling | Low | Low | Template-only — superseded naturally if runtime handles it |

---

## References

- Issue: bradygaster/squad#623
99 changes: 99 additions & 0 deletions packages/squad-cli/templates/skills/error-recovery/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,99 @@
---
name: "error-recovery"
description: "Standard recovery patterns for all squad agents. When something fails, adapt — don't just report the failure."
domain: "reliability, agent-coordination"
confidence: "high"
license: MIT

Copilot AI Mar 28, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Skill frontmatter deviates from the convention used by other skills: it includes license: MIT but is missing the source: field that’s present across existing skill templates. Consider adding a source: entry (e.g., earned/manual/extracted) and dropping license if it’s not consumed anywhere, to keep metadata consistent and avoid unused fields.

Suggested change
license: MIT
source: manual

Copilot uses AI. Check for mistakes.
---

# Error Recovery Patterns

Standard recovery patterns for all squad agents. When something fails, **adapt** — don't just report the failure.

---

## 1. Retry with Backoff

**When:** Transient failures — API timeouts, rate limits, network errors, temporary service unavailability.

**Pattern:**
1. Wait briefly, then retry (start at 2s, double each attempt)
2. Maximum 3 retries before escalating
3. Log each attempt with the error received

**Example:** API call returns 429 Too Many Requests → wait 2s → retry → wait 4s → retry → wait 8s → retry → escalate if still failing.

---

## 2. Fallback Alternatives

**When:** Primary tool or approach fails and an alternative exists.

**Pattern:**
1. Attempt primary approach
2. On failure, identify alternative tool/method
3. Try the alternative with the same intent
4. Document which alternative was used and why

**Example:** Primary CLI tool fails → fall back to direct API call for the same operation.

---

## 3. Diagnose-and-Fix

**When:** Build failures, test failures, linting errors — structured errors with actionable output.

**Pattern:**
1. Read the full error output carefully
2. Identify the root cause from error messages
3. Attempt a targeted fix
4. Re-run to verify the fix
5. Maximum 3 fix-retry cycles before escalating

**Example:** Build fails with a type error → check for missing import → add it → rebuild.

---

## 4. Escalate with Context

**When:** Recovery attempts have been exhausted, or the failure requires human judgment.

**Pattern:**
1. Summarize what was attempted and what failed
2. Include the exact error messages
3. State what you believe the root cause is
4. Suggest next steps or who might be able to help
5. Hand off to the coordinator or the appropriate specialist

**Example:** After 3 failed build attempts → "Build fails on line 42 with null reference. Tried X, Y, Z. Likely a design issue in the Foo module. Recommend the code owner review."

---

## 5. Graceful Degradation

**When:** A non-critical step fails but the overall task can still deliver value.

**Pattern:**
1. Determine if the failed step is critical to the task outcome
2. If non-critical, log the failure and continue
3. Deliver partial results with a clear note of what was skipped
4. Offer to retry the skipped step separately

**Example:** Generating a report with 5 sections — section 3 data source is unavailable → produce the report with 4 sections, note that section 3 was skipped and why.

---

## Applying These Patterns

Each agent should reference these patterns in their charter's `## Error Recovery` section, tailored to their domain. The charter should list the agent's most common failure modes and map each to the appropriate pattern above.

**Selection guide:**

| Failure Type | Primary Pattern | Fallback Pattern |
|---|---|---|
| Network/API transient | Retry with Backoff | Escalate with Context |
| Tool/dependency missing | Fallback Alternatives | Escalate with Context |
| Build/test error | Diagnose-and-Fix | Escalate with Context |
| Auth/permissions | Retry with Backoff | Escalate with Context |
| Non-critical data missing | Graceful Degradation | — |
| Unknown/novel error | Escalate with Context | — |
99 changes: 99 additions & 0 deletions packages/squad-sdk/templates/skills/error-recovery/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,99 @@
---
name: "error-recovery"
description: "Standard recovery patterns for all squad agents. When something fails, adapt — don't just report the failure."
domain: "reliability, agent-coordination"
confidence: "high"
license: MIT

Copilot AI Mar 28, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Skill frontmatter deviates from the convention used by other skills: it includes license: MIT but is missing the source: field that’s present across existing skill templates. Consider adding a source: entry (e.g., earned/manual/extracted) and dropping license if it’s not consumed anywhere, to keep metadata consistent and avoid unused fields.

Suggested change
license: MIT
source: manual

Copilot uses AI. Check for mistakes.
---

# Error Recovery Patterns

Standard recovery patterns for all squad agents. When something fails, **adapt** — don't just report the failure.

---

## 1. Retry with Backoff

**When:** Transient failures — API timeouts, rate limits, network errors, temporary service unavailability.

**Pattern:**
1. Wait briefly, then retry (start at 2s, double each attempt)
2. Maximum 3 retries before escalating
3. Log each attempt with the error received

**Example:** API call returns 429 Too Many Requests → wait 2s → retry → wait 4s → retry → wait 8s → retry → escalate if still failing.

---

## 2. Fallback Alternatives

**When:** Primary tool or approach fails and an alternative exists.

**Pattern:**
1. Attempt primary approach
2. On failure, identify alternative tool/method
3. Try the alternative with the same intent
4. Document which alternative was used and why

**Example:** Primary CLI tool fails → fall back to direct API call for the same operation.

---

## 3. Diagnose-and-Fix

**When:** Build failures, test failures, linting errors — structured errors with actionable output.

**Pattern:**
1. Read the full error output carefully
2. Identify the root cause from error messages
3. Attempt a targeted fix
4. Re-run to verify the fix
5. Maximum 3 fix-retry cycles before escalating

**Example:** Build fails with a type error → check for missing import → add it → rebuild.

---

## 4. Escalate with Context

**When:** Recovery attempts have been exhausted, or the failure requires human judgment.

**Pattern:**
1. Summarize what was attempted and what failed
2. Include the exact error messages
3. State what you believe the root cause is
4. Suggest next steps or who might be able to help
5. Hand off to the coordinator or the appropriate specialist

**Example:** After 3 failed build attempts → "Build fails on line 42 with null reference. Tried X, Y, Z. Likely a design issue in the Foo module. Recommend the code owner review."

---

## 5. Graceful Degradation

**When:** A non-critical step fails but the overall task can still deliver value.

**Pattern:**
1. Determine if the failed step is critical to the task outcome
2. If non-critical, log the failure and continue
3. Deliver partial results with a clear note of what was skipped
4. Offer to retry the skipped step separately

**Example:** Generating a report with 5 sections — section 3 data source is unavailable → produce the report with 4 sections, note that section 3 was skipped and why.

---

## Applying These Patterns

Each agent should reference these patterns in their charter's `## Error Recovery` section, tailored to their domain. The charter should list the agent's most common failure modes and map each to the appropriate pattern above.

**Selection guide:**

| Failure Type | Primary Pattern | Fallback Pattern |
|---|---|---|
| Network/API transient | Retry with Backoff | Escalate with Context |
| Tool/dependency missing | Fallback Alternatives | Escalate with Context |
| Build/test error | Diagnose-and-Fix | Escalate with Context |
| Auth/permissions | Retry with Backoff | Escalate with Context |
| Non-critical data missing | Graceful Degradation | — |
| Unknown/novel error | Escalate with Context | — |
Loading