Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
51 changes: 31 additions & 20 deletions docs/guides/deployment-diagnostics.md
Original file line number Diff line number Diff line change
Expand Up @@ -172,26 +172,36 @@ load-bearing evidence in that case.
### Short-lived crashed pod logs are reaped quickly

`/agent-diagnose` depends on `get_container_logs` to surface the log tail.
[#3547](https://github.com/jwbron/egg/issues/3547) closed most of this gap:
`remove_agent_job` now snapshots the pod's log tail into a Redis-backed
agent-log store just before Job deletion, and `get_container_logs` falls
back to that capture when the live Pod is already gone — the result carries
`"source": "persisted"` plus `captured_at` / `exit_code`. The hard
`logs_unavailable` / `pods "<pod>" not found` case is now limited to
captures that themselves expired (24h TTL) or were never written (the
capture raced the Pod's own teardown, or the Job carried no pipeline
label).

The skill detects a genuine miss and surfaces it in the Top finding — it
As of [#3547](https://github.com/jwbron/egg/issues/3547), `remove_agent_job`
snapshots the pod's log tail into a Redis-backed agent-log store just
before deleting the Job, and `get_container_logs` (both the REST route and
the MCP tool) falls back to that capture once the live Pod is gone — the
response carries `"source": "persisted"` plus `captured_at` / `exit_code`.
This covers the common case: the orchestrator observes the exit and reaps
the Job itself.

The `logs_unavailable` / `pods "<pod>" not found` case is now narrower: it
still occurs when the Pod was removed by some path other than
`remove_agent_job` (e.g. raw k8s `ttlSecondsAfterFinished` GC winning the
race before the orchestrator reaped it), the capture's 24-hour TTL has
expired, the Job carried no pipeline label, or the Redis store itself is
unavailable. Note the capture is strictly best-effort even on the
`remove_agent_job` path: if the pod was already GC'd by the time
`remove_agent_job` ran, the pre-removal log read comes back empty and no
capture is written — so a missing capture does **not** imply the
orchestrator failed to reap the Job.

The skill detects an outright miss and surfaces it in the Top finding — it
does **not** silently return "no classifier match" when logs are simply
gone. You will still get the Job spec, Events, and env keys; the classifier
runs against the reduced evidence.

`get_agent_transcript` is a second, independent source of evidence: it
reads the agent's Claude Code session transcript from the session-state
store (pushed on every event-pod exit, ~6h TTL), so it can survive even
when the agent-log capture itself is missing — but only if the agent got
far enough to push a transcript at least once.
A second, longer-lived evidence source is the `get_agent_transcript` MCP
tool, which reads the agent's Claude Code session transcript from the
session-state store (pushed on every event-pod exit, ~6-hour retention) —
useful for diagnosing *why* an agent exited when even the persisted pod-log
capture is unavailable. It survives independently of the agent-log capture,
but only if the agent got far enough to push a transcript at least once.

**Concurrent BRC phases**: For pipelines running in concurrent execution
mode, the orchestrator now captures a frozen exit snapshot
Expand All @@ -211,10 +221,11 @@ egg-orch phase get <pipeline-id> --json | jq '.data.phase_execution.agent_exits'

The `container_id` in each entry can also be fed directly to
`/agent-diagnose` while the Pod still exists; `last_lines` gives you the
log tail after it doesn't. This means the `short-lived pod, log
unavailable` failure class is mitigated for concurrent phases — provided
the orchestrator's exit-poll observed the exit before the Pod was
GC'd. If the poll lost the race with k8s GC, the log fetch fails
log tail after it doesn't. This is on top of the general agent-log-store
fallback described above — `agent_exits` is concurrent-phase-only and
lives in pipeline state, while the agent-log-store capture applies to any
agent role reaped via `remove_agent_job`, concurrent or not. If neither
path observed the exit before the Pod was GC'd, the log fetch fails
silently and `last_lines` is `[]`; you still have the role, exit code,
and `terminated_at`, but you'll need to fall back to the Job spec and
Events for the log content itself.
Expand Down
14 changes: 12 additions & 2 deletions orchestrator/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -171,8 +171,10 @@ in-cluster agents cannot bypass HITL phase gates. Successful calls log a
| `GET` | `/pipelines/{id}/containers/{cid}` | Get container details |
| `DELETE` | `/pipelines/{id}/containers/{cid}` † | Remove container |
| `POST` | `/pipelines/{id}/containers/{cid}/stop` † | Stop container |
| `GET` | `/pipelines/{id}/containers/{cid}/logs` | Get container logs |
| `GET` | `/pipelines/{id}/containers/{cid}/logs` | Get container logs (falls back to the persisted post-reap capture if the live pod is gone, #3547) |
| `GET` | `/pipelines/{id}/containers/{cid}/health` | Container health check |
| `GET` | `/pipelines/{id}/agent-logs` | List persisted post-reap agent log captures, metadata only (#3547) |
| `GET` | `/pipelines/{id}/agent-logs/{job_name}` | Get one persisted post-reap agent log capture, including the log body (#3547) |

### HITL Decisions

Expand Down Expand Up @@ -203,6 +205,14 @@ in-cluster agents cannot bypass HITL phase gates. Successful calls log a
| `POST` | `/pipelines/{id}/progress` | Emit structured progress event |
| `GET` | `/pipelines/{id}/progress` | Query progress events (filterable by agent, time, limit) |

### Session State

| Method | Path | Description |
|--------|------|-------------|
| `POST` | `/pipelines/{id}/session-state` | Push a role's warm-resume record (session_id, window_occupancy, transcript) |
| `GET` | `/pipelines/{id}/session-state` | Pull a role's warm-resume record, or `found: false` on any miss |
| `GET` | `/pipelines/{id}/session-state/index` | List the pipeline's stored records, metadata only — backs the `get_agent_transcript` MCP tool (#3547) |

### Anchors

| Method | Path | Description |
Expand Down Expand Up @@ -322,7 +332,7 @@ These tools require a `gateway_url` and authenticate via a gateway session. The

### Orchestrator-Backed Tools

`submit_task`, `get_status`, `provide_input`, `list_tasks`, `cancel_task`, `check_health`, `list_containers`, `get_container_logs`, `send_message`, `get_consensus_status`, `get_phase`, `get_pipeline_snapshot`, `validate_config`, `restart_agent`, `restart_phase`, `advance_phase`, `start_phase`, `complete_phase`, `populate_contract`
`submit_task`, `get_status`, `provide_input`, `list_tasks`, `cancel_task`, `check_health`, `list_containers`, `get_container_logs`, `get_agent_transcript`, `send_message`, `get_consensus_status`, `get_phase`, `get_pipeline_snapshot`, `validate_config`, `restart_agent`, `restart_phase`, `advance_phase`, `start_phase`, `complete_phase`, `populate_contract`

The `get_status` tool returns an enriched pipeline status response. In addition to the standard fields (`current_phase`, `status`, `pipeline`, `running_agents`, `completed_agents`, `pending_decisions`, `recent_messages`), the response includes server-computed timing fields:

Expand Down
Loading