guardrails: surface fatal provider errors instead of waiting out the deadline - #87
Conversation
A CLI can report a fatal provider error and keep running, so waiting for exit turns a known failure into a spent deadline. Measured 2026-08-06 with exhausted quotas: codex exits nonzero in seconds, pi prints its 429 with a reset time and keeps running, opencode prints nothing at its default log level and keeps running — that last shape is why an exhausted quota was reported as a bare dispatch_timeout. Report dispatch_cli_error from the error line itself, and prefer a route flag that surfaces the error over discovering it by timeout. The canary is now scoped to routes that expose no error channel at all. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
opencode run says nothing about a fatal provider error at its default log level, which is what made an exhausted quota look like a healthy silent run. --print-logs --log-level ERROR surfaces it on stderr and leaves stdout clean, verified on both a failing and a succeeding run. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
🚧 Files skipped from review as they are similar to previous changes (1)
WalkthroughOpenCode 실행에 오류 로그 옵션을 추가했습니다. 확정적 provider 오류를 조기에 감지하고 ChangesProvider 오류 처리
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@skills/productivity/delegate/references/cli-invocations.md`:
- Line 31: Update the OpenCode entries in executors.json so their dispatch argv
includes the required --print-logs and --log-level ERROR flags, matching the
invocation documented in the CLI reference. Add or extend evaluation coverage
for these executors to verify fatal errors are detected from stderr while normal
results are extracted from stdout.
- Line 76: OpenCode 호출에서 global 로그 플래그의 위치가 잘못되었습니다. OpenCode 라우트의 실행 명령을 수정해
print-logs 및 log-level ERROR를 run 서브커맨드보다 앞에 배치하고, auto 플래그와 나머지 인자는 run 뒤에
유지하십시오.
In `@skills/productivity/delegate/references/dispatch-guardrails.md`:
- Around line 22-24: Replace the spawnSync-based execution around the delegate
runner with an asynchronous supervisor that drains stderr concurrently. Detect
permitted fatal provider errors before process exit, terminate the child process
group immediately, and report dispatch_cli_error with the matching stderr line
and reset time. Update the dispatch guardrails contract to define this pre-exit
error path, while preserving post-exit handling for other failures. Add a test
where the child emits a fatal stderr line and continues running, verifying early
termination and dispatch_cli_error reporting.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro Plus
Run ID: a958188b-2120-4189-8380-68b1d23890ee
📒 Files selected for processing (2)
skills/productivity/delegate/references/cli-invocations.mdskills/productivity/delegate/references/dispatch-guardrails.md
| The prompt is already in argv, so connect stdin to DEVNULL for every current route. In a process API, set the child stdin to DEVNULL. Under rule 2, redirect with `< /dev/null` and never pass the redirect as an argv token. Add a future stdin-consuming route as an explicit exception instead of inheriting an open pipe. | ||
|
|
||
| For the OpenCode routes, `--auto` has the same trust implications as other non-interactive permission bypass flags. | ||
| For the OpenCode routes, `--auto` has the same trust implications as other non-interactive permission bypass flags. The two log flags are not optional either: without them a fatal provider error prints nothing, the run stays silent as if healthy, and the failure only surfaces as a spent deadline. They leave stdout clean, so extraction is unaffected. |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/usr/bin/env bash
set -euo pipefail
command -v opencode
opencode --version
opencode run --help | rg -n -- '--print-logs|--log-level'Repository: sungjunlee/skills
Length of output: 155
🏁 Script executed:
#!/usr/bin/env bash
set -euo pipefail
echo "== repo files around target =="
git ls-files | rg '(^skills/productivity/delegate/references/cli-invocations\.md|opencode|package\.json|README\.md)$' || true
echo "== target excerpt =="
sed -n '55,90p' skills/productivity/delegate/references/cli-invocations.md || true
echo "== opencode references =="
rg -n "opencode|print-logs|stdout|stderr|fatal provider" -S . || trueRepository: sungjunlee/skills
Length of output: 26675
🌐 Web query:
opencode 1.17.18 run --print-logs --log-level ERROR source
💡 Result:
The command opencode 1.17.18 run --print-logs --log-level ERROR source is syntactically malformed based on the official OpenCode CLI documentation [1][2]. In the OpenCode CLI, the run command follows the pattern opencode run [message] [flags] [1][2]. The positional argument [message] is the task or prompt you want the AI agent to execute [1][3]. The string "source" at the end of your command would be interpreted by the CLI as a positional argument (a message to the AI), rather than as a flag or configuration parameter [3][2]. To correctly use the provided flags, the command should be structured as follows: opencode --print-logs --log-level ERROR run "your prompt or task here" Explanation of the components: 1. opencode: The main CLI executable [4][5]. 2. --print-logs: A global flag that forces logs to be output directly to stderr [6][2]. 3. --log-level ERROR: A global flag that sets the verbosity of the logs to only report errors [6][2]. 4. run: The command used to execute a prompt non-interactively [1][2]. 5. "your prompt or task here": The required positional argument (message) for the run command [1][2]. If "source" was intended to be the task/message for the AI, the correct command would be: opencode --print-logs --log-level ERROR run "source" Note that the --log-level and --print-logs flags are global options that must precede the run subcommand in the command structure [6][2]. Using them after the run command or after the message may result in them being ignored or treated as part of the prompt [3][2].
Citations:
- 1: https://opencode.ai/docs/cli/
- 2: https://github.com/mseep-ai/OpenCode-Everything-You-Need-to-Know/blob/main/docs/reference/cli.md
- 3: https://github.com/anomalyco/opencode/blob/dev/packages/opencode/src/cli/cmd/run.ts
- 4: https://github.com/anomalyco/opencode
- 5: https://opencode.ai/docs/
- 6: https://github.com/anomalyco/opencode/blob/9afbdc10/packages/opencode/src/index.ts
OpenCode의 --print-logs를 run 호출 전으로 이동하십시오.
OpenCode 문서에 따르면 --print-logs와 --log-level ERROR는 global flag인데 현재 reference에서는 opencode run --auto --print-logs --log-level ERROR ...로 전달합니다. OpenCode 1.17.18 이상에서 이 인자는 입력 메시지로 처리되어 정상 호출이 깨집니다. opencode --print-logs --log-level ERROR run --auto ...로 배치하십시오.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/productivity/delegate/references/cli-invocations.md` at line 76,
OpenCode 호출에서 global 로그 플래그의 위치가 잘못되었습니다. OpenCode 라우트의 실행 명령을 수정해 print-logs 및
log-level ERROR를 run 서브커맨드보다 앞에 배치하고, auto 플래그와 나머지 인자는 run 뒤에 유지하십시오.
| - Treat a definitive provider error on stderr as terminal. Terminate at once and report `dispatch_cli_error` with that line and any reset time it names, rather than waiting for the process to exit or for the deadline. | ||
| - Prefer a route flag that surfaces such errors over discovering them by timeout. `opencode run` needs `--print-logs --log-level ERROR`, which surfaces quota exhaustion about 36 seconds in, after the CLI's internal retries. | ||
| - Send a bounded canary (≤90 s, `Reply with exactly: OK`) only for a route that stays silent and exposes no error channel. A silent canary condemns the route, not the model. |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift
조기 오류 감지 경로를 실행기에 구현해야 합니다.
skills/productivity/delegate/references/dispatch-guardrails.md Lines 22-24는 프로세스 종료 전에 stderr를 읽고 dispatch_cli_error를 보고하도록 정의합니다. 그러나 제공된 실행 경로인 scripts/run-delegate-eval.mjs Lines 181-203은 spawnSync를 사용합니다. 이 코드는 프로세스가 종료되거나 deadline에 도달한 뒤에만 result.stderr를 확인하므로 즉시 종료와 조기 보고를 수행할 수 없습니다.
또한 현재 실행기는 nonzero 종료를 일반 note로만 반환하고, dispatch-guardrails.md Line 43은 dispatch_cli_error를 종료 후 오류로만 정의합니다. 비동기 supervisor로 변경하고, stderr를 concurrently drain하며, 허용된 fatal 오류를 감지하면 child process group을 종료하고 pre-exit dispatch_cli_error를 보고하도록 계약을 함께 수정하십시오. 프로세스가 fatal stderr를 출력한 뒤 계속 실행되는 테스트도 추가하십시오.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@skills/productivity/delegate/references/dispatch-guardrails.md` around lines
22 - 24, Replace the spawnSync-based execution around the delegate runner with
an asynchronous supervisor that drains stderr concurrently. Detect permitted
fatal provider errors before process exit, terminate the child process group
immediately, and report dispatch_cli_error with the matching stderr line and
reset time. Update the dispatch guardrails contract to define this pre-exit
error path, while preserving post-exit handling for other failures. Add a test
where the child emits a fatal stderr line and continues running, verifying early
termination and dispatch_cli_error reporting.
CodeRabbit: the new terminal-error rule force-kills the process, but the failure-code table still defined dispatch_cli_error as a nonzero exit, so the rule had no code it could legally report. Widen the definition. Also give the eval runner's opencode argv the same log flags — not to mirror the skill's dispatch, which executors.json deliberately does not do, but because the runner otherwise spends its full 30-minute deadline on a dead route and records no cause. It only captures stderr to a file, so the added lines are information, never a failure trigger. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Review disposition (dd86bcb): 1. 2. Move Two live runs confirm they are parsed as flags in that position, not absorbed into the 3. Early-detection path — the second half was a real defect, fixed; the first half declined. |
Verify #87 end-to-end via a real dispatch to opencode-go/glm-5.2 (quota-blocked until ~08-20). With --print-logs --log-level ERROR, the first 'Monthly usage limit reached. Resets in 13 days.' line appears ~11 s in; terminating on that line reports dispatch_cli_error at 13 s instead of spending the 30-minute deadline. Corrects the estimated 'about 36 seconds' surfacing figure to the measured 11-36 s range (two runs on 2026-08-06) and notes the rule is now exercised live, not just documented and installed.
Summary
An exhausted
opencode-goquota was being reported as a baredispatch_timeoutafter the full deadline. Measurement showed the error existed the whole time — the CLI just never printed it.Measured 2026-08-06 with real exhausted quotas (canary:
Reply with exactly: OK):429 … reset at <UTC>, then keeps runningTwo changes:
dispatch-guardrails.md— new Surface fatal errors early section. A fatal provider error does not imply the process exits, so a definitive error line on stderr is terminal: terminate at once and reportdispatch_cli_errorwith that line and any reset time, rather than waiting for exit or for the deadline. The canary is scoped down to routes that expose no error channel at all.cli-invocations.md— the OpenCode row gains--print-logs --log-level ERROR. It surfacesMonthly usage limit reached. Resets in 14 days.about 36 s in (after the CLI's internal 3 retries) instead of never.Net effect on the observed failure: cause known in ~36 s with its reset time, instead of an ambiguous timeout 10–30 min later.
Also worth recording: this invalidates the 2026-08-04 diagnosis of an "opencode-go provider outage" — a monthly workspace limit resetting in 14 days was already exhausted then, and the silent canaries across glm-5.2/glm-5.1/kimi were all that one limit. Exactly the misdiagnosis these changes prevent.
Test plan
npm testgreen on each commit.opencode run --auto --print-logs --log-level ERROR -m opencode-go/glm-5.2→AI_APICallError: Monthly usage limit reachedon stderr at 36 s; process still alive afterward, confirming force-termination is required.opencode/mimo-v2.5-free→ exit 0, stdout exactlyOK, stderr only the startup banner, so output extraction is unaffected.🤖 Generated with Claude Code
Summary by CodeRabbit
버그 수정
문서