Skip to content

guardrails: surface fatal provider errors instead of waiting out the deadline - #87

Merged
sungjunlee merged 3 commits into
mainfrom
guardrails-terminal-error
Aug 6, 2026
Merged

sungjunlee merged 3 commits into
mainfrom
guardrails-terminal-error

Conversation

@sungjunlee

@sungjunlee sungjunlee commented Aug 6, 2026 •

Copy link
Copy Markdown
Owner

Summary

An exhausted opencode-go quota was being reported as a bare dispatch_timeout after the full deadline. Measurement showed the error existed the whole time — the CLI just never printed it.

Measured 2026-08-06 with real exhausted quotas (canary: Reply with exactly: OK):

Route Behavior Cost to learn the cause
codex exits nonzero, reset time in stderr seconds
pi prints 429 … reset at <UTC>, then keeps running seconds to learn, never exits
opencode prints nothing at default log level, keeps running the entire deadline

Two changes:

  • dispatch-guardrails.md — new Surface fatal errors early section. A fatal provider error does not imply the process exits, so a definitive error line on stderr is terminal: terminate at once and report dispatch_cli_error with that line and any reset time, rather than waiting for exit or for the deadline. The canary is scoped down to routes that expose no error channel at all.
  • cli-invocations.md — the OpenCode row gains --print-logs --log-level ERROR. It surfaces Monthly usage limit reached. Resets in 14 days. about 36 s in (after the CLI's internal 3 retries) instead of never.

Net effect on the observed failure: cause known in ~36 s with its reset time, instead of an ambiguous timeout 10–30 min later.

Also worth recording: this invalidates the 2026-08-04 diagnosis of an "opencode-go provider outage" — a monthly workspace limit resetting in 14 days was already exhausted then, and the silent canaries across glm-5.2/glm-5.1/kimi were all that one limit. Exactly the misdiagnosis these changes prevent.

Test plan

  • npm test green on each commit.
  • Failing path: opencode run --auto --print-logs --log-level ERROR -m opencode-go/glm-5.2 → AI_APICallError: Monthly usage limit reached on stderr at 36 s; process still alive afterward, confirming force-termination is required.
  • Success path: same flags on opencode/mimo-v2.5-free → exit 0, stdout exactly OK, stderr only the startup banner, so output extraction is unaffected.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • 버그 수정

    • 치명적인 제공자 오류가 조용히 종료되어 시간 초과로만 표시되던 문제를 개선했습니다.
    • 오류를 더 빠르게 감지하고 터미널에 명확하게 표시합니다.
    • 오류 발생 시 관련 실행 경로의 실패로 정확히 보고합니다.
    • 오류 채널이 제한된 실행 경로에서도 최대 90초 내 상태를 확인해 경로 문제를 구분합니다.
  • 문서

    • 오류 처리 방식과 로그 확인 방법을 안내합니다.
    • 상세 오류 로그를 선택적으로 확인할 수 있으며, 일반 출력에는 영향을 주지 않습니다.

sungjunlee and others added 2 commits August 6, 2026 18:08
A CLI can report a fatal provider error and keep running, so waiting for
exit turns a known failure into a spent deadline. Measured 2026-08-06
with exhausted quotas: codex exits nonzero in seconds, pi prints its 429
with a reset time and keeps running, opencode prints nothing at its
default log level and keeps running — that last shape is why an
exhausted quota was reported as a bare dispatch_timeout.

Report dispatch_cli_error from the error line itself, and prefer a route
flag that surfaces the error over discovering it by timeout. The canary
is now scoped to routes that expose no error channel at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
opencode run says nothing about a fatal provider error at its default
log level, which is what made an exhausted quota look like a healthy
silent run. --print-logs --log-level ERROR surfaces it on stderr and
leaves stdout clean, verified on both a failing and a succeeding run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@coderabbitai

coderabbitai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: fdb575c9-d0c8-4da6-9666-dccac4766e11

📥 Commits

Reviewing files that changed from the base of the PR and between f326a00 and dd86bcb.

📒 Files selected for processing (2)
  • evals/delegate/executors.json
  • skills/productivity/delegate/references/dispatch-guardrails.md
🚧 Files skipped from review as they are similar to previous changes (1)
  • skills/productivity/delegate/references/dispatch-guardrails.md

Walkthrough

OpenCode 실행에 오류 로그 옵션을 추가했습니다. 확정적 provider 오류를 조기에 감지하고 dispatch_cli_error로 보고하는 규칙을 정의했습니다. 오류 채널이 없는 route에는 최대 90초의 OK canary 규칙을 추가했습니다.

Changes

Provider 오류 처리

Layer / File(s) Summary
OpenCode 로그 출력 설정
skills/productivity/delegate/references/cli-invocations.md, evals/delegate/executors.json
OpenCode 실행에 --print-logs와 --log-level ERROR를 추가했습니다. 해당 옵션이 stdout을 변경하지 않는 동작을 설명했습니다.
Provider 오류 조기 보고
skills/productivity/delegate/references/dispatch-guardrails.md
확정적 stderr provider 오류를 즉시 종료하고 dispatch_cli_error로 보고하도록 정의했습니다. 오류 노출용 route flag와 최대 90초 OK canary 규칙을 추가했습니다.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related PRs

  • sungjunlee/skills#7: 동일한 provider dispatch 지침과 OpenCode 실행 문서를 확장한 변경입니다.
  • sungjunlee/skills#81: cli-invocations.md와 executors.json을 함께 수정했지만, 모델 라우팅을 다룹니다.

Poem

토끼가 로그를 켜고 깡충,
오류는 늦기 전에 콕.
조용한 route엔 canary 한 번,
stdout은 깨끗이 그대로.
dispatch_cli_error로 알려요,
당근처럼 빠르게요!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed 제목은 치명적 provider 오류를 deadline 대기 없이 즉시 노출하는 주요 변경 사항을 정확하고 간결하게 설명합니다.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch guardrails-terminal-error

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@skills/productivity/delegate/references/cli-invocations.md`:
- Line 31: Update the OpenCode entries in executors.json so their dispatch argv
includes the required --print-logs and --log-level ERROR flags, matching the
invocation documented in the CLI reference. Add or extend evaluation coverage
for these executors to verify fatal errors are detected from stderr while normal
results are extracted from stdout.
- Line 76: OpenCode 호출에서 global 로그 플래그의 위치가 잘못되었습니다. OpenCode 라우트의 실행 명령을 수정해
print-logs 및 log-level ERROR를 run 서브커맨드보다 앞에 배치하고, auto 플래그와 나머지 인자는 run 뒤에
유지하십시오.

In `@skills/productivity/delegate/references/dispatch-guardrails.md`:
- Around line 22-24: Replace the spawnSync-based execution around the delegate
runner with an asynchronous supervisor that drains stderr concurrently. Detect
permitted fatal provider errors before process exit, terminate the child process
group immediately, and report dispatch_cli_error with the matching stderr line
and reset time. Update the dispatch guardrails contract to define this pre-exit
error path, while preserving post-exit handling for other failures. Add a test
where the child emits a fatal stderr line and continues running, verifying early
termination and dispatch_cli_error reporting.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: a958188b-2120-4189-8380-68b1d23890ee

📥 Commits

Reviewing files that changed from the base of the PR and between 30b6dbf and f326a00.

📒 Files selected for processing (2)
  • skills/productivity/delegate/references/cli-invocations.md
  • skills/productivity/delegate/references/dispatch-guardrails.md

Comment thread skills/productivity/delegate/references/cli-invocations.md
The prompt is already in argv, so connect stdin to DEVNULL for every current route. In a process API, set the child stdin to DEVNULL. Under rule 2, redirect with `< /dev/null` and never pass the redirect as an argv token. Add a future stdin-consuming route as an explicit exception instead of inheriting an open pipe.

For the OpenCode routes, `--auto` has the same trust implications as other non-interactive permission bypass flags.
For the OpenCode routes, `--auto` has the same trust implications as other non-interactive permission bypass flags. The two log flags are not optional either: without them a fatal provider error prints nothing, the run stays silent as if healthy, and the failure only surfaces as a spent deadline. They leave stdout clean, so extraction is unaffected.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

command -v opencode
opencode --version
opencode run --help | rg -n -- '--print-logs|--log-level'

Repository: sungjunlee/skills

Length of output: 155


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

echo "== repo files around target =="
git ls-files | rg '(^skills/productivity/delegate/references/cli-invocations\.md|opencode|package\.json|README\.md)$' || true

echo "== target excerpt =="
sed -n '55,90p' skills/productivity/delegate/references/cli-invocations.md || true

echo "== opencode references =="
rg -n "opencode|print-logs|stdout|stderr|fatal provider" -S . || true

Repository: sungjunlee/skills

Length of output: 26675


🌐 Web query:

opencode 1.17.18 run --print-logs --log-level ERROR source

💡 Result:

The command opencode 1.17.18 run --print-logs --log-level ERROR source is syntactically malformed based on the official OpenCode CLI documentation [1][2]. In the OpenCode CLI, the run command follows the pattern opencode run [message] [flags] [1][2]. The positional argument [message] is the task or prompt you want the AI agent to execute [1][3]. The string "source" at the end of your command would be interpreted by the CLI as a positional argument (a message to the AI), rather than as a flag or configuration parameter [3][2]. To correctly use the provided flags, the command should be structured as follows: opencode --print-logs --log-level ERROR run "your prompt or task here" Explanation of the components: 1. opencode: The main CLI executable [4][5]. 2. --print-logs: A global flag that forces logs to be output directly to stderr [6][2]. 3. --log-level ERROR: A global flag that sets the verbosity of the logs to only report errors [6][2]. 4. run: The command used to execute a prompt non-interactively [1][2]. 5. "your prompt or task here": The required positional argument (message) for the run command [1][2]. If "source" was intended to be the task/message for the AI, the correct command would be: opencode --print-logs --log-level ERROR run "source" Note that the --log-level and --print-logs flags are global options that must precede the run subcommand in the command structure [6][2]. Using them after the run command or after the message may result in them being ignored or treated as part of the prompt [3][2].

Citations:


OpenCode의 --print-logs를 run 호출 전으로 이동하십시오.

OpenCode 문서에 따르면 --print-logs와 --log-level ERROR는 global flag인데 현재 reference에서는 opencode run --auto --print-logs --log-level ERROR ...로 전달합니다. OpenCode 1.17.18 이상에서 이 인자는 입력 메시지로 처리되어 정상 호출이 깨집니다. opencode --print-logs --log-level ERROR run --auto ...로 배치하십시오.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@skills/productivity/delegate/references/cli-invocations.md` at line 76,
OpenCode 호출에서 global 로그 플래그의 위치가 잘못되었습니다. OpenCode 라우트의 실행 명령을 수정해 print-logs 및
log-level ERROR를 run 서브커맨드보다 앞에 배치하고, auto 플래그와 나머지 인자는 run 뒤에 유지하십시오.

Comment on lines +22 to +24
- Treat a definitive provider error on stderr as terminal. Terminate at once and report `dispatch_cli_error` with that line and any reset time it names, rather than waiting for the process to exit or for the deadline.
- Prefer a route flag that surfaces such errors over discovering them by timeout. `opencode run` needs `--print-logs --log-level ERROR`, which surfaces quota exhaustion about 36 seconds in, after the CLI's internal retries.
- Send a bounded canary (≤90 s, `Reply with exactly: OK`) only for a route that stays silent and exposes no error channel. A silent canary condemns the route, not the model.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift

조기 오류 감지 경로를 실행기에 구현해야 합니다.

skills/productivity/delegate/references/dispatch-guardrails.md Lines 22-24는 프로세스 종료 전에 stderr를 읽고 dispatch_cli_error를 보고하도록 정의합니다. 그러나 제공된 실행 경로인 scripts/run-delegate-eval.mjs Lines 181-203은 spawnSync를 사용합니다. 이 코드는 프로세스가 종료되거나 deadline에 도달한 뒤에만 result.stderr를 확인하므로 즉시 종료와 조기 보고를 수행할 수 없습니다.

또한 현재 실행기는 nonzero 종료를 일반 note로만 반환하고, dispatch-guardrails.md Line 43은 dispatch_cli_error를 종료 후 오류로만 정의합니다. 비동기 supervisor로 변경하고, stderr를 concurrently drain하며, 허용된 fatal 오류를 감지하면 child process group을 종료하고 pre-exit dispatch_cli_error를 보고하도록 계약을 함께 수정하십시오. 프로세스가 fatal stderr를 출력한 뒤 계속 실행되는 테스트도 추가하십시오.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@skills/productivity/delegate/references/dispatch-guardrails.md` around lines
22 - 24, Replace the spawnSync-based execution around the delegate runner with
an asynchronous supervisor that drains stderr concurrently. Detect permitted
fatal provider errors before process exit, terminate the child process group
immediately, and report dispatch_cli_error with the matching stderr line and
reset time. Update the dispatch guardrails contract to define this pre-exit
error path, while preserving post-exit handling for other failures. Add a test
where the child emits a fatal stderr line and continues running, verifying early
termination and dispatch_cli_error reporting.

CodeRabbit: the new terminal-error rule force-kills the process, but the
failure-code table still defined dispatch_cli_error as a nonzero exit,
so the rule had no code it could legally report. Widen the definition.

Also give the eval runner's opencode argv the same log flags — not to
mirror the skill's dispatch, which executors.json deliberately does not
do, but because the runner otherwise spends its full 30-minute deadline
on a dead route and records no cause. It only captures stderr to a file,
so the added lines are information, never a failure trigger.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sungjunlee

Copy link
Copy Markdown
Owner Author

Review disposition (dd86bcb):

1. executors.json missing the flags — applied (argv), declined (new evals).
The argv now carries --print-logs --log-level ERROR. Framing differs from the comment though: executors.json deliberately does not replay the skill's dispatch (see the eval README — it holds one harness steady across observation dates). The reason to change it is independent: without the flags the runner spends its full 30-minute deadline on a dead route and records no cause. Safe because the runner only writes stderr to a file for manual assessment, so extra lines are information, never a failure trigger. The "add evaluation coverage" half is declined — that eval infra is deliberately deferred, and this PR is not the place to start it.

2. Move --print-logs before run — declined, the claim is incorrect.
opencode run --help on the installed CLI lists both under the run subcommand's own Options:

--print-logs   print logs to stderr                        [boolean]
--log-level    log level   [string] [choices: "DEBUG","WARN","ERROR",...]

Two live runs confirm they are parsed as flags in that position, not absorbed into the [message..] positional: the same command without them is silent, with them it prints AI_APICallError: Monthly usage limit reached at ~36 s; and on a succeeding run stdout was exactly OK with the prompt intact.

3. Early-detection path — the second half was a real defect, fixed; the first half declined.
Correct catch that dispatch_cli_error was defined as "the process exited nonzero", which the new rule can never satisfy since it force-kills. The definition now reads "exited nonzero, or was terminated on a definitive provider error, and it was not a timeout". Rewriting run-delegate-eval.mjs into an async supervisor is declined: dispatch-guardrails.md governs the skill's dispatch, which is supervised by the agent, not by that script — the eval runner is a separate, deliberately simpler transport.

@sungjunlee
sungjunlee merged commit 294f338 into main Aug 6, 2026
2 checks passed
@sungjunlee
sungjunlee deleted the guardrails-terminal-error branch August 6, 2026 10:45
sungjunlee added a commit that referenced this pull request Aug 7, 2026
Verify #87 end-to-end via a real dispatch to opencode-go/glm-5.2 (quota-blocked
until ~08-20). With --print-logs --log-level ERROR, the first 'Monthly usage
limit reached. Resets in 13 days.' line appears ~11 s in; terminating on that
line reports dispatch_cli_error at 13 s instead of spending the 30-minute
deadline. Corrects the estimated 'about 36 seconds' surfacing figure to the
measured 11-36 s range (two runs on 2026-08-06) and notes the rule is now
exercised live, not just documented and installed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant