Skip to content

fix: surface real codex errors, seed a ChatGPT-auth model, show step failures inline - #8

Merged
hutusi merged 4 commits into
mainfrom
fix/codex-error-surfacing
Jul 27, 2026
Merged

hutusi merged 4 commits into
mainfrom
fix/codex-error-surfacing

Conversation

@hutusi

@hutusi hutusi commented Jul 27, 2026 •

Copy link
Copy Markdown
Contributor

Why

A live requirement-delivery run (real executors: Claude Code implementer via CLAUDE_CODE_OAUTH_TOKEN, Codex reviewer via ChatGPT-account CODEX_HOME) failed its review step with the opaque message "codex reported an error". The backend had stated the exact cause — The 'gpt-5.1-codex' model is not supported when using Codex with a ChatGPT account. — but three mapper defects swallowed it, the run timeline showed nothing but a red icon, and the seed contained no OpenAI model that works under ChatGPT-account auth at all. Diagnosing this required querying run_events by hand.

What

One commit per concern:

  1. fix(executor-codex) — surface real codex errors, stop failing on warnings
    • codex-cli 0.145 emits error items with the text in message; the mapper read only text, dropping the detail.
    • First-error-wins let a non-fatal metadata warning ("Model metadata … Defaulting to fallback") shadow the fatal turn.failed/error that followed — and any recorded error failed the step even when the CLI exited 0 after completing the turn fine.
    • stderr ?? fallback was dead code (stderr is always a string), so a silent CLI death produced an empty error message.
    • Failure is now decided by fatal stream errors (JSON blobs unwrapped to the inner human sentence) plus the exit code; item errors only refine the message; a completed turn with a lone warning succeeds (logged via ctx.logger, previously unused in this package). Fixture scenarios replay the live-captured 0.145 JSONL; the README and design docs now document the error stream, which was previously unspecified.
  2. feat(db) — seed gpt-5.6-sol for ChatGPT-account codex auth
    • Codex model ids are auth-mode-dependent: the gpt-*-codex ids serve API keys but 400 under a ChatGPT-account login, leaving such deployments with no runnable reviewer model. gpt-5.6-sol (the CLI's default model) was verified live under account auth and lands as an openai strong-tier row, priced like gpt-5.1-codex until a list price is confirmed. New CONTRIBUTING recipe records the gotcha.
  3. feat(web) — show step failure detail on the run timeline
    • step.failed only flipped the turn's icon; the error text surfaced solely in the run-level banner, and only after the whole run failed — a retried-then-recovered attempt's message was never shown anywhere. Each attempt's turn now renders its own code: message inline, styled like the run banner.

Verification

  • bun run check, bun test (247 pass / 0 fail, integration suites ran against local Postgres — none skipped), bun run templates:validate, bun run build, bunx commitlint --from main --to HEAD — all green.
  • 3 new executor tests replay the exact JSONL captured from the failing run (backend-rejection surfaces its message, warning-only turn succeeds, silent death gets a non-empty message).

Related

Summary by CodeRabbit

  • New Features

    • Added GPT-5.6 Sol to the built-in OpenAI model catalog.
    • Run timelines now show the exact failure location, error code, and message inline.
  • Bug Fixes

    • Improved Codex failure reporting and handling of non-fatal warnings.
    • Corrected error message extraction from newer Codex responses.
    • Added useful fallback messages when Codex exits without output.

hutusi added 3 commits July 27, 2026 18:54
A live requirement-delivery run failed its review step with the opaque
message "codex reported an error" while the backend had stated the exact
cause (a model id the ChatGPT-account auth mode does not serve). Three
mapper defects hid it:

- codex-cli 0.145 emits error items with the text in `message`; the
  mapper read only `text`, so the detail was dropped.
- First-error-wins let a non-fatal metadata warning item shadow the
  fatal turn.failed/error that followed, and any recorded error failed
  the step even when the CLI exited 0 after completing the turn.
- `stderr ?? fallback` was dead code (stderr is always a string), so a
  silent CLI death produced an empty error message.

Failure is now decided by fatal stream errors (turn.failed / top-level
error, JSON blobs unwrapped to the inner sentence) and the exit code;
item-level errors only refine the message and a completed turn with a
lone warning succeeds (logged via ctx.logger, previously unused here).
Fixture scenarios replay the live-captured 0.145 shapes; README and
design docs document the error stream, which was previously unspecified.
Codex model ids are auth-mode-dependent: the seeded gpt-*-codex ids work
with API keys but the backend rejects them with a 400 under a
ChatGPT-account login (CODEX_HOME auth.json), leaving such deployments
with no runnable reviewer model. gpt-5.6-sol is codex-cli 0.145's
default model and was verified live under account auth; it lands as an
openai strong-tier row (idempotent insert), priced like gpt-5.1-codex
until a list price is confirmed. The seed comment and a new CONTRIBUTING
recipe record the auth-mode gotcha so the next id gets verified before
it ships.
step.failed only flipped the turn's icon to a red cross; the error text
surfaced solely in the run-level banner, and only after the whole run
had failed — a retried-then-recovered attempt's message was never shown
anywhere. Debugging a live codex failure meant querying run_events by
hand to learn what the executor had already reported.

Each attempt re-opens its own turn, so the step.failed payload's
{code, message} now stays attached to the attempt that produced it and
renders inline under the turn, styled like the run-level error banner.
No new i18n keys: the code is a stable identifier and the message is
executor prose.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@coderabbitai

coderabbitai Bot commented Jul 27, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@hutusi, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 47 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 60bf335f-0bfb-4da7-879c-1816810ebe78

📥 Commits

Reviewing files that changed from the base of the PR and between 30d2f36 and c3ce28e.

📒 Files selected for processing (3)
  • CHANGELOG.md
  • CONTRIBUTING.md
  • packages/db/src/seed/index.ts
📝 Walkthrough

Walkthrough

Codex event parsing now separates fatal failures from non-fatal item errors, improves executor fallback messages, adds tests and documentation, renders per-turn failure details in the run timeline, and seeds gpt-5.6-sol in the built-in model catalog.

Changes

Codex updates

Layer / File(s) Summary
Built-in model registration
packages/db/src/seed/index.ts, CONTRIBUTING.md, CHANGELOG.md
Adds gpt-5.6-sol to the idempotent model seed and documents the model-registration workflow and auth-mode behavior.
Codex error collection and execution
packages/executor-codex/src/events.ts, packages/executor-codex/src/executor.ts, packages/executor-codex/test/fixtures/fake-codex.ts, packages/executor-codex/src/executor.test.ts, packages/executor-codex/README.md, docs/design/*, CHANGELOG.md
Separates fatal and item-level errors, unwraps JSON error messages, preserves successful turns after warnings, improves silent-exit fallbacks, and adds scenario coverage and documentation.
Per-turn failure rendering
apps/web/src/features/runs/RunTimeline.tsx
Stores structured step.failed details and renders the failure code and message within the affected turn.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant CodexCLI
  participant CodexEventCollector
  participant createCodexExecutor
  participant RunTimeline
  CodexCLI->>CodexEventCollector: Emit JSONL turn or error events
  CodexEventCollector->>createCodexExecutor: Return fatal or item-level error state
  createCodexExecutor->>RunTimeline: Emit step.failed or step.completed
  RunTimeline->>RunTimeline: Render code and message for failed turn
Loading

Possibly related PRs

  • ainaive/agrippa#5: Introduced the Codex executor adapter and JSONL event mapping extended by these failure-handling changes.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is concise and accurately summarizes the three main changes: Codex error handling, seeding a ChatGPT-auth model, and inline run failure display.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/codex-error-surfacing

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@CONTRIBUTING.md`:
- Line 56: Update the modelRows contribution instructions to list every required
field for a complete model row: provider, provider model id, displayName, tier,
contextWindow, inputCostPerMtok, and outputCostPerMtok. Preserve the existing
guidance to verify the provider model id against the live provider first.

In `@packages/db/src/seed/index.ts`:
- Around line 387-397: Update the GPT-5.6 Sol seed entry in
packages/db/src/seed/index.ts lines 387-397 to use contextWindow 1_050_000,
inputCostPerMtok "5.00", and outputCostPerMtok "30.00". In CHANGELOG.md line 19,
remove or revise the provisional “priced like gpt-5.1-codex” wording to reflect
the official catalog pricing.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3588af7b-64ba-457b-97a6-ffab1a51ac49

📥 Commits

Reviewing files that changed from the base of the PR and between 79f52f0 and 30d2f36.

📒 Files selected for processing (11)
  • CHANGELOG.md
  • CONTRIBUTING.md
  • apps/web/src/features/runs/RunTimeline.tsx
  • docs/design/03-executor-abstraction.md
  • docs/design/09-testing-and-ci.md
  • packages/db/src/seed/index.ts
  • packages/executor-codex/README.md
  • packages/executor-codex/src/events.ts
  • packages/executor-codex/src/executor.test.ts
  • packages/executor-codex/src/executor.ts
  • packages/executor-codex/test/fixtures/fake-codex.ts

Comment thread CONTRIBUTING.md Outdated
Comment thread packages/db/src/seed/index.ts Outdated
Review finding (CodeRabbit, verified against
developers.openai.com/api/docs/models/gpt-5.6-sol): the placeholder
values copied from gpt-5.1-codex undersold the model 4x on input and
3x on output — the budget meter charges real dollars from these
columns, so a run could overshoot its USD cap unnoticed. Now seeded
with the documented 1,050,000-token context window and $5.00/$30.00
per MTok. The CONTRIBUTING recipe also lists every required model-row
field so the next row starts complete.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@hutusi
hutusi merged commit 8ac1a10 into main Jul 27, 2026
5 checks passed
@hutusi
hutusi deleted the fix/codex-error-surfacing branch July 27, 2026 11:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant