Repository navigation
fix: surface real codex errors, seed a ChatGPT-auth model, show step failures inline - #8
Conversation
A live requirement-delivery run failed its review step with the opaque message "codex reported an error" while the backend had stated the exact cause (a model id the ChatGPT-account auth mode does not serve). Three mapper defects hid it: - codex-cli 0.145 emits error items with the text in `message`; the mapper read only `text`, so the detail was dropped. - First-error-wins let a non-fatal metadata warning item shadow the fatal turn.failed/error that followed, and any recorded error failed the step even when the CLI exited 0 after completing the turn. - `stderr ?? fallback` was dead code (stderr is always a string), so a silent CLI death produced an empty error message. Failure is now decided by fatal stream errors (turn.failed / top-level error, JSON blobs unwrapped to the inner sentence) and the exit code; item-level errors only refine the message and a completed turn with a lone warning succeeds (logged via ctx.logger, previously unused here). Fixture scenarios replay the live-captured 0.145 shapes; README and design docs document the error stream, which was previously unspecified.
Codex model ids are auth-mode-dependent: the seeded gpt-*-codex ids work with API keys but the backend rejects them with a 400 under a ChatGPT-account login (CODEX_HOME auth.json), leaving such deployments with no runnable reviewer model. gpt-5.6-sol is codex-cli 0.145's default model and was verified live under account auth; it lands as an openai strong-tier row (idempotent insert), priced like gpt-5.1-codex until a list price is confirmed. The seed comment and a new CONTRIBUTING recipe record the auth-mode gotcha so the next id gets verified before it ships.
step.failed only flipped the turn's icon to a red cross; the error text
surfaced solely in the run-level banner, and only after the whole run
had failed — a retried-then-recovered attempt's message was never shown
anywhere. Debugging a live codex failure meant querying run_events by
hand to learn what the executor had already reported.
Each attempt re-opens its own turn, so the step.failed payload's
{code, message} now stays attached to the attempt that produced it and
renders inline under the turn, styled like the run-level error banner.
No new i18n keys: the code is a stable identifier and the message is
executor prose.
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
|
Warning Review limit reached
Next review available in: 47 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (3)
📝 WalkthroughWalkthroughCodex event parsing now separates fatal failures from non-fatal item errors, improves executor fallback messages, adds tests and documentation, renders per-turn failure details in the run timeline, and seeds ChangesCodex updates
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant CodexCLI
participant CodexEventCollector
participant createCodexExecutor
participant RunTimeline
CodexCLI->>CodexEventCollector: Emit JSONL turn or error events
CodexEventCollector->>createCodexExecutor: Return fatal or item-level error state
createCodexExecutor->>RunTimeline: Emit step.failed or step.completed
RunTimeline->>RunTimeline: Render code and message for failed turn
Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CONTRIBUTING.md`:
- Line 56: Update the modelRows contribution instructions to list every required
field for a complete model row: provider, provider model id, displayName, tier,
contextWindow, inputCostPerMtok, and outputCostPerMtok. Preserve the existing
guidance to verify the provider model id against the live provider first.
In `@packages/db/src/seed/index.ts`:
- Around line 387-397: Update the GPT-5.6 Sol seed entry in
packages/db/src/seed/index.ts lines 387-397 to use contextWindow 1_050_000,
inputCostPerMtok "5.00", and outputCostPerMtok "30.00". In CHANGELOG.md line 19,
remove or revise the provisional “priced like gpt-5.1-codex” wording to reflect
the official catalog pricing.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 3588af7b-64ba-457b-97a6-ffab1a51ac49
📒 Files selected for processing (11)
CHANGELOG.mdCONTRIBUTING.mdapps/web/src/features/runs/RunTimeline.tsxdocs/design/03-executor-abstraction.mddocs/design/09-testing-and-ci.mdpackages/db/src/seed/index.tspackages/executor-codex/README.mdpackages/executor-codex/src/events.tspackages/executor-codex/src/executor.test.tspackages/executor-codex/src/executor.tspackages/executor-codex/test/fixtures/fake-codex.ts
Review finding (CodeRabbit, verified against developers.openai.com/api/docs/models/gpt-5.6-sol): the placeholder values copied from gpt-5.1-codex undersold the model 4x on input and 3x on output — the budget meter charges real dollars from these columns, so a run could overshoot its USD cap unnoticed. Now seeded with the documented 1,050,000-token context window and $5.00/$30.00 per MTok. The CONTRIBUTING recipe also lists every required model-row field so the next row starts complete.
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
Why
A live requirement-delivery run (real executors: Claude Code implementer via
CLAUDE_CODE_OAUTH_TOKEN, Codex reviewer via ChatGPT-accountCODEX_HOME) failed its review step with the opaque message "codex reported an error". The backend had stated the exact cause —The 'gpt-5.1-codex' model is not supported when using Codex with a ChatGPT account.— but three mapper defects swallowed it, the run timeline showed nothing but a red icon, and the seed contained no OpenAI model that works under ChatGPT-account auth at all. Diagnosing this required queryingrun_eventsby hand.What
One commit per concern:
fix(executor-codex)— surface real codex errors, stop failing on warningsmessage; the mapper read onlytext, dropping the detail.turn.failed/errorthat followed — and any recorded error failed the step even when the CLI exited 0 after completing the turn fine.stderr ?? fallbackwas dead code (stderr is always a string), so a silent CLI death produced an empty error message.ctx.logger, previously unused in this package). Fixture scenarios replay the live-captured 0.145 JSONL; the README and design docs now document the error stream, which was previously unspecified.feat(db)— seedgpt-5.6-solfor ChatGPT-account codex authgpt-*-codexids serve API keys but 400 under a ChatGPT-account login, leaving such deployments with no runnable reviewer model.gpt-5.6-sol(the CLI's default model) was verified live under account auth and lands as an openai strong-tier row, priced likegpt-5.1-codexuntil a list price is confirmed. New CONTRIBUTING recipe records the gotcha.feat(web)— show step failure detail on the run timelinestep.failedonly flipped the turn's icon; the error text surfaced solely in the run-level banner, and only after the whole run failed — a retried-then-recovered attempt's message was never shown anywhere. Each attempt's turn now renders its owncode: messageinline, styled like the run banner.Verification
bun run check,bun test(247 pass / 0 fail, integration suites ran against local Postgres — none skipped),bun run templates:validate,bun run build,bunx commitlint --from main --to HEAD— all green.Related
Summary by CodeRabbit
New Features
Bug Fixes