Skip to content

fix(cli): recover from provider context-limit errors by compacting - #14635

Merged
eshurakov merged 2 commits into
mainfrom
eshurakov/sturdy-gorge
Sep 28, 2026
Merged

eshurakov merged 2 commits into
mainfrom
eshurakov/sturdy-gorge

Conversation

@eshurakov

Copy link
Copy Markdown
Contributor

Summary

Classify a provider/gateway context-overflow wording the CLI currently misses, so the existing compact-and-retry recovery runs instead of the session dying at assistant_context_limit.

Why

Cloud Code Reviewer sessions on kilo-auto/free fail with terminal_reason = 'assistant_context_limit'. It is a real overflow, not a misclassification — the CLI log shows:

The request is 265246 tokens long and exceeds this model's context length of 262144 tokens.

The CLI's overflow check (packages/opencode/src/provider/error.ts:192) misses this error:

  • isContextOverflow(m) does not match. The nearest patterns need maximum context length (packages/llm/src/provider-error.ts:9,14); the observed text says exceeds this model's context length of <n> tokens.
  • Status is 400, not 413.
  • The provider code is under error.metadata.provider_code, so body?.error?.code === "context_length_exceeded" is false.

Because the error is never classified as context_overflow, needsCompaction is never set (packages/opencode/src/session/processor.ts) and prompt.ts never compacts and retries.

Change

One added pattern in packages/llm/src/provider-error.ts (wrapped in kilocode_change markers; shared upstream file):

/exceeds (?:this|the) model'?s (?:(?:maximum|max) )?context length(?: of [\d,]+ tokens?)?/i,

This fixes both affected paths, since both go through isContextOverflow:

  • AI SDK errors → parseAPICallError → isContextOverflow(m) (packages/opencode/src/provider/error.ts:192)
  • native route errors → isContextOverflow(body) (packages/llm/src/route/executor.ts:262)

No other behavior change: once matched, context_overflow → MessageV2.ContextOverflowError → processor halt → "compact" → compact-and-retry already works.

Tests

Extended packages/llm/test/provider-error.test.ts with the production wordings. Existing rate-limit exclusion test stays green.

  • bun test test/provider-error.test.ts (packages/llm): 2 pass
  • bun run typecheck (packages/llm): clean
  • bun test test/kilocode/provider/error.test.ts test/kilocode/session-overflow.test.ts (packages/opencode): 61 pass
  • bun run script/check-opencode-annotations.ts --worktree: all annotated

Notes

  • @opencode-ai/llm is private, so only a patch changeset for @kilocode/cli is included.
  • Deliberately not included: a secondary structural check on the gateway envelope (error_type / error.metadata.provider_code). The primary message fix covers the observed case, and the envelope shape is not represented in this repo. It can be added later if a captured response proves a case where the message text is absent.
  • Lowering the auto-compact threshold is not a fix here: KiloSessionOverflow.shouldCompact returns false for tool-result continuations, and the error must be classified as overflow first.

Follow-up (other repo, not here)

After this ships, the cloud pin KILOCODE_CLI_VERSION in services/cloud-agent-next/Dockerfile must be bumped to the release containing the fix, or the Code Reviewer keeps failing.

Classify provider/gateway errors whose message reads "<N> tokens long
and exceeds this model's context length of <M> tokens" as context
overflow, so the existing compact-and-retry recovery runs instead of
the session dying at assistant_context_limit.
@kilo-code-bot

kilo-code-bot Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Summary

Status: No Issues Found | Recommendation: Merge

Incremental review of 55bd657 found no new issues. Dropping the optional (?: of [\d,]+ tokens?)? group from the new overflow pattern (/exceeds (?:this|the) model'?s (?:(?:maximum|max) )?context length/i) is behavior-preserving, since that group was already optional, and it resolves the earlier thousands-separator question by no longer relying on the numeric format. The added 265,246/262,144 case still classifies as overflow, and the pattern stays kilocode_change-marked on a shared upstream file.

Files Reviewed (2 files)
  • packages/llm/src/provider-error.ts
  • packages/llm/test/provider-error.test.ts
Previous Review Summary (commit 0faccb0)

Current summary above is authoritative. Previous snapshots are kept for context only.

Previous review (commit 0faccb0)

Status: No Issues Found | Recommendation: Merge

Verified the new overflow pattern on packages/llm/src/provider-error.ts:34 against the production wordings: exceeds (?:this|the) model'?s (?:(?:maximum|max) )?context length(?: of [\d,]+ tokens?)? matches both quoted messages, stays excluded from the rate-limit/throttle guard, and is correctly marked with kilocode_change since it modifies a shared upstream file. The @kilocode/cli patch changeset is appropriate (@opencode-ai/llm is private).

Files Reviewed (3 files)
  • .changeset/context-overflow-recovery.md
  • packages/llm/src/provider-error.ts
  • packages/llm/test/provider-error.test.ts

Reviewed by deepseek-v4.1-flash · Input: 0 · Output: 0 · Cached: 0

Review guidance: REVIEW.md from base branch main

Comment thread packages/llm/src/provider-error.ts Outdated
/token limit exceeded/i,
// kilocode_change start - providers/gateways report over-long requests as
// "<N> tokens long and exceeds this model's context length of <M> tokens"
/exceeds (?:this|the) model'?s (?:(?:maximum|max) )?context length(?: of [\d,]+ tokens?)?/i,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's all good but I wonder if that translates well in terms of i18n ... the , for thousands is not universally common so that:

  • why do we have a thousand separator at all?
  • once we answer previous point, should we add all other thousand separators in there too?

not a blocker, rather a question, and once answered, I'll approve

@WebReflection WebReflection left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not the right RegExp

/token limit exceeded/i,
// kilocode_change start - providers/gateways report over-long requests as
// "<N> tokens long and exceeds this model's context length"
/exceeds (?:this|the) model'?s (?:(?:maximum|max) )?context length/i,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

there's something off in that model'?s expression:

  • is model OK? here is not without an s after: exceeds this model max context length is false
  • is model or models or model's meant? the RegExp is broken that way, it should be:
/exceeds (?:this|the) model(?:'s|s)? (?:(?:maximum|max) )?context length/i

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we're matching a specific error from the gateway and we don't need to support "model"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is a tiny change that prevents future gotchas worth? I let you decide, I'll approve if the constrain was that short-sighting

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants