Skip to content

fix: increase timeout and add error logging for context summarization - #1363

Closed
oyi77 wants to merge 59 commits into
diegosouzapw:mainfrom
oyi77:fix/context-summarization-timeout
Closed

oyi77 wants to merge 59 commits into
diegosouzapw:mainfrom
oyi77:fix/context-summarization-timeout

Conversation

@oyi77

@oyi77 oyi77 commented Apr 17, 2026

Copy link
Copy Markdown
Contributor

Summary

Fixes OpenCode compaction 500 errors by increasing timeout and adding detailed error logging for context summarization.

Problem

OpenCode /compact command was returning 500 Internal Server Error during session compaction due to:

  1. Insufficient timeout: 120s was too short for large context summarization (MEMORY.md ~35KB)
  2. Silent error handling: Errors were swallowed without logging, making debugging impossible

Solution

Changes to open-sse/services/contextHandoff.ts

  1. Increased timeout: 120s → 300s (5 minutes)
  2. Added Promise.race timeout handling with clear error messages
  3. Added detailed error logging (status + error message + session ID)
  4. Added JSON parse failure logging

Testing

✅ Tested on dev server - /api/v1/responses/compact endpoint working correctly

  • Returns 200 OK with proper SSE streaming
  • No 500 errors during testing
  • Response generated in 6.5s

Impact

Before:

  • ❌ Timeout: 120s (too short)
  • ❌ Errors: Silent (no logs)
  • ❌ Debugging: Impossible

After:

  • ✅ Timeout: 300s (enough time for large context)
  • ✅ Errors: Detailed logging
  • ✅ Debugging: Full visibility

Monitoring

Watch for these log patterns:

  • [context-handoff] Summarization timeout for session - timeout occurred
  • [context-handoff] Summarization failed for session - provider error
  • [context-handoff] Failed to parse handoff JSON - parse failure

diegosouzapw and others added 30 commits April 14, 2026 07:46
…ent literal \n\n artifacts

- Changed regex quantifier from ? to * in combo.ts, comboAgentMiddleware.ts,
  and contextHandoff.ts to greedily strip all JSON-escaped newline sequences
  surrounding <omniModel> tags in SSE streaming chunks
- Added \r to the character class for cross-platform robustness
- Fixed Playwright strict-mode violation in combo-unification.spec.ts
- Bumped OpenAPI version and CHANGELOG to 3.6.6
…w#1187/diegosouzapw#1218, diegosouzapw#1202)

- fix(gemini): strip VS Code JSON Schema extensions from tool schemas (diegosouzapw#1175)
  Add enumDescriptions, markdownDescription, markdownEnumDescriptions,
  enumItemLabels and tags to UNSUPPORTED_SCHEMA_CONSTRAINTS so the Gemini
  sanitizer removes them before forwarding. GitHub Copilot injects these
  non-standard fields into tool definitions, causing Gemini to reject with
  'Unknown name enumDescriptions at functionDeclarations[n].parameters'.

- fix(health-check): unwrap proxy config object before passing to getAccessToken (diegosouzapw#1187 diegosouzapw#1218)
  resolveProxyForConnection() returns { proxy, level, levelId } but the health
  check loop was passing the full wrapper to getAccessToken(), which expects the
  inner config object (.host, .port etc). The proxy dispatcher validated .host
  on the wrapper (undefined) and threw 'Context proxy host is required', silently
  marking every connection as unhealthy every sweep. Fix mirrors the pattern
  already used in chatHelpers.ts: proxyResult?.proxy || null.

- fix(ui): debounce models.dev sync interval slider to save only on release (diegosouzapw#1202)
  The slider's onChange fired updateInterval() on every drag tick, sending a
  PATCH per pixel of movement. Rapid API responses overwrote UI state mid-drag.
  Introduce draftIntervalHours for smooth visual feedback; the PATCH fires
  on onMouseUp / onBlur once the user releases the control.
diegosouzapw#1220, diegosouzapw#1231)

- fix(core): diegosouzapw#1206 inject startup guard against app/ and src/app/ conflict
- fix(health): diegosouzapw#1220 add HEALTHCHECK_STAGGER_MS to prevent token refresh bursting
- fix(proxy): diegosouzapw#1231 prioritize HTTP 429 over quota body heuristics
- fix(sse): diegosouzapw#1211 strip leading double-newlines in responses API stream
Update package versions for the electron app and open-sse package.
Sync llm.txt metadata and feature headings with the 3.6.6 release.
Add guarded outbound fetch helpers with private/local URL blocking,
controlled retries, timeout normalization, and route-level status
propagation for provider validation and model discovery.

Introduce cooldown-aware chat retries with configurable
requestRetry and maxRetryIntervalSec settings, model-scoped cooldown
responses, and improved rate-limit learning from headers and error
bodies so short upstream lockouts can recover automatically.

Also align Antigravity and Codex header handling, require API keys
for Pollinations, validate web runtime env at startup, restore
sanitized Gemini tool names in translated responses, and inject a
synthetic Claude text block when upstream SSE completes empty.
Introduce GLM Thinking as a first-class provider preset with shared GLM
model metadata, pricing, usage sync, dashboard support, and provider
request defaults for higher token budgets and longer timeouts.

Use provider-side /messages/count_tokens when a Claude-compatible
upstream supports it, while preserving estimated fallback behavior for
missing models, missing credentials, and upstream failures.

Also add startup seeding for default model aliases and normalize common
cross-proxy model dialects so canonical slashful model ids do not get
misrouted during resolution.
Add dedicated sync token storage, issuance, revocation, and bundle
download routes backed by stable config bundle versioning and ETag
support.

Expose the v1 websocket handshake route and custom Next server bridge so
OpenAI-compatible websocket traffic can be upgraded and proxied through
the dashboard and API bridge.

Expand compliance auditing with structured metadata, pagination, request
context, auth and provider credential events, and SSRF-blocked
validation logging.
- CHANGELOG: Add WebSocket bridge, GLM Thinking preset, safe outbound
  fetch/SSRF guard, cooldown-aware retries, compliance audit v2, model
  alias seeding, and all Internal Improvements for the 3 new commits
- README: Expand v3.6.x highlights table with 10 new features; add
  SafeOutboundFetch, CooldownAwareRetry, SSRF guard, TPS metric, sync
  tokens, WebSocket bridge to Resilience/Observability/Deployment tables
- ARCHITECTURE: Bump date; add new modules to executive summary, API
  routes, SSE core services, Auth/Security section; add SSRF/Outbound
  guard failure mode (section 6); expand module mapping
- ENVIRONMENT: Add OMNIROUTE_CRYPT_KEY/OMNIROUTE_API_KEY_BASE64 legacy
  aliases, OUTBOUND_SSRF_GUARD_ENABLED, CODEX_CLIENT_VERSION, and
  REQUEST_RETRY/MAX_RETRY_INTERVAL_SEC cooldown retry settings
- FEATURES: Add 6 new feature sections — V1 WebSocket Bridge, Sync
  Tokens & Config Bundle, GLM Thinking Preset, Safe Outbound Fetch &
  SSRF Guard, Cooldown-Aware Retries, Compliance Audit v2
Integrated into release/v3.6.6 — IPv6 proxy test fix
…osouzapw#1200 (diegosouzapw#1256)

Integrated into release/v3.6.6 — Gemini custom model picker fix
…ests (diegosouzapw#1246)

Integrated into release/v3.6.6 — OAuth client_id default fallbacks
…n Chat→Responses translator (diegosouzapw#1245)

Integrated into release/v3.6.6 — max_tokens → max_output_tokens Responses API translation + unit tests
…egosouzapw#1258)

Integrated into release/v3.6.6 — cursor-agent CLI credential source support
…eout behavior (diegosouzapw#1257)

Integrated into release/v3.6.6 — CC-compatible upstream SSE restore + stream timeout fix + README table repair
This integrates rdself's unify-provider-profile-locks PR manually to handle structural conflicts.
RaviTharuma and others added 16 commits April 15, 2026 17:16
…del button positioning, and clarify oauth documentation
Document the latest fixes for Codex routing configuration parsing and
Lobehub provider icon fallback behavior.

Add the note that the remaining JavaScript test files were migrated to
TypeScript ES modules to reflect the completed test stack transition.
Proactively compress oversized contexts before sending to upstream providers,
preventing context_length_exceeded errors. Compression triggers at 85% of
model's context limit using the existing 3-layer compressContext() function.

- Import compressContext, estimateTokens, getTokenLimit from contextManager
- Add compression check after translation, before executor dispatch
- Estimate tokens and compare against 85% threshold of model's context limit
- Apply 3-layer compression (trim tools, compress thinking, purify history)
- Log compression events with before/after token counts and layers applied
- Audit compression events for observability
- Add unit tests verifying integration behavior

Closes diegosouzapw#1290
When purifyHistory() drops oldest messages to fit context window, it can
split tool_use/tool_result pairs — keeping the tool_result but dropping
the tool_use that initiated it. This causes upstream providers to reject
the request with format errors.

Add fixToolPairs() that runs after each purification pass to remove:
- OpenAI format: orphaned role='tool' messages without matching tool_calls ID
- Claude format: orphaned tool_result content blocks without matching tool_use ID

Closes diegosouzapw#1291
- Increase SUMMARIZATION_TIMEOUT_MS from 120s to 300s (5 minutes)
- Add Promise.race timeout handling with clear error messages
- Add detailed error logging (status + error message + session ID)
- Add JSON parse failure logging

Fixes OpenCode compaction 500 errors caused by:
1. Insufficient timeout for large context summarization
2. Silent error handling preventing debugging

Tested on dev server - /responses/compact endpoint working correctly.
@oyi77
oyi77 requested a review from diegosouzapw as a code owner April 17, 2026 12:14

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request upgrades OmniRoute to version 3.6.6, introducing a Perplexity Web provider, a WebSocket bridge, and a configuration synchronization system. It hardens security with an outbound SSRF guard and runtime environment validation, while also enhancing the memory and skills systems. Feedback identifies several critical issues: missing crypto imports in the Cursor and Perplexity executors, a logic error in the memory FTS5 implementation that breaks indexing for new data, and a security vulnerability where encryption failures fall back to plaintext. Additionally, the Perplexity executor contains dangerous JSON string truncation, and the context handoff service has a potential resource leak from an uncleared timeout.

// (cursor-agent imports don't provide a machineId)
const machineId =
credentials.providerSpecificData?.machineId ||
crypto.createHash("sha256").update(accessToken).digest("hex");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The crypto module is used here to call createHash, but it is not imported in this file. This will cause a ReferenceError at runtime. Please add import crypto from "node:crypto"; at the top of the file.

* completions format and Perplexity's internal protocol.
*/

import { BaseExecutor, type ExecuteInput } from "./base.ts";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The crypto module is required for randomUUID() calls later in this file. Please add the import.

Suggested change
import { BaseExecutor, type ExecuteInput } from "./base.ts";
import crypto from "node:crypto";
import { BaseExecutor, type ExecuteInput } from "./base.ts";


-- Step 6: Recreate triggers using memory_id (INTEGER rowid) instead of id (UUID TEXT)
CREATE TRIGGER IF NOT EXISTS memory_fts_ai AFTER INSERT ON memories BEGIN
INSERT INTO memory_fts(rowid, content, key) VALUES (new.memory_id, new.content, new.key);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The memory_id column is not auto-populated during INSERT operations in src/lib/memory/store.ts. Consequently, new memories will have a NULL value for memory_id. This causes the memory_fts_ai trigger to fail (or sync incorrectly) and breaks the JOIN in the retrieval logic. A more robust solution is to use the built-in SQLite rowid in both the triggers and the retrieval query, as it is a guaranteed stable integer for this table structure.

  INSERT INTO memory_fts(rowid, content, key) VALUES (new.rowid, new.content, new.key);

if (config.query && ftsAvailable) {
const ftsQuery =
`SELECT m.* FROM ${tableName} m ` +
`JOIN memory_fts f ON m.memory_id = f.rowid ` +

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

As noted in the migration review, memory_id will be NULL for all new memories because it is not handled in the application's INSERT logic. Joining on memory_id will result in zero matches for new data. Please use rowid instead.

Suggested change
`JOIN memory_fts f ON m.memory_id = f.rowid ` +
`JOIN memory_fts f ON m.rowid = f.rowid ` +

obj.query = "";
}
const json = JSON.stringify(obj);
return json.length > 96000 ? json.slice(-96000) : json;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Slicing a JSON string directly is dangerous as it will likely result in invalid JSON (e.g., cutting off opening braces or quotes). This will cause JSON.parse to fail in the consumer. It is better to truncate the history array or the systemMsg content before calling JSON.stringify.

Comment on lines +276 to +278
const timeoutPromise = new Promise<Response>((_, reject) =>
setTimeout(() => reject(new Error("Summarization timeout")), SUMMARIZATION_TIMEOUT_MS)
);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The setTimeout used in the timeoutPromise is never cleared. This can lead to a resource leak and unnecessary event loop activity, especially if many summarization requests are handled. Additionally, the underlying handleSingleModel call is not aborted when the timeout fires, meaning the provider request continues to consume resources in the background.

Comment thread src/lib/db/encryption.ts
`[Encryption] Encryption failed: ${message}. ` +
`Check your STORAGE_ENCRYPTION_KEY — generate one with: openssl rand -base64 32`
);
return plaintext; // fallback to plaintext rather than crashing

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Falling back to returning the plaintext when encryption fails is a security risk. If a user has configured STORAGE_ENCRYPTION_KEY, they expect their data to be protected. Silently saving it in plaintext defeats the purpose of the feature. It is safer to throw an error and let the caller handle it (or let the application fail loudly).

@oyi77

oyi77 commented Apr 17, 2026

Copy link
Copy Markdown
Contributor Author

Code Review Fixes Applied

All issues from the gemini-code-assist review have been addressed:

✅ High Priority Issues Fixed

  1. Missing crypto imports (cursor.ts, perplexity-web.ts)

    • Added import crypto from "node:crypto"; to both files
    • Fixes runtime ReferenceError when calling createHash() and randomUUID()
  2. Memory FTS5 JOIN logic error (migration 023, retrieval.ts)

    • Changed JOIN from m.memory_id = f.rowid to m.rowid = f.rowid
    • Simplified migration to use SQLite's internal rowid directly (no memory_id column needed)
    • Fixes semantic search returning 0 results for new memories
  3. Encryption security vulnerability (encryption.ts)

    • Changed fallback behavior from returning plaintext to throwing error
    • Prevents silent data exposure when encryption fails
    • Users with STORAGE_ENCRYPTION_KEY now get loud failures instead of silent plaintext storage

✅ Medium Priority Issues Fixed

  1. Dangerous JSON truncation (perplexity-web.ts)

    • Replaced string slicing with proper array truncation before JSON.stringify
    • Truncates history array to 50 items first, then 10 if still too large
    • Prevents invalid JSON from breaking downstream parsing
  2. Timeout resource leak (contextHandoff.ts)

    • Added timeout ID tracking and clearTimeout() calls
    • Clears timeout on both success and error paths
    • Prevents event loop pollution from uncleaned timers

All changes have passed TypeScript type checking and lint validation.

@oyi77

oyi77 commented Apr 17, 2026

Copy link
Copy Markdown
Contributor Author

Additional Fix: Dynamic Context Limits

Added a second commit to make compaction context limits dynamic based on model capabilities:

Problem

The original fix used a hardcoded 8000 token limit for all models, which:

  • Was too small for large context models (Claude Opus 200K, Gemini 1M+)
  • Could still cause "Conversation history too large to compact" errors
  • Didn't leverage the models.dev integration already in OmniRoute

Solution

// Before: Hardcoded limit
const MAX_HISTORY_TOKENS_FOR_SUMMARY = 8000;

// After: Dynamic based on model
const modelContextLimit = getTokenLimit(provider, model);
const maxHistoryTokens = Math.floor(modelContextLimit * 0.4); // 40% safety margin
const effectiveLimit = Math.max(maxHistoryTokens, 8000); // Fallback to 8K

How It Works

  1. Parse model string to extract provider and model name
  2. Query getTokenLimit() which checks models.dev DB for actual context limits
  3. Use 40% of model's context window for history (safety margin for prompt + response)
  4. Fallback to 8000 tokens if model not found in DB

Examples

  • gpt-4o-mini (128K context) → ~51K tokens for history
  • claude-opus-4 (200K context) → ~80K tokens for history
  • gemini-3-pro (1M context) → ~400K tokens for history
  • Unknown model → 8K tokens (safe fallback)

Benefits

✅ Eliminates "history too large" errors for large context models
✅ Automatically adapts to any model's capabilities
✅ Leverages existing models.dev integration
✅ Maintains backward compatibility with 8K fallback

Commit: bc906602

@oyi77

oyi77 commented Apr 17, 2026

Copy link
Copy Markdown
Contributor Author

Third Fix: Handle Mixed Context Sizes in Combos

Added a third commit to handle combos with mixed model context sizes:

Problem

When a combo has models with different context windows:

combo: [gpt-4o-mini (128K), claude-opus-4 (200K), gemini-3-pro (1M)]

The previous implementation only used the first model's limit (128K), which meant:

  • ✅ Works fine if request succeeds on first model
  • ❌ But if it fails over to gemini-3-pro, we've already truncated to 128K
  • ❌ Wasting 872K of available context!

Solution

Calculate the minimum context limit across ALL combo targets:

// Before: Only first model
const summaryModel = relayConfig.handoffModel || options.model;
selectMessagesForSummary(messages, maxMessages, summaryModel);

// After: All combo targets
const modelsToConsider = options.comboTargets || [summaryModel];
selectMessagesForSummary(messages, maxMessages, modelsToConsider);

// Inside selectMessagesForSummary:
let minContextLimit = Infinity;
for (const modelStr of modelArray) {
  const limit = getTokenLimit(provider, model);
  minContextLimit = Math.min(minContextLimit, limit);
}

Example

Combo: [gpt-4o-mini (128K), claude-opus-4 (200K), gemini-3-pro (1M)]

  • Minimum limit: 128K (gpt-4o-mini)
  • History tokens: ~51K (40% of 128K)
  • Result: Compacted history works with ANY model in the combo

Why This Matters

  • ✅ Ensures handoff context is compatible with all fallback models
  • ✅ Prevents "history too large" errors during failover
  • ✅ Conservative approach: uses smallest limit for maximum compatibility
  • ✅ No wasted context if request succeeds on first model anyway

Commit: e72964d4

@oyi77 oyi77 closed this Apr 17, 2026
oyi77 added a commit to oyi77/OmniRoute that referenced this pull request Apr 17, 2026
- Add missing crypto imports in cursor.ts and perplexity-web.ts
- Fix memory FTS5 to use rowid instead of memory_id for correct JOIN
- Fix dangerous JSON string truncation in perplexity-web.ts by truncating history array before stringify
- Fix timeout resource leak in contextHandoff.ts by clearing timeout
- Fix encryption security issue by throwing error instead of falling back to plaintext

Addresses all high and medium priority issues from gemini-code-assist review.
oyi77 added a commit to oyi77/OmniRoute that referenced this pull request Apr 17, 2026
- Add missing crypto imports in cursor.ts and perplexity-web.ts
- Fix memory FTS5 to use rowid instead of memory_id for correct JOIN
- Fix dangerous JSON string truncation in perplexity-web.ts by truncating history array before stringify
- Fix timeout resource leak in contextHandoff.ts by clearing timeout
- Fix encryption security issue by throwing error instead of falling back to plaintext

Addresses all high and medium priority issues from gemini-code-assist review.
@diegosouzapw diegosouzapw mentioned this pull request Apr 30, 2026
@Tr0sT Tr0sT mentioned this pull request Apr 30, 2026
1 of 5 tasks
@diegosouzapw diegosouzapw mentioned this pull request May 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.