Skip to content

fix: Dynamic Context Compaction - #1368

Closed
oyi77 wants to merge 61 commits into
diegosouzapw:mainfrom
oyi77:fix/context-compaction-reapplied
Closed

oyi77 wants to merge 61 commits into
diegosouzapw:mainfrom
oyi77:fix/context-compaction-reapplied

Conversation

@oyi77

@oyi77 oyi77 commented Apr 17, 2026

Copy link
Copy Markdown
Contributor

This PR is a clean replacement for PR #1363, correctly rebased onto the latest main branch.

Work Included:

  1. Code Review Fixes Applied: Added missing crypto imports, corrected memory FTS5 logic, hardened database encryption against fallback failures, fixed dangerous JSON string truncation in the Perplexity executor, and eliminated timeout resource leaks in the context handoff service.
  2. Dynamic Context Limits: Context compaction tokens are now calculated dynamically (40% of context window) using the models.dev API integration instead of a hardcoded 8,000 limit. This ensures safe truncation for large context models like Opus or Gemini 1M.
  3. Handle Mixed Context Sizes in Combos: When routing through a combo, the minimum context limit across all models in the combo is used for summarization, guaranteeing compatibility during fallback scenarios.

diegosouzapw and others added 30 commits April 14, 2026 07:46
…ent literal \n\n artifacts

- Changed regex quantifier from ? to * in combo.ts, comboAgentMiddleware.ts,
  and contextHandoff.ts to greedily strip all JSON-escaped newline sequences
  surrounding <omniModel> tags in SSE streaming chunks
- Added \r to the character class for cross-platform robustness
- Fixed Playwright strict-mode violation in combo-unification.spec.ts
- Bumped OpenAPI version and CHANGELOG to 3.6.6
…w#1187/diegosouzapw#1218, diegosouzapw#1202)

- fix(gemini): strip VS Code JSON Schema extensions from tool schemas (diegosouzapw#1175)
  Add enumDescriptions, markdownDescription, markdownEnumDescriptions,
  enumItemLabels and tags to UNSUPPORTED_SCHEMA_CONSTRAINTS so the Gemini
  sanitizer removes them before forwarding. GitHub Copilot injects these
  non-standard fields into tool definitions, causing Gemini to reject with
  'Unknown name enumDescriptions at functionDeclarations[n].parameters'.

- fix(health-check): unwrap proxy config object before passing to getAccessToken (diegosouzapw#1187 diegosouzapw#1218)
  resolveProxyForConnection() returns { proxy, level, levelId } but the health
  check loop was passing the full wrapper to getAccessToken(), which expects the
  inner config object (.host, .port etc). The proxy dispatcher validated .host
  on the wrapper (undefined) and threw 'Context proxy host is required', silently
  marking every connection as unhealthy every sweep. Fix mirrors the pattern
  already used in chatHelpers.ts: proxyResult?.proxy || null.

- fix(ui): debounce models.dev sync interval slider to save only on release (diegosouzapw#1202)
  The slider's onChange fired updateInterval() on every drag tick, sending a
  PATCH per pixel of movement. Rapid API responses overwrote UI state mid-drag.
  Introduce draftIntervalHours for smooth visual feedback; the PATCH fires
  on onMouseUp / onBlur once the user releases the control.
diegosouzapw#1220, diegosouzapw#1231)

- fix(core): diegosouzapw#1206 inject startup guard against app/ and src/app/ conflict
- fix(health): diegosouzapw#1220 add HEALTHCHECK_STAGGER_MS to prevent token refresh bursting
- fix(proxy): diegosouzapw#1231 prioritize HTTP 429 over quota body heuristics
- fix(sse): diegosouzapw#1211 strip leading double-newlines in responses API stream
Update package versions for the electron app and open-sse package.
Sync llm.txt metadata and feature headings with the 3.6.6 release.
Add guarded outbound fetch helpers with private/local URL blocking,
controlled retries, timeout normalization, and route-level status
propagation for provider validation and model discovery.

Introduce cooldown-aware chat retries with configurable
requestRetry and maxRetryIntervalSec settings, model-scoped cooldown
responses, and improved rate-limit learning from headers and error
bodies so short upstream lockouts can recover automatically.

Also align Antigravity and Codex header handling, require API keys
for Pollinations, validate web runtime env at startup, restore
sanitized Gemini tool names in translated responses, and inject a
synthetic Claude text block when upstream SSE completes empty.
Introduce GLM Thinking as a first-class provider preset with shared GLM
model metadata, pricing, usage sync, dashboard support, and provider
request defaults for higher token budgets and longer timeouts.

Use provider-side /messages/count_tokens when a Claude-compatible
upstream supports it, while preserving estimated fallback behavior for
missing models, missing credentials, and upstream failures.

Also add startup seeding for default model aliases and normalize common
cross-proxy model dialects so canonical slashful model ids do not get
misrouted during resolution.
Add dedicated sync token storage, issuance, revocation, and bundle
download routes backed by stable config bundle versioning and ETag
support.

Expose the v1 websocket handshake route and custom Next server bridge so
OpenAI-compatible websocket traffic can be upgraded and proxied through
the dashboard and API bridge.

Expand compliance auditing with structured metadata, pagination, request
context, auth and provider credential events, and SSRF-blocked
validation logging.
- CHANGELOG: Add WebSocket bridge, GLM Thinking preset, safe outbound
  fetch/SSRF guard, cooldown-aware retries, compliance audit v2, model
  alias seeding, and all Internal Improvements for the 3 new commits
- README: Expand v3.6.x highlights table with 10 new features; add
  SafeOutboundFetch, CooldownAwareRetry, SSRF guard, TPS metric, sync
  tokens, WebSocket bridge to Resilience/Observability/Deployment tables
- ARCHITECTURE: Bump date; add new modules to executive summary, API
  routes, SSE core services, Auth/Security section; add SSRF/Outbound
  guard failure mode (section 6); expand module mapping
- ENVIRONMENT: Add OMNIROUTE_CRYPT_KEY/OMNIROUTE_API_KEY_BASE64 legacy
  aliases, OUTBOUND_SSRF_GUARD_ENABLED, CODEX_CLIENT_VERSION, and
  REQUEST_RETRY/MAX_RETRY_INTERVAL_SEC cooldown retry settings
- FEATURES: Add 6 new feature sections — V1 WebSocket Bridge, Sync
  Tokens & Config Bundle, GLM Thinking Preset, Safe Outbound Fetch &
  SSRF Guard, Cooldown-Aware Retries, Compliance Audit v2
Integrated into release/v3.6.6 — IPv6 proxy test fix
…osouzapw#1200 (diegosouzapw#1256)

Integrated into release/v3.6.6 — Gemini custom model picker fix
…ests (diegosouzapw#1246)

Integrated into release/v3.6.6 — OAuth client_id default fallbacks
…n Chat→Responses translator (diegosouzapw#1245)

Integrated into release/v3.6.6 — max_tokens → max_output_tokens Responses API translation + unit tests
…egosouzapw#1258)

Integrated into release/v3.6.6 — cursor-agent CLI credential source support
…eout behavior (diegosouzapw#1257)

Integrated into release/v3.6.6 — CC-compatible upstream SSE restore + stream timeout fix + README table repair
This integrates rdself's unify-provider-profile-locks PR manually to handle structural conflicts.
diegosouzapw and others added 20 commits April 15, 2026 15:56
…del button positioning, and clarify oauth documentation
Document the latest fixes for Codex routing configuration parsing and
Lobehub provider icon fallback behavior.

Add the note that the remaining JavaScript test files were migrated to
TypeScript ES modules to reflect the completed test stack transition.
Proactively compress oversized contexts before sending to upstream providers,
preventing context_length_exceeded errors. Compression triggers at 85% of
model's context limit using the existing 3-layer compressContext() function.

- Import compressContext, estimateTokens, getTokenLimit from contextManager
- Add compression check after translation, before executor dispatch
- Estimate tokens and compare against 85% threshold of model's context limit
- Apply 3-layer compression (trim tools, compress thinking, purify history)
- Log compression events with before/after token counts and layers applied
- Audit compression events for observability
- Add unit tests verifying integration behavior

Closes diegosouzapw#1290
When purifyHistory() drops oldest messages to fit context window, it can
split tool_use/tool_result pairs — keeping the tool_result but dropping
the tool_use that initiated it. This causes upstream providers to reject
the request with format errors.

Add fixToolPairs() that runs after each purification pass to remove:
- OpenAI format: orphaned role='tool' messages without matching tool_calls ID
- Claude format: orphaned tool_result content blocks without matching tool_use ID

Closes diegosouzapw#1291
- Add missing crypto imports in cursor.ts and perplexity-web.ts
- Fix memory FTS5 to use rowid instead of memory_id for correct JOIN
- Fix dangerous JSON string truncation in perplexity-web.ts by truncating history array before stringify
- Fix timeout resource leak in contextHandoff.ts by clearing timeout
- Fix encryption security issue by throwing error instead of falling back to plaintext

Addresses all high and medium priority issues from gemini-code-assist review.
- Replace hardcoded 8000 token limit with dynamic calculation from models.dev
- Use 40% of model's context window for history (safety margin)
- Parse model string to get provider and model name
- Query getTokenLimit() which checks models.dev DB for actual context limits
- Fallback to 8000 tokens if model not found
- Add warning log when history too large even with single message

This ensures compaction works with any model size:
- Small models (8K context): ~3.2K tokens for history
- Medium models (128K context): ~51K tokens for history
- Large models (1M+ context): ~400K+ tokens for history

Fixes 'Conversation history too large to compact' errors by adapting to model capabilities.
Problem: When a combo has mixed context sizes (e.g., gpt-4o-mini 128K, gemini-3-pro 1M),
using only the first model's limit could cause compaction to fail on fallback models.

Solution:
- Pass all combo target models to selectMessagesForSummary()
- Calculate minimum context limit across all targets
- Use the smallest limit to ensure compatibility with any fallback

Example combo: [gpt-4o-mini (128K), claude-opus (200K), gemini-3-pro (1M)]
- Before: Used 128K limit (first model only)
- After: Uses 128K limit (minimum across all targets)
- Result: Compacted history works with any model in the combo

This ensures handoff context is compatible with all potential fallback models.
@oyi77
oyi77 requested a review from diegosouzapw as a code owner April 17, 2026 13:53

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request upgrades OmniRoute to version 3.6.6, introducing a V1 WebSocket bridge, a configuration sync system with scoped tokens, the GLM Thinking model preset, and the Perplexity Web provider. Security is significantly hardened through a new outbound SSRF guard, runtime environment validation, and a fix for a moderate vulnerability in the follow-redirects dependency. The update also adds hybrid token counting, cooldown-aware retries, and UI enhancements like pagination for memory and skills. Feedback suggests several improvements: increasing the severity of critical security audit events, aborting migrations upon detecting name mismatches, enhancing error reporting in semantic search, optimizing sequential database queries in the memory store, and implementing signal-aware sleeps in the retry logic to ensure immediate cancellation.

Comment on lines +30 to +39
logAuditEvent({
action: "auth.login.misconfigured",
actor: "system",
target: "dashboard-auth",
resourceType: "auth_session",
status: "failed",
ipAddress: auditContext.ipAddress || undefined,
requestId: auditContext.requestId,
metadata: { reason: "missing_jwt_secret" },
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The audit event for a missing JWT secret is logged as 'system' actor. Since this is a critical configuration error preventing authentication, consider logging this with a higher severity or ensuring it triggers an alert, as it indicates a broken security setup.

Comment on lines +115 to +134
function detectNameMismatches(
appliedRecords: Array<{ version: string; name: string }>,
files: Array<{ version: string; name: string; path: string }>
): Array<{ version: string; appliedName: string; diskName: string }> {
const appliedByName = new Map(appliedRecords.map((r) => [r.version, r.name]));
const mismatches: Array<{ version: string; appliedName: string; diskName: string }> = [];

for (const file of files) {
const appliedName = appliedByName.get(file.version);
if (appliedName && appliedName !== file.name) {
mismatches.push({
version: file.version,
appliedName,
diskName: file.name,
});
}
}

return mismatches;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The detectNameMismatches function only logs errors but does not prevent the migration from proceeding. If a renumbering mismatch is detected, it should likely abort the migration process to prevent inconsistent database states.

Comment on lines +181 to 218
case "semantic": {
if (config.query && ftsAvailable) {
const ftsQuery =
`SELECT m.* FROM ${tableName} m ` +
`JOIN memory_fts f ON m.rowid = f.rowid ` +
`WHERE f.memory_fts MATCH ? AND m.${columns.apiKeyId} = ? ` +
`AND (m.${columns.expiresAt} IS NULL OR datetime(m.${columns.expiresAt}) > datetime('now'))` +
(normalizedConfig.scope === "session" && config.sessionId
? ` AND m.${columns.sessionId} = ?`
: "") +
(normalizedConfig.retentionDays > 0
? ` AND datetime(m.${columns.createdAt}) >= datetime(?)`
: "") +
` ORDER BY f.rank LIMIT 100`;
const ftsParams: any[] = [config.query, apiKeyId];
if (normalizedConfig.scope === "session" && config.sessionId) {
ftsParams.push(config.sessionId);
}
if (normalizedConfig.retentionDays > 0) {
const cutoff = new Date(
Date.now() - normalizedConfig.retentionDays * 24 * 60 * 60 * 1000
).toISOString();
ftsParams.push(cutoff);
}
try {
rows = db.prepare(ftsQuery).all(...ftsParams) as MemoryRow[];
} catch {
rows = [];
}
if (rows.length === 0) {
query += ` ORDER BY ${columns.createdAt} DESC LIMIT 100`;
rows = db.prepare(query).all(...params) as MemoryRow[];
}
} else {
query += ` ORDER BY ${columns.createdAt} DESC LIMIT 100`;
rows = db.prepare(query).all(...params) as MemoryRow[];
}
break;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The semantic search implementation uses a try-catch block that silently swallows errors and returns an empty array. This hides potential database issues (e.g., syntax errors in generated SQL) during FTS5 queries. It should log the error or rethrow if the database is in an invalid state.

Comment thread src/lib/memory/store.ts
Comment on lines 260 to 328
whereClauses.push("api_key_id = ?");
params.push(filters.apiKeyId);
whereParams.push(filters.apiKeyId);
}

if (filters.type) {
whereClauses.push("type = ?");
params.push(filters.type);
whereParams.push(filters.type);
}

if (filters.sessionId) {
whereClauses.push("session_id = ?");
params.push(filters.sessionId);
whereParams.push(filters.sessionId);
}

if (typeof filters.query === "string" && filters.query.trim().length > 0) {
const likeQuery = `%${filters.query.trim().toLowerCase()}%`;
whereClauses.push("(LOWER(content) LIKE ? OR LOWER(key) LIKE ?)");
whereParams.push(likeQuery, likeQuery);
}

// Run COUNT query + byType aggregation in a single query
let countQuery = "SELECT COUNT(*) as total FROM memories";
if (whereClauses.length > 0) {
countQuery += " WHERE " + whereClauses.join(" AND ");
}
const countStmt = db.prepare(countQuery);
const countRow = countStmt.get(...whereParams) as { total: number };
const total = countRow.total;

// Build byType aggregation (counts ALL matching rows, not just the page)
let byTypeQuery = "SELECT type, COUNT(*) as count FROM memories";
const byTypeParams: unknown[] = [...whereParams];
if (whereClauses.length > 0) {
byTypeQuery += " WHERE " + whereClauses.join(" AND ");
}
byTypeQuery += " GROUP BY type";
const byTypeStmt = db.prepare(byTypeQuery);
const byTypeRows = byTypeStmt.all(...byTypeParams) as { type: string; count: number }[];
const byType = Object.fromEntries(byTypeRows.map((r) => [r.type, r.count])) as Record<
string,
number
>;

// Calculate effective limit and offset
const effectiveLimit = filters.limit ?? 50;
const effectivePage = filters.page ?? 1;
const effectiveOffset = filters.offset ?? (effectivePage - 1) * effectiveLimit;

// Build SELECT query with pagination
let query = "SELECT * FROM memories";
if (whereClauses.length > 0) {
query += " WHERE " + whereClauses.join(" AND ");
}

// Add ordering and pagination
query += " ORDER BY created_at DESC";
query += " ORDER BY created_at DESC LIMIT ? OFFSET ?";

if (filters.limit !== undefined) {
query += " LIMIT ?";
params.push(filters.limit);
}

if (filters.offset !== undefined) {
if (filters.limit === undefined) {
query += " LIMIT -1";
}
query += " OFFSET ?";
params.push(filters.offset);
}
// Build params for SELECT query (WHERE params + pagination params)
const params = [...whereParams, effectiveLimit, effectiveOffset];

const stmt = db.prepare(query);
const rows = stmt.all(...params);

return (rows as MemoryRow[]).map(rowToMemory);
return {
data: (rows as MemoryRow[]).map(rowToMemory),
total,
byType,
};
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The listMemories function performs multiple database queries (COUNT, byType, and SELECT) sequentially. For large datasets, this could be a performance bottleneck. Consider optimizing these into a single query or using a more efficient approach if the memory table grows significantly.

Comment on lines +276 to +331
for (let attempt = 1; attempt <= retryConfig.attempts; attempt++) {
try {
const executeFetch = () =>
fetchWithTimeout(targetUrl.toString(), {
...fetchOptions,
method,
redirect,
signal,
timeoutMs,
});

const response = proxyConfig
? await runWithProxyContext(proxyConfig, executeFetch)
: await executeFetch();

if (!allowRedirect && response.status >= 300 && response.status < 400) {
const location = response.headers.get("location");
await cancelResponseBody(response);
throw new SafeOutboundFetchError(
`Redirect blocked for ${method} ${targetUrl.toString()} (${response.status})`,
{
code: "REDIRECT_BLOCKED",
url: targetUrl.toString(),
method,
attempts: attempt,
status: response.status,
location,
isRetryable: false,
}
);
}

if (
retryConfig.shouldRetryMethod &&
attempt < retryConfig.attempts &&
retryConfig.statusCodes.has(response.status)
) {
await cancelResponseBody(response);
await sleep(getBackoffDelay(retryConfig.backoffMs, attempt));
continue;
}

return response;
} catch (error) {
const normalizedError = normalizeFetchFailure(error, targetUrl.toString(), method, attempt);
const shouldRetry =
retryConfig.shouldRetryMethod &&
attempt < retryConfig.attempts &&
normalizedError.isRetryable;

if (!shouldRetry) {
throw normalizedError;
}

await sleep(getBackoffDelay(retryConfig.backoffMs, attempt));
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The retry loop uses sleep(getBackoffDelay(...)) which is not interruptible by the signal passed to safeOutboundFetch. If the request is aborted during a backoff sleep, the function will continue to wait until the sleep finishes before checking the signal. Use a signal-aware sleep or setTimeout with signal.addEventListener to ensure immediate cancellation.

@oyi77 oyi77 changed the title fix: Re-apply PR #1363 (Code Review Fixes & Dynamic Context Compaction) fix: Dynamic Context Compaction Apr 17, 2026
@oyi77 oyi77 closed this Apr 17, 2026
@diegosouzapw diegosouzapw mentioned this pull request Apr 30, 2026
@Tr0sT Tr0sT mentioned this pull request Apr 30, 2026
1 of 5 tasks
@diegosouzapw diegosouzapw mentioned this pull request May 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.