Skip to content

fix(oom): prevent per-request memory accumulation (256MB heap) - #2973

Merged
diegosouzapw merged 13 commits into
diegosouzapw:release/v3.8.8from
soyelmismo:fix/cpu-leak-sse-streaming
Jun 1, 2026
Merged

diegosouzapw merged 13 commits into
diegosouzapw:release/v3.8.8from
soyelmismo:fix/cpu-leak-sse-streaming

Conversation

@soyelmismo

Copy link
Copy Markdown
Contributor

Problem

OOM crashes within 5 minutes of intensive use. CPU climbs like a staircase (0.1% → 2-3%) then process dies. Larger contexts cause faster OOM.

Root Cause

persistAttemptLogs in chatCore.ts deep-clones translatedBody (1-5MB with large contexts) 17 times per request via cloneBoundedChatLogPayload. With 5 concurrent requests that's 85-250MB of clones alone — exceeding the 256MB heap.

Fix (2 changes in chatCore.ts)

  1. truncateForLog() — caps request/response bodies logged by persistAttemptLogs at 8KB JSON. Returns a lightweight summary (model, provider, message count) instead of a full deep clone when the body exceeds 8KB.

  2. Memory pressure guard — checks process.memoryUsage().heapUsed at the top of handleChatCore. Returns HTTP 503 when heap exceeds HEAP_PRESSURE_THRESHOLD_MB (default 200MB, configurable via env). Self-healing: no counters to leak, no timers to clean up.

Scope

  • 1 file changed: open-sse/handlers/chatCore.ts (+57, -2)
  • Zero new TypeScript errors (only pre-existing L3101/L3198 ClaudeMessage type mismatch)
  • 256MB heap limit preserved — no Dockerfile changes

Not in this PR

The previous PR #2965 addressed cache eviction (comboMetrics, usage, providerRegistry). This PR addresses the per-request memory accumulation that was the second OOM vector.

branben and others added 11 commits May 30, 2026 21:17
…ab, and 20 tests (#2959)

Integrated into release/v3.8.8
… scopes to all dynamic tool definitions (#2958)

Integrated into release/v3.8.8
…ombo target's providerId over model-inferred provider (#2946)

Integrated into release/v3.8.8
…native Claude OAuth (#2943)

Integrated into release/v3.8.8
Integrated into release/v3.8.8
Two fixes to keep the process stable within the 256MB heap limit:

1. truncateForLog(): caps request/response bodies logged by
   persistAttemptLogs at 8KB JSON. Previously, translatedBody
   (1-5MB with large contexts) was deep-cloned 17x per request.
   With 5 concurrent requests that's 85-250MB of clones alone.
   Now each clone is at most 8KB (summary with model, provider,
   message count instead of full body).

2. Memory pressure guard: checks process.memoryUsage().heapUsed
   at the top of handleChatCore. Returns 503 when heap exceeds
   HEAP_PRESSURE_THRESHOLD_MB (default 200MB, configurable via
   env). Prevents cascading OOM when many large-context requests
   arrive concurrently. Self-healing: no counters to leak.
@soyelmismo
soyelmismo requested a review from diegosouzapw as a code owner May 31, 2026 04:44

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a global memory pressure guard to reject incoming requests with a 503 status when V8 heap usage exceeds a threshold, and adds a truncateForLog utility to truncate large request/response bodies before logging. The review feedback suggests optimizing truncateForLog by using the existing estimateSizeFast utility instead of JSON.stringify to avoid CPU and memory overhead on large objects. Additionally, the reviewer notes that unit tests must be added for these changes to comply with the repository style guide regarding production code modifications.

Comment thread open-sse/handlers/chatCore.ts
Comment thread open-sse/handlers/chatCore.ts
@soyelmismo
soyelmismo marked this pull request as draft May 31, 2026 04:46
1. truncateForLog now uses estimateSizeFast() instead of
   JSON.stringify() to check object size. This avoids serializing
   multi-MB request bodies into strings just to measure them.
   estimateSizeFast walks the object tree directly (safe for
   circular refs, early-exits at 256KB).

2. Add 6 unit tests for the memory management changes:
   - estimateSizeFast: small/large/primitives/circular refs
   - HEAP_PRESSURE_THRESHOLD_MB default value
   - 8KB threshold logic (small vs large payloads)
@soyelmismo

Copy link
Copy Markdown
Contributor Author
imagen this is what im trying to fix. we are not in solana exchange to allow this.

@soyelmismo
soyelmismo marked this pull request as ready for review May 31, 2026 05:05
@soyelmismo
soyelmismo marked this pull request as draft May 31, 2026 05:58
@soyelmismo

Copy link
Copy Markdown
Contributor Author

k this is too aggressive so leave this untouched until i check it later

@diegosouzapw
diegosouzapw changed the base branch from main to release/v3.8.8 June 1, 2026 10:37
…k-sse-streaming

# Conflicts:
#	.env.example
#	.source/browser.ts
#	.source/server.ts
#	docs/frameworks/MCP-SERVER.md
#	open-sse/handlers/chatCore.ts
#	open-sse/mcp-server/server.ts
#	src/app/(dashboard)/dashboard/api-manager/ApiManagerPageClient.tsx
#	tests/unit/api-manager-page-static.test.ts
@diegosouzapw
diegosouzapw marked this pull request as ready for review June 1, 2026 11:06
@diegosouzapw
diegosouzapw merged commit 57dfa25 into diegosouzapw:release/v3.8.8 Jun 1, 2026
2 checks passed
@diegosouzapw

Copy link
Copy Markdown
Owner

Merged into release/v3.8.8 🎉 Thanks @soyelmismo! Solid OOM fix (8KB log-body cap + heap-pressure guard). I pushed one review fix — the 503 body no longer exposes the heap figure (kept it in an internal log instead, per our error-sanitization rule). Ships next release.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b5dc193161

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

// ── Global memory pressure guard ────────────────────────────────────────
// Prevents OOM by rejecting new requests when V8 heap exceeds threshold.
// Self-healing: no counters to leak, no cleanup needed.
const HEAP_PRESSURE_THRESHOLD_MB = parseInt(process.env.HEAP_PRESSURE_THRESHOLD_MB || "200", 10);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Tie heap-pressure threshold to the configured heap limit

With the new hard-coded 200 MB default, any deployment using the repo defaults can start returning 503s while still far from OOM: Docker sets OMNIROUTE_MEMORY_MB=1024 in Dockerfile, and omniroute serve defaults/clamps to 512 MB, but the guard rejects all chat requests once heapUsed crosses 200 MB unless operators discover and set this separate env var. This makes normal high-memory-but-healthy processes unavailable; derive the threshold from the actual V8 heap limit / OMNIROUTE_MEMORY_MB or default it proportionally.

Useful? React with 👍 / 👎.

@soyelmismo

Copy link
Copy Markdown
Contributor Author

oh man i left this as a draft, it was giving me issues D:

diegosouzapw added a commit that referenced this pull request Jun 1, 2026
…env-doc fixes

- bump package.json / open-sse / electron / openapi / llm.txt to 3.8.8
- restructure CHANGELOG: Unreleased -> [3.8.8], dedup broken Notion/MCP block (was 12x)
- add every PR since v3.8.7 that was missing: Quota Share Engine (#2859/#3022/#3032),
  page redesigns (#2827/#2839/#2847/#2849/#2869/#2873), and fixes #2960/#2973/#2984/
  #3021/#3029/#3031/#3035/#3036/#3037/#3039/#3043/#3028; folded #2978/#2988/#3041
- insert [3.8.8] section into all 41 i18n CHANGELOGs + sync llm.txt mirrors
- document OMNIROUTE_PLUGINS_ALLOW_EXEC in .env.example + ENVIRONMENT.md (env-doc-sync gap)
@diegosouzapw diegosouzapw mentioned this pull request Jun 2, 2026
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…souzapw#2973)

Integrated into release/v3.8.8. OOM fix: truncateForLog caps logged bodies at 8KB (prevents multi-MB clone accumulation across log call-sites) + a heap-pressure 503 guard. Applied review fix: the 503 body no longer leaks the heap figure (Hard Rule diegosouzapw#12) — logged internally instead. Thanks @soyelmismo!
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…env-doc fixes

- bump package.json / open-sse / electron / openapi / llm.txt to 3.8.8
- restructure CHANGELOG: Unreleased -> [3.8.8], dedup broken Notion/MCP block (was 12x)
- add every PR since v3.8.7 that was missing: Quota Share Engine (diegosouzapw#2859/diegosouzapw#3022/diegosouzapw#3032),
  page redesigns (diegosouzapw#2827/diegosouzapw#2839/diegosouzapw#2847/diegosouzapw#2849/diegosouzapw#2869/diegosouzapw#2873), and fixes diegosouzapw#2960/diegosouzapw#2973/diegosouzapw#2984/
  diegosouzapw#3021/diegosouzapw#3029/diegosouzapw#3031/diegosouzapw#3035/diegosouzapw#3036/diegosouzapw#3037/diegosouzapw#3039/diegosouzapw#3043/diegosouzapw#3028; folded diegosouzapw#2978/diegosouzapw#2988/diegosouzapw#3041
- insert [3.8.8] section into all 41 i18n CHANGELOGs + sync llm.txt mirrors
- document OMNIROUTE_PLUGINS_ALLOW_EXEC in .env.example + ENVIRONMENT.md (env-doc-sync gap)
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
…souzapw#2973)

Integrated into release/v3.8.8. OOM fix: truncateForLog caps logged bodies at 8KB (prevents multi-MB clone accumulation across log call-sites) + a heap-pressure 503 guard. Applied review fix: the 503 body no longer leaks the heap figure (Hard Rule diegosouzapw#12) — logged internally instead. Thanks @soyelmismo!
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
…env-doc fixes

- bump package.json / open-sse / electron / openapi / llm.txt to 3.8.8
- restructure CHANGELOG: Unreleased -> [3.8.8], dedup broken Notion/MCP block (was 12x)
- add every PR since v3.8.7 that was missing: Quota Share Engine (diegosouzapw#2859/diegosouzapw#3022/diegosouzapw#3032),
  page redesigns (diegosouzapw#2827/diegosouzapw#2839/diegosouzapw#2847/diegosouzapw#2849/diegosouzapw#2869/diegosouzapw#2873), and fixes diegosouzapw#2960/diegosouzapw#2973/diegosouzapw#2984/
  diegosouzapw#3021/diegosouzapw#3029/diegosouzapw#3031/diegosouzapw#3035/diegosouzapw#3036/diegosouzapw#3037/diegosouzapw#3039/diegosouzapw#3043/diegosouzapw#3028; folded diegosouzapw#2978/diegosouzapw#2988/diegosouzapw#3041
- insert [3.8.8] section into all 41 i18n CHANGELOGs + sync llm.txt mirrors
- document OMNIROUTE_PLUGINS_ALLOW_EXEC in .env.example + ENVIRONMENT.md (env-doc-sync gap)
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…souzapw#2973)

Integrated into release/v3.8.8. OOM fix: truncateForLog caps logged bodies at 8KB (prevents multi-MB clone accumulation across log call-sites) + a heap-pressure 503 guard. Applied review fix: the 503 body no longer leaks the heap figure (Hard Rule diegosouzapw#12) — logged internally instead. Thanks @soyelmismo!
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…env-doc fixes

- bump package.json / open-sse / electron / openapi / llm.txt to 3.8.8
- restructure CHANGELOG: Unreleased -> [3.8.8], dedup broken Notion/MCP block (was 12x)
- add every PR since v3.8.7 that was missing: Quota Share Engine (diegosouzapw#2859/diegosouzapw#3022/diegosouzapw#3032),
  page redesigns (diegosouzapw#2827/diegosouzapw#2839/diegosouzapw#2847/diegosouzapw#2849/diegosouzapw#2869/diegosouzapw#2873), and fixes diegosouzapw#2960/diegosouzapw#2973/diegosouzapw#2984/
  diegosouzapw#3021/diegosouzapw#3029/diegosouzapw#3031/diegosouzapw#3035/diegosouzapw#3036/diegosouzapw#3037/diegosouzapw#3039/diegosouzapw#3043/diegosouzapw#3028; folded diegosouzapw#2978/diegosouzapw#2988/diegosouzapw#3041
- insert [3.8.8] section into all 41 i18n CHANGELOGs + sync llm.txt mirrors
- document OMNIROUTE_PLUGINS_ALLOW_EXEC in .env.example + ENVIRONMENT.md (env-doc-sync gap)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants