fix(oom): prevent per-request memory accumulation (256MB heap) - #2973
diegosouzapw merged 13 commits into
Conversation
…ab, and 20 tests (#2959) Integrated into release/v3.8.8
… scopes to all dynamic tool definitions (#2958) Integrated into release/v3.8.8
…2957) Integrated into release/v3.8.8
…mizations (#2951) Integrated into release/v3.8.8
…ombo target's providerId over model-inferred provider (#2946) Integrated into release/v3.8.8
…native Claude OAuth (#2943) Integrated into release/v3.8.8
Integrated into release/v3.8.8
Integrated into release/v3.8.8
Integrated into release/v3.8.8
Integrated into release/v3.8.8
Two fixes to keep the process stable within the 256MB heap limit: 1. truncateForLog(): caps request/response bodies logged by persistAttemptLogs at 8KB JSON. Previously, translatedBody (1-5MB with large contexts) was deep-cloned 17x per request. With 5 concurrent requests that's 85-250MB of clones alone. Now each clone is at most 8KB (summary with model, provider, message count instead of full body). 2. Memory pressure guard: checks process.memoryUsage().heapUsed at the top of handleChatCore. Returns 503 when heap exceeds HEAP_PRESSURE_THRESHOLD_MB (default 200MB, configurable via env). Prevents cascading OOM when many large-context requests arrive concurrently. Self-healing: no counters to leak.
There was a problem hiding this comment.
Code Review
This pull request introduces a global memory pressure guard to reject incoming requests with a 503 status when V8 heap usage exceeds a threshold, and adds a truncateForLog utility to truncate large request/response bodies before logging. The review feedback suggests optimizing truncateForLog by using the existing estimateSizeFast utility instead of JSON.stringify to avoid CPU and memory overhead on large objects. Additionally, the reviewer notes that unit tests must be added for these changes to comply with the repository style guide regarding production code modifications.
1. truncateForLog now uses estimateSizeFast() instead of JSON.stringify() to check object size. This avoids serializing multi-MB request bodies into strings just to measure them. estimateSizeFast walks the object tree directly (safe for circular refs, early-exits at 256KB). 2. Add 6 unit tests for the memory management changes: - estimateSizeFast: small/large/primitives/circular refs - HEAP_PRESSURE_THRESHOLD_MB default value - 8KB threshold logic (small vs large payloads)
|
k this is too aggressive so leave this untouched until i check it later |
…k-sse-streaming # Conflicts: # .env.example # .source/browser.ts # .source/server.ts # docs/frameworks/MCP-SERVER.md # open-sse/handlers/chatCore.ts # open-sse/mcp-server/server.ts # src/app/(dashboard)/dashboard/api-manager/ApiManagerPageClient.tsx # tests/unit/api-manager-page-static.test.ts
|
Merged into |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b5dc193161
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // ── Global memory pressure guard ──────────────────────────────────────── | ||
| // Prevents OOM by rejecting new requests when V8 heap exceeds threshold. | ||
| // Self-healing: no counters to leak, no cleanup needed. | ||
| const HEAP_PRESSURE_THRESHOLD_MB = parseInt(process.env.HEAP_PRESSURE_THRESHOLD_MB || "200", 10); |
There was a problem hiding this comment.
Tie heap-pressure threshold to the configured heap limit
With the new hard-coded 200 MB default, any deployment using the repo defaults can start returning 503s while still far from OOM: Docker sets OMNIROUTE_MEMORY_MB=1024 in Dockerfile, and omniroute serve defaults/clamps to 512 MB, but the guard rejects all chat requests once heapUsed crosses 200 MB unless operators discover and set this separate env var. This makes normal high-memory-but-healthy processes unavailable; derive the threshold from the actual V8 heap limit / OMNIROUTE_MEMORY_MB or default it proportionally.
Useful? React with 👍 / 👎.
|
oh man i left this as a draft, it was giving me issues D: |
…env-doc fixes - bump package.json / open-sse / electron / openapi / llm.txt to 3.8.8 - restructure CHANGELOG: Unreleased -> [3.8.8], dedup broken Notion/MCP block (was 12x) - add every PR since v3.8.7 that was missing: Quota Share Engine (#2859/#3022/#3032), page redesigns (#2827/#2839/#2847/#2849/#2869/#2873), and fixes #2960/#2973/#2984/ #3021/#3029/#3031/#3035/#3036/#3037/#3039/#3043/#3028; folded #2978/#2988/#3041 - insert [3.8.8] section into all 41 i18n CHANGELOGs + sync llm.txt mirrors - document OMNIROUTE_PLUGINS_ALLOW_EXEC in .env.example + ENVIRONMENT.md (env-doc-sync gap)
…souzapw#2973) Integrated into release/v3.8.8. OOM fix: truncateForLog caps logged bodies at 8KB (prevents multi-MB clone accumulation across log call-sites) + a heap-pressure 503 guard. Applied review fix: the 503 body no longer leaks the heap figure (Hard Rule diegosouzapw#12) — logged internally instead. Thanks @soyelmismo!
…env-doc fixes - bump package.json / open-sse / electron / openapi / llm.txt to 3.8.8 - restructure CHANGELOG: Unreleased -> [3.8.8], dedup broken Notion/MCP block (was 12x) - add every PR since v3.8.7 that was missing: Quota Share Engine (diegosouzapw#2859/diegosouzapw#3022/diegosouzapw#3032), page redesigns (diegosouzapw#2827/diegosouzapw#2839/diegosouzapw#2847/diegosouzapw#2849/diegosouzapw#2869/diegosouzapw#2873), and fixes diegosouzapw#2960/diegosouzapw#2973/diegosouzapw#2984/ diegosouzapw#3021/diegosouzapw#3029/diegosouzapw#3031/diegosouzapw#3035/diegosouzapw#3036/diegosouzapw#3037/diegosouzapw#3039/diegosouzapw#3043/diegosouzapw#3028; folded diegosouzapw#2978/diegosouzapw#2988/diegosouzapw#3041 - insert [3.8.8] section into all 41 i18n CHANGELOGs + sync llm.txt mirrors - document OMNIROUTE_PLUGINS_ALLOW_EXEC in .env.example + ENVIRONMENT.md (env-doc-sync gap)
…souzapw#2973) Integrated into release/v3.8.8. OOM fix: truncateForLog caps logged bodies at 8KB (prevents multi-MB clone accumulation across log call-sites) + a heap-pressure 503 guard. Applied review fix: the 503 body no longer leaks the heap figure (Hard Rule diegosouzapw#12) — logged internally instead. Thanks @soyelmismo!
…env-doc fixes - bump package.json / open-sse / electron / openapi / llm.txt to 3.8.8 - restructure CHANGELOG: Unreleased -> [3.8.8], dedup broken Notion/MCP block (was 12x) - add every PR since v3.8.7 that was missing: Quota Share Engine (diegosouzapw#2859/diegosouzapw#3022/diegosouzapw#3032), page redesigns (diegosouzapw#2827/diegosouzapw#2839/diegosouzapw#2847/diegosouzapw#2849/diegosouzapw#2869/diegosouzapw#2873), and fixes diegosouzapw#2960/diegosouzapw#2973/diegosouzapw#2984/ diegosouzapw#3021/diegosouzapw#3029/diegosouzapw#3031/diegosouzapw#3035/diegosouzapw#3036/diegosouzapw#3037/diegosouzapw#3039/diegosouzapw#3043/diegosouzapw#3028; folded diegosouzapw#2978/diegosouzapw#2988/diegosouzapw#3041 - insert [3.8.8] section into all 41 i18n CHANGELOGs + sync llm.txt mirrors - document OMNIROUTE_PLUGINS_ALLOW_EXEC in .env.example + ENVIRONMENT.md (env-doc-sync gap)
…souzapw#2973) Integrated into release/v3.8.8. OOM fix: truncateForLog caps logged bodies at 8KB (prevents multi-MB clone accumulation across log call-sites) + a heap-pressure 503 guard. Applied review fix: the 503 body no longer leaks the heap figure (Hard Rule diegosouzapw#12) — logged internally instead. Thanks @soyelmismo!
…env-doc fixes - bump package.json / open-sse / electron / openapi / llm.txt to 3.8.8 - restructure CHANGELOG: Unreleased -> [3.8.8], dedup broken Notion/MCP block (was 12x) - add every PR since v3.8.7 that was missing: Quota Share Engine (diegosouzapw#2859/diegosouzapw#3022/diegosouzapw#3032), page redesigns (diegosouzapw#2827/diegosouzapw#2839/diegosouzapw#2847/diegosouzapw#2849/diegosouzapw#2869/diegosouzapw#2873), and fixes diegosouzapw#2960/diegosouzapw#2973/diegosouzapw#2984/ diegosouzapw#3021/diegosouzapw#3029/diegosouzapw#3031/diegosouzapw#3035/diegosouzapw#3036/diegosouzapw#3037/diegosouzapw#3039/diegosouzapw#3043/diegosouzapw#3028; folded diegosouzapw#2978/diegosouzapw#2988/diegosouzapw#3041 - insert [3.8.8] section into all 41 i18n CHANGELOGs + sync llm.txt mirrors - document OMNIROUTE_PLUGINS_ALLOW_EXEC in .env.example + ENVIRONMENT.md (env-doc-sync gap)

Problem
OOM crashes within 5 minutes of intensive use. CPU climbs like a staircase (0.1% → 2-3%) then process dies. Larger contexts cause faster OOM.
Root Cause
persistAttemptLogsinchatCore.tsdeep-clonestranslatedBody(1-5MB with large contexts) 17 times per request viacloneBoundedChatLogPayload. With 5 concurrent requests that's 85-250MB of clones alone — exceeding the 256MB heap.Fix (2 changes in
chatCore.ts)truncateForLog()— caps request/response bodies logged bypersistAttemptLogsat 8KB JSON. Returns a lightweight summary (model, provider, message count) instead of a full deep clone when the body exceeds 8KB.Memory pressure guard — checks
process.memoryUsage().heapUsedat the top ofhandleChatCore. Returns HTTP 503 when heap exceedsHEAP_PRESSURE_THRESHOLD_MB(default 200MB, configurable via env). Self-healing: no counters to leak, no timers to clean up.Scope
open-sse/handlers/chatCore.ts(+57, -2)Not in this PR
The previous PR #2965 addressed cache eviction (comboMetrics, usage, providerRegistry). This PR addresses the per-request memory accumulation that was the second OOM vector.