Skip to content

chore(perf): add deterministic request-body heap benchmark (#7847) - #8549

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.49from
MumuTW:chore/heap-benchmark-request-body
Jul 26, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.49from
MumuTW:chore/heap-benchmark-request-body

Conversation

@MumuTW

@MumuTW MumuTW commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

Groundwork for #7847. Reports only — no production code changes, nothing wired into CI.

Why first

#7847 asks for "a regression benchmark that records peak heap for representative 500–800-message, tool-rich requests". The repo currently has no memory baseline at all (bench:compression is the only benchmark), so a clone-reduction change could neither be justified nor regression-guarded. This lands the measurement before the fix.

What it measures

npm run bench:heap-body reproduces the incident shape (3.05 MiB / 729 messages / 86 tools) and attributes retained V8 heap to each copy the chat path makes:

mechanism call site retained x wire
cloneLogPayload (unbounded) chat.ts buildClientRawRequest 3.18 MiB 1.04x
cloneBoundedForLog (bounded) requestLogger.logClientRawRequest 0.04 MiB 0.01x
structuredClone x3 (combo targets) combo.ts attemptBody 9.53 MiB 3.12x
JSON.stringify (token estimate) combo.ts estimateTokens 3.06 MiB 1.00x
per request (sum) 15.81 MiB 5.17x

8 concurrent requests retain 25.42 MiB via the entry clone alone.

It calls the real production helpers, not reimplementations, so changing the log bounds or the clone strategy shows up directly in the numbers.

What the first numbers already say

Two things worth flagging for whoever picks up the fix:

  1. The entry clone is provably oversized. buildClientRawRequest retains 3.18 MiB; everything downstream of it either discards the value (logClientRawRequest is a no-op when the logger is disabled) or re-clones it bounded into 0.04 MiB. That is ~80x more than any consumer keeps. I traced all three consumers (logClientRawRequest, trackPendingRequests clientRequest, recordRejectedRequestUsages requestBody) — every one is observability; none feeds dispatch, translation or the upstream request.

  2. But it is not the biggest mechanism. combo.ts per-target attemptBody clones cost 9.53 MiB at only 3 targets — 3x the entry clone. Any ordering that starts with the entry clone should be honest that it is the cheapest and safest win, not the largest one.

Design notes

  • Deterministic — fixed-seed LCG, no Math.random(). Verified byte-identical across three consecutive runs; without that a before/after delta measures noise rather than the change.
  • Corpus is a separate side-effect-free module so the unit test can import it without booting SQLite (requestLogger transitively opens the DB at import time).
  • Hermetic — DATA_DIR is redirected to a temp dir before importing, so the benchmark never touches the operator real ~/.omniroute store.
  • Node, not bun — --expose-gc and V8 heap accounting are the measurement; another engine heap number would not describe the production runtime. (Per CLAUDE.md, bun stays limited to its allow-listed gate/generator scripts.)
  • --max-retained-mib exits non-zero, so this can become a CI gate once a target is agreed. Left out of CI deliberately for now — there is no agreed budget yet.

Verification

  • tests/unit/heap-benchmark-corpus.test.ts — 5/5 pass: byte-stability across runs and across module instances, incident wire size (~3.05 MiB), agent-request shape, and that the size knobs actually scale
  • typecheck:core, check:docs-sync, check:any-budget:t11 — clean
  • scripts/ is eslint-ignored (same as the existing bench:compression); the new test file is linted and error-free

Inherited base-red (not from this PR)

release/v3.8.49 is red on two gates this branch does not touch: lint (stale suppression count, fixed by #8544) and file-size (providers/page.tsx, tokenHealthCheck.ts — fixed by #8532 / #8524).

@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for this — verified it end to end (checked out the branch, ran the test suite and the benchmark itself).

Confirmed:

  • tests/unit/heap-benchmark-corpus.test.ts passes 5/5, and it's correctly picked up by the tests/unit/*.test.ts glob in test:unit — not orphaned.
  • npm run bench:heap-body reproduces essentially the same numbers as your table (3.18 / 0.06 / 9.53 / 3.06 MiB, ~15.8 MiB per request, ~25.4 MiB at 8-way concurrency) — good determinism.
  • cloneLogPayload, cloneBoundedForLog, buildClientRawRequest, and the combo.ts attemptBody/estimateTokens call sites are exactly what's called in production today — this is measuring the real path, not a model of it.
  • scripts/** is eslint-ignored as you said (matches bench:compression), and neither bench:heap-body nor bench:compression appears in any CI workflow — this is purely on-demand, no CI risk.
  • typecheck:core, any-budget, docs-sync all clean.
  • The file-size and mutation-test-coverage gate failures I saw locally are pre-existing on the release tip and don't touch any file in this diff — confirmed with git diff against the two flagged files, they're untouched here.

One small, non-blocking note for later: the migration bootstrap logs ("[Migration] Skipped executing ...") still leak to stdout when running npm run bench:heap-body, despite the comment saying console output is fully parked during the import — cosmetic only, doesn't affect any of the reported numbers.

Also worth knowing (not something to change here): #8550, #8553 and #8558 are already open and map 1:1 to the three mechanisms in your table, so this reads as the intended groundwork for that sequence rather than a standalone tool.

…apw#7847)

diegosouzapw#7847 reports a 3.05 MiB request (729 messages / 86 tools) reaching ~12,282 MiB of V8
heap, and asks for "a regression benchmark that records peak heap for representative
500-800-message, tool-rich requests" before any fix lands. There is currently no memory
baseline in the repo at all (bench:compression is the only benchmark), so a clone-reduction
change could neither be justified nor regression-guarded.

npm run bench:heap-body attributes retained heap to each copy the chat path makes:

  | mechanism                          | call site                          | retained | x wire |
  | cloneLogPayload (unbounded)        | chat.ts buildClientRawRequest      | 3.18 MiB |  1.04x |
  | cloneBoundedForLog (bounded)       | requestLogger.logClientRawRequest  | 0.04 MiB |  0.01x |
  | structuredClone x3 (combo targets) | combo.ts attemptBody               | 9.53 MiB |  3.12x |
  | JSON.stringify (token estimate)    | combo.ts estimateTokens            | 3.06 MiB |  1.00x |
  | per request (sum)                  |                                    |15.81 MiB |  5.17x |

It measures the real production helpers rather than reimplementations, so a change to the
log bounds or the clone strategy is reflected directly.

Design notes:
- Deterministic: fixed-seed LCG, no Math.random(). Verified byte-identical across three
  consecutive runs — without that, a before/after delta measures noise, not the change.
- Corpus lives in its own side-effect-free module so the unit test can import it without
  booting SQLite (requestLogger transitively opens the DB at import time).
- Hermetic: DATA_DIR is redirected to a temp dir before importing, so the benchmark never
  touches the operator's real ~/.omniroute store.
- Node, not bun: --expose-gc and V8 heap accounting are the measurement; another engine's
  heap number would not describe the production runtime.
- --max-retained-mib exits non-zero, so this can become a CI gate once a target is agreed.

Reports only; wires nothing into CI and changes no production code.
@MumuTW
MumuTW force-pushed the chore/heap-benchmark-request-body branch from 2e15ff7 to 21c81b4 Compare July 25, 2026 14:36
MumuTW added a commit to MumuTW/OmniRoute that referenced this pull request Jul 25, 2026
…ut building it (diegosouzapw#7847)

Two changes with one root cause: several hot paths built a full JSON string only to read
its .length, and one of them silently changed the answer.

1. CORRECTNESS -- combo's fallback-compression trigger

   estimateTokens(JSON.stringify(attemptBody)) took the STRING branch of estimateTokens,
   which is ceil(length / CHARS_PER_TOKEN) over the raw JSON. An inline base64 image is
   then charged as if every character of the data URL were prose. Measured on a 200 KB
   inline image:

     via string (before)  50,039 tokens
     via object (after)    1,231 tokens

   a 40x over-count, tripping fallback compression on requests nowhere near the context
   window. This is the same class diegosouzapw#8368/diegosouzapw#8401 fixed on the request path; the combo call
   site was missed. Passing the object routes through extractImageTokens, which charges
   images structurally. Text-only bodies are unaffected -- verified identical, and pinned
   by a test.

2. ALLOCATION -- jsonLength()

   Adds an exact serialized-length walker: same O(n) scan, no string. Used by
   estimateTokens' object branch and by streamReadinessPolicy (which runs on every
   streaming request and only ever used .length).

   Exactness matters because every consumer feeds a threshold, so this is property-tested
   against JSON.stringify over 4000 generated structures covering escaping, lone
   surrogates, omitted values, non-finite numbers, toJSON, Date, Map, cycles and BigInt.
   Anything outside the plain-JSON subset falls back to JSON.stringify for THAT SUBTREE
   only, so an exotic leaf never forces the message history back onto the allocating path.

Honest scoping of the memory win: the string was always transient, and V8 collects it
efficiently, so this is not 3 MiB of retained heap. Measured allocation churn over 20 calls
on a 3.06 MiB body: 3.1 MiB -> 0.5 MiB, about 6x less. The diegosouzapw#8549 benchmark row for this
mechanism measures a HELD string and therefore overstates it; the correctness fix above is
the larger deliverable here.
@MumuTW

MumuTW commented Jul 25, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto 4053e2314 — the file-size and mutation-coverage reds you saw locally are fixed upstream now (#8561 rebaselined check:file-size, #8538 the stryker registration), alongside #8544 and #8534. Re-verified on the rebased head: heap-benchmark-corpus.test.ts 5/5, check:file-size OK.

On the migration-log leak: confirmed, the console parking doesn't cover the migration bootstrap's own writes. It's outside the measured region so none of the reported numbers move, and I left it rather than widening a benchmark-only PR — say the word if you'd rather it not ship noisy and I'll fold in the fix.

And yes, #8550/#8553/#8558 are the three mechanisms from the table — this was the groundwork, not a standalone tool.

@MumuTW

MumuTW commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor Author

Correction to my comment above: Fast Quality Gates will still be red after the rebase, and it isn't this PR. #8561 fixed check:file-size and thereby unmasked a fifth base-red gate behind it in the same job — check:complexity-ratchets (complexity 2169 > baseline 2130, cognitiveComplexity 956 > 951).

I measured it on pristine detached checkouts with an empty working tree: 4053e2314 (current tip) and 30709255c (the base you reviewed against) both report 2169 / 956 — identical to this PR's head, so it predates #8561 and is not caused by anything here. It was hidden because check:file-size ran earlier in the same bash -e job and short-circuited it. Full measurement table in my comment on #8546.

The gate is in quality.yml's fast-gates job, so it's red for every PR against release/v3.8.49 regardless of content. Everything else on this PR is green and verified locally.

diegosouzapw pushed a commit that referenced this pull request Jul 26, 2026
…ut building it (#7847) (#8558)

Two changes with one root cause: several hot paths built a full JSON string only to read
its .length, and one of them silently changed the answer.

1. CORRECTNESS -- combo's fallback-compression trigger

   estimateTokens(JSON.stringify(attemptBody)) took the STRING branch of estimateTokens,
   which is ceil(length / CHARS_PER_TOKEN) over the raw JSON. An inline base64 image is
   then charged as if every character of the data URL were prose. Measured on a 200 KB
   inline image:

     via string (before)  50,039 tokens
     via object (after)    1,231 tokens

   a 40x over-count, tripping fallback compression on requests nowhere near the context
   window. This is the same class #8368/#8401 fixed on the request path; the combo call
   site was missed. Passing the object routes through extractImageTokens, which charges
   images structurally. Text-only bodies are unaffected -- verified identical, and pinned
   by a test.

2. ALLOCATION -- jsonLength()

   Adds an exact serialized-length walker: same O(n) scan, no string. Used by
   estimateTokens' object branch and by streamReadinessPolicy (which runs on every
   streaming request and only ever used .length).

   Exactness matters because every consumer feeds a threshold, so this is property-tested
   against JSON.stringify over 4000 generated structures covering escaping, lone
   surrogates, omitted values, non-finite numbers, toJSON, Date, Map, cycles and BigInt.
   Anything outside the plain-JSON subset falls back to JSON.stringify for THAT SUBTREE
   only, so an exotic leaf never forces the message history back onto the allocating path.

Honest scoping of the memory win: the string was always transient, and V8 collects it
efficiently, so this is not 3 MiB of retained heap. Measured allocation churn over 20 calls
on a 3.06 MiB body: 3.1 MiB -> 0.5 MiB, about 6x less. The #8549 benchmark row for this
mechanism measures a HELD string and therefore overstates it; the correctness fix above is
the larger deliverable here.
@diegosouzapw
diegosouzapw merged commit 4bf47c9 into diegosouzapw:release/v3.8.49 Jul 26, 2026
5 checks passed
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @MumuTW — merged into release/v3.8.49 via the local merge-train (validated as one combined tree: full test:unit + test:vitest 274/274 on the 32-core box, tip d4b9ce6016). Your commit keeps its authorship. 🚀

HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…ut building it (diegosouzapw#7847) (diegosouzapw#8558)

Two changes with one root cause: several hot paths built a full JSON string only to read
its .length, and one of them silently changed the answer.

1. CORRECTNESS -- combo's fallback-compression trigger

   estimateTokens(JSON.stringify(attemptBody)) took the STRING branch of estimateTokens,
   which is ceil(length / CHARS_PER_TOKEN) over the raw JSON. An inline base64 image is
   then charged as if every character of the data URL were prose. Measured on a 200 KB
   inline image:

     via string (before)  50,039 tokens
     via object (after)    1,231 tokens

   a 40x over-count, tripping fallback compression on requests nowhere near the context
   window. This is the same class diegosouzapw#8368/diegosouzapw#8401 fixed on the request path; the combo call
   site was missed. Passing the object routes through extractImageTokens, which charges
   images structurally. Text-only bodies are unaffected -- verified identical, and pinned
   by a test.

2. ALLOCATION -- jsonLength()

   Adds an exact serialized-length walker: same O(n) scan, no string. Used by
   estimateTokens' object branch and by streamReadinessPolicy (which runs on every
   streaming request and only ever used .length).

   Exactness matters because every consumer feeds a threshold, so this is property-tested
   against JSON.stringify over 4000 generated structures covering escaping, lone
   surrogates, omitted values, non-finite numbers, toJSON, Date, Map, cycles and BigInt.
   Anything outside the plain-JSON subset falls back to JSON.stringify for THAT SUBTREE
   only, so an exotic leaf never forces the message history back onto the allocating path.

Honest scoping of the memory win: the string was always transient, and V8 collects it
efficiently, so this is not 3 MiB of retained heap. Measured allocation churn over 20 calls
on a 3.06 MiB body: 3.1 MiB -> 0.5 MiB, about 6x less. The diegosouzapw#8549 benchmark row for this
mechanism measures a HELD string and therefore overstates it; the correctness fix above is
the larger deliverable here.
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…apw#7847) (diegosouzapw#8549)

diegosouzapw#7847 reports a 3.05 MiB request (729 messages / 86 tools) reaching ~12,282 MiB of V8
heap, and asks for "a regression benchmark that records peak heap for representative
500-800-message, tool-rich requests" before any fix lands. There is currently no memory
baseline in the repo at all (bench:compression is the only benchmark), so a clone-reduction
change could neither be justified nor regression-guarded.

npm run bench:heap-body attributes retained heap to each copy the chat path makes:

  | mechanism                          | call site                          | retained | x wire |
  | cloneLogPayload (unbounded)        | chat.ts buildClientRawRequest      | 3.18 MiB |  1.04x |
  | cloneBoundedForLog (bounded)       | requestLogger.logClientRawRequest  | 0.04 MiB |  0.01x |
  | structuredClone x3 (combo targets) | combo.ts attemptBody               | 9.53 MiB |  3.12x |
  | JSON.stringify (token estimate)    | combo.ts estimateTokens            | 3.06 MiB |  1.00x |
  | per request (sum)                  |                                    |15.81 MiB |  5.17x |

It measures the real production helpers rather than reimplementations, so a change to the
log bounds or the clone strategy is reflected directly.

Design notes:
- Deterministic: fixed-seed LCG, no Math.random(). Verified byte-identical across three
  consecutive runs — without that, a before/after delta measures noise, not the change.
- Corpus lives in its own side-effect-free module so the unit test can import it without
  booting SQLite (requestLogger transitively opens the DB at import time).
- Hermetic: DATA_DIR is redirected to a temp dir before importing, so the benchmark never
  touches the operator's real ~/.omniroute store.
- Node, not bun: --expose-gc and V8 heap accounting are the measurement; another engine's
  heap number would not describe the production runtime.
- --max-retained-mib exits non-zero, so this can become a CI gate once a target is agreed.

Reports only; wires nothing into CI and changes no production code.
@MumuTW
MumuTW deleted the chore/heap-benchmark-request-body branch September 5, 2026 10:18
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ut building it (diegosouzapw#7847) (diegosouzapw#8558)

Two changes with one root cause: several hot paths built a full JSON string only to read
its .length, and one of them silently changed the answer.

1. CORRECTNESS -- combo's fallback-compression trigger

   estimateTokens(JSON.stringify(attemptBody)) took the STRING branch of estimateTokens,
   which is ceil(length / CHARS_PER_TOKEN) over the raw JSON. An inline base64 image is
   then charged as if every character of the data URL were prose. Measured on a 200 KB
   inline image:

     via string (before)  50,039 tokens
     via object (after)    1,231 tokens

   a 40x over-count, tripping fallback compression on requests nowhere near the context
   window. This is the same class diegosouzapw#8368/diegosouzapw#8401 fixed on the request path; the combo call
   site was missed. Passing the object routes through extractImageTokens, which charges
   images structurally. Text-only bodies are unaffected -- verified identical, and pinned
   by a test.

2. ALLOCATION -- jsonLength()

   Adds an exact serialized-length walker: same O(n) scan, no string. Used by
   estimateTokens' object branch and by streamReadinessPolicy (which runs on every
   streaming request and only ever used .length).

   Exactness matters because every consumer feeds a threshold, so this is property-tested
   against JSON.stringify over 4000 generated structures covering escaping, lone
   surrogates, omitted values, non-finite numbers, toJSON, Date, Map, cycles and BigInt.
   Anything outside the plain-JSON subset falls back to JSON.stringify for THAT SUBTREE
   only, so an exotic leaf never forces the message history back onto the allocating path.

Honest scoping of the memory win: the string was always transient, and V8 collects it
efficiently, so this is not 3 MiB of retained heap. Measured allocation churn over 20 calls
on a 3.06 MiB body: 3.1 MiB -> 0.5 MiB, about 6x less. The diegosouzapw#8549 benchmark row for this
mechanism measures a HELD string and therefore overstates it; the correctness fix above is
the larger deliverable here.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…apw#7847) (diegosouzapw#8549)

diegosouzapw#7847 reports a 3.05 MiB request (729 messages / 86 tools) reaching ~12,282 MiB of V8
heap, and asks for "a regression benchmark that records peak heap for representative
500-800-message, tool-rich requests" before any fix lands. There is currently no memory
baseline in the repo at all (bench:compression is the only benchmark), so a clone-reduction
change could neither be justified nor regression-guarded.

npm run bench:heap-body attributes retained heap to each copy the chat path makes:

  | mechanism                          | call site                          | retained | x wire |
  | cloneLogPayload (unbounded)        | chat.ts buildClientRawRequest      | 3.18 MiB |  1.04x |
  | cloneBoundedForLog (bounded)       | requestLogger.logClientRawRequest  | 0.04 MiB |  0.01x |
  | structuredClone x3 (combo targets) | combo.ts attemptBody               | 9.53 MiB |  3.12x |
  | JSON.stringify (token estimate)    | combo.ts estimateTokens            | 3.06 MiB |  1.00x |
  | per request (sum)                  |                                    |15.81 MiB |  5.17x |

It measures the real production helpers rather than reimplementations, so a change to the
log bounds or the clone strategy is reflected directly.

Design notes:
- Deterministic: fixed-seed LCG, no Math.random(). Verified byte-identical across three
  consecutive runs — without that, a before/after delta measures noise, not the change.
- Corpus lives in its own side-effect-free module so the unit test can import it without
  booting SQLite (requestLogger transitively opens the DB at import time).
- Hermetic: DATA_DIR is redirected to a temp dir before importing, so the benchmark never
  touches the operator's real ~/.omniroute store.
- Node, not bun: --expose-gc and V8 heap accounting are the measurement; another engine's
  heap number would not describe the production runtime.
- --max-retained-mib exits non-zero, so this can become a CI gate once a target is agreed.

Reports only; wires nothing into CI and changes no production code.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants