Repository navigation
feat(compression): update vendored GCF (Headroom) codec to spec v3.2 — nested flattening - #6838
Conversation
…(nested flattening)
Homogeneous arrays whose rows carry nested objects/arrays now tabularize
via GCF v3.2 `>`-path flattening instead of a low-yield per-row fallback,
so nested MCP tool-result rows (meta:{...}, tags:[...]) compact like flat
rows. Round-trip stays lossless (order-insensitive deepEqual).
Re-vendored from current gcf-typescript into the Headroom generic-profile
codec (open-sse/services/compression/engines/headroom/gcf/); still zero
runtime deps, MIT, SPDX-marked, generic-profile only. Also folds in two
upstream round-trip-safety fixes: the [N]: inline-array quoting fix and
canonical decimal formatting.
Regression guard: tests/unit/compression/headroom-smartcrusher.test.ts
gains a deep-nested case (two-level object + array-of-objects) asserting
the v3.2 flatten paths and order-insensitive round-trip. Vendored-code
baseline bumps (complexity 2053->2055, cognitive 885->888, decode_generic
no-explicit-any 18->22) each carry an inline _rebaseline_2026_07_10_gcf_v3_2
justification noting the growth is the vendored surface, not new project code.
There was a problem hiding this comment.
Code Review
This pull request updates the vendored GCF (Graph Compact Format) codec to spec v3.2, introducing support for nested object flattening and unflattening using ">"-separated path columns. The review feedback correctly identifies critical prototype pollution vulnerabilities and prototype-chain lookup bugs in both the decoder's unflattenPaths function and the encoder's analyzeFlattenable function. Specifically, using the in operator allows checking inherited properties, and the lack of validation on keys like __proto__, constructor, and prototype can lead to prototype pollution or runtime errors. Addressing these issues is essential to prevent security risks and ensure correct shape matching.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
…ngine-stack table
…ngine-stack table
…tion The v3.2 flatten/unflatten paths (and the pre-existing inline-object parser) built decoded objects with bracket assignment and `key in obj` membership, so a hostile or unusual payload could pollute Object.prototype via a `__proto__` path segment, and any key shadowing an Object.prototype member (`toString`, `constructor`) was misparsed or wrongly flagged duplicate. - Encoder (`analyzeFlattenable`): builds the shape map with `Object.create(null)` and refuses to flatten objects carrying `__proto__`/`constructor`/`prototype` keys (they round-trip whole instead). - Decoder: `unflattenPaths` drops any path with an unsafe segment; a shared `safeAssign` writes a literal `__proto__` key as an own data property (JSON.parse semantics) instead of reassigning the prototype, used at every object-build site; `checkDup` and orphan-merge use `hasOwnProperty` so built-in-named keys are not spuriously treated as duplicates. Also a losslessness fix: objects with keys named `toString`/`constructor`/ `valueOf` now round-trip. Regression guard: prototype-pollution + built-in-key cases in tests/unit/compression/headroom-smartcrusher.test.ts. Prototype pollution is JS/TS-specific; the Go/Python/Rust/Swift/Kotlin SDKs use native maps and are unaffected.
|
/gemini review |
There was a problem hiding this comment.
Code Review
This pull request updates the vendored GCF (Graph Compact Format) codec behind the Headroom compression engine to spec v3.2, enabling nested object flattening and hardening the decoder against prototype pollution. The review feedback identifies several important opportunities to improve robustness and security: guarding the unflattening logic against runtime crashes on malformed inputs, replacing the in operator with Object.prototype.hasOwnProperty.call to prevent prototype chain lookup bugs, resolving a redundant variable shadowing issue in the encoder, and using the strict parseCount helper instead of parseInt for consistent count validation.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
…sweep, unflatten null-guard, strict count) Addresses the second-round review on the vendored codec: - Replace every `key in obj` membership test with `Object.prototype.hasOwnProperty.call(...)` across generic.ts (flatten shape analysis, key-chain resolution, inline-schema/shared-array helpers, row encode) so inherited names (`toString`/`constructor`) never match the prototype chain, and remove a redundant `obj` re-declaration in the ">" field attachment loop. - `unflattenPaths` guards each intermediate segment: a missing OR non-object slot is replaced with a fresh object before traversal, so malformed/hostile input can no longer dereference a primitive and crash. - Use the strict `parseCount` helper (not `parseInt`) for the shared-schema count so malformed counts fail the mismatch check instead of coercing. The decoder grew past the 800-line file-size cap; frozen at 880 in file-size-baseline.json with a justification (vendored file kept faithful to upstream gcf-typescript for clean re-vendoring). Verified: prototype-pollution + hostile-input + built-in-key round-trip probes, 54/54 compression tests, typecheck, lint, cyclomatic/cognitive baselines unchanged, compression-budget.
|
/gemini review |
There was a problem hiding this comment.
Code Review
This pull request updates the vendored GCF (Graph Compact Format) codec behind the Headroom compression engine to spec v3.2, introducing nested object flattening via ">"-prefixed path fields and hardening the decoder against prototype pollution. The quality, complexity, and file-size baselines are updated accordingly. Feedback on the changes highlights a high-severity issue in the shape analysis during flattening: if a nested object is null in some rows, it is still flattened, leading to silent data loss during unflattening. A code suggestion is provided to reject flattening if the nested object is null in any row.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
…row (losslessness)
analyzeFlattenable skipped null values during shape analysis, so a field that
was an object in some rows and null in others was still flattened. On decode,
the null row's leaves resolved as absent ("~") and unflattened to a missing key
instead of null, silently dropping the value (e.g. {meta:{owner:null}} decoded
to {}). analyzeFlattenable now bails (returns null) when the field is null in
any row, routing it through the lossless whole-object attachment path. Applies
at every nesting depth via the existing recursion. Regression guard: null
nested-object cases in tests/unit/compression/headroom-smartcrusher.test.ts.
|
/gemini review |
There was a problem hiding this comment.
Code Review
This pull request updates the vendored GCF (Graph Compact Format) codec behind the Headroom compression engine to spec v3.2, enabling nested object flattening via >-prefixed path columns. The update also includes inline-array quoting fixes, canonical decimal formatting, and robust hardening against prototype pollution. Quality baselines for complexity, file size, and ESLint suppressions have been adjusted to accommodate the updated vendored code, and comprehensive unit tests have been added to verify nested round-tripping and prototype-pollution safety. No review comments were provided, so there is no feedback to evaluate.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
…iRoute into feat/headroom-gcf-v3.2-nested-flattening # Conflicts: # config/quality/complexity-baseline.json
… nulls only The previous fix bailed flattening whenever a nested field was null in any row. That is correct but over-broad: a top-level null round-trips losslessly through flattening (it emits "-" and reconstructs via the all-null rule). Only a null at an intermediate nesting level loses data (its leaves encode as absent "~" and unflatten to a missing key). Bail only when parentPath is non-empty, so top-level nulls keep flattening (compression preserved) while intermediate nulls fall back to the lossless attachment path. Matches GCF conformance fixtures 004/013.
|
Verified: check:complexity, check:cognitive-complexity, and check:compression-budget gates all reproduce cleanly against the claimed numbers (cognitive-complexity matches exactly: 888/888). The prototype-pollution hardening is real and covered by 4 dedicated hostile-input round-trip tests, not just the nested-flatten feature. Whole-repo lint and typecheck are clean. Nice, tightly-scoped vendored-codec update — looks merge-ready. |
Resync onto the current release tip; resolves the quality-baseline conflict by keeping the release's cognitiveComplexity=890 rebaseline, which already supersedes and covers this PR's own +3 growth to 888. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Merged — re-synced onto the release tip (cognitiveComplexity baseline reconciled to 890, which already covers this PR's growth). GCF codec re-vendor validated, 35/35 tests green. Thanks @blackwell-systems! (Remaining red checks are pre-existing base-reds on the release tip — openadapter live-catalog tests + a stale eslint suppression — unrelated to this change.) |
…— nested flattening (diegosouzapw#6838) * feat(compression): update vendored GCF (Headroom) codec to spec v3.2 (nested flattening) Homogeneous arrays whose rows carry nested objects/arrays now tabularize via GCF v3.2 `>`-path flattening instead of a low-yield per-row fallback, so nested MCP tool-result rows (meta:{...}, tags:[...]) compact like flat rows. Round-trip stays lossless (order-insensitive deepEqual). Re-vendored from current gcf-typescript into the Headroom generic-profile codec (open-sse/services/compression/engines/headroom/gcf/); still zero runtime deps, MIT, SPDX-marked, generic-profile only. Also folds in two upstream round-trip-safety fixes: the [N]: inline-array quoting fix and canonical decimal formatting. Regression guard: tests/unit/compression/headroom-smartcrusher.test.ts gains a deep-nested case (two-level object + array-of-objects) asserting the v3.2 flatten paths and order-insensitive round-trip. Vendored-code baseline bumps (complexity 2053->2055, cognitive 885->888, decode_generic no-explicit-any 18->22) each carry an inline _rebaseline_2026_07_10_gcf_v3_2 justification noting the growth is the vendored surface, not new project code. * chore(changelog): add fragment for headroom GCF v3.2 nested flattening (diegosouzapw#6838) * docs(readme): note Headroom handles nested arrays (GCF v3.2) in the engine-stack table * docs(readme): note Headroom handles nested arrays (GCF v3.2) in the engine-stack table * fix(compression): harden vendored GCF decoder against prototype pollution The v3.2 flatten/unflatten paths (and the pre-existing inline-object parser) built decoded objects with bracket assignment and `key in obj` membership, so a hostile or unusual payload could pollute Object.prototype via a `__proto__` path segment, and any key shadowing an Object.prototype member (`toString`, `constructor`) was misparsed or wrongly flagged duplicate. - Encoder (`analyzeFlattenable`): builds the shape map with `Object.create(null)` and refuses to flatten objects carrying `__proto__`/`constructor`/`prototype` keys (they round-trip whole instead). - Decoder: `unflattenPaths` drops any path with an unsafe segment; a shared `safeAssign` writes a literal `__proto__` key as an own data property (JSON.parse semantics) instead of reassigning the prototype, used at every object-build site; `checkDup` and orphan-merge use `hasOwnProperty` so built-in-named keys are not spuriously treated as duplicates. Also a losslessness fix: objects with keys named `toString`/`constructor`/ `valueOf` now round-trip. Regression guard: prototype-pollution + built-in-key cases in tests/unit/compression/headroom-smartcrusher.test.ts. Prototype pollution is JS/TS-specific; the Go/Python/Rust/Swift/Kotlin SDKs use native maps and are unaffected. * fix(compression): apply GCF decoder review hardening (hasOwnProperty sweep, unflatten null-guard, strict count) Addresses the second-round review on the vendored codec: - Replace every `key in obj` membership test with `Object.prototype.hasOwnProperty.call(...)` across generic.ts (flatten shape analysis, key-chain resolution, inline-schema/shared-array helpers, row encode) so inherited names (`toString`/`constructor`) never match the prototype chain, and remove a redundant `obj` re-declaration in the ">" field attachment loop. - `unflattenPaths` guards each intermediate segment: a missing OR non-object slot is replaced with a fresh object before traversal, so malformed/hostile input can no longer dereference a primitive and crash. - Use the strict `parseCount` helper (not `parseInt`) for the shared-schema count so malformed counts fail the mismatch check instead of coercing. The decoder grew past the 800-line file-size cap; frozen at 880 in file-size-baseline.json with a justification (vendored file kept faithful to upstream gcf-typescript for clean re-vendoring). Verified: prototype-pollution + hostile-input + built-in-key round-trip probes, 54/54 compression tests, typecheck, lint, cyclomatic/cognitive baselines unchanged, compression-budget. * fix(compression): do not flatten a nested object that is null in any row (losslessness) analyzeFlattenable skipped null values during shape analysis, so a field that was an object in some rows and null in others was still flattened. On decode, the null row's leaves resolved as absent ("~") and unflattened to a missing key instead of null, silently dropping the value (e.g. {meta:{owner:null}} decoded to {}). analyzeFlattenable now bails (returns null) when the field is null in any row, routing it through the lossless whole-object attachment path. Applies at every nesting depth via the existing recursion. Regression guard: null nested-object cases in tests/unit/compression/headroom-smartcrusher.test.ts. * fix(compression): narrow the null-nested flatten bail to intermediate nulls only The previous fix bailed flattening whenever a nested field was null in any row. That is correct but over-broad: a top-level null round-trips losslessly through flattening (it emits "-" and reconstructs via the all-null rule). Only a null at an intermediate nesting level loses data (its leaves encode as absent "~" and unflatten to a missing key). Bail only when parentPath is non-empty, so top-level nulls keep flattening (compression preserved) while intermediate nulls fall back to the lossless attachment path. Matches GCF conformance fixtures 004/013. --------- Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com>
…— nested flattening (diegosouzapw#6838) * feat(compression): update vendored GCF (Headroom) codec to spec v3.2 (nested flattening) Homogeneous arrays whose rows carry nested objects/arrays now tabularize via GCF v3.2 `>`-path flattening instead of a low-yield per-row fallback, so nested MCP tool-result rows (meta:{...}, tags:[...]) compact like flat rows. Round-trip stays lossless (order-insensitive deepEqual). Re-vendored from current gcf-typescript into the Headroom generic-profile codec (open-sse/services/compression/engines/headroom/gcf/); still zero runtime deps, MIT, SPDX-marked, generic-profile only. Also folds in two upstream round-trip-safety fixes: the [N]: inline-array quoting fix and canonical decimal formatting. Regression guard: tests/unit/compression/headroom-smartcrusher.test.ts gains a deep-nested case (two-level object + array-of-objects) asserting the v3.2 flatten paths and order-insensitive round-trip. Vendored-code baseline bumps (complexity 2053->2055, cognitive 885->888, decode_generic no-explicit-any 18->22) each carry an inline _rebaseline_2026_07_10_gcf_v3_2 justification noting the growth is the vendored surface, not new project code. * chore(changelog): add fragment for headroom GCF v3.2 nested flattening (diegosouzapw#6838) * docs(readme): note Headroom handles nested arrays (GCF v3.2) in the engine-stack table * docs(readme): note Headroom handles nested arrays (GCF v3.2) in the engine-stack table * fix(compression): harden vendored GCF decoder against prototype pollution The v3.2 flatten/unflatten paths (and the pre-existing inline-object parser) built decoded objects with bracket assignment and `key in obj` membership, so a hostile or unusual payload could pollute Object.prototype via a `__proto__` path segment, and any key shadowing an Object.prototype member (`toString`, `constructor`) was misparsed or wrongly flagged duplicate. - Encoder (`analyzeFlattenable`): builds the shape map with `Object.create(null)` and refuses to flatten objects carrying `__proto__`/`constructor`/`prototype` keys (they round-trip whole instead). - Decoder: `unflattenPaths` drops any path with an unsafe segment; a shared `safeAssign` writes a literal `__proto__` key as an own data property (JSON.parse semantics) instead of reassigning the prototype, used at every object-build site; `checkDup` and orphan-merge use `hasOwnProperty` so built-in-named keys are not spuriously treated as duplicates. Also a losslessness fix: objects with keys named `toString`/`constructor`/ `valueOf` now round-trip. Regression guard: prototype-pollution + built-in-key cases in tests/unit/compression/headroom-smartcrusher.test.ts. Prototype pollution is JS/TS-specific; the Go/Python/Rust/Swift/Kotlin SDKs use native maps and are unaffected. * fix(compression): apply GCF decoder review hardening (hasOwnProperty sweep, unflatten null-guard, strict count) Addresses the second-round review on the vendored codec: - Replace every `key in obj` membership test with `Object.prototype.hasOwnProperty.call(...)` across generic.ts (flatten shape analysis, key-chain resolution, inline-schema/shared-array helpers, row encode) so inherited names (`toString`/`constructor`) never match the prototype chain, and remove a redundant `obj` re-declaration in the ">" field attachment loop. - `unflattenPaths` guards each intermediate segment: a missing OR non-object slot is replaced with a fresh object before traversal, so malformed/hostile input can no longer dereference a primitive and crash. - Use the strict `parseCount` helper (not `parseInt`) for the shared-schema count so malformed counts fail the mismatch check instead of coercing. The decoder grew past the 800-line file-size cap; frozen at 880 in file-size-baseline.json with a justification (vendored file kept faithful to upstream gcf-typescript for clean re-vendoring). Verified: prototype-pollution + hostile-input + built-in-key round-trip probes, 54/54 compression tests, typecheck, lint, cyclomatic/cognitive baselines unchanged, compression-budget. * fix(compression): do not flatten a nested object that is null in any row (losslessness) analyzeFlattenable skipped null values during shape analysis, so a field that was an object in some rows and null in others was still flattened. On decode, the null row's leaves resolved as absent ("~") and unflattened to a missing key instead of null, silently dropping the value (e.g. {meta:{owner:null}} decoded to {}). analyzeFlattenable now bails (returns null) when the field is null in any row, routing it through the lossless whole-object attachment path. Applies at every nesting depth via the existing recursion. Regression guard: null nested-object cases in tests/unit/compression/headroom-smartcrusher.test.ts. * fix(compression): narrow the null-nested flatten bail to intermediate nulls only The previous fix bailed flattening whenever a nested field was null in any row. That is correct but over-broad: a top-level null round-trips losslessly through flattening (it emits "-" and reconstructs via the all-null rule). Only a null at an intermediate nesting level loses data (its leaves encode as absent "~" and unflatten to a missing key). Bail only when parentPath is non-empty, so top-level nulls keep flattening (compression preserved) while intermediate nulls fall back to the lossless attachment path. Matches GCF conformance fixtures 004/013. --------- Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com>
Summary
Updates the Headroom engine's vendored GCF codec to spec v3.2 (nested flattening). Homogeneous arrays whose rows carry nested objects or arrays now tabularize into the columnar
[N]{fields}form via>-prefixed path fields, instead of a low-yield per-row fallback. Nested MCP tool-result rows (meta:{...},tags:[...]) now compact like flat rows do. Round-trip stays lossless (order-insensitive, matching the existing oracle). Flat-array behavior is unchanged.Measured savings
Headroom's own codec, pre-update vs this PR, on representative nested tool-result shapes at 50 rows. Compact-JSON baseline;
cl100k_base; every new-codec output verified round-trip lossless.The gain scales with nesting depth. Shallow-nested rows the pre-update codec already handled see a small bump; multi-level nested rows it left nearly uncompressed (k8s pods at 2.8% vs JSON sits below the "only compress if it saves" threshold, so Headroom captured ~nothing there) reach ~32% under v3.2.
o200k_baseis within a point of these figures. Flat-array behavior is unchanged.Related Issues
Validation
npm run lint(clean on the codec)npm run typecheck:corecheck:complexity) and cognitive gate (check:cognitive-complexity) green against the v3.8.47 baselinetests/unit/compression/headroom-smartcrusher.test.ts— 30/30npm run test:coverage(60/60/60/60) — confirm via CITests Added Or Updated
tests/unit/compression/headroom-smartcrusher.test.ts— added a deep-nested regression case: rows with a two-level object (meta.owner.name/meta.owner.team) plus an array-of-objects (items[]). Asserts the encoded header emits the v3.2>-prefixed flatten paths (meta>owner>name) and that the codec round-trips order-insensitively (assert.deepEqual). The pre-existing one-level nested round-trip case is retained.Coverage Notes
The codec gains the v3.2 flatten encode/decode paths (~370 lines) under
open-sse/services/compression/engines/headroom/gcf/. The added deep-nested test exercises those branches so coverage on that directory holds at or above the frozen baseline.Reviewer Notes
open-sse/services/compression/engines/headroom/gcf/(scalar.ts,generic.ts,decode_generic.ts,index.ts), re-vendored from current gcf-typescript. Provenance headers and SPDX marks preserved; still zero runtime dependencies. The graph-profile branch stays intentionally out of scope (this vendored build is generic-profile only and throwsgraph_profile_unsupported)._rebaseline_2026_07_10_gcf_v3_2justification:complexity-baseline.json(2053 -> 2055),quality-baseline.jsoncognitive (885 -> 888), andeslint-suppressions.json(decode_generic.tsno-explicit-any 18 -> 22, the vendored decoder'sanyon dynamic JSON reconstruction). Each was re-measured against the pristinerelease/v3.8.47tip.[N]:inline-array quoting fix (prevents a value likename[3]: a,b,cfrom being misparsed) and canonical decimal formatting. Both are strictly round-trip-safety improvements.