Repository navigation
fix(json): unwrap single-element array wrapper for object schemas - #1152
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthrough
ChangesJSON wrapper repair
Estimated code review effort: 3 (Moderate) | ~20 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Warning There were issues while running some tools. Please review the errors and either fix the tool's configuration or disable the tool if it's a critical failure. 🔧 ESLint
ESLint install timed out. The project may have too many dependencies for the sandbox. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Single Commit Policy - COMPLIANTStatus: Policy requirements met • 1 commit • Valid format • Ready for merge 📊 View validation details📝 Commit Details
✅ Validation Results
🤖 Automated validation by NeuroLink Single Commit Enforcement |
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
There was a problem hiding this comment.
Pull request overview
Hardens coerceJsonToSchema against a native-Anthropic structured-output edge case where an object schema response may arrive as a single-element array ([{...}]), causing callers to receive an array-shaped structuredData even when the schema expects an object.
Changes:
- Add a schema-gated unwrap for single-element array wrappers where the lone element validates against the provided schema.
- Add continuous tests covering the unwrap behavior and ensuring array-typed schemas aren’t over-unwrapped.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
src/lib/utils/json/coerce.ts |
Adds single-element array-wrapper unwrapping logic gated by safeParse. |
test/continuous-test-suite-coerce-nested-unwrap.ts |
Adds tests for the array-wrapper unwrap behavior and related invariants. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| } else if ( | ||
| // Single-element array wrapper: for an OBJECT schema, models sometimes | ||
| // return `[{...}]` instead of `{...}` (seen on the native Anthropic | ||
| // path under escaping stress). Unwrap a lone object element and | ||
| // re-validate — the safeParse gate rejects an incorrect unwrap, so an | ||
| // array schema (which validates the array directly above) is untouched. | ||
| Array.isArray(outcome.value) && | ||
| outcome.value.length === 1 && | ||
| outcome.value[0] !== null && | ||
| typeof outcome.value[0] === "object" && | ||
| !Array.isArray(outcome.value[0]) && | ||
| safeParseable.safeParse(outcome.value[0]).success | ||
| ) { | ||
| schemaValid.push({ | ||
| value: outcome.value[0], | ||
| repaired: true, | ||
| truncated: candidate.truncated, | ||
| }); | ||
| } |
| const inner = { | ||
| summary: "wrapped in an array", | ||
| attachment: null, | ||
| }; | ||
| const r = coerceJsonToSchema(JSON.stringify([inner]), schema); | ||
| assertEqual( | ||
| Array.isArray(r?.structuredData), | ||
| false, | ||
| "result is the object, not an array", | ||
| ); | ||
| assertEqual( | ||
| obj(r).summary, | ||
| "wrapped in an array", | ||
| "object content recovered", | ||
| ); |
Tara-ag
left a comment
There was a problem hiding this comment.
Review Summary
Files reviewed: 2
src/lib/utils/json/coerce.tstest/continuous-test-suite-coerce-nested-unwrap.ts
New issues raised this run: 0
Decision: Approve
The fix is narrow, well-gated by safeParse, and correctly avoids over-unwrapping array schemas and multi-element arrays. The new tests cover the stated native-Anthropic [{...}] regression case and the anti-regression cases.
I did not raise any new inline comments because the two substantive points I would have flagged are already covered by the existing Copilot review threads:
- The array-wrapper branch validates
outcome.value[0]beforedeepUnwrapJsonStringsis applied, so a wrapped object that only becomes schema-valid after nested-string unwrapping is still returned as an array. - The test suite lacks coverage for that combined wrapper + nested-string case.
Those are worth addressing, but per review protocol I am not re-raising already-commented points as blocking criteria. No new security, architectural, or backward-compatibility concerns were introduced.
On the native Anthropic path under escaping stress, models sometimes
return `[{...}]` instead of `{...}` for an object schema. Because a
JS array is `typeof 'object'`, coerceJsonToSchema accepted it as
firstValid and returned an array as structuredData, failing the caller's
object-shaped schema.
When a candidate array does not itself satisfy the schema, and it holds
exactly one non-array object element that DOES satisfy it, unwrap to that
element. Gated by safeParse, mirroring the existing nested-string unwrap:
array-typed schemas keep their array, and multi-element arrays are left
untouched (never silently reduced to their first element).
Adds 3 regression tests (coerce-nested-unwrap 9/9). Refs #635.
821d684 to
fbcb9fb
Compare
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
Tara-ag
left a comment
There was a problem hiding this comment.
Review Summary
Files reviewed: 2
src/lib/utils/json/coerce.tstest/continuous-test-suite-coerce-nested-unwrap.ts
New issues raised this run: 0
I analyzed the single-element array unwrap logic added to coerceJsonToSchema. The change is narrowly scoped, gated by safeParse, and correctly preserves array schemas and multi-element arrays. The regression tests cover the motivating native-Anthropic [{...}] case, the no-over-unwrap array-schema case, and the multi-element rejection case.
I noted the two existing unresolved review comments from the automated reviewer regarding the combined wrapper + nested-string-unwrap scenario. Those are already raised and should be addressed by the author; I am not duplicating them here.
No blocking issues (security, backward compatibility, CLAUDE.md critical-rule violations, or 3+ MAJOR issues) were found in this run. Approving.
|
🎉 This PR is included in version 9.88.9 🎉 The release is available on: Your semantic-release bot 📦🚀 |
On the native Anthropic-direct path a max_tokens cut left `structuredData` as a string instead of the schema object. When output is truncated the root brace never closes, so the balanced-span scan walks past it and matches a bracket pair living INSIDE a string value — `[step 1]` in a shell script becomes `["step 1"]` — reported with `truncated: false`. At other cut points nothing parsed at all, coerceJsonToSchema returned null, and the caller kept the raw text. A prefix sweep over a realistic huge-output payload hit the first case at ~22% of cut points and the second at a handful more. coerce: mark candidates that start at the document's first opening bracket and order those first, so the partial real root beats a span scraped from inside a string; flag every candidate truncated when the root never closes; and when an unclosed root yields nothing trustworthy, back off to the last completed structural boundary and repair from there, returning a PARTIAL object rather than degrading to raw text. consumers: add schemaAccepts and gate structuredData on it — a scalar root, or an experimental_output string that is not an exact raw-text echo, is only published when the caller's schema accepts it. neurolink also re-runs recovery when a provider already produced a schema-rejected string. String-root schemas are unaffected. anthropic: forced-json mode drops text blocks because the payload rides in the synthetic tool's input; when the response is cut short that tool call can be missing entirely, leaving an empty completion. Keep the text as a fallback when no synthetic tool call arrived, so a partial object can still be salvaged. final_result still supersedes it. Adds test:coerce-truncation (8 tests) sweeping every truncation point of a huge-output payload. Fixes juspay#1156. Refs juspay#635, juspay#1152.
On the native Anthropic-direct path a max_tokens cut left `structuredData` as a string instead of the schema object. When output is truncated the root brace never closes, so the balanced-span scan walks past it and matches a bracket pair living INSIDE a string value — `[step 1]` in a shell script becomes `["step 1"]` — reported with `truncated: false`. At other cut points nothing parsed at all, coerceJsonToSchema returned null, and the caller kept the raw text. A prefix sweep over a realistic huge-output payload hit the first case at ~22% of cut points and the second at a handful more. coerce: mark candidates that start at the document's first opening bracket and order those first, so the partial real root beats a span scraped from inside a string; flag every candidate truncated when the root never closes; and when an unclosed root yields nothing trustworthy, back off to the last completed structural boundary and repair from there, returning a PARTIAL object rather than degrading to raw text. consumers: add schemaAccepts and gate structuredData on it — a scalar root, or an experimental_output string that is not an exact raw-text echo, is only published when the caller's schema accepts it. neurolink also re-runs recovery when a provider already produced a schema-rejected string. String-root schemas are unaffected. anthropic: forced-json mode drops text blocks because the payload rides in the synthetic tool's input; when the response is cut short that tool call can be missing entirely, leaving an empty completion. Keep the text as a fallback when no synthetic tool call arrived, so a partial object can still be salvaged. final_result still supersedes it. Adds test:coerce-truncation (8 tests) sweeping every truncation point of a huge-output payload. Fixes juspay#1156. Refs juspay#635, juspay#1152.
On the native Anthropic-direct path a max_tokens cut left `structuredData` as a string instead of the schema object. When output is truncated the root brace never closes, so the balanced-span scan walks past it and matches a bracket pair living INSIDE a string value — `[step 1]` in a shell script becomes `["step 1"]` — reported with `truncated: false`. At other cut points nothing parsed at all, coerceJsonToSchema returned null, and the caller kept the raw text. A prefix sweep over a realistic huge-output payload hit the first case at ~22% of cut points and the second at a handful more. coerce: mark candidates that start at the document's first opening bracket and order those first, so the partial real root beats a span scraped from inside a string; flag every candidate truncated when the root never closes; and when an unclosed root yields nothing trustworthy, back off to the last completed structural boundary and repair from there, returning a PARTIAL object rather than degrading to raw text. consumers: add schemaAccepts and gate structuredData on it — a scalar root, or an experimental_output string that is not an exact raw-text echo, is only published when the caller's schema accepts it. neurolink also re-runs recovery when a provider already produced a schema-rejected string. String-root schemas are unaffected. anthropic: forced-json mode drops text blocks because the payload rides in the synthetic tool's input; when the response is cut short that tool call can be missing entirely, leaving an empty completion. Keep the text as a fallback when no synthetic tool call arrived, so a partial object can still be salvaged. final_result still supersedes it. Adds test:coerce-truncation (8 tests) sweeping every truncation point of a huge-output payload. Fixes juspay#1156. Refs juspay#635, juspay#1152.
On the native Anthropic-direct path a max_tokens cut left `structuredData` as a string instead of the schema object. When output is truncated the root brace never closes, so the balanced-span scan walks past it and matches a bracket pair living INSIDE a string value - `[step 1]` in a shell script becomes `["step 1"]` - reported with `truncated: false`. At other cut points nothing parsed at all, coerceJsonToSchema returned null, and the caller kept the raw text. A prefix sweep over a realistic huge-output payload hit the first case at ~22% of cut points and the second at a handful more. coerce: mark candidates that start at the document's first opening bracket and order those first, so the partial real root beats a span scraped from inside a string; flag every candidate truncated when the root never closes; and when an unclosed root yields nothing trustworthy, back off to the last completed structural boundary and repair from there, returning a PARTIAL object rather than degrading to raw text. consumers: add schemaAccepts and gate structuredData on it - a scalar root, or an experimental_output string that is not an exact raw-text echo, is only published when the caller's schema accepts it. neurolink also re-runs recovery when a provider already produced a schema-rejected string. String-root schemas are unaffected. The shared scalar-recovery policy lives in recoverScalarRoot (with ScalarRecoveryDecision in src/lib/types), used by both neurolink.recoverStructuredData and GenerationHandler.coerceTextMode so it cannot drift. anthropic: forced-json mode drops text blocks because the payload rides in the synthetic tool's input; when the response is cut short that tool call can be missing entirely, leaving an empty completion. Keep the text as a fallback when no synthetic tool call arrived, so a partial object can still be salvaged. final_result still supersedes it. Adds test:coerce-truncation (8 tests) sweeping every truncation point of a huge-output payload, with assertion messages restricted to structural cut positions (never recovered payload values) so a genuine failure reports as FAIL, not SKIP. Adds a test:json-e2e cell on the direct Anthropic path that forces a maxTokens cut and asserts structuredData is a plain object with jsonTruncated === true. Fixes juspay#1156. Refs juspay#635, juspay#1152.
On the native Anthropic-direct path a max_tokens cut left `structuredData` as a string instead of the schema object. When output is truncated the root brace never closes, so the balanced-span scan walks past it and matches a bracket pair living INSIDE a string value - `[step 1]` in a shell script becomes `["step 1"]` - reported with `truncated: false`. At other cut points nothing parsed at all, coerceJsonToSchema returned null, and the caller kept the raw text. A prefix sweep over a realistic huge-output payload hit the first case at ~22% of cut points and the second at a handful more. coerce: mark candidates that start at the document's first opening bracket and order those first, so the partial real root beats a span scraped from inside a string; flag every candidate truncated when the root never closes; and when an unclosed root yields nothing trustworthy, back off to the last completed structural boundary and repair from there, returning a PARTIAL object rather than degrading to raw text. consumers: add schemaAccepts and gate structuredData on it - a scalar root, or an experimental_output string that is not an exact raw-text echo, is only published when the caller's schema accepts it. neurolink also re-runs recovery when a provider already produced a schema-rejected string. String-root schemas are unaffected. The shared scalar-recovery policy lives in recoverScalarRoot (with ScalarRecoveryDecision in src/lib/types), used by both neurolink.recoverStructuredData and GenerationHandler.coerceTextMode so it cannot drift. anthropic: forced-json mode drops text blocks because the payload rides in the synthetic tool's input; when the response is cut short that tool call can be missing entirely, leaving an empty completion. Keep the text as a fallback when no synthetic tool call arrived, so a partial object can still be salvaged. final_result still supersedes it. Adds test:coerce-truncation (8 tests) sweeping every truncation point of a huge-output payload, with assertion messages restricted to structural cut positions (never recovered payload values) so a genuine failure reports as FAIL, not SKIP. Adds a test:json-e2e cell on the direct Anthropic path that forces a maxTokens cut and asserts structuredData is a plain object with jsonTruncated === true. Fixes #1156. Refs #635, #1152.
What
On the native-Anthropic structured-output path under escaping stress, models sometimes return
[{...}]instead of{...}for an object schema. Because a JS array istypeof "object",coerceJsonToSchemaaccepted it asfirstValidand returned an array asstructuredData, failing the caller's object-shaped schema. Caught adversarially during the audit's livetest:json-e2erun (#635).Fix
When a candidate array does not itself satisfy the schema, and it holds exactly one non-array object element that does satisfy it, unwrap to that element. Gated by
safeParse, mirroring the existing nested-string unwrap:Proof
coerce-nested-unwrap9/9 (3 new tests),structured-coerce7/7,json21/21 — no regressions.Refs #635 (feature already shipped; this hardens the fallback edge case).
Summary by CodeRabbit
Bug Fixes
Tests