From 6cafa3c0d987cf2ef4eae59e5c8e18292fbcfefa Mon Sep 17 00:00:00 2001 From: Sachin Sharma Date: Fri, 12 Jun 2026 02:13:52 +0530 Subject: [PATCH] fix(generation): guarantee valid JSON for schema requests + fix huge-text truncation Two related JSON-validity fixes for generate({ schema }), entirely in the SDK. 1) Structured output was wrongly disabled for Vertex+Claude. The tools<->schema exclusion is a Gemini-only API limitation, but the gate keyed on the whole Vertex provider, forcing Vertex+Claude+tools (TARA's production config) into text mode where hand-written JSON broke on a single mis-escaped character. - structuredOutputPolicy.ts: exclusion gated on isGeminiProvider only; isToolsSchemaConflictError detects runtime rejections (e.g. Groq) and transparently retries without structured output. - Provider-agnostic guarantee at the neurolink.ts boundary: content is always valid JSON (balanced-brace scan + jsonrepair) and the parsed object is exposed as result.structuredData -- override providers (vertex/anthropic/ bedrock/google-ai) bypass GenerationHandler, so the guarantee lives at the SDK boundary. 2) Huge structured responses were silently truncated. The native Claude paths hard-coded max_tokens to 4096; past ~16KB the JSON was cut mid-stream and coercion closed it into a valid-but-incomplete object with no signal (the Vertex native path didn't even surface finishReason). - resolveClaudeMaxTokens(): model-aware output ceiling (Sonnet 4.x 64K, Opus 4.x 32K, older models at their published limits); clamps over-large caller values so the native paths never 400. - Vertex native generate surfaces finishReason (max_tokens -> "length"). - coerceJsonToSchema returns { repaired, truncated }; GenerateResult exposes jsonRepaired / jsonTruncated plus a WARN -- truncation observable, never silent. - Anthropic client sets an explicit timeout so the SDK's non-streaming long-request guard doesn't reject a large max_tokens; generate timeout scales when a large output budget is in play. Verified live with tools active and complex Zod schemas (escaping torture, nested, array-heavy, 200-line huge-output with no maxTokens) across Vertex (Claude Sonnet/Opus 4.6 + Gemini 2.5), direct Anthropic (Sonnet/Opus 4.6), Google AI Studio, OpenAI and breadth providers: 33 passed / 0 failed. Huge outputs return complete valid JSON (16-24KB) where the old cap truncated; forced truncation yields jsonTruncated=true + finishReason=length. Deterministic unit suite covers the policy predicate, extractor, coercion and the new flags. --- CLAUDE.md | 3 +- .../2026-06-11-neurolink-json-validity.md | 1185 +++++++++++++++++ package.json | 3 + pnpm-lock.yaml | 6 +- src/lib/core/modules/GenerationHandler.ts | 114 +- .../core/modules/structuredOutputPolicy.ts | 70 + src/lib/neurolink.ts | 69 + src/lib/providers/anthropic.ts | 43 +- src/lib/providers/googleVertex.ts | 19 +- src/lib/types/generate.ts | 67 +- src/lib/types/utilities.ts | 21 + src/lib/utils/json/coerce.ts | 162 +++ src/lib/utils/json/extract.ts | 75 +- src/lib/utils/tokenLimits.ts | 60 + test/continuous-test-suite-json-e2e.ts | 529 ++++++++ test/continuous-test-suite-json.ts | 273 ++++ 16 files changed, 2634 insertions(+), 65 deletions(-) create mode 100644 docs/superpowers/plans/2026-06-11-neurolink-json-validity.md create mode 100644 src/lib/core/modules/structuredOutputPolicy.ts create mode 100644 src/lib/utils/json/coerce.ts create mode 100644 test/continuous-test-suite-json-e2e.ts create mode 100644 test/continuous-test-suite-json.ts diff --git a/CLAUDE.md b/CLAUDE.md index 84a44fc8c..3851ae659 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -26,7 +26,8 @@ These are non-negotiable. Violating them breaks the build or introduces bugs. 1. **Dynamic imports only in registry** — All providers must use dynamic imports inside factory functions in `providerRegistry.ts`. Static imports create circular dependencies. 2. **Types in canonical location** — All type definitions go in `src/lib/types/`. Never create type files inside feature subdirectories. -3. **Gemini tools + JSON schema are mutually exclusive** — Google AI Studio and Vertex AI cannot use tools and `structuredOutput` with a JSON schema simultaneously. It's an API limitation. Design workflows to use one or the other. +3. **Gemini tools + JSON schema are mutually exclusive** — Google AI Studio and Vertex **Gemini** models cannot use tools and `structuredOutput` with a JSON schema simultaneously (a Gemini API limitation). This does **not** apply to Vertex **Claude** models, which support both at once — the exclusion is gated on `isGeminiProvider` in `structuredOutputPolicy.ts`, not on the Vertex provider as a whole. Providers that reject the combination at runtime (e.g. Groq) are detected via `isToolsSchemaConflictError` and transparently retried without structured output. Regardless of provider, `generate({ schema })` is guaranteed to return valid JSON in `content` plus a parsed `structuredData` object (see `coerceJsonToSchema`). + - **Huge-text / truncation:** the native Claude paths (Vertex+Claude, direct Anthropic) must default `max_tokens` to the model's real output ceiling via `resolveClaudeMaxTokens` (Sonnet 4.x → 64K, Opus 4.x → 32K), **never** the legacy hard-coded 4096 that silently truncated large structured responses mid-JSON. The direct Anthropic non-streaming path also passes an explicit request `timeout` so the SDK's "streaming is required for long requests" pre-flight guard doesn't reject a large `max_tokens`. When output still hits the cap, truncation is surfaced — not silent: `coerceJsonToSchema` returns `{ repaired, truncated }`, and `GenerateResult` exposes `jsonRepaired` / `jsonTruncated` (set when `finishReason==="length"` or the recovered JSON came from an unclosed span) plus a WARN log. 4. **CLI ≠ SDK** — CLI can use manual MCP connections; the SDK cannot. Keep concerns separate. 5. **Backward compatibility** — Public SDK API must not break existing callers. 6. **`formatProviderError` must return, never throw** — Any provider error formatter must return the error object, not throw it. diff --git a/docs/superpowers/plans/2026-06-11-neurolink-json-validity.md b/docs/superpowers/plans/2026-06-11-neurolink-json-validity.md new file mode 100644 index 000000000..4c4f42ff5 --- /dev/null +++ b/docs/superpowers/plans/2026-06-11-neurolink-json-validity.md @@ -0,0 +1,1185 @@ +# NeuroLink JSON Validity — Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Make `neurolink.generate({ schema })` guarantee that `result.content` is syntactically valid JSON (and expose the parsed object as `result.structuredData`) for every provider, so consumers (curator/TARA, lighthouse, etc.) never have to parse fragile hand-escaped model text. + +**Architecture:** Three defensive layers, all inside the **neurolink SDK** (nothing in curator): + +1. **Root-cause gate fix** — the tools-vs-schema mutual-exclusion is a _Gemini_ limitation, but `GenerationHandler` applies it to _all_ Vertex (including Vertex+Claude, TARA's production config). Narrow the exclusion to Gemini-only so Vertex+Claude+tools uses AI-SDK `experimental_output` → schema-enforced output → `content = JSON.stringify(validatedObject)` (valid by construction). +2. **Robust text-mode coercion** — for the genuinely-unavoidable text-mode paths (real Gemini+tools, or any provider that returned raw text), parse the model text with a balanced-brace scanner + `jsonrepair` fallback, then re-serialize to canonical JSON. Guarantees syntactic validity even when the model mis-escaped its hand-written JSON. +3. **Expose `structuredData`** — thread the parsed object through `GenerateResult` so consumers can skip re-parsing entirely. + +Plus a correctness fix to the SDK's public `extractJsonStringFromText` (replace its non-greedy regex with the balanced scanner already present in the same file). + +**Tech Stack:** TypeScript (strict, ESM/NodeNext), Vercel AI SDK v6 (`ai@^6`, `@ai-sdk/anthropic@^3`), Zod, `jsonrepair@^3.14.0` (new dep), test harness `test/helpers/harness.ts` run via `npx tsx`. + +**Conventions (from CLAUDE.md — non-negotiable):** no `interface` (use `type`); named exports only; no `any` (use `unknown` + narrowing); types belong in `src/lib/types/` and are imported via the barrel `../types/index.js`; comments only when the _why_ is non-obvious. Run `pnpm run lint` (AST ESLint rules enforce these). + +**Verified facts this plan relies on:** + +- `GenerationHandler.ts` gate: `const useStructuredOutput = wantsStructuredOutput && !(isGoogleProvider && shouldUseTools && Object.keys(tools).length > 0);` where `isGoogleProvider = providerName === "google-ai" || providerName === "vertex"`. +- The file already defines (but does not use here) `isAnthropicProvider = ... || (providerName === "vertex" && modelName?.startsWith("claude-"))`. +- `formatEnhancedResult` sets `content = JSON.stringify(experimental_output)` when present, else strips fences from `generateResult.text` — and **discards** the parsed object. +- `NoObjectGeneratedError` fallback (re-runs without `experimental_output`) already exists → enabling structured output for Vertex+Claude is strictly safe. +- TARA runtime defaults: `neurolink-provider=vertex`, `neurolink-model=claude-sonnet-4-6`, tools registered (curator `registry.ts`). +- `GenerateResult` (src/lib/types/generate.ts) has no `structuredData` field; DTO builder in `neurolink.ts` (`const generateResult: GenerateResult = { content: textResult.content, ... }`) does not set one. +- `options.schema` type is `ValidationSchema = ZodTypeAny | Schema` (Zod schema _or_ AI-SDK JSON schema). +- `jsonrepair` is NOT yet a dependency. +- Tests: `import { defineSuite, assert, assertEqual, assertNotNull } from "./helpers/harness.js"`, `const { test, runSuite } = defineSuite("…")`, run via `npx tsx test/.ts`. `tsx` can import `src/**/*.ts` directly (fast TDD, no build). + +**Edit anchoring:** This branch will be rebased onto `origin/release` (Task 1), which shifts line numbers. **All edits below anchor on unique code strings, never line numbers.** If an anchor string is not found verbatim after rebase, re-grep for the nearest stable substring before editing. + +--- + +## File Structure + +**Create:** + +- `src/lib/core/modules/structuredOutputPolicy.ts` — pure predicate: is the tools/schema exclusion in force for this provider+model? (Gemini-only.) +- `src/lib/utils/json/coerce.ts` — pure `coerceJsonToSchema(text, schema)`: balanced-scan + `jsonrepair` → canonical `{ content, structuredData }` or `null`. +- `test/continuous-test-suite-json.ts` — harness suite covering the policy predicate, the extractor fix, and the coercion (no API calls). + +**Modify:** + +- `src/lib/utils/json/extract.ts` — replace non-greedy regex in `extractJsonStringFromText` with a shared balanced-span scanner; export the scanner for reuse. +- `src/lib/core/modules/GenerationHandler.ts` — (a) use the policy predicate at the gate; (b) in `formatEnhancedResult`, capture `structuredData` and run `coerceJsonToSchema` on the text-mode fallback; (c) add `structuredData` to the returned object. +- `src/lib/types/utilities.ts` — add `JsonCoercionResult` (rule 2: all types live in `src/lib/types/`; exported via the barrel). +- `src/lib/types/generate.ts` — add `structuredData?: unknown` to `GenerateResult`. +- `src/lib/neurolink.ts` — set `structuredData: textResult.structuredData` in the `GenerateResult` DTO builder. +- `package.json` — add `jsonrepair` dependency; add `test:json` script. + +--- + +## Task 1: Rebase branch onto origin/release + +**Files:** none (git only). The branch has 0 commits ahead and is behind several releases; this is a fast-forward with zero conflict risk. + +- [ ] **Step 1: Confirm clean tree and no local commits** + +Run: + +```bash +cd /Users/sachinsharma/Developer/temp/neurolink-fork/feat/json-fix +git fetch origin release +git status --short # expect: empty +git log --oneline origin/release..HEAD # expect: empty (0 commits ahead) +``` + +Expected: working tree clean, no commits ahead. + +- [ ] **Step 2: Rebase (fast-forward) onto origin/release** + +Run: + +```bash +git rebase origin/release +git log --oneline -1 # expect: tip now matches origin/release tip +``` + +Expected: branch advanced to `origin/release` tip; no conflicts. + +- [ ] **Step 3: Install deps (lockfile may have advanced)** + +Run: + +```bash +pnpm install +``` + +Expected: completes without errors. + +--- + +## Task 2: `structuredOutputPolicy` predicate (pure, TDD) + +**Files:** + +- Create: `src/lib/core/modules/structuredOutputPolicy.ts` +- Test: `test/continuous-test-suite-json.ts` + +- [ ] **Step 1: Write the failing test** + +Create `test/continuous-test-suite-json.ts`: + +```ts +#!/usr/bin/env tsx +/** + * Continuous Test Suite: JSON validity (no API). + * + * Covers the structured-output policy predicate, the balanced-brace JSON + * extractor, and schema-coercion of mis-escaped model text. All pure — no + * provider calls — so it runs in CI without keys. + * + * Run: npx tsx test/continuous-test-suite-json.ts + */ +import { + defineSuite, + assert, + assertEqual, + assertNotNull, +} from "./helpers/harness.js"; +import { + isGeminiProvider, + isToolsSchemaExclusionInForce, +} from "../src/lib/core/modules/structuredOutputPolicy.js"; + +const { test, runSuite } = defineSuite("JSON Validity"); + +await test("isGeminiProvider: google-ai is Gemini", () => { + assert( + isGeminiProvider("google-ai", "gemini-2.5-pro") === true, + "google-ai should be Gemini", + ); +}); + +await test("isGeminiProvider: vertex+gemini is Gemini", () => { + assert( + isGeminiProvider("vertex", "gemini-2.5-pro") === true, + "vertex+gemini should be Gemini", + ); +}); + +await test("isGeminiProvider: vertex+claude is NOT Gemini", () => { + assert( + isGeminiProvider("vertex", "claude-sonnet-4-6") === false, + "vertex+claude must not be Gemini", + ); +}); + +await test("isGeminiProvider: anthropic is NOT Gemini", () => { + assert( + isGeminiProvider("anthropic", "claude-sonnet-4-6") === false, + "anthropic must not be Gemini", + ); +}); + +await test("exclusion in force only for Gemini + tools present", () => { + // Vertex+Claude+tools: exclusion must NOT fire (this is the production bug fix). + assertEqual( + isToolsSchemaExclusionInForce("vertex", "claude-sonnet-4-6", true, 5), + false, + "vertex+claude+tools", + ); + // Vertex+Gemini+tools: exclusion fires (real API limitation). + assertEqual( + isToolsSchemaExclusionInForce("vertex", "gemini-2.5-pro", true, 5), + true, + "vertex+gemini+tools", + ); + // Gemini with NO tools: no exclusion. + assertEqual( + isToolsSchemaExclusionInForce("google-ai", "gemini-2.5-pro", false, 0), + false, + "gemini no tools", + ); + // Gemini with shouldUseTools but zero tools registered: no exclusion. + assertEqual( + isToolsSchemaExclusionInForce("google-ai", "gemini-2.5-pro", true, 0), + false, + "gemini zero tools", + ); +}); + +await runSuite(); +``` + +- [ ] **Step 2: Run the test to verify it fails** + +Run: + +```bash +npx tsx test/continuous-test-suite-json.ts +``` + +Expected: FAIL — module `../src/lib/core/modules/structuredOutputPolicy.js` not found (file does not exist yet). + +- [ ] **Step 3: Write the minimal implementation** + +Create `src/lib/core/modules/structuredOutputPolicy.ts`: + +```ts +/** + * Policy for when AI-SDK structured output (experimental_output) must be + * disabled because the provider cannot combine tool calls with JSON-schema + * enforcement. + * + * This is a GEMINI-ONLY API limitation. Anthropic Claude — including when + * hosted on Vertex (modelName starts with "claude-") — supports tools and + * structured output simultaneously, so it must NOT be excluded. The previous + * gate keyed on "any Vertex model", which wrongly disabled structured output + * for Vertex+Claude (the primary production config) and forced fragile + * hand-parsed JSON. + */ + +/** True when the provider+model is a Gemini model (the only family with the tools↔schema conflict). */ +export function isGeminiProvider( + providerName: string, + modelName: string | undefined, +): boolean { + if (providerName === "google-ai") { + return true; + } + if (providerName === "vertex") { + // Vertex hosts both Gemini and Claude. Only non-Claude (Gemini) models + // have the tools↔schema conflict. + return !(modelName?.startsWith("claude-") ?? false); + } + return false; +} + +/** + * True when structured output must be disabled for this call because tools are + * active on a Gemini provider. Mirrors the AI-SDK constraint exactly. + */ +export function isToolsSchemaExclusionInForce( + providerName: string, + modelName: string | undefined, + shouldUseTools: boolean, + toolCount: number, +): boolean { + return ( + isGeminiProvider(providerName, modelName) && shouldUseTools && toolCount > 0 + ); +} +``` + +- [ ] **Step 4: Run the test to verify it passes** + +Run: + +```bash +npx tsx test/continuous-test-suite-json.ts +``` + +Expected: PASS — all 5 tests green. + +- [ ] **Step 5: Commit** + +```bash +git add src/lib/core/modules/structuredOutputPolicy.ts test/continuous-test-suite-json.ts +git commit -m "feat(generation): add Gemini-only structured-output exclusion policy" +``` + +--- + +## Task 3: Use the policy at the GenerationHandler gate + +**Files:** + +- Modify: `src/lib/core/modules/GenerationHandler.ts` + +- [ ] **Step 1: Add the import** + +Add to the import block at the top of `GenerationHandler.ts` (next to other local module imports): + +```ts +import { isToolsSchemaExclusionInForce } from "./structuredOutputPolicy.js"; +``` + +- [ ] **Step 2: Replace the over-broad gate** + +Find (anchor — `callGenerateText`): + +```ts +const useStructuredOutput = + wantsStructuredOutput && + !(isGoogleProvider && shouldUseTools && Object.keys(tools).length > 0); +``` + +Replace with: + +```ts +// The tools↔schema conflict is a Gemini-only API limitation. Vertex+Claude +// supports both simultaneously, so only exclude for actual Gemini models. +const useStructuredOutput = + wantsStructuredOutput && + !isToolsSchemaExclusionInForce( + this.providerName, + this.modelName, + shouldUseTools, + Object.keys(tools).length, + ); +``` + +Note: leave `isGoogleProvider` defined — it is still used elsewhere in this method (thinking config / `providerOptions.google`). Only this gate changes. If ESLint now flags `isGoogleProvider` as unused, that means it had no other use; in that case delete its declaration too. (Verify with `pnpm run lint` in Step 4.) + +- [ ] **Step 3: Type-check** + +Run: + +```bash +pnpm run check +``` + +Expected: no new type errors. + +- [ ] **Step 4: Lint** + +Run: + +```bash +pnpm run lint +``` + +Expected: clean. If `isGoogleProvider` is reported unused, remove its `const isGoogleProvider = ...` declaration and re-run. + +- [ ] **Step 5: Commit** + +```bash +git add src/lib/core/modules/GenerationHandler.ts +git commit -m "fix(generation): enable structured output for Vertex+Claude with tools" +``` + +--- + +## Task 4: Balanced-brace scanner for `extractJsonStringFromText` + +**Files:** + +- Modify: `src/lib/utils/json/extract.ts` +- Test: `test/continuous-test-suite-json.ts` + +- [ ] **Step 1: Add the failing tests** + +Append to `test/continuous-test-suite-json.ts` BEFORE the final `await runSuite();` line: + +````ts +import { extractJsonStringFromText } from "../src/lib/utils/json/extract.js"; + +await test("extractor: returns the full outer object, not the first inner brace", () => { + // The string value contains a "}" — a non-greedy regex would stop early. + const input = 'noise {"a":{"b":"}"},"c":1} trailing'; + const got = extractJsonStringFromText(input); + assertNotNull(got, "should extract an object"); + assertEqual( + JSON.parse(got as string).c, + 1, + "must parse to the full object with c=1", + ); + assertEqual( + JSON.parse(got as string).a.b, + "}", + "nested brace-in-string preserved", + ); +}); + +await test("extractor: prose preamble then object (Vertex+tools text shape)", () => { + const input = 'Here is your result:\n{"summary":"ok","attachment":null}'; + const got = extractJsonStringFromText(input); + assertNotNull(got, "should find object after prose"); + assertEqual(JSON.parse(got as string).summary, "ok", "summary parsed"); +}); + +await test("extractor: fenced json code block", () => { + const input = '```json\n{"x":2}\n```'; + const got = extractJsonStringFromText(input); + assertNotNull(got, "should extract from fence"); + assertEqual(JSON.parse(got as string).x, 2, "x parsed"); +}); +```` + +- [ ] **Step 2: Run to verify the new tests fail** + +Run: + +```bash +npx tsx test/continuous-test-suite-json.ts +``` + +Expected: the "full outer object" test FAILS — the current non-greedy regex returns `{"b":"}"}` (the first balanced-looking inner fragment) or otherwise loses `c:1`. + +- [ ] **Step 3: Implement the shared scanner and rewire the function** + +In `src/lib/utils/json/extract.ts`, add this exported helper near the top (after the imports): + +```ts +/** + * Find the first balanced JSON object/array span starting at or after + * `fromIndex`. Quote- and escape-aware: braces inside string literals do not + * affect depth. Returns the matched substring and the index just past it, or + * null if no balanced span exists. + */ +export function nextBalancedJsonSpan( + text: string, + fromIndex = 0, +): { span: string; end: number } | null { + for (let start = fromIndex; start < text.length; start++) { + const openChar = text[start]; + if (openChar !== "{" && openChar !== "[") { + continue; + } + const closeChar = openChar === "{" ? "}" : "]"; + let depth = 0; + let inString = false; + let escapeNext = false; + for (let i = start; i < text.length; i++) { + const ch = text[i]; + if (escapeNext) { + escapeNext = false; + continue; + } + if (ch === "\\") { + escapeNext = true; + continue; + } + if (ch === '"') { + inString = !inString; + continue; + } + if (inString) { + continue; + } + if (ch === openChar) { + depth++; + } else if (ch === closeChar) { + depth--; + if (depth === 0) { + return { span: text.substring(start, i + 1), end: i + 1 }; + } + } + } + // Unbalanced from this start — try the next opening char. + } + return null; +} +``` + +Then replace the non-greedy candidate loop in `extractJsonStringFromText`. Find (anchor): + +```ts +// Try to find JSON object or array pattern using non-greedy iterative scan. +// Note: [\s\S]*? is non-greedy but can still produce over-spanning matches +// in texts with many braces. This is acceptable as we try-parse each candidate +// and move to the next on failure. A bracket-balancing parser would be more +// precise but significantly more complex for marginal benefit. +const candidateRegex = /(\{[\s\S]*?\}|\[[\s\S]*?\])/g; +let candidate: RegExpExecArray | null; +while ((candidate = candidateRegex.exec(text)) !== null) { + try { + JSON.parse(candidate[1]); + return candidate[1]; + } catch { + // Try next candidate + } +} + +return null; +``` + +Replace with: + +```ts +// Scan for balanced JSON object/array spans (quote/escape aware) and return +// the first one that parses. Unlike a non-greedy regex, this never stops at a +// "}" that lives inside a string value, so nested objects are preserved. +let searchFrom = 0; +for (;;) { + const found = nextBalancedJsonSpan(text, searchFrom); + if (!found) { + break; + } + try { + JSON.parse(found.span); + return found.span; + } catch { + // Not valid JSON — resume scanning just past this opening character. + } + searchFrom = text.indexOf(found.span[0], searchFrom) + 1; +} + +return null; +``` + +- [ ] **Step 4: Run to verify all extractor tests pass** + +Run: + +```bash +npx tsx test/continuous-test-suite-json.ts +``` + +Expected: PASS — including the "full outer object" test. + +- [ ] **Step 5: Type-check + lint** + +Run: + +```bash +pnpm run check && pnpm run lint +``` + +Expected: clean. + +- [ ] **Step 6: Commit** + +```bash +git add src/lib/utils/json/extract.ts test/continuous-test-suite-json.ts +git commit -m "fix(json): use balanced-brace scanner in extractJsonStringFromText" +``` + +--- + +## Task 5: `coerceJsonToSchema` — jsonrepair-backed canonicaliser (TDD) + +**Files:** + +- Modify: `package.json` (add `jsonrepair`) +- Create: `src/lib/utils/json/coerce.ts` +- Test: `test/continuous-test-suite-json.ts` + +- [ ] **Step 1: Add the dependency** + +Run: + +```bash +pnpm add jsonrepair@^3.14.0 +``` + +Expected: `jsonrepair` appears under `dependencies` in `package.json`. + +- [ ] **Step 2: Add the failing tests** + +Append to `test/continuous-test-suite-json.ts` before `await runSuite();`: + +```ts +import { coerceJsonToSchema } from "../src/lib/utils/json/coerce.js"; +import { z } from "zod"; + +const demoSchema = z.object({ + summary: z.string().min(1), + attachment: z + .object({ extension: z.string(), content: z.string() }) + .nullable(), +}); + +await test("coerce: unescaped double-quote inside content is repaired", () => { + // The model wrote a bare " inside the content value — invalid JSON. + const bad = + '{"summary":"done","attachment":{"extension":"txt","content":"He said "hi""}}'; + const out = coerceJsonToSchema(bad, demoSchema); + assertNotNull(out, "should repair + parse"); + // content must be valid JSON now and round-trip through JSON.parse. + const obj = JSON.parse((out as { content: string }).content); + assertEqual(obj.summary, "done", "summary preserved"); +}); + +await test("coerce: raw newline inside string is repaired", () => { + const bad = '{"summary":"line1\nline2","attachment":null}'; + const out = coerceJsonToSchema(bad, demoSchema); + assertNotNull(out, "should repair raw newline"); + const obj = JSON.parse((out as { content: string }).content); + assert( + typeof obj.summary === "string" && obj.summary.includes("line1"), + "summary text retained", + ); +}); + +await test("coerce: already-valid JSON passes through unchanged in meaning", () => { + const good = '{"summary":"ok","attachment":null}'; + const out = coerceJsonToSchema(good, demoSchema); + assertNotNull(out, "valid JSON should coerce"); + assertEqual( + (out as { structuredData: { summary: string } }).structuredData.summary, + "ok", + "object exposed", + ); +}); + +await test("coerce: prose-wrapped object is extracted then canonicalised", () => { + const wrapped = 'Sure! {"summary":"ok","attachment":null} hope that helps'; + const out = coerceJsonToSchema(wrapped, demoSchema); + assertNotNull(out, "should extract embedded object"); + assertEqual( + (out as { content: string }).content[0], + "{", + "content is canonical JSON starting with {", + ); +}); + +await test("coerce: pure prose with no JSON returns null", () => { + const out = coerceJsonToSchema( + "just a friendly hello, no json here", + demoSchema, + ); + assertEqual(out, null, "no JSON object -> null (caller keeps raw text)"); +}); +``` + +- [ ] **Step 3: Run to verify failure** + +Run: + +```bash +npx tsx test/continuous-test-suite-json.ts +``` + +Expected: FAIL — module `../src/lib/utils/json/coerce.js` not found. + +- [ ] **Step 4: Implement `coerceJsonToSchema`** + +Create `src/lib/utils/json/coerce.ts`: + +```ts +/** + * Coerce arbitrary model text into canonical, syntactically-valid JSON. + * + * Used on the text-mode path (providers/models that could not use AI-SDK + * structured output, e.g. real Gemini + tools). The model hand-writes JSON and + * frequently mis-escapes the content field (bare newline, unescaped quote, + * invalid escape like \d). A balanced-brace scan finds the object span; if + * JSON.parse rejects it, jsonrepair fixes common escaping mistakes; the result + * is re-serialised with JSON.stringify so downstream consumers always receive + * valid JSON. + * + * NOTE: jsonrepair is a heuristic. On content where a lone backslash is + * meaningful (regex/script/Windows path) it may drop the backslash, producing + * valid-but-semantically-altered content. This only affects the residual + * text-mode path — the primary Vertex+Claude path uses experimental_output and + * never reaches here. When jsonrepair changes the input we log at debug level + * so the event is observable. + */ +import { jsonrepair } from "jsonrepair"; +import type { + JsonCoercionResult, + ValidationSchema, +} from "../../types/index.js"; +import { logger } from "../../utils/logger.js"; +import { nextBalancedJsonSpan } from "./extract.js"; + +// NOTE: the result type lives in src/lib/types/utilities.ts (repo rule 2 — +// all types in src/lib/types/) and is imported via the barrel. Phase 2 added +// the `repaired`/`truncated` observability flags, so the final shipped shape is: +// export type JsonCoercionResult = { +// content: string; +// structuredData: unknown; +// repaired: boolean; // jsonrepair altered the text to make it parse +// truncated: boolean; // recovered from an unclosed (token-truncated) span +// }; + +/** Narrow a ValidationSchema to "looks like a Zod schema" (has safeParse). */ +function hasSafeParse(schema: ValidationSchema): schema is { + safeParse: (v: unknown) => { success: boolean; data?: unknown }; +} { + return typeof (schema as { safeParse?: unknown }).safeParse === "function"; +} + +/** Parse `candidate` as JSON, repairing common escaping mistakes on failure. */ +function parseOrRepair(candidate: string): unknown | undefined { + try { + return JSON.parse(candidate); + } catch { + // fall through to repair + } + try { + const repaired = jsonrepair(candidate); + const value = JSON.parse(repaired); + if (repaired !== candidate && logger.shouldLog("debug")) { + logger.debug("[coerceJsonToSchema] jsonrepair altered model output", { + originalLength: candidate.length, + repairedLength: repaired.length, + }); + } + return value; + } catch { + return undefined; + } +} + +/** + * Try to produce canonical JSON from `text`. Returns null when no JSON object + * could be recovered (caller should then keep the raw text). + * + * When `schema` is a Zod schema, candidates that satisfy it are preferred; a + * syntactically-valid-but-schema-failing object is still returned (we guarantee + * JSON *validity*, leaving schema/content checks to the caller's own pipeline). + */ +export function coerceJsonToSchema( + text: string, + schema?: ValidationSchema, +): JsonCoercionResult | null { + if (typeof text !== "string" || text.trim().length === 0) { + return null; + } + + let firstValid: unknown | undefined; + let schemaMatch: unknown | undefined; + + let searchFrom = 0; + for (;;) { + const found = nextBalancedJsonSpan(text, searchFrom); + if (!found) { + break; + } + const parsed = parseOrRepair(found.span); + if (parsed !== undefined && parsed !== null && typeof parsed === "object") { + if (firstValid === undefined) { + firstValid = parsed; + } + if (schema && hasSafeParse(schema)) { + if (schema.safeParse(parsed).success) { + schemaMatch = parsed; + break; + } + } else { + // No Zod schema to discriminate — first parseable object wins. + break; + } + } + searchFrom = found.end; + } + + const chosen = schemaMatch ?? firstValid; + if (chosen === undefined) { + return null; + } + // Phase 1 shape shown here; Phase 2 extends the return with the + // `repaired`/`truncated` flags (see the JsonCoercionResult note above and + // the Phase 2 section) — the shipped code returns all four fields. + return { content: JSON.stringify(chosen), structuredData: chosen }; +} +``` + +Note: confirm the logger import path matches the codebase. Find it with: + +```bash +grep -rn "export const logger\|export { logger" src/lib/utils/logger.ts src/lib/**/logger.ts 2>/dev/null | head +``` + +Adjust the `import { logger }` path to the real location if different. + +- [ ] **Step 5: Run to verify pass** + +Run: + +```bash +npx tsx test/continuous-test-suite-json.ts +``` + +Expected: PASS — all coercion tests green. + +- [ ] **Step 6: Type-check + lint** + +Run: + +```bash +pnpm run check && pnpm run lint +``` + +Expected: clean. (`ValidationSchema` must be imported from the barrel `../../types/index.js` per CLAUDE.md rule 13.) + +- [ ] **Step 7: Commit** + +```bash +git add package.json pnpm-lock.yaml src/lib/utils/json/coerce.ts test/continuous-test-suite-json.ts +git commit -m "feat(json): add jsonrepair-backed coerceJsonToSchema for text-mode JSON" +``` + +--- + +## Task 6: Populate `structuredData` + coerce text-mode output in `formatEnhancedResult` + +**Files:** + +- Modify: `src/lib/core/modules/GenerationHandler.ts` + +- [ ] **Step 1: Add the import** + +Add near the other local imports in `GenerationHandler.ts`: + +```ts +import { coerceJsonToSchema } from "../../utils/json/coerce.js"; +``` + +- [ ] **Step 2: Rewrite the structured-output branch to capture `structuredData`** + +Find (anchor — the whole `content` resolution block in `formatEnhancedResult`): + +````ts +let content: string; +if (useStructuredOutput) { + try { + const experimentalOutput = generateResult.experimental_output; + if (experimentalOutput !== undefined) { + content = JSON.stringify(experimentalOutput); + } else { + // Fall back to text parsing + const rawText = generateResult.text || ""; + const strippedText = rawText + .replace(/^```(?:json)?\s*\n?/i, "") + .replace(/\n?```\s*$/i, "") + .trim(); + content = strippedText; + } + } catch (outputError) { + // experimental_output is a getter that can throw NoObjectGeneratedError + // Fall back to text parsing when structured output fails + logger.debug( + "[GenerationHandler] experimental_output threw, falling back to text parsing", + { + error: + outputError instanceof Error + ? outputError.message + : String(outputError), + }, + ); + const rawText = generateResult.text || ""; + const strippedText = rawText + .replace(/^```(?:json)?\s*\n?/i, "") + .replace(/\n?```\s*$/i, "") + .trim(); + content = strippedText; + } +} else { + content = generateResult.text; +} +```` + +Replace with: + +````ts +let content: string; +let structuredData: unknown; +// Strip an outer ```json fence and coerce raw model text into canonical +// JSON. Shared by both text-mode fallbacks below so a mis-escaped +// hand-written object still yields valid JSON. +const coerceTextMode = (rawText: string): string => { + const strippedText = rawText + .replace(/^```(?:json)?\s*\n?/i, "") + .replace(/\n?```\s*$/i, "") + .trim(); + const coerced = coerceJsonToSchema(strippedText, options.schema); + if (coerced) { + structuredData = coerced.structuredData; + return coerced.content; + } + return strippedText; +}; +if (useStructuredOutput) { + try { + const experimentalOutput = generateResult.experimental_output; + if (experimentalOutput !== undefined) { + // AI-SDK already parsed + schema-validated the object. Expose it + // directly and serialise canonically — no hand-parsing needed. + structuredData = experimentalOutput; + content = JSON.stringify(experimentalOutput); + } else { + content = coerceTextMode(generateResult.text || ""); + } + } catch (outputError) { + // experimental_output is a getter that can throw NoObjectGeneratedError. + logger.debug( + "[GenerationHandler] experimental_output threw, falling back to text parsing", + { + error: + outputError instanceof Error + ? outputError.message + : String(outputError), + }, + ); + content = coerceTextMode(generateResult.text || ""); + } +} else { + content = generateResult.text; +} +```` + +- [ ] **Step 3: Add `structuredData` to the returned object** + +Find (anchor — the start of the `formatEnhancedResult` return): + +```ts + return { + content, + usage, + provider: this.providerName, +``` + +Replace with: + +```ts + return { + content, + structuredData, + usage, + provider: this.providerName, +``` + +- [ ] **Step 4: Type-check** + +Run: + +```bash +pnpm run check +``` + +Expected: type error — `structuredData` is not assignable to the return type `EnhancedGenerateResult` yet. This is fixed in Task 7. (If you prefer green-at-every-step, do Task 7 Step 1 now, then return here.) + +- [ ] **Step 5: Commit (after Task 7 makes it green)** + +Defer the commit until Task 7 lands, then: + +```bash +git add src/lib/core/modules/GenerationHandler.ts +git commit -m "feat(generation): expose structuredData and coerce text-mode JSON" +``` + +--- + +## Task 7: Thread `structuredData` through the result type + DTO + +**Files:** + +- Modify: `src/lib/types/generate.ts` +- Modify: `src/lib/neurolink.ts` + +- [ ] **Step 1: Add the field to `GenerateResult`** + +In `src/lib/types/generate.ts`, find (anchor): + +```ts +export type GenerateResult = { + content: string; // Primary output + outputs?: { text: string }; // Future extensible for multi-modal +``` + +Replace with: + +```ts +export type GenerateResult = { + content: string; // Primary output + /** + * Parsed structured object when a `schema` was requested. Populated from + * AI-SDK experimental_output, or from text-mode coercion (balanced-scan + + * jsonrepair). Prefer this over JSON.parse(content) — it never requires the + * caller to re-parse hand-escaped model text. + */ + structuredData?: unknown; + outputs?: { text: string }; // Future extensible for multi-modal +``` + +- [ ] **Step 2: Set it in the DTO builder** + +In `src/lib/neurolink.ts`, find (anchor): + +```ts + const generateResult: GenerateResult = { + content: textResult.content, +``` + +Replace with: + +```ts + const generateResult: GenerateResult = { + content: textResult.content, + structuredData: textResult.structuredData, +``` + +- [ ] **Step 3: Type-check** + +Run: + +```bash +pnpm run check +``` + +Expected: PASS — `formatEnhancedResult` (Task 6) now type-checks against the extended `GenerateResult`, and `textResult.structuredData` resolves. + +- [ ] **Step 4: Lint** + +Run: + +```bash +pnpm run lint +``` + +Expected: clean. + +- [ ] **Step 5: Commit Task 6 + Task 7 together** + +```bash +git add src/lib/types/generate.ts src/lib/neurolink.ts src/lib/core/modules/GenerationHandler.ts +git commit -m "feat(generation): thread structuredData through GenerateResult DTO" +``` + +--- + +## Task 8: Wire the new suite into package.json + full verification + +**Files:** + +- Modify: `package.json` + +- [ ] **Step 1: Add the test script** + +In `package.json` `scripts`, add next to the other `test:*` entries: + +```json + "test:json": "npx tsx test/continuous-test-suite-json.ts", +``` + +- [ ] **Step 2: Run the JSON suite from the script** + +Run: + +```bash +pnpm run test:json +``` + +Expected: PASS — all tests across policy, extractor, and coercion. + +- [ ] **Step 3: Full build (compiles src → dist that consumers import)** + +Run: + +```bash +pnpm run build +``` + +Expected: build succeeds (no TS errors). + +- [ ] **Step 4: Quality gate** + +Run: + +```bash +pnpm run check && pnpm run lint +``` + +Expected: both clean. + +- [ ] **Step 5: Run an existing structured/provider suite that exercises generate() (no regressions)** + +Run (requires Vertex creds; skips gracefully without): + +```bash +pnpm run test:providers-mocked +``` + +Expected: no new failures vs the pre-change baseline. (Mocked suite runs without live keys.) + +- [ ] **Step 6: Commit** + +```bash +git add package.json +git commit -m "test(json): wire continuous-test-suite-json into package scripts" +``` + +--- + +## Task 9 (optional, recommended): Live end-to-end confirmation against Vertex+Claude+tools + +Only if Vertex credentials are available. Confirms the production path now emits valid JSON via `experimental_output` and exposes `structuredData`. + +- [ ] **Step 1: One-off live probe** + +Run: + +```bash +npx tsx -e ' +import { NeuroLink } from "./dist/index.js"; +import { z } from "zod"; +const nl = new NeuroLink(); +const schema = z.object({ summary: z.string(), attachment: z.object({ extension: z.string(), content: z.string() }).nullable() }); +const r = await nl.generate({ + input: { text: "Return a JSON object: summary plus an attachment whose content is a 5-line bash script that prints quotes and a Windows path C:\\\\Users\\\\me. Use extension sh." }, + provider: "vertex", + model: "claude-sonnet-4-6", + schema, +}); +console.log("structuredData present:", r.structuredData !== undefined); +console.log("content parses:", (() => { try { JSON.parse(r.content); return true; } catch { return false; } })()); +' +``` + +Expected: `structuredData present: true` and `content parses: true`. (Before the fix, with tools registered, this path produced raw text and could fail to parse.) + +- [ ] **Step 2: No commit** — this is a manual verification only. + +--- + +## Self-Review + +**Spec coverage:** + +- "Fix everything in neurolink, nothing in curator" → all tasks touch only `src/lib/**` and `test/**` in the neurolink repo. ✓ +- "Apply jsonrepair in neurolink" → Task 5. ✓ +- Root cause (Vertex+Claude wrongly excluded) → Tasks 2-3. ✓ +- "Ensure JSON output is always valid" → experimental_output path (Tasks 3,6) for providers that support it; coercion fallback (Tasks 5-6) for the rest; extractor fix (Task 4) for the public util. ✓ +- "Rebase if required" → Task 1. ✓ +- Expose parsed object so consumers never re-parse → Tasks 6-7 (`structuredData`). ✓ + +**Placeholder scan:** No TBD/TODO; every code step shows complete code; every command shows expected output. The only two "verify the real path" notes (logger import location in Task 5; `isGoogleProvider` possibly-unused in Task 3) are explicit grep/lint checks with defined fallbacks, not placeholders. + +**Type consistency:** `coerceJsonToSchema(text, schema?) → { content, structuredData } | null` used identically in Task 5 (def) and Task 6 (call). `nextBalancedJsonSpan(text, fromIndex) → { span, end } | null` defined in Task 4, consumed in Tasks 4 and 5. `isToolsSchemaExclusionInForce(providerName, modelName, shouldUseTools, toolCount)` defined Task 2, used Task 3. `structuredData?: unknown` added to `GenerateResult` (Task 7) matches its assignment in `formatEnhancedResult` (Task 6) and the DTO builder (Task 7). + +**Residual risks (documented, not gaps):** + +- jsonrepair can semantically alter backslash-bearing content on the _text-mode_ path only (Gemini+tools / non-structured providers). The primary Vertex+Claude path bypasses it. A debug log fires when repair changes the input. Extension-scoped skipping can be added later if telemetry shows real corruption. +- finishReason=length truncation can still yield an incomplete attachment; coercion makes it _valid_ JSON but cannot restore missing bytes. Out of scope for "valid JSON" — handled separately by the caller's truncation notice. + +``` + +``` + +--- + +## Phase 2 — Huge-text truncation (follow-up) + +The Phase 1 residual risk ("finishReason=length can yield an incomplete +attachment") turned out to be the **dominant** real-world failure for large +TARA responses. Root-caused via a 6-probe workflow + live repro. + +### Root cause + +The native Claude paths hard-coded `max_tokens` to **4096**, bypassing +`getSafeMaxTokens` (which would return the 64K provider default): + +- `googleVertex.ts` `executeNativeAnthropicStream` / `...Generate`: `options.maxTokens || 4096` +- `anthropic.ts`: `ANTHROPIC_DEFAULT_MAX_TOKENS = 4096` + +Any structured response larger than ~16 KB was silently truncated mid-JSON. +On truncation the AI SDK skips `parseCompleteOutput` (it only runs on +`finishReason==="stop"`), so the path fell to text-mode coercion, which closed +the dangling JSON into a **valid-but-incomplete** object with **no signal** — +the Vertex native generate path didn't even surface `finishReason`. + +### Fix (all in NeuroLink) + +1. **Model-aware output ceiling** — `resolveClaudeMaxTokens(model, requested)` in + `tokenLimits.ts`: defaults to the model's real max (Sonnet 4.x → 64K, Opus + 4.x → 32K, older models at their published limits), clamps over-large + caller values (avoids 400s on the native paths). Used at both Vertex+Claude + sites and both Anthropic native sites. +2. **Surface `finishReason`** on the Vertex native generate path (map Anthropic + `stop_reason: "max_tokens"` → `"length"`); it previously hard-coded `"stop"`. +3. **Make truncation observable** — `coerceJsonToSchema` returns `{ repaired, +truncated }`; `GenerateResult`/`TextGenerationResult` expose `jsonRepaired` + / `jsonTruncated`; `neurolink.ts` + `GenerationHandler.ts` set the flag when + `finishReason==="length"` and emit a WARN. No more silent data loss. +4. **Anthropic non-streaming guard** — pass an explicit request `timeout` so the + SDK's "streaming is required for long requests" pre-flight throw doesn't + reject a large `max_tokens`; the abort signal stays the real duration bound. + +### Verification + +- Live matrix across Vertex (Claude Sonnet/Opus 4.6 + Gemini 2.5), direct + Anthropic (Sonnet/Opus 4.6), Google AI Studio, OpenAI, and breadth providers: + huge-output (260-line script, **no** `maxTokens`) returns **complete** valid + JSON (20–24 KB) — the old 4096 cap truncated it. +- Dedicated tests on the production cell: "complete (no maxTokens)" and "forced + truncation is observable" (`jsonTruncated=true, finishReason=length`). +- Unit suite covers the `repaired` / `truncated` flags deterministically. + +### Out of scope (flagged, not fixed) + +- **OpenAI per-model default** — `getSafeMaxTokens("openai")` returns the + provider default (128K), which exceeds smaller models' completion limit (e.g. + `gpt-4o-mini` = 16384) and 400s when a caller omits `maxTokens`. This is a + pre-existing issue on a non-Claude path; a model-aware OpenAI ceiling is a + separate follow-up. +- **Auto-continuation** of a truncated JSON generation (stitch partial + resume) + was deliberately not implemented — fragile JSON-stitching that can produce + wrong output is worse than a flagged, raised-ceiling truncation. Raising the + ceiling to ~256 KB output + making any residual truncation observable is the + robust, correct fix. diff --git a/package.json b/package.json index 237517908..7bff57659 100644 --- a/package.json +++ b/package.json @@ -108,6 +108,8 @@ "test:dynamic": "npx tsx test/continuous-test-suite-dynamic.ts", "test:proxy": "npx tsx test/continuous-test-suite-proxy.ts", "test:bugfixes": "npx tsx test/continuous-test-suite-bugfixes.ts", + "test:json": "npx tsx test/continuous-test-suite-json.ts", + "test:json-e2e": "npx tsx test/continuous-test-suite-json-e2e.ts", "test:workflow": "npx tsx test/continuous-test-suite-workflow.ts", "test:hitl": "npx tsx test/continuous-test-suite-hitl.ts", "test:tasks": "npx tsx test/continuous-test-suite-tasks.ts", @@ -341,6 +343,7 @@ "inquirer": "^13.3.0", "jose": "^6.1.3", "json-schema-to-zod": "^2.7.0", + "jsonrepair": "^3.14.0", "nanoid": "^5.1.5", "open": "^11.0.0", "ora": "^9.3.0", diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index ebf2be4cc..4381974a4 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -158,6 +158,9 @@ importers: json-schema-to-zod: specifier: ^2.7.0 version: 2.8.0 + jsonrepair: + specifier: ^3.14.0 + version: 3.14.0 nanoid: specifier: ^5.1.5 version: 5.1.7 @@ -15652,8 +15655,7 @@ snapshots: optionalDependencies: graceful-fs: 4.2.11 - jsonrepair@3.14.0: - optional: true + jsonrepair@3.14.0: {} jszip@3.10.1: dependencies: diff --git a/src/lib/core/modules/GenerationHandler.ts b/src/lib/core/modules/GenerationHandler.ts index 4c30b9102..b93ddc4bd 100644 --- a/src/lib/core/modules/GenerationHandler.ts +++ b/src/lib/core/modules/GenerationHandler.ts @@ -39,6 +39,11 @@ import { extractTokenUsage, } from "../../utils/tokenUtils.js"; import { DEFAULT_MAX_STEPS } from "../constants.js"; +import { + isToolsSchemaConflictError, + isToolsSchemaExclusionInForce, +} from "./structuredOutputPolicy.js"; +import { coerceJsonToSchema } from "../../utils/json/coerce.js"; import type { LanguageModel, ModelMessage, @@ -125,9 +130,16 @@ export class GenerationHandler { (!!options.schema || options.output?.format === "json" || options.output?.format === "structured"); + // The tools↔schema conflict is a Gemini-only API limitation. Vertex+Claude + // supports both simultaneously, so only exclude for actual Gemini models. const useStructuredOutput = wantsStructuredOutput && - !(isGoogleProvider && shouldUseTools && Object.keys(tools).length > 0); + !isToolsSchemaExclusionInForce( + this.providerName, + this.modelName, + shouldUseTools, + Object.keys(tools).length, + ); // Annotate the last tool with cache_control so the full tool-definition // block becomes a cache breakpoint for Anthropic-family providers. @@ -417,25 +429,37 @@ export class GenerationHandler { span.setStatus({ code: SpanStatusCode.OK }); return result; } catch (error) { - // If NoObjectGeneratedError is thrown when using schema + tools together, - // fall back to generating without experimental_output and extract JSON manually - if (error instanceof NoObjectGeneratedError && useStructuredOutput) { + // Fall back to text-mode (no experimental_output) when structured + // output + tools failed, in two cases: + // 1. NoObjectGeneratedError — the SDK couldn't coerce the object. + // 2. The provider rejected json-mode-with-tools outright (e.g. Groq: + // "json mode cannot be combined with tool/function calling"). + // In both cases we retry without structured output and let + // formatEnhancedResult coerce the text response into valid JSON. + const isStructuredOutputConflict = + useStructuredOutput && + (error instanceof NoObjectGeneratedError || + isToolsSchemaConflictError(error)); + if (isStructuredOutputConflict) { span.setAttribute("neurolink.has_fallback", true); // NLK-GAP-007: Record initial failure event before fallback retry span.addEvent("retry.initial_failure", { - "error.message": error.message, + "error.message": + error instanceof Error ? error.message : String(error), "retry.attempt": 1, "retry.reason": - "NoObjectGeneratedError_structured_output_fallback", + error instanceof NoObjectGeneratedError + ? "NoObjectGeneratedError_structured_output_fallback" + : "tools_schema_conflict_structured_output_fallback", }); logger.debug( - "[GenerationHandler] NoObjectGeneratedError caught - falling back to manual JSON extraction", + "[GenerationHandler] structured-output conflict caught - falling back to manual JSON extraction", { provider: this.providerName, model: this.modelName, - error: error.message, + error: error instanceof Error ? error.message : String(error), }, ); @@ -704,23 +728,58 @@ export class GenerationHandler { options.output?.format === "structured"; let content: string; + let structuredData: unknown; + let jsonRepaired = false; + let jsonTruncated = false; + // Strip an outer ```json fence and coerce raw model text into canonical + // JSON. Object/array roots are recovered via balanced-scan + jsonrepair; + // scalar JSON roots (string/number/bool) via plain JSON.parse. When + // nothing JSON-shaped is recoverable, the raw text is returned unchanged, + // structuredData stays unset, and a WARN makes the broken case observable. + const coerceTextMode = (rawText: string): string => { + const strippedText = rawText + .replace(/^```(?:json)?\s*\n?/i, "") + .replace(/\n?```\s*$/i, "") + .trim(); + const coerced = coerceJsonToSchema(strippedText, options.schema); + if (coerced) { + structuredData = coerced.structuredData; + if (coerced.repaired) { + jsonRepaired = true; + } + if (coerced.truncated) { + jsonTruncated = true; + } + return coerced.content; + } + try { + const scalar: unknown = JSON.parse(strippedText); + if (scalar !== null && scalar !== undefined) { + structuredData = scalar; + return strippedText; + } + } catch { + // not JSON at all — fall through to raw text + WARN + } + logger.warn( + "[GenerationHandler] schema requested but no JSON could be recovered from model text; returning raw text", + { provider: this.providerName, model: this.modelName }, + ); + return strippedText; + }; if (useStructuredOutput) { try { const experimentalOutput = generateResult.experimental_output; if (experimentalOutput !== undefined) { + // AI-SDK already parsed + schema-validated the object. Expose it + // directly and serialise canonically — no hand-parsing needed. + structuredData = experimentalOutput; content = JSON.stringify(experimentalOutput); } else { - // Fall back to text parsing - const rawText = generateResult.text || ""; - const strippedText = rawText - .replace(/^```(?:json)?\s*\n?/i, "") - .replace(/\n?```\s*$/i, "") - .trim(); - content = strippedText; + content = coerceTextMode(generateResult.text || ""); } } catch (outputError) { - // experimental_output is a getter that can throw NoObjectGeneratedError - // Fall back to text parsing when structured output fails + // experimental_output is a getter that can throw NoObjectGeneratedError. logger.debug( "[GenerationHandler] experimental_output threw, falling back to text parsing", { @@ -730,17 +789,23 @@ export class GenerationHandler { : String(outputError), }, ); - const rawText = generateResult.text || ""; - const strippedText = rawText - .replace(/^```(?:json)?\s*\n?/i, "") - .replace(/\n?```\s*$/i, "") - .trim(); - content = strippedText; + content = coerceTextMode(generateResult.text || ""); } } else { content = generateResult.text; } + // Tie the coercion repair to the provider's truncation signal: if the + // response stopped on the token cap, treat the structured output as + // truncated regardless of which coerce candidate won, and warn. + if (useStructuredOutput && generateResult.finishReason === "length") { + jsonTruncated = true; + logger.warn( + "[GenerationHandler] Structured output truncated by token cap (finishReason=length); increase maxTokens", + { provider: this.providerName, model: this.modelName }, + ); + } + // Extract usage with support for different formats and reasoning tokens // Note: The AI SDK bundles thinking tokens into promptTokens for Google models. // Separate reasoningTokens tracking will work when/if the AI SDK adds support. @@ -798,8 +863,11 @@ export class GenerationHandler { return { content, + structuredData, usage, finishReason: generateResult.finishReason, + jsonRepaired: jsonRepaired || undefined, + jsonTruncated: jsonTruncated || undefined, provider: this.providerName, model: this.modelName, reasoning, diff --git a/src/lib/core/modules/structuredOutputPolicy.ts b/src/lib/core/modules/structuredOutputPolicy.ts new file mode 100644 index 000000000..50eb0cc2c --- /dev/null +++ b/src/lib/core/modules/structuredOutputPolicy.ts @@ -0,0 +1,70 @@ +/** + * Policy for when AI-SDK structured output (experimental_output) must be + * disabled because the provider cannot combine tool calls with JSON-schema + * enforcement. + * + * This is a GEMINI-ONLY API limitation. Anthropic Claude — including when + * hosted on Vertex (modelName starts with "claude-") — supports tools and + * structured output simultaneously, so it must NOT be excluded. A gate keyed on + * "any Vertex model" wrongly disables structured output for Vertex+Claude (the + * primary production config) and forces fragile hand-parsed JSON. + */ + +/** True when the provider+model is a Gemini model (the only family with the tools↔schema conflict). */ +export function isGeminiProvider( + providerName: string, + modelName: string | undefined, +): boolean { + if (providerName === "google-ai") { + return true; + } + if (providerName === "vertex") { + // Vertex hosts both Gemini and Claude. Only non-Claude (Gemini) models + // have the tools↔schema conflict. + return !(modelName?.startsWith("claude-") ?? false); + } + return false; +} + +/** + * True when structured output must be disabled for this call because tools are + * active on a Gemini provider. Mirrors the AI-SDK constraint exactly. + */ +export function isToolsSchemaExclusionInForce( + providerName: string, + modelName: string | undefined, + shouldUseTools: boolean, + toolCount: number, +): boolean { + return ( + isGeminiProvider(providerName, modelName) && shouldUseTools && toolCount > 0 + ); +} + +/** + * True when a provider error indicates the request was rejected because JSON / + * structured output and tool-calling cannot be combined for that provider + * (e.g. Groq: "json mode cannot be combined with tool/function calling"). + * + * This is the runtime, provider-agnostic complement to the static Gemini gate: + * any provider with the same limitation can be detected from its error message + * and retried without structured output instead of failing the call. + */ +export function isToolsSchemaConflictError(error: unknown): boolean { + const message = + error instanceof Error + ? error.message + : typeof error === "string" + ? error + : ""; + return ( + /cannot be combined with (a )?(tool|function)/i.test(message) || + /json[\s_-]?(mode|schema|object|output)[^.]{0,60}(tool|function)/i.test( + message, + ) || + /(tool|function)[\s-]?call[^.]{0,60}json[\s_-]?(mode|schema)/i.test( + message, + ) || + /response_format[^.]{0,60}(tool|function)/i.test(message) + ); +} diff --git a/src/lib/neurolink.ts b/src/lib/neurolink.ts index 9b095c49d..54bc0d22b 100644 --- a/src/lib/neurolink.ts +++ b/src/lib/neurolink.ts @@ -204,6 +204,7 @@ import { } from "./utils/lifecycleCallbacks.js"; import { resolveLifecycleTimeoutMs } from "./utils/lifecycleTimeout.js"; import { cloneOptionsForCallIsolation } from "./utils/cloneOptions.js"; +import { coerceJsonToSchema } from "./utils/json/coerce.js"; // Factory processing imports import { createCleanStreamOptions, @@ -4467,6 +4468,71 @@ Current user's request: ${currentInput}`; startTime, } = params; + // Provider-agnostic JSON coercion for schema requests. Structured-output + // enforcement makes valid JSON the overwhelming case; for every other + // provider path — including generate() overrides (Vertex, Anthropic, + // Bedrock, Google AI Studio) — object/array roots are recovered here via + // balanced-scan + jsonrepair and scalar JSON roots via plain JSON.parse, + // with the parsed value exposed as `structuredData`. If nothing + // JSON-shaped is recoverable (pure prose), the raw text is returned, + // `structuredData` stays undefined, and a WARN makes the case observable. + // Runs BEFORE the end-of-generation emits below so event consumers see + // the same coerced content/structuredData the caller receives. + if ( + textOptions.schema && + textResult.structuredData === undefined && + typeof textResult.content === "string" + ) { + const coerced = coerceJsonToSchema( + textResult.content, + textOptions.schema, + ); + if (coerced) { + textResult.content = coerced.content; + textResult.structuredData = coerced.structuredData; + if (coerced.repaired) { + textResult.jsonRepaired = true; + } + if (coerced.truncated) { + textResult.jsonTruncated = true; + } + } else { + try { + const scalar: unknown = JSON.parse(textResult.content); + if (scalar !== null && scalar !== undefined) { + textResult.structuredData = scalar; + } + } catch { + logger.warn( + "[NeuroLink] schema requested but no JSON could be recovered from model output; returning raw text", + { provider: textResult.provider, model: textResult.model }, + ); + } + } + } + + // Surface truncation when a schema was requested: either the provider + // reported finishReason="length" or the recovered JSON came from an + // unclosed span. Either way `structuredData` may be incomplete — warn at + // info level so it is observable in production (not just debug logs). + if (textOptions.schema) { + if (textResult.finishReason === "length") { + textResult.jsonTruncated = true; + } + if (textResult.jsonTruncated) { + logger.warn( + "[NeuroLink] Structured output may be truncated (finishReason=length or unclosed JSON); " + + "increase maxTokens to fit the full response.", + { + provider: textResult.provider, + model: textResult.model, + finishReason: textResult.finishReason, + outputTokens: textResult.usage?.output, + }, + ); + } + } + // Skip the top-level `generation:end` emission when the provider already // emitted it from its native generate path (Vertex / Google AI Studio). // Without this guard, native-path providers would surface TWO events @@ -4507,7 +4573,10 @@ Current user's request: ${currentInput}`; const generateResult: GenerateResult = { content: textResult.content, + structuredData: textResult.structuredData, finishReason: textResult.finishReason, + jsonRepaired: textResult.jsonRepaired, + jsonTruncated: textResult.jsonTruncated, provider: textResult.provider, model: textResult.model, usage: textResult.usage diff --git a/src/lib/providers/anthropic.ts b/src/lib/providers/anthropic.ts index 8f2c15ad9..0ca154815 100644 --- a/src/lib/providers/anthropic.ts +++ b/src/lib/providers/anthropic.ts @@ -74,6 +74,7 @@ import { stampNoOutputSpan, } from "../utils/noOutputSentinel.js"; import { convertZodToJsonSchema } from "../utils/schemaConversion.js"; +import { resolveClaudeMaxTokens } from "../utils/tokenLimits.js"; import { createChunkQueue, createDeferredAnalytics, @@ -657,10 +658,19 @@ const mapAnthropicStopReason = ( } }; -// Anthropic's Messages API requires max_tokens on every request. The previous -// @ai-sdk/anthropic implementation defaulted it to 4096 when the caller did -// not specify maxTokens — preserve that wire behavior. -const ANTHROPIC_DEFAULT_MAX_TOKENS = 4096; +// Anthropic's Messages API requires max_tokens on every request. When the +// caller omits it, default to the model's real output ceiling via +// resolveClaudeMaxTokens (e.g. 64K for Sonnet 4.x) instead of the legacy 4096, +// which silently truncated large structured responses mid-JSON. +// +// Client-level request timeout. The Anthropic SDK throws "Streaming is required +// for long requests" from a NON-streaming `messages.create` when `max_tokens` +// is large AND no client-level timeout is configured (it can't estimate a safe +// timeout). Setting an explicit client timeout — equal to the SDK's own default +// for the non-throwing path — suppresses that pre-flight throw so large +// max_tokens (our model-ceiling default) works. Per-request duration is still +// bounded by the abort signal NeuroLink composes for each call. +const ANTHROPIC_CLIENT_TIMEOUT_MS = 600_000; /** * Anthropic Provider v2 - BaseProvider Implementation @@ -797,6 +807,7 @@ export class AnthropicProvider extends BaseProvider { apiKey: "oauth-authenticated", // Placeholder, actual auth is in fetch wrapper // Note: No headers passed - fetch wrapper sets oauth-2025-04-20 beta header fetch: oauthFetch as unknown as typeof globalThis.fetch, + timeout: ANTHROPIC_CLIENT_TIMEOUT_MS, }); logger.debug( "[AnthropicProvider] Anthropic SDK client created with OAuth fetch wrapper", @@ -850,6 +861,7 @@ export class AnthropicProvider extends BaseProvider { defaultHeaders: headers, ...(normalizedBaseURL && { baseURL: normalizedBaseURL }), fetch: createProxyFetch() as unknown as typeof globalThis.fetch, + timeout: ANTHROPIC_CLIENT_TIMEOUT_MS, }); logger.debug("Anthropic Provider initialized with API key", { @@ -1440,7 +1452,7 @@ export class AnthropicProvider extends BaseProvider { const params: Anthropic.Messages.MessageCreateParamsNonStreaming = { model: modelId, messages, - max_tokens: options.maxOutputTokens ?? ANTHROPIC_DEFAULT_MAX_TOKENS, + max_tokens: resolveClaudeMaxTokens(modelId, options.maxOutputTokens), ...(system ? { system } : {}), ...(options.temperature !== undefined && options.temperature !== null ? { temperature: options.temperature } @@ -1456,8 +1468,25 @@ export class AnthropicProvider extends BaseProvider { ...(thinking ? { thinking } : {}), }; + // The 60s anthropic generate default was tuned for the old ~4096 + // max_tokens. Now that the default ceiling is the model's real max, + // a large structured response needs more wall-clock to be produced — + // otherwise the inner controller aborts mid-generation (the AI-SDK + // doGenerate layer doesn't see the caller's `timeout`). Raise the + // floor to 5 min when a large output budget is in play — but only + // when the caller did NOT set an explicit timeout: an explicit value + // is a contract and must never be silently extended. The abort + // signal stays the real bound. + const callerTimeout = (options as { timeout?: number | string }) + .timeout; + const callerSpecifiedTimeout = + callerTimeout !== undefined && callerTimeout !== null; + const generateTimeoutMs = + params.max_tokens > 8192 && !callerSpecifiedTimeout + ? Math.max(getTimeoutForOptions(options), 300_000) + : getTimeoutForOptions(options); const timeoutController = createTimeoutController( - getTimeoutForOptions(options), + generateTimeoutMs, providerName, "generate", ); @@ -1751,7 +1780,7 @@ export class AnthropicProvider extends BaseProvider { const params: Anthropic.Messages.MessageCreateParamsStreaming = { model: modelId, messages: conversation, - max_tokens: options.maxTokens ?? ANTHROPIC_DEFAULT_MAX_TOKENS, + max_tokens: resolveClaudeMaxTokens(modelId, options.maxTokens), stream: true, ...(payload.system ? { system: payload.system } : {}), ...(options.temperature !== undefined && options.temperature !== null diff --git a/src/lib/providers/googleVertex.ts b/src/lib/providers/googleVertex.ts index 1bb62aa3e..e95232aac 100644 --- a/src/lib/providers/googleVertex.ts +++ b/src/lib/providers/googleVertex.ts @@ -54,6 +54,7 @@ import { hasRestrictedOutputLimit, RESTRICTED_OUTPUT_TOKEN_LIMIT, } from "../utils/modelDetection.js"; +import { resolveClaudeMaxTokens } from "../utils/tokenLimits.js"; import { validateApiKey, createVertexProjectConfig, @@ -3066,7 +3067,11 @@ export class GoogleVertexProvider extends BaseProvider { const requestParams: Parameters[0] = { model: modelName, - max_tokens: options.maxTokens || 4096, + // Default to the model's real output ceiling (e.g. 64K for Sonnet 4.x) + // instead of the legacy 4096, which silently truncated large structured + // responses mid-JSON. resolveClaudeMaxTokens also clamps over-large + // caller values so the native Vertex path never 400s. + max_tokens: resolveClaudeMaxTokens(modelName, options.maxTokens), messages: messages as Parameters< typeof client.messages.stream >[0]["messages"], @@ -3744,7 +3749,8 @@ export class GoogleVertexProvider extends BaseProvider { const requestParams = { model: modelName, - max_tokens: options.maxTokens || 4096, + // Default to the model's real output ceiling (see stream path note). + max_tokens: resolveClaudeMaxTokens(modelName, options.maxTokens), messages, ...(tools && tools.length > 0 && { tools }), ...(useFinalResultTool && { tool_choice: { type: "any" as const } }), @@ -3770,6 +3776,10 @@ export class GoogleVertexProvider extends BaseProvider { }> = []; let totalInputTokens = 0; let totalOutputTokens = 0; + // Track the final Anthropic stop_reason so we can surface finishReason + // (notably "length" on token truncation) — the legacy native path always + // reported "stop", hiding truncation from callers. + let lastStopReason: string | null | undefined; const currentMessages = [...messages]; while (step < maxSteps) { @@ -3793,6 +3803,7 @@ export class GoogleVertexProvider extends BaseProvider { // Update token counts totalInputTokens += response.usage?.input_tokens || 0; totalOutputTokens += response.usage?.output_tokens || 0; + lastStopReason = response.stop_reason; // Check if we need to handle tool use const toolUseBlocks = ( @@ -4004,6 +4015,10 @@ export class GoogleVertexProvider extends BaseProvider { const result: EnhancedGenerateResult = { content: finalText, + // Surface truncation: Anthropic "max_tokens" → unified "length" so the + // SDK boundary can flag/observe incomplete structured output. Anything + // else (end_turn / stop_sequence / tool_use) is a normal stop. + finishReason: lastStopReason === "max_tokens" ? "length" : "stop", provider: this.providerName, model: modelName, usage: { diff --git a/src/lib/types/generate.ts b/src/lib/types/generate.ts index e2122b8ab..b9b07404c 100644 --- a/src/lib/types/generate.ts +++ b/src/lib/types/generate.ts @@ -270,30 +270,32 @@ export type GenerateOptions = { /** * Zod schema for structured output validation * - * @important Google Gemini Limitation - * Google Vertex AI and Google AI Studio cannot combine function calling with - * structured output. You MUST use `disableTools: true` when using schemas with - * Google providers. - * - * Error without disableTools: "Function calling with a response mime type: - * 'application/json' is unsupported" - * - * This is a documented Google API limitation, not a NeuroLink bug. - * All frameworks (LangChain, Vercel AI SDK, Agno, Instructor) use this approach. + * @important Google GEMINI limitation (Gemini models only) + * Gemini models (Google AI Studio, and Vertex GEMINI models) cannot combine + * function calling with schema-enforced structured output — a Gemini API + * limitation ("Function calling with a response mime type: + * 'application/json' is unsupported"). Vertex CLAUDE models and all other + * providers support tools + schema simultaneously. + * + * You do NOT need to set `disableTools` yourself: when the combination is + * impossible, NeuroLink automatically falls back to text-mode JSON coercion + * (see `coerceJsonToSchema`), and `disableTools: true` remains available as + * an explicit override. * * @example * ```typescript - * // ✅ Correct for Google providers + * // ✅ Vertex + Claude: tools AND schema together are fully supported * const result = await neurolink.generate({ * schema: MySchema, * provider: "vertex", - * disableTools: true // Required for Google + * model: "claude-sonnet-4-6", * }); * - * // ✅ No restriction for other providers + * // ✅ Gemini + tools: SDK auto-falls back to coerced text-mode JSON * const result = await neurolink.generate({ * schema: MySchema, - * provider: "openai" // Works without disableTools + * provider: "google-ai", + * model: "gemini-2.5-pro", * }); * ``` * @@ -321,16 +323,18 @@ export type GenerateOptions = { /** * Disable tool execution (including built-in tools) * - * @required For Google Gemini providers when using schemas - * Google Vertex AI and Google AI Studio require this flag when using - * structured output (schemas) due to Google API limitations. + * Optional with schemas: the tools↔schema exclusion applies only to Google + * GEMINI models (Google AI Studio / Vertex Gemini — a Gemini API + * limitation), and NeuroLink handles it automatically by falling back to + * text-mode JSON coercion. Vertex CLAUDE models support tools + schema + * together. Set this only when you explicitly want a tool-free call. * * @example * ```typescript - * // Required for Google providers with schemas + * // Explicit override: schema-only call with no tools at all * await neurolink.generate({ * schema: MySchema, - * provider: "vertex", + * provider: "google-ai", * disableTools: true * }); * ``` @@ -618,6 +622,13 @@ export type AdditionalMemoryUser = { */ export type GenerateResult = { content: string; // Primary output + /** + * Parsed structured object when a `schema` was requested. Populated from + * AI-SDK experimental_output, or from text-mode coercion (balanced-scan + + * jsonrepair). Prefer this over JSON.parse(content) — it never requires the + * caller to re-parse hand-escaped model text. + */ + structuredData?: unknown; outputs?: { text: string }; // Future extensible for multi-modal /** @@ -708,6 +719,18 @@ export type GenerateResult = { // Finish reason from the AI provider (e.g., "stop", "length", "tool-calls") finishReason?: string; + /** + * True when the schema JSON in `content`/`structuredData` was repaired from + * malformed model text (jsonrepair ran). The result is still valid JSON. + */ + jsonRepaired?: boolean; + /** + * True when the schema JSON appears truncated — the model hit the output + * token cap (finishReason="length") or the recovered object came from an + * unclosed span. `structuredData` may be incomplete; raise `maxTokens`. + */ + jsonTruncated?: boolean; + // Usage and performance usage?: TokenUsage; responseTime?: number; @@ -1224,7 +1247,13 @@ export type TextGenerationOptions = { */ export type TextGenerationResult = { content: string; + /** Parsed structured object when a `schema` was requested (see GenerateResult.structuredData). */ + structuredData?: unknown; finishReason?: string; + /** True when the schema JSON was repaired from malformed model text. */ + jsonRepaired?: boolean; + /** True when the schema JSON appears truncated (output hit the token cap). */ + jsonTruncated?: boolean; provider?: string; model?: string; usage?: TokenUsage; diff --git a/src/lib/types/utilities.ts b/src/lib/types/utilities.ts index 47786a0c8..a0b102897 100644 --- a/src/lib/types/utilities.ts +++ b/src/lib/types/utilities.ts @@ -313,3 +313,24 @@ export type StepToolResult = { result?: unknown; error?: string; }; + +// ============================================================================= +// JSON COERCION (from utils/json/coerce.ts) +// ============================================================================= + +/** + * Result of coercing arbitrary model text into canonical, valid JSON. + * `content` is a JSON.stringify of the recovered object; `structuredData` is + * the parsed object itself. + */ +export type JsonCoercionResult = { + content: string; + structuredData: unknown; + /** True when jsonrepair altered the model text to make it parse. */ + repaired: boolean; + /** + * True when the recovered object came from a truncated (unclosed) span — + * the response likely hit the output-token cap and data may be incomplete. + */ + truncated: boolean; +}; diff --git a/src/lib/utils/json/coerce.ts b/src/lib/utils/json/coerce.ts new file mode 100644 index 000000000..607aea53c --- /dev/null +++ b/src/lib/utils/json/coerce.ts @@ -0,0 +1,162 @@ +/** + * Coerce arbitrary model text into canonical, syntactically-valid JSON. + * + * Used on the text-mode path (providers/models that could not use AI-SDK + * structured output, e.g. real Gemini + tools). The model hand-writes JSON and + * frequently mis-escapes the content field (bare newline, unescaped quote, + * invalid escape like \d). A balanced-brace scan finds the object span; if + * JSON.parse rejects it, jsonrepair fixes common escaping mistakes; the result + * is re-serialised with JSON.stringify so downstream consumers always receive + * valid JSON. + * + * NOTE: jsonrepair is a heuristic. On content where a lone backslash is + * meaningful (regex/script/Windows path) it may drop the backslash, producing + * valid-but-semantically-altered content. This only affects the residual + * text-mode path — the primary Vertex+Claude path uses experimental_output and + * never reaches here. When jsonrepair changes the input we log at debug level + * so the event is observable. + */ +import { jsonrepair } from "jsonrepair"; +import type { + JsonCoercionResult, + ValidationSchema, +} from "../../types/index.js"; +import { logger } from "../logger.js"; +import { nextBalancedJsonSpan } from "./extract.js"; + +/** True when the schema exposes a Zod-style `safeParse` we can validate with. */ +function hasSafeParse(schema: ValidationSchema): boolean { + return typeof (schema as { safeParse?: unknown }).safeParse === "function"; +} + +/** + * Parse `candidate` as JSON, repairing common escaping mistakes on failure. + * Returns the parsed value plus whether jsonrepair had to alter the text. + */ +function parseOrRepair( + candidate: string, +): { value: unknown; repaired: boolean } | undefined { + try { + return { value: JSON.parse(candidate), repaired: false }; + } catch { + // fall through to repair + } + try { + const repaired = jsonrepair(candidate); + const value = JSON.parse(repaired); + if (repaired !== candidate && logger.shouldLog("debug")) { + logger.debug("[coerceJsonToSchema] jsonrepair altered model output", { + originalLength: candidate.length, + repairedLength: repaired.length, + }); + } + return { value, repaired: repaired !== candidate }; + } catch { + return undefined; + } +} + +/** + * Try to produce canonical JSON from `text`. Returns null when no JSON object + * could be recovered (caller should then keep the raw text). + * + * When `schema` is a Zod schema, candidates that satisfy it are preferred; a + * syntactically-valid-but-schema-failing object is still returned (we guarantee + * JSON *validity*, leaving schema/content checks to the caller's own pipeline). + */ +export function coerceJsonToSchema( + text: string, + schema?: ValidationSchema, +): JsonCoercionResult | null { + if (typeof text !== "string" || text.trim().length === 0) { + return null; + } + + // Ordered candidate substrings, best-formed first: + // 1. every balanced object/array span (clean, common case) + // 2. first "{" or "[" to last "}" or "]" (drops surrounding prose; lets + // jsonrepair fix escaping inside) — root ARRAYS matter for array schemas + // 3. first "{" or "[" to end of text (TRUNCATED output — + // finishReason=length — where the closing bracket was cut off; + // jsonrepair closes it) + // `truncated` marks the first-open-to-end candidate: it is only reachable + // when no balanced span and no first-to-last span matched, i.e. there was no + // closing bracket at all — the signature of token-truncated output. + const candidates: Array<{ text: string; truncated: boolean }> = []; + let searchFrom = 0; + for (;;) { + const found = nextBalancedJsonSpan(text, searchFrom); + if (!found) { + break; + } + candidates.push({ text: found.span, truncated: false }); + searchFrom = found.end; + } + const openIndexes = [text.indexOf("{"), text.indexOf("[")].filter( + (i) => i >= 0, + ); + const firstOpen = openIndexes.length > 0 ? Math.min(...openIndexes) : -1; + const lastClose = Math.max(text.lastIndexOf("}"), text.lastIndexOf("]")); + if (firstOpen >= 0 && lastClose > firstOpen) { + candidates.push({ + text: text.slice(firstOpen, lastClose + 1), + truncated: false, + }); + } + if (firstOpen >= 0) { + candidates.push({ text: text.slice(firstOpen), truncated: true }); + } + + let firstValid: + | { value: unknown; repaired: boolean; truncated: boolean } + | undefined; + let schemaMatch: + | { value: unknown; repaired: boolean; truncated: boolean } + | undefined; + const seen = new Set(); + for (const candidate of candidates) { + if (seen.has(candidate.text)) { + continue; + } + seen.add(candidate.text); + const outcome = parseOrRepair(candidate.text); + if ( + outcome === undefined || + outcome.value === null || + typeof outcome.value !== "object" + ) { + continue; + } + const record = { + value: outcome.value, + repaired: outcome.repaired, + truncated: candidate.truncated, + }; + if (firstValid === undefined) { + firstValid = record; + } + if (schema && hasSafeParse(schema)) { + const safeParseable = schema as { + safeParse: (v: unknown) => { success: boolean }; + }; + if (safeParseable.safeParse(outcome.value).success) { + schemaMatch = record; + break; + } + } else { + // No Zod schema to discriminate — first parseable object wins. + break; + } + } + + const chosen = schemaMatch ?? firstValid; + if (chosen === undefined) { + return null; + } + return { + content: JSON.stringify(chosen.value), + structuredData: chosen.value, + repaired: chosen.repaired, + truncated: chosen.truncated, + }; +} diff --git a/src/lib/utils/json/extract.ts b/src/lib/utils/json/extract.ts index ab4fc26bd..44001627a 100644 --- a/src/lib/utils/json/extract.ts +++ b/src/lib/utils/json/extract.ts @@ -7,6 +7,56 @@ import { parseJsonOrNull } from "./safeParse.js"; +/** + * Find the first balanced JSON object/array span starting at or after + * `fromIndex`. Quote- and escape-aware: braces inside string literals do not + * affect depth. Returns the matched substring and the index just past it, or + * null if no balanced span exists. + */ +export function nextBalancedJsonSpan( + text: string, + fromIndex = 0, +): { span: string; end: number } | null { + for (let start = fromIndex; start < text.length; start++) { + const openChar = text[start]; + if (openChar !== "{" && openChar !== "[") { + continue; + } + const closeChar = openChar === "{" ? "}" : "]"; + let depth = 0; + let inString = false; + let escapeNext = false; + for (let i = start; i < text.length; i++) { + const ch = text[i]; + if (escapeNext) { + escapeNext = false; + continue; + } + if (ch === "\\") { + escapeNext = true; + continue; + } + if (ch === '"') { + inString = !inString; + continue; + } + if (inString) { + continue; + } + if (ch === openChar) { + depth++; + } else if (ch === closeChar) { + depth--; + if (depth === 0) { + return { span: text.substring(start, i + 1), end: i + 1 }; + } + } + } + // Unbalanced from this start — try the next opening char. + } + return null; +} + /** * Extract JSON string from text that may contain surrounding content. * @@ -47,20 +97,23 @@ export function extractJsonStringFromText(text: string): string | null { } } - // Try to find JSON object or array pattern using non-greedy iterative scan. - // Note: [\s\S]*? is non-greedy but can still produce over-spanning matches - // in texts with many braces. This is acceptable as we try-parse each candidate - // and move to the next on failure. A bracket-balancing parser would be more - // precise but significantly more complex for marginal benefit. - const candidateRegex = /(\{[\s\S]*?\}|\[[\s\S]*?\])/g; - let candidate: RegExpExecArray | null; - while ((candidate = candidateRegex.exec(text)) !== null) { + // Scan for balanced JSON object/array spans (quote/escape aware) and return + // the first one that parses. Unlike a non-greedy regex, this never stops at a + // "}" that lives inside a string value, so nested objects are preserved. + let searchFrom = 0; + for (;;) { + const found = nextBalancedJsonSpan(text, searchFrom); + if (!found) { + break; + } try { - JSON.parse(candidate[1]); - return candidate[1]; + JSON.parse(found.span); + return found.span; } catch { - // Try next candidate + // Not valid JSON — resume scanning just past this opening character so a + // valid inner object/array can still be found. } + searchFrom = found.end - found.span.length + 1; } return null; diff --git a/src/lib/utils/tokenLimits.ts b/src/lib/utils/tokenLimits.ts index e61ca7e25..cccbc9792 100644 --- a/src/lib/utils/tokenLimits.ts +++ b/src/lib/utils/tokenLimits.ts @@ -101,6 +101,66 @@ export function getSafeMaxTokens( return requestedMaxTokens; } +/** + * Maximum output tokens supported by a given Anthropic Claude model. + * + * The native Vertex+Claude and native Anthropic message paths send `max_tokens` + * straight to the Anthropic API, which returns 400 if the value exceeds the + * model's published output ceiling. (The AI-SDK path clamps automatically; + * these native paths do not.) This table lets those paths default to the + * model's real ceiling — 64K for Sonnet/Haiku 4.x, 32K for Opus 4.x — instead of + * the legacy 4096 that silently truncated large structured responses. + * + * Unknown identifiers fall back to a safe modern floor (8192). + */ +export function getClaudeMaxOutputTokens(model: string | undefined): number { + const m = (model ?? "").toLowerCase(); + // Claude 4.x family: Opus 4.x = 32K, Sonnet/Haiku 4.x = 64K. + if (/opus[-_.]?4/.test(m)) { + return 32000; + } + if (/sonnet[-_.]?4/.test(m) || /haiku[-_.]?4/.test(m)) { + return 64000; + } + // Claude 3.7 Sonnet supports 64K output. + if (/3[-_.]?7[-_.]?sonnet/.test(m)) { + return 64000; + } + // Claude 3.5 Sonnet / Haiku → 8192. + if (/3[-_.]?5[-_.]?(sonnet|haiku)/.test(m)) { + return 8192; + } + // Claude 3 Opus / Sonnet / Haiku → 4096. + if (/claude-3-(opus|sonnet|haiku)/.test(m) || /3[-_.]?opus/.test(m)) { + return 4096; + } + // Bare family aliases (latest of a family) → assume the modern ceiling. + if (m.includes("opus")) { + return 32000; + } + if (m.includes("sonnet") || m.includes("haiku")) { + return 64000; + } + return 8192; +} + +/** + * Resolve the `max_tokens` to send on a native Anthropic/Claude request: honour + * the caller's value but clamp it to the model's published ceiling, and default + * to that ceiling when the caller did not specify one. Prevents both silent + * truncation (the legacy 4096 default) and 400s from over-large requests. + */ +export function resolveClaudeMaxTokens( + model: string | undefined, + requested?: number, +): number { + const ceiling = getClaudeMaxOutputTokens(model); + if (requested !== undefined && requested !== null && requested > 0) { + return Math.min(requested, ceiling); + } + return ceiling; +} + /** * Validate if maxTokens is safe for a provider/model combination */ diff --git a/test/continuous-test-suite-json-e2e.ts b/test/continuous-test-suite-json-e2e.ts new file mode 100644 index 000000000..8d7db6c86 --- /dev/null +++ b/test/continuous-test-suite-json-e2e.ts @@ -0,0 +1,529 @@ +#!/usr/bin/env tsx +/** + * Continuous Test Suite: JSON validity — END-TO-END (live). + * + * The real verification of the JSON-validity + huge-text fix. It makes LIVE + * calls to actual models across providers, with TOOLS ACTIVE (the exact + * condition that used to disable structured output on Vertex+Claude and force + * fragile hand-parsed JSON), using complex Zod schemas — including: + * - an escaping torture test (code with quotes, backslashes, newlines), + * - deeply-nested config and array-heavy schemas, + * - a HUGE-OUTPUT case (no maxTokens → relies on the model's real output + * ceiling) that used to truncate at the legacy 4096 native-path default. + * + * For every (provider, model, schema) cell it asserts: + * 1. result.content parses as JSON (syntactic validity) + * 2. result.structuredData is present and == content (parsed object exposed) + * 3. the parsed object satisfies the Zod schema (shape — strict cells) + * 4. huge-output cells: result is NOT truncated (the maxTokens fix) + * + * Plus two dedicated huge-text tests on the production cell (Vertex+Claude): + * - complete: no maxTokens → large, valid, non-truncated structured output + * - forced truncation: tiny maxTokens → still valid JSON AND jsonTruncated=true + * (truncation is observable, never silent) + * + * Cells that fail for provider reasons (missing key, quota, unknown model, + * region) are SKIPPED, not failed — only a genuine JSON-validity break fails. + * + * Run: npx tsx test/continuous-test-suite-json-e2e.ts + * npx tsx test/continuous-test-suite-json-e2e.ts --only=vertex,openai + */ +process.env.NEUROLINK_DISABLE_BUILTIN_TOOLS = "true"; +import "dotenv/config"; +import { z } from "zod"; +import { tool } from "ai"; +import { NeuroLink } from "../dist/index.js"; +import { defineSuite, assert, Skip } from "./helpers/harness.js"; + +const { test, runSuite } = defineSuite("JSON Validity E2E (live)"); +const nl = new NeuroLink(); + +// Optional filter: --only=vertex,openai +const onlyArg = process.argv.find((a) => a.startsWith("--only=")); +const onlyProviders = onlyArg + ? onlyArg + .slice("--only=".length) + .split(",") + .map((s) => s.trim()) + : null; + +// A harmless tool that forces tools-active (triggers the tools↔schema gate) +// without giving the model anything dangerous to call. +const pingTool = { + ping: tool({ + description: + "Health check. Returns 'pong'. Do not call unless explicitly asked.", + inputSchema: z.object({ value: z.string().describe("anything") }), + execute: async () => "pong", + }), +}; + +// The shape the SDK returns for a schema request (after this fix). +type SchemaResult = { + content: string; + structuredData?: unknown; + finishReason?: string; + jsonTruncated?: boolean; + jsonRepaired?: boolean; +}; + +// ── Complex schemas ───────────────────────────────────────────────────────── + +// (A) TARA-like: free-form content field is the escaping danger zone. +// All fields are required — OpenAI strict structured output rejects optional +// keys. The escaping/size danger zone is `content`. +const agentSchema = z.object({ + summary: z.string().min(1).max(4000), + attachment: z + .object({ + filename: z.string().min(1), + extension: z.string().min(1), + mimetype: z.string().min(1), + content: z.string().min(1), + }) + .nullable(), +}); + +// (B) Deeply nested config with arrays, enums, numbers, nested objects. +const nestedConfigSchema = z.object({ + project: z.object({ + name: z.string(), + version: z.string(), + private: z.boolean(), + }), + settings: z.object({ + flags: z.array(z.string()), + limits: z.object({ min: z.number(), max: z.number() }), + mode: z.enum(["dev", "staging", "prod"]), + }), + tags: z.array(z.string()).min(1), +}); + +// (C) Array-of-objects with a total. +const arrayHeavySchema = z.object({ + items: z + .array( + z.object({ + id: z.number(), + label: z.string(), + enabled: z.boolean(), + }), + ) + .min(2), + total: z.number(), +}); + +// ── Prompts ───────────────────────────────────────────────────────────────── + +const STRESS_PROMPT = [ + "Produce a structured response with a non-empty `summary` and an `attachment`.", + "The attachment must be a SHELL script (extension `sh`, mimetype `text/x-shellscript`, filename `probe`).", + "The script content MUST include, across multiple lines:", + ' - a line printing a double-quoted phrase: echo "hello world"', + " - a Windows-style path with backslashes: C:\\Users\\test\\app", + ' - a grep with a backslash regex: grep -E "[0-9]\\\\d+" file.txt', + " - at least 5 lines total.", + "Return the script verbatim inside attachment.content.", +].join("\n"); + +const NESTED_PROMPT = + "Return a project configuration object. project: name 'neurolink', version '9.69.3', private true. " + + "settings: flags ['a','b','c'], limits min 1 max 100, mode 'prod'. tags: ['sdk','ai','json']."; + +const ARRAY_PROMPT = + "Return an object with `items` (an array of at least 3 objects, each {id:number, label:string, enabled:boolean}) " + + "and `total` = the number of items. Use ids 1,2,3 and any labels."; + +// Forces a LARGE attachment.content — comfortably more than the legacy 4096 +// output-token default (which silently truncated this mid-JSON), yet well under +// the model's real ceiling. With the fix, this returns complete, valid JSON. +const HUGE_OUTPUT_PROMPT = [ + "Produce a structured response.", + "summary: one sentence describing a deployment probe script.", + "attachment: filename 'probe', extension 'sh', mimetype 'text/x-shellscript'.", + "attachment.content MUST be a bash script with AT LEAST 200 numbered lines.", + "Each line must be a DISTINCT, descriptive echo, for example:", + ' echo "[step 7] validating service health for region ap-south-1 — retry budget remaining: 3"', + "Number the steps 1..200 with varied, realistic operational messages. No blank lines, no comments.", + "Return the FULL script verbatim in attachment.content — do not abbreviate, summarise, or use '...'.", +].join("\n"); + +// ── Matrix ────────────────────────────────────────────────────────────────── + +type SchemaCase = { + name: string; + schema: z.ZodTypeAny; + prompt: string; + /** Per-call output cap. Omit to rely on the SDK's provider default. */ + maxTokens?: number; + /** Per-call timeout (ms). A genuinely large generation needs headroom; the + * Vertex native path uses a 5-min internal bound, so match it for fairness. */ + timeout?: number; + /** When true, assert the result is NOT truncated (the huge-text fix). */ + expectComplete?: boolean; + /** Soft, logged-only content checks (not hard assertions). */ + soft?: (parsed: unknown) => string[]; +}; + +const SCHEMA_CASES: SchemaCase[] = [ + { + name: "agent/escaping-stress", + schema: agentSchema, + prompt: STRESS_PROMPT, + maxTokens: 2500, + soft: (p) => { + const obj = p as z.infer; + const c = obj.attachment?.content ?? ""; + return [ + `content length=${c.length}`, + `has double-quote=${c.includes('"')}`, + `preserves backslash=${c.includes("\\")}`, + `multiline=${c.includes("\n")}`, + ]; + }, + }, + { + name: "nested-config", + schema: nestedConfigSchema, + prompt: NESTED_PROMPT, + maxTokens: 2500, + }, + { + name: "array-heavy", + schema: arrayHeavySchema, + prompt: ARRAY_PROMPT, + maxTokens: 2500, + }, + { + // No maxTokens → exercises the per-model default (64K for Sonnet 4.x, etc). + // Before the fix the Vertex+Claude native path defaulted to 4096 and + // truncated this mid-JSON. + name: "huge-output", + schema: agentSchema, + prompt: HUGE_OUTPUT_PROMPT, + timeout: 300_000, // 5 min — a 260-line generation needs real wall-clock time + expectComplete: true, + soft: (p) => { + const obj = p as z.infer; + const c = obj.attachment?.content ?? ""; + const lines = c ? c.split("\n").length : 0; + return [`content length=${c.length}`, `lines=${lines}`]; + }, + }, +]; + +type Cell = { + provider: string; + model?: string; + /** core cells run all schema cases; breadth cells run a subset. */ + tier: "core" | "breadth"; + /** + * When true, schema conformance is a HARD assertion (the provider enforces + * structured output). When false, conformance is logged but not asserted — + * weak models may emit valid-but-non-conforming JSON, which is a model + * limitation, not an SDK JSON-validity failure. Valid-JSON + structuredData + * consistency + non-truncation are always hard-asserted on core cells. + */ + strictSchema: boolean; +}; + +const CELLS: Cell[] = [ + // ── core: latest production-grade models, all schema cases, strict ── + // THE production cell (Vertex + Claude + tools). + { + provider: "vertex", + model: "claude-sonnet-4-6", + tier: "core", + strictSchema: true, + }, + // Latest Claude on Vertex. + { + provider: "vertex", + model: "claude-opus-4-6", + tier: "core", + strictSchema: true, + }, + // Gemini on Vertex (text-mode tools↔schema path). + { + provider: "vertex", + model: "gemini-2.5-pro", + tier: "core", + strictSchema: true, + }, + // Latest Gemini on Vertex. + { + provider: "vertex", + model: "gemini-3-pro-preview", + tier: "core", + strictSchema: true, + }, + // Direct Anthropic. + { + provider: "anthropic", + model: "claude-sonnet-4-6", + tier: "core", + strictSchema: true, + }, + { + provider: "anthropic", + model: "claude-opus-4-6", + tier: "core", + strictSchema: true, + }, + // Google AI Studio (text-mode). + { + provider: "google-ai", + model: "gemini-2.5-flash", + tier: "core", + strictSchema: true, + }, + // OpenAI. + { + provider: "openai", + model: "gpt-4o-mini", + tier: "core", + strictSchema: true, + }, + + // ── breadth: valid-JSON + structuredData are hard-asserted; schema + // conformance is soft-logged (live model/infra variance under back-to-back + // load must not flake the SDK-guarantee verification). ── + { + provider: "vertex", + model: "gemini-2.5-flash", + tier: "breadth", + strictSchema: false, + }, + { + provider: "google-ai", + model: "gemini-3-pro-preview", + tier: "breadth", + strictSchema: false, + }, + { + provider: "mistral", + model: "mistral-large-latest", + tier: "breadth", + strictSchema: false, + }, + { provider: "xai", model: "grok-3", tier: "breadth", strictSchema: false }, + { + provider: "groq", + model: "llama-3.3-70b-versatile", + tier: "breadth", + strictSchema: false, + }, + { + provider: "deepseek", + model: "deepseek-chat", + tier: "breadth", + strictSchema: false, + }, + { + provider: "openrouter", + model: "openai/gpt-4o-mini", + tier: "breadth", + strictSchema: false, + }, + { provider: "bedrock", tier: "breadth", strictSchema: false }, // provider default model +]; + +// Provider/model/auth errors → SKIP (not a JSON-validity failure). +// Deliberately narrow: only infra-shaped statuses (429/5xx) are skippable — +// generic 4xx / "bad request" would mask real request-construction bugs. +function isInfraError(message: string): boolean { + return /api key|apikey|credential|security token|unauthor|permission|quota|rate.?limit|too many requests|429|not found|unknown model|model.*not|region|ENOTFOUND|ECONNREFUSED|ECONNRESET|socket hang|fetch failed|network|timeout|deadline|unavailable|overloaded|throttl|capacity|exhausted|billing|credits|insufficient|payment|402|access|forbidden|invalid.*model|does not exist|status (429|500|502|503|504)|internal server|service unavailable|max_tokens is too large|supports at most|completion tokens|maximum context|context length/i.test( + message, + ); +} + +async function runCell(cell: Cell, sc: SchemaCase): Promise { + let res: SchemaResult; + try { + res = (await nl.generate({ + input: { text: sc.prompt }, + provider: cell.provider, + ...(cell.model ? { model: cell.model } : {}), + schema: sc.schema, + tools: pingTool, + disableTools: false, + temperature: 0, + ...(sc.maxTokens ? { maxTokens: sc.maxTokens } : {}), + ...(sc.timeout ? { timeout: sc.timeout } : {}), + })) as SchemaResult; + } catch (err) { + const msg = err instanceof Error ? err.message : String(err); + if (isInfraError(msg)) { + throw new Skip(`${cell.provider}: ${msg.slice(0, 120)}`); + } + throw err; + } + + // (1) content must be valid JSON — neurolink's core guarantee, HARD for all. + let parsed: unknown; + try { + parsed = JSON.parse(res.content); + } catch (e) { + throw new Error( + `content is NOT valid JSON (${(e as Error).message}). head: ${JSON.stringify( + res.content.slice(0, 240), + )}`, + { cause: e }, + ); + } + + // (2) structuredData present and == content — HARD for ALL cells (an SDK + // boundary guarantee, never model-dependent: the SDK exposes the parsed + // value for object, array, and scalar JSON roots alike). + const sdPresent = + res.structuredData !== undefined && res.structuredData !== null; + const sdConsistent = + sdPresent && JSON.stringify(res.structuredData) === JSON.stringify(parsed); + assert( + sdPresent, + "result.structuredData is missing (the parsed value was not exposed)", + ); + assert( + sdConsistent, + "result.structuredData does not match content (inconsistent exposure)", + ); + + // (3) schema conformance — HARD only where the provider enforces structured + // output. Logged for weak breadth models (valid JSON is still guaranteed). + const sp = sc.schema.safeParse(parsed); + if (cell.strictSchema) { + assert( + sp.success, + `schema validation failed: ${JSON.stringify( + sp.success ? [] : sp.error.issues.slice(0, 3), + )}`, + ); + } else if (!sp.success) { + console.log( + ` · ${cell.provider}/${sc.name}: valid JSON but NON-conforming shape (model limitation): ${JSON.stringify( + sp.error.issues.slice(0, 1), + )}`, + ); + } + + // (4) huge-output: with no maxTokens, the SDK default must NOT truncate — + // HARD on core cells (the maxTokens fix). On breadth, log only. + if (sc.expectComplete) { + const truncated = + res.finishReason === "length" || res.jsonTruncated === true; + if (cell.tier === "core") { + assert( + !truncated, + `huge output was TRUNCATED (finishReason=${res.finishReason}, jsonTruncated=${res.jsonTruncated}) — the maxTokens default did not protect a large response`, + ); + } else if (truncated) { + console.log( + ` · ${cell.provider}/${sc.name}: truncated (finishReason=${res.finishReason}, jsonTruncated=${res.jsonTruncated})`, + ); + } + } + + // Soft, logged-only observations (e.g. backslash preservation, sizes). + if (sc.soft) { + const notes = sc.soft(parsed); + console.log(` · ${cell.provider}/${sc.name}: ${notes.join(" ")}`); + } +} + +// ── Register matrix tests ─────────────────────────────────────────────────── +for (const cell of CELLS) { + if (onlyProviders && !onlyProviders.includes(cell.provider)) { + continue; + } + // core runs all cases (incl. the slow huge-output); breadth runs only the + // escaping-stress case — breadth providers don't exercise the changed + // native-Claude cap path, and keeping the run tractable avoids load-flakes. + const cases = + cell.tier === "core" + ? SCHEMA_CASES + : SCHEMA_CASES.filter((c) => c.name === "agent/escaping-stress"); + for (const sc of cases) { + const label = `${cell.provider}${cell.model ? `:${cell.model}` : ""} — ${sc.name}`; + await test(label, async () => { + await runCell(cell, sc); + }); + } +} + +// ── Dedicated huge-text tests on the production cell (Vertex+Claude) ───────── + +await test("vertex:claude-sonnet-4-6 — huge output is complete (no maxTokens)", async () => { + let res: SchemaResult; + try { + res = (await nl.generate({ + input: { text: HUGE_OUTPUT_PROMPT }, + provider: "vertex", + model: "claude-sonnet-4-6", + schema: agentSchema, + tools: pingTool, + disableTools: false, + temperature: 0, + timeout: 300_000, // 5 min — a complete 260-line script takes real time + // NO maxTokens — must default to the model ceiling, not the old 4096. + })) as SchemaResult; + } catch (err) { + const msg = err instanceof Error ? err.message : String(err); + if (isInfraError(msg)) { + throw new Skip(`vertex: ${msg.slice(0, 120)}`); + } + throw err; + } + const parsed = JSON.parse(res.content) as z.infer; + assert( + res.structuredData !== undefined && res.structuredData !== null, + "structuredData missing", + ); + assert( + res.finishReason !== "length" && res.jsonTruncated !== true, + `huge output truncated (finishReason=${res.finishReason}, jsonTruncated=${res.jsonTruncated})`, + ); + const len = parsed.attachment?.content?.length ?? 0; + // The legacy 4096 cap truncated near ~12-16KB; a complete 260-line script is + // well beyond that. Assert a generous floor so model brevity doesn't flake. + assert( + len > 4000, + `expected a large complete script, got attachment.content length=${len}`, + ); + console.log(` · complete huge output: attachment.content length=${len}`); +}); + +await test("vertex:claude-sonnet-4-6 — forced truncation is observable, never silent", async () => { + let res: SchemaResult; + try { + res = (await nl.generate({ + input: { text: HUGE_OUTPUT_PROMPT }, + provider: "vertex", + model: "claude-sonnet-4-6", + schema: agentSchema, + tools: pingTool, + disableTools: false, + temperature: 0, + maxTokens: 200, // force a cut-off mid-JSON (260-line script won't fit) + })) as SchemaResult; + } catch (err) { + const msg = err instanceof Error ? err.message : String(err); + if (isInfraError(msg)) { + throw new Skip(`vertex: ${msg.slice(0, 120)}`); + } + throw err; + } + // Even under forced truncation the SDK coerces to valid JSON (jsonrepair + // closes the cut-off span), so parsing is asserted UNCONDITIONALLY; and the + // truncation must be OBSERVABLE — flagged via jsonTruncated/finishReason, + // never silently returned as if complete. + JSON.parse(res.content); // must be valid JSON, truncated or not + assert( + res.jsonTruncated === true || res.finishReason === "length", + `truncation was not surfaced (jsonTruncated=${res.jsonTruncated}, finishReason=${res.finishReason})`, + ); + console.log( + ` · forced truncation flagged: jsonTruncated=${res.jsonTruncated}, finishReason=${res.finishReason}`, + ); +}); + +await runSuite(); diff --git a/test/continuous-test-suite-json.ts b/test/continuous-test-suite-json.ts new file mode 100644 index 000000000..f05b0d31c --- /dev/null +++ b/test/continuous-test-suite-json.ts @@ -0,0 +1,273 @@ +#!/usr/bin/env tsx +/** + * Continuous Test Suite: JSON validity — pure helpers (no API). + * + * Deterministic unit coverage for the building blocks of the JSON-validity fix: + * - structured-output policy predicate (Gemini-only tools↔schema exclusion) + * - balanced-brace JSON extractor (nested brace-in-string safety) + * - schema coercion of mis-escaped model text (jsonrepair-backed) + * + * These run without provider keys. The END-TO-END behaviour (real models + * hand-writing JSON across providers) is verified by + * continuous-test-suite-json-e2e.ts, which makes live calls. + * + * Run: npx tsx test/continuous-test-suite-json.ts + */ +import { z } from "zod"; +import { + defineSuite, + assert, + assertEqual, + assertNotNull, +} from "./helpers/harness.js"; +import { + isGeminiProvider, + isToolsSchemaConflictError, + isToolsSchemaExclusionInForce, +} from "../src/lib/core/modules/structuredOutputPolicy.js"; +import { extractJsonStringFromText } from "../src/lib/utils/json/extract.js"; +import { coerceJsonToSchema } from "../src/lib/utils/json/coerce.js"; + +const { test, runSuite } = defineSuite("JSON Validity (pure)"); + +// ── Policy predicate ──────────────────────────────────────────────────────── +await test("isGeminiProvider: google-ai is Gemini", () => { + assert( + isGeminiProvider("google-ai", "gemini-2.5-pro") === true, + "google-ai is Gemini", + ); +}); +await test("isGeminiProvider: vertex+gemini is Gemini", () => { + assert( + isGeminiProvider("vertex", "gemini-2.5-pro") === true, + "vertex+gemini is Gemini", + ); +}); +await test("isGeminiProvider: vertex+claude is NOT Gemini", () => { + assert( + isGeminiProvider("vertex", "claude-sonnet-4-6") === false, + "vertex+claude is not Gemini", + ); +}); +await test("isGeminiProvider: anthropic is NOT Gemini", () => { + assert( + isGeminiProvider("anthropic", "claude-sonnet-4-6") === false, + "anthropic is not Gemini", + ); +}); +await test("exclusion fires only for Gemini + tools present", () => { + // The production bug fix: Vertex+Claude+tools must NOT be excluded. + assertEqual( + isToolsSchemaExclusionInForce("vertex", "claude-sonnet-4-6", true, 5), + false, + "vertex+claude+tools", + ); + // Real API limitation: Vertex+Gemini+tools IS excluded. + assertEqual( + isToolsSchemaExclusionInForce("vertex", "gemini-2.5-pro", true, 5), + true, + "vertex+gemini+tools", + ); + // No tools → never excluded. + assertEqual( + isToolsSchemaExclusionInForce("google-ai", "gemini-2.5-pro", false, 0), + false, + "gemini no tools", + ); + assertEqual( + isToolsSchemaExclusionInForce("google-ai", "gemini-2.5-pro", true, 0), + false, + "gemini zero tools", + ); + // Non-Gemini providers never excluded. + assertEqual( + isToolsSchemaExclusionInForce("openai", "gpt-4.1", true, 5), + false, + "openai+tools", + ); +}); + +await test("conflict detector: recognises Groq json+tools rejection", () => { + assert( + isToolsSchemaConflictError( + new Error( + "Groq error: json mode cannot be combined with tool/function calling", + ), + ), + "Groq message should be a tools↔schema conflict", + ); + assert( + isToolsSchemaConflictError( + "response_format is not supported with function calling", + ), + "response_format+function message should match", + ); + assertEqual( + isToolsSchemaConflictError(new Error("rate limit exceeded")), + false, + "unrelated error must not match", + ); + assertEqual( + isToolsSchemaConflictError(undefined), + false, + "undefined must not match", + ); +}); + +// ── Balanced-brace extractor ──────────────────────────────────────────────── +await test("extractor: returns full outer object, not first inner brace", () => { + const input = 'noise {"a":{"b":"}"},"c":1} trailing'; + const got = extractJsonStringFromText(input); + assertNotNull(got, "should extract an object"); + assertEqual(JSON.parse(got as string).c, 1, "full object with c=1"); + assertEqual( + JSON.parse(got as string).a.b, + "}", + "nested brace-in-string preserved", + ); +}); +await test("extractor: prose preamble then object (Vertex+tools text shape)", () => { + const input = 'Here is your result:\n{"summary":"ok","attachment":null}'; + const got = extractJsonStringFromText(input); + assertNotNull(got, "should find object after prose"); + assertEqual(JSON.parse(got as string).summary, "ok", "summary parsed"); +}); +await test("extractor: fenced json code block", () => { + const input = '```json\n{"x":2}\n```'; + const got = extractJsonStringFromText(input); + assertNotNull(got, "should extract from fence"); + assertEqual(JSON.parse(got as string).x, 2, "x parsed"); +}); + +// ── Schema coercion (jsonrepair-backed) ───────────────────────────────────── +const demoSchema = z.object({ + summary: z.string().min(1), + attachment: z + .object({ extension: z.string(), content: z.string() }) + .nullable(), +}); + +await test("coerce: unescaped internal quotes are repaired", () => { + // Realistic model error: forgot to escape an internal quoted phrase. + const bad = '{"summary":"He said "hi" today","attachment":null}'; + const out = coerceJsonToSchema(bad, demoSchema); + assertNotNull(out, "should repair + parse"); + const obj = JSON.parse((out as { content: string }).content); + assert( + typeof obj.summary === "string" && obj.summary.includes("hi"), + "summary text retained", + ); +}); +await test("coerce: trailing comma + missing brace are repaired", () => { + const bad = '{"summary":"ok","attachment":null,'; + const out = coerceJsonToSchema(bad, demoSchema); + assertNotNull(out, "should close + drop trailing comma"); + assertEqual( + JSON.parse((out as { content: string }).content).summary, + "ok", + "summary parsed", + ); +}); +await test("coerce: invalid \\d escape never throws and yields valid JSON", () => { + // jsonrepair turns an invalid \d escape into "d" (drops the backslash) — a + // documented residual risk for code content on the text-mode path only; the + // primary structured-output path never reaches jsonrepair. The contract here + // is only: never throw, and any returned content must be valid JSON. + const out = coerceJsonToSchema( + '{"summary":"use \\d+ here","attachment":null}', + demoSchema, + ); + if (out) { + JSON.parse((out as { content: string }).content); // must not throw + } + assert(true, "coercion did not throw on invalid escape"); +}); +await test("coerce: raw newline inside string is repaired", () => { + const bad = '{"summary":"line1\nline2","attachment":null}'; + const out = coerceJsonToSchema(bad, demoSchema); + assertNotNull(out, "should repair raw newline"); + const obj = JSON.parse((out as { content: string }).content); + assert( + typeof obj.summary === "string" && obj.summary.includes("line1"), + "summary text retained", + ); +}); +await test("coerce: already-valid JSON exposes parsed object", () => { + const good = '{"summary":"ok","attachment":null}'; + const out = coerceJsonToSchema(good, demoSchema); + assertNotNull(out, "valid JSON should coerce"); + assertEqual( + (out as { structuredData: { summary: string } }).structuredData.summary, + "ok", + "object exposed", + ); +}); +await test("coerce: prose-wrapped object is canonicalised", () => { + const wrapped = 'Sure! {"summary":"ok","attachment":null} hope that helps'; + const out = coerceJsonToSchema(wrapped, demoSchema); + assertNotNull(out, "should extract embedded object"); + assertEqual( + (out as { content: string }).content[0], + "{", + "content is canonical JSON", + ); +}); +await test("coerce: pure prose with no JSON returns null", () => { + const out = coerceJsonToSchema( + "just a friendly hello, no json here", + demoSchema, + ); + assertEqual(out, null, "no JSON object -> null"); +}); + +// ── Truncation / repair flags (huge-text observability) ───────────────────── +await test("coerce: truncated (unclosed) JSON sets truncated + repaired", () => { + // Output cut off mid-string (finishReason=length shape): no closing brace. + const truncated = + '{"summary":"big report","attachment":{"extension":"txt","content":"line1\nline2 and the rest was cut'; + const out = coerceJsonToSchema(truncated, demoSchema); + assertNotNull(out, "should still recover valid JSON from truncated text"); + const r = out as { truncated: boolean; repaired: boolean; content: string }; + assert(r.truncated === true, "truncated flag set for an unclosed span"); + assert(r.repaired === true, "repaired flag set (jsonrepair closed it)"); + JSON.parse(r.content); // must remain valid JSON +}); + +await test("coerce: clean JSON sets neither truncated nor repaired", () => { + const out = coerceJsonToSchema( + '{"summary":"ok","attachment":null}', + demoSchema, + ); + assertNotNull(out, "clean JSON coerces"); + const r = out as { truncated: boolean; repaired: boolean }; + assertEqual(r.truncated, false, "clean JSON is not truncated"); + assertEqual(r.repaired, false, "clean JSON is not repaired"); +}); + +await test("coerce: truncated root array is repaired and flagged", () => { + const out = coerceJsonToSchema('["alpha","beta","gam', z.array(z.string())); + assertNotNull(out, "truncated array should be recovered"); + const r = out as { content: string; truncated: boolean }; + assert( + Array.isArray(JSON.parse(r.content)), + "recovered content parses to an array", + ); + assertEqual(r.truncated, true, "unclosed array flagged truncated"); +}); + +await test("coerce: prose-wrapped root array is extracted", () => { + const out = coerceJsonToSchema( + 'Here you go: ["a","b","c"] enjoy', + z.array(z.string()), + ); + assertNotNull(out, "prose-wrapped array should be extracted"); + const r = out as { content: string; truncated: boolean }; + assertEqual( + JSON.stringify(JSON.parse(r.content)), + JSON.stringify(["a", "b", "c"]), + "array extracted verbatim", + ); + assertEqual(r.truncated, false, "balanced array is not truncated"); +}); + +await runSuite();