Skip to content

fix(gateway): cache thought_signature with correct ID - #1448

Merged
steebchen merged 1 commit into
mainfrom
fix/gemini-thought-signature-caching
Jan 14, 2026
Merged

steebchen merged 1 commit into
mainfrom
fix/gemini-thought-signature-caching

Conversation

@steebchen

@steebchen steebchen commented Jan 14, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Fix Gemini 3 multi-turn conversations failing with thought_signature missing error
  • Cache thoughtSignature in Redis using the same tool_call ID that's sent to clients
  • Remove duplicate caching in extract-tool-calls.ts that was using a mismatched ID

Problem

The gateway was generating tool_call IDs in two different places using Date.now():

  • transform-streaming-to-openai.ts - ID sent to client
  • extract-tool-calls.ts - ID used for Redis caching

Since these were called milliseconds apart, the IDs didn't match. When clients sent back multi-turn conversations, the Redis lookup failed and Gemini rejected the request.

Test plan

  • Build passes
  • Test multi-turn conversation with Gemini 3 Flash Preview using tools

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Chores
    • Optimized tool call processing and caching infrastructure to enhance system reliability and multi-turn conversation performance.

✏️ Tip: You can customize this high-level summary in your review settings.

The gateway was generating tool_call IDs in two places using
Date.now(), causing a mismatch between the ID sent to clients
and the ID used for Redis caching. This broke multi-turn
conversations with Gemini 3 models that require thought_signature.

- Add Redis caching in transform-streaming-to-openai.ts where
  the correct ID is generated
- Remove duplicate caching from extract-tool-calls.ts

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings January 14, 2026 18:37
@coderabbitai

coderabbitai Bot commented Jan 14, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Redis caching for thought_signature is being relocated between two tool-call processing modules. The extraction module removes caching logic, while the streaming transformation module assumes this responsibility by caching thought_signature with a newly structured toolCallId format.

Changes

Cohort / File(s) Summary
Caching removal
apps/gateway/src/chat/tools/extract-tool-calls.ts
Removed Redis caching and error logging for thought_signature. Retains comment indicating caching occurs elsewhere and tool_call ID generation happens in a different module.
Caching addition
apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
Added Redis caching of thought_signature using a structured toolCallId (name_timestamp_index format) with 1-day expiration. Imported redisClient from @llmgateway/cache. Updated inline ID generation for function calls to use new toolCallId format across Google AI Studio path.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • PR #1313: Both PRs modify thought_signature storage/propagation for tool_calls and handle toolCall ID generation logic.
  • PR #1222: Both PRs adjust how Google/Vertex thought_signature data is cached and handled in the tool-call pipeline.
  • PR #860: Directly modifies transform-streaming-to-openai.ts to change function/tool call production with ID generation logic.
🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'fix(gateway): cache thought_signature with correct ID' directly addresses the main change: moving Redis caching to use the correct tool call ID to fix multi-turn Gemini conversations.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a critical bug where Gemini 3 multi-turn conversations with tool calls were failing due to mismatched tool call IDs in Redis caching. The gateway was generating tool call IDs in two different places using Date.now(), causing the IDs sent to clients to differ from the IDs used as Redis cache keys.

Changes:

  • Moved thoughtSignature Redis caching to transform-streaming-to-openai.ts where the actual client-facing tool call ID is generated
  • Removed duplicate Redis caching logic from extract-tool-calls.ts that was using mismatched IDs
  • Added explanatory comments about where caching occurs

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

File Description
apps/gateway/src/chat/tools/transform-streaming-to-openai.ts Added Redis caching of thoughtSignature using the correct tool call ID that gets sent to clients
apps/gateway/src/chat/tools/extract-tool-calls.ts Removed duplicate Redis caching that was using a different tool call ID, added comment explaining the caching location
Comments suppressed due to low confidence (1)

apps/gateway/src/chat/tools/extract-tool-calls.ts:53

  • The ID generated here using Date.now() creates a mismatch with the ID generated in transform-streaming-to-openai.ts (line 428). This ID is used for internal tracking in the streamingToolCalls array (chat.ts line 3658), while a different ID is sent to clients. For Google providers that send complete tool calls in single chunks this may not cause issues, but it creates unnecessary confusion and potential for bugs. Consider removing this ID generation entirely and relying on the ID from the transformed data, or document why two different IDs are needed.
							id: part.functionCall.name + "_" + Date.now() + "_" + index,

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

@steebchen
steebchen enabled auto-merge January 14, 2026 18:41

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
apps/gateway/src/chat/tools/extract-tool-calls.ts (1)

45-53: Fix ID generation mismatch for Google tool calls between extractToolCalls and transform-streaming-to-openai.

The comment correctly notes that Redis caching of thought_signature happens in transform-streaming-to-openai.ts. However, this creates a critical issue: extractToolCalls() generates its own ID using Date.now() at line 53:

id: part.functionCall.name + "_" + Date.now() + "_" + index,

Meanwhile, transform-streaming-to-openai.ts generates a different ID at line 428 with a separate Date.now() call and caches the thought_signature using that ID (line 455).

The extracted tool calls are returned as toolResults and later appear in messages for subsequent requests. When the code attempts to retrieve the cached thought_signature (line 2418 in chat.ts), it uses the extractToolCalls ID, which won't match the Redis cache key created by transform-streaming-to-openai.

This breaks thought_signature retrieval for multi-turn conversations with Google providers. Either:

  1. Remove Google provider ID generation from extractToolCalls and delegate to the transform function
  2. Coordinate ID generation between both modules to ensure consistency
🧹 Nitpick comments (1)
apps/gateway/src/chat/tools/transform-streaming-to-openai.ts (1)

427-465: Core fix looks correct — same ID now used for client and Redis cache.

The fix properly addresses the root cause by:

  1. Generating toolCallId once (line 427-428)
  2. Using it for the tool call sent to clients (line 430)
  3. Using the same ID for Redis cache key (line 455)

The fire-and-forget pattern with .catch() is acceptable for non-critical caching.

Consider extracting the magic number 86400 to a named constant for clarity:

🔧 Optional: Extract expiration to named constant
+const THOUGHT_SIGNATURE_CACHE_TTL_SECONDS = 86400; // 1 day
+
 // Cache thoughtSignature in Redis for server-side retrieval in multi-turn conversations
 // This is especially important when OpenAI SDKs don't preserve extra_content/provider_extra
 if (sig) {
 	redisClient
 		.setex(
 			`thought_signature:${toolCallId}`,
-			86400, // 1 day expiration
+			THOUGHT_SIGNATURE_CACHE_TTL_SECONDS,
 			sig,
 		)
📜 Review details

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 6740fa5 and 3eee084.

📒 Files selected for processing (2)
  • apps/gateway/src/chat/tools/extract-tool-calls.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
🧰 Additional context used
📓 Path-based instructions (7)
**/*.{ts,tsx}

📄 CodeRabbit inference engine (CLAUDE.md)

**/*.{ts,tsx}: Never use any or as any unless absolutely necessary in TypeScript code
For database reads: Use db().query.<table>.findMany() or db().query.<table>.findFirst()

Files:

  • apps/gateway/src/chat/tools/extract-tool-calls.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
**/*.{ts,tsx,js,jsx,json,md}

📄 CodeRabbit inference engine (CLAUDE.md)

Always use tabs for indentation

Files:

  • apps/gateway/src/chat/tools/extract-tool-calls.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
**/*.{ts,tsx,js,jsx}

📄 CodeRabbit inference engine (CLAUDE.md)

**/*.{ts,tsx,js,jsx}: Always use top-level import, never use require or dynamic imports
No unnecessary code comments

Files:

  • apps/gateway/src/chat/tools/extract-tool-calls.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
apps/{gateway,api}/**/*.{ts,tsx}

📄 CodeRabbit inference engine (CLAUDE.md)

Use Hono framework with Zod validation and OpenAPI documentation for backend APIs

Files:

  • apps/gateway/src/chat/tools/extract-tool-calls.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
**/*.{js,ts,tsx,jsx}

📄 CodeRabbit inference engine (AGENTS.md)

Always use top-level import, never use require or dynamic imports

Files:

  • apps/gateway/src/chat/tools/extract-tool-calls.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
{apps/api,apps/gateway,packages/db}/**/*.ts

📄 CodeRabbit inference engine (AGENTS.md)

{apps/api,apps/gateway,packages/db}/**/*.ts: Use Drizzle ORM with latest object syntax for database operations
For database reads: Use db().query.<table>.findMany() or db().query.<table>.findFirst()

Files:

  • apps/gateway/src/chat/tools/extract-tool-calls.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
apps/{gateway,api}/src/**/*.ts

📄 CodeRabbit inference engine (AGENTS.md)

apps/{gateway,api}/src/**/*.ts: Use Hono for backend framework in Gateway and API services
Use Zod schemas for validation in Hono services

Files:

  • apps/gateway/src/chat/tools/extract-tool-calls.ts
  • apps/gateway/src/chat/tools/transform-streaming-to-openai.ts
🧬 Code graph analysis (1)
apps/gateway/src/chat/tools/transform-streaming-to-openai.ts (1)
packages/logger/src/index.ts (1)
  • logger (181-181)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (11)
  • GitHub Check: Agent
  • GitHub Check: e2e-shards (5)
  • GitHub Check: e2e-shards (1)
  • GitHub Check: e2e-shards (2)
  • GitHub Check: e2e-shards (4)
  • GitHub Check: e2e-shards (3)
  • GitHub Check: build / run
  • GitHub Check: test / run
  • GitHub Check: generate / run
  • GitHub Check: lint / run
  • GitHub Check: autofix
🔇 Additional comments (2)
apps/gateway/src/chat/tools/transform-streaming-to-openai.ts (1)

1-1: LGTM on the new import.

The redisClient import is correctly added to support the new caching logic.

apps/gateway/src/chat/tools/extract-tool-calls.ts (1)

61-67: LGTM on retaining extra_content for client passback.

The thoughtSignature is correctly included in extra_content to allow clients that properly handle it to pass it back directly, serving as a fallback alongside the server-side Redis cache.

✏️ Tip: You can disable this entire section by setting review_details to false in your review settings.

@steebchen
steebchen added this pull request to the merge queue Jan 14, 2026
Merged via the queue into main with commit 84155d7 Jan 14, 2026
19 of 20 checks passed
@steebchen
steebchen deleted the fix/gemini-thought-signature-caching branch January 14, 2026 18:50
rcogal pushed a commit that referenced this pull request Jan 15, 2026
## Summary
- Fix Gemini 3 multi-turn conversations failing with `thought_signature`
missing error
- Cache `thoughtSignature` in Redis using the same tool_call ID that's
sent to clients
- Remove duplicate caching in `extract-tool-calls.ts` that was using a
mismatched ID

## Problem
The gateway was generating tool_call IDs in two different places using
`Date.now()`:
- `transform-streaming-to-openai.ts` - ID sent to client  
- `extract-tool-calls.ts` - ID used for Redis caching

Since these were called milliseconds apart, the IDs didn't match. When
clients sent back multi-turn conversations, the Redis lookup failed and
Gemini rejected the request.

## Test plan
- [x] Build passes
- [ ] Test multi-turn conversation with Gemini 3 Flash Preview using
tools

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Optimized tool call processing and caching infrastructure to enhance
system reliability and multi-turn conversation performance.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude <noreply@anthropic.com>
steebchen pushed a commit that referenced this pull request Aug 14, 2026
## Summary

`parseProviderResponse` (Google branch) caches Gemini thought signatures
in
Redis under the id it emits — `${name}_${shortid(24)}`, as
`thought_signature:<id>` — so the signature can be re-injected when the
client
replays the call next turn. The `n > 1` branch of
`transformResponseToOpenai`
(`apps/gateway/src/chat/tools/transform-response-to-openai.ts:454-456`)
regenerated tool calls from the raw parts with
`${name}_${candidateIndex}_${fcIndex}`
ids and dropped `extra_content`, so the id the client echoes back never
matches
the cached key. The next-turn lookup in `chat.ts`
(`redisClient.get("thought_signature:" + toolCall.id)`) misses, the
signature
is not re-injected, and Gemini rejects the replay with *"Corrupted
thought
signature"* — the exact failure the shortid fix (#1448) eliminated on
the
streaming path. The `n > 1` branch is the one path left behind: the
single-candidate path reuses parse's `toolResults` and is consistent.

This is the same "forgotten sibling branch" pattern as #3492 (our
previous
fix) and #3504 (the maintainer's own follow-up).

## Fix

Two changes in the multi-candidate branch:

- **Candidate 0** reuses the tool calls from `toolResults` (what
`parseProviderResponse` already emitted and cached), so the id the
client
  receives is exactly the one the signature is cached under.
- **Candidates 1+** (which parse does not process) get the same
treatment
parse applies to candidate 0: a unique `${name}_${shortid(24)}` id, the
  inline `extra_content.google.thought_signature`, and the
`setex("thought_signature:<id>", 86400, sig)` Redis entry under the
emitted
  id.

Fallback keeps the branch safe: if `toolResults` is absent for candidate
0
(e.g. no tool calls parsed), the shortid path applies.

## Tests

`transform-response-to-openai.spec.ts`:

- The existing multi-candidate test no longer pins the broken
`${name}_${candidateIndex}_${fcIndex}` ids; it asserts the id scheme and
  that candidates are distinct.
- New test: `keeps multi-candidate Google tool_call ids in the
thought_signature
  cache` — mocks `@llmgateway/cache` (`setex`) and asserts:
- choice 0 reuses the parsed tool calls verbatim (id + inline
signature);
- choice 1 gets a unique id, `extra_content` with its own signature, and
a
`setex("thought_signature:<id>", 86400, sig)` call under the emitted id;
  - no `setex` is written under choice 0's id from the transform (parse
    already did that).

## Verification

Every command below was run and its output captured. The record is
reproducible — commands included so you can re-run them.

**red — probe against main (c1cd97b), real functions via esbuild,
before the
fix** (harness preserved at the hub):

```
parse toolResults[0].id      : get_weather_xxxxxxxxxxxxxxxxxxxxxxxx
                             + extra_content.google.thought_signature
Redis writes made by parse   : [{"key":"thought_signature:get_weather_xxxxxxxxxxxxxxxxxxxxxxxx", …}]
transform choice 0 tool_calls: id get_weather_0_0, extra_content undefined
RESULT: MISMATCH — next-turn GET thought_signature:get_weather_0_0 misses;
signature lost, Gemini 3 rejects replay
=== control (n=1) ===
CONTROL RESULT: MATCH — single-candidate path is consistent
```

**green — same probe after the fix**:

```
choice 0 tool_calls: id get_weather_xxxxxxxxxxxxxxxxxxxxxxxx
                     extra_content.google.thought_signature sig-candidate-0
choice 1 tool_calls: id get_weather_xxxxxxxxxxxxxxxxxxxxxxxx
                     extra_content.google.thought_signature sig-candidate-1
Redis writes: thought_signature:get_weather_xxx = sig-candidate-0
              thought_signature:get_weather_xxx = sig-candidate-1
RESULT: MATCH (no bug)
CONTROL RESULT: MATCH — single-candidate path is consistent
```

(The deterministic shortid stub makes both candidate ids equal in the
probe;
the real `shortid` yields unique ids per candidate, asserted in the
spec.)

**green — vitest, full spec**:

```
$ pnpm exec vitest run apps/gateway/src/chat/tools/transform-response-to-openai.spec.ts --no-file-parallelism
 ✓ apps/gateway/src/chat/tools/transform-response-to-openai.spec.ts (14 tests) 24ms
 Test Files  1 passed (1)
      Tests  14 passed (14)
[exit code 0]
```

**Area suite** — `vitest run apps/gateway/src/chat/tools`: **593/603
tests,
37/38 files pass**. The only failing file is
`openai-content-filter.spec.ts` (10 tests), a service-harness spec whose
fetch/Redis mocks time out in this environment; it imports only
`openai-content-filter.ts`, which this change does not touch (same class
of
harness failures documented in previous PRs here).

<details><summary>Environment</summary>

```
so: Windows 11 (AMD64)
node --version: v24.14.1
pnpm --version: 9.15.9 (via npx; repo packageManager is pnpm@10.30.3)
vitest: 4.1.8
base: c1cd97b (origin/main @ 08-09)
fix commit: 7e76bf9 (2 files, +177/−13)
```

</details>

**Gates:** eslint ✅ on both files (after `import/order` fix) · prettier
✅ ·
`git diff --check` clean · gateway typecheck ✅ (builds in `build:core`,
turbo 12/12) · `pnpm-lock.yaml` untouched.

**Collisions:** no open PR touches the multi-candidate Google tool_call
ids /
thought_signature path (checked live 2026-08-11). #3486 (reasoning
tokens)
and #3233 (Claude on Azure) also touch `transform-response-to-openai.ts`
but
in unrelated areas (usage tokens / new provider branch); no semantic
overlap.

Not verified: a live multi-turn round-trip against Gemini (no provider
keys on
this machine). The id↔Redis-key chain is covered end-to-end by the
spec's
`setex` assertions and the probe's recorded Redis writes.

---

Disclosure: an AI coding assistant helped locate the drift and draft the
test
scaffolding. The reproduction, the red→green cycle and this write-up
were
verified by running the code; I own the change and will follow up on
review
comments.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved handling of Google responses containing multiple candidates
and tool calls.
* Preserved candidate-specific thought signatures for reliable follow-up
processing.
* Prevented tool-call identifiers and cached data from being overwritten
across candidates.
  * Added graceful logging when caching metadata cannot be completed.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Codebuff <noreply@codebuff.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants