Skip to content

fix(openai-shim): stop tool-result filler stall on local OpenAI backends - #2062

Closed
jatmn wants to merge 10 commits into
Twigpine:mainfrom
jatmn:fix/issue-2059-tool-results-filler
Closed

jatmn wants to merge 10 commits into
Twigpine:mainfrom
jatmn:fix/issue-2059-tool-results-filler

Conversation

@jatmn

@jatmn jatmn commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Gate the synthetic [Tool results received] assistant boundary so it is no longer inserted on every tool → user conversion. Local OpenAI-compatible hosts (loopback, RFC1918, CGNAT, Docker internal, reserved TLDs) and Ollama never get it; Mistral cloud / CLAUDE_CODE_USE_MISTRAL / Mistral-class model ids on non-local routes still can for Jinja sequencing.
  • Strip prior placeholder-only assistant echoes from outbound history, clear inherited CLAUDE_CODE_USE_MISTRAL on providerOverride, and treat placeholder-only model replies as a continuation nudge.
  • Document llama.cpp / vLLM mitigation notes in docs/advanced-setup.md.

Fixes #2059
Fixes #2039

Related incomplete work: #1977 (still conflicting / changes-requested; this PR is a focused root-cause fix on current main).

Impact

  • user-facing impact: Local / self-hosted OpenAI-compatible sessions (llama.cpp, Ollama, vLLM on local-classified hosts) no longer silently end a turn after echoing [Tool results received].
  • developer/maintainer impact: shouldInjectToolResultSemanticBoundary is the single decision point; isLocalProviderUrl now matches cacheMetrics reserved-TLD / CGNAT classification more closely.

Testing

  • bun run build
  • bun run smoke
  • bun run check
  • focused tests: bun test ./src/services/api/openaiShim/messageConversion.test.ts ./src/services/api/providerConfig.local.test.ts ./src/__tests__/bugfixes.test.ts ./src/services/api/openaiShim.test.ts (371 pass)

Notes

  • provider/model path tested: unit/integration coverage for Mistral inject-on, local/Qwen inject-off, snip reminder history, placeholder echo nudge/scrub, providerOverride env isolation. No live llama.cpp server in this environment.
  • screenshots attached (if UI changed): n/a
  • follow-up work or known limitations:
    • Remote public hosts with Mistral-class model ids still inject (intentional for Jinja).
    • Stale CLAUDE_CODE_USE_MISTRAL=1 on a remote non-Mistral primary route still injects (product flag semantics); overrides clear the flag.
    • --print may briefly stream an echoed placeholder before the continuation nudge recovers.
    • Final reviewed head: will match PR head after push verification.

Summary by CodeRabbit

  • New Features

    • Expanded guidance and configuration for local OpenAI-compatible providers (e.g., llama.cpp, vLLM, and other local servers).
    • Added smarter handling for synthetic tool-result “boundary” injection, including configurable placeholder behavior.
  • Bug Fixes

    • Prevented repeated tool-result placeholder echoes from being treated as meaningful continuation signals.
    • Improved continuation nudging logic when the tool-result placeholder appears mid-session.
  • Documentation

    • Updated advanced setup notes with new provider examples and troubleshooting tips for repeated short phrases.

@coderabbitai

coderabbitai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 45feebaa-4e63-4034-a8c0-39a5b371a58b

📥 Commits

Reviewing files that changed from the base of the PR and between 6bc3f61 and 7e1cea7.

📒 Files selected for processing (9)
  • docs/advanced-setup.md
  • src/__tests__/bugfixes.test.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/providerConfig.ts
  • src/utils/continuation.ts
📜 Recent review details
🧰 Additional context used
📓 Path-based instructions (5)
**/*.{ts,tsx}

📄 CodeRabbit inference engine (AGENTS.md)

TypeScript code in this repository must use strict mode and ESM imports.

**/*.{ts,tsx}: Add or update tests when a TypeScript or TSX change affects behavior.
Run the relevant TypeScript validation checks for changed code, including bun run typecheck and, when applicable, bun run typecheck:type-tests.

Files:

  • src/__tests__/bugfixes.test.ts
  • src/utils/continuation.ts
  • src/services/api/openaiShim.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/providerConfig.ts
**/*

📄 CodeRabbit inference engine (CONTRIBUTING.md)

**/*: Preserve existing repository patterns unless intentionally refactoring them.
Keep changes small, readable, and focused; avoid broad rewrites or unrelated cleanup.
Do not reformat unrelated files, and keep comments useful and concise.
Update documentation when setup, commands, or user-facing behavior changes.
Review AI-generated changes for correctness, style consistency, unnecessary noise, and adherence to project architecture before submitting them.
Provider changes must follow the documented integration patterns in docs/integrations/overview.md and the focused guides under docs/integrations/how-to/.
When changing provider behavior, avoid breaking third-party providers and test the exact provider/model path changed when possible.
Provider pull requests must explicitly identify affected providers, limitations, and follow-up work.
Run the narrowest meaningful validation command for the touched area, and ensure relevant CI checks pass before merging.
Use bun install to install dependencies and the repository's Bun scripts for building, testing, smoke testing, and development.
Dependency changes require a concrete project benefit such as a bug fix, security issue, or approved feature; preference alone is insufficient.
Do not change the project's language, core runtime, or dependency stack, or introduce a new runtime, without prior maintainer agreement.
Keep each pull request focused on one issue or clearly scoped improvement and avoid bundling unrelated fixes, features, or refactors.

Files:

  • src/__tests__/bugfixes.test.ts
  • docs/advanced-setup.md
  • src/utils/continuation.ts
  • src/services/api/openaiShim.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/providerConfig.ts

⚙️ CodeRabbit configuration file

**/*: Apply the OpenClaude maintainer review rubric from AGENTS.md. Review the current diff, not stale discussion context. Separate real blockers from suggestions. Do not request changes for vague style churn. Treat approval as merge-ready from CodeRabbit's side, pending required human review and GitHub Checks. If checks are failing or unavailable, say so clearly instead of implying the PR is fully ready.

Files:

  • src/__tests__/bugfixes.test.ts
  • docs/advanced-setup.md
  • src/utils/continuation.ts
  • src/services/api/openaiShim.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/providerConfig.ts
{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}

⚙️ CodeRabbit configuration file

{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}: Review tests for meaningful coverage of the changed behavior, isolation of global/env/config state, async cleanup, fake timers, provider profile leaks, and Windows-compatible assumptions. Block when risky runtime changes lack focused regression coverage or tests assert implementation details while missing the user-visible behavior.

Files:

  • src/__tests__/bugfixes.test.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/providerConfig.local.test.ts
{README.md,CONTRIBUTING.md,docs/**,.github/pull_request_template.md}

⚙️ CodeRabbit configuration file

{README.md,CONTRIBUTING.md,docs/**,.github/pull_request_template.md}: Review docs for accuracy against current code behavior. Flag security or provider claims that overpromise, stale install commands, missing setup caveats, and instructions that could push users toward unsafe credential handling. Keep purely wording-level suggestions non-blocking.

Files:

  • docs/advanced-setup.md
{src/services/api/**,src/integrations/**,src/utils/model/**,src/utils/provider*.ts,src/commands/provider/**}

⚙️ CodeRabbit configuration file

{src/services/api/**,src/integrations/**,src/utils/model/**,src/utils/provider*.ts,src/commands/provider/**}: Review provider routing, model selection, env precedence, auth/token handling, OpenAI-compatible shims, retries, proxy behavior, and outbound HTTP behavior with high scrutiny. Block on silent default changes, hidden fallback expansion, credential reuse mistakes, hardcoded provider assumptions, or new network reach that is not intentional and documented.

Files:

  • src/services/api/openaiShim.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/providerConfig.ts
🔇 Additional comments (10)
src/services/api/providerConfig.ts (1)

43-120: LGTM!

Also applies to: 276-282, 304-306, 643-651, 662-673

src/services/api/providerConfig.local.test.ts (1)

12-13: LGTM!

Also applies to: 81-96, 105-215

src/services/api/openaiShim/messageConversion.ts (1)

1-5: LGTM!

Also applies to: 24-30, 161-169, 228-232, 283-309

src/services/api/openaiShim.ts (2)

110-111: LGTM!

Also applies to: 770-771, 1296-1301


1049-1058: 🔒 Security & Privacy

No issue here: baseUrl and model are passed explicitly. shouldInjectToolResultSemanticBoundary() can’t fall back to MISTRAL_BASE_URL or MISTRAL_MODEL on this path, so the extra env scrubbing isn’t needed.

			> Likely an incorrect or invalid review comment.
src/services/api/openaiShim/messageConversion.test.ts (1)

139-140: LGTM!

Also applies to: 164-252

src/services/api/openaiShim.test.ts (1)

46-46: LGTM!

Also applies to: 557-557, 8364-8367, 8502-8593

src/utils/continuation.ts (1)

2-2: LGTM!

Also applies to: 108-118, 141-146

src/__tests__/bugfixes.test.ts (1)

114-126: LGTM!

docs/advanced-setup.md (1)

228-247: LGTM!


📝 Walkthrough

Walkthrough

OpenClaude now applies provider-aware tool-result boundary injection, suppresses placeholder echoes in converted history, detects echoed placeholders as continuation signals, expands local endpoint classification, and documents local OpenAI-compatible server configuration.

Changes

Semantic tool-result boundary handling

Layer / File(s) Summary
Provider routing and local endpoint classification
src/services/api/providerConfig.ts, src/services/api/providerConfig.local.test.ts
Adds provider-aware placeholder detection and injection rules, expands local-host classification, and tests local, remote, Mistral, Ollama, and model-selection cases.
OpenAI message conversion and routing
src/services/api/openaiShim.ts, src/services/api/openaiShim/messageConversion.ts, src/services/api/openaiShim/*.test.ts
Makes boundary injection opt-in, filters placeholder-only assistant echoes, wires provider-specific routing, and updates conversion and override tests.
Placeholder echo continuation handling
src/utils/continuation.ts, src/__tests__/bugfixes.test.ts, docs/advanced-setup.md
Detects placeholder-only responses as continuation signals, adds regression coverage, and documents local OpenAI-compatible server configuration and repetition controls.

Estimated code review effort: 4 (Complex) | ~45 minutes

Suggested reviewers: 0xfandom, kevincodex1, chioarub

🚥 Pre-merge checks | ✅ 6 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 30.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (6 passed)
Check name Status Explanation
Title check ✅ Passed The title is concise and accurately reflects the main openai-shim fix for local OpenAI backends.
Description check ✅ Passed The description follows the template sections and covers summary, impact, testing, and notes.
Linked Issues check ✅ Passed The changes address both linked issues by removing the synthetic marker from history and adding continuation and regression coverage.
Out of Scope Changes check ✅ Passed No clear unrelated changes appear; docs, shim, and tests all support the placeholder-boundary fix.
Risk Surface Disclosed ✅ Passed Risk surface is disclosed: docs/comments flag local/Ollama/providerOverride stall risk, and the fix gates it to Mistral-only paths with no blocker added.
No Hidden Policy Change ✅ Passed Policy-sensitive routing/trust changes are explicit in dedicated helpers, tests, and docs; no hidden default change is buried in cleanup.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/services/api/providerConfig.ts`:
- Around line 662-682: Remove the mappedDotted declaration and its associated
IPv4 conversion and return logic from the hostname classification code. Keep the
mappedHex handling unchanged, as it processes the normalized IPv4-mapped IPv6
form used by current callers.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 7b4ff0a1-3ac2-40d6-accd-0ca26b2e3643

📥 Commits

Reviewing files that changed from the base of the PR and between c2030bb and 6bc3f61.

📒 Files selected for processing (9)
  • docs/advanced-setup.md
  • src/__tests__/bugfixes.test.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/providerConfig.ts
  • src/utils/continuation.ts
📜 Review details
⏰ Context from checks skipped due to timeout. (2)
  • GitHub Check: smoke-and-tests (22)
  • GitHub Check: smoke-and-tests (24.11.x)
🧰 Additional context used
📓 Path-based instructions (5)
**/*.{ts,tsx}

📄 CodeRabbit inference engine (AGENTS.md)

TypeScript code in this repository must use strict mode and ESM imports.

**/*.{ts,tsx}: Add or update tests when a TypeScript or TSX change affects behavior.
Run the relevant TypeScript validation checks for changed code, including bun run typecheck and, when applicable, bun run typecheck:type-tests.

Files:

  • src/__tests__/bugfixes.test.ts
  • src/utils/continuation.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/openaiShim.ts
  • src/services/api/providerConfig.ts
**/*

📄 CodeRabbit inference engine (CONTRIBUTING.md)

**/*: Preserve existing repository patterns unless intentionally refactoring them.
Keep changes small, readable, and focused; avoid broad rewrites or unrelated cleanup.
Do not reformat unrelated files, and keep comments useful and concise.
Update documentation when setup, commands, or user-facing behavior changes.
Review AI-generated changes for correctness, style consistency, unnecessary noise, and adherence to project architecture before submitting them.
Provider changes must follow the documented integration patterns in docs/integrations/overview.md and the focused guides under docs/integrations/how-to/.
When changing provider behavior, avoid breaking third-party providers and test the exact provider/model path changed when possible.
Provider pull requests must explicitly identify affected providers, limitations, and follow-up work.
Run the narrowest meaningful validation command for the touched area, and ensure relevant CI checks pass before merging.
Use bun install to install dependencies and the repository's Bun scripts for building, testing, smoke testing, and development.
Dependency changes require a concrete project benefit such as a bug fix, security issue, or approved feature; preference alone is insufficient.
Do not change the project's language, core runtime, or dependency stack, or introduce a new runtime, without prior maintainer agreement.
Keep each pull request focused on one issue or clearly scoped improvement and avoid bundling unrelated fixes, features, or refactors.

Files:

  • src/__tests__/bugfixes.test.ts
  • docs/advanced-setup.md
  • src/utils/continuation.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/openaiShim.ts
  • src/services/api/providerConfig.ts

⚙️ CodeRabbit configuration file

**/*: Apply the OpenClaude maintainer review rubric from AGENTS.md. Review the current diff, not stale discussion context. Separate real blockers from suggestions. Do not request changes for vague style churn. Treat approval as merge-ready from CodeRabbit's side, pending required human review and GitHub Checks. If checks are failing or unavailable, say so clearly instead of implying the PR is fully ready.

Files:

  • src/__tests__/bugfixes.test.ts
  • docs/advanced-setup.md
  • src/utils/continuation.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/openaiShim.ts
  • src/services/api/providerConfig.ts
{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}

⚙️ CodeRabbit configuration file

{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}: Review tests for meaningful coverage of the changed behavior, isolation of global/env/config state, async cleanup, fake timers, provider profile leaks, and Windows-compatible assumptions. Block when risky runtime changes lack focused regression coverage or tests assert implementation details while missing the user-visible behavior.

Files:

  • src/__tests__/bugfixes.test.ts
  • src/services/api/providerConfig.local.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/openaiShim.test.ts
{README.md,CONTRIBUTING.md,docs/**,.github/pull_request_template.md}

⚙️ CodeRabbit configuration file

{README.md,CONTRIBUTING.md,docs/**,.github/pull_request_template.md}: Review docs for accuracy against current code behavior. Flag security or provider claims that overpromise, stale install commands, missing setup caveats, and instructions that could push users toward unsafe credential handling. Keep purely wording-level suggestions non-blocking.

Files:

  • docs/advanced-setup.md
{src/services/api/**,src/integrations/**,src/utils/model/**,src/utils/provider*.ts,src/commands/provider/**}

⚙️ CodeRabbit configuration file

{src/services/api/**,src/integrations/**,src/utils/model/**,src/utils/provider*.ts,src/commands/provider/**}: Review provider routing, model selection, env precedence, auth/token handling, OpenAI-compatible shims, retries, proxy behavior, and outbound HTTP behavior with high scrutiny. Block on silent default changes, hidden fallback expansion, credential reuse mistakes, hardcoded provider assumptions, or new network reach that is not intentional and documented.

Files:

  • src/services/api/providerConfig.local.test.ts
  • src/services/api/openaiShim/messageConversion.test.ts
  • src/services/api/openaiShim.test.ts
  • src/services/api/openaiShim/messageConversion.ts
  • src/services/api/openaiShim.ts
  • src/services/api/providerConfig.ts
🔇 Additional comments (9)
src/services/api/providerConfig.ts (1)

304-306: 🚀 Performance & Scalability

CGNAT range now treated as "local" everywhere isLocalProviderUrl is consulted.

Classifying 100.64.0.0/10 as local is reasonable for Tailscale/self-hosted setups, but isLocalProviderUrl gates more than the semantic-boundary decision — it also drives stream_options omission, fast-path skipping of tool-history compression/stable-stringify, and local retry-base-URL promotion in openaiShim.ts. CGNAT space is also used by some ISPs for public-facing NAT, so a legitimately remote (non-self-hosted) endpoint reachable via a CGNAT-assigned address would silently get the "local" fast path treatment across all of these unrelated behaviors. Likely an acceptable tradeoff given the target use case, but worth confirming the wider blast radius is intended rather than scoped only to this feature.

src/services/api/providerConfig.local.test.ts (1)

12-13: LGTM!

Also applies to: 81-93, 102-212

src/services/api/openaiShim/messageConversion.ts (1)

1-5: LGTM!

Also applies to: 24-30, 161-169, 228-232, 283-309

src/services/api/openaiShim.ts (1)

110-111: LGTM!

Also applies to: 770-771, 1137-1146, 1479-1484

src/services/api/openaiShim/messageConversion.test.ts (1)

139-140: LGTM!

Also applies to: 164-176, 177-201, 203-222, 224-235, 237-252

src/services/api/openaiShim.test.ts (1)

46-46: LGTM!

Also applies to: 557-557, 8364-8367, 8502-8593

src/utils/continuation.ts (1)

2-2: LGTM!

Also applies to: 108-118, 141-146

src/__tests__/bugfixes.test.ts (1)

114-126: LGTM!

docs/advanced-setup.md (1)

236-242: 📐 Maintainability & Code Quality | ⚡ Quick win

Doc list of local TLDs is incomplete vs. code.

isLocalProviderUrl in providerConfig.ts also treats .internal and .intranet suffixes as local, but this doc only mentions .local / .localhost / .home.arpa / .lan. A self-hoster reasoning about whether their hostname (e.g. foo.internal) avoids the Mistral placeholder should see the full list. Also, .lan isn't actually an RFC/ICANN-reserved TLD (only .local/.localhost/.home.arpa are RFC-reserved and .internal is ICANN-reserved as of 2024) — worth softening that wording too.

As per path instructions, "Review docs for accuracy against current code behavior."

📝 Suggested fix
-OpenClaude never inserts the Mistral/Devstral tool-boundary placeholder
-(`[Tool results received]`) for local OpenAI-compatible hosts (loopback,
-RFC1918, CGNAT/`100.64/10`, `host.docker.internal`, reserved TLDs such as
-`.local` / `.localhost` / `.home.arpa` / `.lan`) or Ollama endpoints. That
+OpenClaude never inserts the Mistral/Devstral tool-boundary placeholder
+(`[Tool results received]`) for local OpenAI-compatible hosts (loopback,
+RFC1918, CGNAT/`100.64/10`, `host.docker.internal`, local-only suffixes such
+as `.local` / `.localhost` / `.home.arpa` / `.lan` / `.internal` /
+`.intranet`) or Ollama endpoints. That

Source: Path instructions

Comment thread src/services/api/providerConfig.ts Outdated
@jatmn

jatmn commented Jul 29, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed CI + CodeRabbit on 4e82b45c:

  1. Smoke/unit failure — classifying .internal / .intranet as local made https://my-proxy.internal/v1 report as Local OpenAI-compatible, breaking StartupScreen GLM fallback tests. Those suffixes are no longer treated as local in isLocalProviderUrl (corporate/custom proxies stay generic). RFC-ish TLDs (.local / .localhost / .home.arpa / .lan) remain local.
  2. CodeRabbit — removed the unreachable dotted ::ffff:a.b.c.d branch; URL.hostname already normalizes to hextet form.

Focused tests + bun run smoke pass locally on this head.

coderabbitai[bot]
coderabbitai Bot previously approved these changes Jul 29, 2026
jatmn added 10 commits July 29, 2026 06:21
…ackends

Gate the synthetic "[Tool results received]" assistant boundary to Mistral/Devstral routes so local OpenAI-compatible models stop echoing it as end_turn, and nudge continuation if a placeholder-only reply still appears.
Prevent a Mistral parent session from re-injecting the tool-result semantic placeholder into a non-Mistral providerOverride route.
Block the Mistral semantic placeholder on local OpenAI-compatible hosts, strip prior placeholder-only echoes from history, share the placeholder constant for continuation recovery, and cover snip-reminder tool boundaries.
Reuse the shared placeholder echo normalizer when stripping history, and skip injection on Ollama endpoints whose hostname is not loopback.
Include mixtral/magistral/mathstral in the tool-boundary injection heuristic so non-local Mixtral routes keep Jinja-compatible sequencing.
Treat host.docker.internal and Tailscale CGNAT addresses as local so Mistral-class model ids on those backends do not re-enable placeholder injection, and strengthen the providerOverride regression test.
Unwrap ::ffff:a.b.c.d hostnames in isLocalProviderUrl so dual-stack loopback/RFC1918/CGNAT URLs keep the same tool-boundary exclusion as IPv4.
URL parsers expand ::ffff:127.0.0.1 to ::ffff:7f00:1; classify both forms as local for tool-boundary exclusion.
Align isLocalProviderUrl with cacheMetrics reserved-TLD classification so .localhost/.home.arpa/.lan hosts do not re-enable placeholder injection.
Stop treating corporate .internal/.intranet hosts as local so StartupScreen
proxy fallbacks stay generic, and drop the unreachable dotted IPv4-mapped branch.
@jatmn
jatmn force-pushed the fix/issue-2059-tool-results-filler branch from 4e82b45 to 7e1cea7 Compare July 29, 2026 13:23
@leftrk

leftrk commented Aug 9, 2026

Copy link
Copy Markdown

Confirmed on a cloud OpenAI-compatible gateway too — not just local backends.

Running against the opencode.ai gateway (https://opencode.ai/zen/go/v1) with deepseek-v4-flash[1m] (a reasoning model), the stall reproduces exactly as described — 4 times in one session:

# tool_result → end_turn latency output tokens stop_reason assistant content
1 3.0s 6 end_turn [Tool results received]
2 3.3s 6 end_turn [Tool results received]
3 4.3s 6 end_turn [Tool results received]
4 6.0s 6 end_turn [Tool results received]

Every stall: the model echoes the injected placeholder verbatim as its entire reply (6 output tokens) and returns end_turn — turn ends with zero work done. Healthy turns in the same session produce 175–51,004 output tokens, so this is unambiguous.

Since shouldInjectToolResultSemanticBoundary() returns false for this route (non-local host, non-Mistral model id, no CLAUDE_CODE_USE_MISTRAL), the fix covers cloud gateways like opencode.ai, not only llama.cpp/Ollama. This matches the exact root cause in #2039/#2059.

PR is currently a draft and has a merge conflict with main — happy to help rebase if useful. Any ETA on moving this to ready-for-review?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants