Staging validation: sf/MLX client-tools healing at 1f7c065e0 - #248
danielhanchen wants to merge 4 commits into
Conversation
PR 6801 made response-side tool-call healing default-on for the client-tool passthrough, but only on the GGUF path: the passthrough branch in /v1/chat/completions is gated on using_gguf, and the safetensors section never reads payload.tools, so a client-tools request against a safetensors or MLX model silently dropped the tool schemas and returned prose with no tool_calls. Add the missing leg. When a non-GGUF model is loaded, the request declares client tools (or carries tool-role history), server-side tools are off, and the template supports tools, the route now: - renders the tools into the chat template for a single turn via the existing backend.generate_chat_response(..., tools=...) seam (worker templating already accepts role=tool and assistant.tool_calls messages, normalized with _openai_messages_for_passthrough); - non-streaming: promotes text-form calls with heal_openai_message, honors the opt-in nudge single retry (nudge_should_retry / nudge_messages), caps healed calls when parallel_tool_calls=false (covers the nudge retry too), and sets finish_reason=tool_calls with content null on a pure tool-call turn; - streaming: derives deltas from the worker's cumulative snapshots and feeds StreamToolCallHealer, emitting healed tool-call deltas and the correct finish chunk, guarded against repeated or shrinking snapshots. heal_gate semantics are identical to the GGUF passthrough: default on, auto_heal_tool_calls=false or UNSLOTH_DISABLE_TOOL_CALL_HEALING=1 relays verbatim, tool_choice narrows promotion, undeclared names stay text. MLX rides the same orchestrator seam, so both local backends gain the behavior. CompletionMessage.content becomes Optional so a promoted pure tool-call turn matches the OpenAI contract (content null when only tool_calls return). Adds tests/test_sf_client_tools_passthrough.py (22 cases: healing, gating, opt-outs, streaming deltas, tool-role history, dict-arguments history, forced tool_choice, parallel cap, usage, nudge on/off/double-failure, generator error hygiene, disconnect reset, empty output, MLX path).
…monitor reply Four review follow-ups on the safetensors/MLX client-tool passthrough leg: - tool_choice="none" keeps the tool-history templating but no longer advertises the tools, so a forced final-answer turn is not prompted into emitting markup that the (correctly disabled) healer would relay as prose. Mirrors the GGUF passthrough where llama-server honors tool_choice itself. - OpenAI "developer" messages fold into a single leading system message via _set_or_prepend_system_message before templating; local templates reject the role and the fallback formatter drops it. - A nudge retry that fails or is cancelled after the original answer exists falls back to the first response instead of surfacing a 500, matching the GGUF nudge path. - The API monitor records the healed tool call summary instead of the raw markup on a promoted turn. Adds four regression tests.
for more information, see https://pre-commit.ci
…g, stream monitor parity - A forced tool_choice function is now the only schema rendered into the local template, so the advertised tools and the healer allowlist can no longer disagree (llama-server enforces tool_choice itself on the GGUF path). - Content-part lists are flattened to their text parts before templating. Remote image URLs are not decodable locally, so such requests reached this path with part lists that raise inside apply_chat_template on text-only templates; the plain non-GGUF path has always flattened them. - The streaming monitor entry is now fed from the healed events the client actually receives, recording promoted calls as the [tool_calls] summary the non-streaming path records.
There was a problem hiding this comment.
Code Review
This pull request implements client-tool passthrough healing for the safetensors and MLX backends, aligning them with the GGUF implementation. It enables the promotion of text-form tool calls back into structured tool calls, supports nudge retries for unparseable markup, caps parallel tool calls, and ensures the API monitor records healed calls rather than raw XML. Additionally, a comprehensive test suite has been added to verify these behaviors. Feedback on the changes suggests simplifying a redundant dictionary get-access to a direct key lookup for consistency.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| ], | ||
| ) | ||
| ) | ||
| _fn = value.get("function") or {} |
There was a problem hiding this comment.
On line 668, value["function"] is accessed directly via a key lookup. If "function" were missing from value, a KeyError would have already been raised there. Therefore, using value.get("function") or {} on line 673 is redundant and inconsistent. We can simplify this to direct key access for consistency and clarity.
| _fn = value.get("function") or {} | |
| _fn = value["function"] |
|
All 46 checks passed at 1f7c065, including the Mac M1 MLX jobs. Validation complete, closing without merge as usual. |
CI validation run for unslothai#6870 at head 1f7c065 (review fixes: forced tool_choice templating, content-part flattening, streaming monitor parity, nudge retry fallback, developer folding). Will be closed after checks complete, not merged.