Skip to content

Fix Mistral tool parser for the [ARGS]-marker format (Ministral 3, Devstral Small 2) - #631

Merged
waybarrios merged 5 commits into
waybarrios:mainfrom
mabaeyens:fix/mistral-args-token-tool-parsing
Aug 3, 2026
Merged

waybarrios merged 5 commits into
waybarrios:mainfrom
mabaeyens:fix/mistral-args-token-tool-parsing

Conversation

@mabaeyens

Copy link
Copy Markdown
Contributor

Summary

Ministral 3 14B and Devstral Small 2 (Dec 2025 tokenizers) emit tool calls as
[TOOL_CALLS]<name>[ARGS]<json arguments> — confirmed directly in their
chat_template.jinja:

{{- '[TOOL_CALLS]' + tool['function']['name'] + '[ARGS]' + arguments }}

The mistral tool parser only knew about the old bracket-JSON-array format and a
name{args} format with no separator, so it mishandled this one in two ways:

  • Non-streaming (extract_tool_calls): split on the first {, so the parsed
    function name came back as "get_weather[ARGS]" instead of "get_weather".
  • Streaming (_parse_streaming_tool_delta): re-classified every incoming delta
    independently
    , using a "does this chunk start with JSON punctuation" heuristic,
    with no memory of having already passed the name/arguments boundary. Bare-word
    JSON string fragments (e.g. city, Paris) have no distinguishing leading
    punctuation, so once already inside the arguments blob they got misclassified as
    more of the function name — reconstructing garbage from any standard OpenAI-style
    delta accumulator (e.g. name ending up as "get_weather[ARGS]city":"Paris"}" and
    arguments as '[ARGS]{"').

Fix

  • Recognize the [ARGS] marker explicitly in both code paths (checked before
    falling back to the older {-only split, so older Mistral checkpoints are
    unaffected).
  • Give the streaming parser persistent per-tool-call state (_args_started,
    _name_buffer), reset whenever a new tool call starts. The name/arguments
    boundary is decided exactly once — text is buffered until the marker is found
    (this also handles the marker itself being split across two deltas), then every
    subsequent delta for that tool call is unconditionally arguments, never
    re-classified.

Testing

Verified against a live vllm-mlx serve mlx-community/Ministral-3-14B-Instruct-2512-4bit --enable-auto-tool-choice --tool-call-parser mistral instance, both streaming and non-streaming:

Before (streaming, delta-by-delta):

name: "get" / "_" / "weather"
arguments: "[ARGS]" / "{\""
name: "city" / "\":" / "\"" / "Paris" / "\"}"   # <- wrong, these are arguments

After:

name: "get_weather"                              # single clean chunk
arguments: "{\"" / "city" / "\":" / " \"" / "Paris" / "\"}"   # reconstructs to {"city": "Paris"}

Added two regression tests replaying the exact delta sequences captured from that
server (including a marker-split-across-deltas case), plus a non-streaming format
test. Full existing suite passes: 113 passed.

Test plan

  • pytest tests/test_tool_parsers.py -k mistral — new + existing Mistral tests pass
  • pytest tests/test_tool_parsers.py — full suite, 113 passed, no regressions
  • Live streaming + non-streaming smoke test against Ministral-3-14B-Instruct-2512-4bit

TimotejLabsky added a commit to TimotejLabsky/vllm-mlx that referenced this pull request Jul 7, 2026
…k of upstream waybarrios#631

Dec-2025 Mistral tokenizers (Devstral Small 2, Ministral 3) emit
[TOOL_CALLS]name[ARGS]{json}; the parser read the name as
"name[ARGS]" and shredded streaming deltas, so Devstral Small 2 tool
calling was fully broken. Older formats untouched. PATCHES.md waybarrios#42.

Cherry-pick of waybarrios#631
(mabaeyens); retire on the next rebase past its merge.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

@Thump604 Thump604 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The [ARGS] boundary handling is directionally correct, but the buffered streaming path can omit the tool-call id entirely.

Using the live delta shape from the PR ([TOOL_CALLS]get, _, weather, [ARGS], arguments), the first three calls return None. When [ARGS] arrives, execution is in the current_tool_id >= 0 branch, which emits the function name without an id; subsequent argument deltas also omit it. Therefore the reconstructed OpenAI stream has a name and arguments but no tool-call ID.

Please retain one generated ID per active tool call and emit it with the first structured delta, even when the marker is split or arrives after the initial [TOOL_CALLS] delta. Add a regression assertion that the accumulated streamed call contains exactly one non-empty, stable ID. The existing focused tests pass because they currently collect only function.name and function.arguments.

mabaeyens added a commit to mabaeyens/vllm-mlx that referenced this pull request Jul 10, 2026
Same fix as upstream PR waybarrios#629 (mvmories) for issue waybarrios#628: mlx-lm's
stream_generate exhausts without ever claiming finished, so the
epilogue yielded finished=True with finish_reason=None on natural
EOS. Only the max_tokens cutoff path stamped a reason ("length").
Local-only commit (not for the open waybarrios#631 PR) so Mira can pick up
the already-verified upstream fix before it merges.
mabaeyens added a commit to mabaeyens/vllm-mlx that referenced this pull request Jul 10, 2026
…KEN delta

The [ARGS]-marker buffering defers the name/arguments boundary past the
delta that contains [TOOL_CALLS], so the id generated in that branch was
frequently never attached to any delta: the middle-of-call branch that
ends up carrying the actual name/arguments never emitted an id at all.
Clients that correlate streamed tool-call deltas by id would see a
call with no id.

Now the id is generated once when a tool call starts and attached to
whichever delta is the first to carry real content, in either branch,
exactly once.

Addresses review feedback from PR waybarrios#631:
waybarrios#631 (review)
@mabaeyens
mabaeyens requested a review from Thump604 July 10, 2026 14:44

@Thump604 Thump604 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The requested tool-call ID correction is now implemented correctly: the parser generates one ID when the call starts, emits it with the first structured name/arguments delta, and the new test proves exactly one non-empty ID. Focused Mistral tests pass (11 passed).

The branch now also contains two unrelated SimpleEngine commits:

  • d408d70 — natural-stop finish_reason="stop", which belongs to #629;
  • edc07d4 — system-KV-cache closeout synthesis, which is a separate engine behavior change.

Those commits add vllm_mlx/engine/simple.py to this parser PR and broaden both its behavior and review surface. Please rebase/drop those two commits so #631 contains only the original Mistral [ARGS] parser change plus the ID follow-up (d8d15b8). After that, rerun the focused parser tests and CI; the parser-specific requested change is otherwise satisfied.

…vstral Small 2)

Ministral 3 14B and Devstral Small 2 (Dec 2025 tokenizers) emit tool calls as
[TOOL_CALLS]<name>[ARGS]<json arguments> — confirmed directly in their
chat_template.jinja: '[TOOL_CALLS]' + name + '[ARGS]' + arguments. The
mistral parser only knew about the old bracket-JSON-array format and a
name{args} format with no separator, so it mishandled this one in two ways:

- Non-streaming (extract_tool_calls): split on the first "{", so the
  function name came back as "get_weather[ARGS]" instead of "get_weather".
- Streaming (_parse_streaming_tool_delta): re-classified every incoming
  delta independently using a "does this chunk start with JSON punctuation"
  heuristic, with no memory of having already passed the name/arguments
  boundary. Bare-word JSON string fragments (e.g. "city", "Paris") have no
  distinguishing leading punctuation, so they got misclassified as more of
  the function name once already inside the arguments blob — reconstructing
  garbage from any standard OpenAI-style delta accumulator.

Fix: recognize the [ARGS] marker explicitly in both code paths, and give
the streaming parser persistent per-tool-call state (_args_started,
_name_buffer) so the name/arguments boundary is only decided once — buffered
until the marker is found (also handles the marker being split across two
deltas), then every subsequent delta is unconditionally arguments.

Verified against a live vllm-mlx server (--tool-call-parser mistral) with
mlx-community/Ministral-3-14B-Instruct-2512-4bit, both streaming and
non-streaming. Added two regression tests replaying the exact delta
sequences captured from that server, plus a non-streaming format test.
Full existing suite (113 tests) still passes.
…KEN delta

The [ARGS]-marker buffering defers the name/arguments boundary past the
delta that contains [TOOL_CALLS], so the id generated in that branch was
frequently never attached to any delta: the middle-of-call branch that
ends up carrying the actual name/arguments never emitted an id at all.
Clients that correlate streamed tool-call deltas by id would see a
call with no id.

Now the id is generated once when a tool call starts and attached to
whichever delta is the first to carry real content, in either branch,
exactly once.

Addresses review feedback from PR waybarrios#631:
waybarrios#631 (review)
@mabaeyens
mabaeyens force-pushed the fix/mistral-args-token-tool-parsing branch from d8d15b8 to dfe6f46 Compare July 10, 2026 18:53
@mabaeyens

Copy link
Copy Markdown
Contributor Author

Rebased: dropped d408d70 (finish_reason=stop, belongs to #629) and edc07d4 (system-KV-cache closeout) — branch now contains only the [ARGS] parser fix (2ecb0f0, rebased from aa14cbb) plus the tool-call ID follow-up (dfe6f46, rebased from d8d15b8). Focused Mistral tests: 11 passed.

@mabaeyens
mabaeyens requested a review from Thump604 July 10, 2026 18:55

@Thump604 Thump604 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed the current dfe6f46 head after the requested cleanup.

Both prior blockers are resolved: the streamed tool-call ID is emitted exactly once with the first structured delta, and the unrelated SimpleEngine commits are no longer in this branch. The current diff is limited to the Mistral [ARGS] parser behavior and focused tests.

I also reran the focused suite locally: 11 passed. The newly authorized current-head CI run is green. This is ready from my review.

@waybarrios

Copy link
Copy Markdown
Owner

Thanks for this PR, the [ARGS] support is genuinely well structured and the id follow-up is solid. We ran a deeper review pass on the current head (dfe6f46) with several independent reviewers and reproduced some edge cases against the parser directly. The happy paths all behave as described in the body, and the streaming reconstruction for the [ARGS] format is correct. That said, we found five issues we would love to see addressed, two of them regressions against the old behavior.

1. Silent response loss when the boundary marker never arrives (regression, confirmed)

Source: _parse_streaming_tool_delta in vllm_mlx/tool_parsers/mistral_tool_parser.py:295-318, interacting with the suppression in server.py:6264-6266 (OpenAI streaming) and the end-of-stream fallback at server.py:6330-6369 which only re-emits parsed tool calls.

Once [TOOL_CALLS] is seen, every delta is withheld until [ARGS] or { shows up in the buffer. If that marker never arrives (max_tokens truncation in the middle of the name, or a model that emits [TOOL_CALLS] and then plain prose), every delta returns None and the server suppresses it. The client ends up with an empty completion, and the buffered text is discarded.

Reproduction, running the actual parser from the PR head:

deltas = ["[TOOL_CALLS]get", "_weather", " then", " prose", " continues"]
# every delta returns None
# _name_buffer final: '[TOOL_CALLS]get_weather then prose continues'  -> never emitted

The buffer also grows without a bound (one find over the whole buffer per delta, so quadratic overall, on the event loop). After 100 prose deltas it already holds 412 characters with nothing emitted.

The old code classified those fragments as a name and streamed them, so nothing was lost. Suggested direction: bound the name-phase buffer (emit the accumulated text as content once it exceeds a window, similar to what qwen_tool_parser.py does with marker-suffix-only buffering) and flush it when the stream ends without a marker.

2. Marker ordering regression for the legacy format (regression, confirmed)

Source: the [ARGS] branch is checked before the { split in extract_tool_calls (mistral_tool_parser.py:113-114), and the streaming path scans markers in fixed order (ARGS_TOKEN, "{") (line 301-302) instead of taking the earliest position.

A legacy-format call whose JSON arguments contain the literal string [ARGS] is now parsed at the wrong boundary. The base implementation handled this correctly.

Actual output from the PR head:

p.extract_tool_calls('[TOOL_CALLS]get_weather{"k": "[ARGS]"}')
# -> name: 'get_weather{"k": "', arguments: '"}'   (base gave name 'get_weather', arguments '{"k": "[ARGS]"}')

Suggested direction: pick the boundary by earliest position, for example min of find("[ARGS]") and find("{"), in both paths, and add regression tests for arguments containing the literal marker.

3. Tool-call smuggling through markers inside JSON strings (confirmed)

Source: split("[TOOL_CALLS]") at mistral_tool_parser.py:101 splits on every occurrence, including inside quoted JSON string values, and the new [ARGS] branch (lines 112-125) passes the arguments through without validating they parse as JSON.

If model output contains a marker sequence inside an argument value, the parser emits a second, clean, dispatchable call that the model never generated:

p.extract_tool_calls('[TOOL_CALLS]get_weather[ARGS]{"city":"[TOOL_CALLS]rm[ARGS]{"f":1}')
# -> call 1: name 'get_weather', arguments '{"city":"'
# -> call 2: name 'rm', arguments '{"f":1}'      <- forged, valid JSON, dispatchable

Before this PR the forged name would have come out as rm[ARGS], which matches no registered function and fails safe. Suggested direction: only split on markers that are outside quoted strings, validate that the arguments portion parses as JSON, and reject calls whose name does not match a plain identifier pattern.

4. The legacy streaming path has no reconstruction coverage

Source: the three new streaming tests in tests/test_tool_parsers.py only exercise the [ARGS] boundary. The rewritten _parse_streaming_tool_delta also owns the legacy { streaming path, and mutating it (for example removing "{" from the marker scan) passes all current tests while legacy streams would silently never emit name or arguments.

Suggested direction: add a streaming delta replay for the legacy format asserting the reconstructed name and arguments.

5. The id regression test passes against the old implementation

Source: tests/test_tool_parsers.py:855-890 (test_mistral_streaming_args_token_has_stable_id). Traced against the base: the old per-delta heuristic emitted the id on the BOT_TOKEN delta, which also produces exactly one id, so the test passes on the unfixed code. It only fails on intermediate states, and it never pins which delta carries the id or the id format.

Suggested direction: give [TOOL_CALLS] its own first delta in the replay so an id tied to the BOT_TOKEN delta emits nothing and fails, and/or assert that the id-bearing delta is the first one with real content.

One thing we checked and is fine

The new per-call state (name buffer, args latch, tool-call id) is safe on current main: the streaming paths build a fresh parser instance per request via _build_tool_parser, so concurrent streams do not share state. Worth pinning with a test at some point, but not a blocker.

The core fix is correct and the tests you added are well specified for the happy path. Happy to help with any of the five items above, and please let us know if any of the reproductions look different on your side.

@waybarrios

Copy link
Copy Markdown
Owner

We went ahead and implemented the fixes for the five findings from the review comment above, on top of the current head. Two new commits landed on the branch:

Commit 1 — Fix Mistral parser marker boundary, JSON-aware splitting and bounded name buffering

  • The [ARGS] marker is now the boundary only when it precedes the first { (both non-streaming and streaming), so legacy calls whose JSON arguments contain the literal "[ARGS]" substring keep the { boundary.
  • The non-streaming split on [TOOL_CALLS] is now JSON-aware: markers inside quoted string values are argument data and never split into a forged second call. The [ARGS] branch also rejects calls whose arguments do not parse as JSON, and names that are not plain identifiers.
  • Streaming arguments are latched after the boundary and never re-scanned for new [TOOL_CALLS] markers.
  • The withheld name buffer is bounded (256 chars); if the boundary never arrives, the withheld text is flushed as content instead of being silently lost.

Commit 2 — Cover the legacy streaming path and pin the tool-call id to the first content delta

  • New delta-by-delta replay test for the legacy { streaming boundary.
  • The stable-id test now splits [TOOL_CALLS] into its own first delta and asserts the id rides the delta carrying the complete name; it fails against the old implementation.

Verification: the full tests/test_tool_parsers.py suite passes locally (120 tests, including all pre-existing Mistral tests), ruff and black are clean, and the previous reproduction cases (legacy [ARGS]-in-JSON, forged call, marker-less truncation) now behave as described in the fixes. CI will re-run on the updated head.

@Thump604 @janhilgard could you take another look when you have a moment? Happy to adjust anything that does not match your expectations.

@mabaeyens

mabaeyens commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks for picking these up, and for pushing the fixes instead of handing them back. I re-ran your five reproductions against a9daebf and they behave the way you describe: the legacy [ARGS]-in-JSON case keeps the { boundary, the forged call is rejected on both the streaming and non-streaming paths, and the bounded name buffer flushes the withheld text instead of losing it.

Two things came up while I was checking, both introduced by the two new commits rather than pre-existing, and the first of them regresses against main as well. Putting the runs in front of you before this merges.

1. Parallel calls collapse into a single call on the streaming path

Delta replay of two consecutive [ARGS] calls, accumulating per index the way a client would:

deltas: [TOOL_CALLS] / get_weather / [ARGS] / {"city":"Madrid"}
        [TOOL_CALLS] / get_time    / [ARGS] / {"tz":"CET"}

a9daebf   indices emitted: [0]
          index 0 arguments: '{"city":"Madrid"}[TOOL_CALLS]get_time[ARGS]{"tz":"CET"}'
                             json.loads -> Extra data

dfe6f46   indices emitted: [0, 1]
          index 0 arguments: '{"city":"Madrid"}'   index 1 arguments: '{"tz":"CET"}'

It is an ordering thing: the if self._args_started early return at mistral_tool_parser.py:293 runs ahead of the new-call detection at line 310, so a second [TOOL_CALLS] can no longer open index 1. The non-streaming path takes the same input and returns both calls correctly, so this is streaming only.

The part that made me want to flag it before merge rather than after: it also catches the legacy brace format, which main already streams correctly. Same replay with get_weather{"city":"Madrid"} and get_time{"tz":"CET"}:

main (0dd1157)   indices [0, 1], both arguments valid JSON
dfe6f46          indices [0, 1], both arguments valid JSON
a9daebf          indices [0],    arguments '{"city":"Madrid"}[TOOL_CALLS]get_time{"tz":"CET"}'

So for the legacy format this is a regression against main, not only against the earlier state of this branch. The [ARGS] format is a different story: main never parsed it, so there is no regression there, only the improvement in this PR not landing in full.

On the note that the end-of-stream re-parse recovers calls arriving after the boundary: I do not think it can reach this case. That fallback is guarded by not tool_calls_detected, and tool_calls_detected is set as soon as any delta carries tool_calls (server.py:6232 in my checkout, fallback at 6294-6300). Index 0 has already gone out by then, so the fallback is skipped and the client is left holding the malformed arguments string.

The one-line version of the fix does not work, and your own suite says so. Letting a delta that contains the marker fall through:

if self._args_started and self.BOT_TOKEN not in delta_text:

restores [0, 1] for the parallel case, but the smuggling replay then also emits [0, 1], and TestStreamingParsing::test_mistral_streaming_marker_in_arguments_does_not_reset fails on it (119 passed, 1 failed):

assert '{"city": "' == '{"city": "[TOOL_CALLS]rm"}'

which is the coupling in one line: the early return that swallows a second call is the same thing that keeps a marker inside a quoted value as data. Telling those two apart needs quote state carried across argument deltas, the streaming counterpart of what _split_on_tool_call_markers now does for the whole string. Happy to write that as a commit on this branch if it is useful, or leave it to you.

Either way, parallel [ARGS] calls have no streaming coverage at the moment, which is why this passes CI.

2. An odd number of double quotes before the marker hides the tool call

_split_on_tool_call_markers tracks string state from index 0, but the text ahead of the first [TOOL_CALLS] is prose, not JSON. A single unbalanced quote there leaves in_string true when the marker arrives, so the marker is never a split point:

'The column is called "id.[TOOL_CALLS]get_schema[ARGS]{"table":"users"}'

a9daebf   tools_called=False, the whole string comes back as content
dfe6f46   tools_called=True,  get_schema {"table":"users"}

Same outcome with the legacy [{...}] format. With the quote balanced, or with no content at all before the marker, the two revisions agree.

Starting the scan at the first marker rather than at index 0 fixes it and keeps the forged call rejected:

token = self.BOT_TOKEN
# The text before the first marker is prose, not JSON.
first = text.find(token)
if first == -1:
    return [text]
parts: list[str] = [text[:first]]
start = first + len(token)
in_string = False
escaped = False
i = start

with the loop below unchanged. With that applied: the odd-quote case parses again, the non-streaming smuggling case is still rejected, the streaming smuggling replay still emits [0] only, and tests/test_tool_parsers.py is 120 of 120.

All of the above on a9daebf, Python 3.13.12, pytest 9.1.1, macOS 26.6 arm64. The unpatched head is 120 of 120 here too, matching what you saw.

TimotejLabsky added a commit to TimotejLabsky/vllm-mlx that referenced this pull request Aug 3, 2026
…k of upstream waybarrios#631

Dec-2025 Mistral tokenizers (Devstral Small 2, Ministral 3) emit
[TOOL_CALLS]name[ARGS]{json}; the parser read the name as
"name[ARGS]" and shredded streaming deltas, so Devstral Small 2 tool
calling was fully broken. Older formats untouched. PATCHES.md waybarrios#42.

Cherry-pick of waybarrios#631
(mabaeyens); retire on the next rebase past its merge.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@waybarrios

Copy link
Copy Markdown
Owner

Apologies, my previous comment was deleted by mistake. Re-posting the findings:

  1. Two consecutive tool calls in a streamed response collapse into one. After the first call's boundary, extract_tool_calls_streaming returns early on _args_started, so a new [TOOL_CALLS] delta is emitted as index-0 arguments instead of opening index 1. The arguments become malformed JSON and the end-of-stream re-parse cannot recover them. This regresses the legacy { format (main streamed it correctly) and the original version of this PR, which detected [TOOL_CALLS] before the mid-call branch and streamed consecutive calls as separate indices. No streaming test covers two consecutive calls.

  2. _split_on_tool_call_markers tracks quote state from the start of the text, but prose before the first marker is not JSON. An odd number of double quotes there (e.g. The column is called "id.) leaves in_string set, so the marker never splits and the call is silently dropped. Starting the scan at the first marker avoids this.

Also, no PATCHES.md or other .md documentation is needed here; we keep the repo clean of extra doc files.

…r own indices and prose quotes no longer hide a call

Found two regressions from the earlier [ARGS] fix while stress-testing the streaming path.

The streaming parser used to return early once the first call's boundary was seen, so a second [TOOL_CALLS] never got its own index: two calls collapsed into one malformed blob of arguments. Now the parser tracks JSON quote state across argument deltas. A [TOOL_CALLS] outside a string opens the next index; one inside a quoted value (e.g. {"city": "[TOOL_CALLS]rm"}) stays argument data, so the smuggling guard still holds.

The non-streaming split also started counting quotes at the very start of the text, but everything before the first marker is prose, not JSON. An odd number of double quotes there left in_string set and the call was silently dropped. The scan now starts at the first marker.

Added five regression tests covering both formats, the legacy brace format, and a call starting mid-delta. They all fail on the previous head and pass now. Tool-parser suite is green (125 tests), ruff and black are clean.

@waybarrios waybarrios left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made small changes to deal with some findings. Now it is ready to work

@waybarrios
waybarrios merged commit 57e91a9 into waybarrios:main Aug 3, 2026
9 checks passed
TimotejLabsky added a commit to TimotejLabsky/vllm-mlx that referenced this pull request Aug 10, 2026
…ios#42/waybarrios#43 retired

Upstream's two new commits are our own two cherry-picks merging: waybarrios#631
(mistral [ARGS] parser, 57e91a9) and waybarrios#562 (gpt-oss harmony tool calls,
b998776). Both are now in the base, so patches waybarrios#42 and waybarrios#43 retire.

Neither auto-dropped — both PRs gained review hardening after we took
them at head 98d4f83, so the merged versions are strict supersets.
Dropped explicitly and upstream's taken; both tool_parsers files are now
byte-identical to upstream/main. Net gain includes a real fix our
cherry-pick lacked: waybarrios#631's JSON-string-aware splitting stops a
[TOOL_CALLS] marker inside a quoted argument value from forging a second
dispatchable call.

One hand-merge, same region as the original: upstream's waybarrios#562 hands the
tool parser _strip_harmony_analysis_blocks(output_text) rather than raw
output_text; patch #27's fold-not-drop block re-applied after it. Patch
waybarrios#47's _explicit_reasoning_markers_present collided only on placement —
both helpers kept.

Retirement audit of the remaining cherry-picks (upstream state checked
live): waybarrios#41/waybarrios#626, waybarrios#44/waybarrios#552, waybarrios#45/waybarrios#551, waybarrios#49/waybarrios#634 all still OPEN upstream —
keep. waybarrios#46/waybarrios#497 is now CLOSED without merging, so it is permanently ours
rather than pending-retirement; status and tracking entry corrected.

Suite green: 2586 passed / 29 skipped / 26 deselected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TimotejLabsky added a commit to TimotejLabsky/vllm-mlx that referenced this pull request Aug 18, 2026
…ios#42/waybarrios#43 retired

Upstream's two new commits are our own two cherry-picks merging: waybarrios#631
(mistral [ARGS] parser, 57e91a9) and waybarrios#562 (gpt-oss harmony tool calls,
b998776). Both are now in the base, so patches waybarrios#42 and waybarrios#43 retire.

Neither auto-dropped — both PRs gained review hardening after we took
them at head 98d4f83, so the merged versions are strict supersets.
Dropped explicitly and upstream's taken; both tool_parsers files are now
byte-identical to upstream/main. Net gain includes a real fix our
cherry-pick lacked: waybarrios#631's JSON-string-aware splitting stops a
[TOOL_CALLS] marker inside a quoted argument value from forging a second
dispatchable call.

One hand-merge, same region as the original: upstream's waybarrios#562 hands the
tool parser _strip_harmony_analysis_blocks(output_text) rather than raw
output_text; patch #27's fold-not-drop block re-applied after it. Patch
waybarrios#47's _explicit_reasoning_markers_present collided only on placement —
both helpers kept.

Retirement audit of the remaining cherry-picks (upstream state checked
live): waybarrios#41/waybarrios#626, waybarrios#44/waybarrios#552, waybarrios#45/waybarrios#551, waybarrios#49/waybarrios#634 all still OPEN upstream —
keep. waybarrios#46/waybarrios#497 is now CLOSED without merging, so it is permanently ours
rather than pending-retirement; status and tracking entry corrected.

Suite green: 2586 passed / 29 skipped / 26 deselected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TimotejLabsky added a commit to TimotejLabsky/vllm-mlx that referenced this pull request Aug 23, 2026
…ios#42/waybarrios#43 retired

Upstream's two new commits are our own two cherry-picks merging: waybarrios#631
(mistral [ARGS] parser, 57e91a9) and waybarrios#562 (gpt-oss harmony tool calls,
b998776). Both are now in the base, so patches waybarrios#42 and waybarrios#43 retire.

Neither auto-dropped — both PRs gained review hardening after we took
them at head 98d4f83, so the merged versions are strict supersets.
Dropped explicitly and upstream's taken; both tool_parsers files are now
byte-identical to upstream/main. Net gain includes a real fix our
cherry-pick lacked: waybarrios#631's JSON-string-aware splitting stops a
[TOOL_CALLS] marker inside a quoted argument value from forging a second
dispatchable call.

One hand-merge, same region as the original: upstream's waybarrios#562 hands the
tool parser _strip_harmony_analysis_blocks(output_text) rather than raw
output_text; patch #27's fold-not-drop block re-applied after it. Patch
waybarrios#47's _explicit_reasoning_markers_present collided only on placement —
both helpers kept.

Retirement audit of the remaining cherry-picks (upstream state checked
live): waybarrios#41/waybarrios#626, waybarrios#44/waybarrios#552, waybarrios#45/waybarrios#551, waybarrios#49/waybarrios#634 all still OPEN upstream —
keep. waybarrios#46/waybarrios#497 is now CLOSED without merging, so it is permanently ours
rather than pending-retirement; status and tracking entry corrected.

Suite green: 2586 passed / 29 skipped / 26 deselected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TimotejLabsky added a commit to TimotejLabsky/vllm-mlx that referenced this pull request Aug 27, 2026
…ios#42/waybarrios#43 retired

Upstream's two new commits are our own two cherry-picks merging: waybarrios#631
(mistral [ARGS] parser, 57e91a9) and waybarrios#562 (gpt-oss harmony tool calls,
b998776). Both are now in the base, so patches waybarrios#42 and waybarrios#43 retire.

Neither auto-dropped — both PRs gained review hardening after we took
them at head 98d4f83, so the merged versions are strict supersets.
Dropped explicitly and upstream's taken; both tool_parsers files are now
byte-identical to upstream/main. Net gain includes a real fix our
cherry-pick lacked: waybarrios#631's JSON-string-aware splitting stops a
[TOOL_CALLS] marker inside a quoted argument value from forging a second
dispatchable call.

One hand-merge, same region as the original: upstream's waybarrios#562 hands the
tool parser _strip_harmony_analysis_blocks(output_text) rather than raw
output_text; patch #27's fold-not-drop block re-applied after it. Patch
waybarrios#47's _explicit_reasoning_markers_present collided only on placement —
both helpers kept.

Retirement audit of the remaining cherry-picks (upstream state checked
live): waybarrios#41/waybarrios#626, waybarrios#44/waybarrios#552, waybarrios#45/waybarrios#551, waybarrios#49/waybarrios#634 all still OPEN upstream —
keep. waybarrios#46/waybarrios#497 is now CLOSED without merging, so it is permanently ours
rather than pending-retirement; status and tracking entry corrected.

Suite green: 2586 passed / 29 skipped / 26 deselected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TimotejLabsky added a commit to TimotejLabsky/vllm-mlx that referenced this pull request Sep 23, 2026
…ios#42/waybarrios#43 retired

Upstream's two new commits are our own two cherry-picks merging: waybarrios#631
(mistral [ARGS] parser, 57e91a9) and waybarrios#562 (gpt-oss harmony tool calls,
b998776). Both are now in the base, so patches waybarrios#42 and waybarrios#43 retire.

Neither auto-dropped — both PRs gained review hardening after we took
them at head 98d4f83, so the merged versions are strict supersets.
Dropped explicitly and upstream's taken; both tool_parsers files are now
byte-identical to upstream/main. Net gain includes a real fix our
cherry-pick lacked: waybarrios#631's JSON-string-aware splitting stops a
[TOOL_CALLS] marker inside a quoted argument value from forging a second
dispatchable call.

One hand-merge, same region as the original: upstream's waybarrios#562 hands the
tool parser _strip_harmony_analysis_blocks(output_text) rather than raw
output_text; patch #27's fold-not-drop block re-applied after it. Patch
waybarrios#47's _explicit_reasoning_markers_present collided only on placement —
both helpers kept.

Retirement audit of the remaining cherry-picks (upstream state checked
live): waybarrios#41/waybarrios#626, waybarrios#44/waybarrios#552, waybarrios#45/waybarrios#551, waybarrios#49/waybarrios#634 all still OPEN upstream —
keep. waybarrios#46/waybarrios#497 is now CLOSED without merging, so it is permanently ours
rather than pending-retirement; status and tracking entry corrected.

Suite green: 2586 passed / 29 skipped / 26 deselected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
TimotejLabsky added a commit to TimotejLabsky/vllm-mlx that referenced this pull request Sep 25, 2026
…ios#42/waybarrios#43 retired

Upstream's two new commits are our own two cherry-picks merging: waybarrios#631
(mistral [ARGS] parser, 57e91a9) and waybarrios#562 (gpt-oss harmony tool calls,
b998776). Both are now in the base, so patches waybarrios#42 and waybarrios#43 retire.

Neither auto-dropped — both PRs gained review hardening after we took
them at head 98d4f83, so the merged versions are strict supersets.
Dropped explicitly and upstream's taken; both tool_parsers files are now
byte-identical to upstream/main. Net gain includes a real fix our
cherry-pick lacked: waybarrios#631's JSON-string-aware splitting stops a
[TOOL_CALLS] marker inside a quoted argument value from forging a second
dispatchable call.

One hand-merge, same region as the original: upstream's waybarrios#562 hands the
tool parser _strip_harmony_analysis_blocks(output_text) rather than raw
output_text; patch #27's fold-not-drop block re-applied after it. Patch
waybarrios#47's _explicit_reasoning_markers_present collided only on placement —
both helpers kept.

Retirement audit of the remaining cherry-picks (upstream state checked
live): waybarrios#41/waybarrios#626, waybarrios#44/waybarrios#552, waybarrios#45/waybarrios#551, waybarrios#49/waybarrios#634 all still OPEN upstream —
keep. waybarrios#46/waybarrios#497 is now CLOSED without merging, so it is permanently ours
rather than pending-retirement; status and tracking entry corrected.

Suite green: 2586 passed / 29 skipped / 26 deselected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants