Skip to content

common/chat : tighten bare function parsing for Qwen models - #26793

Merged
aldehir merged 1 commit into
ggml-org:masterfrom
aldehir:fix-function-trigger
Aug 11, 2026
Merged

common/chat : tighten bare function parsing for Qwen models#26793
aldehir merged 1 commit into
ggml-org:masterfrom
aldehir:fix-function-trigger

Conversation

@aldehir

@aldehir aldehir commented Aug 9, 2026

Copy link
Copy Markdown
Contributor

Overview

Tighten the grammar trigger to only trigger on complete <function=name> sequences. This path is only there to support the older Qwen3-Coder models, and using <function as the trigger inadvertently starts constraining against valid content generation e.g. #include <functional> when tools are provided.

Fixes #26787

Requirements

@aldehir
aldehir merged commit ba360ef into ggml-org:master Aug 11, 2026
24 of 27 checks passed
gabe-l-hart added a commit to gabe-l-hart/llama.cpp that referenced this pull request Aug 12, 2026
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* origin/master: (383 commits)
  cmake :  introduce semantic versioning  (ggml-org#26839)
  gguf : harden loader against malformed tensor dims and metadata types (ggml-org#25596)
  kleidiai: Add runtime feature detection mechanism for aarch64/kleidiai (ggml-org#26076)
  model : disallow integer dflash sliding_window_pattern (ggml-org#26900)
  sync : ggml
  cmake : add config version support (ggml/1582)
  server : support slot save/restore with media inputs (ggml-org#26640)
  ui: add read_media tool (ggml-org#25877)
  opencl: default FA c8 cluster width to 16 on X1E (ggml-org#26433)
  tests : update speculative params (ggml-org#26925)
  vulkan: add TQ2_0 (ternary) support (ggml-org#25850)
  wavtokenizer-dec : bound posnet/convnext block_count against n_layer_all (ggml-org#26892)
  convert : handle per_layer_config in Gemma4 (transformers 5.15) (ggml-org#26882)
  opencl: use flat mv q5_k when weight exceeds image1d_buffer_t limit (ggml-org#26880)
  chat : fix muse-glimmer detection of tool calls after EOM (ggml-org#26879)
  ci : add missing release check (ggml-org#26923)
  CUDA: only disable CUDA graphs when mul_mat_id actually needs a stream sync (ggml-org#26802)
  cuda : add warp-per-row wkv7 kernel for single-token decode (ggml-org#26111)
  spec : update speculative-simple (ggml-org#26904)
  chat : tighten bare function parsing for Qwen models (ggml-org#26793)
  ...
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Aug 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval bug: grammar : Unexpected empty grammar stack when trigger matches inside a longer word (e.g. "function" in "functional")

3 participants