Skip to content

fix(ci): kimi-k2.6 SMG reasoning parser (kimi_k25 → kimi_thinking) - #1868

Merged
key4ng merged 1 commit into
mainfrom
fix/tau2-kimi-smg
Jul 2, 2026
Merged

key4ng merged 1 commit into
mainfrom
fix/tau2-kimi-smg

Conversation

@key4ng

@key4ng key4ng commented Jul 2, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

In the nightly A/B, kimi-k2.6 SMG scored 0 vs vLLM 100. Not a tool-parse bug — SMG emits well-formed tool calls (correct name + args). The SMG agent systematically over-confirms and never fires the mutating tool before the user-sim stops → DB unchanged → reward 0 (all tasks).

Solution

Root cause confirmed on the B200 by inspecting Kimi-K2.6's native chat_template.jinja: it prefills <think> at the generation prompt (every assistant turn starts inside a <think>…</think> reasoning block). That's an always-in-reasoning model. The matrix ran SMG with --reasoning-parser kimi_k25 (always_in_reasoning: false — expects the model to emit <think> itself, which never happens), so SMG mis-splits reasoning vs content and the agent degrades. vLLM's kimi_k2 parser handles the prefill, hence 100.

Fix: kimi-k2.6 smg_reason kimi_k25 → kimi_thinking (always_in_reasoning: true, same <think>/</think>), in both the tau2 and bfcl matrices.

Test Plan

  • Confirmed the template prefill + delimiter (<think>) on the B200 via apply_chat_template.
  • Validate live: workflow_dispatch only=kimi-k2.6 (after merge) — expect SMG pass^k to recover toward parity with vLLM.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Bug Fixes
    • Improved nightly evaluation setup for one model configuration, helping avoid incorrect parsing and scoring behavior during automated runs.
    • Updated the reasoning/tool parsing behavior used in that workflow leg for more reliable results.

…ing (template prefills <think>)

Signed-off-by: key4ng <rukeyang@gmail.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Note

Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported.

@github-actions github-actions Bot added the ci CI/CD configuration changes label Jul 2, 2026
@coderabbitai

coderabbitai Bot commented Jul 2, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

This PR updates the kimi-k2.6 leg's SMG reasoning parser setting from kimi_k25 to kimi_thinking in two nightly CI workflow files: nightly-bfcl.yml and nightly-tau2.yml.

Changes

SMG reasoning parser configuration update

Layer / File(s) Summary
Update smg_reason for kimi-k2.6 leg
.github/workflows/nightly-bfcl.yml, .github/workflows/nightly-tau2.yml
Changed the smg_reason value for the kimi-k2.6 matrix leg from kimi_k25 to kimi_thinking in both nightly workflow files, altering the reasoning-parser mode used for that leg's scoring/tool-parsing.

Estimated code review effort: 1 (Trivial) | ~3 minutes

Possibly related PRs

  • lightseekorg/smg#1724: Directly modifies the same nightly BFCL workflow's model matrix that this PR further tweaks.
  • lightseekorg/smg#1791: Adjusts the same SMG reasoning-parser configuration for the kimi-k2.6 leg in the nightly BFCL/Tau setup.
  • lightseekorg/smg#1834: Tweaks the same nightly workflow's generated model matrix SMG reasoning-parser configuration.

Suggested reviewers: CatherineSue, XinyueZhang369, slin1237, claude

Poem

A tiny hop, a tiny switch,
kimi_k25 gave way to a better hitch,
Now kimi_thinking leads the race,
Two YAML files in one embrace,
This bunny thumps with quiet glee! 🐇

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the ci change and the parser switch for kimi-k2.6, matching the main update in the PR.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/tau2-kimi-smg

Comment @coderabbitai help to get the list of available commands.

@claude

claude Bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 The PR description doesn't fully follow
PULL_REQUEST_TEMPLATE.md:

  • Missing header: ## Changes

Please update the PR description so reviewers have the context they need.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good. The kimi_thinking parser (always_in_reasoning: true, <think>/</think>) is the correct match for K2.6's template-prefilled reasoning. Both workflow files updated consistently. No issues found.

@key4ng key4ng changed the title fix(tau2,bfcl): kimi-k2.6 SMG reasoning parser (kimi_k25 → kimi_thinking) fix(ci): kimi-k2.6 SMG reasoning parser (kimi_k25 → kimi_thinking) Jul 2, 2026
@key4ng

key4ng commented Jul 2, 2026

Copy link
Copy Markdown
Member Author

Note for merging without all check passed: This change only modified CI, no real code change. Want to merge it before the tau2 nightly bench triggered

@key4ng
key4ng merged commit c871927 into main Jul 2, 2026
29 of 36 checks passed
@key4ng
key4ng deleted the fix/tau2-kimi-smg branch July 2, 2026 06:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant