Skip to content

fix(agent): resolve kimi-k3 context window to 1M in fallback table - #67874

Closed
ericmourant wants to merge 1 commit into
NousResearch:mainfrom
ericmourant:fix/kimi-k3-context-window
Closed

ericmourant wants to merge 1 commit into
NousResearch:mainfrom
ericmourant:fix/kimi-k3-context-window

Conversation

@ericmourant

Copy link
Copy Markdown

Motivation

Kimi K3 (Moonshot AI, 2.8T MoE) ships with a 1M-token context window — models.dev agrees: kimi-for-coding/k3, moonshotai/kimi-k3, and moonshotai-cn/kimi-k3 all report context: 1048576.

However, DEFAULT_CONTEXT_LENGTHS in agent/model_metadata.py had only the K2-generation catch-all "kimi": 262144. Longest-prefix matching resolves kimi-k3 → kimi → 256K, so for configurations where the models.dev provider lookup misses — e.g. provider: kimi-for-coding with base_url: https://api.moonshot.ai/v1 — the stale fallback wins and Hermes budgets/compacts at a quarter of K3's actual working memory.

Change

One line: "kimi-k3": 1048576 ahead of the "kimi" catch-all (longest-first matching makes it win). Fallback-table only; no change to live-metadata resolution paths.

Verification

agent.model_metadata.get_model_context_length('kimi-k3', base_url='https://api.moonshot.ai/v1')
  pre-patch:  262144
  post-patch: 1048576

kimi-k2.6 and other K2-family ids still resolve via the kimi catch-all to 262144, unchanged.

Risk

None identified — additive, more-specific fallback entry; live metadata (models.dev / provider endpoints) still takes precedence where available.

The DEFAULT_CONTEXT_LENGTHS catch-all "kimi": 262144 (written for the K2
generation) was the only Kimi entry, so longest-prefix matching resolved
kimi-k3 to 256K. For configurations whose base_url does not reach a
models.dev provider entry (e.g. provider=kimi-for-coding with
base_url=https://api.moonshot.ai/v1), the fallback won and Hermes
budgeted/compacted at a quarter of K3's actual 1M-token window.

Add "kimi-k3": 1048576 ahead of the catch-all, matching models.dev
(kimi-for-coding/k3, moonshotai/kimi-k3, moonshotai-cn/kimi-k3 all
report context=1048576).

Verified with agent.model_metadata.get_model_context_length:
  pre-patch:  262144
  post-patch: 1048576
@alt-glitch alt-glitch added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint provider/kimi Kimi / Moonshot P3 Low — cosmetic, nice to have needs-decision Awaiting maintainer decision before any implementation labels Jul 20, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Related to #67685: both add a Kimi K3 fallback alias, but they use materially different context values (1,048,576 here vs. 1,000,000 there). #67115 is a separate endpoint-gated repair; maintainer consolidation is needed.

@teknium1

Copy link
Copy Markdown
Collaborator

Resolved on main via PR #68108, which salvaged the materially identical earlier submission #67685 (first-submitted 2026-07-19) with that author's attribution preserved. Note the final landed value is 1,048,576 (1 Mi, matching models.dev / OpenRouter live metadata and the endpoint-scoped override) rather than 1,000,000. Thanks for the contribution — the independent confirmation was useful.

@teknium1 teknium1 closed this Jul 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint needs-decision Awaiting maintainer decision before any implementation P3 Low — cosmetic, nice to have provider/kimi Kimi / Moonshot type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants