Skip to content

feat(together-ai): add GLM-5.1, Kimi K2.6, DeepSeek V4 Pro - #2094

Merged
steebchen merged 3 commits into
mainfrom
together-models
Apr 26, 2026
Merged

steebchen merged 3 commits into
mainfrom
together-models

Conversation

@steebchen

@steebchen steebchen commented Apr 26, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Add Together AI provider mappings for GLM-5.1, Kimi K2.6, and DeepSeek V4 Pro (MiniMax M2.7 already had one with matching pricing).
  • Mark gemma-2-27b-it on Together AI as deactivated 2026-04-25.
  • Deactivate remaining llama-3.1-8b-instruct provider mappings (aws-bedrock, nebius, inference.net, cerebras, novita) and the sole llama-4-scout mapping (together-ai) as of 2026-04-25 — together-ai for llama-3.1-8b-instruct already had 2026-03-27.
  • Tweak together-ai entries based on e2e results: reasoningOutput: "omit" for DeepSeek V4 Pro (Together AI doesn't expose reasoning), test: "skip" for Kimi K2.6 (no streaming content returned).

Together AI pricing

Model Input Cached Output
GLM-5.1 $1.40 — $4.40
Kimi K2.6 $1.20 $0.20 $4.50
DeepSeek V4 Pro $2.10 $0.20 $4.40

Test plan

  • pnpm --filter @llmgateway/models build
  • TEST_MODELS="together-ai/glm-5.1,together-ai/deepseek-v4-pro" pnpm test:e2e — all pass
  • Together AI Kimi K2.6 e2e skipped via test: "skip"

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added DeepSeek V4 Pro with streaming, reasoning, and JSON output.
    • Added Moonshot Kimi K2.6 with streaming, reasoning, and vision.
    • Added Zhipu GLM-5.1 with streaming, reasoning, and JSON output.
    • Added additional provider options for GPT-OSS 120B and 20B with large context, streaming, tools, and reasoning.
  • Updates

    • Meta Llama 3.1 and Llama 4 Scout providers marked deactivated (2026-04-25).

Copilot AI review requested due to automatic review settings April 26, 2026 17:40
@coderabbitai

coderabbitai Bot commented Apr 26, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: a7b8f468-829a-4ea2-b324-032e953843c3

📥 Commits

Reviewing files that changed from the base of the PR and between ec080cb and 52eed64.

📒 Files selected for processing (1)
  • packages/models/src/models/openai.ts

Walkthrough

Adds multiple Together AI provider entries across model catalogs (deepseek, moonshot, zai, openai) and marks several Meta provider variants as deactivated by adding deactivatedAt: new Date("2026-04-25") to specific metaModels entries.

Changes

Cohort / File(s) Summary
Together AI provider additions
packages/models/src/models/deepseek.ts, packages/models/src/models/moonshot.ts, packages/models/src/models/zai.ts, packages/models/src/models/openai.ts
Introduces new together-ai provider objects for deepseek-v4-pro, kimi-k2.6, GLM-5.1, and gpt-oss-* models. Each entry specifies token/request pricing, context/output limits, and capability flags (streaming, reasoning, tools, vision, jsonOutput/jsonOutputSchema) per model.
Meta model deactivations
packages/models/src/models/meta.ts
Adds deactivatedAt: new Date("2026-04-25") to several providers for llama-3.1-8b-instruct (aws-bedrock, nebius, inference.net, cerebras, novita) and to llama-4-scout (together-ai).

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main changes: adding Together AI provider mappings for three new models (GLM-5.1, Kimi K2.6, DeepSeek V4 Pro).
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch together-models

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

steebchen and others added 2 commits April 27, 2026 00:40
Also deactivate gemma-2-27b-it on Together AI as of 2026-04-25.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Mark llama-3.1-8b-instruct and llama-4-scout providers as deactivated
2026-04-25 across providers without an existing date.

Adjust together-ai entries: omit reasoning output for DeepSeek V4 Pro
(not exposed) and skip e2e for Kimi K2.6 (no streaming content).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: af7405de4e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

requestPrice: 0,
contextSize: 262144,
maxOutput: 32768,
streaming: true,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Mark Kimi K2.6 Together mapping as non-streaming

This mapping advertises streaming: true, so apps/gateway/src/chat/chat.ts will allow streamed requests for it, but this same provider entry is already tagged test: "skip" due to known streaming failures (no streaming content returned in the commit context). As shipped, clients using stream: true can be routed to a configuration that returns empty streamed responses instead of being rejected early; setting streaming: false (or deactivating the mapping until fixed) would prevent broken runtime behavior.

Useful? React with 👍 / 👎.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds/updates Together AI provider mappings for several newly supported models and deactivates a set of provider mappings that have been retired.

Changes:

  • Add Together AI mappings (with pricing/capabilities) for GLM-5.1, Kimi K2.6, and DeepSeek V4 Pro.
  • Mark several existing provider mappings as deactivated as of 2026-04-25 (Gemma 2 27B IT on Together AI; multiple Llama 3.1 8B Instruct mappings; Together AI’s Llama 4 Scout mapping).
  • Apply provider-specific behavior tweaks (e.g., omit reasoning output for Together AI DeepSeek V4 Pro; skip e2e for Together AI Kimi K2.6).

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
packages/models/src/models/zai.ts Adds Together AI mapping for GLM-5.1 with pricing and capability flags.
packages/models/src/models/moonshot.ts Adds Together AI mapping for Kimi K2.6 (tests skipped).
packages/models/src/models/deepseek.ts Adds Together AI mapping for DeepSeek V4 Pro, including cached pricing and reasoning-output behavior.
packages/models/src/models/google.ts Marks Together AI Gemma 2 27B IT mapping as deactivated.
packages/models/src/models/meta.ts Deactivates remaining Llama 3.1 8B Instruct mappings and Together AI’s Llama 4 Scout mapping.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines 331 to +335
streaming: true,
vision: false,
tools: true,
jsonOutput: false,
deactivatedAt: new Date("2026-04-25"),

Copilot AI Apr 26, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For the llama-4-scout model’s Together AI mapping, the configured modelName in this provider entry appears to point to a Llama 3.1 405B Turbo model rather than a Llama 4 Scout variant. Even though this mapping is now deactivated, this mismatch is confusing and can lead to incorrect routing/analytics if the mapping is ever referenced; please verify the intended Together model slug and either correct modelName (and any related pricing/context fields) or rename/remove the model definition so the id/name and modelName align.

Copilot uses AI. Check for mistakes.
JSON output mode is unreliable on Together AI for both, so disabled.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 52eed64412

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

vision: false,
tools: true,
jsonOutput: false,
deactivatedAt: new Date("2026-04-25"),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Deactivate the correct Together Scout model mapping

This deactivatedAt change is applied to the llama-4-scout provider record whose modelName is meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo, so after 2026-04-25 requests targeting together-ai/meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo will be rejected as deactivated. That disables the 3.1-405B mapping rather than the actual Scout model mapping (meta-llama/llama-4-scout-17b-16e-instruct), so the retirement is being enforced on the wrong entry.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants