Skip to content
Merged
Show file tree
Hide file tree
Changes from 17 commits
Commits
Show all changes
46 commits
Select commit Hold shift + click to select a range
5363916
update: init m2
Oct 31, 2025
261fe5c
update: docs and config
Oct 31, 2025
ac4613c
update: init minimax-m2 test
Oct 31, 2025
3421fe7
update: fix tests
Oct 31, 2025
3a5df7a
update: use partial_rotary_factor
Nov 4, 2025
cb17f62
update: some fix
Nov 5, 2025
f6775d8
fix: import Unpack from processing_utils
Nov 5, 2025
73904ee
update: apply suggestions from code review
rogeryoungh Nov 6, 2025
6b7e397
update: remove MiniMaxM2DecoderLayer and MiniMaxM2MLP
Nov 6, 2025
1565721
update: remove use_qk_norm
Nov 13, 2025
fa09301
update: remove unused use_qk_norm
Nov 13, 2025
11fdb58
update: update config and attention
Nov 24, 2025
93b7598
update: add to tokenization_auto and remove unused test
Nov 24, 2025
5de5179
Merge branch 'main' into minimax-m2
rogeryoungh Nov 24, 2025
f5219ca
update: fix decoder layer and experts
Nov 24, 2025
ef9a1f9
update: fix docs
Nov 24, 2025
a0eea92
update: make ci happy
Nov 24, 2025
2dbfc3b
refactor: use mapping
Nov 25, 2025
640fb9f
update: remove unused comments
Nov 25, 2025
8fe4d2b
Merge branch 'main' into minimax-m2
rogeryoungh Nov 26, 2025
17c68c8
Merge branch 'main' into minimax-m2
rogeryoungh Dec 5, 2025
8ba23a6
update: fix rope_params and router
Dec 5, 2025
8a76b78
update: remove rope_theta
Dec 8, 2025
bcc0aa1
update: test_load_balancing_loss
Dec 8, 2025
5a82172
Merge branch 'main' into minimax-m2
rogeryoungh Dec 8, 2025
47c84e2
update: docs
Dec 8, 2025
fdb6807
update: fix default theta
Dec 9, 2025
1f07a00
update to proper default values, proper config rope, simplified modular
vasqu Dec 10, 2025
25a8b5c
fix docs
vasqu Dec 10, 2025
7ea9a2b
modular fixup
vasqu Dec 10, 2025
13ed0be
Merge branch 'main' into minimax-m2
vasqu Dec 10, 2025
9eebced
review comments
vasqu Dec 11, 2025
668dbef
Merge branch 'main' into minimax-m2
vasqu Dec 11, 2025
90acf7c
update slow tests
vasqu Dec 11, 2025
e4463e1
style
vasqu Dec 11, 2025
fde9591
fp32 strict
vasqu Dec 11, 2025
7b6824f
revert the flag
vasqu Dec 11, 2025
dda58d9
Merge branch 'main' into minimax-m2
vasqu Dec 16, 2025
04c0b5f
Merge branch 'main' into minimax-m2
vasqu Jan 7, 2026
8703235
sync with latest changes
vasqu Jan 7, 2026
53c19ac
fixup buffer init
vasqu Jan 7, 2026
28f26fd
add cache exception to minimax m2 as we have a naming clash
vasqu Jan 7, 2026
cca9ee7
fix dtype issue in gate
vasqu Jan 7, 2026
3d6d671
Merge branch 'main' into minimax-m2
vasqu Jan 8, 2026
b867712
lift fp8 test restriction and apply new linter rules
vasqu Jan 8, 2026
252d02d
update docs
vasqu Jan 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions docs/source/en/_toctree.yml
Original file line number Diff line number Diff line change
Expand Up @@ -592,6 +592,8 @@
title: MegatronGPT2
- local: model_doc/minimax
title: MiniMax
- local: model_doc/minimax_m2
title: MiniMax-M2
- local: model_doc/ministral
title: Ministral
- local: model_doc/mistral
Expand Down
2 changes: 2 additions & 0 deletions docs/source/en/model_doc/minimax.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,8 @@ rendered properly in your Markdown viewer.

# MiniMax

> [MiniMax-M2](https://huggingface.co/docs/transformers/en/model_doc/minimax_m2) was released on 2025‑10‑27. We recommend using MiniMax‑M2 for most use cases due to better overall performance.

## Overview

The MiniMax-Text-01 model was proposed in [MiniMax-01: Scaling Foundation Models with Lightning Attention](https://huggingface.co/papers/2501.08313) by MiniMax, Aonian Li, Bangwei Gong, Bo Yang, Boji Shan, Chang Liu, Cheng Zhu, Chunhao Zhang, Congchao Guo, Da Chen, Dong Li, Enwei Jiao, Gengxin Li, Guojun Zhang, Haohai Sun, Houze Dong, Jiadai Zhu, Jiaqi Zhuang, Jiayuan Song, Jin Zhu, Jingtao Han, Jingyang Li, Junbin Xie, Junhao Xu, Junjie Yan, Kaishun Zhang, Kecheng Xiao, Kexi Kang, Le Han, Leyang Wang, Lianfei Yu, Liheng Feng, Lin Zheng, Linbo Chai, Long Xing, Meizhi Ju, Mingyuan Chi, Mozhi Zhang, Peikai Huang, Pengcheng Niu, Pengfei Li, Pengyu Zhao, Qi Yang, Qidi Xu, Qiexiang Wang, Qin Wang, Qiuhui Li, Ruitao Leng, Shengmin Shi, Shuqi Yu, Sichen Li, Songquan Zhu, Tao Huang, Tianrun Liang, Weigao Sun, Weixuan Sun, Weiyu Cheng, Wenkai Li, Xiangjun Song, Xiao Su, Xiaodong Han, Xinjie Zhang, Xinzhu Hou, Xu Min, Xun Zou, Xuyang Shen, Yan Gong, Yingjie Zhu, Yipeng Zhou, Yiran Zhong, Yongyi Hu, Yuanxiang Fan, Yue Yu, Yufeng Yang, Yuhao Li, Yunan Huang, Yunji Li, Yunpeng Huang, Yunzhi Xu, Yuxin Mao, Zehan Li, Zekang Li, Zewei Tao, Zewen Ying, Zhaoyang Cong, Zhen Qin, Zhenhua Fan, Zhihang Yu, Zhuo Jiang, Zijia Wu.
Expand Down
82 changes: 82 additions & 0 deletions docs/source/en/model_doc/minimax_m2.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
<!--Copyright 2025 the HuggingFace Team. All rights reserved.

Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.


⚠️ Note that this file is in Markdown but contain specific syntax for our doc-builder (similar to MDX) that may not be rendered properly in your Markdown viewer.

-->


# MiniMax-M2

## Overview

MiniMax-M2 is a compact, fast, and cost-effective MoE model (230 billion total parameters with 10 billion active parameters) built for elite performance in coding and agentic tasks, all while maintaining powerful general intelligence. With just 10 billion activated parameters, MiniMax-M2 provides the sophisticated, end-to-end tool use performance expected from today's leading models, but in a streamlined form factor that makes deployment and scaling easier than ever.

For more details refer to the [release blog post](https://www.minimax.io/news/minimax-m2).

## Usage examples

```python
from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig

model = AutoModelForCausalLM.from_pretrained("MiniMaxAI/MiniMax-M2", device_map="auto")

tokenizer = AutoTokenizer.from_pretrained("MiniMaxAI/MiniMax-M2")

generation_config = GenerationConfig.from_pretrained("MiniMaxAI/MiniMax-M2")

messages = [
{"role": "user", "content": "What is your favourite condiment?"},
{"role": "assistant", "content": "Well, I'm quite partial to a good squeeze of fresh lemon juice. It adds just the right amount of zesty flavour to whatever I'm cooking up in the kitchen!"},
{"role": "user", "content": "Do you have mayonnaise recipes?"}
]

model_inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to("cuda")

generated_ids = model.generate(model_inputs, max_new_tokens=100, generation_config=generation_config)

response = tokenizer.batch_decode(generated_ids)[0]

print(response)
```

## MiniMaxM2Config

[[autodoc]] MiniMaxM2Config

## MiniMaxM2ForCausalLM

[[autodoc]] MiniMaxM2ForCausalLM
- forward

## MiniMaxM2ForQuestionAnswering

[[autodoc]] MiniMaxM2ForQuestionAnswering
- forward

## MiniMaxM2Model

[[autodoc]] MiniMaxM2Model
- forward

## MiniMaxM2ForSequenceClassification

[[autodoc]] MiniMaxM2ForSequenceClassification
- forward

## MiniMaxM2ForTokenClassification

[[autodoc]] MiniMaxM2ForTokenClassification
- forward
1 change: 1 addition & 0 deletions src/transformers/models/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -220,6 +220,7 @@
from .mgp_str import *
from .mimi import *
from .minimax import *
from .minimax_m2 import *
from .ministral import *
from .mistral import *
from .mistral3 import *
Expand Down
2 changes: 2 additions & 0 deletions src/transformers/models/auto/configuration_auto.py
Original file line number Diff line number Diff line change
Expand Up @@ -256,6 +256,7 @@
("mgp-str", "MgpstrConfig"),
("mimi", "MimiConfig"),
("minimax", "MiniMaxConfig"),
("minimax_m2", "MiniMaxM2Config"),
("ministral", "MinistralConfig"),
("mistral", "MistralConfig"),
("mistral3", "Mistral3Config"),
Expand Down Expand Up @@ -698,6 +699,7 @@
("mgp-str", "MGP-STR"),
("mimi", "Mimi"),
("minimax", "MiniMax"),
("minimax_m2", "MiniMax-M2"),
("ministral", "Ministral"),
("mistral", "Mistral"),
("mistral3", "Mistral3"),
Expand Down
5 changes: 5 additions & 0 deletions src/transformers/models/auto/modeling_auto.py
Comment thread
rogeryoungh marked this conversation as resolved.
Original file line number Diff line number Diff line change
Expand Up @@ -256,6 +256,7 @@ class _BaseModelWithGenerate(PreTrainedModel, GenerationMixin):
("mgp-str", "MgpstrForSceneTextRecognition"),
("mimi", "MimiModel"),
("minimax", "MiniMaxModel"),
("minimax_m2", "MiniMaxM2Model"),
("ministral", "MinistralModel"),
("mistral", "MistralModel"),
("mistral3", "Mistral3Model"),
Expand Down Expand Up @@ -692,6 +693,7 @@ class _BaseModelWithGenerate(PreTrainedModel, GenerationMixin):
("mbart", "MBartForCausalLM"),
("megatron-bert", "MegatronBertForCausalLM"),
("minimax", "MiniMaxForCausalLM"),
("minimax_m2", "MiniMaxM2ForCausalLM"),
("ministral", "MinistralForCausalLM"),
("mistral", "MistralForCausalLM"),
("mixtral", "MixtralForCausalLM"),
Expand Down Expand Up @@ -1228,6 +1230,7 @@ class _BaseModelWithGenerate(PreTrainedModel, GenerationMixin):
("mbart", "MBartForSequenceClassification"),
("megatron-bert", "MegatronBertForSequenceClassification"),
("minimax", "MiniMaxForSequenceClassification"),
("minimax_m2", "MiniMaxM2ForSequenceClassification"),
("ministral", "MinistralForSequenceClassification"),
("mistral", "MistralForSequenceClassification"),
("mixtral", "MixtralForSequenceClassification"),
Expand Down Expand Up @@ -1322,6 +1325,7 @@ class _BaseModelWithGenerate(PreTrainedModel, GenerationMixin):
("mbart", "MBartForQuestionAnswering"),
("megatron-bert", "MegatronBertForQuestionAnswering"),
("minimax", "MiniMaxForQuestionAnswering"),
("minimax_m2", "MiniMaxM2ForQuestionAnswering"),
("ministral", "MinistralForQuestionAnswering"),
("mistral", "MistralForQuestionAnswering"),
("mixtral", "MixtralForQuestionAnswering"),
Expand Down Expand Up @@ -1434,6 +1438,7 @@ class _BaseModelWithGenerate(PreTrainedModel, GenerationMixin):
("markuplm", "MarkupLMForTokenClassification"),
("megatron-bert", "MegatronBertForTokenClassification"),
("minimax", "MiniMaxForTokenClassification"),
("minimax_m2", "MiniMaxM2ForTokenClassification"),
("ministral", "MinistralForTokenClassification"),
("mistral", "MistralForTokenClassification"),
("mixtral", "MixtralForTokenClassification"),
Expand Down
7 changes: 7 additions & 0 deletions src/transformers/models/auto/tokenization_auto.py
Original file line number Diff line number Diff line change
Expand Up @@ -439,6 +439,13 @@
"GPT2TokenizerFast" if is_tokenizers_available() else None,
),
),
(
"minimax_m2",
(
"GPT2Tokenizer" if is_sentencepiece_available() else None,
"GPT2TokenizerFast" if is_tokenizers_available() else None,
),
),
(
"ministral",
(
Expand Down
29 changes: 29 additions & 0 deletions src/transformers/models/minimax_m2/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# coding=utf-8
# Copyright 2025 the HuggingFace Team. All rights reserved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

from typing import TYPE_CHECKING

from ...utils import _LazyModule
from ...utils.import_utils import define_import_structure


if TYPE_CHECKING:
from .configuration_minimax_m2 import *
from .modeling_minimax_m2 import *
else:
import sys

_file = globals()["__file__"]
sys.modules[__name__] = _LazyModule(__name__, _file, define_import_structure(_file), module_spec=__spec__)
Loading