MiniMax-M1-80k is a 456B-parameter hybrid-attention reasoning model from MiniMaxAI. According to the model card, its MoE design with lightning attention targets long-context reasoning with an 80K thinking budget.
Publisher
MiniMaxAI
Task
text-generation
Model type
minimax_m1
License
apache-2.0
Library
transformers
Publication status
Approved for indexing
Model overview
MiniMax-M1-80k is published by MiniMaxAI as a text-generation reasoning model. The captured configuration identifies MiniMaxM1ForCausalLM, and Safetensors metadata reports about 456B parameters. According to the model card, it is a hybrid Mixture-of-Experts model with lightning attention, developed from MiniMax-Text-01 with an 80K thinking budget for long-input, extended-reasoning tasks.
Recorded capabilities
Hybrid MoE with lightning attention
According to the model card, the hybrid Mixture-of-Experts design with lightning attention activates 45.9B of 456B parameters per token and uses a fraction of the FLOPs of comparable reasoning models at long generations.
RL-trained reasoning with CISPO
According to the model card, the model is trained with large-scale reinforcement learning using the CISPO algorithm across mathematical reasoning and software-engineering environments.
Reported benchmark comparisons
According to the model card, experiments show the models outperforming strong open-weight models such as DeepSeek-R1 and Qwen3-235B on software engineering, tool use, and long-context tasks.
Function-calling support
According to the model card, the model outputs structured function-call parameters and ships with a dedicated function-call guide.
Use cases in the source record
Long-context reasoning tasks such as competition mathematics, coding, and software-engineering evaluations described in the card.
Agentic tool-use workflows built on the documented function-calling capability and function-call guide.
Reasoning deployments tuned with the recommended temperature 1.0, top_p 0.95, and task-specific system prompts.
Limitations and unknowns
Context length, FLOPs comparisons, and benchmark advantages are publisher-reported claims and were not independently verified.
No VRAM or hardware requirement was extracted from this record.
Provider state is historical snapshot data, not independently refreshed current availability.
MiniMax-M1 1. Model Overview We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combined with a lightning attention mechanism. The model is developed based on our previous MiniMax-Text-01 model , which contains a total of 456 billion parameters with 45.9 billion parameters activated per token. Consistent with MiniMax-Text-01, the M1 model natively supports a context length of 1 million tokens, 8x the context size of DeepSeek R1. Furthermore, the lightning attention mechanism in MiniMax-M1 enables efficient scali…