MiniMax

MiniMax

Freemium

Intelligence with everyone.

LLM
Open-weight

Published 27 September 2026

Scores

Popularity3/5

A fast-iterating model line (M2 through M3 within 2026) with growing recognition, but still behind the Western developer mindshare of more established open-weight labs like DeepSeek and Qwen.

Learning Curve3/5

Using the hosted API is as simple as any OpenAI-compatible endpoint, while self-hosting via vLLM/SGLang or GGUF quantization requires the same infrastructure familiarity as any other large open-weight model.

Flexibility4/5

Deployable via a hosted API, self-hosted through vLLM/SGLang, or run locally through GGUF quantizations on llama.cpp/Ollama/LM Studio, with native multimodal input across text, image, audio, and video.

Performance5/5

Leads the August 2026 open-weight intelligence-index ranking and scores 69.4% on SWE-Bench Verified for the M2 line, on par with proprietary frontier models on coding tasks specifically.

Portability4/5

Self-hostable via vLLM/SGLang and widely available as GGUF quantizations, though the custom minimax-community license is a real consideration for portability compared to a fully permissive MIT/Apache license.

About MiniMax

MiniMax is a Shanghai-based AI company that builds language, speech, video, and music models. Its M-series language models are released with open weights, and the current flagship is MiniMax-M3, a mixture-of-experts model with about 428 billion total and 23 billion active parameters, a one-million-token context window, and native multimodal input covering images, video, and computer-use agent tasks.

M3's main technical feature is MiniMax Sparse Attention (MSA), which makes long contexts much cheaper to process than the previous generation, with large speedups in both prefill and decoding at full context. MiniMax reports strong agentic and coding benchmark results, and the model is widely used in coding agents and long-document workloads.

The weights are published on Hugging Face under MiniMax's own community licence rather than MIT or Apache 2.0, so commercial users should read its terms. Self-hosting is well supported: official vLLM and SGLang recipes, Hugging Face Transformers compatibility, and community GGUF quantisations for llama.cpp, Ollama, and LM Studio, though the full model needs a multi-GPU server.

The hosted API is OpenAI-compatible and priced per million tokens, with a permanent discount off list price, higher rates for very long prompts, cheap cache reads, and an optional priority tier for faster admission. For coding agents, MiniMax sells the Token Plan, a monthly subscription with five-hour and weekly quotas that works in Claude Code, Codex, Cursor, and other tools. MiniMax also offers separate speech, video (Hailuo), and music products.

Key Features

  • MiniMax-M3: ~428B total / ~23B active parameter MoE
  • 1M-token context with MiniMax Sparse Attention
  • Native multimodal input, including computer-use agent tasks
  • Open weights under the minimax-community licence
  • Self-hosting via vLLM, SGLang, Transformers, and GGUF quantisations
  • OpenAI-compatible API with cache reads and a priority tier
  • Token Plan subscription for Claude Code, Codex, and Cursor

Pros

  • Frontier-level open-weight model with very long context
  • Well-supported self-hosting with official vLLM and SGLang recipes
  • Low API prices compared with proprietary frontier models
  • Native multimodal input without a separate vision model
  • Flat-rate Token Plan for coding agents

Cons

  • Custom community licence, not MIT or Apache 2.0
  • Self-hosting the full model needs substantial GPU hardware
  • Smaller Western developer mindshare than DeepSeek or Qwen
  • Headline benchmarks mix vendor-reported and third-party figures

MiniMax Pricing

Freemium
Open weightsFree
  • · Download MiniMax-M3 from Hugging Face
  • · minimax-community licence
  • · Serve with vLLM, SGLang, Transformers, or GGUF builds
Token Plan Plus$22/monthly
  • · For personal projects and prototyping
  • · 5-hour rolling and weekly quotas
  • · Works in Claude Code, Codex, Cursor, and more
Token Plan Max$55/monthly
  • · For daily coding with agents and multimodal work
  • · Higher 5-hour and weekly quotas
Token Plan Ultra$132/monthly
  • · For heavy agent workflows and long sessions
  • · Highest 5-hour and weekly quotas
API: MiniMax-M3Contact sales
  • · $0.30 per 1M input, $1.20 per 1M output up to 512K input (50% off list)
  • · $0.60 input and $2.40 output above 512K input
  • · Cache reads from $0.06 per 1M tokens
  • · Priority service tier at 1.5x standard pricing

Tech Stacks with MiniMax

OpenCode Dev Stack

Developer

OpenCode as the coding agent, in the terminal, the desktop app, an IDE extension, or the browser, running open-weight models such as DeepSeek, GLM, and Kimi, with your pick of GitHub or GitLab for version control.

LLM:
Version Control:
Server add-on:
Remote Access add-on:
Session Persistence add-on:
Terminal add-on:
Code Review add-on:
Model Aggregator add-on:
CI/CD add-on:
Containerization add-on:

Tools Related to MiniMax

Alternatives to MiniMax(8)

MiniMax M3 (68.8) and GLM-5.1 (66.9) sit directly next to each other on the August 2026 open-weight intelligence-index ranking, with MiniMax now ranking ahead. The closest head-to-head comparison for this addition.

MiniMax and DeepSeek are both self-hostable, open-weight Chinese-lab model families competing on the same open-weight leaderboards, with MiniMax M3 currently ranking ahead on the August 2026 intelligence-index.

MiniMax and Qwen are both self-hostable, open-weight Chinese-lab model families with strong coding-benchmark results, competing directly on the same open-weight leaderboards.

MiniMax M3 offers a 1M-token context and native multimodal input; Mistral provides a range from edge models to a MoE flagship with European hosting options. MiniMax for very long context, Mistral for EU availability and size range.

Both are Chinese open-weight families built for long context and agentic coding: Kimi is known for bilingual Chinese and English performance, MiniMax M3 adds native multimodal input with a 1M-token context. Compare them on your own coding benchmarks.

MiMo and MiniMax are both recent open-weight MoE families topping open-source agentic leaderboards; usually evaluated head-to-head when selecting an open model.

Learning Resources

No resources yet — check back soon.

Vendor

Tags

Self-hostableFree TierAI-poweredWeb

Details

Maintained
Yes