Kimi

Kimi

Freemium

Moonshot AI's open-weight LLM family: long context, agentic coding, and bilingual performance.

LLM
Open-weight

Published 29 May 2026 · Last updated 27 September 2026

Scores

Popularity3/5

Very popular in China with a large consumer user base on kimi.ai. Growing international profile — K2 repo reached 10.7k GitHub stars rapidly, and the Kimi K2 family is among the most downloaded on Hugging Face. Still behind Llama and GPT-series internationally.

Learning Curve2/5

The Moonshot AI API is fully OpenAI-compatible — developers already familiar with the OpenAI SDK can switch by changing the base URL and API key. Open-weight models follow standard Hugging Face download patterns. The main friction is English documentation lagging behind Chinese-language resources.

Flexibility4/5

Open-weight releases (K2, K2.5, K2.6, Kimi-VL) allow full self-hosting and customization. The managed API covers chat, vision, and tool use. No fine-tuning on the managed API limits flexibility there, but the open weights compensate. Multiple access paths (API, chat, CLI, Hugging Face).

Performance5/5

Kimi K2 and K2.6 are top-tier on SWE-bench Verified and agentic coding benchmarks — competitive with GPT-4o and Claude Sonnet. K1.5 matched o1 on math at launch. 256K context is a standout for long-document tasks.

Portability4/5

Open-weight models are fully portable — download from Hugging Face, run on any compatible GPU infrastructure. The API uses OpenAI-compatible schema, reducing switching costs. Larger models (1T parameters) require significant hardware to self-host at full capacity.

About Kimi

Kimi is the model family and AI assistant from Moonshot AI, a Beijing-based lab. Its models are known for very long context windows, strong agentic coding, and bilingual Chinese and English performance, and Moonshot publishes the weights of its major releases on Hugging Face as well as selling API access and consumer subscriptions.

The flagship is Kimi K3, a large sparse mixture-of-experts model with a context window of about one million tokens and native vision input. Alongside it, the K2 line remains in use: Kimi K2.6 is a general multimodal model, and K2.7-Code is a coding specialist that runs in thinking mode only, with a faster highspeed variant for lower-latency serving. The K2 models have 256K-token context windows.

The weights are open but not all under the same terms. The K2 checkpoints use a modified MIT licence, while K3's licence adds conditions for the largest users: companies that sell model access as a service and earn more than twenty million dollars a year must sign a separate agreement with Moonshot, and very large consumer products must display "Kimi K3". Internal use is exempt.

Developers reach the models through the OpenAI-compatible API at platform.kimi.ai, priced per million tokens with automatic context caching that cuts the cost of repeated prompts, and through cloud platforms such as Amazon Bedrock. Kimi Code is Moonshot's terminal coding agent. The kimi.com assistant has a free tier and paid memberships, which Moonshot has restructured several times; Kimi Code is included only on the higher paid tiers.

Key Features

  • Kimi K3 flagship: sparse mixture-of-experts, ~1M-token context, native vision
  • K2.6 general model and K2.7-Code coding specialist with 256K context
  • Open weights on Hugging Face (modified MIT; K3 adds large-operator terms)
  • OpenAI-compatible API at platform.kimi.ai
  • Automatic context caching with discounted cache-hit pricing
  • Kimi Code terminal agent for agentic coding
  • Available on Amazon Bedrock as well as Moonshot's own API

Pros

  • Frontier-level open-weight models for agentic coding
  • Downloadable weights for self-hosting and research
  • OpenAI-compatible API makes switching from OpenAI simple
  • Very long context for large codebases and documents
  • Strong bilingual Chinese and English performance
  • Context caching lowers the cost of repeated prompts

Cons

  • K3's licence requires a deal with Moonshot for large Model-as-a-Service operators
  • Self-hosting the largest models needs a multi-GPU cluster
  • English documentation lags the Chinese-language resources
  • Consumer membership tiers have changed repeatedly, which makes plans hard to compare
  • Many overlapping model names (K2.5, K2.6, K2.7-Code, highspeed variants)
  • No fine-tuning through the public API

Kimi Pricing

Freemium
Open weightsFree
  • · K2 checkpoints under a modified MIT licence
  • · K3 weights free except for large Model-as-a-Service operators
  • · Self-hosted from Hugging Face
kimi.com FreeFree
  • · Free tier of the Kimi assistant
  • · Limited agent and chat usage
kimi.com MembershipContact sales
  • · Four paid membership tiers, monthly or annual
  • · More agent usage and longer conversations on higher tiers
  • · Kimi Code on the higher tiers, with a weekly usage cap
API: Kimi K3Contact sales
  • · $3.00 per 1M input tokens ($0.30 on cache hit)
  • · $15.00 per 1M output tokens
  • · ~1M-token context, vision input
API: Kimi K2.6Contact sales
  • · $0.95 per 1M input tokens ($0.16 on cache hit)
  • · $4.00 per 1M output tokens
  • · 256K context, multimodal
API: Kimi K2.7-CodeContact sales
  • · $0.95 per 1M input tokens ($0.19 on cache hit)
  • · $4.00 per 1M output tokens
  • · 256K context, thinking mode only
API: K2.7-Code HighspeedContact sales
  • · $1.90 per 1M input tokens ($0.38 on cache hit)
  • · $8.00 per 1M output tokens
  • · Lower-latency variant of K2.7-Code

Tech Stacks with Kimi

Orca Dev Stack

Developer

Orca runs several coding agents side by side, each in its own git worktree, with your pick of agents and models underneath and GitHub or GitLab for version control.

Coding Agent:
LLM:
Version Control:
Server add-on:
Remote Access add-on:
Code Review add-on:
CI/CD add-on:

Herdr Dev Stack

Developer

Herdr keeps several coding agents running in persistent terminal sessions on your laptop or a server, with your pick of agents and models and GitHub or GitLab for version control.

Coding Agent:
LLM:
Version Control:
Server add-on:
Remote Access add-on:
Terminal add-on:
Code Review add-on:
CI/CD add-on:

Paseo Dev Stack

Developer

Paseo runs coding agents on a machine you control, a laptop, home server, or VPS, and lets you steer them from desktop, web, phone, or terminal, with your pick of agents, models, and git host.

Coding Agent:
LLM:
Version Control:
Server add-on:
Remote Access add-on:
Code Review add-on:
CI/CD add-on:

Tools Related to Kimi

Alternatives to Kimi(8)

Kimi is Moonshot AI's long-context family with strong agentic coding and Chinese and English performance; Mistral is Europe's open-weight family, available via its API, major clouds, or self-hosting. Mistral suits EU data needs, Kimi long-context coding.

Both are Chinese open-weight families built for long context and agentic coding: Kimi is known for bilingual Chinese and English performance, MiniMax M3 adds native multimodal input with a 1M-token context. Compare them on your own coding benchmarks.

Kimi and Qwen are open-weight LLM alternatives from Chinese AI labs. Kimi differentiates with extremely long context (up to 1M tokens); Qwen offers a wider range of model sizes and better multilingual coverage.

Kimi and GLM are Chinese open-weight LLM alternatives from Moonshot AI and Zhipu AI (Z.ai) respectively. Both are bilingual; Kimi differentiates with long context (up to 1M tokens) and strong reasoning via K1.5/K2.

Kimi and Meta Llama are open-weight LLM alternatives. Kimi differentiates with very long context windows and strong reasoning; Llama has a much larger global ecosystem and fine-tune community.

Kimi and Google Gemma are open-weight LLM alternatives. Kimi leads on context length and reasoning; Gemma leads on small-model efficiency suitable for edge and on-device deployment.

Vendor

Tags

Open SourceWeb

Details

Maintained
Yes