
GLM
FreemiumOpen-weight bilingual LLMs from Z.ai, with a free Flash API tier and a coding plan.
Published 29 May 2026 · Last updated 27 September 2026
Scores
Popularity3/5
ChatGLM-6B repo has 41k+ GitHub stars, making it one of the most-starred Chinese LLM repos. However, globally GLM has lower mindshare than DeepSeek (92k+ stars on R1) or Qwen. Dominant within the Chinese developer community; growing but limited international adoption.
Learning Curve3/5
The OpenAI-compatible API at open.bigmodel.cn is easy to adopt for any developer familiar with OpenAI. Ollama support (`ollama run glm-4.7`) makes local inference a one-command setup. The difficulty is in navigating the sprawling model family — GLM-4, GLM-4.5, GLM-4.6, GLM-4.7, GLM-Z1, GLM-5 — and the fact that much documentation is in Chinese first.
Flexibility4/5
Apache 2.0 open weights for all major variants, multiple size options (9B to 32B self-hosted; up to 744B via API), reasoning-specialist (Z1), vision (GLM-4V), and all-tools variants, plus the managed API. Solid flexibility, though fewer self-hostable sizes above 32B and less third-party cloud availability than DeepSeek or Qwen.
Performance4/5
GLM-4 All Tools benchmarks favourably against GPT-4 All Tools; GLM-Z1 reasoning models are strong at 32B; GLM-5.1 reportedly matched GPT-5.4 on coding benchmarks. However, independent Western benchmark coverage is sparser than for DeepSeek or Qwen, and the largest frontier models are proprietary.
Portability4/5
Apache 2.0 weights on Hugging Face under zai-org/, supported by Ollama, vLLM, SGLang, GGUF/llama.cpp — solid portability for the open-weight sizes. However, no major Western managed cloud (Bedrock, Azure) support as of mid-2026, and the frontier models are API-only with China-hosted infrastructure.
About GLM
GLM is the large language model family from Z.ai, the company formerly known as Zhipu AI, which grew out of Tsinghua University research. It is known for strong bilingual Chinese and English performance, agentic coding, and a steady cadence of releases whose weights are published on Hugging Face under the zai-org organisation.
The flagship is GLM-5.3, a large mixture-of-experts model with a one-million-token context window and up to 128K tokens of output. It is text-only and always reasons, with low, high, and max effort levels. GLM-5.3-Flash is a smaller, much cheaper sibling that adds native image, video, and file input. The GLM-4.x models remain available as mid-range and budget options, and specialised vision, OCR, and speech-recognition models round out the lineup.
The API is OpenAI-compatible (api.z.ai internationally, open.bigmodel.cn in China) and priced per million tokens, with cached input billed at a fraction of the normal rate. Several Flash models are completely free, which makes GLM a common choice for prototyping. The GLM Coding Plan is a flat monthly subscription for using GLM inside coding agents such as Claude Code, Cline, Kilo Code, and OpenCode, metered in credits over five-hour and weekly windows.
Licences vary by model: GLM-5.3-Flash and the GLM-4.x releases are MIT, while GLM-5.3 ships under Z.ai's own licence, so check the model card before commercial self-hosting. The API is hosted mainly in China, which matters for data-residency requirements, and English documentation sometimes trails the Chinese.
Key Features
- GLM-5.3 flagship: 1M-token context, 128K output, three reasoning effort levels
- GLM-5.3-Flash: low-cost model with image, video, and file input
- Free API models (GLM-4.7-Flash, GLM-4.5-Flash)
- Open weights on Hugging Face (zai-org), many under MIT
- OpenAI-compatible API with discounted cached input
- GLM Coding Plan for Claude Code, Cline, Kilo Code, and OpenCode
- Specialised vision, OCR, and speech-recognition models
Pros
- Strong bilingual Chinese and English models
- Free Flash models for prototyping and light workloads
- Open weights for self-hosting, with permissive MIT licences on many models
- Low per-token prices compared with Western frontier APIs
- OpenAI-compatible API makes switching simple
- Flat-rate coding plan works inside popular coding agents
Cons
- API hosted mainly in China, a concern for EU and US data residency
- GLM-5.3 uses Z.ai's own licence rather than MIT
- English documentation sometimes lags the Chinese
- Less third-party tooling and community content outside China than DeepSeek or Qwen
- Fast release cadence complicates version pinning
GLM Pricing
Freemium- · GLM-4.7-Flash and GLM-4.5-Flash
- · Free input, cached input, and output
- · GLM-5.3, GLM-5.3-Flash, GLM-4.x and more on Hugging Face (zai-org)
- · MIT for GLM-5.3-Flash and GLM-4.x; GLM-5.3 under Z.ai's licence
- · 2,000 credits per 5 hours, 10,000 per week
- · Vision, web search, web reader, and Zread MCP tools
- · Discounts for quarterly and yearly billing
- · $0.60 per 1M input tokens ($0.11 cached)
- · $2.20 per 1M output tokens
- · $0.20 per 1M input tokens ($0.03 cached)
- · $1.10 per 1M output tokens
- · 12,000 credits per 5 hours, 60,000 per week
- · For frequent, complex coding work
- · See z.ai for current pricing
- · 28,000 credits per 5 hours, 140,000 per week
- · For the heaviest agent workloads
- · See z.ai for current pricing
- · $1.40 per 1M input tokens ($0.26 cached)
- · $4.40 per 1M output tokens
- · 1M context, 128K output
- · $0.15 per 1M input tokens ($0.03 cached)
- · $0.50 per 1M output tokens
- · Image, video, and file input
Tech Stacks with GLM
OpenCode Dev Stack
DeveloperOpenCode as the coding agent, in the terminal, the desktop app, an IDE extension, or the browser, running open-weight models such as DeepSeek, GLM, and Kimi, with your pick of GitHub or GitLab for version control.
Cline Dev Stack
DeveloperCline as the coding agent, in the terminal CLI or the Cline Desktop app, running GLM or another model of your choice, with GitHub or GitLab for version control.
ZCode Dev Stack
DeveloperZCode as the desktop coding agent, running GLM on Z.ai's Coding Plan or models such as Kimi, MiniMax, and Claude, with GitHub or GitLab for version control.
Tools Related to GLM
Works well with GLM(3)
OpenCode supports Z.AI's API and coding-plan subscription as a GLM provider.
Unsloth maintains quantized builds and fine-tuning support for GLM models.
Ollama can pull and run GLM's open-weight models locally with a single command, exposing them through its OpenAI-compatible API.
Integrates with GLM(1)
Z.AI's coding harness, built to run the GLM model family.
Alternatives to GLM(8)
GLM-5.1 (66.9) and MiniMax M3 (68.8) sit directly next to each other on the August 2026 open-weight intelligence-index ranking, with MiniMax now ranking ahead. The closest head-to-head comparison for this addition.
GLM is Z.ai's bilingual Chinese and English family with an OpenAI-compatible API and free Flash models; Gemma is Google DeepMind's family built for self-hosting and on-device use up to 31B. GLM for hosted bilingual use, Gemma for compact self-hosted models.
GLM and Meta Llama are open-weight LLM alternatives. GLM specialises in Chinese-English bilingual tasks with strong API support from Zhipu AI (Z.ai); Llama has the larger global open-source ecosystem.
MiMo and GLM are both open-weight model families; teams comparing open models weigh them against each other on benchmark standing and cost.
GLM and Qwen are Chinese open-weight LLM alternatives. Both are bilingual and self-hostable; Qwen has broader international adoption while GLM is tightly integrated with Zhipu AI's (Z.ai) commercial API.
GLM and Kimi are Chinese open-weight LLM alternatives — both bilingual (Chinese/English) and backed by Chinese AI labs (Zhipu/Z.ai and Moonshot AI). GLM has a longer open-source history; Kimi differentiates with very long context windows.