GLM

GLM

Freemium

Open-weight bilingual LLMs from Z.ai, with a free Flash API tier and a coding plan.

LLM
Open-weight

Published 29 May 2026 · Last updated 27 September 2026

Scores

Popularity3/5

ChatGLM-6B repo has 41k+ GitHub stars, making it one of the most-starred Chinese LLM repos. However, globally GLM has lower mindshare than DeepSeek (92k+ stars on R1) or Qwen. Dominant within the Chinese developer community; growing but limited international adoption.

Learning Curve3/5

The OpenAI-compatible API at open.bigmodel.cn is easy to adopt for any developer familiar with OpenAI. Ollama support (`ollama run glm-4.7`) makes local inference a one-command setup. The difficulty is in navigating the sprawling model family — GLM-4, GLM-4.5, GLM-4.6, GLM-4.7, GLM-Z1, GLM-5 — and the fact that much documentation is in Chinese first.

Flexibility4/5

Apache 2.0 open weights for all major variants, multiple size options (9B to 32B self-hosted; up to 744B via API), reasoning-specialist (Z1), vision (GLM-4V), and all-tools variants, plus the managed API. Solid flexibility, though fewer self-hostable sizes above 32B and less third-party cloud availability than DeepSeek or Qwen.

Performance4/5

GLM-4 All Tools benchmarks favourably against GPT-4 All Tools; GLM-Z1 reasoning models are strong at 32B; GLM-5.1 reportedly matched GPT-5.4 on coding benchmarks. However, independent Western benchmark coverage is sparser than for DeepSeek or Qwen, and the largest frontier models are proprietary.

Portability4/5

Apache 2.0 weights on Hugging Face under zai-org/, supported by Ollama, vLLM, SGLang, GGUF/llama.cpp — solid portability for the open-weight sizes. However, no major Western managed cloud (Bedrock, Azure) support as of mid-2026, and the frontier models are API-only with China-hosted infrastructure.

About GLM

GLM is the large language model family from Z.ai, the company formerly known as Zhipu AI, which grew out of Tsinghua University research. It is known for strong bilingual Chinese and English performance, agentic coding, and a steady cadence of releases whose weights are published on Hugging Face under the zai-org organisation.

The flagship is GLM-5.3, a large mixture-of-experts model with a one-million-token context window and up to 128K tokens of output. It is text-only and always reasons, with low, high, and max effort levels. GLM-5.3-Flash is a smaller, much cheaper sibling that adds native image, video, and file input. The GLM-4.x models remain available as mid-range and budget options, and specialised vision, OCR, and speech-recognition models round out the lineup.

The API is OpenAI-compatible (api.z.ai internationally, open.bigmodel.cn in China) and priced per million tokens, with cached input billed at a fraction of the normal rate. Several Flash models are completely free, which makes GLM a common choice for prototyping. The GLM Coding Plan is a flat monthly subscription for using GLM inside coding agents such as Claude Code, Cline, Kilo Code, and OpenCode, metered in credits over five-hour and weekly windows.

Licences vary by model: GLM-5.3-Flash and the GLM-4.x releases are MIT, while GLM-5.3 ships under Z.ai's own licence, so check the model card before commercial self-hosting. The API is hosted mainly in China, which matters for data-residency requirements, and English documentation sometimes trails the Chinese.

Key Features

  • GLM-5.3 flagship: 1M-token context, 128K output, three reasoning effort levels
  • GLM-5.3-Flash: low-cost model with image, video, and file input
  • Free API models (GLM-4.7-Flash, GLM-4.5-Flash)
  • Open weights on Hugging Face (zai-org), many under MIT
  • OpenAI-compatible API with discounted cached input
  • GLM Coding Plan for Claude Code, Cline, Kilo Code, and OpenCode
  • Specialised vision, OCR, and speech-recognition models

Pros

  • Strong bilingual Chinese and English models
  • Free Flash models for prototyping and light workloads
  • Open weights for self-hosting, with permissive MIT licences on many models
  • Low per-token prices compared with Western frontier APIs
  • OpenAI-compatible API makes switching simple
  • Flat-rate coding plan works inside popular coding agents

Cons

  • API hosted mainly in China, a concern for EU and US data residency
  • GLM-5.3 uses Z.ai's own licence rather than MIT
  • English documentation sometimes lags the Chinese
  • Less third-party tooling and community content outside China than DeepSeek or Qwen
  • Fast release cadence complicates version pinning

GLM Pricing

Freemium
API: Free Flash modelsFree
  • · GLM-4.7-Flash and GLM-4.5-Flash
  • · Free input, cached input, and output
Open weightsFree
  • · GLM-5.3, GLM-5.3-Flash, GLM-4.x and more on Hugging Face (zai-org)
  • · MIT for GLM-5.3-Flash and GLM-4.x; GLM-5.3 under Z.ai's licence
Coding Plan Lite$18/monthly
  • · 2,000 credits per 5 hours, 10,000 per week
  • · Vision, web search, web reader, and Zread MCP tools
  • · Discounts for quarterly and yearly billing
API: GLM-4.7 / GLM-4.6Contact sales
  • · $0.60 per 1M input tokens ($0.11 cached)
  • · $2.20 per 1M output tokens
API: GLM-4.5-AirContact sales
  • · $0.20 per 1M input tokens ($0.03 cached)
  • · $1.10 per 1M output tokens
Coding Plan ProContact sales
  • · 12,000 credits per 5 hours, 60,000 per week
  • · For frequent, complex coding work
  • · See z.ai for current pricing
Coding Plan MaxContact sales
  • · 28,000 credits per 5 hours, 140,000 per week
  • · For the heaviest agent workloads
  • · See z.ai for current pricing
API: GLM-5.3Contact sales
  • · $1.40 per 1M input tokens ($0.26 cached)
  • · $4.40 per 1M output tokens
  • · 1M context, 128K output
API: GLM-5.3-FlashContact sales
  • · $0.15 per 1M input tokens ($0.03 cached)
  • · $0.50 per 1M output tokens
  • · Image, video, and file input

Tech Stacks with GLM

OpenCode Dev Stack

Developer

OpenCode as the coding agent, in the terminal, the desktop app, an IDE extension, or the browser, running open-weight models such as DeepSeek, GLM, and Kimi, with your pick of GitHub or GitLab for version control.

LLM:
Version Control:
Server add-on:
Remote Access add-on:
Session Persistence add-on:
Terminal add-on:
Code Review add-on:
Model Inference add-on:
Model Aggregator add-on:
CI/CD add-on:
Containerization add-on:

Cline Dev Stack

Developer

Cline as the coding agent, in the terminal CLI or the Cline Desktop app, running GLM or another model of your choice, with GitHub or GitLab for version control.

LLM:
Version Control:
Server add-on:
Remote Access add-on:
Session Persistence add-on:
Terminal add-on:
Code Review add-on:
Model Inference add-on:
Model Aggregator add-on:
CI/CD add-on:
Containerization add-on:

ZCode Dev Stack

Developer

ZCode as the desktop coding agent, running GLM on Z.ai's Coding Plan or models such as Kimi, MiniMax, and Claude, with GitHub or GitLab for version control.

LLM:
Version Control:
Server add-on:
Remote Access add-on:
Code Review add-on:
CI/CD add-on:
Containerization add-on:

Tools Related to GLM

Alternatives to GLM(8)

GLM-5.1 (66.9) and MiniMax M3 (68.8) sit directly next to each other on the August 2026 open-weight intelligence-index ranking, with MiniMax now ranking ahead. The closest head-to-head comparison for this addition.

GLM is Z.ai's bilingual Chinese and English family with an OpenAI-compatible API and free Flash models; Gemma is Google DeepMind's family built for self-hosting and on-device use up to 31B. GLM for hosted bilingual use, Gemma for compact self-hosted models.

GLM and Meta Llama are open-weight LLM alternatives. GLM specialises in Chinese-English bilingual tasks with strong API support from Zhipu AI (Z.ai); Llama has the larger global open-source ecosystem.

MiMo and GLM are both open-weight model families; teams comparing open models weigh them against each other on benchmark standing and cost.

GLM and Qwen are Chinese open-weight LLM alternatives. Both are bilingual and self-hostable; Qwen has broader international adoption while GLM is tightly integrated with Zhipu AI's (Z.ai) commercial API.

GLM and Kimi are Chinese open-weight LLM alternatives — both bilingual (Chinese/English) and backed by Chinese AI labs (Zhipu/Z.ai and Moonshot AI). GLM has a longer open-source history; Kimi differentiates with very long context windows.

Vendor

Tags

Open SourceSelf-hostableWeb

Details

Maintained
Yes