Qwen

Qwen

Freemium

Alibaba's open-weight LLM family, from 0.6B to 2.4T parameters.

LLM
Open-weight

Published 29 May 2026 · Last updated 27 September 2026

Scores

Popularity4/5

Qwen3 repo has ~27K GitHub stars; Qwen3-Coder ~16K; QwenLM org collectively has tens of thousands of stars across repositories. Dominant in Chinese AI ecosystem and rapidly growing global adoption. Strong presence on Hugging Face, Ollama library, and all major inference providers.

Learning Curve2/5

Open weights + OpenAI-compatible API means onboarding is straightforward for any developer familiar with the OpenAI SDK. Ollama makes local inference a one-command setup. Navigating the large model family and choosing the right variant adds some complexity, but the barrier is low overall.

Flexibility5/5

Apache 2.0 licence, sizes from 0.6B to 480B+, dense and MoE architectures, specialised coding and vision variants, toggleable reasoning mode, and multiple access paths (self-hosted, DashScope, third-party APIs). Maximum flexibility for any use case.

Performance5/5

Qwen3-Coder ranks among the very best open-weight coding models; Qwen3 MoE variants compete with GPT-4-class models on benchmarks; QwQ-32B rivals leading reasoning specialists; Qwen-VL leads vision-language open-weight benchmarks. Across all domains, Qwen is at or near the open-weight frontier.

Portability5/5

Apache 2.0 weights on Hugging Face run on Ollama, vLLM, llama.cpp, LM Studio, and MLX with no licence restrictions. Same models available via DashScope and third-party APIs. Maximum portability — switch providers or go self-hosted any time.

About Qwen

Qwen is Alibaba's family of large language models, released both as open weights and through Alibaba Cloud's hosted API. The flagship Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model with about 95B active parameters, native text, image, and video input, and context extensible to 1M tokens; unusually for a Max-class model, Alibaba also published its weights as Qwen3.8-2.4T-A95B, alongside a smaller Qwen3.8-27B. The range reaches down to sub-1B models for phones, so there is a checkpoint for nearly every deployment size.

Qwen3.8-Flash-Next previews the architecture planned for the next generation: a 125B-total, 6B-active mixture-of-experts model with an added n-gram memory component and hybrid linear and sparse attention, with a native 262K context that extrapolates to about 1M tokens. For coding, Qwen3-Coder comes as a 480B-total, 35B-active flagship for agentic coding and the much lighter Qwen3-Coder-Next (80B total, 3B active) for local coding agents.

Most open-weight releases use the Apache 2.0 license and run on Ollama, vLLM, SGLang, or Hugging Face Transformers. Hosted access runs through Alibaba Cloud Model Studio, an OpenAI-compatible API with Max, Plus, Flash, and Turbo tiers, a free quota for new accounts, and separate international (Singapore) and China-mainland pricing, with mainland rates substantially lower.

Key Features

  • Qwen3.8-Max: 2.4T-parameter MoE (~95B active), text, image, and video input, up to 1M context
  • Open weights for the flagship (Qwen3.8-2.4T-A95B) and a companion Qwen3.8-27B
  • Qwen3.8-Flash-Next: 125B-total / 6B-active preview of the architecture for Qwen's next major release
  • Qwen3-Coder-480B-A35B for agentic coding; Qwen3-Coder-Next (80B / 3B active) for local agents
  • Native 262K context, extrapolatable to about 1M tokens
  • Apache 2.0 on most open-weight releases
  • OpenAI-compatible hosted API through Alibaba Cloud Model Studio
  • Free token quota for new API accounts

Pros

  • Strong multilingual coverage, especially Chinese-English bilingual work
  • Apache 2.0 on most models, permissive for commercial products and fine-tuning
  • Sizes from phone-scale to a 2.4T flagship, all from one family
  • MoE models activate a small fraction of parameters, keeping inference cheap
  • Fast release cadence with steady quality gains
  • Model Studio's OpenAI-compatible API keeps migration effort low

Cons

  • Chinese-company origin raises data-governance concerns for some compliance regimes
  • Documentation sometimes appears in Chinese before English
  • Fast model turnover makes version pinning and long-term support harder to plan
  • Self-hosting or fine-tuning the large MoE models needs distributed GPU infrastructure
  • International API rates are well above China-mainland rates for the same models

Qwen Pricing

Freemium
Self-hosted (open weights)Free
  • · Apache 2.0 on most releases
  • · Run Qwen3.8, Qwen3.8-Flash-Next, or Qwen3-Coder via Ollama, vLLM, SGLang, or Transformers
API free trial (new accounts)Free
  • · 1M free tokens per model over 90 days for new accounts
  • · Covers Model Studio's Flash, Plus, and Max tiers
Model Studio API — Qwen-FlashContact sales
  • · $0.10/M input, $0.40/M output (international)
Model Studio API — Qwen3.7-Max (legacy)Contact sales
  • · $2.50/M input, $7.50/M output (international)
  • · Previous flagship, still available
Model Studio API — Qwen3.8-MaxContact sales
  • · $2.00/M input, $6.00/M output (international)
  • · Current flagship: 2.4T parameters, about 95B active
Model Studio API — Qwen-PlusContact sales
  • · $0.40/M input, $1.20/M output up to 256K context
  • · $1.20/M input, $3.60/M output from 256K to 1M context
Model Studio API — Qwen-TurboContact sales
  • · $0.05/M input; $0.20/M output, or $0.50/M in thinking mode (international)

Tech Stacks with Qwen

OpenCode Dev Stack

Developer

OpenCode as the coding agent, in the terminal, the desktop app, an IDE extension, or the browser, running open-weight models such as DeepSeek, GLM, and Kimi, with your pick of GitHub or GitLab for version control.

LLM:
Version Control:
Server add-on:
Remote Access add-on:
Session Persistence add-on:
Terminal add-on:
Code Review add-on:
Model Inference add-on:
Model Aggregator add-on:
CI/CD add-on:
Containerization add-on:

Cline Dev Stack

Developer

Cline as the coding agent, in the terminal CLI or the Cline Desktop app, running GLM or another model of your choice, with GitHub or GitLab for version control.

LLM:
Version Control:
Server add-on:
Remote Access add-on:
Session Persistence add-on:
Terminal add-on:
Code Review add-on:
Model Inference add-on:
Model Aggregator add-on:
CI/CD add-on:
Containerization add-on:

Qwen Code Dev Stack

Developer

Qwen Code as the coding agent, in the terminal, its desktop app, an editor, or a browser, running Qwen or models such as GLM, Kimi, and DeepSeek, with GitHub or GitLab for version control.

LLM:
Version Control:
Server add-on:
Remote Access add-on:
Session Persistence add-on:
Terminal add-on:
Code Review add-on:
Model Inference add-on:
Model Aggregator add-on:
CI/CD add-on:
Containerization add-on:

Tools Related to Qwen

Works well with Qwen(5)

Cline supports Alibaba Qwen models as a documented cloud provider.

Unsloth offers optimized fine-tuning recipes and quantized releases for Qwen models.

Ollama can pull and run Qwen's open-weight models locally with a single command, exposing them through its OpenAI-compatible API.

Qwen's open-weight models, including its coding variants, run locally in LM Studio as GGUF or MLX builds and can be served through its OpenAI-compatible API.

Qwen's open-weight models run on vLLM for high-throughput self-hosting, and Qwen's own model cards document the vLLM deployment commands.

Integrates with Qwen(1)

Qwen Code is the Qwen team's open-source terminal agent for these models: it calls them through Alibaba Cloud Model Studio or a local Ollama or vLLM server, and hosted use needs a paid plan or API key.

Alternatives to Qwen(8)

MiMo and Qwen are both open-weight model families with strong agentic/coding results — competing picks for a self-hostable or API-served open model.

Qwen and DeepSeek are the two leading Chinese open-weight LLM families. Qwen leads on multilingual tasks and coding (Qwen-Coder); DeepSeek-R1 leads on reasoning benchmarks and has lower API pricing.

Qwen and Mistral are popular open-weight LLM alternatives — both offer wide model size ranges and are available on Ollama and major cloud APIs. Qwen leads on multilingual (especially Asian languages) and coding; Mistral leads on European languages.

Qwen and MiniMax are both self-hostable, open-weight Chinese-lab model families with strong coding-benchmark results, competing directly on the same open-weight leaderboards.

Qwen and Meta Llama are leading open-weight LLM alternatives. Qwen excels at multilingual (especially Chinese) tasks and has a broad model range from 0.6B to 480B+; Llama has the larger Western ecosystem and fine-tune community.

Gemma is Google's compact open family for on-device and self-hosted use; Qwen spans sub-1B models to a 2.4T flagship across reasoning, coding, and vision, mostly under Apache 2.0. Gemma for small models, Qwen for range and scale.

Vendor

Alibaba Cloud

Alibaba Cloud

Website →

Tags

Open SourceSelf-hostableWeb

Details

Maintained
Yes