vLLM

vLLM

Open Source

Easy, fast, and cheap LLM serving for everyone.

APIs & Infrastructure
AI Runtime & Serving

Published 1 October 2026

Scores

Popularity4/5

Around 93K GitHub stars and the engine underneath many hosted inference services, so anyone self-hosting models knows it, while developers who only call hosted APIs rarely touch it directly.

Learning Curve4/5

Starting a server is one command, but running it well in production means understanding GPU memory, KV-cache sizing, quantization, and multi-GPU parallelism, plus the Kubernetes layer around it.

Flexibility5/5

Serves hundreds of model architectures with configurable quantization, parallelism, LoRA adapters, structured outputs, and speculative decoding, and can be embedded as a Python library or run as a server.

Performance5/5

PagedAttention and continuous batching set the throughput bar that other open-source engines are measured against, and the V1 engine cut scheduling overhead further.

Portability5/5

Apache-2.0, runs on GPUs and accelerators from several vendors, installs anywhere Python or Docker runs, and exposes a standard OpenAI-style API.

About vLLM

vLLM is an open-source LLM inference and serving engine that started at UC Berkeley's Sky Computing Lab and is now a PyTorch Foundation project with thousands of contributors. It takes open-weight models from Hugging Face, such as Llama, Qwen, DeepSeek, Mistral, Gemma, and gpt-oss, and serves them at high throughput behind an OpenAI-compatible API, so applications and agent frameworks built for hosted providers can point at a self-hosted endpoint instead.

Its speed comes from memory and scheduling techniques it popularised. PagedAttention manages the KV cache like virtual memory pages, which cuts wasted GPU memory and lets far more requests share a card, and continuous batching adds new requests to running batches instead of waiting for a batch to finish. The re-architected V1 engine adds near-zero-overhead prefix caching, chunked prefill, speculative decoding, structured outputs, multi-LoRA serving, quantization formats such as FP8, AWQ, and GPTQ, and tensor, pipeline, and expert parallelism for models that span many GPUs or nodes.

vLLM runs on NVIDIA and AMD GPUs, Google TPUs, AWS Trainium, and Intel Gaudi, with community plugins for other accelerators. It is installed with pip or run from official Docker images, and on Kubernetes it serves as the engine inside larger stacks such as llm-d and many managed inference platforms. Companies including Amazon, Meta, Stripe, and Roblox run it in production, and several cloud inference providers build on it.

The project is Apache-2.0 licensed and free. Its core maintainers founded Inferact, a company that funds development and offers commercial support. vLLM targets datacenter serving: for running a model on a laptop, Ollama or LM Studio are simpler, while vLLM is the usual choice once many concurrent users hit the same model.

Key Features

  • OpenAI-compatible API server for chat, completions, and embeddings
  • PagedAttention KV-cache management and continuous batching
  • Prefix caching, chunked prefill, and speculative decoding
  • Tensor, pipeline, and expert parallelism across GPUs and nodes
  • FP8, AWQ, GPTQ, and other quantization formats
  • Multi-LoRA serving from a single base model
  • Runs on NVIDIA, AMD, TPU, Trainium, and Gaudi hardware

Pros

  • Among the highest throughput of any open-source engine for many concurrent users
  • Supports new open-weight models within days of release
  • Drop-in OpenAI-compatible endpoint works with existing SDKs and agent frameworks
  • Broad hardware support avoids tying a deployment to one GPU vendor

Cons

  • Needs datacenter-class GPUs and real ops work, unlike Ollama or LM Studio on a laptop
  • Fast release cadence means flags and defaults change between versions
  • Tuning memory, parallelism, and batching for a given model takes experimentation
  • Single-user latency on small models is no better than lighter local runtimes

vLLM Pricing

Open Source

Tools Related to vLLM

Works well with vLLM(5)

Open WebUI connects to vLLM's OpenAI-compatible endpoint, giving a team a browser chat interface over models it serves on its own GPUs.

Modal's serverless GPUs are a common place to run vLLM without owning hardware: a vLLM server is defined in Python and scales to zero between requests.

vLLM serves Meta's Llama models from their Hugging Face weights behind an OpenAI-compatible API, the usual route for running Llama at production throughput.

DeepSeek's open-weight checkpoints can be self-hosted on vLLM, which spreads the large models across several GPUs with tensor parallelism and exposes them through an OpenAI-compatible API.

Qwen's open-weight models run on vLLM for high-throughput self-hosting, and Qwen's own model cards document the vLLM deployment commands.

Integrates with vLLM(2)

Alternatives to vLLM(3)

vLLM is a datacenter serving engine that batches many concurrent requests across GPUs; Ollama runs a model on a laptop or single server with one command. Use Ollama for local and small-scale use, vLLM once many users hit the same model.

vLLM lets you serve open-weight models on GPUs you run yourself; Groq serves them as a hosted API on its own LPU hardware. vLLM gives control over models and data, Groq removes the operations work.

vLLM serves open-weight models to many concurrent users on GPU servers; LM Studio is a desktop app for one person to download and chat with local models. They mark the production and personal ends of self-hosted inference.

Tags

PythonOpen SourceSelf-hostableMachine Learning

Details

Maintained
Yes
Primary language
Python
Domain
ML / AI
GitHub stars
93k
Stars updated
2026-09-30