vLLM Inference Server
AdvancedAi AgentsA self-hosted, OpenAI-compatible API for open-weight models, served from your own GPU server.
Published 6 October 2026 · Last updated 6 October 2026
About vLLM Inference Server
vLLM is an inference engine: it loads the weights of an open model onto GPUs and serves them over an OpenAI-compatible HTTP API, so any application written for a hosted provider can point at your server instead. Its speed comes from continuous batching and paged attention memory, which let many requests share one GPU without waiting for each other, and it runs more than 200 model architectures, including quantized formats such as FP8, INT4, and AWQ. The reason to run it yourself is steady, shared load: a GPU server costs the same whether it answers ten requests or ten thousand, so at volume a fixed price beats per-token billing, and prompts and fine-tuned weights stay on hardware you control.
The official Docker image runs one container: vllm/vllm-openai, started with the NVIDIA runtime, the host's shared memory, a mounted model cache, and a model name, and listening on port 8000. There is no database and no queue; the container is stateless apart from the model files it downloads once and caches. One server instance serves one model, so a second model means a second instance behind a router. The hardware is the real requirement: Linux and an NVIDIA card with compute capability 7.5 or higher, such as a T4, L4, A100, or H100, with AMD and Intel GPUs also supported, and enough card memory to hold the weights plus the context cache.
The server's --api-key flag protects only the /v1 family of paths, and other routes, including an /invocations endpoint that accepts the same inference requests, stay open to anyone who can reach the port. The project's own security guidance is therefore to put a reverse proxy in front that allows only the endpoints you mean to serve, to keep any multi-node traffic on an isolated network, and to expose nothing but the API port. A tunnel is the alternative when the clients run somewhere you do not control.
Scaling past one card is configuration. Tensor parallelism splits one model across several GPUs in a server, and the project's Kubernetes production stack adds a router, cache-aware routing, and Helm deployment for a cluster. Teams usually add a gateway in front for keys and budgets, and a chat interface such as Open WebUI when people, not only applications, will talk to the model; both are optional.
Key Features
- ✓vLLM from the official vllm/vllm-openai Docker image: one stateless container with no database
- ✓An OpenAI-compatible API on port 8000 for chat, completions, and embeddings
- ✓Continuous batching and paged attention memory, so many users share one GPU
- ✓More than 200 model architectures, with FP8, INT4, GPTQ, and AWQ quantization
- ✓Tensor, pipeline, data, and expert parallelism for models that need several GPUs
- ✓NVIDIA, AMD, and Intel GPU support, with Docker images for CUDA, ROCm, and Intel XPU
- ✓A Kubernetes production stack with Helm, cache-aware routing, and tracing
- ✓Apache-2.0 licensed, from the UC Berkeley Sky Computing Lab
When to Use vLLM Inference Server
- →A private, shared model endpoint for an engineering team's agents and coding tools
- →Replacing a per-token API bill with a fixed GPU server once volume is steady
- →Serving a fine-tuned or custom open-weight model that no hosted provider carries
- →Keeping prompts and documents on hardware you control for compliance reasons
- →The back end for a gateway or a chat interface that many people use at once
Pros
- Apache 2.0 licensed engine with batching built for many users sharing a GPU
- Applications keep an OpenAI-style client, so moving between your server and a hosted provider is a base URL change
- Runs any open-weight model with a supported architecture, including your own fine-tunes
- Scales from one card to several GPUs to a Kubernetes cluster without changing engine
- A fixed monthly cost at steady volume, instead of billing that grows with every token
Cons
- A GPU server is a fixed cost whether or not anyone is calling it, and sits far above the price of a small VPS
- The --api-key flag covers only part of the API, so any exposure beyond a trusted network needs a reverse proxy with an allowlist
- One server instance serves one model, so several models mean several instances and a router
- Memory planning is yours: the weights plus the context cache have to fit the card
- The host needs NVIDIA drivers and the container toolkit before the first container starts
LLM Options for vLLM Inference Server
The default, and the model in vLLM's own Docker quickstart. Qwen comes in the widest spread of sizes, from under 1B to very large mixture-of-experts models, so there is a version for almost any card. In 16-bit form the weights take about two bytes per parameter, so an 8B model needs roughly 16 GB of card memory and a 32B model about 64 GB, before the context cache. Most releases are Apache 2.0 and support tools and long contexts. A common pick for coding and for multilingual use.
The best-known open-weight family. At 16 bits an 8B model needs roughly 16 GB of card memory and a 70B model about 140 GB, which means several GPUs working together through tensor parallelism or a quantized copy on one large card. The weights are gated on Hugging Face: accept Meta's licence there and pass an access token to the container, which the quickstart's HF_TOKEN variable is for. The licence is Meta's own community licence, not Apache 2.0.
Open models from a European lab, many under the Apache 2.0 licence. Mistral Small, a 24B model, needs roughly 48 GB for the weights alone at 16 bits, which leaves no room for the context cache on a 48 GB card, so use an 80 GB card, an 8-bit copy on a 48 GB one, or a 4-bit copy on a 24 GB one, and it handles images and tool calls. A reasonable choice when you want a capable mid-sized model or prefer a European vendor.
DeepSeek's reasoning and general models are among the most capable open weights, and the weights are MIT licensed. The full-size models are very large mixture-of-experts networks that need a multi-GPU node with tensor or expert parallelism, which is a data-centre setup and far from one card. The distilled versions, 8B and 14B, fit a single card and write out their reasoning before the answer, which helps with maths and logic and makes replies longer.
Google's open-weight family, with sizes that suit a single card and support for images as well as text. A 12B model needs roughly 24 GB of card memory at 16 bits, which leaves no room for the context cache on a 24 GB card, so it wants a 48 GB card or an 8-bit copy, and the larger sizes want quantization or a bigger card. Read the licence terms of the exact release before building on it, since Google has changed them between generations.
These are highlighted picks. To see all the tools, check the LLM category.
Hosting Options for vLLM Inference Server
The default and the lowest fixed price here: dedicated GPU servers with no virtualisation layer between the container and the card. The entry GEX45 has a 24 GB NVIDIA card and 64 GB of RAM for about €214 a month plus a one-time setup fee of about the same, enough for an 8B model in 16-bit form, or a 24B-class model quantized to 4 bits. The GEX131 carries a 96 GB card and 256 GB of RAM at about €889 a month, which holds a 70B-class model once it is quantized. Either is a flat monthly cost whether or not anyone is calling it, so it pays off for steady shared use. Docker and the NVIDIA container toolkit are yours to install.
For teams already on AWS. A g6e.xlarge has one NVIDIA L40S with 48 GB of card memory, 4 vCPUs, and 32 GiB of RAM for about $1.86 an hour, around $1,360 a month if left running, and the larger P-series instances carry A100 and H100 cards for bigger models. GPU instance types can need a quota increase before the first launch. Stopping the instance stops the compute charge while the volume keeps the downloaded weights, and the vLLM production stack documents Kubernetes deployments on AWS for the step past one server.
The same on Google Cloud: a g2-standard-4 with one NVIDIA L4 and 24 GB of card memory costs about $0.71 an hour, around $516 a month, which suits a small model, and an a2-highgpu-1g with an A100 40 GB, 12 vCPUs, and 85 GB of RAM costs about $3.67 an hour in US regions, around $2,680 a month. GPU quotas are set per project and region, so request them early. It fits when the applications calling the endpoint already run in the same project, so requests stay on the internal network, and the vLLM production stack covers GKE.
An NC24ads A100 v4 virtual machine gives one 80 GB A100, 24 vCPUs, and 220 GiB of RAM for about $3.67 an hour on demand, around $2,680 a month if left running. Spot capacity is far cheaper, about $0.68 an hour, but it can be evicted at any time, which is fine for batch work and wrong for an endpoint people depend on. It fits organizations whose data-residency rules name an Azure region: the weights, the prompts, and the traffic all stay inside it.
These are highlighted picks. To see all the tools, check the Hosting & Cloud category.
Reverse Proxy Options for vLLM Inference Server
The family default: it discovers the vLLM container from Docker labels and renews certificates itself. Write the router so that only the API paths you serve, /v1, reach the container and everything else returns 404, because vLLM's --api-key covers only part of the API and leaves routes such as /invocations open. Raise the response timeouts, since streamed completions and long generations run for minutes and a proxy that cuts the connection breaks clients. Traefik forwards the Authorization header untouched, so the API key keeps working.
Automatic HTTPS and the shortest config: one handle block for /v1/* that forwards to port 8000, and a closing respond 404 for every other path, which keeps the unauthenticated extra routes off the network. Caddy flushes event streams as they arrive, which suits streamed tokens. The API key is one shared secret for every client, so add Caddy's own authentication or a gateway in front if different teams need separate keys.
The proxy many servers already run. A location block for /v1/ that proxies to port 8000, with a default return 404 for everything else, keeps the open routes unreachable. Turn proxy_buffering off so streamed tokens are not held back, raise proxy_read_timeout well past the 60-second default for long generations, and lift client_max_body_size when prompts carry large documents. vLLM's own docs include an NGINX least-connections example for balancing several instances, which fits when one model runs on more than one GPU server.
vLLM Inference Server Add-ons
Each addition below extends this stack with a capability the base stack works fine without. None are required: include the ones your product actually needs when building this stack, and skip the rest.
Model Aggregator Add-ons
Add a model aggregator when you want one API key and one bill for models from many providers, with automatic fallback when one of them is down, instead of setting up each provider separately.
vLLM serves one model per server instance, and its FAQ says several models on one port are not supported, so a second model means a second server behind a routing layer. LiteLLM is that layer: the entry uses the hosted_vllm/ prefix with the server's address, and the gateway adds virtual keys, budgets, rate limits, and a fallback to a hosted provider when the GPU box is down. Applications then never need vLLM's own address. LiteLLM Self-Hosted is a stack of its own.
These are highlighted picks. To see all the tools, check the AI Model Aggregators category.
Chat Interface Add-ons
Add a chat interface when people, not only applications, will talk to the model: a web app with accounts, conversation history, and a model picker that connects to the API you already serve.
The interface vLLM's own docs describe: run it in Docker and set the OpenAI API base URL to the server's /v1 address, and the served model appears at the top of the model picker. It adds accounts, conversation history, and document chat with built-in retrieval. Its licence keeps the branding in place unless you have 50 or fewer users or an enterprise licence. Point it at the reverse proxy or the gateway rather than at port 8000, so the interface and every other client use the same endpoint.
A self-hosted, MIT-licensed chat interface with the broadest multi-provider support, which suits a team that already mixes hosted and self-hosted models. It connects through a custom endpoint in its librechat.yaml file, with the vLLM server's /v1 address and the API key, using its generic OpenAI-compatible support; vLLM is not named in its docs. It runs its own services for accounts and conversation storage, so it is a heavier deployment than the model server beside it.
A self-hosted, MIT-licensed interface built around chatting with your own documents: each workspace has its own document set and vector settings. It reaches vLLM through its generic OpenAI-compatible provider, set to the server's /v1 address and key, since its docs do not name vLLM. A desktop app skips Docker entirely, which suits one person trying the endpoint, while the Docker version serves a team.
Tunnel Add-ons
Add a tunnel when you're self-hosting without a static IP or can't open inbound ports — a home server, a VPS behind restrictive network policies, or anywhere a reverse proxy alone can't reach the internet.
A public HTTPS address through an outbound-only connection, for clients that run somewhere you do not control: an application on someone else's platform can call a tunnel hostname while no inbound port is open on the server. Point the tunnel at the reverse proxy, not at port 8000, so the endpoint allowlist still applies. Free, with Cloudflare Access in front if you want a login layer. A request that gets no answer within about 125 seconds fails with a 524 and non-Enterprise plans cannot raise that, so stream long generations. A stable hostname needs a domain on Cloudflare.
The quick-start tunnel, for trying the endpoint from a laptop: one command gives a public HTTPS address and no domain is needed. The free plan includes three endpoints, 1 GB of transfer, 20,000 requests, and an interstitial page on browser visits, and streamed completions use transfer quickly, so it suits a trial and not the address every application calls. Paid plans start at $10 a month, and sustained use belongs on pay-as-you-go from $20 a month or on Cloudflare Tunnel.
Frequently Asked Questions about vLLM Inference Server
Is there a managed vLLM, and why run your own?
There is no vLLM cloud: it is a community project that started at UC Berkeley, and anyone can run it. The managed equivalents are inference providers that serve open models behind the same API and bill per token, such as Groq or Hugging Face's hosted endpoints. The trade is money against control. A provider costs nothing while idle and nothing to operate, but the bill grows with every token, and your prompts and weights sit on its infrastructure. A GPU server is a fixed monthly price, from about €214 for a 24 GB card to about €889 for a 96 GB one, so it wins only when the load is steady and shared. It also runs models no provider carries, including your own fine-tunes.
How much GPU memory does a model need?
The card must hold the weights plus the context cache. At 16-bit precision the weights take about two bytes per parameter: roughly 16 GB for an 8B model, 64 GB for a 32B model, and 140 GB for a 70B model. Quantization to 8 or 4 bits cuts that by half or three quarters, and by default vLLM claims about 92% of the card. These are estimates, since vLLM publishes no per-model figures. When one card is not enough, tensor parallelism splits the model across several GPUs in a server; the default is one.
What needs backing up, and how do upgrades work?
Almost nothing, because the container is stateless. The only large files are the model weights in the mounted Hugging Face cache, which download again if the volume is lost, so keep your own copy only of fine-tuned weights or adapters nobody else hosts. Pin the image to a version tag instead of latest: the quickstart uses latest, which moves with every release, and a new release can change defaults or kernel and quantization support, so test your model after each upgrade. The configuration to keep in version control is the launch command: model, context length, parallelism, and API key.
Is it actually free? Licensing and model terms.
The engine is Apache 2.0, with no paid edition and no feature gating, so the cost is the GPU server and the time to run it. The weights are separate: each model has its own licence. Licences run from Apache 2.0 and MIT to Meta's own community licence for Llama, and each model entry above says which applies. Gated models need an account on Hugging Face and an accepted licence before they download. Read the licence of the exact model before building a product on it, because families change terms between releases.
vLLM or Ollama: which should serve my model?
They suit different loads. Ollama is a single-machine runtime: it pulls a model by name, runs on a laptop or a modest card, and pairs with a chat interface in minutes, which is right for one person or a small team. vLLM is built for many concurrent requests on a datacenter GPU: continuous batching keeps the card busy across users, models come straight from Hugging Face, and it has no model library or interface of its own. If a handful of people chat, Ollama is simpler. If agents and applications send requests all day, vLLM serves more of them from the same card. Self-Hosted AI with Ollama and Open WebUI is the stack for the first case.
Stacks Related to vLLM Inference Server
Strapi Self-Hosted
InfrastructureSelf-hosted Strapi: an open-source headless CMS with PostgreSQL, on a server you control.
Grafana Self-Hosted
InfrastructureSelf-hosted Grafana and Prometheus: metrics dashboards and alerting on a server you run.
Airflow Self-Hosted
InfrastructureSelf-hosted Apache Airflow: the standard data-pipeline scheduler on your infrastructure.
n8n Self-Hosted
InfrastructureSelf-hosted n8n on your own server, with full control over the database, the host, and how it's exposed to the internet.
Scores
Popularity4/5
The most common engine for serving open-weight models in production, with a very large GitHub following and a contributor base across many companies and universities. It is a tool for ML and platform engineers more than a household name.
Learning Curve2/5
The API is the OpenAI one, but running it is GPU operations: drivers and the container toolkit, matching a model to card memory, tuning context length and parallelism, and a reverse proxy to cover the routes the API key leaves open. The first request is easy; a server that stays healthy under load is the work.
Flexibility5/5
Hundreds of model architectures, quantization formats, and several kinds of parallelism, on NVIDIA, AMD, and Intel hardware, from one card to a Kubernetes cluster. The launch command is the whole configuration, and the same endpoint serves applications, gateways, and chat interfaces.
Performance5/5
Continuous batching and paged attention are the reason people choose vLLM: many concurrent requests share a GPU with high throughput instead of queuing one at a time. The ceiling is the hardware, and one instance serves one model.
Portability4/5
Apache 2.0 and the OpenAI request format on the front, so applications move between vLLM and a hosted provider by changing a base URL. The engine is tied to specific GPU and driver stacks, and models are tuned to the card they run on, which makes moving hardware a re-test.
Tools in the vLLM Inference Server Stack
vLLM Inference Server Pricing
The engine is free, so the fixed cost is the GPU server, and the card decides the price: about €214 a month for a dedicated server with a 24 GB card, about €889 for a 96 GB card, and about $1,360 to $2,700 a month for a cloud GPU instance left running. Most open model weights download free, though their licences vary. Compared with a hosted provider, the bill stops growing with tokens, but it does not shrink when nobody is calling.
The card decides the price. A dedicated server with a 24 GB card is about €214 a month and one with a 96 GB card about €889 at Hetzner. Cloud GPU instances run about $1.86 an hour for a 48 GB L40S on AWS and about $3.67 an hour for an A100 on Google Cloud or Azure, roughly $1,360 to $2,700 a month if left running. Spot and reserved terms lower the hourly rate, spot at the risk of eviction.
The engine, its OpenAI-compatible server, and the Kubernetes production stack are open source, with no paid edition.
Most open models download free from Hugging Face. Each has its own licence, and gated ones need an account and an accepted licence first.
The reverse proxies and tunnels here are free, open-source software or free tiers, and TLS certificates come from Let's Encrypt.
vLLM Inference Server System Requirements
- GPU
- NVIDIA with compute capability 7.5 or higher (T4, RTX 20 series, L4, A100, H100, B200), or a supported AMD or Intel GPU
- GPU memory
- About 2 GB per billion parameters at 16-bit precision, plus room for the context cache; about half that at 8 bits and a quarter at 4 bits
- Disk
- Space for the model files, about the size of the weights, plus the Docker image
- OS
- Linux with Docker and the NVIDIA container toolkit; Python 3.10 to 3.13 for an install without Docker
No official requirements published — tekyous guidance based on the bundle's services.
The GPU and OS rows are quoted from vLLM's installation page (Linux, Python 3.10 to 3.13, NVIDIA compute capability 7.5 or higher, plus supported AMD and Intel GPUs). vLLM publishes no memory floor, so the memory row is an estimate: two bytes per parameter is the size of 16-bit weights, and the context cache comes on top, which depends on the context length and the number of concurrent requests. By default the server claims about 92% of the card's memory, and the tensor parallel size defaults to 1.