LiteLLM Self-Hosted

AdvancedAi Agents

A self-hosted gateway that gives every application one OpenAI-compatible endpoint, with virtual keys, budgets, and spend tracking you control.

Published 6 October 2026 · Last updated 6 October 2026

Core Tools
PostgreSQL
PostgreSQL
Docker
Docker
Prometheus
Prometheus
LiteLLM
LiteLLM
Hosting
Hetzner
Hostinger
DigitalOcean
Amazon Web Services
Google Cloud Platform
+2
Reverse Proxy
Traefik
Caddy
NGINX
Self-Hosted PaaS
Coolify
Dokploy

About LiteLLM Self-Hosted

LiteLLM is an AI gateway. Your applications and agents call one OpenAI-compatible address, and the gateway translates each request for the right provider, whether that is Anthropic, OpenAI, Gemini, Bedrock, Azure, or a model you run yourself. Around that routing it adds virtual API keys with budgets and rate limits per team, user, and key, spend tracking, fallbacks between deployments, and an admin interface. The reason to run it yourself is that every prompt and every provider credential passes through it: a self-hosted gateway keeps both on infrastructure you administer, and adds no per-token charge on top of what the providers bill.

The official Docker Compose file runs three services: the LiteLLM proxy on port 4000, PostgreSQL for keys, teams, budgets, and spend logs, and Prometheus for metrics. PostgreSQL is what makes keys, budgets, and spend tracking work, so it is a hard dependency. Redis is not in the file; it becomes the shared state once more than one proxy instance runs, holding rate-limit counters, router state, and a response cache. The vendor's production guidance is 1 vCPU and 4 GB of memory per proxy instance, which puts the whole bundle on an 8 GB server, and Kubernetes with the Helm chart, or the Terraform modules for the large clouds, is the route once a single server is no longer enough.

The API and the admin interface share port 4000. The Compose file publishes that port on every interface, and it also publishes PostgreSQL on 5432 and Prometheus on 9090 with a default database password, so it is a starting point, not a hardened deployment: bind or close the two extra ports, replace the password, and set the master key first. Put a reverse proxy with TLS in front of port 4000, because provider keys and bearer tokens cross it on every request, or install LiteLLM from a self-hosted PaaS template that brings its own proxy. A tunnel is the alternative when the clients run somewhere you do not control.

Upstream models are configuration, not code. Each entry in the config file or the admin interface names a provider, a model, and a credential, and the gateway exposes it under whatever name your applications call. Several entries can share one name to balance load or fall back across providers, and a model you run yourself is one more entry, which is how local and hosted models end up behind one address. Tracing prompts and costs in an external tool is optional, since the bundled Prometheus covers the gateway's own metrics rather than request contents.

Key Features

  • ✓LiteLLM in Docker from the official Compose file: the proxy, PostgreSQL, and Prometheus
  • ✓One OpenAI-compatible endpoint in front of more than 100 model providers
  • ✓Virtual API keys with budgets, rate limits, and spend tracking per key, team, and user
  • ✓Load balancing and fallbacks across several deployments of the same model
  • ✓An admin interface on the same port for keys, models, and spend
  • ✓Prometheus metrics in the open-source edition, plus callbacks to external tracing tools

When to Use LiteLLM Self-Hosted

  • →One internal endpoint for every team, with a budget and a rate limit per key instead of shared provider keys
  • →Switching or mixing model providers without changing application code, with fallbacks when one provider is down
  • →Putting a model server you run and hosted APIs behind the same address
  • →Cost allocation: spend per team, project, or user from one place
  • →Keeping provider credentials out of application code and developer laptops

Pros

  • MIT-licensed core with virtual keys, budgets, load balancing, and Prometheus metrics, and no per-token fee
  • Applications keep one OpenAI-style client whatever the provider behind it
  • Prompts and provider credentials stay on infrastructure you operate
  • Standard pieces everywhere: Docker, PostgreSQL, Helm, and official Terraform modules for two large clouds

Cons

  • A service every model call depends on: the vendor's production guidance is two or more instances behind a load balancer
  • PostgreSQL is a hard dependency, and Redis joins it as soon as there is more than one instance
  • SSO beyond five users, SCIM, JWT and OIDC authentication, and audit logs need a paid Enterprise license
  • The official Compose file publishes the database and metrics ports with a default password, so it needs hardening before it faces the internet

Hosting Options for LiteLLM Self-Hosted

Hetzner

Deploy LiteLLM Self-Hosted on Hetzner

The default. A server in the 8 GB class costs about €11 a month on the shared-vCPU line, which holds the proxy at its recommended 4 GB with PostgreSQL and Prometheus beside it. LiteLLM publishes no host-specific guide, so this is the general Docker route: the Compose file with a real master key and database password, the database and metrics ports closed or bound to loopback, and a proxy in front of port 4000. One server is a single point of failure for every model call your applications make, so plan a second instance once the traffic matters.

Hostinger

Deploy LiteLLM Self-Hosted on Hostinger

The budget entry. The KVM 2 plan carries 2 vCPUs, 8 GB of RAM, and 100 GB of NVMe at $8.99 a month on a two-year term, $14.99 at the standard rate, which fits the proxy and its two companions. Root access is full and the panel adds weekly backups, but TLS, upgrades, and closing the extra ports stay your job. Run one proxy worker here: the docs size each worker at a core and 4 GB, so 8 GB covers one worker plus PostgreSQL and Prometheus, and more workers or heavier traffic mean moving up a plan.

DigitalOcean

Deploy LiteLLM Self-Hosted on DigitalOcean

Per-second billing and snapshots from the panel. The 8 GB Basic Droplet with 4 vCPUs is $48 a month and holds one proxy worker with PostgreSQL and Prometheus beside it, while the 4 GB plan at $24 meets the proxy's own memory floor but leaves PostgreSQL and Prometheus almost no room, so it suits a trial only. Managed PostgreSQL can take over the database later, which also removes the password-protected database port from the server. Place the Droplet in the region nearest your applications, since the gateway adds a hop to every call.

Amazon Web Services

Deploy LiteLLM Self-Hosted on Amazon Web Services

The official infrastructure-as-code route: LiteLLM maintains a Terraform module that runs the gateway, backend, and admin interface as separate ECS Fargate services behind a load balancer, with Aurora PostgreSQL and ElastiCache Redis, and the Helm chart covers EKS. Bedrock models sit in the same account, so those calls can use IAM roles instead of stored provider keys. It costs several times a single server and removes the operations: a managed database, a shared cache, and horizontal scaling out of the box.

Google Cloud Platform

Deploy LiteLLM Self-Hosted on Google Cloud Platform

The same official route on Google's side: a maintained Terraform module that runs the services on Cloud Run with Cloud SQL for PostgreSQL and Memorystore, or the Helm chart on GKE. Vertex AI models are available from the same project, and Cloud Run scales the gateway with traffic. It fits when the applications calling the gateway already run in Google Cloud, so requests stay on the internal network.

Railway

Deploy LiteLLM Self-Hosted on Railway

A container platform with a one-click template linked from LiteLLM's own README and docs: it deploys the proxy as an always-on service, and the docs say to set PORT=4000. The template runs the proxy alone, so add Railway's PostgreSQL to the project and point DATABASE_URL at it, or keys, budgets, and spend tracking have nowhere to live, and add Redis there once a second instance runs. Railway terminates TLS itself, so the reverse proxy or PaaS choice doesn't apply. Billing is usage-based on top of the $5 a month Hobby plan, with memory at about $10 per GB a month, so a proxy holding the vendor's recommended 4 GB plus a database lands in the tens of dollars, more than a small VPS in exchange for no server to patch.

Render

Deploy LiteLLM Self-Hosted on Render

A container platform with a Deploy to Render button in LiteLLM's README: the repository's blueprint runs the proxy as a web service from the main-stable image, generates the master key, and health-checks /health/liveliness. It provisions no database, so add Render Postgres and set DATABASE_URL, and a Key Value instance for Redis once there is more than one proxy. TLS ends at the platform, so the reverse proxy or PaaS choice doesn't apply. The vendor's 4 GB per-instance guidance means the 2 CPU and 4 GB instance at $85 a month plus the database; the free instance sleeps after 15 minutes and its free database expires after 30 days, so both suit a demo only. Pin a version tag instead of main-stable before real use.

These are highlighted picks. To see all the tools, check the Hosting & Cloud category.

Reverse Proxy Options for LiteLLM Self-Hosted

Traefik

LiteLLM Self-Hosted with Traefik

The family default and the least work here: it discovers the proxy container from Docker labels and renews certificates itself. LiteLLM serves its API and admin interface from the same port, 4000, so one router covers both. Raise the idle and response timeouts, because a proxy that cuts a connection mid-answer breaks streamed completions and slow reasoning models. The Compose file also publishes PostgreSQL and Prometheus on every interface, so close those mappings before adding any proxy.

Caddy

LiteLLM Self-Hosted with Caddy

Automatic HTTPS from a few lines of config: one site block that forwards to port 4000 covers the API and the admin interface. Caddy flushes event streams as they arrive, which suits streamed completions, and its default timeouts are generous for long generations. Provider keys and bearer tokens cross this hostname on every request, so the TLS it manages is the main reason to put it in front of the gateway at all.

NGINX

LiteLLM Self-Hosted with NGINX

The proxy many servers already run. Two of its defaults fight a gateway: response buffering, which delays streamed tokens until proxy_buffering is turned off for this location, and the 60-second read timeout, which cuts long generations unless proxy_read_timeout is raised. Forward the Host and X-Forwarded headers so the admin interface builds correct links, and raise client_max_body_size if clients send large prompts or files.

Self-Hosted PaaS Options for LiteLLM Self-Hosted

Coolify

LiteLLM Self-Hosted with Coolify

A free, self-hosted deploy platform with a LiteLLM template in its service library. The template runs the litellm-database image on the main-stable tag with PostgreSQL and Redis, generates the master key and the admin login, switches to production mode, turns telemetry off, and wires Redis in as the response cache and router state, which is closer to the vendor's production advice than the bare Compose file. It runs no Prometheus, and main-stable moves with every release, so pin a version tag before real use. Coolify's own proxy handles the domain and TLS.

Dokploy

LiteLLM Self-Hosted with Dokploy

Free and self-hosted, with a LiteLLM template that runs the proxy beside a PostgreSQL 16 container, generates the master key and the admin password, and serves the service on your domain with HTTPS and Traefik routing. The template uses the main-latest image tag, which the vendor advises against for production, so change it to a version tag first. It has no Redis or Prometheus, which is fine for one instance. Domains and TLS come with the platform, so no separate reverse proxy is added next to it.

LiteLLM Self-Hosted Add-ons

Each addition below extends this stack with a capability the base stack works fine without. None are required: include the ones your product actually needs when building this stack, and skip the rest.

LLM Observability Add-ons

Add LLM observability when you want to see every model call, tool call, and token cost inside a run, so a wrong answer can be traced to the step that caused it.

Langfuse

LiteLLM Self-Hosted with Langfuse

The tracing backend LiteLLM supports with a built-in callback: set the public and secret keys and the host, and every request through the gateway becomes a trace with its cost and latency, with a team's or a key's traffic routable to its own Langfuse project. A self-hosted Langfuse server has to be recent enough for the SDK the proxy uses, 3.63.0 or newer for SDK v4, so upgrade it before upgrading LiteLLM. The bundled Prometheus covers gateway metrics, not prompts, which is the gap this fills.

LangSmith

LiteLLM Self-Hosted with LangSmith

A second callback for teams already tracing with LangChain's platform: the proxy sends each request to a LangSmith project, set with an API key and a project name, and a custom endpoint can be pointed at with LANGSMITH_BASE_URL. It is a hosted, closed-source service, with self-hosting limited to the Enterprise plan, so prompts leave your server unless you choose that plan. The free Developer plan covers one seat and 5,000 traces a month.

Model Inference Add-ons

Add model inference when you want an open-weight model in the mix: on your own hardware for privacy, or on a hosted inference provider for speed and low per-token prices.

Ollama

LiteLLM Self-Hosted with Ollama

A model you run on your own machine becomes one more entry: name it ollama_chat/<model> with the Ollama server as api_base, which defaults to port 11434. The gateway then gives it a virtual key, a budget, and a fallback to a hosted model like any other entry. Tool calling depends on the model, since LiteLLM notes that not every Ollama model supports function calls and falls back to JSON mode, so test the agent's tools first. When the gateway runs in a container, localhost points at the container, so use the host's address.

vLLM

LiteLLM Self-Hosted with vLLM

For open-weight models served to many users from a GPU server: the entry uses the hosted_vllm/ prefix and the vLLM server's OpenAI-compatible address as api_base, and chat, embedding, and reranking endpoints are supported. The gateway adds what the inference server lacks, which is per-key budgets, rate limits, and fallback to a hosted provider when the GPU box is down. It needs a datacenter-class GPU and someone to operate it.

Groq

LiteLLM Self-Hosted with Groq

A hosted inference provider that runs open-weight models on its own chips: set GROQ_API_KEY and name models groq/<model>. Its speed makes it a natural fast tier or first fallback in the gateway's routing, because every agent step waits on a response. Billing is per token with a free rate-limited tier, and other hosted providers of open models connect the same way.

These are highlighted picks. To see all the tools, check the AI Runtime & Serving category.

Tunnel Add-ons

Add a tunnel when you're self-hosting without a static IP or can't open inbound ports — a home server, a VPS behind restrictive network policies, or anywhere a reverse proxy alone can't reach the internet.

Cloudflare Tunnel

LiteLLM Self-Hosted with Cloudflare Tunnel

A public HTTPS address through an outbound-only connection, for clients that run somewhere you do not control: an application on someone else's platform can call a tunnel hostname while no inbound port is open on the server. Free, with Cloudflare Access in front if the admin interface should get its own login layer. Point the tunnel at port 4000 only, never at the database or metrics ports the Compose file publishes. A proxied request that gets no answer within about 125 seconds fails with a 524, and non-Enterprise plans cannot raise that, so stream long generations. A stable hostname needs a domain on Cloudflare.

ngrok

LiteLLM Self-Hosted with ngrok

The quick-start tunnel, for trying the gateway from a laptop: one command gives a public HTTPS address and no domain is needed. The free plan includes three endpoints, 1 GB of transfer, 20,000 requests, and an interstitial page on browser visits, which long completions and a busy gateway outgrow quickly, so it suits a trial and not the address every application calls. Paid plans start at $10 a month, and sustained use belongs on pay-as-you-go from $20 a month or on Cloudflare Tunnel.

Caching Add-ons

Add caching when the same reads or computations repeat and you want them answered from memory: a cache holds hot data, shared counters, and session state in front of a database or an upstream API.

Redis

LiteLLM Self-Hosted with Redis

The shared state a gateway needs once it runs on more than one server. One proxy instance works without it, but with two or more, the vendor's production guide makes Redis 7.0 or newer essential: it holds the rate-limit counters, the router's view of which deployments are healthy, and a response cache that every instance reads, so a budget or a limit means the same thing whichever instance answers. A cached response to a repeated identical request skips the provider call, so it saves tokens as well as time. Around 1,000 requests a second the guide adds a Redis buffer for spend writes so PostgreSQL does not deadlock. Redis is cache and counters here, not durable state. Run it as a container next to the proxy, or use the managed services the official Terraform modules provision.

Frequently Asked Questions about LiteLLM Self-Hosted

Is there a managed LiteLLM, or do I have to self-host it?

There is no managed LiteLLM: its own site lists an open-source edition and an Enterprise edition, and both run on infrastructure you operate. So self-hosting is the product, not a cost-saving alternative to a hosted plan. What it buys is control: prompts and provider credentials pass only through your server, you contract with each provider directly, and nothing is added per token. What it costs is operating a service that every model call depends on. The hosted counterpart for the same job is a different product, OpenRouter, compared below. The free edition is enough to start: SSO is free for up to five users, and Enterprise is priced by quote, sized to request volume rather than tokens.

What size server does the gateway need?

LiteLLM's production guide sizes each proxy instance at 1 vCPU and 4 GB of memory, as both request and limit, and scales both with the worker count: eight workers on one instance want 8 vCPU and 32 GB. The 4 GB floor is real, because the database engine inside the proxy keeps its peak memory instead of returning it. That figure covers the proxy alone. With PostgreSQL and Prometheus on the same host, which is what the Compose file starts, an 8 GB server is the practical size for one instance. For production the guide recommends at least two instances behind a load balancer, and Redis from version 7 once there is more than one, so they share rate-limit counters and router state. Disk has no official floor; it grows with the spend logs PostgreSQL keeps.

Where does state live, and what needs backing up?

Almost everything durable is in PostgreSQL: virtual keys, teams, budgets, spend logs, and any models added through the admin interface, whose provider credentials are stored encrypted with the salt key. That salt key is the one thing to treat like a database backup, because the docs state it cannot be rotated once models have been added, and a restored database without it holds credentials nobody can read. The config file belongs in version control. Redis, when present, is cache and counters, so losing it costs warm responses and not data. Prometheus keeps 15 days of metrics in the Compose file. Nothing here is backed up for you: pg_dump the database on a schedule and store the salt key next to the backup, not in it.

How do upgrades and the Enterprise license work?

Releases ship as container images, and the vendor says to pin a version tag in production, because the main-stable tag the Compose file uses moves with every release. Database migrations run when the proxy starts, which is fine for one instance; with several, the docs advise setting DISABLE_SCHEMA_UPDATE on the serving pods and running migrations as a separate job so replicas do not race. The code outside the enterprise directory is MIT-licensed, and virtual keys, budgets, load balancing, and Prometheus metrics are all in it. Enterprise is a separate license sold by quote, and it also covers secret-manager integration, IP allowlists, and multi-region deployment.

LiteLLM or OpenRouter: which gateway should I use?

They do the same job from opposite sides. OpenRouter is hosted: one account, one prepaid balance, hundreds of models behind one endpoint, and a fee on credit purchases, 5.5% on the standard plan, with nothing to run. LiteLLM is the gateway you operate: you hold your own provider accounts and keys, pay each provider directly with no markup, and get budgets, virtual keys, and routing rules of your own, but you also carry its uptime and upgrades. OpenRouter fits one developer or a small team that wants every model on one bill. LiteLLM fits when keys and spend must be governed per team, when prompts cannot go through a third party, or when a model you run yourself belongs behind the same address. The two combine: OpenRouter can be one upstream provider inside LiteLLM.

Scores

Popularity4/5

The most widely used open-source LLM gateway, with tens of thousands of GitHub stars and a large number of container pulls, and a place in many teams' agent and application stacks. Templates exist on two deploy platforms and as official Terraform modules, though it is a platform-team tool more than a household name.

Learning Curve3/5

The idea is simple, since an application only changes its base URL, but running the gateway is not: a config file, PostgreSQL, a master key and a salt key that cannot be rotated, and hardening of ports the Compose file leaves open. Budgets and virtual keys take reading before they behave the way a team expects.

Flexibility5/5

Any provider, local or hosted, is one more entry, and routing, fallbacks, budgets, and logging callbacks are all configuration. The deployment bends as well: Compose, a deploy platform, Helm on any cloud, or the maintained Terraform modules for AWS and GCP.

Performance4/5

The vendor reports single-digit-millisecond overhead at a thousand requests a second, and the production guide covers workers, a shared cache, and a spend-write buffer for higher loads. Every call crosses the gateway and a database write path, and one instance is a single point of failure, which keeps it off the top score.

Portability5/5

An MIT-licensed core on PostgreSQL and standard containers, speaking the OpenAI request format on both sides. Moving off it means pointing applications back at a provider's base URL and exporting keys and spend from one database, and moving between Compose, Kubernetes, and a cloud is configuration.

Tools in the LiteLLM Self-Hosted Stack

Databases

DevOps & CI/CD

Observability & Monitoring

AI Infrastructure

Hosting (choose one)

Reverse Proxy or PaaS (choose one)

Reverse Proxy

Self-Hosted PaaS

Add-ons (optional — add any, or none)

LLM Observability

Model Inference

Tunnel

Caching

LiteLLM Self-Hosted Pricing

From ~€11/mo Free to start

The software is free, so the fixed cost is a small server: the vendor recommends 1 vCPU and 4 GB of memory per proxy instance, and with PostgreSQL and Prometheus beside it an 8 GB server is the practical size, from about €11 a month. The real spend is model usage, which each provider bills at its own prices; LiteLLM adds no per-token fee. Enterprise features, such as SSO beyond five users, need a paid license quoted by sales.

Server (VPS or cloud)$9–48/mo

The vendor sizes each proxy instance at 1 vCPU and 4 GB of memory. With PostgreSQL and Prometheus on the same host, an 8 GB server is the practical size: about €11 a month at Hetzner, $8.99 at Hostinger on a two-year term, and $48 for an 8 GB cloud server at DigitalOcean. Production adds a second instance, a managed database, and a shared cache.

LiteLLMFree (MIT core)

The proxy, virtual keys, budgets, load balancing, and Prometheus metrics are MIT-licensed. SSO beyond five users, SCIM, JWT and OIDC authentication, and audit logs need an Enterprise license, priced by quote and never per token.

Model providersPay per token

Every request is billed by the provider it reaches, at that provider's own prices, through your own account. The gateway's spend tracking shows the same figures per key, team, and user.

Exposure (optional)Free

The reverse proxies, deploy platforms, and tunnels here are free, open-source software or free tiers, and TLS certificates come from Let's Encrypt.

LiteLLM Self-Hosted System Requirements

source
CPU
1 vCPU per proxy instance, one more for each extra worker
RAM
4 GB per proxy instance, as both request and limit
Disk
No official floor; PostgreSQL grows with spend logs, and Prometheus keeps 15 days of metrics in the Compose file
OS
Any Linux with Docker for the Compose path; Kubernetes with the Helm chart for production

Quoted from LiteLLM's production guide: each pod gets 1 vCPU and 4 GiB of memory as both requests and limits, scaled with the worker count (a container running 8 workers needs 8 vCPU and 32 GiB). The 4 GiB floor exists because the Prisma query engine's resident memory acts as a high-water mark that is not returned to the operating system. The guide sizes the proxy only. The 8 GB server named in the stack's copy is an estimate for the proxy plus PostgreSQL and Prometheus on one host, not a vendor figure. Production guidance is at least two instances behind a load balancer, with Redis 7.0 or newer once there is more than one.