Self-Hosted AI with Ollama and Open WebUI

IntermediateAi Agents

Your own private AI chat: open-weight models served by Ollama, with Open WebUI as the interface.

Published 2 October 2026

Core Tools
Docker
Docker
Ollama
Ollama
Open WebUI
Open WebUI
Database
SQLite
PostgreSQL
LLM
Google Gemma
Qwen
Meta Llama
DeepSeek
Mistral

About Self-Hosted AI with Ollama and Open WebUI

This stack is a private AI assistant that runs on hardware you own. Ollama downloads open-weight models and serves them on your machine, and Open WebUI puts a chat interface in front of them that works much like the hosted chatbots: conversations with history, file and image uploads, web search, voice, and a model picker. Prompts, documents, and answers stay on the machine. There is no subscription and no per-token bill, and it keeps working without an internet connection.

The official Compose file starts two containers: Ollama, which holds the model files, and Open WebUI, which keeps accounts, chats, and settings in a built-in SQLite database, with PostgreSQL as the alternative for a larger installation. The first person to sign up becomes the administrator and later accounts wait for approval, so one instance can serve a household or a team from their browsers. Open WebUI also searches your own documents: uploaded files are embedded into a built-in vector store, and a dedicated vector database such as Qdrant can be added when the collection outgrows it.

The model decides what hardware you need. Google Gemma is the default, and its 12B version is a download of about 8 GB that fits a single consumer graphics card or a recent Mac, while Qwen, Meta Llama, DeepSeek, and Mistral models are pulled by name from the same interface and switched per chat. Larger models answer better and need more memory, and a model that doesn't fit in the graphics card's memory spills over to the processor and slows to a crawl. When local hardware runs out, a model gateway such as OpenRouter or LiteLLM can be added to put hosted models in the same picker, next to the local ones.

On the machine it runs on, the interface is a local address in the browser, so nothing has to be published: no reverse proxy, no TLS certificate, no open port. Reaching it from a phone or sharing it with other people is optional and comes in steps. Tailscale can be added for private access from your own devices, Cloudflare Tunnel or a reverse proxy such as Caddy for a public address with HTTPS, and a self-hosted PaaS if you would rather deploy from a dashboard. If you have no suitable computer, the same setup runs on a rented GPU server, at a monthly price well above a chat subscription.

Key Features

  • ✓Ollama and Open WebUI from the official Compose file: two containers and no separate database to run
  • ✓A browser chat interface with history, file uploads, web search, voice, and a model picker
  • ✓Models pulled and updated by name from the admin panel: Google Gemma by default, or Qwen, Meta Llama, DeepSeek, or Mistral
  • ✓User accounts with roles and an approval queue, so one instance serves several people
  • ✓Chat with your own documents through the built-in vector store
  • ✓Works fully offline once the models are downloaded
  • ✓Hosted models can sit in the same picker through any OpenAI-compatible connection

When to Use Self-Hosted AI with Ollama and Open WebUI

  • →A private alternative to a hosted chatbot for notes, drafts, and questions you would rather not send to a provider
  • →Chatting with confidential documents that have to stay on your own hardware
  • →One shared AI assistant for a family, a classroom, or a small team
  • →Trying and comparing open-weight models side by side before building on one
  • →A local model endpoint that coding agents and other tools on your network can use

Pros

  • No subscription and no per-token cost: after the hardware, usage is free
  • Prompts, files, and chat history never leave your machine
  • Both tools are free to run, and models are swapped with a click
  • The same interface takes hosted models too, so local and cloud can be mixed in one place

Cons

  • Answer quality and speed are capped by your hardware: a model needs roughly its file size in graphics or unified memory
  • Local models trail the frontier hosted models on hard reasoning and long tasks
  • Open WebUI's licence is not open source in the usual sense: its branding has to stay unless you have 50 or fewer users or an enterprise licence
  • Ollama's default context window is small on graphics cards under 24 GB, so long documents are cut off until you raise it
  • Updates, backups, and access control are yours once other people use it

Database Options for Self-Hosted AI with Ollama and Open WebUI

SQLite

Self-Hosted AI with Ollama and Open WebUI with SQLite

Open WebUI's built-in store, and the default: accounts, chats, and settings live in one file inside the data volume, with nothing to install or tune. The project's docs call it fine for personal use, a home lab, and small teams, as long as the file sits on a local SSD and only one instance runs. Backing it up means copying the volume.

PostgreSQL

Self-Hosted AI with Ollama and Open WebUI with PostgreSQL

The step for a larger installation. Open WebUI switches to PostgreSQL through one connection setting, and its docs require it before running more than one instance or putting the data on network storage. It can hold the document embeddings as well, through pgvector, which the Open WebUI team maintains as a first-class vector store. Decide early, because Open WebUI does not move existing data from SQLite to PostgreSQL.

These are highlighted picks. To see all the tools, check the Databases category.

LLM Options for Self-Hosted AI with Ollama and Open WebUI

Google Gemma

Self-Hosted AI with Ollama and Open WebUI powered by Google Gemma

The default. Gemma 4 is Google's open-weight family, and on Ollama it comes in sizes that match home hardware: an edge model of about 7 GB for laptops, a 12B model of about 8 GB, and 26B and 31B models of 16 to 20 GB for a 24 GB graphics card. It reads images as well as text, supports tool calling and a thinking mode, and has a context window of up to 256K tokens. A sound first model for general chat and questions about documents.

Qwen

Self-Hosted AI with Ollama and Open WebUI powered by Qwen

Alibaba's Qwen family has the widest spread of sizes on Ollama, from about 1 GB for the smallest Qwen 3.5 model to very large ones, so there is a version for almost any machine. The 9B model is a 6.6 GB download, and the newer 27B and 35B models want a 24 GB card or more. Most releases are Apache 2.0, and the current ones support vision, tools, and thinking. A common pick for coding help and for multilingual use.

Meta Llama

Self-Hosted AI with Ollama and Open WebUI powered by Meta Llama

The best-known open-weight family, and still the most downloaded on Ollama. Llama 3.2 in its 3B size is a 2 GB download that runs on almost anything, which makes it a handy small model for quick tasks. Llama 3.3 at 70B is a 43 GB download for a workstation with two large graphics cards or a high-memory Mac, and the Llama 4 models are larger still. The weights come under Meta's own community licence.

DeepSeek

Self-Hosted AI with Ollama and Open WebUI powered by DeepSeek

DeepSeek-R1 is the reasoning model that made local thinking models popular, and it is among the most downloaded on Ollama. It writes out its reasoning before the answer, which helps with maths, logic, and step-by-step problems and makes replies slower. The small sizes are distilled versions: 8B is a 5.2 GB download and 14B is 9 GB, while the full model is far beyond home hardware. The weights are MIT licensed.

Mistral

Self-Hosted AI with Ollama and Open WebUI powered by Mistral

Open models from a European lab, under the Apache 2.0 licence. Mistral Small 3.2, a 24B model, is a 15 GB download that suits a 24 GB card and handles images and tool calls, and the Ministral 3 models start at 3 GB for modest hardware. A reasonable choice when you want a capable mid-sized model, or prefer a European vendor.

These are highlighted picks. To see all the tools, check the LLM category.

Self-Hosted AI with Ollama and Open WebUI Add-ons

Each addition below extends this stack with a capability the base stack works fine without. None are required: include the ones your product actually needs when building this stack, and skip the rest.

Vector Db Add-ons

Add a vector database when the agent needs to find documents, notes, or past conversations by meaning instead of exact keywords.

Qdrant

Self-Hosted AI with Ollama and Open WebUI with Qdrant

A dedicated vector database for document chat once the built-in store is not enough, run as one more container beside Open WebUI and selected with a single setting. It filters by metadata during the search and supports multitenancy, which matters when many users each have their own knowledge bases. Open WebUI's docs say the stores its own team maintains are the built-in one and pgvector on PostgreSQL, and that Qdrant support is community-maintained.

Pinecone

Self-Hosted AI with Ollama and Open WebUI with Pinecone

A fully managed vector database with nothing to run: Open WebUI connects with an API key and stores document embeddings in Pinecone's cloud. That removes a service from your machine, but the text of your documents then leaves it in chunks, which undoes the reason many people run this stack. The Starter plan is free. Like Qdrant, it is a community-maintained integration.

Hosting Add-ons

Add hosting when you want this to run somewhere other than its default home: off a platform's built-in hosting, or off your own computer and onto a rented server.

Hetzner

Self-Hosted AI with Ollama and Open WebUI with Hetzner

A dedicated GPU server for when no computer at home fits the job. The entry model has a 24 GB NVIDIA card and 64 GB of RAM for about €214 a month plus a one-time setup fee of about the same, with unlimited traffic. That runs the 26B to 31B models comfortably for a team. It is a fixed monthly cost whether or not anyone is chatting, so it pays off for shared, daily use. Docker and the NVIDIA container toolkit are yours to install.

DigitalOcean

Self-Hosted AI with Ollama and Open WebUI with DigitalOcean

GPU Droplets billed by the hour, which suits occasional use: start one, work, and destroy it. The smallest has a 20 GB NVIDIA card at $0.76 an hour, about $550 if left running for a month. Keep the Open WebUI data on a volume or a snapshot so chats and models survive between sessions. A regular small Droplet can also run Open WebUI alone and point it at hosted models.

Amazon Web Services

Self-Hosted AI with Ollama and Open WebUI with Amazon Web Services

For teams already on AWS. A g4dn.xlarge instance has a 16 GB NVIDIA T4 at about $0.53 an hour, and a g6.xlarge has a 24 GB L4 at about $0.80, both before storage. Stopping the instance when it is idle stops the compute charge while the disk keeps your models and chats. Open WebUI documents a deployment on ECS for the interface, and an Application Load Balancer can handle HTTPS.

Google Cloud Platform

Self-Hosted AI with Ollama and Open WebUI with Google Cloud Platform

The equivalent on Google Cloud: a g2-standard-4 machine with one NVIDIA L4 card costs about $0.71 an hour, or around $516 for a full month. Open WebUI also has a Cloud Run guide for running the interface as a serverless container in front of a model endpoint, which fits when the models are hosted elsewhere and only the chat layer is yours.

These are highlighted picks. To see all the tools, check the Hosting & Cloud category.

Model Aggregator Add-ons

Add a model aggregator when you want one API key and one bill for models from many providers, with automatic fallback when one of them is down, instead of setting up each provider separately.

OpenRouter

Self-Hosted AI with Ollama and Open WebUI with OpenRouter

Hosted models in the same picker as the local ones. Open WebUI connects to OpenRouter as an OpenAI-compatible provider with one API key, which brings in hundreds of models from many labs, billed at provider prices with a 5.5% fee on credit purchases. Useful for the questions a local model can't handle, and for comparing a local answer with a frontier one in the same chat. Those prompts do leave your machine.

LiteLLM

Self-Hosted AI with Ollama and Open WebUI with LiteLLM

A gateway you run yourself, as a third container. Open WebUI connects to it like any OpenAI-compatible endpoint, and LiteLLM holds the provider keys and applies budgets per user or per key, which a shared instance needs before a team is given paid models. It also puts local and hosted models behind one address for other tools. The open-source edition is free and uses PostgreSQL for its records.

These are highlighted picks. To see all the tools, check the AI Model Aggregators category.

Tunnel Add-ons

Add a tunnel when you're self-hosting without a static IP or can't open inbound ports — a home server, a VPS behind restrictive network policies, or anywhere a reverse proxy alone can't reach the internet.

Tailscale

Self-Hosted AI with Ollama and Open WebUI with Tailscale

The private route, and the easiest way to use a home setup from a phone. Tailscale gives the machine a stable name on your own encrypted network, and its serve command puts Open WebUI on an HTTPS address with a valid certificate that only your devices can open. HTTPS matters here because browsers refuse microphone access without it, so voice input depends on it. Open WebUI can also sign users in from their Tailscale identity. The Personal plan is free.

Cloudflare Tunnel

Self-Hosted AI with Ollama and Open WebUI with Cloudflare Tunnel

What Open WebUI's docs recommend for a public address without opening a port: cloudflared connects outbound from your machine, and Cloudflare serves the interface on your domain with HTTPS. It works behind a home router. Put Cloudflare Access in front when only named people should reach the login page, and keep Ollama's own port off the tunnel. Free, for a domain managed by Cloudflare.

Reverse Proxy Add-ons

Add a reverse proxy when the service should be reachable at its own web address: it terminates HTTPS on your domain and can put a login in front of an interface that is otherwise kept private.

Caddy

Self-Hosted AI with Ollama and Open WebUI with Caddy

Automatic HTTPS from a few lines of config, and the proxy Open WebUI's docs suggest when you want certificates handled for you. It streams answers and passes WebSockets without extra settings. Set Open WebUI's public URL and allowed origin to the HTTPS address, or real-time features fail quietly. It needs a domain pointing at a machine the internet can reach, which a home connection often is not.

NGINX

Self-Hosted AI with Ollama and Open WebUI with NGINX

The proxy for full control, with a detailed Open WebUI guide. Three settings matter: turn proxy buffering off, or streamed answers arrive with broken formatting; pass the WebSocket upgrade headers; and raise the read timeout to several minutes, because a local model can take that long to answer. Certificates come from Let's Encrypt through Certbot or NGINX Proxy Manager.

Traefik

Self-Hosted AI with Ollama and Open WebUI with Traefik

Routes to the Open WebUI container from Docker labels and renews certificates itself, convenient on a server that already runs other containers. Open WebUI has no Traefik guide of its own, and the same rules apply as for any proxy: WebSocket support, no response buffering, long timeouts, and the public URL set in Open WebUI. It is the proxy Coolify and Dokploy use underneath.

Self Hosted Paas Add-ons

Add a self-hosted PaaS when you would rather deploy and update from a dashboard than from the command line: it installs the service from a template and brings its own reverse proxy and certificates.

Coolify

Self-Hosted AI with Ollama and Open WebUI with Coolify

A free, self-hosted deploy dashboard with a ready-made Ollama with Open WebUI service: both official images, two volumes, and the interface on a public domain with HTTPS through Coolify's proxy, while Ollama's port stays internal. The only login is Open WebUI's own, and the first person to sign up becomes the administrator, so create that account the moment it is deployed. The template sets no secret key, so add one or every redeploy signs everyone out, and it runs the rolling main image, so pin a version for shared use. A graphics card has to be passed to the Ollama container by hand.

Dokploy

Self-Hosted AI with Ollama and Open WebUI with Dokploy

Free and self-hosted, with an Open WebUI template that runs the official image on a public domain behind Dokploy's bundled Traefik and generates the secret key that keeps logins valid across restarts. Ollama is in the template as a commented-out service, with its graphics card settings, next to web browsing and image generation services that are off by default too. As on Coolify, the login is Open WebUI's own: the first account created becomes the administrator, so claim it at once. No separate reverse proxy is added next to it.

Frequently Asked Questions about Self-Hosted AI with Ollama and Open WebUI

Why run models yourself instead of using a hosted chatbot or Ollama's cloud?

Three reasons: privacy, cost shape, and independence. Nothing you type or upload is sent to a provider, the bill is hardware you already own plus electricity and not a monthly fee per person, and it works offline and keeps working whatever a vendor changes. What you give up is the top of the quality range: the frontier hosted models are stronger than anything a home machine runs, and faster. There is a middle path. Ollama's own cloud runs larger open models through the same tools, on a free tier and a Pro plan at $20 a month, and Open WebUI can show those next to your local models. Many people keep private work local and send the hard questions out. Open WebUI itself has no hosted version from its makers, only an enterprise licence for large organisations.

What hardware do I need?

It depends on the model, and neither project publishes a minimum. The working rule is that a model needs about its download size in graphics memory, or in unified memory on a Mac, plus headroom for the conversation. Gemma's 12B model wants about 8 GB of memory, so a 12 GB graphics card or a 16 GB Mac. The 26B to 31B models want a 24 GB card or a 32 GB Mac. With no graphics card it still runs on the processor, slowly, and small 2 to 3 GB models are the practical limit. Open WebUI is light next to the model. On a Mac, Ollama's Docker instructions cover only NVIDIA and AMD cards, so the usual setup is the Ollama app running natively, where it uses the Mac's graphics chip, with Open WebUI in Docker finding it automatically.

What needs backing up, and how do updates work?

Open WebUI's data volume is the one that matters: the database with accounts and chats, uploaded files, and the built-in document index. A dated archive of that volume is the backup the docs describe. The Ollama volume holds only model files, which can be downloaded again. An update is an image swap for each container. Open WebUI's docs suggest the rolling tag for personal use and a pinned version once other people depend on the instance, and warn that database migrations run one way: after a bad update, going back means restoring a backup taken before it. One setting to fix early is the secret key. Without a fixed one, every recreated container signs everybody out.

How do I use it from my phone or share it with other people?

In steps, from private to public. On your home network, other devices can open the interface at the computer's local address. For access from anywhere that only you have, add Tailscale: your devices join a private network and the interface gets an HTTPS address nobody else can reach. For family or a team without Tailscale on every device, Cloudflare Tunnel or a reverse proxy such as Caddy gives it a public address. Then accounts matter: the first user is the administrator, new sign-ups wait in a pending queue, and Open WebUI's hardening guide is worth reading before the address is public. Keep Ollama's own port private in every case, since it has no login. Past 50 users, Open WebUI's licence requires its branding to stay in place.

Can a coding agent use these models?

Yes. Ollama serves an API on the same machine that coding agents can call, and its launch command starts Claude Code, Codex, OpenCode, and others against a local or Ollama cloud model with no manual configuration. Two limits apply. Coding agents need a large context window, and Ollama's docs recommend at least 64,000 tokens for them, far above its default on cards under 24 GB, so the setting has to be raised and the memory has to be there. And an open model on one graphics card is noticeably weaker at long, multi-file tasks than the hosted models those agents were built for. Open WebUI is not involved: the agent talks to Ollama directly. The Local LLM Dev Stack covers this setup, with the runtime, the agent, and the models that can code, and Claude Code, Codex, and OpenCode each have a developer stack of their own for the rest of the workflow.

Scores

Popularity5/5

Ollama is the most-used local model runtime and Open WebUI the most-used self-hosted chat interface, at around 180,000 and 150,000 GitHub stars, and the pairing is the standard starting point for running models at home.

Learning Curve3/5

One Compose file and a browser get a first chat running, and everyday use is a chat window. Getting it to run well takes more: graphics drivers and the container toolkit, matching model size to memory, raising the context length, and securing it once others have access.

Flexibility4/5

Any model in Ollama's library, hosted models through OpenAI-compatible connections, a plugin system, several vector stores, and a choice of database. The interface and the runtime themselves are fixed, and Open WebUI's licence limits rebranding.

Performance3/5

Speed and quality are set by the hardware and the model: a mid-sized model on a 24 GB card answers at a comfortable pace, a model that spills onto the processor does not. Ollama serves a handful of users well and is not built for high-throughput serving.

Portability4/5

Models are standard open-weight files that other runtimes can load, chats sit in SQLite or PostgreSQL, and Open WebUI speaks to any OpenAI-compatible backend, so either half can be replaced. Moving chat history to a different interface is manual.

Tools in the Self-Hosted AI with Ollama and Open WebUI Stack

DevOps & CI/CD

AI Infrastructure

Database (choose one)

LLM (choose one)

Add-ons (optional — add any, or none)

Vector Db

Hosting

Model Aggregator

Tunnel

Reverse Proxy

Self Hosted Paas

Self-Hosted AI with Ollama and Open WebUI Pricing

Free on hardware you own

Ollama and Open WebUI are free to run, and so are the models, so on a computer you already own the cost is electricity. The real spend is hardware when yours is too small: a graphics card with 24 GB of memory is the usual target. Renting is the expensive route, from about €214 a month for a dedicated GPU server or under a dollar an hour for a cloud GPU you switch off between sessions. Hosted models through Ollama's cloud or a model gateway are optional and billed by usage.

Ollama and Open WebUIFree

Ollama is MIT-licensed. Open WebUI is free under its own licence, which requires its branding to stay above 50 users unless an enterprise licence is held.

ModelsFree (open weights)

Downloaded once from Ollama's library, 2 to 20 GB each. Licences vary by model family.

Hardware (your own computer)Owned, plus electricity

About 8 GB of graphics or unified memory runs Gemma's 12B model, and 24 GB runs the 26B to 31B models.

Rented GPU server (optional)From ~€214/mo, or ~$0.50–0.80/hour

A dedicated server with a 24 GB card at a fixed monthly price, or a cloud GPU with 16 to 24 GB billed by the hour.

Remote access (optional)Free

Tailscale's Personal plan and Cloudflare Tunnel are free, and the reverse proxies and self-hosted deploy platforms are free, open-source software.

Self-Hosted AI with Ollama and Open WebUI System Requirements

RAM
About 8 GB of graphics or unified memory for a 12B model; 24 GB for a 26B to 31B model
GPU
NVIDIA, AMD, or Apple Silicon; a processor alone works for small models, slowly
Disk
About 2 GB for the Open WebUI image, plus 2 to 20 GB per model
OS
Linux, macOS, or Windows with Docker

No official requirements published — tekyous guidance based on the bundle's services.

Neither Ollama nor Open WebUI publishes minimum requirements, so these are estimates. They follow from the model sizes in Ollama's library (the default 12B model is about an 8 GB download, the 26B and 31B models 16 to 20 GB) and the rule that a model needs roughly its file size in graphics or unified memory, plus room for the context window. Ollama's default context is 4k tokens below 24 GiB of VRAM and 32k from 24 to 48 GiB. Open WebUI's standard image is a 1.66 GB download.