[{"data":1,"prerenderedAt":894},["ShallowReactive",2],{"categories-init":3,"stack-vllm-self-hosted":4,"stack-res-vllm-self-hosted":893},true,{"stack_id":5,"slug":6,"name":7,"tagline":8,"long_description":9,"key_features":10,"use_cases":19,"pros":25,"cons":31,"cover_image_url":37,"scores":38,"options":52,"additions":487,"option_groups":668,"multi_select_option_types":669,"tools_by_category":670,"related_stacks":761,"faqs":837,"pricing":853,"system_requirements":872,"experience_level":887,"project_type":888,"stack_type_slug":769,"stack_type_icon_url":770,"published_date":521,"last_updated_date":521,"seo_meta":889},246,"vllm-self-hosted","vLLM Inference Server","A self-hosted, OpenAI-compatible API for open-weight models, served from your own GPU server.","vLLM is an inference engine: it loads the weights of an open model onto GPUs and serves them over an OpenAI-compatible HTTP API, so any application written for a hosted provider can point at your server instead. Its speed comes from **continuous batching and paged attention memory**, which let many requests share one GPU without waiting for each other, and it runs more than 200 model architectures, including quantized formats such as FP8, INT4, and AWQ. The reason to run it yourself is steady, shared load: a GPU server costs the same whether it answers ten requests or ten thousand, so at volume a fixed price beats per-token billing, and prompts and fine-tuned weights stay on hardware you control.\n\nThe official Docker image runs **one container**: vllm\u002Fvllm-openai, started with the NVIDIA runtime, the host's shared memory, a mounted model cache, and a model name, and listening on port 8000. There is no database and no queue; the container is stateless apart from the model files it downloads once and caches. One server instance serves one model, so a second model means a second instance behind a router. The hardware is the real requirement: Linux and an NVIDIA card with compute capability 7.5 or higher, such as a T4, L4, A100, or H100, with AMD and Intel GPUs also supported, and enough card memory to hold the weights plus the context cache.\n\nThe server's `--api-key` flag protects only the \u002Fv1 family of paths, and other routes, including an \u002Finvocations endpoint that accepts the same inference requests, stay open to anyone who can reach the port. The project's own security guidance is therefore **to put a reverse proxy in front that allows only the endpoints you mean to serve**, to keep any multi-node traffic on an isolated network, and to expose nothing but the API port. A tunnel is the alternative when the clients run somewhere you do not control.\n\n**Scaling past one card is configuration.** Tensor parallelism splits one model across several GPUs in a server, and the project's Kubernetes production stack adds a router, cache-aware routing, and Helm deployment for a cluster. Teams usually add a gateway in front for keys and budgets, and a chat interface such as Open WebUI when people, not only applications, will talk to the model; both are optional.",[11,12,13,14,15,16,17,18],"vLLM from the official vllm\u002Fvllm-openai Docker image: one stateless container with no database","An OpenAI-compatible API on port 8000 for chat, completions, and embeddings","Continuous batching and paged attention memory, so many users share one GPU","More than 200 model architectures, with FP8, INT4, GPTQ, and AWQ quantization","Tensor, pipeline, data, and expert parallelism for models that need several GPUs","NVIDIA, AMD, and Intel GPU support, with Docker images for CUDA, ROCm, and Intel XPU","A Kubernetes production stack with Helm, cache-aware routing, and tracing","Apache-2.0 licensed, from the UC Berkeley Sky Computing Lab",[20,21,22,23,24],"A private, shared model endpoint for an engineering team's agents and coding tools","Replacing a per-token API bill with a fixed GPU server once volume is steady","Serving a fine-tuned or custom open-weight model that no hosted provider carries","Keeping prompts and documents on hardware you control for compliance reasons","The back end for a gateway or a chat interface that many people use at once",[26,27,28,29,30],"Apache 2.0 licensed engine with batching built for many users sharing a GPU","Applications keep an OpenAI-style client, so moving between your server and a hosted provider is a base URL change","Runs any open-weight model with a supported architecture, including your own fine-tunes","Scales from one card to several GPUs to a Kubernetes cluster without changing engine","A fixed monthly cost at steady volume, instead of billing that grows with every token",[32,33,34,35,36],"A GPU server is a fixed cost whether or not anyone is calling it, and sits far above the price of a small VPS","The --api-key flag covers only part of the API, so any exposure beyond a trusted network needs a reverse proxy with an allowlist","One server instance serves one model, so several models mean several instances and a router","Memory planning is yours: the weights plus the context cache have to fit the card","The host needs NVIDIA drivers and the container toolkit before the first container starts",null,{"popularity":39,"learning_curve":42,"flexibility":45,"performance":48,"portability":50},{"score":40,"reasoning":41},4,"The most common engine for serving open-weight models in production, with a very large GitHub following and a contributor base across many companies and universities. It is a tool for ML and platform engineers more than a household name.",{"score":43,"reasoning":44},2,"The API is the OpenAI one, but running it is GPU operations: drivers and the container toolkit, matching a model to card memory, tuning context length and parallelism, and a reverse proxy to cover the routes the API key leaves open. The first request is easy; a server that stays healthy under load is the work.",{"score":46,"reasoning":47},5,"Hundreds of model architectures, quantization formats, and several kinds of parallelism, on NVIDIA, AMD, and Intel hardware, from one card to a Kubernetes cluster. The launch command is the whole configuration, and the same endpoint serves applications, gateways, and chat interfaces.",{"score":46,"reasoning":49},"Continuous batching and paged attention are the reason people choose vLLM: many concurrent requests share a GPU with high throughput instead of queuing one at a time. The ceiling is the hardware, and one instance serves one model.",{"score":40,"reasoning":51},"Apache 2.0 and the OpenAI request format on the front, so applications move between vLLM and a hosted provider by changing a base URL. The engine is tied to specific GPU and driver stacks, and models are tuned to the card they run on, which makes moving hardware a re-test.",{"agent_framework":53,"vector_db":57,"database":61,"orm":65,"authentication":69,"analytics":73,"model_inference":77,"coding_agent":81,"llm":85,"language":242,"frontend_framework":246,"cms":250,"hosting":254,"reverse_proxy":397,"self_hosted_paas":483},{"name":37,"tools":54,"descriptions":55,"aliases":56,"see_all":37},[],{},{},{"name":37,"tools":58,"descriptions":59,"aliases":60,"see_all":37},[],{},{},{"name":37,"tools":62,"descriptions":63,"aliases":64,"see_all":37},[],{},{},{"name":37,"tools":66,"descriptions":67,"aliases":68,"see_all":37},[],{},{},{"name":37,"tools":70,"descriptions":71,"aliases":72,"see_all":37},[],{},{},{"name":37,"tools":74,"descriptions":75,"aliases":76,"see_all":37},[],{},{},{"name":37,"tools":78,"descriptions":79,"aliases":80,"see_all":37},[],{},{},{"name":37,"tools":82,"descriptions":83,"aliases":84,"see_all":37},[],{},{},{"name":86,"tools":87,"descriptions":232,"aliases":238,"see_all":239},"LLM",[88,136,162,186,209],{"tool_id":89,"name":90,"slug":91,"tooltip_description":92,"logo_url":93,"logo_bg":94,"pricing_model":95,"learning_curve_score":43,"popularity_score":40,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":100,"subcategory":103,"categories":107,"subcategories":110,"flexibility_score":46,"performance_score":46,"portability_score":46,"is_featured":112,"tags":113,"score_reasonings":128,"published_date":134,"last_updated_date":135},162,"Qwen","qwen","Alibaba's LLM family, from sub-1B models to a 2.4-trillion-parameter flagship, covering reasoning, coding, vision, and long context, mostly under Apache 2.0 open weights with a hosted API on Alibaba Cloud.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fqwen.svg","dark",{"slug":96,"display_name":97,"description":98},"freemium","Freemium","A free tier is available; additional features, usage limits, or managed hosting require a paid plan.","open",{"category_id":101,"name":86,"slug":102},19,"llm",{"subcategory_id":104,"name":105,"slug":106},52,"Open-weight","open-weight",[108],{"category_id":101,"name":86,"slug":102,"is_primary":3,"display_order":109},0,[111],{"subcategory_id":104,"name":105,"slug":106,"category_id":101,"is_primary":3,"display_order":109},false,[114,119,124],{"tag_id":115,"name":116,"slug":117,"tag_type":118},11,"Open Source","open-source","feature",{"tag_id":120,"name":121,"slug":122,"tag_type":123},40,"Web","web","platform",{"tag_id":125,"name":126,"slug":127,"tag_type":118},12,"Self-hostable","self-hostable",{"learning_curve":129,"flexibility":130,"performance":131,"popularity":132,"portability":133},"Open weights + OpenAI-compatible API means onboarding is straightforward for any developer familiar with the OpenAI SDK. Ollama makes local inference a one-command setup. Navigating the large model family and choosing the right variant adds some complexity, but the barrier is low overall.","Apache 2.0 licence, sizes from 0.6B to 480B+, dense and MoE architectures, specialised coding and vision variants, toggleable reasoning mode, and multiple access paths (self-hosted, DashScope, third-party APIs). Maximum flexibility for any use case.","Qwen3-Coder ranks among the very best open-weight coding models; Qwen3 MoE variants compete with GPT-4-class models on benchmarks; QwQ-32B rivals leading reasoning specialists; Qwen-VL leads vision-language open-weight benchmarks. Across all domains, Qwen is at or near the open-weight frontier.","Qwen3 repo has ~27K GitHub stars; Qwen3-Coder ~16K; QwenLM org collectively has tens of thousands of stars across repositories. Dominant in Chinese AI ecosystem and rapidly growing global adoption. Strong presence on Hugging Face, Ollama library, and all major inference providers.","Apache 2.0 weights on Hugging Face run on Ollama, vLLM, llama.cpp, LM Studio, and MLX with no licence restrictions. Same models available via DashScope and third-party APIs. Maximum portability — switch providers or go self-hosted any time.","2026-05-29","2026-09-27",{"tool_id":137,"name":138,"slug":139,"tooltip_description":140,"logo_url":141,"logo_bg":94,"pricing_model":142,"learning_curve_score":145,"popularity_score":46,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":146,"subcategory":147,"categories":148,"subcategories":150,"flexibility_score":46,"performance_score":40,"portability_score":46,"is_featured":112,"tags":152,"score_reasonings":156,"published_date":134,"last_updated_date":135},154,"Meta Llama","meta-llama","Meta's open-weight LLM family, from compact 1B edge models to the Llama 4 mixture-of-experts models Scout and Maverick; download the weights and self-host, or call them through dozens of managed API providers.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fmeta-llama.svg",{"slug":143,"display_name":116,"description":144},"open_source","Source code is publicly available and free to use, modify, and distribute. No paid plans from the project itself.",3,{"category_id":101,"name":86,"slug":102},{"subcategory_id":104,"name":105,"slug":106},[149],{"category_id":101,"name":86,"slug":102,"is_primary":3,"display_order":109},[151],{"subcategory_id":104,"name":105,"slug":106,"category_id":101,"is_primary":3,"display_order":109},[153,154,155],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":157,"flexibility":158,"performance":159,"popularity":160,"portability":161},"Using Llama via a managed API (Groq, Together AI, Bedrock) is as straightforward as any REST API. Self-hosting is significantly harder: choosing the right quantisation, setting up vLLM or Ollama, managing GPU memory, and configuring serving all require infrastructure expertise.","Unmatched flexibility: download weights, fine-tune on proprietary data, deploy on any hardware (from a MacBook with Ollama to a multi-GPU cluster), integrate with any framework, or call from any of dozens of API providers. No vendor controls your access or pricing.","Llama 3.1 405B and Llama 4 Maverick are competitive with frontier proprietary models on most benchmarks. Llama 4 Scout and 8B\u002F70B variants punch above their weight for their parameter counts. However, the very top of the performance leaderboard is still held by closed models.","The most popular open-weight LLM family by a wide margin — backbone of the open-source AI ecosystem, most downloaded model family on Hugging Face, ~59K GitHub stars, and the default choice for self-hosted AI deployments.","Maximum portability: run locally on CPU with llama.cpp, on any GPU cloud, via any of 10+ managed API providers, or even in-browser with WebGPU. No vendor lock-in whatsoever.",{"tool_id":163,"name":164,"slug":165,"tooltip_description":166,"logo_url":167,"logo_bg":168,"pricing_model":169,"learning_curve_score":43,"popularity_score":40,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":170,"subcategory":171,"categories":172,"subcategories":174,"flexibility_score":46,"performance_score":40,"portability_score":46,"is_featured":3,"tags":176,"score_reasonings":180,"published_date":134,"last_updated_date":135},163,"Mistral","mistral","Europe's leading open-weight LLM family, from small edge models to a frontier MoE flagship, available through Mistral's API, its Vibe assistant, the major clouds, or self-hosted.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fmistral.svg","white",{"slug":96,"display_name":97,"description":98},{"category_id":101,"name":86,"slug":102},{"subcategory_id":104,"name":105,"slug":106},[173],{"category_id":101,"name":86,"slug":102,"is_primary":3,"display_order":109},[175],{"subcategory_id":104,"name":105,"slug":106,"category_id":101,"is_primary":3,"display_order":109},[177,178,179],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":181,"flexibility":182,"performance":183,"popularity":184,"portability":185},"The la Plateforme API is OpenAI-compatible, so any developer already familiar with the OpenAI Python SDK can switch with a one-line URL change. Running open-weight models locally via Ollama is also beginner-friendly. The main complexity is the breadth of model choices and understanding trade-offs between MoE and dense architectures.","Unmatched spectrum — from 3B edge models to 675B MoE frontier models; open-weight for full fine-tuning and self-hosting; API for managed access; third-party cloud for enterprise compliance. Codestral covers code, Mathstral covers math, Pixtral covers vision. Configurable reasoning effort in Small 4. Few LLM families offer this range.","Mixtral MoE models punch well above their weight in inference efficiency. Mistral Large 3 is competitive with frontier models on benchmarks. However, on raw capability benchmarks the top Mistral models are generally tier-2 behind GPT-4o and Claude Opus. The efficiency advantage is real and often decisive for cost-sensitive deployments.","Mistral 7B and Mixtral 8x7B were among the most downloaded open-weight models of 2023-2024. The lab is widely recognised as Europe's premier LLM provider, has $14B valuation, and models are available on every major cloud. Strong in enterprise and research circles.","Open-weight models are maximally portable — download once, run anywhere, no ongoing API dependency. Apache 2.0 permits commercial use and modification. Proprietary models (Mistral Large, Pixtral Large) carry some lock-in, but the API is OpenAI-compatible so switching costs are low.",{"tool_id":187,"name":188,"slug":189,"tooltip_description":190,"logo_url":191,"logo_bg":94,"pricing_model":192,"learning_curve_score":145,"popularity_score":46,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":193,"subcategory":194,"categories":195,"subcategories":197,"flexibility_score":46,"performance_score":46,"portability_score":46,"is_featured":112,"tags":199,"score_reasonings":203,"published_date":134,"last_updated_date":135},161,"DeepSeek","deepseek","Chinese open-weight LLM family under the MIT license, led by DeepSeek-V4.1-Flash, offering frontier-level results at very low API prices, with full self-hosting through vLLM, SGLang, and Ollama.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fdeepseek.svg",{"slug":96,"display_name":97,"description":98},{"category_id":101,"name":86,"slug":102},{"subcategory_id":104,"name":105,"slug":106},[196],{"category_id":101,"name":86,"slug":102,"is_primary":3,"display_order":109},[198],{"subcategory_id":104,"name":105,"slug":106,"category_id":101,"is_primary":3,"display_order":109},[200,201,202],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":204,"popularity":205,"portability":206,"flexibility":207,"performance":208},"The DeepSeek API (OpenAI-compatible) is trivial to call for developers already using OpenAI. Running distilled models via Ollama is also beginner-friendly. The steep part is self-hosting the full 671B V3\u002FR1 models — that requires multi-GPU infrastructure, tensor parallelism configuration in vLLM, and knowledge of quantisation trade-offs.","One of the fastest open-source repository growth stories in GitHub history. DeepSeek-R1 reached ~92K stars within weeks; total deepseek-ai org stars exceeded 170K by end of 2025. Massive developer adoption, significant industry impact, and strong Hugging Face download counts.","MIT license, Hugging Face weights, Ollama\u002FvLLM\u002FLM Studio support, and availability on Amazon Bedrock, Fireworks, Together AI, and OpenRouter. No lock-in whatsoever.","Open MIT weights with no use restrictions, publishable on any infrastructure, fine-tuneable with standard tools, and accessible via multiple APIs (direct, Bedrock, Together AI, Fireworks). Maximum possible flexibility for an LLM.","DeepSeek-V3 and R1 match or exceed GPT-4o on coding (HumanEval, SWE-bench), math (AIME, MATH-500), and reasoning benchmarks. R1 matches OpenAI o1. V4-Pro (2026) rivals the world's top closed models. Exceptional especially on coding tasks.",{"tool_id":210,"name":211,"slug":212,"tooltip_description":213,"logo_url":214,"logo_bg":94,"pricing_model":215,"learning_curve_score":43,"popularity_score":40,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":216,"subcategory":217,"categories":218,"subcategories":220,"flexibility_score":46,"performance_score":40,"portability_score":46,"is_featured":112,"tags":222,"score_reasonings":226,"published_date":134,"last_updated_date":135},158,"Google Gemma","google-gemma","Google DeepMind's family of open-weight LLMs, from compact edge models to a 31B dense flagship and a 26B MoE, built for self-hosting, on-device use, and fine-tuning under Apache 2.0.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fgoogle-gemma.svg",{"slug":143,"display_name":116,"description":144},{"category_id":101,"name":86,"slug":102},{"subcategory_id":104,"name":105,"slug":106},[219],{"category_id":101,"name":86,"slug":102,"is_primary":3,"display_order":109},[221],{"subcategory_id":104,"name":105,"slug":106,"category_id":101,"is_primary":3,"display_order":109},[223,224,225],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":227,"flexibility":228,"performance":229,"popularity":230,"portability":231},"Running Gemma locally is accessible to any developer: `ollama run gemma3` requires no account, no API key, and no GPU on smaller models. The Hugging Face Transformers integration is industry-standard. Fine-tuning via LoRA or full fine-tuning adds moderate complexity but is well-documented.","Multiple sizes (1B to 31B), multiple architectures (dense and MoE), multimodal variants (PaliGemma), code-specialised variants (CodeGemma), domain fine-tunes (MedGemma), and embedding models (EmbeddingGemma). Deployable on any hardware — phones, laptops, consumer GPUs, cloud VMs. Runs with Ollama, vLLM, llama.cpp, JAX, PyTorch, Keras, and Transformers. Apache 2.0 (Gemma 4) allows full fine-tuning and redistribution.","Gemma models consistently top their weight class on standard benchmarks. Gemma 2 27B rivalled 70B-class models at release; Gemma 3 27B set new small-model records on MMLU and reasoning. Gemma 4 31B reaches 85.2% MMLU Pro and 89.2% AIME 2026. However, top-of-market reasoning (GPT-5, Claude Opus, Gemini 3.1 Pro) still exceeds the 31B ceiling.","The google-deepmind\u002Fgemma GitHub repository has ~5.1K stars; Gemma models on Hugging Face have millions of downloads. Widely used in on-device AI, edge deployment, and fine-tuning research. Strong adoption in the open-source ML community, though Meta Llama commands a larger absolute user base and mindshare.","Maximum portability — Gemma models run on CPUs, consumer GPUs, server GPUs, TPUs, mobile phones (E2B\u002FE4B variants), and Raspberry Pi-class hardware. GGUF quantisation via llama.cpp enables deployment anywhere. No cloud lock-in; weights are self-contained.",{"qwen":233,"meta-llama":234,"mistral":235,"deepseek":236,"google-gemma":237},"The default, and the model in vLLM's own Docker quickstart. Qwen comes in the widest spread of sizes, from under 1B to very large mixture-of-experts models, so there is a version for almost any card. In 16-bit form the weights take about two bytes per parameter, so an 8B model needs roughly 16 GB of card memory and a 32B model about 64 GB, before the context cache. Most releases are Apache 2.0 and support tools and long contexts. A common pick for coding and for multilingual use.","The best-known open-weight family. At 16 bits an 8B model needs roughly 16 GB of card memory and a 70B model about 140 GB, which means several GPUs working together through tensor parallelism or a quantized copy on one large card. The weights are gated on Hugging Face: accept Meta's licence there and pass an access token to the container, which the quickstart's HF_TOKEN variable is for. The licence is Meta's own community licence, not Apache 2.0.","Open models from a European lab, many under the Apache 2.0 licence. Mistral Small, a 24B model, needs roughly 48 GB for the weights alone at 16 bits, which leaves no room for the context cache on a 48 GB card, so use an 80 GB card, an 8-bit copy on a 48 GB one, or a 4-bit copy on a 24 GB one, and it handles images and tool calls. A reasonable choice when you want a capable mid-sized model or prefer a European vendor.","DeepSeek's reasoning and general models are among the most capable open weights, and the weights are MIT licensed. The full-size models are very large mixture-of-experts networks that need a multi-GPU node with tensor or expert parallelism, which is a data-centre setup and far from one card. The distilled versions, 8B and 14B, fit a single card and write out their reasoning before the answer, which helps with maths and logic and makes replies longer.","Google's open-weight family, with sizes that suit a single card and support for images as well as text. A 12B model needs roughly 24 GB of card memory at 16 bits, which leaves no room for the context cache on a 24 GB card, so it wants a 48 GB card or an 8-bit copy, and the larger sizes want quantization or a bigger card. Read the licence terms of the exact release before building on it, since Google has changed them between generations.",{},{"kind":240,"slug":102,"name":86,"href":241},"category","\u002Ftools\u002Fcategories\u002Fllm",{"name":37,"tools":243,"descriptions":244,"aliases":245,"see_all":37},[],{},{},{"name":37,"tools":247,"descriptions":248,"aliases":249,"see_all":37},[],{},{},{"name":37,"tools":251,"descriptions":252,"aliases":253,"see_all":37},[],{},{},{"name":255,"tools":256,"descriptions":389,"aliases":394,"see_all":395},"Hosting",[257,289,332,364],{"tool_id":104,"name":258,"slug":259,"tooltip_description":260,"logo_url":261,"logo_bg":168,"pricing_model":262,"learning_curve_score":43,"popularity_score":145,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":266,"subcategory":269,"categories":273,"subcategories":275,"flexibility_score":145,"performance_score":40,"portability_score":40,"is_featured":112,"tags":277,"score_reasonings":283,"published_date":134,"last_updated_date":135},"Hetzner","hetzner","German cloud and dedicated server provider offering VPS, bare-metal servers, and storage at prices far below the big clouds, with data centers in the EU, the US, and Singapore.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fhetzner.svg",{"slug":263,"display_name":264,"description":265},"paid","Paid","No meaningful free tier — a subscription or one-time purchase is required to use the tool.",{"category_id":46,"name":267,"slug":268},"Hosting & Cloud","hosting-cloud",{"subcategory_id":270,"name":271,"slug":272},42,"VPS & Servers","vps-and-servers",[274],{"category_id":46,"name":267,"slug":268,"is_primary":3,"display_order":109},[276],{"subcategory_id":270,"name":271,"slug":272,"category_id":46,"is_primary":3,"display_order":109},[278,279],{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":280,"name":281,"slug":282,"tag_type":118},21,"Multi-region","multi-region",{"learning_curve":284,"flexibility":285,"performance":286,"portability":287,"popularity":288},"Standard VPS setup; any developer comfortable with Linux can deploy in an afternoon.","Standard VPS; configure anything on Linux; no managed abstraction layer constraining choices.","Competitive hardware at low cost; dedicated and VPS servers deliver strong baseline performance.","Standard Linux VPS; highly portable with no managed-service abstraction to escape.","Popular among European developers and cost-conscious teams; strong reputation for value.",{"tool_id":290,"name":291,"slug":292,"tooltip_description":293,"logo_url":294,"logo_bg":94,"pricing_model":295,"learning_curve_score":40,"popularity_score":46,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":299,"subcategory":300,"categories":303,"subcategories":305,"flexibility_score":46,"performance_score":46,"portability_score":145,"is_featured":112,"tags":307,"score_reasonings":326,"published_date":134,"last_updated_date":135},50,"Amazon Web Services","aws","World's largest cloud platform with 200+ services. Market leader in infrastructure as a service (IaaS) and platform as a service (PaaS).","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Faws.png",{"slug":296,"display_name":297,"description":298},"usage_based","Usage-Based","Pricing scales with consumption: API calls, data volume, compute time, or similar metered units.",{"category_id":46,"name":267,"slug":268},{"subcategory_id":120,"name":301,"slug":302},"Cloud Providers","cloud-providers",[304],{"category_id":46,"name":267,"slug":268,"is_primary":3,"display_order":109},[306],{"subcategory_id":120,"name":301,"slug":302,"category_id":46,"is_primary":3,"display_order":109},[308,312,316,317,321],{"tag_id":309,"name":310,"slug":311,"tag_type":118},14,"Serverless","serverless",{"tag_id":313,"name":314,"slug":315,"tag_type":118},20,"Auto-scaling","auto-scaling",{"tag_id":280,"name":281,"slug":282,"tag_type":118},{"tag_id":318,"name":319,"slug":320,"tag_type":118},24,"Docker Compatible","docker-compatible",{"tag_id":322,"name":323,"slug":324,"tag_type":325},36,"CI\u002FCD","ci-cd","use_case",{"learning_curve":327,"flexibility":328,"performance":329,"portability":330,"popularity":331},"Vast service catalog; IAM, VPC, and networking alone can take months to master.","Hundreds of composable services with full IaC support; virtually unlimited architectural freedom.","Global infrastructure; EC2, Lambda, and CloudFront deliver performance at any scale.","Broad API surface and proprietary services create meaningful lock-in; migration is possible but costly.","Market-leading cloud provider; used by the majority of production deployments globally.",{"tool_id":333,"name":334,"slug":335,"tooltip_description":336,"logo_url":337,"logo_bg":94,"pricing_model":338,"learning_curve_score":145,"popularity_score":145,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":339,"subcategory":340,"categories":341,"subcategories":343,"flexibility_score":46,"performance_score":46,"portability_score":145,"is_featured":112,"tags":345,"score_reasonings":358,"published_date":134,"last_updated_date":135},51,"Google Cloud Platform","gcp","Google's cloud platform with strong data, analytics, AI, and Kubernetes capabilities. The third-largest cloud provider, after AWS and Azure.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fgcp.svg",{"slug":296,"display_name":297,"description":298},{"category_id":46,"name":267,"slug":268},{"subcategory_id":120,"name":301,"slug":302},[342],{"category_id":46,"name":267,"slug":268,"is_primary":3,"display_order":109},[344],{"subcategory_id":120,"name":301,"slug":302,"category_id":46,"is_primary":3,"display_order":109},[346,350,351,352,353,354],{"tag_id":347,"name":348,"slug":349,"tag_type":118},13,"Free Tier","free-tier",{"tag_id":309,"name":310,"slug":311,"tag_type":118},{"tag_id":313,"name":314,"slug":315,"tag_type":118},{"tag_id":280,"name":281,"slug":282,"tag_type":118},{"tag_id":318,"name":319,"slug":320,"tag_type":118},{"tag_id":355,"name":356,"slug":357,"tag_type":325},25,"Machine Learning","machine-learning",{"learning_curve":359,"flexibility":360,"performance":361,"portability":362,"popularity":363},"Slightly more approachable than AWS; BigQuery and Cloud Run have excellent documentation.","Deep service catalog; Cloud Run, GKE, and BigQuery combine freely for any architecture.","BigQuery's columnar engine handles petabyte queries in seconds; Cloud Run scales fast.","Similar to AWS; proprietary managed services create meaningful switching costs.","Strong in data and AI workloads; second or third cloud by market share in most segments.",{"tool_id":365,"name":366,"slug":367,"tooltip_description":368,"logo_url":369,"logo_bg":94,"pricing_model":370,"learning_curve_score":145,"popularity_score":40,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":371,"subcategory":372,"categories":373,"subcategories":375,"flexibility_score":46,"performance_score":40,"portability_score":145,"is_featured":112,"tags":377,"score_reasonings":383,"published_date":134,"last_updated_date":135},54,"Microsoft Azure","azure","Microsoft's cloud platform with 200+ products and services. The second-largest cloud provider by revenue, after AWS.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fazure.svg",{"slug":296,"display_name":297,"description":298},{"category_id":46,"name":267,"slug":268},{"subcategory_id":120,"name":301,"slug":302},[374],{"category_id":46,"name":267,"slug":268,"is_primary":3,"display_order":109},[376],{"subcategory_id":120,"name":301,"slug":302,"category_id":46,"is_primary":3,"display_order":109},[378,379,380,381,382],{"tag_id":309,"name":310,"slug":311,"tag_type":118},{"tag_id":313,"name":314,"slug":315,"tag_type":118},{"tag_id":280,"name":281,"slug":282,"tag_type":118},{"tag_id":318,"name":319,"slug":320,"tag_type":118},{"tag_id":322,"name":323,"slug":324,"tag_type":325},{"learning_curve":384,"performance":385,"portability":386,"flexibility":387,"popularity":388},"Familiar for Microsoft ecosystem teams; CLI and portal are well-documented.","Enterprise SLA with globally distributed data centers; some managed services have higher latency than GCP.","Similar to AWS and GCP; proprietary services create lock-in for enterprise workloads.","Enterprise service breadth; Active Directory, DevOps, and AI services compose extensively.","Dominant in enterprise environments; Microsoft ecosystem adoption drives broad corporate use.",{"hetzner":390,"aws":391,"gcp":392,"azure":393},"The default and the lowest fixed price here: dedicated GPU servers with no virtualisation layer between the container and the card. The entry GEX45 has a 24 GB NVIDIA card and 64 GB of RAM for about €214 a month plus a one-time setup fee of about the same, enough for an 8B model in 16-bit form, or a 24B-class model quantized to 4 bits. The GEX131 carries a 96 GB card and 256 GB of RAM at about €889 a month, which holds a 70B-class model once it is quantized. Either is a flat monthly cost whether or not anyone is calling it, so it pays off for steady shared use. Docker and the NVIDIA container toolkit are yours to install.","For teams already on AWS. A g6e.xlarge has one NVIDIA L40S with 48 GB of card memory, 4 vCPUs, and 32 GiB of RAM for about $1.86 an hour, around $1,360 a month if left running, and the larger P-series instances carry A100 and H100 cards for bigger models. GPU instance types can need a quota increase before the first launch. Stopping the instance stops the compute charge while the volume keeps the downloaded weights, and the vLLM production stack documents Kubernetes deployments on AWS for the step past one server.","The same on Google Cloud: a g2-standard-4 with one NVIDIA L4 and 24 GB of card memory costs about $0.71 an hour, around $516 a month, which suits a small model, and an a2-highgpu-1g with an A100 40 GB, 12 vCPUs, and 85 GB of RAM costs about $3.67 an hour in US regions, around $2,680 a month. GPU quotas are set per project and region, so request them early. It fits when the applications calling the endpoint already run in the same project, so requests stay on the internal network, and the vLLM production stack covers GKE.","An NC24ads A100 v4 virtual machine gives one 80 GB A100, 24 vCPUs, and 220 GiB of RAM for about $3.67 an hour on demand, around $2,680 a month if left running. Spot capacity is far cheaper, about $0.68 an hour, but it can be evicted at any time, which is fine for batch work and wrong for an endpoint people depend on. It fits organizations whose data-residency rules name an Azure region: the weights, the prompts, and the traffic all stay inside it.",{},{"kind":240,"slug":268,"name":267,"href":396},"\u002Ftools\u002Fcategories\u002Fhosting-cloud",{"name":398,"tools":399,"descriptions":478,"aliases":482,"see_all":37},"Reverse Proxy",[400,429,454],{"tool_id":401,"name":402,"slug":403,"tooltip_description":404,"logo_url":405,"logo_bg":168,"pricing_model":406,"learning_curve_score":145,"popularity_score":40,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":407,"subcategory":410,"categories":414,"subcategories":416,"flexibility_score":40,"performance_score":40,"portability_score":46,"is_featured":112,"tags":418,"score_reasonings":423,"published_date":135,"last_updated_date":37},189,"Traefik","traefik","A cloud-native reverse proxy that auto-discovers routes from Docker container labels, updating its configuration live as containers start and stop.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Ftraefik.svg",{"slug":96,"display_name":97,"description":98},{"category_id":125,"name":408,"slug":409},"DevOps & CI\u002FCD","devops-cicd",{"subcategory_id":411,"name":412,"slug":413},57,"Web Servers & Reverse Proxies","web-servers-reverse-proxies",[415],{"category_id":125,"name":408,"slug":409,"is_primary":3,"display_order":109},[417],{"subcategory_id":411,"name":412,"slug":413,"category_id":125,"is_primary":3,"display_order":109},[419,420,421,422],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":318,"name":319,"slug":320,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":424,"flexibility":425,"performance":426,"popularity":427,"portability":428},"Basic Docker label routing is quick to pick up, but understanding the provider model, middleware chains, and dashboard takes more upfront investment than Caddy's flat config file, especially for anyone unfamiliar with label-driven tooling.","Native providers for Docker, Kubernetes, Swarm, Consul, and ECS plus a composable middleware system cover most container-routing scenarios, though its module ecosystem is narrower than NGINX's after two decades of third-party extensions.","Built in Go with an efficient routing engine, it performs comparably to Caddy for typical self-hosted and mid-scale container workloads.","It's the de facto reverse proxy for the Docker\u002FKubernetes self-hosting community and powers routing inside Coolify and Dokploy, with a large and actively growing GitHub following, though still behind NGINX's decades-long install base.","MIT licensed, ships as a single binary and official Docker image, and runs identically across any VPS, Kubernetes cluster, or cloud provider with no lock-in.",{"tool_id":430,"name":431,"slug":432,"tooltip_description":433,"logo_url":434,"logo_bg":94,"pricing_model":435,"learning_curve_score":436,"popularity_score":145,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":437,"subcategory":438,"categories":439,"subcategories":441,"flexibility_score":145,"performance_score":40,"portability_score":46,"is_featured":112,"tags":443,"score_reasonings":448,"published_date":135,"last_updated_date":37},188,"Caddy","caddy","A modern web server and reverse proxy written in Go, best known for provisioning and renewing HTTPS certificates automatically with zero configuration.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fcaddy.svg",{"slug":143,"display_name":116,"description":144},1,{"category_id":125,"name":408,"slug":409},{"subcategory_id":411,"name":412,"slug":413},[440],{"category_id":125,"name":408,"slug":409,"is_primary":3,"display_order":109},[442],{"subcategory_id":411,"name":412,"slug":413,"category_id":125,"is_primary":3,"display_order":109},[444,445,446,447],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":318,"name":319,"slug":320,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":449,"flexibility":450,"performance":451,"popularity":452,"portability":453},"A working reverse proxy with HTTPS is typically a few lines of Caddyfile, and automatic certificate provisioning removes the TLS setup step entirely, making it one of the easiest reverse proxies to get running for the first time.","The Caddyfile and JSON config cover most reverse-proxy and static-serving needs well, but its plugin ecosystem and edge-case module coverage are considerably smaller than NGINX's after two decades of third-party modules.","Built in Go with a modern, concurrent architecture and native HTTP\u002F3 support, it performs well for typical self-hosted and small-to-mid traffic workloads, though it hasn't accumulated NGINX's extreme-scale benchmark track record.","It's a well-known, growing choice in the self-hosted community specifically for its automatic-HTTPS pitch, but its adoption and star count remain well behind NGINX's decades-long dominance.","Apache 2.0 licensed, distributed as a single static binary with an official Docker image, and runs identically on any VPS or OS with zero vendor lock-in.",{"tool_id":455,"name":456,"slug":457,"tooltip_description":458,"logo_url":459,"logo_bg":94,"pricing_model":460,"learning_curve_score":145,"popularity_score":46,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":461,"subcategory":462,"categories":463,"subcategories":465,"flexibility_score":46,"performance_score":46,"portability_score":46,"is_featured":112,"tags":467,"score_reasonings":472,"published_date":135,"last_updated_date":37},187,"NGINX","nginx","The most widely deployed web server and reverse proxy on the internet, known for event-driven performance under high concurrency.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fnginx.svg",{"slug":96,"display_name":97,"description":98},{"category_id":125,"name":408,"slug":409},{"subcategory_id":411,"name":412,"slug":413},[464],{"category_id":125,"name":408,"slug":409,"is_primary":3,"display_order":109},[466],{"subcategory_id":411,"name":412,"slug":413,"category_id":125,"is_primary":3,"display_order":109},[468,469,470,471],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":318,"name":319,"slug":320,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":473,"flexibility":474,"performance":475,"popularity":476,"portability":477},"Basic reverse-proxy blocks are simple to copy-paste, but real production setups (upstream health checks, TLS automation, rewrite rules, rate limiting) require understanding NGINX's directive-based config language in depth, and there's no built-in automatic HTTPS to lean on.","A vast module ecosystem and a fully declarative config language let it act as a reverse proxy, load balancer, cache, static file server, or WAF front-end, with third-party modules covering nearly any edge case.","Its event-driven, non-blocking architecture was purpose-built to solve the C10k problem, and it remains a top performer in every major reverse-proxy benchmark for throughput and memory efficiency under concurrent load.","It has been the most widely deployed web server on the public internet for years, appears in essentially every stack survey, and has the largest surrounding tutorial and community knowledge base of any web server.","The open-source core is free, packaged for every OS and as an official Docker image, and runs identically on any VPS or cloud provider with no vendor lock-in.",{"traefik":479,"caddy":480,"nginx":481},"The family default: it discovers the vLLM container from Docker labels and renews certificates itself. Write the router so that only the API paths you serve, \u002Fv1, reach the container and everything else returns 404, because vLLM's --api-key covers only part of the API and leaves routes such as \u002Finvocations open. Raise the response timeouts, since streamed completions and long generations run for minutes and a proxy that cuts the connection breaks clients. Traefik forwards the Authorization header untouched, so the API key keeps working.","Automatic HTTPS and the shortest config: one handle block for \u002Fv1\u002F* that forwards to port 8000, and a closing respond 404 for every other path, which keeps the unauthenticated extra routes off the network. Caddy flushes event streams as they arrive, which suits streamed tokens. The API key is one shared secret for every client, so add Caddy's own authentication or a gateway in front if different teams need separate keys.","The proxy many servers already run. A location block for \u002Fv1\u002F that proxies to port 8000, with a default return 404 for everything else, keeps the open routes unreachable. Turn proxy_buffering off so streamed tokens are not held back, raise proxy_read_timeout well past the 60-second default for long generations, and lift client_max_body_size when prompts carry large documents. vLLM's own docs include an NGINX least-connections example for balancing several instances, which fits when one model runs on more than one GPU server.",{},{"name":37,"tools":484,"descriptions":485,"aliases":486,"see_all":37},[],{},{},{"model_aggregator":488,"chat_interface":529,"tunnel":614},{"name":489,"tools":490,"descriptions":522,"aliases":524,"preface":525,"see_all":526},"Model Aggregator",[491],{"tool_id":492,"name":493,"slug":494,"tooltip_description":495,"logo_url":496,"logo_bg":94,"pricing_model":497,"learning_curve_score":43,"popularity_score":40,"hosting_assignment_type":498,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":499,"subcategory":502,"categories":506,"subcategories":508,"flexibility_score":46,"performance_score":40,"portability_score":46,"is_featured":112,"tags":510,"score_reasonings":515,"published_date":135,"last_updated_date":521},221,"LiteLLM","litellm","The open-source AI gateway you run yourself: one OpenAI-compatible endpoint in front of more than 100 model providers, with virtual keys, budgets, and spend tracking. It is the self-hosted counterpart to OpenRouter.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Flitellm.png",{"slug":96,"display_name":97,"description":98},"deployable",{"category_id":280,"name":500,"slug":501},"AI Infrastructure","ai-infrastructure",{"subcategory_id":503,"name":504,"slug":505},63,"AI Model Aggregators","ai-model-aggregators",[507],{"category_id":280,"name":500,"slug":501,"is_primary":3,"display_order":109},[509],{"subcategory_id":503,"name":504,"slug":505,"category_id":280,"is_primary":3,"display_order":109},[511,512,513,514],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":355,"name":356,"slug":357,"tag_type":325},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":516,"flexibility":517,"performance":518,"popularity":519,"portability":520},"Requires standing up and maintaining real infrastructure, PostgreSQL, Redis, the proxy server itself, meaningfully more setup than a zero-ops hosted aggregator like OpenRouter.","Full control over routing logic, guardrails, virtual keys, and spend limits across 100+ providers, the most configurable option in the aggregator genre precisely because you own the deployment.","8ms P95 latency at 1,000 RPS and a Rust core deliver genuinely fast proxy performance, though real-world throughput depends on the self-hosted infrastructure backing it.","56K+ GitHub stars and production adoption at Stripe, Google ADK, Greptile, and OpenHands make it the clear self-hosted leader in this genre.","MIT licensed, genuinely self-hostable anywhere with Docker or the provided Terraform modules, no vendor lock-in to a hosted service.","2026-10-06",{"litellm":523},"vLLM serves one model per server instance, and its FAQ says several models on one port are not supported, so a second model means a second server behind a routing layer. LiteLLM is that layer: the entry uses the hosted_vllm\u002F prefix with the server's address, and the gateway adds virtual keys, budgets, rate limits, and a fallback to a hosted provider when the GPU box is down. Applications then never need vLLM's own address. LiteLLM Self-Hosted is a stack of its own.",{},"Add a model aggregator when you want one API key and one bill for models from many providers, with automatic fallback when one of them is down, instead of setting up each provider separately.",{"kind":527,"slug":505,"name":504,"href":528},"subcategory","\u002Ftools\u002Fcategories\u002Fai-infrastructure\u002Fai-model-aggregators",{"name":530,"tools":531,"descriptions":608,"aliases":612,"preface":613,"see_all":37},"Chat Interface",[532,560,584],{"tool_id":533,"name":534,"slug":535,"tooltip_description":536,"logo_url":537,"logo_bg":168,"pricing_model":538,"learning_curve_score":43,"popularity_score":40,"hosting_assignment_type":498,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":539,"subcategory":540,"categories":544,"subcategories":546,"flexibility_score":40,"performance_score":145,"portability_score":40,"is_featured":112,"tags":548,"score_reasonings":553,"published_date":135,"last_updated_date":559},219,"Open WebUI","open-webui","The most widely deployed self-hosted chat UI (149K+ GitHub stars), a feature-rich frontend for Ollama or any OpenAI-compatible API with RAG, RBAC, and enterprise auth — doesn't serve inference itself, just the interface to talk to whatever does.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fopen-webui.png",{"slug":143,"display_name":116,"description":144},{"category_id":280,"name":500,"slug":501},{"subcategory_id":541,"name":542,"slug":543},64,"AI Chat Interfaces","ai-chat-interfaces",[545],{"category_id":280,"name":500,"slug":501,"is_primary":3,"display_order":109},[547],{"subcategory_id":541,"name":542,"slug":543,"category_id":280,"is_primary":3,"display_order":109},[549,550,551,552],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":355,"name":356,"slug":357,"tag_type":325},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"flexibility":554,"performance":555,"portability":556,"popularity":557,"learning_curve":558},"Connects to Ollama or any OpenAI-compatible backend, plus RAG, web search, and a plugin architecture cover most self-hosted AI-chat use cases without needing a different tool.","Performance is largely a function of the backend it's paired with rather than Open WebUI itself, the UI layer adds minimal overhead on top of whatever is serving inference.","Genuinely self-hostable via Docker, pip, or Helm with database choice (Postgres or SQLite), though the custom license's branding requirement is a real constraint above 50 users.","A self-hosted, ChatGPT-style interface for local models served through Ollama or llama.cpp, with the largest community of any project in that space. Self-hosting your own chat UI is a specific privacy-minded habit within the broader AI-tooling world, not something most developers have set up themselves.","Docker Compose or a single `docker run` gets a working chat UI in minutes when paired with an existing Ollama install, though the deeper feature set (RAG, plugins, enterprise auth) takes more setup.","2026-10-01",{"tool_id":561,"name":562,"slug":563,"tooltip_description":564,"logo_url":565,"logo_bg":94,"pricing_model":566,"learning_curve_score":145,"popularity_score":40,"hosting_assignment_type":498,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":567,"subcategory":568,"categories":569,"subcategories":571,"flexibility_score":46,"performance_score":40,"portability_score":46,"is_featured":112,"tags":573,"score_reasonings":578,"published_date":135,"last_updated_date":37},223,"LibreChat","librechat","Self-hosted, MIT-licensed chat UI with the broadest multi-provider support of the self-hosted field, plus a sandboxed code interpreter, agents with MCP, and per-user token spend tracking.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Flibrechat.png",{"slug":143,"display_name":116,"description":144},{"category_id":280,"name":500,"slug":501},{"subcategory_id":541,"name":542,"slug":543},[570],{"category_id":280,"name":500,"slug":501,"is_primary":3,"display_order":109},[572],{"subcategory_id":541,"name":542,"slug":543,"category_id":280,"is_primary":3,"display_order":109},[574,575,576,577],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":355,"name":356,"slug":357,"tag_type":325},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":579,"flexibility":580,"performance":581,"popularity":582,"portability":583},"Docker Compose gets a basic instance running quickly, but configuring the full breadth of supported providers, agents, and the code interpreter takes real setup time.","The widest multi-provider support of any self-hosted chat UI, plus agents\u002FMCP, a sandboxed multi-language code interpreter, and generative UI cover the broadest range of use cases in its category.","Sandboxed code execution across 8 languages and per-user token tracking demonstrate genuine production-grade engineering, though real-world scaling depends on the self-hosted infrastructure behind it.","42K+ GitHub stars make it a well-established peer to Open WebUI and AnythingLLM, particularly known for multi-provider breadth in comparisons.","MIT licensed with no branding requirements, self-hostable via Docker Compose, cloud one-click templates, or Kubernetes, and portable across the widest set of LLM providers of the three.",{"tool_id":585,"name":586,"slug":587,"tooltip_description":588,"logo_url":589,"logo_bg":168,"pricing_model":590,"learning_curve_score":43,"popularity_score":40,"hosting_assignment_type":498,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":591,"subcategory":592,"categories":593,"subcategories":595,"flexibility_score":46,"performance_score":145,"portability_score":46,"is_featured":112,"tags":597,"score_reasonings":602,"published_date":135,"last_updated_date":559},222,"AnythingLLM","anythingllm","Self-hosted chat UI built around RAG from the ground up: MIT licensed, with per-workspace document sets and vector DB settings, plus a desktop app that skips Docker entirely.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fanythingllm.svg",{"slug":143,"display_name":116,"description":144},{"category_id":280,"name":500,"slug":501},{"subcategory_id":541,"name":542,"slug":543},[594],{"category_id":280,"name":500,"slug":501,"is_primary":3,"display_order":109},[596],{"subcategory_id":541,"name":542,"slug":543,"category_id":280,"is_primary":3,"display_order":109},[598,599,600,601],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":355,"name":356,"slug":357,"tag_type":325},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"flexibility":603,"performance":604,"popularity":605,"portability":606,"learning_curve":607},"Per-workspace document\u002Fvector-DB\u002Fretrieval configuration, 25+ LLM providers, and 7 vector database backends make it the most configurable RAG-first chat UI in its category.","Suitable for solo and small-team RAG workloads, but multi-user administration and scaling are less proven at organizational scale than Open WebUI's more mature RBAC.","64K+ GitHub stars make it a genuine, well-established peer to Open WebUI and LibreChat, though smaller than Open WebUI's community by a wide margin.","MIT licensed with no branding requirements or user thresholds, genuinely self-hostable via Docker or desktop app, and portable across LLM providers and vector databases.","A native desktop app for Mac\u002FWindows\u002FLinux gets a solo user running in minutes with zero Docker or terminal knowledge required; the self-hosted Docker path for teams takes more setup, especially configuring per-workspace vector DBs.",{"open-webui":609,"librechat":610,"anythingllm":611},"The interface vLLM's own docs describe: run it in Docker and set the OpenAI API base URL to the server's \u002Fv1 address, and the served model appears at the top of the model picker. It adds accounts, conversation history, and document chat with built-in retrieval. Its licence keeps the branding in place unless you have 50 or fewer users or an enterprise licence. Point it at the reverse proxy or the gateway rather than at port 8000, so the interface and every other client use the same endpoint.","A self-hosted, MIT-licensed chat interface with the broadest multi-provider support, which suits a team that already mixes hosted and self-hosted models. It connects through a custom endpoint in its librechat.yaml file, with the vLLM server's \u002Fv1 address and the API key, using its generic OpenAI-compatible support; vLLM is not named in its docs. It runs its own services for accounts and conversation storage, so it is a heavier deployment than the model server beside it.","A self-hosted, MIT-licensed interface built around chatting with your own documents: each workspace has its own document set and vector settings. It reaches vLLM through its generic OpenAI-compatible provider, set to the server's \u002Fv1 address and key, since its docs do not name vLLM. A desktop app skips Docker entirely, which suits one person trying the endpoint, while the Docker version serves a team.",{},"Add a chat interface when people, not only applications, will talk to the model: a web app with accounts, conversation history, and a model picker that connects to the API you already serve.",{"name":615,"tools":616,"descriptions":663,"aliases":666,"preface":667,"see_all":37},"Tunnel",[617,642],{"tool_id":618,"name":619,"slug":620,"tooltip_description":621,"logo_url":622,"logo_bg":94,"pricing_model":623,"learning_curve_score":145,"popularity_score":145,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":624,"subcategory":625,"categories":629,"subcategories":631,"flexibility_score":40,"performance_score":40,"portability_score":43,"is_featured":112,"tags":633,"score_reasonings":636,"published_date":135,"last_updated_date":521},258,"Cloudflare Tunnel","cloudflare-tunnel","Cloudflare Tunnel creates an outbound-only encrypted connection from your server to Cloudflare's network, exposing services to the internet without opening inbound ports. Named tunnels are the production mode, and no-account Quick Tunnels are for testing.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fcloudflare.svg",{"slug":96,"display_name":97,"description":98},{"category_id":125,"name":408,"slug":409},{"subcategory_id":626,"name":627,"slug":628},70,"Tunneling & Secure Access","tunneling-secure-access",[630],{"category_id":125,"name":408,"slug":409,"is_primary":3,"display_order":109},[632],{"subcategory_id":626,"name":627,"slug":628,"category_id":125,"is_primary":3,"display_order":109},[634,635],{"tag_id":347,"name":348,"slug":349,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":637,"flexibility":638,"performance":639,"popularity":640,"portability":641},"Getting a tunnel running requires a Cloudflare account, a domain on Cloudflare DNS, and configuring ingress rules for each hostname — more setup than a single-binary tunnel tool, though well documented and a common homelab pattern once learned.","Ingress rules can route many hostnames to different local services from one tunnel, Zero Trust Access adds per-hostname authentication, and the same daemon handles HTTP alongside TCP\u002FUDP via WARP.","Traffic rides Cloudflare's global Anycast network end to end, giving it the same routing and edge performance as the rest of Cloudflare's CDN.","A standard recommendation in self-hosting and homelab communities for exposing services without port forwarding, though it competes with several other tunnel tools rather than being the default name people reach for first.","Tightly coupled to a Cloudflare account and DNS zone — migrating off Cloudflare means replacing the entire exposure mechanism, not just swapping a config value.",{"tool_id":643,"name":644,"slug":644,"tooltip_description":645,"logo_url":646,"logo_bg":168,"pricing_model":647,"learning_curve_score":436,"popularity_score":40,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":648,"subcategory":649,"categories":650,"subcategories":652,"flexibility_score":145,"performance_score":145,"portability_score":40,"is_featured":112,"tags":654,"score_reasonings":657,"published_date":135,"last_updated_date":37},259,"ngrok","ngrok is a globally distributed reverse proxy that creates secure public URLs for a local server in seconds, commonly used for local dev previews, webhook testing, and exposing self-hosted services.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fngrok.svg",{"slug":96,"display_name":97,"description":98},{"category_id":125,"name":408,"slug":409},{"subcategory_id":626,"name":627,"slug":628},[651],{"category_id":125,"name":408,"slug":409,"is_primary":3,"display_order":109},[653],{"subcategory_id":626,"name":627,"slug":628,"category_id":125,"is_primary":3,"display_order":109},[655,656],{"tag_id":347,"name":348,"slug":349,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"learning_curve":658,"flexibility":659,"performance":660,"popularity":661,"portability":662},"A single CLI command produces a working public URL with no account configuration, DNS setup, or domain ownership required to get started.","Covers HTTP, TCP, and TLS tunnels plus Kubernetes ingress, with edge traffic policies for auth and rate limiting, though the deepest features sit behind paid plans.","Runs on a globally distributed edge network with solid latency for typical dev-preview and webhook workloads, without the scale of the largest CDN networks.","The name most developers reach for first when they need to expose a local server — long-standing mindshare in webhook testing and dev-preview workflows specifically.","A standalone client binary that works against any local service or cloud origin with no DNS zone or vendor account tie-in beyond ngrok itself, making it easy to drop into any project.",{"cloudflare-tunnel":664,"ngrok":665},"A public HTTPS address through an outbound-only connection, for clients that run somewhere you do not control: an application on someone else's platform can call a tunnel hostname while no inbound port is open on the server. Point the tunnel at the reverse proxy, not at port 8000, so the endpoint allowlist still applies. Free, with Cloudflare Access in front if you want a login layer. A request that gets no answer within about 125 seconds fails with a 524 and non-Enterprise plans cannot raise that, so stream long generations. A stable hostname needs a domain on Cloudflare.","The quick-start tunnel, for trying the endpoint from a laptop: one command gives a public HTTPS address and no domain is needed. The free plan includes three endpoints, 1 GB of transfer, 20,000 requests, and an interstitial page on browser visits, and streamed completions use transfer quickly, so it suits a trial and not the address every application calls. Paid plans start at $10 a month, and sustained use belongs on pay-as-you-go from $20 a month or on Cloudflare Tunnel.",{},"Add a tunnel when you're self-hosting without a static IP or can't open inbound ports — a home server, a VPS behind restrictive network policies, or anywhere a reverse proxy alone can't reach the internet.",{},[],{"Hosting & Cloud":671,"DevOps & CI\u002FCD":684,"LLM":716,"AI Infrastructure":730},[672],{"tool_id":104,"name":258,"slug":259,"tooltip_description":260,"logo_url":261,"logo_bg":168,"pricing_model":673,"learning_curve_score":43,"popularity_score":145,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":674,"subcategory":675,"categories":676,"subcategories":678,"flexibility_score":145,"performance_score":40,"portability_score":40,"is_featured":112,"tags":680,"score_reasonings":683,"published_date":134,"last_updated_date":135},{"slug":263,"display_name":264,"description":265},{"category_id":46,"name":267,"slug":268},{"subcategory_id":270,"name":271,"slug":272},[677],{"category_id":46,"name":267,"slug":268,"is_primary":3,"display_order":109},[679],{"subcategory_id":270,"name":271,"slug":272,"category_id":46,"is_primary":3,"display_order":109},[681,682],{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":280,"name":281,"slug":282,"tag_type":118},{"learning_curve":284,"flexibility":285,"performance":286,"portability":287,"popularity":288},[685],{"tool_id":686,"name":687,"slug":688,"tooltip_description":689,"logo_url":690,"logo_bg":94,"pricing_model":691,"learning_curve_score":40,"popularity_score":46,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":692,"subcategory":693,"categories":697,"subcategories":699,"flexibility_score":46,"performance_score":40,"portability_score":46,"is_featured":3,"tags":701,"score_reasonings":710,"published_date":134,"last_updated_date":135},72,"Docker","docker","Container platform for packaging applications and their dependencies into portable images that run the same on a laptop, in CI, and in production, with Docker Desktop, Compose, and Docker Hub.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fdocker.svg",{"slug":96,"display_name":97,"description":98},{"category_id":125,"name":408,"slug":409},{"subcategory_id":694,"name":695,"slug":696},33,"Containerization","containerization",[698],{"category_id":125,"name":408,"slug":409,"is_primary":3,"display_order":109},[700],{"subcategory_id":694,"name":695,"slug":696,"category_id":125,"is_primary":3,"display_order":109},[702,703,704,705,706],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":318,"name":319,"slug":320,"tag_type":118},{"tag_id":322,"name":323,"slug":324,"tag_type":325},{"tag_id":707,"name":708,"slug":709,"tag_type":123},43,"Cross-platform","cross-platform",{"learning_curve":711,"performance":712,"portability":713,"flexibility":714,"popularity":715},"Container images, networking, volumes, and multi-stage builds all need deliberate learning.","Container overhead is minimal; near-native performance for most workloads.","Open standard; containers built with Docker run on any container-compatible platform.","Any runtime, any architecture; multi-stage builds and Compose profiles support complex systems.","The standard for containerization; present in virtually every modern software project.",[717],{"tool_id":89,"name":90,"slug":91,"tooltip_description":92,"logo_url":93,"logo_bg":94,"pricing_model":718,"learning_curve_score":43,"popularity_score":40,"hosting_assignment_type":37,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":719,"subcategory":720,"categories":721,"subcategories":723,"flexibility_score":46,"performance_score":46,"portability_score":46,"is_featured":112,"tags":725,"score_reasonings":729,"published_date":134,"last_updated_date":135},{"slug":96,"display_name":97,"description":98},{"category_id":101,"name":86,"slug":102},{"subcategory_id":104,"name":105,"slug":106},[722],{"category_id":101,"name":86,"slug":102,"is_primary":3,"display_order":109},[724],{"subcategory_id":104,"name":105,"slug":106,"category_id":101,"is_primary":3,"display_order":109},[726,727,728],{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":120,"name":121,"slug":122,"tag_type":123},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"learning_curve":129,"flexibility":130,"performance":131,"popularity":132,"portability":133},[731],{"tool_id":732,"name":733,"slug":734,"tooltip_description":735,"logo_url":736,"logo_bg":94,"pricing_model":737,"learning_curve_score":40,"popularity_score":40,"hosting_assignment_type":498,"hosting_provider_restriction":99,"hosting_target_restriction":99,"hosting_compatible_tool_ids":37,"parent_tool_id":37,"category":738,"subcategory":739,"categories":743,"subcategories":745,"flexibility_score":46,"performance_score":46,"portability_score":46,"is_featured":112,"tags":747,"score_reasonings":755,"published_date":559,"last_updated_date":37},298,"vLLM","vllm","Open-source, high-throughput inference and serving engine for large language models, exposing an OpenAI-compatible API. It is the most common way to self-host open-weight models in production on GPU and accelerator clusters.","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fvllm.svg",{"slug":143,"display_name":116,"description":144},{"category_id":280,"name":500,"slug":501},{"subcategory_id":740,"name":741,"slug":742},60,"AI Runtime & Serving","ai-runtime-serving",[744],{"category_id":280,"name":500,"slug":501,"is_primary":3,"display_order":109},[746],{"subcategory_id":740,"name":741,"slug":742,"category_id":280,"is_primary":3,"display_order":109},[748,752,753,754],{"tag_id":436,"name":749,"slug":750,"tag_type":751},"Python","python","technology",{"tag_id":115,"name":116,"slug":117,"tag_type":118},{"tag_id":125,"name":126,"slug":127,"tag_type":118},{"tag_id":355,"name":356,"slug":357,"tag_type":325},{"learning_curve":756,"flexibility":757,"performance":758,"popularity":759,"portability":760},"Starting a server is one command, but running it well in production means understanding GPU memory, KV-cache sizing, quantization, and multi-GPU parallelism, plus the Kubernetes layer around it.","Serves hundreds of model architectures with configurable quantization, parallelism, LoRA adapters, structured outputs, and speculative decoding, and can be embedded as a Python library or run as a server.","PagedAttention and continuous batching set the throughput bar that other open-source engines are measured against, and the V1 engine cut scheduling overhead further.","Around 93K GitHub stars and the engine underneath many hosted inference services, so anyone self-hosting models knows it, while developers who only call hosted APIs rarely touch it directly.","Apache-2.0, runs on GPUs and accelerators from several vendors, installs anywhere Python or Docker runs, and exposes a standard OpenAI-style API.",[762,782,806,824],{"stack_id":763,"slug":764,"name":765,"tagline":766,"experience_level":767,"project_type":768,"stack_type_slug":769,"stack_type_icon_url":770,"score_popularity":145,"score_learning_curve":43,"catalog_display_order":37,"published_date":37,"last_updated_date":37,"core_tool_previews":771},156,"strapi-self-hosted","Strapi Self-Hosted","Self-hosted Strapi: an open-source headless CMS with PostgreSQL, on a server you control.","beginner","website","infrastructure","https:\u002F\u002Fassets.tekyous.dev\u002Ficons\u002Fstack-types\u002Finfrastructure.svg",[772,776,777],{"tool_id":120,"slug":773,"name":774,"logo_url":775,"logo_bg":94},"postgresql","PostgreSQL","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fpostgresql.svg",{"tool_id":686,"slug":688,"name":687,"logo_url":690,"logo_bg":94},{"tool_id":778,"slug":779,"name":780,"logo_url":781,"logo_bg":168},88,"strapi","Strapi","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fstrapi.svg",{"stack_id":783,"slug":784,"name":785,"tagline":786,"experience_level":787,"project_type":788,"stack_type_slug":769,"stack_type_icon_url":770,"score_popularity":40,"score_learning_curve":43,"catalog_display_order":37,"published_date":37,"last_updated_date":37,"core_tool_previews":789},184,"grafana-self-hosted","Grafana Self-Hosted","Self-hosted Grafana and Prometheus: metrics dashboards and alerting on a server you run.","intermediate","dashboard",[790,795,796,801],{"tool_id":791,"slug":792,"name":793,"logo_url":794,"logo_bg":168},124,"sqlite","SQLite","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fsqlite.svg",{"tool_id":686,"slug":688,"name":687,"logo_url":690,"logo_bg":94},{"tool_id":797,"slug":798,"name":799,"logo_url":800,"logo_bg":94},91,"grafana","Grafana","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fgrafana.svg",{"tool_id":802,"slug":803,"name":804,"logo_url":805,"logo_bg":168},94,"prometheus","Prometheus","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fprometheus.svg",{"stack_id":533,"slug":807,"name":808,"tagline":809,"experience_level":787,"project_type":810,"stack_type_slug":769,"stack_type_icon_url":770,"score_popularity":46,"score_learning_curve":40,"catalog_display_order":37,"published_date":37,"last_updated_date":37,"core_tool_previews":811},"apache-airflow-self-hosted","Airflow Self-Hosted","Self-hosted Apache Airflow: the standard data-pipeline scheduler on your infrastructure.","data_pipeline",[812,813,818,819],{"tool_id":120,"slug":773,"name":774,"logo_url":775,"logo_bg":94},{"tool_id":814,"slug":815,"name":816,"logo_url":817,"logo_bg":94},41,"redis","Redis","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fredis.svg",{"tool_id":686,"slug":688,"name":687,"logo_url":690,"logo_bg":94},{"tool_id":820,"slug":821,"name":822,"logo_url":823,"logo_bg":94},9,"apache-airflow","Apache Airflow","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fapache-airflow.svg",{"stack_id":825,"slug":826,"name":827,"tagline":828,"experience_level":787,"project_type":829,"stack_type_slug":769,"stack_type_icon_url":770,"score_popularity":40,"score_learning_curve":145,"catalog_display_order":37,"published_date":37,"last_updated_date":37,"core_tool_previews":830},150,"n8n-self-hosted","n8n Self-Hosted","Self-hosted n8n on your own server, with full control over the database, the host, and how it's exposed to the internet.","automation",[831,832,836],{"tool_id":120,"slug":773,"name":774,"logo_url":775,"logo_bg":94},{"tool_id":833,"slug":834,"name":834,"logo_url":835,"logo_bg":94},22,"n8n","https:\u002F\u002Fassets.tekyous.dev\u002Flogos\u002Ftools\u002Fn8n.svg",{"tool_id":686,"slug":688,"name":687,"logo_url":690,"logo_bg":94},[838,841,844,847,850],{"question":839,"answer":840},"Is there a managed vLLM, and why run your own?","There is no vLLM cloud: it is a community project that started at UC Berkeley, and anyone can run it. The managed equivalents are inference providers that serve open models behind the same API and bill per token, such as Groq or Hugging Face's hosted endpoints. The trade is money against control. A provider costs nothing while idle and nothing to operate, but the bill grows with every token, and your prompts and weights sit on its infrastructure. A GPU server is a fixed monthly price, from about €214 for a 24 GB card to about €889 for a 96 GB one, so it wins only when the load is steady and shared. It also runs models no provider carries, including your own fine-tunes.",{"question":842,"answer":843},"How much GPU memory does a model need?","The card must hold the weights plus the context cache. At 16-bit precision the weights take about two bytes per parameter: roughly 16 GB for an 8B model, 64 GB for a 32B model, and 140 GB for a 70B model. Quantization to 8 or 4 bits cuts that by half or three quarters, and by default vLLM claims about 92% of the card. These are estimates, since vLLM publishes no per-model figures. When one card is not enough, tensor parallelism splits the model across several GPUs in a server; the default is one.",{"question":845,"answer":846},"What needs backing up, and how do upgrades work?","Almost nothing, because the container is stateless. The only large files are the model weights in the mounted Hugging Face cache, which download again if the volume is lost, so keep your own copy only of fine-tuned weights or adapters nobody else hosts. Pin the image to a version tag instead of latest: the quickstart uses latest, which moves with every release, and a new release can change defaults or kernel and quantization support, so test your model after each upgrade. The configuration to keep in version control is the launch command: model, context length, parallelism, and API key.",{"question":848,"answer":849},"Is it actually free? Licensing and model terms.","The engine is Apache 2.0, with no paid edition and no feature gating, so the cost is the GPU server and the time to run it. The weights are separate: each model has its own licence. Licences run from Apache 2.0 and MIT to Meta's own community licence for Llama, and each model entry above says which applies. Gated models need an account on Hugging Face and an accepted licence before they download. Read the licence of the exact model before building a product on it, because families change terms between releases.",{"question":851,"answer":852},"vLLM or Ollama: which should serve my model?","They suit different loads. Ollama is a single-machine runtime: it pulls a model by name, runs on a laptop or a modest card, and pairs with a chat interface in minutes, which is right for one person or a small team. vLLM is built for many concurrent requests on a datacenter GPU: continuous batching keeps the card busy across users, models come straight from Hugging Face, and it has no model library or interface of its own. If a handful of people chat, Ollama is simpler. If agents and applications send requests all day, vLLM serves more of them from the same card. Self-Hosted AI with Ollama and Open WebUI is the stack for the first case.",{"summary":854,"starting_cost_label":855,"has_free_tier":3,"line_items":856},"The engine is free, so the fixed cost is the GPU server, and the card decides the price: about €214 a month for a dedicated server with a 24 GB card, about €889 for a 96 GB card, and about $1,360 to $2,700 a month for a cloud GPU instance left running. Most open model weights download free, though their licences vary. Compared with a hosted provider, the bill stops growing with tokens, but it does not shrink when nobody is calling.","From ~€214\u002Fmo",[857,861,864,868],{"label":858,"cost":859,"note":860},"Server (GPU, any provider)","€214–2,700\u002Fmo","The card decides the price. A dedicated server with a 24 GB card is about €214 a month and one with a 96 GB card about €889 at Hetzner. Cloud GPU instances run about $1.86 an hour for a 48 GB L40S on AWS and about $3.67 an hour for an A100 on Google Cloud or Azure, roughly $1,360 to $2,700 a month if left running. Spot and reserved terms lower the hourly rate, spot at the risk of eviction.",{"label":733,"cost":862,"note":863},"Free (Apache 2.0)","The engine, its OpenAI-compatible server, and the Kubernetes production stack are open source, with no paid edition.",{"label":865,"cost":866,"note":867},"Model weights","Free to download","Most open models download free from Hugging Face. Each has its own licence, and gated ones need an account and an accepted licence first.",{"label":869,"cost":870,"note":871},"Exposure","Free","The reverse proxies and tunnels here are free, open-source software or free tiers, and TLS certificates come from Let's Encrypt.",{"is_official":112,"source_url":37,"items":873,"notes":886},[874,877,880,883],{"label":875,"value":876},"GPU","NVIDIA with compute capability 7.5 or higher (T4, RTX 20 series, L4, A100, H100, B200), or a supported AMD or Intel GPU",{"label":878,"value":879},"GPU memory","About 2 GB per billion parameters at 16-bit precision, plus room for the context cache; about half that at 8 bits and a quarter at 4 bits",{"label":881,"value":882},"Disk","Space for the model files, about the size of the weights, plus the Docker image",{"label":884,"value":885},"OS","Linux with Docker and the NVIDIA container toolkit; Python 3.10 to 3.13 for an install without Docker","The GPU and OS rows are quoted from vLLM's installation page (Linux, Python 3.10 to 3.13, NVIDIA compute capability 7.5 or higher, plus supported AMD and Intel GPUs). vLLM publishes no memory floor, so the memory row is an estimate: two bytes per parameter is the size of 16-bit weights, and the context cache comes on top, which depends on the context length and the number of concurrent requests. By default the server claims about 92% of the card's memory, and the tensor parallel size defaults to 1.","advanced","ai_agents",{"title":890,"description":891,"og_image":37,"canonical":892},"vLLM Inference Server: Tools, Pricing & How to Deploy | Tekyous","A self-hosted, OpenAI-compatible API for open-weight models, served from your own GPU… Compare vLLM Inference Server tools, pricing & how to deploy on Tekyous.","https:\u002F\u002Ftekyous.dev\u002Fstacks\u002Fvllm-self-hosted",[],1791281195456]