llama.cpp

llama.cpp

Open Source

LLM inference in C/C++.

AI Infrastructure
AI Runtime & Serving

Published 2 October 2026

Scores

Popularity4/5

Around 130K GitHub stars and the engine behind much of the local AI ecosystem, well known to anyone running models locally, while developers who only call hosted APIs rarely use it directly.

Learning Curve3/5

A one-line install and `llama serve -hf` get a model running in minutes, but getting the most out of it means picking a quantization, learning offload and context flags, and sometimes compiling for a specific GPU backend.

Flexibility5/5

Exposes nearly every inference parameter, runs as a CLI, a server, or an embedded library, serves several models through router mode, and speaks both OpenAI and Anthropic API formats.

Performance4/5

Best-in-class speed for single-user inference on consumer hardware and Apple Silicon, though batched multi-user throughput on GPU servers trails engines built for it, such as vLLM.

Portability5/5

MIT-licensed with no dependencies, runs on macOS, Windows, Linux, and Android across every major GPU vendor and plain CPUs, and GGUF files work in any compatible runtime.

About llama.cpp

llama.cpp is an open-source LLM inference engine written in plain C/C++ with no external dependencies. It loads models in GGUF, the quantized file format the project created, and runs them on laptops, desktops, and servers. Quantization from 1.5-bit to 8-bit lets a model that would need a data-center GPU at full precision fit into the memory of an ordinary machine, which is why most open-weight models appear on Hugging Face as GGUF files within days of release.

The engine treats Apple Silicon as a first-class target through Metal and also supports NVIDIA (CUDA), AMD (HIP), Intel (SYCL), Vulkan, and plain CPUs with AVX and ARM optimizations, including hybrid CPU and GPU offloading for models larger than video memory. Vision language models run through the same binaries.

Its tools run from a single llama command. llama cli chats with a model in the terminal, while llama serve (also built as llama-server) exposes an OpenAI-compatible HTTP API, an Anthropic Messages endpoint, and a built-in web chat UI, and the -hf flag downloads a model straight from Hugging Face. Router mode lets one server load, unload, and switch between several models on demand. Prebuilt binaries, Docker images, and package-manager installs are available, and compiling from source unlocks hardware-specific builds. The team also publishes Llama, a small menu-bar app for Mac and Windows that runs the same engine and picks the settings for each model.

llama.cpp is also the engine underneath much of local AI: LM Studio runs GGUF models on it, and Ollama builds on the same GGML library. People who choose it directly do so for control: exact quantization, sampling, and GPU-offload settings, no background daemon, and support for new model architectures before the wrappers catch up.

The project is MIT-licensed. Its creator, Georgi Gerganov, and the ggml.ai team joined Hugging Face in February 2026; the repositories stayed open source under the same license and the team kept maintaining them full-time.

Key Features

  • Runs GGUF models with 1.5-bit to 8-bit quantization
  • Metal, CUDA, HIP, SYCL, Vulkan, and optimized CPU backends
  • Hybrid CPU and GPU offloading for models larger than video memory
  • llama-server with OpenAI-compatible and Anthropic Messages APIs
  • Built-in web chat UI served by llama-server
  • Router mode to load and switch between several models without restarts
  • Vision language model support through the same tools
  • Prebuilt binaries, Docker images, and direct model downloads from Hugging Face

Pros

  • Runs capable models on ordinary laptops and consumer GPUs thanks to aggressive quantization
  • Fine-grained control over quantization, sampling, context size, and GPU offload that wrappers hide
  • New model architectures usually land here before Ollama or LM Studio pick them up
  • No daemon and no account, just a binary and a model file

Cons

  • Steeper setup than Ollama or LM Studio: choosing the right build and learning the command-line flags takes time
  • No curated model library like Ollama's: users pick a GGUF repository and quantization on Hugging Face themselves
  • Aimed at one user or a small team; vLLM handles many concurrent requests on GPU clusters far better
  • Fast release cadence means flags and defaults change often

llama.cpp Pricing

Open Source

Tech Stacks with llama.cpp

OpenCode Dev Stack

Developer

OpenCode as the coding agent, in the terminal, the desktop app, an IDE extension, or the browser, running open-weight models such as DeepSeek, GLM, and Kimi, with your pick of GitHub or GitLab for version control.

LLM:
Version Control:
Server add-on:
Remote Access add-on:
Session Persistence add-on:
Terminal add-on:
Code Review add-on:
Model Inference add-on:
Model Aggregator add-on:
CI/CD add-on:
Containerization add-on:

Local LLM Dev Stack

Developer

A coding agent that works against a model on your own hardware. A local runtime such as Ollama or LM Studio serves an open-weight coding model such as Qwen or Mistral's Devstral, and an agent such as OpenCode or Cline does the editing. No code leaves the machine and there is no per-token bill.

Coding Agent:
LLM:
Version Control:
Model Inference:
Server add-on:
Remote Access add-on:
Terminal add-on:
Code Review add-on:
Model Aggregator add-on:
CI/CD add-on:

Tools Related to llama.cpp

Integrates with llama.cpp(1)

llama.cpp downloads GGUF models straight from the Hugging Face Hub with its -hf flag, and Hugging Face model pages list llama.cpp among the local apps for running a model. The llama.cpp team is part of Hugging Face.

Alternatives to llama.cpp(3)

Ollama adds a model library, a pull command, and a managed background server on top of the same GGML engine family; llama.cpp is the engine itself, with every quantization, sampling, and GPU-offload setting exposed. Pick Ollama for convenience, llama.cpp for control and the newest model support.

LM Studio is a closed-source desktop app that finds, downloads, and chats with local models and runs GGUF files on llama.cpp under the hood; llama.cpp is the open-source engine used directly from the command line or as a server. Pick LM Studio for a graphical app, llama.cpp for scripting and headless machines.

llama.cpp runs quantized GGUF models on laptops, CPUs, and single consumer GPUs for one user or a small team; vLLM batches many concurrent requests across datacenter GPUs. Use llama.cpp for local and edge inference, vLLM once many users hit the same model.

Tags

Open SourceSelf-hostableMachine LearningCross-platform

Details

Maintained
Yes
Primary language
C++
Domain
ML / AI
GitHub stars
130k
Stars updated
2026-10-01