LM Studio

LM Studio

Freemium

Discover, download, and run local LLMs.

APIs & Infrastructure
AI Runtime & Serving

Published 1 October 2026

Scores

Popularity4/5

One of the two default answers, alongside Ollama, whenever someone asks how to run an LLM locally, with a large user base on Mac and Windows beyond professional developers.

Learning Curve1/5

A graphical installer, in-app model search, and automatic quantization choices mean a first local model runs within minutes without touching a terminal or config file.

Flexibility4/5

Serves any GGUF or MLX model through OpenAI- and Anthropic-compatible APIs, SDKs, a CLI, and a headless daemon, with MCP tools and remote access through LM Link, though engine internals are not open for modification.

Performance4/5

MLX on Apple Silicon and llama.cpp elsewhere give fast single-user inference for the hardware, while throughput for many simultaneous requests is well below datacenter engines such as vLLM.

Portability4/5

Runs on all three desktop operating systems, uses open model formats from Hugging Face, and exposes standard APIs, so models and client code move freely to Ollama or vLLM.

About LM Studio

LM Studio is a desktop app for running LLMs on your own computer. It bundles a model browser that searches Hugging Face, a download manager that picks quantizations sized for the machine's memory, a chat interface, and the inference engines needed to run the model, so a user can go from install to a working local model without a terminal. It runs on macOS, Windows, and Linux, using llama.cpp for GGUF models everywhere and Apple's MLX engine on Apple Silicon Macs.

For developers it doubles as a local inference server. Any loaded model can be exposed through OpenAI-compatible and Anthropic-compatible endpoints, so existing SDKs, agent frameworks, and coding tools work by changing the base URL. It also ships a REST API, Python and TypeScript SDKs, structured output and tool-calling support, and an MCP client that lets local models use external tools. The lms CLI and the headless llmster daemon run the same engine on servers and CI machines without the GUI.

LM Link connects several machines running LM Studio or llmster over an encrypted Tailscale-based mesh, so a laptop can use models loaded on a GPU workstation elsewhere as if they were local. Bionic, a newer agent mode, adds optional cloud-hosted open models billed by the token from prepaid credits, for tasks that outgrow local hardware.

The app itself is free for home and work use, with an Enterprise plan for centralized deployment controls. It is closed-source, although its SDKs and MLX engine are open source. Compared with Ollama, LM Studio leads with a graphical app and model discovery while Ollama is CLI- and server-first; for serving many concurrent users in production, teams move to vLLM.

Key Features

  • Hugging Face model browser with quantization picks sized for available memory
  • llama.cpp engine for GGUF models and MLX engine on Apple Silicon
  • Local OpenAI-compatible and Anthropic-compatible API server
  • Python and TypeScript SDKs, REST API, and the lms CLI
  • MCP client support for tool use by local models
  • Headless llmster daemon for servers without a GUI
  • LM Link encrypted access to models on other machines

Pros

  • The easiest way for a non-specialist to run a local model: install, pick a model, chat
  • Strong performance on Apple Silicon thanks to the native MLX engine
  • Drop-in local endpoint for tools that already speak the OpenAI or Anthropic API
  • Free for commercial use without a licence fee

Cons

  • The desktop app is closed-source, unlike Ollama or llama.cpp
  • Built for one user or a small team, not high-concurrency production serving
  • Large models need substantial RAM or VRAM, and speed depends heavily on local hardware
  • Enterprise pricing is not published

LM Studio Pricing

Freemium
FreeFree
  • · Full desktop app and local inference for personal and work use
  • · Local API server, SDKs, CLI, and llmster
Bionic cloud modelsContact sales
  • · Optional cloud-hosted open models paid from prepaid credits
  • · Per-token rates vary by model
EnterpriseContact sales
  • · Centralized deployment and admin controls
  • · Custom quote from sales

Tools Related to LM Studio

Works well with LM Studio(4)

Open WebUI can use LM Studio's local OpenAI-compatible server as a backend, adding a multi-user browser chat interface over models loaded in LM Studio.

Qwen's open-weight models, including its coding variants, run locally in LM Studio as GGUF or MLX builds and can be served through its OpenAI-compatible API.

LM Studio downloads quantized Llama builds sized for the machine's memory and runs them locally through llama.cpp, or MLX on Apple Silicon.

Gemma's small open-weight models fit on consumer laptops, and LM Studio downloads and runs them locally with a chat interface and a local API server.

Integrates with LM Studio(3)

LM Studio's built-in model browser searches and downloads GGUF and MLX models directly from the Hugging Face Hub, and Hugging Face model pages offer LM Studio as a local app to open them in.

LiteLLM has an LM Studio provider, so a LiteLLM gateway can route requests to models running in LM Studio's local server alongside hosted APIs.

AnythingLLM lists LM Studio as a built-in LLM and embedding provider, so its workspaces and RAG can run on models loaded in LM Studio's local server.

Alternatives to LM Studio(3)

LM Studio leads with a graphical app for finding, downloading, and chatting with local models; Ollama is open source and CLI- and server-first, which suits scripts and headless machines.

vLLM serves open-weight models to many concurrent users on GPU servers; LM Studio is a desktop app for one person to download and chat with local models. They mark the production and personal ends of self-hosted inference.

LM Studio runs open-weight models privately on your own computer at whatever speed the hardware allows; Groq serves them as a very fast hosted API, with no local setup but with data leaving the machine.

Vendor

EL

Element Labs

Website →

Tags

Self-hostableFree TierMachine LearningCross-platform

Details

Maintained
Yes