Local LLM Dev Stack
AdvancedA coding agent on a model that runs on your own machine, with no code sent to a provider.
Published 2 October 2026
About Local LLM Dev Stack
This stack runs a coding agent against a model on your own hardware. A local runtime serves an open-weight model on the machine, and the agent talks to that address in place of a provider's API, so prompts, code, and file contents never leave it. There is no subscription and no per-token bill, it works offline, and nothing changes when a vendor changes its prices or its rules. The price is quality and speed: an open model that fits one graphics card is weaker on long, multi-file tasks than the hosted models these agents were designed around.
The runtime is the piece you operate. Ollama is the default: a background service with a model library, and a launch command that starts OpenCode, Claude Code, or Codex already pointed at a local model. LM Studio does the same job from a desktop app and uses Apple's MLX engine on Apple Silicon Macs, llama.cpp is the bare engine for people who want control over every setting, and Unsloth also fine-tunes the models it runs. All four expose the same kind of local server, so the agent does not care which one is behind it.
OpenCode is the default agent because it is open source, tied to no model vendor, and documents these runtimes itself. Cline, Qwen Code, Codex, and Claude Code work as well, each through a setting or an endpoint that its vendor or the runtime documents, and several agents can share one local model. The model matters more than the agent. Qwen's coding models are the default, with Mistral's Devstral, Google Gemma, and DeepSeek as alternatives, and all of them need a large context window to work with an agent, far above what the runtimes set out of the box.
By default everything runs on your own computer, which has to stay on while the agent works. A common layout splits it in two: the model on a desktop with a strong graphics card, and the agent on a laptop that reaches it over the home network, or over Tailscale when you are away. OpenCode's web interface serves the same sessions to a phone browser. There is no vendor cloud to hand work to, which is the point of the stack. The model can also live on a rented GPU server, which keeps it available around the clock at a monthly price well above a coding subscription.
Code lives on GitHub by default, or on GitLab when source control has to stay on your own infrastructure too. A model gateway such as OpenRouter can be added as a fallback for the tasks a local model can't finish, and a faster terminal, an AI code reviewer, and CI/CD are common additions.
Key Features
- ✓A coding agent working against a model on your own hardware, with no code sent to a provider
- ✓A choice of local runtime: Ollama by default, or LM Studio, llama.cpp, or Unsloth
- ✓One agent or several on the same model: OpenCode by default, with Cline, Qwen Code, Codex, or Claude Code
- ✓Open coding models that fit one graphics card: Qwen's coder models, Devstral, Gemma, and DeepSeek
- ✓Works offline, with no subscription and no per-token cost
- ✓The model can sit on a second, stronger machine and be reached over your network
- ✓Hosted models through a gateway as an optional fallback for hard tasks
When to Use Local LLM Dev Stack
- →Working on code that is not allowed to leave your machine or your company's network
- →Coding on a plane, a train, or anywhere without a reliable connection
- →Routine edits, tests, and refactors without spending tokens
- →Comparing open coding models side by side on your own repository
- →One shared model machine that several developers' agents connect to
Pros
- Nothing leaves your hardware: prompts, code, and file contents stay local
- No subscription, no usage limits, and no per-token bill
- Every layer is replaceable: the runtime, the model, the agent, and the git host
- The same local model also serves chat interfaces and other tools on your network
Cons
- Open models on one graphics card are weaker than hosted frontier models on long, multi-file tasks
- Real hardware is needed: about 24 GB of graphics memory, or a Mac with 32 GB or more, for the coding models
- The context window has to be raised well above the runtime's default, which costs memory and speed
- No vendor cloud sessions or phone app: away from the desk, a machine you own has to stay on
- Claude Code and Codex reach local models through compatibility layers, so some of their features are missing
Coding Agent Options for Local LLM Dev Stack
OpenCode is the default here: an open-source agent with no model vendor behind it, in the terminal, a desktop app, and a web interface. Its docs cover Ollama, LM Studio, and llama.cpp as providers, and Ollama's launch command sets it up in one step. Ollama's docs put its context need at 64,000 tokens or more, and its plan and build agents can each use a different local model. Free, MIT licensed.
Cline is an agent that lives in VS Code and other editors, with Ollama and LM Studio as built-in providers. Its docs are frank about local use: a compact-prompt setting for small models, and memory tiers from 16 GB for small quantized models to 64 GB and more for large ones. A good pick when you want to watch and approve each edit in the editor. Free and open source, and with a local model there is nothing else to pay for.
Qwen Code is Alibaba's open-source terminal agent, built around the Qwen models that are this stack's default weights. It reaches a local runtime as an OpenAI-compatible provider with a local address, and has a timeout setting for slow local servers. The natural pairing when the model is a Qwen coder, since agent and model come from the same team. Free and open source.
OpenAI's Codex CLI has a mode for open models that points it at a local Ollama or LM Studio server, and both runtimes document the setup. It keeps Codex's approval flow while the model runs on your machine. Codex is built around OpenAI's own models, so a local model gets a prompt tuned for a different family, and it is worth testing on a real task before settling on it. The CLI is free and open source.
Claude Code can run against a local model through the Anthropic-compatible endpoints that Ollama, LM Studio, and llama.cpp provide, and Ollama's launch command or Unsloth's start command sets it up. Anthropic supports Claude Code with Claude models only, so this route belongs to the runtimes: Ollama's docs note that hosted web search and some tool controls are not fully supported. Pick it when you already work in Claude Code and want the same interface offline.
These are highlighted picks. To see all the tools, check the AI Coding Agents category.
LLM Options for Local LLM Dev Stack
Qwen is the default here. The Qwen 3.6 coding build at 27B is an 18 GB download and Qwen3-Coder at 30B is 19 GB, both with a 256K context limit, sized for a 24 GB graphics card or a 32 GB Mac. OpenCode's docs name the Qwen-Coder family as the one to try when local tool calls misbehave. Most releases are Apache 2.0, and every agent in this stack can run them.
Mistral's Devstral models are built for agentic coding. Devstral Small 2, a 24B model, is a 15 GB download with a context limit of 384K tokens, which leaves more room on a 24 GB card than the Qwen coders do. The weights are Apache 2.0 and come from a European lab. A strong second model to keep downloaded for comparison on your own repository.
Google's Gemma 4 is a general model that also codes. Its 26B size is a 16 GB download and 31B is 19 GB, with a 256K context limit, tool calling, and image input, which helps when a task includes a screenshot or a diagram. The 12B model at about 8 GB is the fallback for a graphics card with 12 or 16 GB, with shorter and simpler tasks in mind.
The DeepSeek models in Ollama's library that fit one graphics card are older ones: DeepSeek-Coder V2 at 16B is a 9 GB download, and the distilled DeepSeek-R1 sizes add a visible reasoning step. They suit cards with 12 to 16 GB. DeepSeek also publishes the weights of its current models under the MIT licence, with quantized builds from the community, so check a build's size against your memory before downloading it.
These are highlighted picks. To see all the tools, check the LLM category.
Version Control Options for Local LLM Dev Stack
GitHub is the default here: the largest pull-request ecosystem, and where AI code reviewers and CI run. Keeping the model local does not keep the repository local, so code pushed to GitHub is on GitHub's servers. For most teams that is the existing arrangement and only the model stays private.
GitLab bundles source control, merge requests, and CI/CD in one platform you can host yourself, the pick when the reason for a local model is that code must not leave your infrastructure at all. The agents only need git for local work, so everyday sessions are the same on either host.
Model Inference Options for Local LLM Dev Stack
Ollama is the default here: a background service that downloads models by name and serves them on a local address. Its launch command starts OpenCode, Claude Code, or Codex already configured for a local model, and Cline has it as a built-in provider. Its default context window is small, and its docs recommend 64,000 tokens for coding tools, so that setting has to be raised. Its cloud models use the same setup when local hardware runs short. Free and MIT licensed.
LM Studio is the desktop-app route: a model browser that picks a build sized for your memory, a chat window for testing, and a local server with OpenAI-compatible and Anthropic-compatible endpoints. It documents Claude Code and Codex setups itself, and on Apple Silicon it runs models through Apple's MLX engine. LM Link lets a laptop use models loaded on another machine. Free for personal and work use.
llama.cpp is the engine most local tools are built on, used directly. Its server offers OpenAI-compatible and Anthropic-compatible endpoints, downloads models straight from Hugging Face, and exposes every setting: the quantization, how much of the model sits on the graphics card, the context size. OpenCode documents it as a provider. Pick it for full control and a small footprint, at the cost of managing models yourself. Free, MIT licensed.
Unsloth runs local models from a desktop app or a web interface and serves them on an OpenAI-compatible address, with one command that connects OpenCode, Claude Code, or Codex to the loaded model. What sets it apart is training: the same tool fine-tunes a model on your own code or data and then serves the result. Pick it when a tuned model is where this is heading. Free and open source.
These are highlighted picks. To see all the tools, check the AI Runtime & Serving category.
Local LLM Dev Stack Add-ons
Each addition below extends this stack with a capability the base stack works fine without. None are required: include the ones your product actually needs when building this stack, and skip the rest.
Server Add-ons
Add a server when you want the agent to keep running somewhere other than your own laptop, reachable at any time and not tied to your machine staying on. Some agents also offer their own managed cloud sessions as an alternative to self-hosting; check the stack's own description for details.
Add a Hetzner GPU server when no machine at home can hold the model: the entry dedicated server has a 24 GB NVIDIA card and 64 GB of RAM for about €214 a month plus a one-time setup fee, enough for the 15 to 19 GB coding models. The model then answers around the clock, on hardware that is rented and not local.
Add a DigitalOcean GPU Droplet for occasional sessions: the smallest has a 20 GB NVIDIA card at $0.76 an hour, about $550 if left on for a month, so it pays to destroy it after work and keep the models on a volume. A fit for trying larger models before buying a graphics card.
These are highlighted picks. To see all the tools, check the Hosting & Cloud category.
Remote Access Add-ons
Add remote access when you want to reach an agent running on another machine — a VPS or a home server — without exposing it to anyone but you.
Add Tailscale to reach the model machine from a laptop or phone anywhere: both join a private network, and the agent points at the desktop's or the server's address as if it were local, with nothing exposed to the internet. LM Studio's own LM Link is built on the same kind of mesh. The personal tier is free.
Terminal Add-ons
Add a terminal when you want a faster, more configurable place to run the agent than your OS default — most agent CLIs live here all day.
A GPU-accelerated terminal with native macOS and Linux integration, one of the terminals OpenCode's docs name, and a solid default for keeping an agent and a runtime's logs open side by side.
A minimal, GPU-accelerated terminal that does little beyond drawing text fast; pick it when a multiplexer or window manager already handles tabs and splits, and every bit of graphics memory should go to the model.
A GPU-accelerated terminal with built-in tabs, splits, and multiplexing configured in Lua, able to stand in for both the terminal and tmux, with the agent in one pane and the model server in another.
A GPU-accelerated terminal with its own graphics protocol and a scripting layer called kittens, for developers who want deep customization of the window their terminal agent lives in.
Code Review Add-ons
Add code review when you want an AI reading every pull request before it merges: it comments on the diff so bugs, security issues, and inconsistencies surface before a human has to catch them.
Add CodeRabbit when you want a second opinion on what a local model wrote: it installs as a GitHub or GitLab app, summarizes each pull request, and leaves line-level comments before anyone merges. It is a hosted service, so the diff is sent to it, which matters if the model stayed local for privacy.
Add Greptile when review should read the whole repository and not just the diff: it indexes the codebase so its pull request comments carry wider context. It is a hosted service by default, and its Enterprise plan adds the option to self-host it on your own infrastructure, which suits a setup built to keep code in house.
Model Aggregator Add-ons
Add a model aggregator when you want one API key and one bill for models from many providers, with automatic fallback when one of them is down, instead of setting up each provider separately.
Add OpenRouter as the way out for tasks a local model can't finish: one key for hosted models, including the larger siblings of the open models used here, selected in the agent like any other provider. Those requests do leave your machine, at provider prices plus 5.5% on credit purchases, so keep it for work that is allowed to.
Add LiteLLM when several developers share the setup: a self-hosted proxy that puts the local runtime and any hosted models behind one address, with a key and a budget per person. The agents connect to it as an OpenAI-compatible provider, and nothing leaves your network unless a hosted model is configured behind it. Free to self-host.
These are highlighted picks. To see all the tools, check the AI Model Aggregators category.
CI/CD Add-ons
Add CI/CD when you want a dedicated pipeline for running tests, linting, or multi-stage builds before a deploy goes out. Many hosting platforms already redeploy automatically on every push on their own — a CI/CD tool adds the most value on top of that by gating the deploy on a passing test suite, and matters even more when the hosting choice does not auto-deploy at all, such as a self-hosted server.
Add GitHub Actions when the repository lives on GitHub and every push the agent makes should run the tests. An automatic test run is worth more here than with a hosted model, since a smaller model makes more mistakes that only the tests catch.
These are highlighted picks. To see all the tools, check the CI/CD Pipelines category.
Frequently Asked Questions about Local LLM Dev Stack
Can the agent keep working when my laptop is closed?
Only if the machine doing the work stays on. There is no vendor cloud in this stack: the model and the agent both run on hardware you control. The usual answer is two machines. A desktop with the graphics card runs the model and stays on, and the laptop runs the agent and reaches the model over the home network, or over Tailscale from anywhere. With the laptop closed, the model keeps serving, but the agent session on the laptop stops. For work that continues while you are away, the agent has to run on the always-on machine as well, and OpenCode's web interface then shows its sessions in a phone browser. A rented GPU server does the same job without hardware at home, at a much higher monthly cost.
Which agents work with which runtime?
All five agents work with Ollama, and most work with the others. OpenCode documents Ollama, LM Studio, and llama.cpp as providers. Cline has Ollama and LM Studio built in. Qwen Code takes any runtime that offers an OpenAI-compatible address, which all four do. Codex's open-model mode targets Ollama and LM Studio. Claude Code needs an Anthropic-compatible endpoint, which Ollama, LM Studio, and llama.cpp provide. Ollama and Unsloth each have a single command that starts OpenCode, Claude Code, or Codex already connected, which is also how those three reach Unsloth. Two cautions apply to the vendor agents: Anthropic supports Claude Code with Claude models only, and Codex is tuned for OpenAI's models, so both lose some features and reliability on a local model. The models are a separate choice and every runtime here can load the same open weights.
Is a local model good enough to replace a coding subscription?
For some work, not for all of it. A coding model that fits one graphics card handles contained tasks well: writing a function, adding tests, explaining code, a refactor inside a few files. It falls behind hosted frontier models on long tasks that span many files, where it loses track, repeats a failed edit, or calls a tool wrongly. It is also slower. Most people who run this stack keep both: the local model for routine and private work, and a hosted model for the hard problems, either through a subscription agent or through a gateway from the same agent. If privacy is the reason, the question is different, and a weaker model that keeps the code in house is the right trade.
What hardware do I need, and what does it cost?
The coding models here are 15 to 19 GB downloads, and a model needs about its file size in graphics memory plus room for the context window. In practice that means a graphics card with 24 GB, or an Apple Silicon Mac with 32 GB or more of unified memory. Cline's docs give a similar ladder for system memory: 16 to 32 GB for small or quantized models, 32 to 64 GB for mid-size coding models, and more for large ones. A card with 12 or 16 GB runs smaller models for simpler tasks. The software is all free, so the cost is the hardware and the electricity. Renting is possible, from about €214 a month for a dedicated GPU server. The Pricing section has the details.
How large a context window do I need?
Larger than the runtime gives you by default, and this is the setting that decides whether the stack works at all. A coding agent sends its instructions, its tool definitions, and the files it is reading on every turn. Ollama's docs recommend at least 64,000 tokens for coding tools, while its default on a card with less than 24 GB is about 4,000, at which point the agent's own instructions no longer fit and tool calls fail. LM Studio's docs ask for more than about 25,000 for Claude Code and Codex. Every extra token of context uses memory on top of the model itself, so the real choice is between a larger model with a tight context and a smaller one with room to read. For agent work, the room usually matters more.
Can I run several coding agents at the same time on one local model?
Yes, in two ways. The simple one is separate sessions that you manage yourself: OpenCode in one terminal, Cline in the editor, Claude Code in another terminal, each pointed at the same local address. The other is an agentic dev environment, and how it connects depends on the kind. Herdr and Orca run every agent in a real terminal, so whatever starts an agent on the local model in your own terminal, a runtime's launch command or the agent's own settings, works the same there, and Orca can store that command per agent. T3 Code and Paseo start the agents themselves, so the local address goes into their provider settings as a custom endpoint, a setup both document. The limit is the hardware, not the software. By default Ollama answers one request per model at a time and queues the rest, and each extra parallel slot multiplies the memory the context needs, so on one graphics card the agents mostly take turns.
Stacks Related to Local LLM Dev Stack
Qwen Code Dev Stack
DeveloperAlibaba's open-source coding agent in the terminal, desktop, or browser, running Qwen, GLM, Kimi, or any model you choose.
OpenCode Dev Stack
DeveloperOpenCode in the terminal, desktop app, or browser, running open-weight models like DeepSeek, GLM, and Kimi.
Scores
Popularity3/5
Running open models locally is widely tried and Ollama is among the most-starred AI projects, but for day-to-day coding most developers still use hosted models, and local setups are the choice of the privacy-minded and the curious.
Learning Curve4/5
Beyond learning the agent, there is a runtime to operate, a model to size against the hardware, and a context window to raise before anything works. Ollama's launch command shortens the first run, but getting good results takes tuning.
Flexibility5/5
Four runtimes, five agents, any open-weight model, and two git hosts, each replaceable without touching the rest, with a hosted model available through a gateway when needed.
Performance2/5
A model that fits one graphics card answers more slowly than a hosted one and handles long, multi-file tasks less reliably. Contained tasks go well, and results depend heavily on the hardware and the context size.
Portability5/5
Every part is open source or an open-weight file, the runtimes speak the same local API, and nothing is tied to an account or a plan. The setup moves to another machine by copying models and settings.
Tools in the Local LLM Dev Stack Stack
Local LLM Dev Stack Pricing
Every tool in this stack is free: the runtimes, the agents, and the open-weight models. The cost is hardware. A graphics card with 24 GB of memory, or a Mac with 32 GB or more, runs the coding models, and electricity is the only running cost. Renting a GPU server starts at about €214 a month, or under a dollar an hour for a cloud GPU you switch off between sessions. A hosted model through a gateway is optional and billed per token.
Ollama, llama.cpp, and Unsloth are open source and LM Studio is free to use. OpenCode, Cline, Qwen Code, and the Codex CLI are open source.
Downloaded once, 9 to 19 GB each for the coding models. Licences vary by family, mostly Apache 2.0 or MIT.
About 24 GB of graphics memory, or a Mac with 32 GB or more, for the 15 to 19 GB coding models.
The free tiers cover individual developers; GitHub Team is $4 a month and GitLab Premium $29 per seat.
OpenRouter charges provider prices plus 5.5% on credit purchases; LiteLLM is free to self-host.
A dedicated server with a 24 GB card and 64 GB of RAM at a fixed monthly price, or a cloud GPU with 20 GB billed by the hour.
Tailscale's Personal plan is free.