Hermes Agent Self-Hosted
IntermediateAi AgentsA self-improving AI agent on your own server, with its skills and memory kept in a folder you own.
Published 2 October 2026
About Hermes Agent Self-Hosted
Hermes Agent is an open-source agent from Nous Research built around a learning loop. After a difficult task it writes down what it worked out as a reusable skill, improves that skill the next time it is used, and keeps notes about you and your projects between conversations. Self-hosting means that everything it learns, the skills, the memory files, and the full conversation history, builds up in one folder on a machine you control. You reach it from the messaging apps you already use, including Telegram, Discord, Slack, WhatsApp, and Signal, or from a terminal.
The official image runs two supervised services and no database: the gateway, which holds the connections to your chat apps and runs scheduled jobs, and a web dashboard for configuration. State is a set of files and one SQLite database in a single mounted folder, so the image is disposable and an upgrade is a pull and a restart. The application code inside the container is read-only, which keeps the agent's self-improvement to its skills, memory, and settings: it cannot rewrite the program it runs on.
The dashboard stays on the server's loopback address, because it can read your API keys and run agent commands. You open it through an SSH tunnel, so by default there is no reverse proxy to set up and no TLS certificate to manage, and the common chat platforms connect outbound. Tailscale can be added to reach the dashboard from your own devices, and Cloudflare Tunnel when a channel that delivers messages by webhook needs a public address. Publishing the dashboard on a domain is optional too, through a reverse proxy such as Traefik or a self-hosted PaaS with a one-click template, and Hermes then insists on a login: a dashboard bound beyond loopback refuses to start without one.
The model is the main running cost and a separate choice from the agent. DeepSeek is the default here, in keeping with Hermes' pitch of a capable agent on a small budget, and GLM, MiniMax, Claude, or OpenAI replaces it with one command and no restart. A model gateway can do that job too: OpenRouter is Hermes' default provider, and Nous Research's own subscription, Nous Portal, is the route its docs recommend, with hundreds of models under one login. Other optional pieces are a local model through Ollama or LM Studio on hardware that can run one, and a cloud sandbox such as Daytona or Modal, so the commands the agent runs execute away from the server.
Key Features
- ✓Hermes Agent in Docker from the official image: a gateway and a web dashboard, with no database
- ✓A learning loop that turns solved problems into reusable skills and maintains them over time
- ✓Reached from Telegram, Discord, Slack, WhatsApp, Signal, email, and a terminal
- ✓State as files and one SQLite database in a single mounted folder
- ✓A model of your choice: DeepSeek by default, or GLM, MiniMax, Claude, or OpenAI, switched without a restart
- ✓Approval prompts for dangerous commands and pairing codes for unknown senders
- ✓A backup command that produces a consistent archive while the agent keeps running
When to Use Hermes Agent Self-Hosted
- →An assistant that gets better at your recurring tasks the longer it runs
- →Scheduled work delivered to a chat: daily briefings, monitoring, reports
- →A shared assistant for a team in Telegram, Slack, or Discord, with an allowlist of who may use it
- →Keeping an agent's memory, skills, and credentials on your own machine instead of a hosted service
- →Running an always-on agent on a low-cost model to keep the monthly bill small
Pros
- Free and MIT-licensed, with no paid tier required: the cost is a small server and model usage
- What the agent learns stays with you, as readable files you can inspect, edit, and back up
- Runs on the smallest VPS plans or a single-board computer, because the model runs elsewhere
- The model, the host, and where the agent's commands execute each change without touching the rest
Cons
- Updates, backups, and server security are yours to handle
- Skills the agent writes for itself change its behaviour over time, which makes it less predictable than a fixed setup
- The dashboard holds your API keys and can run commands, so a mistake in exposing it is serious
- A local model needs a context window of at least 64,000 tokens, which rules out small machines
LLM Options for Hermes Agent Self-Hosted
The default. DeepSeek is a built-in Hermes provider set up with one API key, and its Flash model costs $0.15 per million input tokens and $0.60 per million output off-peak, with weekday peak hours billed at double. At those prices an agent that works through the day, background skill reviews included, costs a few dollars a month. The weights are open, so the same model can later run on your own hardware. The API is operated by a Chinese company, which rules it out under some data policies.
Z.ai's GLM models, a built-in provider that Hermes' docs list among its first-class options. GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output and has a 1M-token context, and the Flash version costs $0.15 and $0.50 and accepts images and files. Pick it as a second low-cost provider for Hermes to fall back on, or for long tasks that fill the context window.
MiniMax's models connect with an API key or through a browser sign-in to a MiniMax account, in which case no key is stored on the server. On the API, MiniMax-M3 costs $0.30 per million input tokens and $1.20 per million output. MiniMax also sells a flat Token Plan from $22 a month with rolling quotas, which puts a ceiling on what an always-on agent spends. The weights are open as well.
The frontier choice, for when task quality matters more than cost. Hermes calls Claude through the Anthropic API with a key, at $2 per million input tokens and $10 per million output on the current Sonnet model. Signing in with a Claude subscription is narrower than it looks: Hermes' docs say it works only on a Claude Max plan with purchased extra-usage credits, and that Claude Pro cannot be used at all. Plan on the API key.
GPT models, by API key or by signing in with a ChatGPT or Codex subscription, which Hermes supports through a device-code login and which runs the Codex models. On the API the mid-tier model costs $2 per million input tokens and $10 per million output, and the smallest $0.10 and $0.50. Hermes' docs note that how subscription use counts against plan limits isn't documented, so watch the quota at first.
These are highlighted picks. To see all the tools, check the LLM category.
Hosting Options for Hermes Agent Self-Hosted
A plain root VPS. Hermes publishes no host-specific guides, so this is the general Docker route: pull the official image, mount one data folder, and start the gateway. A server with 4 GB of RAM costs around €5 to €6 a month and sits at the top of Hermes' recommended range, which covers browser automation too. One warning from its Docker guide applies here by name: a provider's browser console can corrupt pasted commands, so connect over SSH to run them.
The low-effort VPS. Hostinger's catalog has a one-click Hermes Agent template that deploys the Docker container and puts logs, restarts, and updates in its Docker Manager. The entry KVM 1 plan has 4 GB of RAM at an introductory $6.49 a month, renewing at $11.99, and the 8 GB plan is $8.99. You keep root access, so firewall rules and copies of the data folder are still your job.
A cloud VPS billed per second up to a monthly cap, with a 1-Click Hermes Agent Droplet in its Marketplace that boots Ubuntu with the agent preinstalled. The 2 GB Droplet at $12 a month meets the recommended memory, and the 1 GB one at about $6 is enough only without browser tools. As with any marketplace image, check which version it ships and what its firewall allows before putting keys on it. Snapshots give a whole-server backup from the control panel.
Your own hardware, for the agent only. The official image is built and tested for arm64, so the same Docker setup runs on a Pi 4 or 5 with a 64-bit OS, and a board with 4 GB leaves room for browser tools. The model stays at a provider. A Pi is far too slow to run a model with the 64,000-token context Hermes needs, so Ollama and LM Studio belong on a stronger machine. A Pi 5 with 4 GB costs about $110 once, plus a few dollars of electricity a month. Put the data folder on an SSD, since sessions and skills are written to it all day.
These are highlighted picks. To see all the tools, check the Hosting & Cloud category.
Hermes Agent Self-Hosted Add-ons
Each addition below extends this stack with a capability the base stack works fine without. None are required: include the ones your product actually needs when building this stack, and skip the rest.
Code Sandbox Add-ons
Add a code sandbox when the agent runs code it wrote itself, and that code should execute in an isolated environment away from your server and its credentials.
Hermes can run the agent's shell commands in a Daytona workspace in place of the server, set with one line of configuration and an API key. The sandbox is stopped when idle and resumed on the next session, so installed tools and files persist while you pay only for running time. That keeps untrusted commands away from the machine that holds your keys. Disk is capped at 10 GiB. New accounts get $200 in compute credits, then billing is per second.
The same idea on Modal: each task gets an isolated cloud sandbox with the CPU, memory, and disk you set, and its filesystem is snapshotted on cleanup and restored the next time. Choose it when the work needs more compute than the server has. Modal's Starter plan includes $30 of compute a month. Nous Portal subscribers can also use Modal through Nous's managed gateway, with no Modal account of their own.
These are highlighted picks. To see all the tools, check the Agent Sandboxes & Code Execution category.
Model Inference Add-ons
Add model inference when you want an open-weight model in the mix: on your own hardware for privacy, or on a hosted inference provider for speed and low per-token prices.
Runs open-weight models and serves them on an OpenAI-compatible address that Hermes uses as a custom endpoint. One setting matters: Ollama starts models with a small context window by default, and Hermes refuses anything under 64,000 tokens, so the context length has to be raised. That in turn needs a GPU with plenty of memory or a recent Mac, which is why the model belongs on a stronger machine than a small VPS or a Raspberry Pi, with the agent pointing at it over the network.
A desktop app with a graphical model browser and a local server on port 1234. It has its own entry in Hermes' model picker, which lists the models LM Studio has downloaded. Hermes' docs advise a context length of at least 64,000 tokens and, when the machine can't hold that, a smaller model over a shorter context. It suits an agent on a home Mac or PC, or a desktop on the same network as the server. Free for personal and work use.
These are highlighted picks. To see all the tools, check the AI Runtime & Serving category.
Model Aggregator Add-ons
Add a model aggregator when you want one API key and one bill for models from many providers, with automatic fallback when one of them is down, instead of setting up each provider separately.
Hermes' default provider: one key, or a browser sign-in, for hundreds of models from many labs. It is the easy way to try several models on the same agent, and Hermes can pass routing preferences along, such as choosing the cheapest provider for a model. Models are billed at provider prices with a 5.5% fee on credit purchases. Nous Research's own Nous Portal subscription does a similar job and adds web search and other tools under the same login.
A gateway you run yourself, as one more container on the same Docker network. Hermes connects to it as a custom endpoint on port 4000, and LiteLLM holds the provider keys, balances load, and falls back between models from one config file. Its budget limits per key are a useful brake on an agent that runs unattended. The open-source edition is free, and spend tracking needs a PostgreSQL database.
These are highlighted picks. To see all the tools, check the AI Model Aggregators category.
Tunnel Add-ons
Add a tunnel when you're self-hosting without a static IP or can't open inbound ports — a home server, a VPS behind restrictive network policies, or anywhere a reverse proxy alone can't reach the internet.
Private access to the dashboard from your own devices. Hermes' docs call it the clean option: bind the dashboard to the server's Tailscale address, switch on the built-in username and password login, and only devices on your tailnet can reach it. Tailscale Serve works as well and adds HTTPS on a tailnet hostname. The desktop app can then drive the agent on the server as a remote gateway. The Personal plan is free.
For channels that deliver messages by webhook. The official WhatsApp Cloud API, Microsoft Teams, SMS, and LINE all need a public HTTPS address to call, and Hermes' docs recommend Cloudflare Tunnel for it: free, no port forwarding, and no inbound firewall rule. Telegram, Discord, and the unofficial WhatsApp bridge work without it. Keep the tunnel to the webhook port, since the docs say never to put a password-protected dashboard on the open internet.
Reverse Proxy Add-ons
Add a reverse proxy when the service should be reachable at its own web address: it terminates HTTPS on your domain and can put a login in front of an interface that is otherwise kept private.
The proxy Hermes' Docker guide names for running in a second container: it routes to the dashboard from Docker labels and renews certificates itself. Hermes does not trust a proxy on the Docker network by default, so the dashboard's public URL and the proxy's exact address go into its configuration, and any bind beyond loopback needs a login provider or the dashboard will not start. It is also what Coolify and Dokploy run underneath.
Automatic HTTPS from a two-line config, the least work for one dashboard on one domain. The dashboard's live chat runs over a WebSocket, which Caddy passes without extra settings, and Hermes sends a keepalive every 20 seconds so an idle chat isn't cut off. For an address on the open internet, Hermes' docs recommend its OAuth sign-in over the built-in username and password, which they describe as meant for a trusted network.
The proxy many servers already run, with certificates through Certbot. It needs the WebSocket upgrade headers for the dashboard's chat and terminal, and Hermes' docs single out manual NGINX setups as the case where forwarded headers go missing, which is fixed by setting the dashboard's full public URL in the configuration. The login rule is the same as with any proxy here: a provider has to be configured before the dashboard will listen beyond loopback.
Self Hosted Paas Add-ons
Add a self-hosted PaaS when you would rather deploy and update from a dashboard than from the command line: it installs the service from a template and brings its own reverse proxy and certificates.
A free, self-hosted deploy dashboard with two Hermes Agent services in its catalog, both on the official image: one publishes Hermes' own dashboard, the other a community web chat. Each gets a domain with HTTPS through Coolify's proxy and a generated password, and provider keys go in as environment variables. Check three defaults before real use. The dashboard is protected by Hermes' username and password login, which its docs say belongs behind a VPN and not on the open internet, where the OAuth sign-in is the one to use. The template lets anyone who messages the bot use it, while the docs call for an allowlist or pairing in production. And it pins an image version that may be behind.
Free and self-hosted, with a Hermes template that runs the official image behind Dokploy's bundled Traefik: the dashboard gets a domain, HTTPS, and a generated password through Hermes' built-in login. Hermes' docs keep that login for a trusted network, so for an address on the open internet switch it to the OAuth sign-in. The template also turns on the OpenAI-compatible API and publishes it on a second domain behind a key, which is only useful when other software calls the agent, so remove that domain otherwise. It pins an image version and sets a 4 GB memory limit, so review both. No separate reverse proxy is added next to it.
Frequently Asked Questions about Hermes Agent Self-Hosted
Why self-host Hermes Agent instead of using Hermes Cloud?
Hermes Cloud is Nous Research's hosted version: a dedicated, always-on instance with its own workspace and dashboard, deployed in one click. A Medium instance with 2 GB of RAM costs $0.56 a day while running, about $17 a month, and a Large one $1.09 a day, with model and tool usage billed on top. A small VPS that runs the same agent costs $5 to $7 a month, so self-hosting is cheaper, and the skills, memory, and credentials stay on a machine you control. You also get the full choice of where commands run, local models, and your own network. What Hermes Cloud buys is no server to patch, no backups to arrange, and no exposure mistakes to make. If you would rather not run Linux, it is the better choice, and the software is the same either way.
What size server does Hermes Agent need?
Hermes' Docker guide gives 1 GB of memory, one core, and 500 MB of disk as the minimum, and recommends 2 to 4 GB, two cores, and 2 GB or more of disk. Browser automation decides it: without browser tools 1 GB is sufficient, and with them the guide says to allow at least 2 GB. The data folder grows with sessions and skills, so leave room. None of this includes the model, which runs at the provider. A local model is a separate machine with a GPU or a large-memory Mac, because Hermes needs a 64,000-token context window, and a small VPS or a Raspberry Pi can't serve one at a usable speed.
What needs backing up, and how do updates work?
One folder holds everything: settings, API keys, the agent's identity file, memories, skills, scheduled jobs, and the session database. The backup command writes a zip of it that is consistent even while the agent is running, credentials included, so store it as a secret. Restoring on a new machine is one import command. Updating is a pull of the new image and a restart. The container migrates the configuration on startup and first saves timestamped copies of the settings and key files. Two cautions from the docs: never run two gateway containers against the same data folder, and back up before switching release channels, because a newer version can change the data format and switching back does not undo it.
Do I need a reverse proxy or Coolify to run Hermes Agent?
No. The agent works through your chat apps over outbound connections, and its dashboard is a settings screen that stays on the server's loopback address, opened through an SSH tunnel. That is how the official Compose file ships it. Tailscale is the simplest step up when you want the dashboard on your phone. A reverse proxy or a PaaS is for a dashboard on its own domain, for example to use the desktop app against a remote server or to share the agent with a team. Hermes requires a login for that, and its docs are direct about the limits: the username and password login is for a trusted network, and anything on the open internet should use its OAuth sign-in. Coolify and Dokploy both have Hermes templates on the official image, and both publish the dashboard behind that username and password login, so change it before leaving the address open to the internet. Coolify's also lets every chat user in by default, which should be changed too.
Hermes Agent or OpenClaw: which should I self-host?
Both are free, MIT-licensed agents that live on a small server and answer in your chat apps, and each is the other's closest alternative. Hermes Agent's difference is the learning loop: it saves what it works out as skills, maintains them in the background, and builds a profile of how you work, so it suits someone who will hand it the same kinds of task for months. It also has a first-party hosted version, Hermes Cloud, and can send its commands to cloud sandboxes. OpenClaw has the larger community and more hosting guides, and its abilities are installed and configured, not learned, which makes it more predictable. Hermes ships a command that imports an existing OpenClaw setup, so trying it after OpenClaw doesn't mean starting over.
Stacks Related to Hermes Agent Self-Hosted
Strapi Self-Hosted
InfrastructureSelf-hosted Strapi: an open-source headless CMS with PostgreSQL, on a server you control.
Grafana Self-Hosted
InfrastructureSelf-hosted Grafana and Prometheus: metrics dashboards and alerting on a server you run.
Airflow Self-Hosted
InfrastructureSelf-hosted Apache Airflow: the standard data-pipeline scheduler on your infrastructure.
n8n Self-Hosted
InfrastructureSelf-hosted n8n on your own server, with full control over the database, the host, and how it's exposed to the internet.
Scores
Popularity5/5
Hermes Agent passed 250,000 GitHub stars within months of its 2026 launch and is among the best-known self-hosted personal agents, with one-click templates from several hosts and deploy platforms.
Learning Curve3/5
Chatting with the agent needs nothing beyond a messaging app, and setup is a guided wizard. Running it yourself adds Docker, an SSH tunnel, a model key, and channel setup, and getting value from it means learning how its skills and memory change over time.
Flexibility5/5
Dozens of model providers, seven places its commands can execute, a wide range of messaging channels, installable and self-written skills, plugins, and scheduled jobs. The host, the model, and the terminal backend each swap independently.
Performance4/5
The gateway is a light process that fits in 1 GB of memory without browser tools and idles between messages. How quickly and how well it completes a task depends on the model behind it and on the skills it has built up.
Portability5/5
MIT-licensed, with skills, memory, and sessions stored as files and SQLite in one folder that a single backup command moves to any Linux host. Skills follow an open format, and the model is a configuration value.
Tools in the Hermes Agent Self-Hosted Stack
DevOps & CI/CD
Agentic AI
Add-ons (optional — add any, or none)
Code Sandbox
Model Inference
Model Aggregator
Tunnel
Reverse Proxy
Self Hosted Paas
Hermes Agent Self-Hosted Pricing
Hermes Agent and Docker are free, so the fixed cost is the server: $5 to $12 a month for a small VPS, or a one-time purchase for hardware at home. Model usage is the part that varies. On the default low-cost model it is $0.15 per million input tokens and $0.60 per million output, a few dollars a month for regular use, and about $2 and $10 on the frontier models. Hermes Cloud, the hosted alternative, is about $17 a month for a comparable instance before model usage.
1 GB of RAM is the minimum and 2 to 4 GB is recommended, with at least 2 GB when browser tools are in use. Hardware at home is a one-time cost plus electricity.
Open source, with no paid tier needed to self-host. Docker is free as well.
From $0.15 per million input tokens and $0.60 per million output on the default model to about $2 and $10 on frontier models. Usually the largest line.
An SSH tunnel costs nothing, and Tailscale's Personal plan and Cloudflare Tunnel are free. The reverse proxies and the self-hosted deploy platforms are free, open-source software as well.
Hermes Agent Self-Hosted System Requirements
source- CPU
- 1 core minimum; 2 cores recommended
- RAM
- 1 GB minimum; 2 to 4 GB recommended
- Disk
- 500 MB minimum for the data folder; 2 GB or more recommended
- OS
- Any Linux with Docker, on amd64 or arm64
From the resource limits table in Hermes' Docker guide, which recommends 2 to 4 GB of memory. 1 GB is sufficient without browser tools, and the guide says to allow at least 2 GB with them. The data folder grows with sessions and skills. The model runs at the provider and isn't part of these figures; a local model needs separate, much larger hardware.