Python AI Agent
AdvancedAi AgentsA code-first AI agent in Python: an agent framework for the loop, an API in front, and a database for memory.
Published 2 October 2026 · Last updated 2 October 2026
About Python AI Agent
This stack is for an agent you write yourself: Python code in which a language model decides the next step, calls your tools, reads the results, and repeats until the task is done. An agent framework supplies that loop. LangGraph is the default, and it models the agent as a graph of steps with a typed state, so you control where the model has freedom and where the path is fixed. Pydantic AI, CrewAI, and LangChain fill the same role in different styles, and the rest of the stack stays the same whichever you choose.
FastAPI serves the agent as an HTTP API. It streams tokens and tool events to the caller as they happen, and it starts longer runs in the background so a request doesn't have to stay open for minutes. The same service holds what the agent's tools need: API clients, credentials, and the business logic the model is allowed to call.
PostgreSQL is the agent's memory, not an application database full of users and orders. With LangGraph, each conversation is a thread and a checkpoint is saved after every step, so a run that crashes or waits for a person's approval resumes exactly where it stopped. A second store keeps facts that should outlive one conversation. The same database can hold embeddings through the pgvector extension, which is enough retrieval for many agents before a dedicated vector database is worth running.
The model is a separate choice from the framework. Claude is the default, and OpenAI, Gemini, or DeepSeek is swapped in by changing the model name and the API key, since all four frameworks are model-agnostic. Model usage is also where most of the money goes: the server and database cost a few dollars a month, while token spend grows with every step the agent takes. Common optional extras are tracing with Langfuse or LangSmith to see what each run did, a sandbox such as E2B for code the agent writes, and a vector database once pgvector is no longer enough.
Key Features
- ✓An agent loop on a framework: LangGraph by default, or Pydantic AI, CrewAI, or LangChain
- ✓FastAPI endpoints that stream tokens and tool events while a run is in progress
- ✓With LangGraph, the default framework, run state is checkpointed to PostgreSQL after every step, so a run resumes after a crash or a restart
- ✓Human-in-the-loop pauses: a run can stop for approval and continue later from the same point
- ✓Model-agnostic: Claude by default, with OpenAI, Gemini, or DeepSeek a configuration change away
- ✓Vector search in the same database through the pgvector extension
When to Use Python AI Agent
- →Support and operations agents that look things up and act through your internal APIs
- →Research and analysis agents that search, read sources, and write a report
- →Long-running workflows that pause for a person to approve a step before continuing
- →Multi-agent systems where specialised agents hand work to each other
- →An agent API that a web or mobile app calls as its backend
Pros
- Every step is code you can read, test, and version, with no visual builder between you and the logic
- Each layer swaps on its own: framework, model, database, and host
- Python has the widest choice of agent frameworks, model SDKs, and data libraries
- One PostgreSQL database covers run state, long-term memory, and vector search for many projects
Cons
- No user interface included: this is a backend, and a chat or admin frontend is a separate build
- Agent frameworks change quickly, so upgrades can mean reworking code
- Model spend is hard to predict, because one request can turn into dozens of model calls
- Finding out why an agent gave a wrong answer usually takes a tracing tool, which is one more thing to set up
Database Options for Python AI Agent
The default: a PostgreSQL instance next to the API, on the same server or as the host's own database service. It stores the agent's threads, checkpoints, and long-term memory, and with the pgvector extension installed it also holds embeddings. Most of the load is a small write after each step, so a modest instance goes a long way.
Serverless PostgreSQL that scales to zero when no run is active, which fits an agent that is busy in bursts and idle in between. The free plan covers a prototype, and paid usage is billed at $0.106 per compute-unit hour and $0.35 per GB a month with no minimum. pgvector is available on every plan, and branching gives a test run its own copy of real conversation state. The first query after an idle period waits for the database to wake.
Managed PostgreSQL with a dashboard for browsing the agent's threads and memory tables, and with auth and file storage on hand if the project later grows a user-facing app. pgvector is built in. The free plan pauses a project after a week without activity, so anything long-lived belongs on Pro, from $25 a month. The frameworks connect over the standard connection string and need nothing Supabase-specific.
These are highlighted picks. To see all the tools, check the Databases category.
LLM Options for Python AI Agent
The default. Claude holds up well over long tool-calling runs and follows detailed instructions across many steps, which is most of what an agent does. The current Sonnet model costs $2 per million input tokens and $10 per million output, and Haiku at $1 and $5 suits simple steps. Prompt caching cuts the price of the instructions and tool definitions an agent resends on every step.
GPT models have the broadest support across agent frameworks and third-party tools, and many framework examples are written against them first. The mid-tier model is priced like Claude's, at $2 per million input tokens and $10 per million output, and the smallest tier at $0.10 and $0.50 is a cheap way to run high-volume simple steps.
Gemini's Flash and Flash-Lite models are priced below the mid-tier models from Anthropic and OpenAI, and they take images, audio, and video as input, which helps when the agent reads screenshots or scanned documents. Several of them have a rate-limited free tier in Google AI Studio, enough to build and test an agent before paying anything.
The low-cost choice. DeepSeek's API follows the OpenAI format and supports tool calls, and its Flash model costs $0.15 per million input tokens and $0.60 per million output off-peak, with weekday peak hours billed at double. The weights are open, so the same model can later run on your own hardware or at a hosted inference provider. The first-party API is run by a Chinese company, which rules it out under some data policies.
These are highlighted picks. To see all the tools, check the LLM category.
Hosting Options for Python AI Agent
A container platform that runs the FastAPI service as an always-on process, which suits an agent: runs last minutes, responses stream, and background work continues after the request returns. A PostgreSQL service is added to the same project from a template. The Hobby plan is $5 a month including $5 of usage, and memory is billed at about $10 per GB a month, so a small agent API usually stays near the minimum.
Runs the container on machines you place in regions close to your users or your model provider, billed per second while they run. Machines can stop when idle and start on the next request, which cuts the cost of an agent that is used only now and then, at the price of a short wake-up. A small always-on machine with 1 GB of RAM costs under $10 a month. Fly's Managed Postgres is priced for production, so small projects often pair it with a serverless database.
The same always-on service model with a flat price per instance instead of metered usage: $7 a month for the smallest always-on instance and $25 for one with 2 GB of RAM. Background workers and cron jobs are separate service types, useful when agent runs move off the request path. The free instance sleeps after 15 minutes without traffic and takes about a minute to wake, which is fine for a demo and not for an agent that people wait on.
For teams already on AWS. The container runs on ECS with Fargate, PostgreSQL on RDS, and Claude and other models are available through Amazon Bedrock under the same account and bill. Bedrock AgentCore Runtime is a second route built for agents: it hosts LangGraph or CrewAI code with an isolated session per user and sessions lasting up to eight hours, billed by consumption. Expect more setup than a container platform, and tens of dollars a month once a load balancer and a database are running.
A plain VPS for running everything yourself: the FastAPI container, PostgreSQL, and anything else the agent needs, on one machine. A server with 4 GB of RAM costs around €5 to €6 a month, the lowest fixed price here and enough for the API and the database together. You handle deployment, TLS, backups, and updates, and there is no managed database unless you bring one from another provider.
These are highlighted picks. To see all the tools, check the Hosting & Cloud category.
Agent Framework Options for Python AI Agent
The default. LangGraph models the agent as a graph: nodes are steps, edges decide what runs next, and a typed state object passes through all of them. You choose where the model is free to pick the next tool and where the path is fixed. Its Postgres checkpointer saves state after every step, which is what makes resuming, approval pauses, and replaying a past run work. It asks for more code up front than the other three.
Built by the Pydantic team, so it feels like FastAPI: tools are typed Python functions, dependencies are injected, and the model's output is validated against a schema and requested again when it doesn't fit. Pick it when the agent is one model with a set of tools and structured results matter more than a custom control flow. Runs that must survive a restart go through one of its durable execution integrations, such as Temporal or DBOS.
Organises the work as a crew: several agents, each with a role, a goal, and its own tools, that pass tasks between them. Flows wrap those crews in event-driven steps when the order has to be fixed. It suits jobs that split naturally into roles, such as a researcher, a writer, and a reviewer. Flow state is saved to SQLite by default, so check how persistence fits your deployment before relying on resumed runs.
The higher-level route from the company behind LangGraph. Its create_agent function gives a working tool-calling agent in a few lines and runs on LangGraph underneath, so checkpointing to PostgreSQL works the same way. Prebuilt middleware covers common needs such as summarising long conversations, retries, and human approval. Start here for a standard agent, and drop down to LangGraph when you need to draw the control flow yourself.
These are highlighted picks. To see all the tools, check the Agent Frameworks category.
Python AI Agent Add-ons
Each addition below extends this stack with a capability the base stack works fine without. None are required: include the ones your product actually needs when building this stack, and skip the rest.
Vector Db Add-ons
Add a vector database when the agent needs to find documents, notes, or past conversations by meaning instead of exact keywords.
An open-source vector database that runs as one container beside the API or as a managed cluster, with a free small cluster on Qdrant Cloud for testing. It filters on metadata during the search itself, which matters when an agent must only see one customer's documents. Reach for it when the collection grows into the millions of vectors, or when pgvector queries on the main PostgreSQL database start to slow the agent down.
A fully managed, serverless vector database with nothing to run: you create an index and pay for storage, reads, and writes. Hosted embedding and reranking models are part of the service, so the agent needs no separate embedding provider. The Starter plan is free for prototypes. It is closed source and cloud only, so it fits a team that would rather not operate a database, and not one whose data has to stay on its own servers.
The simplest way to try retrieval. Chroma runs in-process or as a single local server, embeds documents for you, and needs a few lines of Python, which is why so many agent tutorials start with it. Chroma Cloud is the managed version with the same API and usage-based pricing. A good first step when you want a separate store but aren't ready to size and run one.
Llm Observability Add-ons
Add LLM observability when you want to see every model call, tool call, and token cost inside a run, so a wrong answer can be traced to the step that caused it.
Open-source tracing and evaluation with integrations for LangGraph, LangChain, Pydantic AI, and CrewAI, so every model call and tool call in a run shows up as a nested trace with its cost and latency. The cloud Hobby plan is free up to 50,000 units a month, and Core is $29. It can also be self-hosted for free, which keeps prompts and outputs on your own infrastructure.
LangChain's own platform, and the closest fit when the agent is built on LangGraph or LangChain: tracing is switched on with environment variables, and Studio steps through a graph visually. The Developer plan is free for one seat with 5,000 traces a month, and Plus is $39 per seat. It is a hosted, closed-source service, and self-hosting is limited to the Enterprise plan.
Code Sandbox Add-ons
Add a code sandbox when the agent runs code it wrote itself, and that code should execute in an isolated environment away from your server and its credentials.
Sandboxes made for running code a model wrote: each one is a Firecracker microVM with its own kernel, started in a fraction of a second from the Python SDK. The Code Interpreter SDK returns output, errors, and charts to the agent, which covers the usual data-analysis tool. The Hobby plan includes free credits and sessions up to an hour, and Pro is $150 a month for sessions up to 24 hours.
Sandboxes that start in under 100 milliseconds and can stay alive with no time limit, so an agent can keep a working directory, a Git checkout, or a long process between steps. They are built from any Docker image, with file, Git, and process APIs in the Python SDK. New accounts get $200 in compute credits, then billing is per second. The platform has been closed source since mid-2026.
A serverless compute platform whose Sandboxes run generated code in gVisor-isolated containers, defined and launched from Python. Pick it when the agent's code needs a GPU, or when the same project also runs batch jobs or serves a model, since all of it shares one platform. The Starter plan includes $30 of compute a month, and usage is billed per second.
These are highlighted picks. To see all the tools, check the Agent Sandboxes & Code Execution category.
Model Inference Add-ons
Add model inference when you want an open-weight model in the mix: on your own hardware for privacy, or on a hosted inference provider for speed and low per-token prices.
A hosted inference provider that runs open-weight models such as Llama, Qwen, and Kimi on its own LPU chips, with very fast output and an OpenAI-compatible API. Speed matters for agents because every step waits on a model response. Billing is per token, with a free rate-limited tier. Other hosted providers, such as Together AI and Fireworks AI, connect the same way.
The hub where open-weight models are published, with two ways to run them. Inference Providers routes one API key to many hosted providers at their own prices, and Inference Endpoints gives a model its own dedicated GPU server billed by the hour. Pick it when the model you want is a specific fine-tune from the Hub, not one of the popular models every provider carries.
Runs an open-weight model on your own machine with one command and serves it on an OpenAI-compatible API, so the agent talks to a local address instead of a paid endpoint. Good for development without token costs and for data that must not leave the server. The model has to support tool calling, and small local models are weaker at multi-step tool use than the large hosted ones, so test the agent's hardest task first.
The serving engine for running an open-weight model on your own GPUs for many users at once. It batches concurrent requests and exposes an OpenAI-compatible API, so the agent code doesn't change. It earns its place when volume makes per-token pricing more expensive than a GPU server, or when policy requires the model inside your own cloud. It needs a datacenter-class GPU and someone to operate it.
These are highlighted picks. To see all the tools, check the AI Runtime & Serving category.
Model Aggregator Add-ons
Add a model aggregator when you want one API key and one bill for models from many providers, with automatic fallback when one of them is down, instead of setting up each provider separately.
One API key and one prepaid balance for hundreds of models from many providers, behind an OpenAI-compatible endpoint that every framework here can call. Useful for comparing models on the same agent without opening an account with each provider, and for automatic fallback when one provider is down. Models are billed at the providers' prices, with a 5.5% fee when you buy credits.
An open-source gateway you host yourself, placed between the agent and the model providers. It gives one OpenAI-compatible endpoint, budgets and spend tracking per key, and fallback rules, while requests still go from your server to each provider under your own API keys. It stores keys and spend in PostgreSQL, which this stack already runs, and it is one more service to deploy and keep running.
These are highlighted picks. To see all the tools, check the AI Model Aggregators category.
CI/CD Add-ons
Add CI/CD when you want a dedicated pipeline for running tests, linting, or multi-stage builds before a deploy goes out. Many hosting platforms already redeploy automatically on every push on their own — a CI/CD tool adds the most value on top of that by gating the deploy on a passing test suite, and matters even more when the hosting choice does not auto-deploy at all, such as a self-hosted server.
Runs the tests on every push and deploys the container to the chosen host. For an agent the useful extra is an evaluation step: replay a set of saved inputs against the new prompt or model, and fail the build when the answers get worse.
These are highlighted picks. To see all the tools, check the CI/CD Pipelines category.
Containerization Add-ons
Add containerization when you want the app packaged the same way across local development, staging, and production, or need to deploy somewhere that isn't a managed serverless platform.
These are highlighted picks. To see all the tools, check the Containerization category.
Frequently Asked Questions about Python AI Agent
Which agent framework should I start with?
Start from the shape of the agent. If it is one model with a set of tools and you care about typed, validated output, Pydantic AI is the least code and feels like FastAPI. If the run has branches, loops, approval steps, or several agents handing off work, LangGraph gives you explicit control and saves state to PostgreSQL after every step. LangChain's create_agent is the quick route to a standard tool-calling agent and runs on LangGraph underneath, so you can move down a level later without changing libraries. CrewAI fits work that divides into roles, such as a researcher and a writer. Switching later means rewriting the agent code, but the API, the database, and the host stay as they are.
Do I need a vector database, or is PostgreSQL enough?
For most agents PostgreSQL is enough to begin with. The pgvector extension stores embeddings in the database this stack already runs, and with an index it works well into the millions of vectors. That keeps retrieval, run state, and your own tables in one place with one backup. Add a dedicated vector database when the collection is much larger, when you filter heavily on metadata at high query rates, or when search load starts competing with the agent's own writes. Qdrant is the usual self-hosted choice, Pinecone the fully managed one, and Chroma the simplest to try. If searching documents is the whole job, the Python RAG App stack is built around ingestion and retrieval, with the vector store as a decision you make up front.
What drives the model bill, and how do I keep it down?
Token spend is the bill that grows, not hosting. An agent resends its instructions, its tool definitions, and the conversation so far on every step, so a ten-step run costs far more than ten single questions. Three things help most. Prompt caching, which Claude and OpenAI both discount heavily for repeated input. Routing simple steps to a small model and keeping the larger one for planning. And a cap on the number of steps a run may take. If you add tracing with Langfuse or LangSmith, it shows the cost of each run and which step is the expensive one. The Pricing section has the per-token rates.
How is this different from the TypeScript AI Agent stack?
The TypeScript AI Agent stack puts the agent and its chat interface in one Next.js app, which is the faster route when the product is a conversation in a browser. This stack is a backend only: FastAPI serves the agent as an API, and any interface is a separate project that calls it. In exchange, Python has the wider choice of agent frameworks, evaluation tools, and data libraries, and a long-lived FastAPI process suits background agents, heavy retrieval, and data analysis better than serverless functions do. Both stacks use PostgreSQL and the same models, so the choice is mostly about language and where the interface lives.
Can I add a frontend to this stack?
Yes. The stack is an API, so the interface is a separate piece that calls it, and there are three common routes. A single-page app in React or Vue talks to the FastAPI endpoints, which is the shape of the Python Web (FastAPI + React) and Vue + FastAPI stacks. A Next.js app can render the chat with ready-made hooks while the agent stays in Python, because Pydantic AI can stream in the format those hooks read. Or the same FastAPI app can serve HTML templates with HTMX for a simple internal screen. If the chat interface is the product and you would rather write everything in one language, the TypeScript AI Agent stack includes it.
Stacks Related to Python AI Agent
Python RAG App
ProjectAnswer questions from your own documents: ingest, index, search by meaning, then let a model write the reply.
Svelte + FastAPI
ProjectSvelte SPA frontend with FastAPI backend: minimal JavaScript output meets Python API performance.
FastAPI Backend
ProjectHigh-performance Python REST API with automatic OpenAPI docs and PostgreSQL.
Python Web (FastAPI + React)
ProjectFastAPI backend with React frontend for Python-first web applications.
Scores
Popularity4/5
Python is the main language for agent development, and LangGraph and LangChain are among the most used frameworks in it, with FastAPI and PostgreSQL as common companions. Building agents in code is still a specialised activity next to mainstream web development.
Learning Curve4/5
FastAPI and PostgreSQL are familiar ground for a Python developer, but the agent layer is not: graph state, checkpoints, tool schemas, prompt design, and evaluation all have to be learned, and the frameworks' APIs are still moving.
Flexibility5/5
Every layer is code and every layer swaps independently: four agent frameworks, four model providers, three ways to run PostgreSQL, and hosts from a container platform to a plain VPS. Nothing about the agent's logic is fixed by a visual builder or a vendor runtime.
Performance3/5
FastAPI's async handlers and PostgreSQL add little overhead, but response time is set by the model: each step waits on a model call, and a run takes several. Streaming keeps the interface responsive without making the run itself faster.
Portability4/5
The frameworks, FastAPI, and PostgreSQL are open source and run on any host that takes a container, and the model is reached through an API that can be pointed at another provider. The agent code itself is written against one framework, so changing framework means a rewrite of that layer.
Tools in the Python AI Agent Stack
Backend Frameworks
Programming Languages
Add-ons (optional — add any, or none)
Vector Db
Llm Observability
Code Sandbox
Model Inference
Model Aggregator
CI/CD
Containerization
Python AI Agent Pricing
The agent frameworks, FastAPI, and PostgreSQL are open source and free. A small always-on server with its database costs roughly $5 to $15 a month on a container platform or a VPS. The model is the line that grows: mid-tier models from Anthropic and OpenAI cost $2 per million input tokens and $10 per million output, small models a tenth of that or less, and an agent makes many calls per task. Tracing, sandboxes, and a vector database are optional and have free tiers.
LangGraph, Pydantic AI, CrewAI, LangChain, and FastAPI are free to use.
About $2 per million input tokens and $10 per million output on mid-tier models; small models from $0.10 and $0.50. Usually the largest line.
An always-on container or a small server. A free instance that sleeps when idle is enough for a demo.
Free on the same server or on a serverless free plan; managed plans run from usage-based to $25 a month.