Python RAG App
AdvancedAi AgentsAnswer questions from your own documents: ingest, index, search by meaning, then let a model write the reply.
Published 2 October 2026
About Python RAG App
This stack is for an app that answers questions from your own documents: a help center, a contract archive, internal wikis, or a folder of PDFs. The model never sees the whole collection. For each question the app finds the few passages that matter and hands only those to the model, which is what retrieval-augmented generation (RAG) means. LlamaIndex is the default framework and is built around that job, with loaders for files and APIs, parsers that split documents into chunks, indexes, and query engines that tie retrieval to a model. LangChain is the alternative when retrieval is one part of a wider agent.
Two steps run before any question arrives. Ingestion reads each document, cuts it into chunks, turns every chunk into an embedding (a list of numbers that places its meaning in space), and writes the result to the vector store. PostgreSQL with the pgvector extension is the default store, and Qdrant, Pinecone, or Chroma take over when the collection or the search load outgrows it. Scanned files and complex tables are the hard part, and LlamaIndex's own LlamaParse service handles them at a per-page price.
At question time, FastAPI embeds the question, searches the store, and streams the answer with the passages it used, so a reader can check the source. Metadata filters keep each user or customer inside their own documents, and hybrid search mixes keyword matching with meaning when exact terms such as product codes matter.
Two models are involved, not one: a language model that writes the answer, and an embedding model that builds the index. Anthropic offers no embedding model, so Claude is paired with another provider's embeddings such as Voyage AI, OpenAI, or Gemini, and changing the embedding model later means re-embedding every document. Common optional extras are tracing to see which passages a bad answer was built from, a gateway in front of the model providers, and a local runtime for private data.
Key Features
- ✓Ingestion with LlamaIndex by default, or LangChain: loaders, chunking, and embeddings written into the vector store
- ✓A vector store you choose: PostgreSQL with pgvector by default, or Qdrant, Pinecone, or Chroma
- ✓FastAPI endpoints that stream the answer together with the source passages it used
- ✓Metadata filtering, so each user or tenant only searches their own documents
- ✓Hybrid search that combines keyword matching with search by meaning
- ✓Model-agnostic: Claude writes the answers by default, with OpenAI or Gemini a configuration change away
When to Use Python RAG App
- →Support assistants that answer from a help center, manuals, and past tickets
- →Internal knowledge search across wikis, policies, and shared drives
- →Document Q&A over contracts, reports, or research papers with cited sources
- →A retrieval layer that an agent or another app calls through an API
- →Search over a product catalogue or a code base by meaning, not exact words
Pros
- Retrieval is the default framework's main job, so ingestion, chunking, and query engines are built in
- One PostgreSQL database can hold documents, metadata, and vectors until the collection grows large
- Every layer swaps on its own: framework, vector store, answer model, and host
- Answers carry their sources, which makes wrong ones easier to catch
Cons
- Answer quality depends on chunking, embeddings, and retrieval settings that need tuning against real questions
- Changing the embedding model means re-embedding the whole collection
- Parsing scanned files and complex tables takes extra tooling and extra cost
- No user interface included: this is a backend, and a search or chat frontend is a separate build
LLM Options for Python RAG App
The default for writing answers. Claude follows instructions to answer only from the supplied passages and to say when they do not contain the answer, which is the behaviour a RAG app depends on, and it handles long retrieved context. The current Sonnet model costs $2 per million input tokens and $10 per million output. Anthropic offers no embedding model, so the index needs embeddings from another provider, and Anthropic points to Voyage AI.
GPT models answer from retrieved passages as well as most, and LlamaIndex's examples are often written against OpenAI first. The mid-tier model is priced like Claude's at $2 per million input tokens and $10 per million output, and the smallest tier at $0.10 and $0.50 handles simple questions cheaply. One account can also provide the embedding model, which keeps the app to a single provider and a single bill.
Gemini reads very long inputs and images, which helps when the retrieved material is a scanned page or a chart. The Flash models cost less than the mid-tier models from Anthropic and OpenAI, with a rate-limited free tier for building. The same provider offers Gemini Embedding 2, free on its free tier and $0.20 per million text tokens on the paid one, so a single key can cover both the index and the answers.
These are highlighted picks. To see all the tools, check the LLM category.
Hosting Options for Python RAG App
Runs the FastAPI service as an always-on container, and adds a PostgreSQL service with pgvector from a template in the same project. Ingestion jobs can run as a second service on the same plan, so a large upload never competes with questions for memory. The Hobby plan is $5 a month including $5 of usage, and memory costs about $10 per GB a month, so a small RAG API stays near the minimum until the index grows.
A flat price per instance: $7 a month for the smallest always-on one and $25 for 2 GB of RAM, which is about what parsing large PDFs needs. Background workers are a separate service type, a natural home for ingestion. Render's managed PostgreSQL supports the pgvector extension, so the database can live on the same platform. The free instance sleeps after 15 minutes without traffic, so it suits a demo, not an app people query all day.
For teams already on AWS or with documents that must stay in one account. The container runs on ECS with Fargate, and Amazon RDS for PostgreSQL supports pgvector, so the same database serves documents and vectors. Claude and other models are available through Amazon Bedrock under the same bill. Expect more setup than a container platform, and tens of dollars a month once a load balancer and a database are running.
A plain VPS that runs everything on one machine: the API, PostgreSQL with pgvector, and a self-hosted Qdrant if you add it. A server with 4 GB of RAM costs around €5 to €6 a month, which holds a collection of a few hundred thousand chunks. You handle deployment, TLS, backups, and updates yourself, and the vector index and the application compete for the same memory.
These are highlighted picks. To see all the tools, check the Hosting & Cloud category.
Vector Db Options for Python RAG App
The default: the pgvector extension stores embeddings in a PostgreSQL database next to your document records, so a search can filter on ordinary columns and a backup covers everything. HNSW indexes make search fast, and vector indexes cover up to 2,000 dimensions, so a model with larger vectors needs shortened vectors or the half-precision type. Hybrid search uses PostgreSQL's built-in full-text search. It stays comfortable well into the millions of chunks, and splitting a table per tenant keeps filtered searches accurate.
An open-source vector database that runs as one container beside the API, or as a managed cluster with a free 1 GB cluster for testing. It applies metadata filters during the search itself, and it supports hybrid search, which suits an app where each customer may only see their own documents. Choose it when the collection reaches many millions of chunks or when search load starts to slow the main database. Qdrant also offers a Hybrid Cloud that runs on your own servers.
A fully managed, serverless vector database with nothing to operate: create an index and pay for storage and for reads and writes. The Starter plan is free but limited to one AWS region and a small amount of storage, Builder is $20 a month, and Standard has a $50 monthly minimum. Hosted embedding and reranking models are part of the service. It is closed source and cloud only, so it suits a team that wants no database to run and whose documents may leave its own servers.
The quickest store to try: Chroma runs inside your Python process or as one local server, and most RAG tutorials begin with it. Chroma Cloud is the managed version, with $5 of free credits and then usage-based billing for storage ($0.33 per GiB a month) and for writes. Good for a prototype and a small collection that one server holds. For production with several app instances, plan to move to a store built for concurrent access.
Agent Framework Options for Python RAG App
The default, and the one built for retrieval. LlamaIndex covers the whole path: loaders for files and services, parsers that split documents into chunks, ingestion pipelines that skip documents that have not changed, and query engines that combine retrieval, reranking, and the answer. It works with all four vector stores here, with metadata filtering and hybrid search. It also has agents, so a retrieval engine can later become one tool among several. The framework is free, and LlamaParse is a separate paid service for hard documents.
The broader framework: chains, tools, and agents, with retrievers and vector store integrations for the same four stores. Pick it when RAG is one part of a larger agent and the rest of the team already uses LangChain, or when you want LangSmith tracing with no extra setup. Its ingestion helpers are lighter than LlamaIndex's, so expect to write more of the document handling yourself. The vectors in the store stay valid when you switch frameworks, as long as the embedding model is the same.
These are highlighted picks. To see all the tools, check the Agent Frameworks category.
Python RAG App Add-ons
Each addition below extends this stack with a capability the base stack works fine without. None are required: include the ones your product actually needs when building this stack, and skip the rest.
Llm Observability Add-ons
Add LLM observability when you want to see every model call, tool call, and token cost inside a run, so a wrong answer can be traced to the step that caused it.
Open-source tracing that records each question as a trace: the query, the passages retrieved, the prompt sent, and the answer with its cost and latency. That is the fastest way to tell a retrieval problem from a model problem. It integrates with LlamaIndex and LangChain. The cloud Hobby plan is free up to 50,000 units a month, Core is $29, and it can be self-hosted for free so documents stay on your own servers.
LangChain's own tracing platform, and the closest fit when LangChain is the chosen framework: tracing switches on with environment variables, and the trace shows each retriever call and the documents it returned. The Developer plan is free for one seat with 5,000 traces a month, and Plus is $39 per seat. It is a hosted, closed-source service, and self-hosting is limited to the Enterprise plan.
Model Inference Add-ons
Add model inference when you want an open-weight model in the mix: on your own hardware for privacy, or on a hosted inference provider for speed and low per-token prices.
Runs open-weight models on your own machine behind a local address, for the answers and for the embeddings, so private documents never reach an outside API. It costs nothing per token and works well for development. A small local model answers worse than a large hosted one, and embeddings from a local model are not interchangeable with another provider's, so build the whole index with the model you will keep.
The hub for open embedding models and fine-tuned answer models. Inference Providers routes one API key to many hosted providers, and Inference Endpoints gives a model its own dedicated GPU billed by the hour. Pick it when a specific open embedding model suits your language or domain better than the general ones from the big providers.
A hosted provider that runs open-weight models such as Llama and Qwen on its own chips with very fast output, through an OpenAI-compatible API. In RAG it shortens the wait for the answer, since the model reads the retrieved passages and replies in a fraction of the usual time. Billing is per token, with a free rate-limited tier. It serves the answering model, so keep the embeddings with the provider you indexed with.
These are highlighted picks. To see all the tools, check the AI Runtime & Serving category.
Model Aggregator Add-ons
Add a model aggregator when you want one API key and one bill for models from many providers, with automatic fallback when one of them is down, instead of setting up each provider separately.
One API key and one prepaid balance for hundreds of answer models from many providers, behind an OpenAI-compatible endpoint. Useful for testing which model answers your questions best on the same retrieved passages, and for automatic fallback when one provider is down. Models are billed at the providers' prices, with a 5.5% fee when you buy credits.
An open-source gateway you host yourself, between the app and the model providers. It offers one OpenAI-compatible endpoint for answers and for embeddings, with budgets and spend tracking per key and fallback rules, while requests still go to each provider under your own keys. It stores its keys and spend in PostgreSQL, which this stack already runs, and it is one more service to deploy.
These are highlighted picks. To see all the tools, check the AI Model Aggregators category.
CI/CD Add-ons
Add CI/CD when you want a dedicated pipeline for running tests, linting, or multi-stage builds before a deploy goes out. Many hosting platforms already redeploy automatically on every push on their own — a CI/CD tool adds the most value on top of that by gating the deploy on a passing test suite, and matters even more when the hosting choice does not auto-deploy at all, such as a self-hosted server.
Runs the tests on every push and deploys the container to the chosen host. For a RAG app the useful extra is an evaluation step: replay a saved set of questions with known good source documents, and fail the build when retrieval misses them after a change to chunking, embeddings, or prompts.
These are highlighted picks. To see all the tools, check the CI/CD Pipelines category.
Containerization Add-ons
Add containerization when you want the app packaged the same way across local development, staging, and production, or need to deploy somewhere that isn't a managed serverless platform.
Packages the FastAPI service and its Python dependencies, including the document parsers, into one image that runs the same on a laptop and on the server. Docker Compose starts the API and a PostgreSQL image with pgvector together for local development, and every host in this stack can deploy the image.
These are highlighted picks. To see all the tools, check the Containerization category.
Frequently Asked Questions about Python RAG App
Do I need a vector database, or is PostgreSQL with pgvector enough?
Start with pgvector. It keeps the vectors, the document records, and your access rules in one database with one backup, and with an HNSW index it searches millions of chunks quickly. Filtering is where it needs care: a query that filters on a column and searches vectors can return fewer results than asked for, and splitting a table per tenant is pgvector's own advice for multi-tenant apps. Move to Qdrant when the collection reaches tens of millions of chunks, when you filter heavily at high query rates, or when search load slows the rest of the application. Pinecone is the same step with nothing to operate, and Chroma is the one to try in an afternoon. Because LlamaIndex talks to all four through the same interface, moving later means re-running ingestion, not rewriting the app.
Should I use LlamaIndex or LangChain for retrieval?
For an app whose main job is answering from documents, LlamaIndex is the shorter path: loaders, chunking, ingestion pipelines that skip unchanged files, and query engines are part of it. LangChain fits when retrieval is one step in a wider agent that also calls tools, or when the team already knows it and wants LangSmith tracing. The two are not exclusive: a LlamaIndex query engine can be wrapped as a tool inside a LangChain or LangGraph agent. Switching means rewriting the pipeline code, but the stored vectors stay valid if you keep the same embedding model, and the API, database, and host do not change.
Which embedding model should I use, and can I change it later?
The embedding model decides how well a question finds the right passage, and it is separate from the model that writes the answer. Claude has no embedding model of its own, so pair it with Voyage AI, which Anthropic recommends, or with OpenAI or Gemini. Gemini Embedding 2 is free on its free tier and $0.20 per million tokens after that, and an open model through Ollama or Hugging Face costs nothing per token but needs hardware. Choose carefully: vectors from different models cannot be compared, so changing the model means re-embedding every document, and long vectors above 2,000 dimensions need shortening to be indexed by pgvector. Test two models on twenty real questions before indexing the full collection.
What does a RAG app cost to run, and how do I keep it down?
There are two bills. Indexing is a one-off charge per document: its tokens go through the embedding model once, and again only when the file changes, which an ingestion pipeline that skips unchanged files keeps small. Questions are the recurring cost, because every answer sends the retrieved passages to the model along with the question, so a longer context costs more on each call. Retrieving fewer, better chunks, adding a reranker, and sending simple questions to a small model cut it most. Parsing scanned documents with LlamaParse adds a per-page charge. The Pricing section has the per-token rates.
How is this different from the Python AI Agent stack?
Python AI Agent is built around an agent loop: a model decides which tool to call next, and run state is saved so a run can pause and resume. This stack is built around retrieval: documents go in through an ingestion pipeline, and each question follows one path of search, then answer. That is why the vector store is a decision to make up front here and an extra you can skip there. Choose this one for search and question answering over a document collection. Choose Python AI Agent when the model has to take actions, and add retrieval to it later. The two combine: a LlamaIndex query engine can be one of the tools an agent calls, and both stacks share FastAPI and PostgreSQL.
Stacks Related to Python RAG App
Python AI Agent
ProjectA code-first AI agent in Python: an agent framework for the loop, an API in front, and a database for memory.
Svelte + FastAPI
ProjectSvelte SPA frontend with FastAPI backend: minimal JavaScript output meets Python API performance.
FastAPI Backend
ProjectHigh-performance Python REST API with automatic OpenAPI docs and PostgreSQL.
Python Web (FastAPI + React)
ProjectFastAPI backend with React frontend for Python-first web applications.
Scores
Popularity3/5
RAG is among the most common ways to put a model to work on private data, and LlamaIndex and pgvector are widely used for it. The tooling and its best practices still change quickly, and fewer teams run it in production than build web backends.
Learning Curve4/5
FastAPI and PostgreSQL are familiar to a Python developer, but retrieval adds its own vocabulary: chunking, embeddings, indexes, reranking, and evaluation. Getting good answers takes tuning against real questions, not just wiring the pieces together.
Flexibility5/5
Four vector stores, two frameworks, three answer models, and hosts from a container platform to a VPS all swap independently. The one constraint is that the embedding model is fixed once documents are indexed.
Performance3/5
pgvector with an HNSW index answers in milliseconds into the millions of chunks, and a dedicated store scales further. Response time is set by the model that writes the answer, and by how many passages it has to read.
Portability4/5
LlamaIndex, FastAPI, and PostgreSQL are open source and run on any host that takes a container, and the stored vectors move between the four stores by re-running ingestion. Changing the embedding model means re-embedding the whole collection.
Tools in the Python RAG App Stack
Backend Frameworks
Programming Languages
Add-ons (optional — add any, or none)
Llm Observability
Model Inference
Model Aggregator
CI/CD
Containerization
Python RAG App Pricing
LlamaIndex, FastAPI, and PostgreSQL with pgvector are open source and free. A small always-on server with its database costs roughly $5 to $15 a month on a container platform or a VPS. Model usage is the line that grows: mid-tier models from Anthropic and OpenAI cost $2 per million input tokens and $10 per million output, and every question sends retrieved passages along with it. Embeddings add a small one-off cost per document. Parsing scanned files with LlamaParse and tracing are optional and have free tiers.
LlamaIndex, LangChain, and FastAPI are free to use.
About $2 per million input tokens and $10 per million output on mid-tier answer models; embeddings from $0.20 per million tokens, or free on a free tier or a local model. Questions are the line that grows.
Free with pgvector on your own PostgreSQL, or on the free plans of Qdrant, Pinecone, and Chroma. Pinecone Builder is $20 and Standard has a $50 minimum.
An always-on container or a small server with enough memory for parsing; a free instance that sleeps is enough for a demo.