Retrieval-augmented generation (RAG) is how most production LLM apps give a model access to your own data. Most tutorials stop after a toy demo. This roadmap goes in the order I would learn it, from the core ideas to what real teams add when RAG has to be reliable.
Phase 1: RAG fundamentals
Understand what RAG is and why it exists, then the building blocks: embeddings, chunking, vector databases, retrieval and context augmentation.
- OpenAI Embeddings guide ↗
- Pinecone: Retrieval-Augmented Generation ↗
- Cohere RAG guide ↗
- freeCodeCamp RAG course (YouTube) ↗
- Andrej Karpathy: Let's build GPT from scratch (YouTube) ↗
Phase 2: LangChain
Many production RAG apps use LangChain or LlamaIndex. For LangChain, learn chains, prompt templates, retrievers, memory, agents, tools and LangSmith for tracing.
- LangChain documentation ↗
- LangChain for LLM Application Development (DeepLearning.AI) ↗
- LangChain crash course (YouTube) ↗
Phase 3: LlamaIndex
LlamaIndex is designed around RAG. Cover document loaders, indexing, query engines, retrieval pipelines and RAG evaluation.
Phase 4: Pinecone
A managed vector database. Learn embeddings, similarity search, metadata filtering, hybrid search and index optimisation.
Phase 5: ChromaDB
A good local vector database for beginners. Learn collections, embeddings, querying and metadata filters.
Phase 6: Agentic RAG
This is where the field is moving: agents that plan, call tools and retrieve in several steps instead of one lookup. Learn tool calling, planning, multi-step retrieval, memory and MCP.
Phase 7: GraphRAG
Microsoft's GraphRAG builds a knowledge graph from your documents, which helps with questions that span many sources. Learn knowledge graphs, entity extraction, relationship mapping, graph traversal and community detection.
What production RAG adds
Real systems spend most of their effort outside the demo path:
- Retrieval quality: hybrid search, reranking and query expansion.
- Observability: LangSmith, OpenTelemetry and Arize Phoenix to see what was retrieved and why.
- Evaluation: RAGAS and TruLens to measure answers instead of guessing.
- Guardrails: hallucination detection, PII protection and output validation.
If you want help designing or reviewing a RAG system for your product, book a free call using the button below.