RAG or Long Context? Picking the Right Tool
Sep 2026 · 1 min read
Context windows keep growing, so a fair question is whether retrieval-augmented generation is still worth the trouble. My answer: sometimes yes, sometimes no.
When long context wins
If your material is small, say a handful of documents, just put it all in the prompt. There is no index to build, no chunking to tune, and no retrieval misses. Start here.
When retrieval still wins
- The corpus is far larger than any window.
- Content changes often, and you do not want to re-send everything each time.
- Cost and latency matter, since sending less text is cheaper and faster.
- You need to show sources, and retrieval gives you natural citations.
Where RAG usually goes wrong
Poor chunking, embeddings that match on topic but not on the actual question, and no way to tell when the right passage was not retrieved. Fixing these is mostly unglamorous data work, not model work.
A practical path
Begin with everything in context. Move to retrieval only when size, cost, or freshness forces you to, and measure both approaches on the same test questions.