Retrieval-Augmented Generation
Systems that retrieve external evidence at query time and supply it to a generative language component before producing an answer.
- Revision
- 1
- Created by
- SCIENDIA Knowledge Desk
- Updated by
- SCIENDIA Knowledge Desk
- Last updated
- 19.08.2026 09:05
Built by the community
Members can improve this article. Every saved change remains visible in the revision ledger.
Overview
Retrieval-augmented generation combines information retrieval with text generation so an answer can draw on documents outside a model's fixed training data. The approach is useful when knowledge changes frequently, belongs to a private collection or requires traceable evidence. Retrieval can improve factual grounding, but it does not automatically make every generated statement correct.
Technical foundations
Sparse retrieval scores exact term matches, while dense retrieval maps queries and passages into vectors whose similarity can capture paraphrase. Hybrid systems combine both signals and rerank candidates with a more expensive cross-encoder. Passage boundaries determine whether evidence stays coherent and fits within the context budget. Generation then estimates an answer conditional on retrieved text, but attention to a passage is not a proof of entailment. Some systems decompose questions, retrieve iteratively and maintain provenance across intermediate claims.
How it works
Documents are collected, cleaned, split into passages and represented by lexical indexes, vector embeddings or both. A user query retrieves candidates that may be reranked before selected passages enter the generation context. The generator follows instructions to answer using that evidence and attach citations. More advanced systems rewrite queries, perform several retrieval steps or route questions to structured databases and tools.
Measurement and research methods
Evaluation begins with a question set paired with relevant passages and supported answers. Retrieval uses recall at k, ranking measures and coverage of all necessary evidence. Generation is assessed claim by claim for correctness, entailment, citation precision, completeness and calibrated abstention. Adversarial tests introduce outdated, duplicated, poisoned or permission-restricted documents. Production monitoring records index version, retrieval results and prompts while respecting privacy. Latency and cost are measured across ingestion, embedding, search, reranking and generation rather than only the final call.
Key ideas
- Retrieval recall sets an upper bound on whether the generator can see the evidence needed for an answer.
- A cited passage may be relevant without actually supporting the claim placed next to it.
- Chunking, metadata, access control and freshness are core system design choices rather than preprocessing details.
Current research frontier
Research explores retrieval over tables, images, knowledge graphs and scientific data, along with agents that choose between search and computation. Long-context models may reduce but do not eliminate retrieval needs because selection, freshness and permissions remain valuable. Methods for citation verification and uncertainty attempt to reject unsupported sentences after generation. Open challenges include learning from user feedback without amplifying popularity bias, defending against indirect prompt injection and creating benchmarks where evidence changes over time rather than remaining frozen in a test set.
Why it matters
RAG supports enterprise search, scientific assistants, customer support and question answering over regulated collections. It allows source updates without retraining an entire model and can provide reviewers with the evidence used for each response.
Limits and open questions
Poor retrieval, stale indexes and contradictory sources can still produce fluent errors. Prompt injection hidden in retrieved documents can redirect the generator, and private passages may leak without end-to-end permission checks. Evaluation must separately measure retrieval, claim support, citation accuracy, abstention and latency instead of relying on subjective answer quality alone.
Explore through connected concepts
This article is indexed with 20 technical tags. Select a tag to explore the Wiki by concept.