生成语法
计算机科学
自然语言生成
钥匙(锁)
数据科学
自然语言
分类学(生物学)
人工智能
语言模型
自然语言理解
生成模型
自然(考古学)
领域(数学)
背景(考古学)
开放式研究
信息集成
信息系统
数据建模
数据存取
信息模型
数据集成
信息抽取
数据驱动
答疑
数据类型
数据检索
计算语言学
知识管理
文本生成
新兴技术
人机交互
最佳实践
机器学习
语义学(计算机科学)
深度学习
数据收集
情报检索
训练集
信息技术
研究计划
作者
Cheng, Mingyue,Luo, Yucong,Ouyang, Jie,Liu, Qi,Liu, Huijie,Li, Li,Yu, Shuo,Zhang, Bohou,Cao, Jiawei,Ma, Jie,Wang, Daoyu,Chen, Enhong
标识
DOI:10.48550/arxiv.2503.10677
摘要
Retrieval-Augmented Generation (RAG) has gained significant attention in recent years for its potential to enhance natural language understanding and generation by combining large-scale retrieval systems with generative models. RAG leverages external knowledge sources, such as documents, databases, or structured data, to improve model performance and generate more accurate and contextually relevant outputs. This survey aims to provide a comprehensive overview of RAG by examining its fundamental components, including retrieval mechanisms, generation processes, and the integration between the two. We discuss the key characteristics of RAG, such as its ability to augment generative models with dynamic external knowledge, and the challenges associated with aligning retrieved information with generative objectives. We also present a taxonomy that categorizes RAG methods, ranging from basic retrieval-augmented approaches to more advanced models incorporating multi-modal data and reasoning capabilities. Additionally, we review the evaluation benchmarks and datasets commonly used to assess RAG systems, along with a detailed exploration of its applications in fields such as question answering, summarization, and information retrieval. Finally, we highlight emerging research directions and opportunities for improving RAG systems, such as enhanced retrieval efficiency, model interpretability, and domain-specific adaptations. This paper concludes by outlining the prospects for RAG in addressing real-world challenges and its potential to drive further advancements in natural language processing.