计算机科学
正确性
推论
水准点(测量)
语言模型
任务(项目管理)
引用
科学文献
人工智能
人类语言
数据科学
情报检索
自然语言处理
数据操作语言
计算模型
作者
Akari Asai,Jacqueline He,Rulin Shao,Weijia Shi,Amanpreet Singh,Joseph Chee Chang,Kyle Shih-Huang Lo,Luca Soldaini,Sergey Feldman,Mike D’Arcy,David Wadden,Matt Latzke,Jenna Sparks,Jena D. Hwang,Varsha Kishore,Minyang Tian,Pan Ji,Shengyan Liu,Hao Tong,Bohao Wu
出处
期刊:Nature
[Nature Portfolio]
日期:2026-02-04
卷期号:650 (8103): 857-863
被引量:20
标识
DOI:10.1038/s41586-025-10072-4
摘要
Scientific progress depends on the ability of researchers to synthesize the growing body of literature. Can large language models (LLMs) assist scientists in this task? Here we introduce OpenScholar, a specialized retrieval-augmented language model (LM)1 that answers scientific queries by identifying relevant passages from 45 million open-access papers and synthesizing citation-backed responses. To evaluate OpenScholar, we develop ScholarQABench, the first large-scale multi-domain benchmark for literature search, comprising 2,967 expert-written queries and 208 long-form answers across computer science, physics, neuroscience and biomedicine. Despite being a smaller open model, OpenScholar-8B outperforms GPT-4o by 6.1% and PaperQA2 by 5.5% in correctness on a challenging multi-paper synthesis task from the new ScholarQABench. Although GPT-4o hallucinates citations 78–90% of the time, OpenScholar achieves citation accuracy on par with human experts. OpenScholar’s data store, retriever and self-feedback inference loop improve off-the-shelf LMs: for instance, OpenScholar-GPT-4o improves the correctness of GPT-4o by 12%. In human evaluations, experts preferred OpenScholar-8B and OpenScholar-GPT-4o responses over expert-written ones 51% and 70% of the time, respectively, compared with 32% for GPT-4o. We open-source all artefacts, including our code, models, data store, datasets and a public demo. A specialized, open-source, retrieval-augmented language model is introduced for answering scientific queries and synthesizing literature, the responses of which are shown to be preferred by human evaluations over expert-written answers.
科研通智能强力驱动
Strongly Powered by AbleSci AI