Federated knowledge retrieval elevates large language model performance on biomedical benchmarks

计算机科学 情报检索 数据科学 语言模型 人工智能 钥匙(锁) 万维网 知识建模 数据建模 知识抽取 数据检索 查询语言 自然语言处理
作者
Janet Joy,Andrew I Su
出处
期刊:GigaScience [University of Oxford]
卷期号:15
标识
DOI:10.1093/gigascience/giag007
摘要

BACKGROUND: Large language models (LLMs) have significantly advanced natural language processing in biomedical research; however, their reliance on implicit, statistical representations often results in factual inaccuracies or hallucinations, posing significant concerns in high-stakes biomedical contexts. RESULTS: To overcome these limitations, we developed BioThings Explorer-Retrieval-Augmented Generation (BTE-RAG), a Retrieval-Augmented Generation framework that integrates the reasoning capabilities of advanced language models with explicit mechanistic evidence sourced from BTE, an API federation of more than sixty authoritative biomedical knowledge sources. We systematically evaluated BTE-RAG in comparison to traditional LLM-only methods across three benchmark datasets that we created from DrugMechDB. These datasets specifically targeted gene-centric mechanisms (798 questions), metabolite effects (201 questions), and drug-biological process relationships (842 questions). On the gene-centric task, BTE-RAG increased accuracy from 51 to 75.8% for GPT-4o mini and from 69.8 to 78.6% for GPT-4o. In metabolite-focused questions, the proportion of responses with cosine similarity scores of at least 0.90 rose by 82% for GPT-4o mini and 77% for GPT-4o. While overall accuracy was consistent in the drug-biological process benchmark, the retrieval method enhanced response concordance, producing a greater than 10% increase in high-agreement answers (from 129 to 144) using GPT-4o. We additionally evaluated BTE-RAG alongside GeneGPT-based models on the GeneTuring gene-disease association benchmark and on our mechanistic gene benchmark, demonstrating that the BTE-RAG layer consistently improves accuracy relative to alternative approaches. CONCLUSION: Federated knowledge retrieval provides transparent improvements in accuracy for LLMs, establishing BTE-RAG as a valuable and practical tool for mechanistic exploration and translational biomedical research.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
刚刚
ffallen完成签到,获得积分10
刚刚
zzz完成签到,获得积分10
1秒前
1秒前
1秒前
1秒前
夏问安发布了新的文献求助10
2秒前
鲨鱼也蛀牙完成签到,获得积分10
4秒前
苹果明轩发布了新的文献求助10
4秒前
汉堡包应助时老采纳,获得10
4秒前
ffallen发布了新的文献求助10
5秒前
6秒前
6秒前
阿坤完成签到,获得积分10
6秒前
ohahaha发布了新的文献求助10
6秒前
ticsadis完成签到,获得积分10
6秒前
无敌猫猫头完成签到,获得积分10
6秒前
ABC完成签到,获得积分10
7秒前
踏实谷蓝完成签到 ,获得积分10
7秒前
安平发布了新的文献求助10
8秒前
共享精神应助哈哈采纳,获得20
8秒前
8秒前
8秒前
科研通AI6.3应助hally采纳,获得10
8秒前
上官若男应助杨超越采纳,获得10
8秒前
黑浩源完成签到,获得积分10
8秒前
慕斯发布了新的文献求助10
9秒前
chuanfu完成签到,获得积分0
9秒前
充电宝应助苹果明轩采纳,获得10
9秒前
小疙瘩发布了新的文献求助10
10秒前
10秒前
10秒前
10秒前
ABC发布了新的文献求助10
11秒前
defef完成签到,获得积分10
11秒前
11秒前
CipherSage应助执着的觅露采纳,获得10
11秒前
Hello应助YXM1采纳,获得10
11秒前
完美世界应助乎乎采纳,获得10
12秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
模型平均及其应用 900
Nondestructive Testing Handbook: Vol. 4, Thermal and Infrared Testing (IR), 4th ed 800
作者名:Kristopher P. Plain,悉尼大学的,目前只能查到其四篇论文,想找到其博士论文 590
Évora na Idade Média 555
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Structural Analysis 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7351577
求助须知:如何正确求助?哪些是违规求助? 8963079
关于积分的说明 19040562
捐赠科研通 7000870
什么是DOI,文献DOI怎么找? 3221314
关于科研通互助平台的介绍 2385823
邀请新用户注册赠送积分活动 2201751