计算机科学
文档
背景(考古学)
医疗保健
人工智能
情报检索
自然语言处理
嵌入
健康信息学
德国的
统一医学语言系统
语言模型
判决
基线(sea)
文字嵌入
数据科学
领域(数学分析)
机器学习
信息系统
上下文模型
梅德林
限制
作者
Kamyar Arzideh,Henning Schäfer,Ahmad Idrissi-Yaghir,Cynthia Sabrina Schmidt,Bahadir Eryilmaz,Mikel Bahn,Amin T Turki,Olivia B. Pollok,Eva Maria Hartmann,Philipp Winnekens,Katarzyna Borys,Johannes Haubold,Felix Nensa,René Hosch
摘要
By leveraging a comprehensive real-world dataset spanning multiple medical specialties and using large language models for synthetic question generation, we successfully created and validated domain-specific embedding models. These models can improve medical IR in large-scale search spaces and perform competitively in constrained RAG applications. By publishing the models trained on pseudonymized data, other health care institutions can integrate or adapt these embedding models to their needs. This work establishes a reproducible framework for developing domain-specific clinical embedding models, with the potential to improve data retrieval in medical settings.
科研通智能强力驱动
Strongly Powered by AbleSci AI