麦克内马尔试验
医学诊断
医学
中国
优势比
心理学
内科学
病理
统计
地理
数学
考古
作者
Wei Zhong,Yifan Liu,Yan Liu,Kai Yang,HuiMin Gao,Haibin Yan,Wentong Hao,Yousheng Yan,Chenghong Yin
摘要
ChatGPT-4o demonstrated superior diagnostic performance for rare diseases. While Llama3.1:8b demonstrates viability for localized deployment in resource-constrained English diagnostic workflows, Chinese applications require larger models to achieve comparable diagnostic accuracy. This urgency is heightened by the release of open-source models like DeepSeek-R1, which may see rapid adoption without thorough validation. Successful clinical implementation of LLMs requires 3 core elements: model parameterization, user language, and pretraining data. The integration of RAG significantly enhanced open-source LLM accuracy for rare disease diagnosis, although caution remains warranted for low-parameter reasoning models showing substantial performance limitations. We recommend hospital IT departments and policymakers prioritize language relevance in model selection and consider integrating RAG with curated knowledge bases to enhance diagnostic utility in constrained settings, while exercising caution with low-parameter models.
科研通智能强力驱动
Strongly Powered by AbleSci AI