正确性
计算机科学
基于案例的推理
演绎推理
自动推理
可信赖性
人工智能
投票
基于模型的推理
水准点(测量)
稳健性(进化)
自然语言处理
推理系统
言语推理
逻辑推理
分析推理
推理心理学
语言模型
非单调逻辑
质量(理念)
答疑
定性推理
语义推理机
多数决原则
实践理性
机器学习
机会主义推理
作者
Zaifu Zhan,Shuang Zhou,Rui Zhang
标识
DOI:10.1093/jamia/ocag042
摘要
OBJECTIVE: To enhance the accuracy, interpretability, and robustness of large language models (LLMs) in medical question answering (MedQA). MATERIALS AND METHODS: We designed a multi-agent peer-reviewed reasoning method in which multiple LLM agents independently generate chain-of-thought (CoT) reasoning with candidate answers, then act as peer reviewers to evaluate each other's reasoning for factual correctness and logical soundness. The highest-rated reasoning chain is selected to produce the final answer. Experiments were conducted with 5 state-of-the-art LLMs (Llama-3.1-8B, Qwen2.5-7B, Phi-4, DeepSeek-LLM-7B, and GPT-oss-20B) on 3 benchmark datasets: HeadQA, MedQA-USMLE, and PubMedQA. Performance was compared against single-model CoT reasoning and CoT-based majority voting. RESULTS: Peer-reviewed reasoning consistently outperformed both baselines. The best model combination achieved an average accuracy of 0.820 across datasets, exceeding the strongest single model (0.777) and majority voting ensembles (up to 0.789). The method also scaled effectively with more participating models, while peer assessments reliably distinguished high- from low-quality reasoning chains. CONCLUSION: The proposed multi-agent peer-reviewed reasoning method enables LLMs to act as both solvers and evaluators, yielding superior performance in MedQA. By emphasizing reasoning quality rather than answer agreement alone, this approach improves accuracy, interpretability, and robustness, offering a promising direction for trustworthy biomedical AI systems.
科研通智能强力驱动
Strongly Powered by AbleSci AI