计算机科学
人工智能
透明度(行为)
强化学习
过度拟合
推论
机器学习
一般化
可信赖性
领域(数学分析)
自然语言
机会主义推理
语言模型
答疑
基于模型的推理
非单调逻辑
医学诊断
语言理解
训练集
共谋
自然语言处理
监督学习
作者
Jiazhen Pan,Liu, Che,Junde Wu,Fenglin Liu,Jiayuan Zhu,Hongwei Li,Chen Chen,Cheng Ouyang,Daniel Rueckert
标识
DOI:10.48550/arxiv.2502.19634
摘要
Reasoning is a critical frontier for advancing medical image analysis, where transparency and trustworthiness play a central role in both clinician trust and regulatory approval. Although Medical Visual Language Models (VLMs) show promise for radiological tasks, most existing VLMs merely produce final answers without revealing the underlying reasoning. To address this gap, we introduce MedVLM-R1, a medical VLM that explicitly generates natural language reasoning to enhance transparency and trustworthiness. Instead of relying on supervised fine-tuning (SFT), which often suffers from overfitting to training distributions and fails to foster genuine reasoning, MedVLM-R1 employs a reinforcement learning framework that incentivizes the model to discover human-interpretable reasoning paths without using any reasoning references. Despite limited training data (600 visual question answering samples) and model parameters (2B), MedVLM-R1 boosts accuracy from 55.11% to 78.22% across MRI, CT, and X-ray benchmarks, outperforming larger models trained on over a million samples. It also demonstrates robust domain generalization under out-of-distribution tasks. By unifying medical image analysis with explicit reasoning, MedVLM-R1 marks a pivotal step toward trustworthy and interpretable AI in clinical practice. Inference model is available at: https://huggingface.co/JZPeterPan/MedVLM-R1.
科研通智能强力驱动
Strongly Powered by AbleSci AI