计算机科学
分割
管道运输
人工智能
医学影像学
图像分割
计算机视觉
自然语言处理
人机交互
工程类
环境工程
作者
Sachin M Sabariram,Singapogu Ravikiran B.,S. E.,M. Kavin,G. Saravanan
标识
DOI:10.1109/icmlas64557.2025.10968301
摘要
Medical imaging plays a pivotal role in modern diagnostics, yet existing AI-based pipelines face significant challenges, including reliance on proprietary datasets, limited modality support, and inadequate explainability. This research introduces a scalable, end-to-end pipeline that integrates anomaly detection, segmentation, and diagnostic report generation to address these limitations. By leveraging a VisualLanguage Model (VLM) trained on context-rich image-text pairs generated from open-source label-only datasets using a fine-tuned Large Language Model (LLM), the system ensures scalability across modalities and protects patient privacy by eliminating reliance on proprietary medical data. Interactive, text-driven segmentation is achieved through the integration of a CLIP-based Segment Anything Model (CLIP-SAM), enabling clinicians to refine anomaly boundaries with high precision. A fine-tuned LLaMA-based LLM, employing Chain-of-Thought reasoning, generates detailed, actionable diagnostic reports tailored to clinical needs. Experimental evaluations on MRI, CT, and X-ray datasets demonstrate an Area Under the Curve (AUC) of 0.91 for anomaly detection, a Dice Similarity Coefficient (DSC) of 84.5 % for segmentation, and a BLEU score of $\mathbf{8 5. 6 \%}$ for report quality, significantly surpassing state-of-the-art baselines. This robust and explainable framework establishes a new standard for automated medical imaging analysis, supporting diverse imaging modalities and aligning AIgenerated insights with clinical decision-making.
科研通智能强力驱动
Strongly Powered by AbleSci AI