计算机科学
分割
医学影像学
人工智能
模态(人机交互)
图像分割
水准点(测量)
模式
磁共振成像
基础(证据)
编码器
桥接(联网)
医学物理学
可信赖性
计算机视觉
情报检索
图像(数学)
机器学习
分类学(生物学)
多模态
图像处理
数据科学
适应(眼睛)
封面(代数)
出处
期刊:Bioengineering
[Multidisciplinary Digital Publishing Institute]
日期:2026-07-13
卷期号:13 (7): 803-803
标识
DOI:10.3390/bioengineering13070803
摘要
Vision-language foundation models have transformed medical image segmentation over the past three years. These models pair large image encoders with text prompts, so a single model can segment many anatomical structures, lesion types, and imaging modalities through natural language. This survey reviews vision-language foundation models designed for medical image segmentation. We describe the technical background from contrastive vision-language pretraining to the Segment Anything Model and its medical variants. We propose a three-part taxonomy that covers text-prompt-guided models, large-language-model-embedded architectures, and hybrid frameworks. We examine adaptation strategies such as full fine-tuning, Low-Rank Adaptation, adapters, and prompt engineering. We organize the literature by modality and cover computed tomography, magnetic resonance imaging, pathology, chest radiography, and ultrasound. We discuss clinical uses such as organ segmentation, tumor delineation, and radiotherapy planning. We summarize evaluation metrics and benchmark datasets. We identify four open challenges: prompt dependence, mask hallucination, slow volumetric inference, and limited annotated data. We close with a research roadmap for trustworthy deployment, multimodal pretraining, and clinical integration.
科研通智能强力驱动
Strongly Powered by AbleSci AI