计算机科学
水准点(测量)
自动汇总
人工智能
语义鸿沟
语义学(计算机科学)
构造(python库)
机器学习
桥(图论)
医学影像学
自然语言处理
图像(数学)
特征提取
可视化
训练集
情报检索
模式识别(心理学)
监督学习
深度学习
统一医学语言系统
语义相似性
异常
图像分割
分割
方向(向量空间)
计算机视觉
数据挖掘
作者
Haoran Lai,Zihang Jiang,Qingsong Yao,Rongsheng Wang,Zhiyang He,Xiaodong Tao,Weifu Lv,Wei Wei,S. Kevin Zhou
标识
DOI:10.1109/jbhi.2025.3629096
摘要
3D medical images such as computed tomography are widely used in clinical practice, offering a great potential for automatic diagnosis. Supervised learning-based approaches have achieved significant progress but rely heavily on extensive manual annotations, limited by the availability of training data and the diversity of abnormality types. Vision-language alignment (VLA) offers a promising alternative by enabling zero-shot learning without additional annotations. However, we empirically discover that the visual and textural embeddings after alignment endeavors from existing VLA methods form two well-separated clusters, presenting a wide gap to be bridged. To bridge this gap, we propose a Bridged Semantic Alignment (BrgSA) framework. First, we utilize a large language model to perform semantic summarization of reports, extracting high-level semantic information. Second, we design a Cross-Modal Knowledge Interaction module that leverages a cross-modal knowledge bank as a semantic bridge, facilitating interaction between the two modalities, narrowing the gap, and improving their alignment. To comprehensively evaluate our method, we construct a benchmark dataset that includes 15 underrepresented abnormalities as well as utilize two existing benchmark datasets. Experimental results demonstrate that BrgSA achieves state-of-the-art performances on both public benchmark datasets and our custom-labeled dataset, with significant improvements in zero-shot diagnosis of underrepresented abnormalities.
科研通智能强力驱动
Strongly Powered by AbleSci AI