情态动词
计算机科学
弹丸
人工智能
医学影像学
计算机视觉
图像(数学)
计算机图形学(图像)
材料科学
高分子化学
冶金
作者
Zhenwei Wang,Qiule Sun,Bingbing Zhang,Weijian Su,Pengfei Wang,Jianxin Zhang,Qiang Zhang
出处
期刊:
日期:2024-12-03
卷期号:: 3799-3804
被引量:5
标识
DOI:10.1109/bibm62325.2024.10821863
摘要
Few-shot learning has become a key technical solution for addressing the challenges of limited data and difficult annotation acquisition in medical image classification. However, relying solely on a single image modality proves inadequate for capture conceptual categories. This paper proposes a novel medical image classification paradigm based on a multi-modal foundation model, called PM2. In addition to the image modality, PM2 introduces supplementary text input (prompt) to further describe images or conceptual categories and facilitate cross-modal few-shot learning. We empirically studied five different prompting schemes under this new paradigm. Furthermore, linear probing in multi-modal models only takes class token as input, ignoring the rich statistical data contained in high-level visual tokens. Therefore, we alternately perform linear classification on the feature distributions of visual tokens and class token. To effectively extract statistical information, we use global covariance pool with efficient matrix power normalization to aggregate the visual tokens. We then combine two classification heads: one for handling image class token and prompt representations encoded by the text encoder, and the other for classifying the feature distributions of visual tokens. Experiments on two medical datasets demonstrate that regardless of the prompting scheme, our method PM2 outperforms its counterparts, achieving state-of-the-art performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI