计算机科学
胶质瘤
情态动词
人工智能
自然语言处理
医学
癌症研究
化学
高分子化学
作者
Chaoyu Shi,Xia Zhang,Runzhen Zhao,Wen Zhang,Fei Chen
标识
DOI:10.1038/s41598-025-88458-7
摘要
Pretraining has laid the foundation for the recent success of deep learning in multimodal medical image analysis. However, existing methods often overlook the semantic structure embedded in modality-specific representations, and supervised pretraining requires a carefully designed, time-consuming two-stage annotation process. To address this, we propose a novel semantic structure-preserving consistency method, named "Review of Free-Text Reports for Preserving Multimodal Semantic Structure" (RFPMSS). During the semantic structure training phase, we learn multiple anchors to capture the semantic structure of each modality, and sample-sample relationships are represented by associating samples with these anchors, forming modality-specific semantic relationships. For comprehensive modality alignment, RFPMSS extracts supervision signals from patient examination reports, establishing global alignment between images and text. Evaluations on datasets collected from Shanxi Provincial Cancer Hospital and Shanxi Provincial People's Hospital demonstrate that our proposed cross-modal supervision using free-text image reports and multi-anchor allocation achieves state-of-the-art performance under highly limited supervision. Code: https://github.com/shichaoyu1/RFPMSS.
科研通智能强力驱动
Strongly Powered by AbleSci AI