计算机科学
人工智能
机器学习
医学影像学
图形
可视化
域适应
适应(眼睛)
卷积神经网络
利用
深度学习
任务分析
发电机(电路理论)
医学诊断
变压器
人工神经网络
监督学习
临床实习
领域(数学分析)
自然语言处理
资源(消歧)
语义学(计算机科学)
空间关系
模式识别(心理学)
计算机视觉
训练集
特征提取
标记数据
强化学习
作者
Yu Bai,Liang Bai,Xian Yang,Jiye Liang
标识
DOI:10.1109/tmi.2026.3656965
摘要
Adapting Vision Transformers (ViTs) for medical imaging is constrained by the scarcity of data and high-quality annotations, hindering effective training and robust generalization. Visual prompt learning offers a parameterefficient solution for domain adaptation, but its success depends on accurate and task-relevant semantic guidance- a resource rarely available in real-world clinical practice despite its proven benefits. This motivates the need for mechanisms that can automatically extract reliable semantic cues from existing clinical data. To this end, we propose Graph-Enhanced Visual Prompting (GEVP), the first framework to incorporate cross-modal graph learning into prompt generation for medical imaging. GEVP models image patches and report tokens as graph nodes, captures their spatial and semantic relations via a graph neural network, and produces semantically rich prompts. These prompts are injected into a frozen ViT backbone, guiding attention to diagnostically relevant regions without heavy fine-tuning. A consistent downstream prediction mechanism leverages the pretrained prompt generator to handle both report-available and report-absent settings. Experiments on six public downstream datasets show GEVP surpasses strong prompt- and adapter-based baselines by up to +9.65% F1 on imbalanced tasks and delivers superior unseen disease classification.
科研通智能强力驱动
Strongly Powered by AbleSci AI