计算机科学
人工智能
判别式
编码器
卷积神经网络
计算机视觉
深度学习
特征提取
特征学习
模式识别(心理学)
特征(语言学)
背景(考古学)
空间分析
人工神经网络
医学影像学
机器学习
可视化
数据建模
稳健性(进化)
机器视觉
目标检测
面子(社会学概念)
变压器
一般化
作者
Fangxiang Feng,Y. L. Xin,Jinghao Xu
标识
DOI:10.1109/icecai66283.2025.11171317
摘要
Parkinson’s disease (PD) is a rapidly progressing neuro-degenerative disorder, and its early detection is a very challenging task. As a crucial modality for PD detection, 3D structural Magnetic Resonance Imaging (MRI) images have been widely adopted in deep learning methods. However, conventional 3D Convolutional Neural Networks (CNNs), despite their promising performance, lack the capacity to capture long-range dependencies. Meanwhile, although pre-trained large-scale Vision-Language Models (VLMs), like BioMedCLIP, offer powerful generalization capabilities and excel in modeling long-range dependencies due to their Transformer-based architecture, they face challenges in adapting to 3D medical data because of spatial information loss in 2D projections. To overcome these limitations, we propose a novel adaptive fine-tuning framework, namely Hybrid3D-ViTA, which integrates a lightweight 3D ResNet-10 for local spatial feature extraction with the pre-trained visual encoder of BioMedCLIP for global context modeling. Additionally, we introduce a CT Adapter for feature alignment and a Patch Attention module for discriminative region enhancement. Experiments conducted on the PPMI dataset demonstrate that our method achieves state-of-the-art performance in PD diagnosis.
科研通智能强力驱动
Strongly Powered by AbleSci AI