计算机科学
人工智能
计算机视觉
融合
变压器
模式识别(心理学)
工程类
哲学
语言学
电压
电气工程
作者
Tahir Hussain,Hayaru Shouno,Abid Hussain,Dostdar Hussain,Muhammad Ismail,Tatheer Hussain Mir,Fang-Rong Hsu,Taukir Alam,Shabnur Anonna Akhy
出处
期刊:IEEE Access
[Institute of Electrical and Electronics Engineers]
日期:2025-01-01
卷期号:13: 54040-54068
被引量:73
标识
DOI:10.1109/access.2025.3554184
摘要
The rapid advancement of medical imaging technologies requires the development of advanced, automated, and interpretable diagnostic tools for clinical decision-making. Although convolutional neural networks (CNNs) have shown significant promise in medical image analysis, they have limitations in capturing the global context and lack interpretability, thereby hindering their clinical adoption. This study presents EFFResNet-ViT, a novel hybrid deep learning (DL) model designed to address these challenges by combining EfficientNet-B0 and ResNet-50 CNN backbones with a vision transformer (ViT) module. The proposed architecture employs a feature fusion strategy to integrate the local feature extraction strengths of CNNs with the global dependency modeling capabilities of transformers. The extracted features are further refined through a post-transformer CNN and a global average pooling layer to enhance the classification performance. To improve interpretability, EFFResNet-ViT incorporates Grad-CAM visualization techniques to highlight regions contributing to classification decisions and employs t-distributed stochastic neighbor embedding for feature space analysis, providing insights into class separability. The proposed model was evaluated on two benchmark datasets: brain tumor (BT) CE-MRI for BT classification and a retinal image dataset for ophthalmological diagnosis. EFFResNet-ViT achieved state-of-the-art performance, with accuracies of 99.31% and 92.54% on the BT CE-MRI and retinal datasets, respectively. Comparative analyses demonstrate the superior classification performance and interpretability of EFFResNet-ViT over existing ViT and CNN-based hybrid models. The explainable design of EFFResNet-ViT addresses the critical need for transparency in artificial intelligence-driven medical diagnostics, facilitating its potential integration into clinical workflows to improve decision-making and patient outcomes.
科研通智能强力驱动
Strongly Powered by AbleSci AI