人工智能
糖尿病性视网膜病变
计算机科学
医学影像学
模式识别(心理学)
二元分类
眼底(子宫)
医学诊断
视网膜病变
医学
村上
数据挖掘
特征提取
计算机视觉
特征(语言学)
诊断准确性
异常
疾病
可靠性(半导体)
机器学习
验光服务
作者
Jaffer Shah Jahangeer,P. Shanmugavadivu
标识
DOI:10.1109/ictbig68706.2025.11323789
摘要
This paper reports a crucial failure mode for Vision Transformers in medical imaging through zero-shot evaluation of DINOv2 on diabetic retinopathy (DR) detection. Through zero-shot evaluation, and using an identical experimental protocol through a unified processing pipeline, we observe a catastrophic drop of 42.33 % accuracy between two standard diabetic retinopathy datasets: we see DINOv2 achieves 96.18 % binary classification accuracy on APTOS (3,662 images) but only 53.85 % on IDRiD (597 images). This striking drop in performance occurs despite both datasets representing the same clinical task of 5-grade diabetic retinopathy severity classification from fundus photographs. A unified analysis shows that this apparent “success” on APTOS is attributable to the model's ability to classify Grade 0 (no DR) with an F 1 -score of 95.71 %. In contrast, the model only achieves$F 1$-scores of$43-52 \%$for minority disease grades, illustrating that DINOv2 does not adequately detect minority disease grades on either dataset. On IDRiD, the model totally fails to detect Grade 1 diabetic retinopathy (0 % precision and recall) and we find that attention mechanisms exhibit less than 1.2 % intersection over union overlap with clinical annotation of lesions. Analysis of the feature space revealed that separability is extremely poor, showing an Adjusted Rand Index of 0.092 on IDRiD, compared to 0.231 on APTOS. This data highlights that high accuracy on a single dataset for medical imaging cannot be generalized to another dataset, even for the same disease and imaging modality, which has serious implications for the reliability of zero-shot foundation model performance in direct clinical applications.
科研通智能强力驱动
Strongly Powered by AbleSci AI