认知障碍
听力学
二元分类
认知
诊断准确性
人工智能
特征(语言学)
计算机科学
学习迁移
曲线下面积
语音识别
考试(生物学)
接收机工作特性
认知测验
疾病
人工神经网络
模式识别(心理学)
心理学
机器学习
灵敏度(控制系统)
医学
深度学习
曲线下面积
认知训练
作者
Yun Ho Choi,Hyung-Jun Kim,Suhun Hong,Chaneun Baek,Bohyun Wang,YongSoo Shim,Yun Jeong Hong,Seonjeong Byun,In‐Uk Song,Seunghee Na,Wang-Yeon Won,S.S. Park,Seon Young Ryu,Changtae Hahn,Hae‐Eun Shin,A-Hyun Cho,Eun Ye Lim,Hyun Kook Lim,Dong Woo Kang,Han Jo Kim
标识
DOI:10.1177/13872877251401202
摘要
BackgroundSpeech abnormalities are recognized as early indicators of Alzheimer's disease (AD) and mild cognitive impairment (MCI).ObjectiveTo determine whether deep-learning models trained on mel-spectrograms of brief speech tasks can (i) discriminate individuals with MCI and AD from cognitively normal controls (NC) and (ii) estimate cognitive status with clinically useful accuracy.MethodsSpeech from 594 participants (185 NC, 231 MCI, 178 AD) was recorded through a mobile application that included 11 cognitive-linguistic tasks. Audio was converted into mel-spectrogram images and processed using a VGG16-based deep-learning model with transfer learning and fine-tuning of block 5. Task-specific feature vectors were extracted, concatenated, and used to train a deep neural network. The dataset was split into training, validation, and test sets (3:1:1), and five-split cross-validation was performed.ResultsThe model demonstrated an overall accuracy of 72.4% in classifying NC from the abnormal group (MCI and AD), with sensitivity and specificity of 72.5% and 72.2%, respectively, a balanced accuracy of 72.4%, and an AUC of 0.997. In binary classifications, the model achieved 82.9% accuracy (balanced accuracy 82.9%, AUC 0.992) for NC versus AD, 70.7% accuracy (balanced accuracy 70.3%, AUC 0.956) for NC versus MCI, and 77.5% accuracy (balanced accuracy 78.9%, AUC 0.889) for MCI versus AD. Tasks such as serial subtraction, storytelling, and picture description contributed most to classification performance, indicating their effectiveness in capturing cognitive deficits.ConclusionsMel-spectrogram-based deep-learning analysis of speech shows promise as a rapid, non-invasive, and language-independent screening tool for early cognitive impairment, with potential advantages over traditional assessments such as the Mini-Mental State Examination.