人工智能
预处理器
机器学习
随机森林
计算机科学
二元分类
人工神经网络
管道(软件)
统计分类
朴素贝叶斯分类器
特征提取
语音识别
深度学习
决策树
数据预处理
语音障碍
痉挛性发音困难
贝叶斯网络
支持向量机
可扩展性
模式识别(心理学)
二进制数
工作流程
临床决策支持系统
作者
Ganesh Bhutkar,Aarti Agarkar,Manthan Khawse,Soham Kanhe,Shravan Katore,Sujal Kandgave
标识
DOI:10.1109/icsscna68616.2026.11546972
摘要
Voice disorders are frequently underdiagnosed due to the subjective and resource-intensive nature of traditional clinical assessments. This work proposes a non-invasive, machine learning based system for early detection and classification of voice pathologies, which is particularly beneficial in settings where access to specialized clinical expertise is limited and vocal health is critical for daily communication and professional use. Using the Saarbrücken Voice Database (SVD), a two-step classification framework is designed hat first performs healthy versus pathological voice detection and then conducts multi-class classification across six clinically significant pathology categories. Audio preprocessing involves FFmpeg-based conversion and extraction of MFCCs, spectral features, jitter, shimmer, pitch, and formants. The evaluation covered several architectures, spanning classic models like Random Forest and LightGBM alongside neural network frameworks such as BiLSTM, 1D and 2D CNNs, LSTM, and hybrid CNN+LSTM setups. Results demonstrate that LightGBM achieves the best performance in binary classification, while the fine-tuned CNN+LSTM hybrid model provides superior multi-class accuracy. This pipeline lays the foundation for a realtime AI-assisted diagnostic tool for scalable clinical deployment.
科研通智能强力驱动
Strongly Powered by AbleSci AI