语音识别
计算机科学
构音障碍
Mel倒谱
特征提取
字错误率
解码方法
特征(语言学)
模式识别(心理学)
人工智能
隐马尔可夫模型
词(群论)
声学模型
说话人识别
性格(数学)
单词识别
语音处理
倒谱
语音编码
基础(拓扑)
滤波器(信号处理)
序列(生物学)
语音活动检测
作者
Faila Nadhifatul Aryza,Syahroni Hidayat
标识
DOI:10.3991/ijoe.v22i04.59849
摘要
Speech recognition for individuals with dysarthria remains challenging due to unstable acoustic signals, high temporal variability, and frequent articulatory distortions, all of which hinder the ability of acoustic models to consistently capture phonetic patterns. This study aims to identify the most effective feature extraction strategy among three approaches, namely Wav2Vec 2.0, MFCC combined with Wav2Vec 2.0, and Wavelet-MFCC combined with Wav2Vec 2.0, evaluated using the UA-Speech dataset. All models were trained using the Wav2Vec 2.0 Base architecture with a CTC decoding mechanism to map audio signals to character sequences in an end-to-end manner. The experimental results demonstrate that the MFCC-Wav2Vec 2.0 combination yields the best performance, achieving a Word Error Rate (WER) of 0.2990. These findings indicate that combining traditional acoustic features with self-supervised representations yields a more robust speech recognition system for dysarthric speech.
科研通智能强力驱动
Strongly Powered by AbleSci AI