增强和替代通信
语音识别
计算机科学
构音障碍
字错误率
电话
混乱
光学(聚焦)
语法
自然语言处理
心理学
语言学
哲学
物理
光学
精神科
精神分析
作者
T. A. Mariya Celin,G. Anushiya Rachel,T. Nagarajan,P. Vijayalakshmi
标识
DOI:10.1109/tnsre.2018.2887089
摘要
An augmentative and alternative speech communication (AASC) aid comprises a speech recognition system and a speech synthesis system. The main challenge in developing such an aid for dysarthric speakers lies in handling errors in the text derived from the recognition system. These errors (substitution, deletion, and insertion) may be due to inability of a dysarthric speaker to utter certain phones (articulatory error) or due to inaccuracy of the models trained (modeling error). Most existing AASC approaches only focus on the articulatory errors and the ones that do address both errors, and do not differentiate between them. However, this paper performs a three-level cascaded analysis to identify and distinguish between these errors, as differentiating these errors will aid in appropriately handling them. Furthermore, analyses in the paper are independent of the syntax of utterances. Based on these analyses, weighted phone confusion transducers are formulated and used to correct erroneous text from the recognition system. The corrected text is finally synthesized by a text-to-speech synthesis system. The proposed AASC is observed to significantly reduce a word error rate of severe dysarthric speakers from 100% to 41.52%, moderate from 61.85% to 18.08%, and mild from 12.23% to 8.55%.
科研通智能强力驱动
Strongly Powered by AbleSci AI