计算机科学
代理(统计)
语言模型
抄写(语言学)
任务(项目管理)
自然语言处理
语音识别
人工智能
机器学习
语言学
工程类
哲学
系统工程
作者
Katrin Tomanek,Jimmy Tobin,Subhashini Venugopalan,Richard Cave,Katie Seaver,Jordan R. Green,Rus Heywood
出处
期刊:
日期:2024-03-18
卷期号:: 10846-10850
被引量:3
标识
DOI:10.1109/icassp48485.2024.10447177
摘要
Automatic Speech Recognition (ASR) systems, despite significant advances in recent years, still have much room for improvement particularly in the recognition of disordered speech. Even so, erroneous transcripts from ASR models can help people with disordered speech be better understood, especially if the transcription doesn't significantly change the intended meaning. Evaluating the efficacy of ASR for this use case requires a methodology for measuring the impact of transcription errors on the intended meaning and comprehensibility. Human evaluation is the gold standard for this, but it can be laborious, slow, and expensive. In this work, we tune and evaluate large language models for this task and find them to be a much better proxy for human evaluators than other metrics commonly used. We further present a case-study using the presented approach to assess the quality of personalized ASR models to make model deployment decisions and correctly set user expectations for model quality as part of our trusted tester program.
科研通智能强力驱动
Strongly Powered by AbleSci AI