复调
计算机科学
变压器
抄写(语言学)
语音识别
工程类
电气工程
声学
语言学
物理
哲学
电压
作者
María Alfaro-Contreras,Antonio Ríos-Vila,Jose J. Valero-Mas,Jorge Calvo-Zaragoza
出处
期刊:
日期:2024-03-18
卷期号:: 706-710
被引量:1
标识
DOI:10.1109/icassp48485.2024.10447162
摘要
End-to-end Audio-to-Score (A2S) transcription aims to derive a score that represents the music content of an audio recording in a single step. While current state-of-the-art methods, which rely on Convolutional Recurrent Neural Networks trained with the Connectionist Temporal Classification loss function, have shown promising results under constrained circumstances, these approaches still exhibit fundamental limitations, especially when dealing with complex sequence modeling tasks, such as polyphonic music. To address these conditions, this work introduces an alternative learning scheme based on a Transformer decoder, specifically tailored for A2S by incorporating a two-dimensional positional encoding to preserve frequency-time relationships when processing the audio signal. The results obtained over three datasets of polyphonic string music confirm the adequacy of the method, which improves the transcription rate by an average of 44% compared to previous approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI