An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation

计算机科学 语音识别 保险丝(电气) 灵活性(工程) 深度学习 语音增强 语音处理 人工智能 视听 可视化 多媒体 降噪 数学 统计 电气工程 工程类
作者
Daniel Michelsanti,Zheng‐Hua Tan,Shixiong Zhang,Yong Xu,Yu Meng,Dong Yu,Jesper Jensen
出处
期刊:IEEE/ACM transactions on audio, speech, and language processing [Institute of Electrical and Electronics Engineers]
卷期号:29: 1368-1396 被引量:210
标识
DOI:10.1109/taslp.2021.3066303
摘要

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been tackled using signal processing and machine learning techniques applied to the available acoustic signals. Since the visual aspect of speech is essentially unaffected by the acoustic environment, visual information from the target speakers, such as lip movements and facial expressions, has also been used for speech enhancement and speech separation systems. In order to efficiently fuse acoustic and visual information, researchers have exploited the flexibility of data-driven approaches, specifically deep learning, achieving strong performance. The ceaseless proposal of a large number of techniques to extract features and fuse multimodal information has highlighted the need for an overview that comprehensively describes and discusses audio-visual speech enhancement and separation based on deep learning. In this paper, we provide a systematic survey of this research topic, focusing on the main elements that characterise the systems in the literature: acoustic features; visual features; deep learning methods; fusion techniques; training targets and objective functions. In addition, we review deep-learning-based methods for speech reconstruction from silent videos and audio-visual sound source separation for non-speech signals, since these methods can be more or less directly applied to audio-visual speech enhancement and separation. Finally, we survey commonly employed audio-visual speech datasets, given their central role in the development of data-driven approaches, and evaluation methods, because they are generally used to compare different systems and determine their performance.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
萌萌完成签到,获得积分10
刚刚
小平完成签到,获得积分10
刚刚
zzy发布了新的文献求助10
刚刚
刚刚
1秒前
1秒前
14nuo发布了新的文献求助10
1秒前
1秒前
齐777发布了新的文献求助20
1秒前
Micah完成签到,获得积分10
1秒前
Gina发布了新的文献求助10
1秒前
爱听歌语海完成签到,获得积分10
1秒前
Nxxxxxx发布了新的文献求助10
1秒前
思源应助背后思卉采纳,获得10
2秒前
搜集达人应助郝好采纳,获得10
2秒前
Owen应助小赖想睡觉采纳,获得10
2秒前
闫江洁完成签到,获得积分20
2秒前
不爱上班发布了新的文献求助10
2秒前
烟花应助嘟噜嘟噜采纳,获得10
3秒前
大个应助zz采纳,获得10
3秒前
李健的粉丝团团长应助shuo采纳,获得10
3秒前
高大诗筠完成签到,获得积分10
3秒前
3秒前
CipherSage应助侯伯军采纳,获得10
4秒前
炙热秋翠发布了新的文献求助10
4秒前
4秒前
cy发布了新的文献求助10
4秒前
4秒前
5秒前
斯文败类应助静影沉璧采纳,获得10
5秒前
小丑岩发布了新的文献求助10
5秒前
6秒前
6秒前
LiChen发布了新的文献求助20
6秒前
hanbin完成签到,获得积分10
6秒前
April_03完成签到 ,获得积分10
6秒前
hu应助缓慢含烟采纳,获得80
6秒前
科研通AI6.2应助Tobin采纳,获得10
6秒前
Criminology34应助wchwei123采纳,获得30
6秒前
cyf完成签到,获得积分10
7秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Rosenblum, Global Change Biology 500
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
DIPPR Project 801 - Full Version 380
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7767141
求助须知:如何正确求助?哪些是违规求助? 9310796
关于积分的说明 20319334
捐赠科研通 7352050
什么是DOI,文献DOI怎么找? 3315202
关于科研通互助平台的介绍 2464641
邀请新用户注册赠送积分活动 2329850