Speech representation learning: Learning bidirectional encoders with single-view, multi-view, and multi-task methods

作者
Qingming Tang
出处
期刊:Cornell University - arXiv [Cornell University]
标识
DOI:10.48550/arxiv.2308.00129
摘要

This thesis focuses on representation learning for sequence data over time or space, aiming to improve downstream sequence prediction tasks by using the learned representations. Supervised learning has been the most dominant approach for training deep neural networks for learning good sequential representations. However, one limiting factor to scale supervised learning is the lack of enough annotated data. Motivated by this challenge, it is natural to explore representation learning methods that can utilize large amounts of unlabeled and weakly labeled data, as well as an additional data modality. I describe my broad study of representation learning for speech data. Unlike most other works that focus on a single learning setting, this thesis studies multiple settings: supervised learning with auxiliary losses, unsupervised learning, semi-supervised learning, and multi-view learning. Besides different learning problems, I also explore multiple approaches for representation learning. Though I focus on speech data, the methods described in this thesis can also be applied to other domains. Overall, the field of representation learning is developing rapidly. State-of-the-art results on speech related tasks are typically based on Transformers pre-trained with large-scale self-supervised learning, which aims to learn generic representations that can benefit multiple downstream tasks. Since 2020, large-scale pre-training has been the de facto choice to achieve good performance. This delayed thesis does not attempt to summarize and compare with the latest results on speech representation learning; instead, it presents a unique study on speech representation learning before the Transformer era, that covers multiple learning settings. Some of the findings in this thesis can still be useful today.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
houniao发布了新的文献求助10
刚刚
1秒前
凡凡fan完成签到,获得积分20
1秒前
同尘发布了新的文献求助10
2秒前
huanmong发布了新的文献求助10
4秒前
cdercder的应助被文艺大侠采纳,获得10
4秒前
6秒前
充电宝的应助被huanmong采纳,获得10
9秒前
10秒前
852的应助被秀秀秀采纳,获得10
11秒前
Jocelyn完成签到,获得积分10
12秒前
古月学术完成签到,获得积分10
13秒前
13秒前
南瓜不说话完成签到,获得积分10
14秒前
简单7879完成签到,获得积分10
14秒前
凡凡fan发布了新的文献求助10
16秒前
16秒前
研友_VZG7GZ的应助被instill采纳,获得10
17秒前
17秒前
lxr完成签到,获得积分20
17秒前
18秒前
AA完成签到,获得积分10
19秒前
yang完成签到 ,获得积分10
19秒前
20秒前
可爱的函函的应助被曾经冰露采纳,获得10
20秒前
20秒前
传奇3的应助被lxr采纳,获得10
21秒前
格兰德法泽尔完成签到,获得积分10
21秒前
爆米花的应助被AA采纳,获得10
22秒前
22秒前
22秒前
qiaoyu完成签到 ,获得积分10
23秒前
王欣发布了新的文献求助10
24秒前
英姑的应助被houniao采纳,获得10
24秒前
ZZZzzzz完成签到 ,获得积分10
24秒前
25秒前
阿北完成签到,获得积分10
26秒前
xingyong发布了新的文献求助10
26秒前
instill发布了新的文献求助10
27秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
Computational Chemical Reaction Engineering: Modeling, Simulation, and Design with MATLAB 600
Organizational Behavior 510
Management and the Arts 510
A Will for the Machine: Computerization, Automation, and the Arts in South Africa 400
Decentring Leadership 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 内科学 物理 有机化学 化学工程 生物化学 复合材料 光电子学 细胞生物学 心理学 量子力学 催化作用 物理化学 电极
热门帖子
关注 科研通微信公众号,转发送积分 7808433
求助须知:如何正确求助?哪些是违规求助? 9340928
关于积分的说明 20504324
捐赠科研通 7400692
什么是DOI,文献DOI怎么找? 3328820
关于科研通互助平台的介绍 2475533
邀请新用户注册赠送积分活动 2347140