A Multimodal Approach for Deep‐Learning Classification of Vocal Fold Pathologies in Stroboscopy

频闪仪 人工智能 计算机科学 语音识别 折叠(高阶函数) 医学 模式治疗法 多模态 发声 声带 自然语言处理
作者
Sruthi Surapaneni,Rachel B. Kutler,Sean A. Setzen,Yeo Eun Kim,Peter Yao,Sana H. Siddiqui,Michael J. Pitman,Lucian Sulica,Olivier Elemento,Pegah Khosravi,Anaïs Rameau
出处
期刊:Laryngoscope [Wiley]
卷期号:136 (6): 2503-2510 被引量:1
标识
DOI:10.1002/lary.70355
摘要

OBJECTIVE: To develop and validate a multimodal deep-learning classifier trained on stroboscopic image, voice, and clinicodemographic data, differentiating between three different vocal fold (VF) states: healthy (HVF), unilateral paralysis (UVFP), and VF lesions, including benign and malignant pathologies. METHODS: Patients with UVFP (n = 54), VF lesions (n = 42), and HVF (n = 41) were retrospectively identified. Image frames and voice samples were extracted from stroboscopic videos. Clinicodemographic variables were collected from the electronic health record. Patient-level data was independently divided into training (80%) and testing (20%). Visual features were extracted using a transformer DINOv2 and acoustic features were extracted using Librosa. All three feature modalities were combined using a custom multilayer perceptron. Unimodality models using only image or only voice data were trained for comparison. Accuracy and F1 scores were used to validate the models. RESULTS: On a hold-out test set, the multimodal classifier demonstrated stronger performance (76.9% accuracy) compared to the image classifier (61.5% accuracy) and audio classifier (65.4% accuracy). On an external dataset, the multimodal classifier accuracy dropped to 45%, though still an improvement compared to accuracies of 42% and 31% for the video-only and audio-only modalities, respectively. CONCLUSIONS: In this proof-of-concept study, we successfully developed a multimodal dataset and classifier for VF pathology, demonstrating the potential of combining stroboscopic frames, voice and text data. The multimodal classifier achieved higher accuracy than the image-only model and audio-only models. Future models should validate these findings on larger datasets.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
鲜于雁山发布了新的文献求助10
1秒前
之之完成签到,获得积分10
1秒前
1秒前
Xue发布了新的文献求助10
2秒前
童77完成签到 ,获得积分10
2秒前
华夫饼完成签到 ,获得积分10
3秒前
闪闪元霜完成签到 ,获得积分10
4秒前
5秒前
专注宫苴完成签到,获得积分20
5秒前
xiaozhang完成签到,获得积分10
5秒前
6秒前
开放涔雨完成签到,获得积分10
6秒前
朱建强发布了新的文献求助10
6秒前
7秒前
8秒前
8秒前
8秒前
9秒前
嘎嘎嘎嘎完成签到,获得积分10
9秒前
9秒前
爱咋咋地完成签到,获得积分10
9秒前
ding应助自然白安采纳,获得10
9秒前
Epiphany发布了新的文献求助10
10秒前
10秒前
10秒前
xiaozhang发布了新的文献求助10
11秒前
科研通AI6.4应助我炸了采纳,获得10
11秒前
开放涔雨发布了新的文献求助10
11秒前
juyu发布了新的文献求助10
11秒前
11秒前
完美世界应助镜中眠采纳,获得10
11秒前
霜序发布了新的文献求助10
11秒前
chen发布了新的文献求助10
12秒前
Jasper应助guo采纳,获得10
13秒前
13秒前
14秒前
nannan发布了新的文献求助10
14秒前
14秒前
天真的音完成签到,获得积分10
14秒前
意忆发布了新的文献求助10
15秒前
高分求助中
Les chinois de jakarta: temples et vie collective 1000
Autoparametric Resonance in Mechanical Systems 1000
Social Psychology 800
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 800
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7649641
求助须知:如何正确求助?哪些是违规求助? 9221911
关于积分的说明 19797954
捐赠科研通 7215431
什么是DOI,文献DOI怎么找? 3278202
关于科研通互助平台的介绍 2439022
邀请新用户注册赠送积分活动 2276718