Encoding laparoscopic image to words using vision transformer for distortion classification and ranking in laparoscopic videos

计算机科学 变压器 人工智能 失真(音乐) 计算机视觉 排名(信息检索) 图像(数学) 情报检索 电信 量子力学 物理 电压 放大器 带宽(计算)
作者
Nouar AlDahoul,Hezerul Abdul Karim,Mhd Adel Momo,Myles Joshua Toledo Tan,Jamie Ledesma Fermin
出处
期刊:Multimedia Tools and Applications [Springer Science+Business Media]
标识
DOI:10.1007/s11042-024-19089-9
摘要

Abstract Laparoscopic videos are tools used by surgeons to insert narrow tubes into the abdomen and keep the skin without large incisions. The videos captured by a camera are prone to numerous distortions such as uneven illumination, motion blur, defocus blur, smoke, and noise which have impact on visual quality. Automatic detection and identification of distortions are significant to enhance the quality of laparoscopic videos to avoid errors during surgery. The video quality assessment includes two stages: classification of distortions affecting the video frames to identify their types and ranking of distortions to estimate the intensity levels. The dataset generated in ICIP2020 challenge including laparoscopic videos was utilized for training, validation, and testing the proposed solution. The difficulty of this dataset is caused by having five categories of distortions and four levels of severity. Additionally, the availability of multiple distortion categories in one video is considered the most challenging part of this dataset. The work presented in this paper contributes to solve the multi-label distortion classification and ranking problem. This paper aims to enhance the performance of distortion classification solutions. Vision transformer which is a deep learning model was used to extract informative features by transferring learning and representation from the general domain to the medical domain (laparoscopic videos). Additionally, six parallel multilayer perceptron (MLP) classifiers were added and attached to vision transformer for distortion classification and ranking. The experiment showed that the proposed solution outperforms existing distortion classification methods in terms of average accuracy (89.7%), average single distortion F1 score (94.18%), and average of both single and multiple distortions F1 score (96.86%). Moreover, it can also rank the distortions with an average accuracy of 79.22% and average F1 score of 78.44%. Hence, the high performance of the method proposed in this paper opens the door to integrate our solution in the intelligent video enhancement system.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
fang完成签到,获得积分10
刚刚
阿高完成签到,获得积分20
刚刚
会飞的鱼完成签到,获得积分10
刚刚
上官若男的应助被超级冥王星采纳,获得10
刚刚
飞飞的应助被静默采纳,获得10
刚刚
小肥才发布了新的文献求助10
刚刚
1秒前
不吵完成签到,获得积分10
2秒前
Carmelo发布了新的文献求助10
2秒前
雪白巨人发布了新的文献求助10
2秒前
Yholy完成签到,获得积分10
3秒前
3秒前
3秒前
坚定背包完成签到,获得积分10
3秒前
3秒前
ccv完成签到,获得积分10
3秒前
Jim_Stone完成签到,获得积分10
4秒前
ZengZeng_完成签到,获得积分10
4秒前
yolo完成签到,获得积分10
4秒前
5秒前
5秒前
852的应助被Tsuki采纳,获得10
6秒前
SciGPT的应助被123采纳,获得10
7秒前
可爱的函函的应助被bodhi采纳,获得10
8秒前
8秒前
8秒前
耍酷靖荷完成签到,获得积分10
9秒前
初衷未央发布了新的文献求助10
9秒前
okayu发布了新的文献求助10
9秒前
10秒前
雪白巨人完成签到,获得积分10
10秒前
可爱的函函的应助被mimilv采纳,获得10
10秒前
清和发布了新的文献求助10
11秒前
Carmelo完成签到,获得积分10
12秒前
12秒前
Jal578完成签到,获得积分10
13秒前
15秒前
15秒前
15秒前
Nole的应助被echo采纳,获得10
15秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
Organizational Behavior 510
Management and the Arts 510
Convergent and bidirectional strategies towards the total synthesis of hemibrevetoxin B 300
Geschichtliche Grundbegriffe (GGB), Band 5: Pro–Soz 300
Die Religion in Geschichte und Gegenwart (RGG), 4. Auflage, Band 7: R–S 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 内科学 物理 有机化学 化学工程 生物化学 复合材料 光电子学 细胞生物学 心理学 量子力学 催化作用 物理化学 电极
热门帖子
关注 科研通微信公众号,转发送积分 7799453
求助须知:如何正确求助?哪些是违规求助? 9334507
关于积分的说明 20467570
捐赠科研通 7390549
什么是DOI,文献DOI怎么找? 3326026
关于科研通互助平台的介绍 2473113
邀请新用户注册赠送积分活动 2343540