Multi-Modal Cross-Attention-Guided Network for Audio-Visual Quality Evaluation via Visual Saliency and Mel-Spectrum Features

计算机科学 人工智能 突出 质量(理念) 公制(单位) 特征(语言学) 卷积神经网络 滑动窗口协议 源代码 特征提取 可视化 模式识别(心理学) 维数之咒 视频质量 图像质量 编码(集合论) 计算机视觉 人工神经网络 相似性(几何) 音频信号 钥匙(锁) 深度学习 特征学习 数据挖掘 主观视频质量 质量得分
作者
Junhao Lin,Yueli Cui,Chenli Fang,Binghong Pan,Chencheng Pan,Gangyi Jiang,Shiqing Zhang,Siwei Ma,Qi Tian
出处
期刊:IEEE Transactions on Circuits and Systems for Video Technology [Institute of Electrical and Electronics Engineers]
卷期号:36 (5): 6783-6798 被引量:3
标识
DOI:10.1109/tcsvt.2026.3652641
摘要

The quality evaluation of audio-visual (A/V) content has become increasingly critical in modern multimedia communication systems. Traditional single-modality quality evaluation methods and existing dedicated A/V quality models often fail to accurately assess the quality of A/V signals. To address this challenge, we propose a novel multi-modal cross-attention guided network specifically designed for A/V quality evaluation. By leveraging visual saliency and Mel-spectrum features, our network aims to achieve accurate and comprehensive quality evaluation. Specifically, distorted video frames are first converted into saliency maps, from which perceptually salient patches are selectively extracted and fed into a Convolutional Neural Network (CNN) for intra-frame visual feature extraction. Concurrently, the distorted audio signal is transformed into a Mel-spectrum, and time-frequency patches are extracted via sliding window techniques for CNN-based audio feature extraction. To effectively integrate these features and capture the long-term dependencies across consecutive A/V segments, we design a multi-modal cross-attention module that explicitly models complex inter-modal interactions. The resulting representations are then passed through a series of fully-connected (FC) layers for dimensionality reduction, ultimately deriving the quality score. Extensive experiments on three publicly available A/V quality datasets indicate that our metric outperforms the traditional quality metrics and newly-developed A/V quality metrics. The source code will be released at https://github.com/Jour3141/avqa.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
吴可盈发布了新的文献求助30
1秒前
杜xin完成签到,获得积分10
2秒前
吴奉基发布了新的文献求助10
3秒前
研友_惊鸿发布了新的文献求助10
4秒前
4秒前
4秒前
年刺猬发布了新的文献求助10
4秒前
5秒前
洁净的孤萍完成签到,获得积分20
6秒前
6秒前
cc应助小稻草人的小幸运采纳,获得10
7秒前
852应助笑然采纳,获得10
7秒前
璐璐发布了新的文献求助10
7秒前
8秒前
8秒前
mmk发布了新的文献求助10
8秒前
灰烬使者发布了新的文献求助10
9秒前
小小大杰哥完成签到 ,获得积分10
9秒前
10秒前
10秒前
10秒前
11秒前
11秒前
11秒前
12秒前
Orange应助522采纳,获得10
13秒前
15秒前
cc科研发布了新的文献求助10
15秒前
16秒前
DAI发布了新的文献求助10
17秒前
Demons发布了新的文献求助10
17秒前
20秒前
LH完成签到,获得积分10
21秒前
yjy完成签到,获得积分10
21秒前
LVZHIPENG发布了新的文献求助10
21秒前
21秒前
SULI发布了新的文献求助10
21秒前
22秒前
圆满发布了新的文献求助10
22秒前
24秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
An Introduction to Foreign Language Learning and Teaching 750
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 500
What is the Future of Psychotherapy in Digital Age? Technology, AI Bots, and Psychotherapy after Covid 444
煤炭地下气化渗流燃烧方法的研究 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7632002
求助须知:如何正确求助?哪些是违规求助? 9206365
关于积分的说明 19744385
捐赠科研通 7201289
什么是DOI,文献DOI怎么找? 3274729
关于科研通互助平台的介绍 2436616
邀请新用户注册赠送积分活动 2271356