计算机科学
人工智能
突出
质量(理念)
公制(单位)
特征(语言学)
卷积神经网络
滑动窗口协议
源代码
特征提取
可视化
模式识别(心理学)
维数之咒
视频质量
图像质量
编码(集合论)
计算机视觉
人工神经网络
相似性(几何)
音频信号
钥匙(锁)
深度学习
特征学习
数据挖掘
主观视频质量
质量得分
作者
Junhao Lin,Yueli Cui,Chenli Fang,Binghong Pan,Chencheng Pan,Gangyi Jiang,Shiqing Zhang,Siwei Ma,Qi Tian
标识
DOI:10.1109/tcsvt.2026.3652641
摘要
The quality evaluation of audio-visual (A/V) content has become increasingly critical in modern multimedia communication systems. Traditional single-modality quality evaluation methods and existing dedicated A/V quality models often fail to accurately assess the quality of A/V signals. To address this challenge, we propose a novel multi-modal cross-attention guided network specifically designed for A/V quality evaluation. By leveraging visual saliency and Mel-spectrum features, our network aims to achieve accurate and comprehensive quality evaluation. Specifically, distorted video frames are first converted into saliency maps, from which perceptually salient patches are selectively extracted and fed into a Convolutional Neural Network (CNN) for intra-frame visual feature extraction. Concurrently, the distorted audio signal is transformed into a Mel-spectrum, and time-frequency patches are extracted via sliding window techniques for CNN-based audio feature extraction. To effectively integrate these features and capture the long-term dependencies across consecutive A/V segments, we design a multi-modal cross-attention module that explicitly models complex inter-modal interactions. The resulting representations are then passed through a series of fully-connected (FC) layers for dimensionality reduction, ultimately deriving the quality score. Extensive experiments on three publicly available A/V quality datasets indicate that our metric outperforms the traditional quality metrics and newly-developed A/V quality metrics. The source code will be released at https://github.com/Jour3141/avqa.
科研通智能强力驱动
Strongly Powered by AbleSci AI