级联
计算机科学
人工智能
比例(比率)
计算机视觉
计算机图形学(图像)
工程类
物理
化学工程
量子力学
作者
Junkai Wang,Dazhong Ma,Qingchen Wang,Jie Wang
标识
DOI:10.1109/tce.2025.3552695
摘要
Applying high-precision 3D reconstruction for driver facial expression recognition can effectively tackle difficulties such as facial occlusion that are commonly encountered when using 2D image-based detection methods. However, challenges such as weak inter-correlation among multi-scale features as well as cost volumes, and insufficient attention to critical features in deep learning-based multi-view stereo (MVS) methods still persist. Therefore, a cross-scale attention-based MVS (CSA-MVSNet) model is proposed to address these challenges. Specifically, a channel-spatial cross-scale attention strategy is designed to enhance the interdependence of features between encoder and decoder. In this strategy, channel cross-scale attention is employed to address the convolutional neural network’s limitations in handling long-range dependencies, while spatial cross-scale attention enhances the correlation among multi-scale features. Meanwhile, a feature enhancement attention module is proposed to facilitate information exchange among sub-features. Furthermore, a multi-scale cost volume fusion method is presented to propagate low-scale cost volume information to higher scales progressively, ensuring the effective transfer of information across cost volumes from fine to coarse levels. Extensive experiments on DTU dataset demonstrate that the proposed model achieved improved accuracy and completeness in reconstruction compared to existing methods, validating its effectiveness. Moreover, the model offers theoretical potential for achieving high-accuracy driver facial expression recognition.
科研通智能强力驱动
Strongly Powered by AbleSci AI