计算机科学
一致性(知识库)
人工智能
姿势
双眼视差
计算机视觉
估计
双眼视觉
工程类
系统工程
标识
DOI:10.1109/isdh64927.2024.00044
摘要
In 3D human pose estimation, binocular vision typically relies on stereo matching to obtain depth information and calculates 3D keypoints using the disparity principle. However, the high computational cost of stereo matching limits its application in real-time scenarios. Alternatively, multi-view consistency can estimate 3D keypoints across multiple views, but with only two views, the constraints are insufficient, often resulting in suboptimal keypoint accuracy. To address these challenges, this paper proposes a binocular 3D pose estimation network based on attention mechanisms, which effectively integrates the advantages of multi-view consistency and disparity methods. Unlike traditional binocular stereo vision systems, our approach leverages a high-precision 2D pose estimation network and dynamically fuses multi-source information via attention mechanisms, thus avoiding the complexity of stereo matching and significantly reducing computational overhead. Additionally, by utilizing multi-view disparity information, we further optimize the estimation of 3D keypoints, improving the accuracy of pose reconstruction. Experiments on the MADS dataset show that DMCPose achieves an average 3D MPJPE of 102.4 mm, outperforming several existing methods, validating its advantages in both accuracy and efficiency.
科研通智能强力驱动
Strongly Powered by AbleSci AI