计算机科学
人工智能
计算机视觉
姿势
杠杆(统计)
正规化(语言学)
三维姿态估计
动作识别
语义学(计算机科学)
单眼
关节式人体姿态估计
噪音(视频)
特征(语言学)
模式识别(心理学)
特征提取
特征向量
卷积(计算机科学)
运动估计
频域
领域(数学分析)
核(代数)
由运动产生的结构
虚假关系
判别式
不变(物理)
可视化
光流
人类视觉系统模型
领域知识
作者
Liyuan Shi,Shuhua Wu,S. Yang,Weibin Qiu,Qiang Dong,Jiarui Zhao
摘要
Abstract Although significant progress has been made in monocular video‐based 3D human pose estimation, existing methods lack guidance from fine‐grained high‐level prior knowledge such as action semantics and camera viewpoints, leading to significant challenges for pose reconstruction accuracy under scenarios with severely missing visual features, i.e., complex occlusion situations. We identify that the 3D human pose estimation task fundamentally constitutes a canonical inverse problem, and propose a motion‐semantics‐based diffusion(MS‐Diff) framework to address this issue by incorporating high‐level motion semantics with spectral feature regularization to eliminate interference noise in complex scenes and improve estimation accuracy. Specifically, we design a Multimodal Diffusion Interaction (MDI) module that incorporates motion semantics including action categories and camera viewpoints into the diffusion process, establishing semantic‐visual feature alignment through a cross‐modal mechanism to resolve pose ambiguities and effectively handle occlusions. Additionally, we leverage a Spectral Convolutional Regularization (SCR) module that implements adaptive filtering in the frequency domain to selectively suppress noise components. Extensive experiments on large‐scale public datasets Human3.6M and MPI‐INF‐3DHP demonstrate that our method achieves state‐of‐the‐art performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI