计算机科学
姿势
先验概率
人工智能
概化理论
机器学习
变压器
特征学习
图形
三维姿态估计
水准点(测量)
深度学习
特征提取
特征(语言学)
模式识别(心理学)
人工神经网络
学习迁移
卷积神经网络
监督学习
自编码
计算机视觉
利用
特征向量
任务分析
深层神经网络
任务(项目管理)
关节式人体姿态估计
可视化
实体造型
作者
Tingting Liu,Ji Gan,Jiaxu Leng,Shuang Li,Lei Cheng Chen,Xinbo Gao
标识
DOI:10.1109/tmm.2026.3651129
摘要
3D hand pose estimation is crucial for many human-computer interaction applications. However, existing deep neural networks (DNNs) for 3D hand pose estimation suffer from poor generalizability due to data scarcity and a lack of domain-specific knowledge. In contrast, humans remain far better than DNNs at learning; Humans require fewer samples for learning new concepts under the guidance of their prior knowledge. Inspired by this, we propose a graph-enhanced CLIP to deliver visual-semantic priors to DNNs, and provide refined domain-specific knowledge for better 3D hand pose estimation. Specifically, we first introduce a pre-trained CLIP to guide the hand estimation model in learning the semantic-aware visual features, and text-free contrastive learning is proposed to effectively transfer high-level visual-semantic priors from the pre-trained large multimodal models. Notably, our strategy is data-agnostic and avoids designing hand-crafted text prompts for various visual inputs. Second, we introduce novel graph Transformers to refine the domain-specific knowledge by fully exploiting the local adjacent relations of hand joints and capturing the global structure representations of hand poses. The introduced graph Transformers are supposed to further refine the generalized CLIP feature for the downstream task (i.e., hand pose estimation) with better performance. Experiments show that our proposed graph-enhanced CLIP achieves state-of-the-art performances on benchmark datasets, demonstrating its effectiveness for 3D hand pose estimation. The source code is available at https://github.com/TLiu2832/TRVSP-GE-CLIP.
科研通智能强力驱动
Strongly Powered by AbleSci AI