计算机科学
骨架(计算机编程)
动作(物理)
动作识别
自然语言处理
人工智能
模式识别(心理学)
计算机视觉
人机交互
程序设计语言
班级(哲学)
量子力学
物理
作者
Liyuan Liu,Shaojie Zhang,Yonghao Dang,Xubo Zhang,Jianqin Yin
标识
DOI:10.1109/robio64047.2024.10907305
摘要
Skeleton-based action recognition has a wide range of applications in intelligent robotics, such as assistance for the elderly and disabled, human-robot interaction, and surveillance. Inspired by the success of vision-language pre-trained models, some works attempt to employ text prompts of actions as the prior knowledge to improve the performance of models. However, these text prompts focus more on spatial information and ignore the temporal knowledge of actions, while skeleton features contain temporal and spatial information. This not only leads to a semantic gap between visual and textual modalities but also limits the model from learning temporal cues of actions. To address this issue, we propose a Visual-Language Aligned spatial-temporal Feature Enhancement framework (VLA-FE), which applies both temporal and spatial text prompts of human actions to guide the model in learning more discriminative representations. VLA-FE adopts a two-branch multi-modal training strategy, which is decomposed into spatial-aware and temporal-aware processes. Moreover, We apply two contrastive losses correspondingly to guide this process. Experiments show that the proposed method achieves promising performance on the NTU RGB+D and NW-UCLA datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI