亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

Toward an Effective Action-Region Tracking Framework for Fine-Grained Video Action Recognition

动作(物理) 跟踪(教育) 计算机科学 动作识别 计算机视觉 人工智能 心理学 物理 教育学 量子力学 班级(哲学)
作者
Baoli Sun,Yihan Wang,Xinzhu Ma,Zhihui Wang,Kun Lu,Zhiyong Wang
出处
期刊:IEEE transactions on neural networks and learning systems [Institute of Electrical and Electronics Engineers]
卷期号:: 1-15
标识
DOI:10.1109/tnnls.2025.3602089
摘要

Fine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture coarse-grained motion patterns but struggle to identify subtle details in local regions evolving over time. In this work, we introduce the action-region tracking (ART) framework, a novel solution leveraging a query-response mechanism to discover and track the dynamics of distinctive local details, enabling distinguishing similar actions effectively. Specifically, we propose a region-specific semantic activation module that employs discriminative and text-constrained semantics serve as queries to capture the most action-related region responses in each video frame, facilitating interaction among spatial and temporal dimensions with corresponding video features. The captured region responses are then organized into action tracklets, which characterize the region-based action dynamics by linking related responses across different video frames in a coherent sequence. The text-constrained queries are designed to expressly encode nuanced semantic representations derived from the textual descriptions of action labels, as extracted by the language branches within visual language models. To optimize generated action tracklets, we design a multilevel tracklet contrastive constraint among multiple region responses at spatial and temporal levels, which can effectively distinguish individual region responses in each video frame (spatial level) and establish the correlation of similar region responses between adjacent video frames (temporal level). In addition, we implement a task-specific fine-tuning mechanism to refine textual semantics during training. This ensures that the semantic representations encoded by vision language models (VLMs) are not only preserved but also optimized for specific task preferences. Comprehensive experiments on several widely used action recognition benchmarks, i.e., FineGym, Diving48, NTURGB-D, Kinetics, and Something-Something, clearly demonstrate the superiority to previous state-of-the-art baselines.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
超帅的若剑完成签到,获得积分10
7秒前
24秒前
25秒前
雪白丸子完成签到,获得积分10
26秒前
ahai发布了新的文献求助10
30秒前
huoo完成签到 ,获得积分10
32秒前
橙子发布了新的社区帖子
35秒前
38秒前
Hi完成签到,获得积分10
39秒前
强健的千柔完成签到,获得积分10
50秒前
踏实凡儿完成签到 ,获得积分10
55秒前
加贝完成签到 ,获得积分10
59秒前
zhoushishan完成签到,获得积分20
1分钟前
1分钟前
hh完成签到 ,获得积分10
1分钟前
花小研发布了新的文献求助10
1分钟前
动听一德完成签到,获得积分10
1分钟前
俏皮友桃完成签到,获得积分10
1分钟前
1分钟前
1分钟前
1分钟前
GingerF应助yangjian采纳,获得50
1分钟前
隐形依秋完成签到,获得积分10
1分钟前
景阁发布了新的文献求助10
1分钟前
凝土完成签到 ,获得积分10
1分钟前
molihuakai应助景阁采纳,获得10
1分钟前
思源应助橙子采纳,获得10
1分钟前
诚心荟完成签到,获得积分10
2分钟前
瘦瘦的如冰完成签到,获得积分10
2分钟前
meeteryu完成签到,获得积分10
2分钟前
2分钟前
77发布了新的文献求助10
2分钟前
2分钟前
科研通AI6.4应助杜阳辉采纳,获得10
2分钟前
2分钟前
2分钟前
小梦完成签到,获得积分10
2分钟前
Leedesweet完成签到,获得积分10
2分钟前
2分钟前
在水一方应助落伍少年采纳,获得10
2分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Navigating Normative Orders. Interdisciplinary Perspectives 800
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7738644
求助须知:如何正确求助?哪些是违规求助? 9287711
关于积分的说明 20184613
捐赠科研通 7316595
什么是DOI,文献DOI怎么找? 3305973
关于科研通互助平台的介绍 2458296
邀请新用户注册赠送积分活动 2315849