人工智能
计算机视觉
分割
计算机科学
骨架(计算机编程)
动作(物理)
模式
图像分割
可视化
动作识别
模式识别(心理学)
模态(人机交互)
尺度空间分割
沟通
作者
Junwen Wu,H. ZHANG,Ruijie Li,Zhifang Wang,Sha Liu,Feng He,Jianzhi Wang,Long Xu,Jingjing Zhang
标识
DOI:10.1109/tase.2026.3673914
摘要
Precise and objective quantification of non-human primates (NHPs) actions is vital for neuroscience research and for the phenotypic analysis of neuropsychiatric disorder models. However, existing computational ethology methods primarily focus on coarse-grained action recognition at the video-clip level and lack frame-level temporal segmentation tools for analyzing continuous action streams. This severely restricts the in-depth exploration of action dynamics, including duration, transition patterns, and rhythmicity. To address this issue, this study proposes a multi-modal, multi-stage action segmentation framework. The framework synergistically leverages the complementary strengths of different modalities by deeply fusing kinematic information from skeleton data, appearance features from RGB video, and dense motion information from optical flow, enabling fine-grained frame-wise action analysis of NHPs. To facilitate this task, MonActSeg was introduced as the first fine-grained, frame-by-frame annotated benchmark specifically designed for NHP temporal action segmentation. It encompasses a total of 1.37 million labeled frames across seven core actions. Evaluated on the MonActSeg benchmark under the video-level protocol, the proposed framework demonstrated superior performance, achieving a frame-wise accuracy (Acc) of 91.08%, an edit distance (Edit) of 85.94%, and a segmental F1@50 of 81.23%. This work provides an analytical tool and a standardized benchmark for computational ethology, enabling fine-grained quantification of NHPs behavioral phenotypes.
科研通智能强力驱动
Strongly Powered by AbleSci AI