计算机科学
人工智能
杠杆(统计)
动作识别
特征(语言学)
事件(粒子物理)
生成语法
机器学习
特征提取
动作(物理)
模式识别(心理学)
神经形态工程学
判别式
模式
尖峰神经网络
人工神经网络
语义特征
视觉对象识别的认知神经科学
模态(人机交互)
功能(生物学)
能量(信号处理)
短时记忆
隐马尔可夫模型
活动识别
语义学(计算机科学)
生成模型
深度学习
深层神经网络
自然语言处理
语音识别
特征学习
作者
Ziliang Ren,Jiaqi Chen,Fuxiang Wu,Qieshi Zhang,Jun Cheng
标识
DOI:10.1109/tmm.2026.3654377
摘要
Event-based human action recognition has gained increasing attention due to its efficiency in dynamic scenarios. Contemporary methodologies for event-based action recognition predominantly treat the problem as a one-hot classification task, which limits their ability to leverage the semantic relationships among various actions. To address this limitation, we propose a Spiking Event-Text Feature Fusion (SETFF) framework, which enhances recognition performance by integrating event and text modalities through a dual-stream architecture. SETFF leverages generative large language models to produce action descriptions, serving as semantic prompts that guide event feature learning. Specifically, a contrastive loss function is employed to align the features of both modalities, enriching the model's capacity to distinguish intricate and subtle actions. Extensive experiments on neuromorphic datasets, including PAF, DailyAction-DVS, DVS128 Gesture, Bullying10K, and UCF101-DVS, demonstrate that SETFF achieves state-of-the-art accuracy, with top-1 accuracy rates of up to 99.65% on the DailyAction-DVS dataset and 98.39% on the PAF dataset. Experimental results underscore the effectiveness of multimodal fusion in SNNs, advancing event-based action recognition while preserving the energy efficiency characteristic of SNNs.
科研通智能强力驱动
Strongly Powered by AbleSci AI