计算机科学
尖峰神经网络
子网
异步通信
事件(粒子物理)
人工神经网络
人工智能
机制(生物学)
光学(聚焦)
语音识别
计算机网络
量子力学
认识论
光学
物理
哲学
作者
Qianhui Liu,Dong Xing,Lang Feng,Huajin Tang,Gang Pan
出处
期刊:
日期:2022-04-27
卷期号:: 8922-8926
被引量:21
标识
DOI:10.1109/icassp43922.2022.9746865
摘要
Human brain can effectively integrate visual and auditory information. Dynamic Vision Sensor (DVS) and Dynamic Audio Sensor (DAS) are event-based sensors imitating the mechanism of human retina and cochlea. Since the sensors record the visual and auditory input as asynchronous discrete events, they are inherently suitable to cooperate with the spiking neural network (SNN). Existing works of SNNs for processing events mainly focus on unimodality, however, audiovisual multimodal SNNs are still limited. In this paper, we propose an end-to-end event-based multimodal spiking neural network. The network consists of visual and auditory unimodal subnetworks and a novel attention-based cross-modal subnetwork for fusion. The attention mechanism measures the significance of each modality and allocates the weights to two modalities. We evaluate our proposed multimodal network on an event-based audiovisual joint dataset (MNIST-DVS and N-TIDIGITS datasets). Experimental results show the performance improvement of this multimodal network and the effectiveness of our proposed attention mechanism.
科研通智能强力驱动
Strongly Powered by AbleSci AI