模式
计算机科学
卷积神经网络
人工智能
背景(考古学)
融合
RGB颜色模型
传感器融合
机器学习
深度学习
模式识别(心理学)
哲学
社会学
古生物学
生物
语言学
社会科学
作者
Konrad Gadzicki,Razieh Khamsehashari,Christoph Zetzsche
标识
DOI:10.23919/fusion45008.2020.9190246
摘要
Combining machine learning in neural networks with multimodal fusion strategies offers an interesting potential for classification tasks but the optimum fusion strategies for many applications have yet to be determined. Here we address this issue in the context of human activity recognition, making use of a state-of-the-art convolutional network architecture (Inception I3D) and a huge dataset (NTU RGB+D). As modalities we consider RGB video, optical flow, and skeleton data. We determine whether the fusion of different modalities can provide an advantage as compared to uni-modal approaches, and whether a more complex early fusion strategy can outperform the simpler late-fusion strategy by making use of statistical correlations between the different modalities. Our results show a clear performance improvement by multi-modal fusion and a substantial advantage of an early fusion strategy.
科研通智能强力驱动
Strongly Powered by AbleSci AI