计算机科学
人工智能
深度学习
稳健性(进化)
卷积神经网络
情绪识别
面部表情
新颖性
特征提取
语音识别
特征(语言学)
特征学习
帧(网络)
模态(人机交互)
人工神经网络
情绪分类
深层神经网络
情感计算
模式识别(心理学)
深信不疑网络
机器学习
面部识别系统
隐马尔可夫模型
计算机视觉
作者
Cheng-Kai Lu,Chien-Wei Lu,Guan Bo Lin
标识
DOI:10.1109/taffc.2025.3621086
摘要
This paper presents a lightweight multimodal deep learning framework for real-time emotion recognition on resource-constrained companion robots, exemplified by Zenbo Junior II. The framework integrates a customized GhostNet with Triplet Attention Modules (TAM) and a Frame Attention Network (FAN) for spatio-temporal facial feature encoding, and employs a depth-optimized one-dimensional convolutional neural network (1D-CNN) for compact speech representation. Decision-level fusion based on the geometric mean enhances robustness to noisy modality predictions. The proposed model comprises 0.92 million parameters and requires 0.77 billion floating-point operations (GFLOPs), achieving 97.56% accuracy on the RAVDESS dataset and 82.33% on CREMA-D. In contrast to existing approaches that optimize accuracy at the expense of computational efficiency, the proposed method demonstrates a balance of accuracy, efficiency, and deployability. These results highlight both the novelty and the feasibility of the framework for real-time emotion monitoring in healthcare and human-robot interaction.
科研通智能强力驱动
Strongly Powered by AbleSci AI