计算机科学
智能手表
动作识别
动作(物理)
人工智能
人机交互
活动识别
可穿戴计算机
计算机视觉
语音识别
特征提取
嵌入式系统
信号处理
模式识别(心理学)
移动设备
作者
Jiale Wang,彩香 楠,Xinyi Dai,Ming Xia,Chuang Shi,Wu Chen
标识
DOI:10.1109/jiot.2026.3686886
摘要
Human Activity Recognition (HAR) has become increasingly important in healthcare, smart homes, and human–computer interaction applications. However, traditional vision-based approaches suffer from privacy concerns and high deployment costs, while single-sensor methods are often limited in robustness and generalization capability. To address these challenges, this study proposes a multimodal HAR framework that integrates a 2×2 Wi-Fi Channel State Information (CSI) array with wearable inertial sensors. The proposed 2×2 CSI array enables synchronized multi-channel acquisition and fusion, improving signal stability and reducing packet loss in complex indoor environments. Meanwhile, accelerometer and gyroscope data are collected from a smartwatch and combined with CSI signals to construct a comprehensive multimodal representation. A hierarchical deep learning architecture, termed M2HAR-Net, is designed to effectively extract and fuse spatial–temporal features from heterogeneous modalities, capturing complementary motion characteristics. To enhance computational efficiency and real-time responsiveness, an ablation study on PCA-based frequency-domain reduction demonstrates that retaining only the first principal component preserves discriminative information while significantly reducing inference overhead. Experimental results on a dataset collected from four participants show that the proposed method achieves an overall accuracy of 98.77% across nine daily activities. Comparative evaluations with several state-of-the-art models further validate the effectiveness and robustness of the proposed framework. These findings indicate that efficient multimodal fusion can substantially improve HAR performance, while future work will focus on cross-environment generalization and large-scale validation.
科研通智能强力驱动
Strongly Powered by AbleSci AI