计算机科学
活动识别
人工智能
瓶颈
计算机视觉
特征提取
特征(语言学)
信道状态信息
计算复杂性理论
噪音(视频)
频道(广播)
人工神经网络
模式识别(心理学)
残余物
钥匙(锁)
软件部署
推论
深度学习
无线
国家(计算机科学)
多径传播
视频处理
实时计算
离散余弦变换
传感器融合
作者
Zhiyuan Yang,Yang Yan-kan
摘要
ABSTRACT Human Activity Recognition (HAR) plays a critical role in intelligent surveillance, health monitoring and human‐computer interaction. Although existing HAR methods rely on vision and wireless signals, each presents some inevitable limitations. Vision‐based HAR methods face privacy issues and struggle in complex environments, whereas WiFi Channel State Information (CSI)‐based HAR methods are vulnerable to noise and multipath effects. Although multimodal fusion of video and WiFi CSI can enhance recognition performance, it also introduces new challenges, including complex processing, high computational costs, and inference delays. To address these challenges, we propose WiSk, a lightweight multimodal fusion method for HAR. Our method converts video data into skeletal images using convolutional neural networks (CNNs) to extract skeletal points from video data. Then, we design a unified inverted residual bottleneck structure to efficiently extract features from WiFi CSI and video skeleton images. Furthermore, a lightweight parallel design facilitates spatial feature cross‐fusion and temporal feature extraction to improve activity recognition accuracy and robustness. By leveraging WiFi signals and skeleton images, WiSk effectively mitigates the limitations of vision‐only methods under poor lighting or occlusion while reducing data complexity and enhancing privacy. Experimental results demonstrate that WiSk achieves 99.58% accuracy in good lighting and 98.75% in low‐light or occluded environments. Notably, WiSk exhibits a low computational complexity of 20.55 M floating‐point operations, enabling deployment in resource‐constrained environments.
科研通智能强力驱动
Strongly Powered by AbleSci AI