计算机科学
人机交互
多通道交互
感知
心理学
神经科学
作者
Xin Peng,Song Wang,Tong Wu,Hao Long
标识
DOI:10.1109/cscwd64889.2025.11033301
摘要
Multimodal interaction enhances human-computer interaction by integrating inputs such as voice, gestures, and images, making it more natural. However, challenges still exist in multimodal data fusion, personalized interaction experiences, and assistive functionalities. This paper focuses on the problem of multimodal interaction fusion. First, it defines and designs multi-channel interaction methods for voice, gestures, and eye movements in immersive environments, citing application examples. A feature fusion method based on the transformer architecture is proposed. Additionally, multiple $(\mathrm{n}=7)$ researchers with varying experience in immersive environment applications were invited to participate in the design and use. A hierarchical segmentation of interaction intentions and the identification of user interaction intentions were achieved based on multimodal feature fusion. Finally, integrating the feature fusion and intention recognition models, a smart interaction assistant application was designed and developed. Feedback from the researchers was used to evaluate the entire method, confirming the feasibility of the approach and application presented in this paper.
科研通智能强力驱动
Strongly Powered by AbleSci AI