计算机科学
变压器
理解力
多媒体
建筑
人工智能
互联网
机器学习
工程类
万维网
电气工程
艺术
视觉艺术
电压
程序设计语言
作者
Sandeep Mandia,Kuldeep Singh,Rajendra Mitharwal
标识
DOI:10.1109/ipas55744.2022.10052945
摘要
Availability of the internet and quality of content attracted more learners to online platforms that are stimulated by COVID-19. Students of different cognitive capabilities join the learning process. However, it is challenging for the instructor to identify the level of comprehension of the individual learner, specifically when they waver in responding to feedback. The learner's facial expressions relate to content comprehension and engagement. This paper presents use of the vision transformer (ViT) to model automatic estimation of student engagement by learning the end-to-end features from facial images. The ViT architecture is used to enlarge the receptive field of the architecture by exploiting the multi-head attention operations. The model is trained using various loss functions to handle class imbalance. The ViT is evaluated on Dataset for Affective States in E-Environments (DAiSEE); it outperformed frame level baseline result by approximately 8% and the other two video level benchmarks by 8.78% and 2.78% achieving an overall accuracy of 55.18%. In addition, ViT with focal loss was also able to produce well distribution among classes except for one minority class.
科研通智能强力驱动
Strongly Powered by AbleSci AI