计算机科学
术语
变压器
人工智能
自然语言处理
机器学习
自动化方法
评价方法
自动化
计算语言学
语言模型
计分系统
语言理解
作者
Hui Jin,Cynthia Lima,Limin Wang
摘要
Abstract Although AI transformer models have demonstrated notable capability in automated scoring, it is difficult to examine how and why these models fall short in scoring some responses. This study investigated how transformer models’ language processing and quantification processes can be leveraged to enhance the accuracy of automated scoring. Automated scoring was applied to five science items. Results indicate that including item descriptions prior to student responses provides additional contextual information to the transformer model, allowing it to generate automated scoring models with improved performance. These automated scoring models achieved scoring accuracy comparable to human raters. However, they struggle to evaluate responses that contain complex scientific terminology and to interpret responses that contain unusual symbols, atypical language errors, or logical inconsistencies. These findings underscore the importance of the efforts from both researchers and teachers in advancing the accuracy, fairness, and effectiveness of automated scoring.
科研通智能强力驱动
Strongly Powered by AbleSci AI