印为红字的
判别式
人工智能
计算机科学
中心性
机器学习
骨料(复合)
等级间信度
协议
心理学
自然语言处理
可靠性(半导体)
探测理论
评分规则
解耦(概率)
统计
风险评估
工具箱
二次方程
计量经济学
众包
项目反应理论
模式识别(心理学)
认知心理学
信号(编程语言)
数据挖掘
规范化(社会学)
社会心理学
心理测量学
标识
DOI:10.1177/01466216261471171
摘要
The rapid adoption of Large Language Models (LLMs) in educational assessment has reshaped scoring practices, yet evaluation remains tethered to aggregate reliability metrics like Quadratic Weighted Kappa, which obscure discrimination and rater effects. This study applies Signal Detection Theory to evaluate eight state-of-the-art LLMs (including Claude 3.5 Haiku, DeepSeek-V3, Gemini 3 Flash, GPT-4o, and Grok 4.1) against expert human raters across 1,726 essays. By decoupling discrimination from response criteria, I provide a diagnostic analysis of AI scoring behavior. Results indicate that human raters exhibit significantly superior evaluative precision, with average discrimination estimates approximately double those of the AI models. Furthermore, LLMs are prone to pronounced centrality effects and score compression, systematically failing to award the highest rubric tiers. These findings demonstrate that low human-machine agreement stems from both a deficit in discriminative accuracy and systematic shifts in response criteria. Ultimately, this research provides a robust framework for calibrating and selecting AI scoring systems based on specific pedagogical goals and fairness requirements.
科研通智能强力驱动
Strongly Powered by AbleSci AI