心理学
拉什模型
写作评估
评定量表
组内相关
比例(比率)
感知
分歧(语言学)
英语作为外语
分类
计分系统
自然语言处理
外语
认知心理学
项目反应理论
应用心理学
语言学
人格
学术写作
等级间信度
语法
控制(管理)
语言能力
人工智能
可靠性(半导体)
语言评估
班级(哲学)
社会心理学
计算机科学
心理测量学
计算语言学
英语
任务分析
作者
Yewon Lee,Myunghwan Hwang
标识
DOI:10.1177/02655322261425759
摘要
This study investigates ChatGPT’s performance as an Automated Writing Evaluation (AWE) system by comparing its scoring with that of human raters and examining learners’ perceptions of its feedback. Six ChatGPT models were developed using different prompt configurations. Sixty English writing samples produced by Korean university English as a Foreign Language (EFL) learners were evaluated by two human raters and the six models. A multifaceted Rasch model, Spearman’s correlation, and intraclass correlation were used to examine reliability, severity, and bias. Learners’ perspectives on the models’ feedback were collected through open-ended surveys and analyzed thematically. The results indicate that prompt design plays a central role in shaping ChatGPT’s scoring behavior. Prompts combining Chain-of-Thought reasoning with Fill-in-the-blank scaffolding were associated with higher scoring consistency, while predefined personas and few-shot exemplars tended to moderate scoring severity. However, no stable patterns were observed for either bias or rating scale use, suggesting that prompt design alone cannot fully control domain-level bias. In particular, reasoning-intensive writing domains showed substantial divergence from human judgment, highlighting the need for human oversight. In parallel, learners generally viewed ChatGPT’s feedback positively, while also noting areas for improvement. Overall, the study demonstrates the potential of prompt-calibrated ChatGPT-based AWE as a supplementary tool for writing assessment and instruction.
科研通智能强力驱动
Strongly Powered by AbleSci AI