印为红字的
链条(单位)
计算机科学
数学教育
心理学
认知科学
人工智能
多媒体
物理
天文
作者
Wenbo Xu,Murizah Kassim,Wai Lam Hoo,Wudao Yang,Tianrong Xu
出处
期刊:International Journal of Modern Physics C
[World Scientific]
日期:2025-06-20
卷期号:37 (06)
被引量:1
标识
DOI:10.1142/s0129183125420136
摘要
Automated essay scoring (AES) systems increasingly leverage fine-tuned large language models (LLMs) to enhance scoring accuracy and feedback generation. However, current LLM-based AES approaches often lack interpretability and consistent alignment with human scoring rubrics, limiting their practical adoption in educational settings. This study proposes QwenScore+, a novel framework that integrates rubric-aware Chain-of-Thought (CoT) prompting with reinforcement learning from human feedback (RLHF) to improve the transparency, quality and educational alignment of automated feedback. QwenScore+ is evaluated on a proprietary IELTS writing dataset comprising over 5000 essays annotated with trait-level scores and expert-written feedback. Experimental results demonstrate that QwenScore[Formula: see text] significantly outperforms strong baselines such as BERT, GPT-3.5 and GPT-4 in feedback generation, achieving higher BLEU, ROUGE-L and cosine similarity scores, alongside improvements in trait-level scoring measured by quadratic weighted kappa (QWK). Furthermore, rubric-aligned CoT prompting enables the generation of feedback that better mirrors human reasoning patterns, as confirmed through automatic metrics and human evaluations. These findings highlight the potential of combining explainable reasoning strategies with human-aligned reward optimization to develop more transparent, reliable and pedagogically valuable AES systems for real-world applications.
科研通智能强力驱动
Strongly Powered by AbleSci AI