形成性评价
印为红字的
一致性(知识库)
教育技术
计算机科学
同行反馈
位于
数学教育
终结性评价
班级(哲学)
心理学
主动学习(机器学习)
科学教育
教育评估
工程教育
学业成绩
电子投票
精英
教学设计
社会技术系统
自我效能感
学习评价
教学方法
风险学生
高等教育
作者
Osman Yaşar,Andriy Kashyrskyy,Charles Xie,Dylan Bulseco
标识
DOI:10.1016/j.ijaied.2026.100013
摘要
Large language models (LLMs) have shown potential not only as content generators but as evaluators capable of providing nuanced feedback. However, much of the current application of LLMs in education treats them as static graders rather than dynamic participants in formative assessment processes. This study explores how rubric-guided prompting and role-aware feedback simulations can enable LLMs to approximate human evaluative reasoning across dimensions critical to design-based learning. Using situated learning theory, iterative design pedagogy, and cognitive models of scientific and engineering thinking, the research developed a framework wherein LLMs were trained to align with expert judgment. A stratified sample of student design artifacts was evaluated across different roles (instructor, peer reviewer, grant reviewer) using targeted prompting. Feedback outputs were coded for tone and evaluation focus. Rubric engineering was found to substantially improve LLM-human agreement in cognitively complex categories. LLMs demonstrated role-sensitive feedback variation, and final rubric-tuned LLM ratings achieved high consistency with human ratings (Cronbach’s Alpha > 0.75). Figures and tables illustrate how role-specific emphasis and tone were reliably modulated. When properly scaffolded, LLMs can serve as dynamic co-evaluators and rubric co-design partners. These findings advance the use of AI from automation to pedagogical emulation, offering scalable, reflective feedback ecosystems for design-rich learning environments.
科研通智能强力驱动
Strongly Powered by AbleSci AI