Improving mathematics assessment readability: Do large language models help?

可读性计算机科学阅读（过程）数学教育自然语言处理理解力考试（生物学）人工智能阅读理解写作评估劣势心理学语言学程序设计语言古生物学哲学生物

作者

Nirmal Patel,Pooja Nagpal,Tirth Shah,Archana Sharma,Shrey Malvi,Derek Lomas

出处

期刊：Journal of Computer Assisted Learning [Wiley]
日期：2023-01-23 卷期号：39 (3): 804-822 被引量：2

链接

tudelft.nl tudelft.nldoi.org

标识

DOI：10.1111/jcal.12776

摘要

Abstract Background Readability metrics provide us with an objective and efficient way to assess the quality of educational texts. We can use the readability measures for finding assessment items that are difficult to read for a given grade level. Hard‐to‐read math word problems can put some students at a disadvantage if they are behind in their literacy learning. Despite their math abilities, these students can perform poorly on difficult‐to‐read word problems because of their poor reading skills. Less readable math tests can create equity issues for students who are relatively new to the language of assessment. Less readable test items can also affect the assessment's construct validity by partially measuring reading comprehension. Objectives This study shows how large language models help us improve the readability of math assessment items. Methods We analysed 250 test items from grades 3 to 5 of EngageNY, an open‐source curriculum. We used the GPT‐3 AI system to simplify the text of these math word problems. We used text prompts and the few‐shot learning method for the simplification task. Results and Conclusions On average, GPT‐3 AI produced output passages that showed improvements in readability metrics, but the outputs had a large amount of noise and were often unrelated to the input. We used thresholds over text similarity metrics and changes in readability measures to filter out the noise. We found meaningful simplifications that can be given to item authors as suggestions for improvement. Takeaways GPT‐3 AI is capable of simplifying hard‐to‐read math word problems. The model generates noisy simplifications using text prompts or few‐shot learning methods. The noise can be filtered using text similarity and readability measures. The meaningful simplifications AI produces are sound but not ready to be used as a direct replacement for the original items. To improve test quality, simplifications can be suggested to item authors at the time of digital question authoring.

求助该文献

最长约 10秒，即可获得该文献文件

Improving mathematics assessment readability: Do large language models help?

今日热心研友