Auto-scoring Student Responses with Images in Mathematics

Auto-scoring Student Responses with Images in Mathematics
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Sami Baral;Anthony F. Botelho;Abhishek Santhanam;Ashish Gurung;Li Cheng;Neil Heffernan
Sami Baral;Anthony F. Botelho;Abhishek Santhanam;Ashish Gurung;Li Cheng;Neil Heffernan
中科院分区:
--
文献类型:
--
作者:
Sami Baral;Anthony F. Botelho;Abhishek Santhanam;Ashish Gurung;Li Cheng;Neil Heffernan

文献摘要

相似文献

教师经常依靠使用一系列开放式问题来评估学生对数学概念的理解。除了学生开放式作业的传统概念之外,通常以文本简答或论文回答的形式,数字,表格,数字线,图表和象形图的使用是数学中常见的开放式作业的其他例子。虽然自然语言处理和机器学习领域的最新发展已经导致了自动化方法来为学生的开放式作业评分,但这些方法在很大程度上仅限于文本答案。一些基于计算机的学习系统允许学生拍摄手写作品的照片,并将这些图像包含在他们对开放式问题的回答中。然而,有了这个,几乎没有现有的解决方案支持学生手写或绘制的问题答案的自动评分。在这项工作中,我们建立在现有的自动评分文本学生答案的方法基础上,并探索使用OpenAI/CLIP,这是一种旨在表示图像和文本的深度学习嵌入方法,以及光学字符识别(OCR)来提高模型性能。我们评估了我们的方法在包含基于文本和图像的响应的学生开放响应数据集上的性能,并发现在控制其他答案级别特征时,在图像存在的情况下模型误差减少。
Teachers often rely on the use of a range of open-ended problems to assess students’ understanding of mathematical concepts. Beyond traditional conceptions of student open-ended work, commonly in the form of textual short-answer or essay responses, the use of figures, tables, number lines, graphs, and pictographs are other examples of open-ended work common in mathematics. While recent developments in areas of natural language processing and machine learning have led to automated methods to score student open-ended work, these methods have largely been limited to textual answers. Several computer-based learning systems allow students to take pictures of hand-written work and include such images within their answers to open-ended questions. With that, however, there are few-to-no existing solutions that support the auto-scoring of student hand-written or drawn answers to questions. In this work, we build upon an existing method for auto-scoring textual student answers and explore the use of OpenAI/CLIP, a deep learning embedding method designed to represent both images and text, as well as Optical Character Recognition (OCR) to improve model performance. We evaluate the performance of our method on a dataset of student open-responses that contains both text-and image-based responses, and find a reduction of model error in the presence of images when controlling for other answer-level features.