Applying natural language processing to automatically assess student conceptual understanding from textual responses

Applying natural language processing to automatically assess student conceptual understanding from textual responses
复制标题

应用自然语言处理自动评估学生对文本响应的概念理解

DOI:
--
复制
发表时间:
2021
影响因子:
4.1
通讯作者:
W. Boles
W. Boles
中科院分区:
教育学3区
文献类型:
--
作者:
Rick Somers;Sam Cunningham;W. Boles

文献摘要

被引文献

相似文献

在这项研究中,我们应用自然语言处理(NLP)技术,在一个教育环境中,评估它们对于从学生的简短回答中自动评估学生的概念理解的有用性。评估理解能力提供了对学生概念理解的洞察和反馈,这在自动评分中经常被忽视。学生和教育工作者受益于自动化形成性评估,特别是在在线教育和大群体中,在需要时提供对概念理解的见解。我们选择了Electra-Small、Roberta-BASE、XLNet-BASE和Albert-BASE-v2NLP机器学习模型来确定学生辩解的自由文本效度和对他们回答的置信度水平。这两条信息为学生的概念理解和理解的性质提供了关键的见解。我们开发了一个自由文本效度集成,使用高性能的NLP模型来评估学生辩护的有效性,准确率在91.46%到98.66%之间。此外,我们提出了一个通用的、非特定问题的回答置信度模型,该模型可以将回答分为高置信度或低置信度,准确率从93.07%到99.46%不等。随着这些模型适用于小数据集的强劲表现,教育工作者有很好的机会在自己的课堂上实施这些技术。 对实践或政策的影响: 学生的概念理解可以通过自然语言处理从简答式回答中自动准确地提取出来,以评估其理解的水平和性质。 教育工作者和学生可以通过概念理解的自动评估,在需要时收到关于概念理解的反馈,而不需要传统的形成性评估的开销。 教育工作者可以对概念理解模型进行准确的自动化评估,学生回答的简短问题不超过100个。
In this study, we applied natural language processing (NLP) techniques, within an educational environment, to evaluate their usefulness for automated assessment of students’ conceptual understanding from their short answer responses. Assessing understanding provides insight into and feedback on students’ conceptual understanding, which is often overlooked in automated grading. Students and educators benefit from automated formative assessment, especially in online education and large cohorts, by providing insights into conceptual understanding as and when required. We selected the ELECTRA-small, RoBERTa-base, XLNet-base and ALBERT-base-v2 NLP machine learning models to determine the free-text validity of students’ justification and the level of confidence in their responses. These two pieces of information provide key insights into students’ conceptual understanding and the nature of their understanding. We developed a free-text validity ensemble using high performance NLP models to assess the validity of students’ justification with accuracies ranging from 91.46% to 98.66%. In addition, we proposed a general, non-question-specific confidence-in-response model that can categorise a response as high or low confidence with accuracies ranging from 93.07% to 99.46%. With the strong performance of these models being applicable to small data sets, there is a great opportunity for educators to implement these techniques within their own classes. Implications for practice or policy: Students’ conceptual understanding can be accurately and automatically extracted from their short answer responses using NLP to assess the level and nature of their understanding. Educators and students can receive feedback on conceptual understanding as and when required through the automated assessment of conceptual understanding, without the overhead of traditional formative assessment. Educators can implement accurate automated assessment of conceptual understanding models with fewer than 100 student responses for their short response questions.