Automated Scoring for Reading Comprehension via In-context BERT Tuning

Automated Scoring for Reading Comprehension via In-context BERT Tuning
复制标题

通过上下文 BERT 调优对阅读理解进行自动评分

DOI:
10.1007/978-3-031-11644-5_69
复制
发表时间:
2022
期刊:
International Conference on Artificial Intelligence in Education
影响因子:
--
通讯作者:
Lan, Andrew S.
Lan, Andrew S.
中科院分区:
--
文献类型:
--
作者:
Fernandez, Nigel;Ghosh, Aritra;Liu, Naiming;Wang, Zichao;Choffin, Benoit;Baraniuk, Richard G.;Lan, Andrew S.

文献摘要

参考文献

被引文献

相似文献

对开放式学生回答的自动评分有可能大大减少人类评分的工作量。自动评分的最新进展利用了来自BERT等预训练语言模型的文本表示。现有的方法为每个项目/问题训练一个单独的模型,适用于像文章评分这样的场景,其中项目可能彼此不同。然而,这些方法有两个局限性:1)它们不能在阅读理解等场景中利用项目链接,其中多个项目可能共享一篇阅读文章;2)它们是不可伸缩的,因为在大型语言模型中,为每个项目存储一个模型是很困难的。我们报告了我们的(大奖获奖)解决方案,以国家教育进步评估(NAEP)阅读理解的自动评分挑战。我们的方法是上下文BERT微调,为所有项目生成一个单一的共享评分模型,该模型具有精心设计的输入结构,以提供每个项目的上下文信息。我们的实验证明了我们的方法优于现有方法的有效性。我们还进行了定性分析,并讨论了我们的方法的局限性。(论文的完整版本可在https://arxiv.org/abs/2205.09864找到,我们的实现可在https://github.com/ni9elf/automated-scoring找到)
Automated scoring of open-ended student responses has the potential to significantly reduce human grader effort. Recent advances in automated scoring leverage textual representations from pre-trained language models like BERT. Existing approaches train a separate model for each item/question, suitable for scenarios like essay scoring where items can be different from one another. However, these approaches have two limitations: 1) they fail to leverage item linkage for scenarios such as reading comprehension where multiple items may share a reading passage; 2) they are not scalable since storing one model per item is difficult with large language models. We report our (grand prize-winning) solution to the National Assessment of Education Progress (NAEP) automated scoring challenge for reading comprehension. Our approach, in-context BERT fine-tuning, produces a single shared scoring model for all items with a carefully designed input structure to provide contextual information on each item. Our experiments demonstrate the effectiveness of our approach which outperforms existing methods. We also perform a qualitative analysis and discuss the limitations of our approach. (Full version of the paper can be found at: https://arxiv.org/abs/2205.09864 Our implementation can be found at: https://github.com/ni9elf/automated-scoring)
DOI: 10.18653/v1/2020.coling-main.535
发表时间: 2020-12
期刊: --
影响因子: --
作者:
Masaki Uto;Yikuan Xie;M. Ueno
通讯作者: Masaki Uto;Yikuan Xie;M. Ueno
DOI: 10.1007/978-3-030-52240-7_61
发表时间: 2020-06-10
期刊: Artificial Intelligence in Education
影响因子: --
作者:
Uto M;Uchida Y
通讯作者: Uchida Y
改进学生数学开放式回答的自动评分
DOI: --
发表时间: 2021
期刊: Educational Data Mining EDM 2021
影响因子: --
作者:
Baral, Sami;Botelho, Anthony;Erickson, John;Benachamardi, Priyanka;Heffernan, Neil
通讯作者: Heffernan, Neil
大学辍学预测模型是否应该包含受保护的属性?
DOI: --
发表时间: 2021
期刊: ACM Conference on Learning @ Scale
影响因子: --
作者:
Renzhe Yu;Hansol Lee;René F. Kizilcec
通讯作者: René F. Kizilcec
写作中文本证据使用自动评分方法的公平性评估
DOI: --
发表时间: 2021
期刊: International Conference on Artificial Intelligence in Education
影响因子: --
作者:
D. Litman;Haoran Zhang;R. Correnti;L. Matsumura;E. Wang
通讯作者: E. Wang