Classifying Math Knowledge Components via Task-Adaptive Pre-Trained BERT

Classifying Math Knowledge Components via Task-Adaptive Pre-Trained BERT
复制标题

通过任务自适应预训练 BERT 对数学知识成分进行分类

DOI:
10.1007/978303078292433
复制
发表时间:
2021
期刊:
22nd International Conference on Artificial Intelligence in Education.
影响因子:
--
通讯作者:
Shen, J.T.
Shen, J.T.
中科院分区:
--
文献类型:
--
作者:
Shen, J.T.

文献摘要

相似文献

标有适当知识成分 (KC) 的教育内容对于教师或内容组织者特别有用。然而,手动标记教育内容是劳动密集型的并且容易出错。为了应对这一挑战,先前的研究提出了基于机器学习的解决方案来自动标记教育内容,但成效有限。在这项工作中,我们通过以下方式显着改进了先前的研究:(1)扩展输入类型以包括 KC 描述、教学视频标题和问题描述(即三种类型的预测任务),(2)将预测的粒度从 198 个 KC 标签加倍到 385 个 KC 标签(即更实际的设置,但更难的多项式分类问题),(3)使用任务自适应将预测精度提高 0.5-2.3%预训练的 BERT,优于 6 个基线,并且 (4) 提出了一种简单的评估措施,通过该措施我们可以恢复 56-73% 的错误预测 KC 标签。实验中的所有代码和数据集均可在:https://github.com/tbs17/TAPT-BERT
Educational content labeled with proper knowledge components (KCs) are particularly useful to teachers or content organizers. However, manually labeling educational content is labor intensive and error-prone. To address this challenge, prior research proposed machine learning based solutions to auto-label educational content with limited success. In this work, we significantly improve prior research by (1) expanding the input types to include KC descriptions, instructional video titles, and problem descriptions (i.e., three types of prediction task), (2) doubling the granularity of the prediction from 198 to 385 KC labels (i.e., more practical setting but much harder multinomial classification problem), (3) improving the prediction accuracies by 0.5–2.3% using Task-adaptive Pre-trained BERT, outperforming six baselines, and (4) proposing a simple evaluation measure by which we can recover 56–73% of mispredicted KC labels. All codes and data sets in the experiments are available at: https://github.com/tbs17/TAPT-BERT