Using Past Data to Warm Start Active Machine Learning: Does Context Matter?

Using Past Data to Warm Start Active Machine Learning: Does Context Matter?
复制标题

使用过去的数据来热启动主动机器学习:上下文重要吗?

DOI:
10.1145/3448139.3448154
复制
发表时间:
2021
期刊:
LAK21: 11th International Learning Analytics and Knowledge Conference
影响因子:
--
通讯作者:
Heffernan, Neil
Heffernan, Neil
中科院分区:
--
文献类型:
--
作者:
Karumbaiah, Shamya;Lan, Andrew;Nagpal, Sachit;Baker, Ryan S.;Botelho, Anthony;Heffernan, Neil

文献摘要

参考文献

被引文献

相似文献

尽管学生在虚拟学习环境中的活动产生了丰富的数据,但在学习分析中使用监督机器学习受到标记数据可用性的限制,这些数据可能难以收集复杂的教育结构。在以前的研究中,机器学习的一个子领域称为主动学习(AL),旨在提高数据标记效率。AL训练一个模型,并并行使用它来选择下一个数据样本,以从人类专家那里获得标记。由于教育结构和数据的复杂性,AL遭受了冷启动问题,即模型无法获得足够的数据来选择最佳的下一个样本进行学习。在本文中,我们探讨了使用过去的数据来热启动AL训练过程。我们还批判性地研究了过去数据收集的不同背景(城市化)的影响。为此,我们使用真实的影响标签,通过在中学数学课堂上的人类观察收集到的,以模拟开发的AL为基础的检测器从事浓度。我们实验了两种AL方法(不确定性采样,L-MMSE)和随机采样的数据选择。我们的研究结果表明,使用过去的数据来热启动AL培训可能是有效的一些方法的基础上,目标人群的城市化。我们提供了建议的数据选择方法和数量的过去的数据使用时,热启动AL培训在城市和郊区的学校。
Despite the abundance of data generated from students’ activities in virtual learning environments, the use of supervised machine learning in learning analytics is limited by the availability of labeled data, which can be difficult to collect for complex educational constructs. In a previous study, a subfield of machine learning called Active Learning (AL) was explored to improve the data labeling efficiency. AL trains a model and uses it, in parallel, to choose the next data sample to get labeled from a human expert. Due to the complexity of educational constructs and data, AL has suffered from the cold-start problem where the model does not have access to sufficient data yet to choose the best next sample to learn from. In this paper, we explore the use of past data to warm start the AL training process. We also critically examine the implications of differing contexts (urbanicity) in which the past data was collected. To this end, we use authentic affect labels collected through human observations in middle school mathematics classrooms to simulate the development of AL-based detectors of engaged concentration. We experiment with two AL methods (uncertainty sampling, L-MMSE) and random sampling for data selection. Our results suggest that using past data to warm start AL training could be effective for some methods based on the target population's urbanicity. We provide recommendations on the data selection method and the quantity of past data to use when warm starting AL training in the urban and suburban schools.
主动学习用于学生情绪检测
DOI: --
发表时间: 2019
期刊: Educational Data Mining
影响因子: --
作者:
Tsung;R. Baker;Christoph Studer;N. Heffernan;Andrew S. Lan
通讯作者: Andrew S. Lan
教育数据挖掘模型的总体有效性:情感检测的案例研究
DOI: 10.1111/bjet.12156
发表时间: 2014
期刊: Br. J. Educ. Technol.
影响因子: --
作者:
Jaclyn L. Ocumpaugh;R. Baker;S. M. Gowda;N. Heffernan;Cristina Heffernan
通讯作者: Cristina Heffernan
学习情感动态数据的重新分析和综合
DOI: 10.1109/taffc.2021.3086118
发表时间: 2023
影响因子: 11.2
作者:
Shamya Karumbaiah;R. Baker;Jaclyn L. Ocumpaugh;J. Andres
通讯作者: J. Andres
DOI: 10.1504/ijlt.2009.028807
发表时间: 2009
期刊: Int. J. Learn. Technol.
影响因子: --
作者:
Scott W. McQuiggan;James C. Lester
通讯作者: James C. Lester
基于计算机的学习系统中的文化:挑战和机遇
DOI: 10.35542/osf.io/ad39g
发表时间: 2020
期刊: 2023 6th International Conference on Information and Communications Technology (ICOIACT)
影响因子: --
作者:
R. Baker;Erin Walker;A. Ogan;Michael A. Madaio
通讯作者: Michael A. Madaio