Multi-Armed Bandits for Intelligent Tutoring Systems

Multi-Armed Bandits for Intelligent Tutoring Systems
复制标题

DOI:
10.5281/zenodo.3554667
复制
发表时间:
2013-10
期刊:
ArXiv
影响因子:
--
通讯作者:
B. Clément;Didier Roy;Pierre-Yves Oudeyer;M. Lopes
B. Clément;Didier Roy;Pierre-Yves Oudeyer;M. Lopes
中科院分区:
其他
文献类型:
--
作者:
B. Clément;Didier Roy;Pierre-Yves Oudeyer;M. Lopes

文献摘要

被引文献

相似文献

我们提出了一种智能辅导系统的方法,自适应个性化的学习活动序列,以最大限度地提高学生获得的技能,考虑到有限的时间和激励资源。在给定的时间点,系统向学生提出使他们进步更快的活动。我们介绍了两种算法,依赖于经验估计的学习进度,RiARiT使用的信息,每个练习的难度和ZPDES,使用少得多的知识的问题。该系统基于三种方法的结合。首先,它利用了最近的内在动机的学习模式,将其转换为积极的教学,依靠经验估计的学习进展提供的具体活动,以特定的学生。其次,它使用最先进的多臂Bandit(MAB)技术来有效地管理这个优化过程的探索/开发挑战。第三,它利用专家知识来约束和引导MAB的初始探索,同时只需要专家的粗略指导信息,并允许系统处理其知识中的教学差距。该系统在一个场景中进行评估,其中7-8岁的学童学习如何在操纵金钱的同时分解数字。系统的实验与模拟学生,然后在400名学童的人口的用户研究的结果。
We present an approach to Intelligent Tutoring Systems which adaptively personalizes sequences of learning activities to maximize skills acquired by students, taking into account the limited time and motivational resources. At a given point in time, the system proposes to the students the activity which makes them progress faster. We introduce two algorithms that rely on the empirical estimation of the learning progress, RiARiT that uses information about the difficulty of each exercise and ZPDES that uses much less knowledge about the problem. The system is based on the combination of three approaches. First, it leverages recent models of intrinsically motivated learning by transposing them to active teaching, relying on empirical estimation of learning progress provided by specific activities to particular students. Second, it uses state-of-the-art Multi-Arm Bandit (MAB) techniques to efficiently manage the exploration/exploitation challenge of this optimization process. Third, it leverages expert knowledge to constrain and bootstrap initial exploration of the MAB, while requiring only coarse guidance information of the expert and allowing the system to deal with didactic gaps in its knowledge. The system is evaluated in a scenario where 7-8 year old schoolchildren learn how to decompose numbers while manipulating money. Systematic experiments are presented with simulated students, followed by results of a user study across a population of 400 school children.