Predicting academic performance by considering student heterogeneity

Predicting academic performance by considering student heterogeneity
复制标题

DOI:
10.1016/j.knosys.2018.07.042
复制
发表时间:
2018-12-01
影响因子:
8.8
通讯作者:
Long, Qi
Long, Qi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Helal, Sumyea;Li, Jiuyong;Long, Qi

文献摘要

被引文献

相似文献

预测学生学业成绩的能力对于任何旨在提高学生表现和毅力的教育机构都很有价值。根据生成的预测,可以更及时地为被确定存在学业保留或表现风险的学生提供支持。这项研究利用从澳大利亚大学收集的数据创建了不同的分类模型来预测学生的表现。这些数据包括学生注册详细信息以及大学学习管理系统 (LMS) 生成的活动数据。注册数据包含学生信息,例如社会人口特征、大学录取依据(例如通过入学考试或过去的经历)和出勤类型(例如全日制与非全日制)。 LMS 数据记录学生对其在线学习活动的参与度。这项研究的一个重要贡献是在构建预测模型时考虑了学生的异质性。这是基于观察到具有不同社会人口特征或学习模式的学生可能表现出不同的学习动机。实验验证了这样的假设:使用学生子群体中的实例训练的模型优于使用所有数据实例构建的模型。此外,实验表明,考虑入学和课程活动特征有助于更准确地识别弱势学生。实验表明,没有任何一种方法在所有方面都表现出优越的性能。然而,基于规则和基于树的方法生成具有更高可解释性的模型,使它们对于设计有效的学生支持更有用。
The capacity to predict student academic outcomes is of value for any educational institution aiming to improve student performance and persistence. Based on the generated predictions, students identified as being at risk of academic retention or performance can be provided support in a more timely manner. This study creates different classification models for predicting student performance, using data collected from an Australian university. The data include student enrolment details as well as the activity data generated from the university learning management system (LMS). The enrolment data contain student information such as socio-demographic features, university admission basis (e.g. via entry exam or past experience) and attendance type (e.g. full-time vs. part-time). The LMS data record student engagement with their online learning activities. An important contribution of this study is the consideration of student heterogeneity in constructing the predictive models. This is based on the observation that students with different socio-demographic features or study modes may exhibit varying learning motivations. The experiments validated the hypothesis that the models trained with instances in student sub-populations outperform those constructed using all data instances. Furthermore, the experiments revealed that considering both enrolment and course activity features aids in identifying vulnerable students more precisely. The experiments determined that no individual method exhibits superior performance in all aspects. However, the rule-based and tree-based methods generate models with higher interpretability, making them more useful for designing effective student support.