Better Data Beat Big Data

Better Data Beat Big Data
复制标题

更好的数据击败大数据

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Ambarish Joshi
Ambarish Joshi
中科院分区:
--
文献类型:
--
作者:
M. Yudelson;Stephen E. Fancsali;Steven Ritter;Susan R. Berman;Tristan Nixon;Ambarish Joshi

文献摘要

被引文献

相似文献

学生学习模式的普遍性是一个非常理想的特征。随着新生与教育系统的互动,高度预测的模型,调整到以前学习者的数据量越来越多,大概允许这样的系统提供一个更个性化的,最佳的学习路径,提供更好的反馈,并提供更有效的学习体验。然而,任何大的学生/用户群体将是异质的,并且可能由特定学习模型可能适合的可辨别的子群体组成。学生亚群可能在认知因素、教学水平和质量以及许多其他环境和非认知因素方面有所不同。“大数据”和广泛部署的教育软件(包括卡内基学习的认知导师(CLCT)智能辅导系统)的时代,为分析学习者与教育系统交互期间收集的越来越多的数据提供了机会。这些数据涵盖了广泛的学习者,使研究人员能够调查越来越具有代表性的学生群体的结构。在这项工作中,我们调查发现学生子群体从“大数据”。使用CLCT一年的数据,我们检验了一个假设,即常用的学生亚群分层(例如,学校位置、社会人口因素)提供了有意义地划分学生的方法。我们发现,与其寻找应该被区别对待的不同子群体,不如特定的学习者子群体提供特别“高质量”的数据,并且从该子群体中学习的模型甚至在预测其他模型所训练的子群体的学生学习时也优于所有其他模型。这样,“更好的数据打败了大数据”。
Generalizability of models of student learning is a highly desirable feature. As new students interact with educational systems, highly predictive models, tuned to increasing amounts of data from previous learners, presumably allow such systems to provide a more individualized, optimal learning path, give better feedback, and provide a more effective learning experience. However, any large student/user population will be heterogeneous and likely consist of discernable sub-populations for which specific models of learning may be appropriate. Student subpopulations may differ with respect to cognitive factors, the level and quality of instruction, and many other environmental and noncognitive factors. The era of both “big data” and widely deployed educational software, including Carnegie Learning’s Cognitive Tutor (CLCT) intelligent tutoring system, presents opportunities to analyze increasingly large volumes of data collected during learners’ interactions with educational systems. These data cover a broad spectrum of learners, allowing researchers to investigate the structure of an increasingly representative student population. In this work, we investigate discovering student sub-populations from “big data.” Using a year’s worth of data from CLCT, we test the hypothesis that commonly used stratifications of student subpopulations (e.g., school location, socio-demographic factors) offer ways to meaningfully partition learners. We discover that, rather than finding distinct subpopulations that should be treated differently, a particular sub-population of learners provides especially “high quality” data and that models learned from this sub-population outperform all other models even when predicting student learning for the sub-population on which other models were trained. In this way, “better data beat big data.”