A Meta-Analysis of Machine Learning-Based Science Assessments: Factors Impacting Machine-Human Score Agreements

A Meta-Analysis of Machine Learning-Based Science Assessments: Factors Impacting Machine-Human Score Agreements
复制标题

DOI:
10.1007/s10956-020-09875-z
复制
发表时间:
2020-11-19
影响因子:
4.4
通讯作者:
Nehm, Ross H.
Nehm, Ross H.
中科院分区:
教育学2区
文献类型:
--
作者:
Zhai, Xiaoming;Shi, Lehong;Nehm, Ross H.

文献摘要

被引文献

相似文献

机器学习(ML)已越来越多地用于科学评估,以促进自动评分工作,尽管取得了不同程度的成功(即机器-人评分协议的大小[MHAs])。在这个不断发展的领域,很少有研究对影响MHA差异的因素进行实证研究,从而限制了机器评分能力的提高及其在科学教育中的广泛应用。我们对110项mha研究进行了荟萃分析,以确定对评分成功最重要的因素(即高科恩kappa [kappa])。我们实证研究了影响MHA大小的六个因素:算法、学科领域、评估格式、结构、学校水平和机器监督类型。我们对110个MHAs的分析显示kappa存在很大的异质性(考虑到权重,平均值= 0.64,范围= 0.09 - 0.97)。采用三水平随机效应模型,MHA评分异质性可以由出版物内部(即评估任务水平:82.6%)和出版物之间(即个体研究水平:16.7%)的变异性来解释。我们的研究结果还表明,这六个因素对得分成功程度有显著的调节作用。其中,算法和学科领域的影响显著大于其他因素,表明技术特征和评估外部特征可能是改进MHAs和基于ml的科学评估的主要目标。
Machine learning (ML) has been increasingly employed in science assessment to facilitate automatic scoring efforts, although with varying degrees of success (i.e., magnitudes of machine-human score agreements [MHAs]). Little work has empirically examined the factors that impact MHA disparities in this growing field, thus constraining the improvement of machine scoring capacity and its wide applications in science education. We performed a meta-analysis of 110 studies of MHAs in order to identify the factors most strongly contributing to scoring success (i.e., high Cohen's kappa [kappa]). We empirically examined six factors proposed as contributors to MHA magnitudes: algorithm, subject domain, assessment format, construct, school level, and machine supervision type. Our analyses of 110 MHAs revealed substantial heterogeneity in kappa(mean=.64; range = .09-.97, taking weights into consideration). Using three-level random-effects modeling, MHA score heterogeneity was explained by the variability both within publications (i.e., the assessment task level: 82.6%) and between publications (i.e., the individual study level: 16.7%). Our results also suggest that all six factors have significant moderator effects on scoring success magnitudes. Among these, algorithm and subject domain had significantly larger effects than the other factors, suggesting that technical features and assessment external features might be primary targets for improving MHAs and ML-based science assessments.