On the Validity of Machine Learning-based Next Generation Science Assessments: A Validity Inferential Network

On the Validity of Machine Learning-based Next Generation Science Assessments: A Validity Inferential Network
复制标题

基于机器学习的下一代科学评估的有效性:有效性推理网络

DOI:
10.1007/s10956-020-09879-9
复制
发表时间:
2021
影响因子:
4.4
通讯作者:
J. Pellegrino
J. Pellegrino
中科院分区:
教育学2区
文献类型:
--
作者:
X. Zhai;J. Krajcik;J. Pellegrino

文献摘要

参考文献

被引文献

相似文献

这项研究提供了一个可靠的有效性推理网络,以指导基于机器学习的下一代科学评估(NGSA)的开发、解释和使用。鉴于机器学习(ML)已被广泛应用于构建响应,论文,模拟,教育游戏和跨学科评估的自动评分,以推进学生科学学习的证据收集和推理,我们认为,由于ML的参与,科学评估会出现额外的有效性问题。这些新出现的有效性问题可能无法通过为非科学或非ML评估开发的先前有效性框架来解决。因此,我们研究了ML给科学评估带来的变化,并确定了基于ML的NGSA的七个关键有效性问题:歪曲感兴趣结构的潜在风险,可能涉及更多变量的潜在混杂因素,分数的解释和使用与设计的学习目标之间的不一致,分数的解释和使用与实际学习质量之间的不一致,机器分数与规则之间的不一致,机器算法模型的推广能力有限,机器算法模型的外推能力有限。基于确定的七个有效性问题,我们提出了一个有效性推理网络,以解决基于ML的NGSA的认知,教学和推理有效性。为了演示该网络的实用性,我们展示了使用七步ML框架开发的基于ML的下一代科学评估的示例。我们阐述了我们如何使用有效性推理网络,以确保负责任的评估设计,以及有效的解释和使用机器分数。
This study provides a solid validity inferential network to guide the development, interpretation, and use of machine learning-based next-generation science assessments (NGSAs). Given that machine learning (ML) has been broadly implemented in the automatic scoring of constructed responses, essays, simulations, educational games, and interdisciplinary assessments to advance the evidence collection and inference of student science learning, we contend that additional validity issues arise for science assessments due to the involvement of ML. These emerging validity issues may not be addressed by prior validity frameworks developed for either non-science or non-ML assessments. We thus examine the changes brought in by ML to science assessments and identify seven critical validity issues of ML-based NGSAs: potential risk of misrepresenting the construct of interest, potential confounders due to that more variables may involve, nonalignment between interpretation and use of scores and designed learning goals, nonalignment between interpretation and use of scores and actual learning quality, nonalignment between machine scores and rubrics, limited generalizable ability of machine algorithmic models, and limited extrapolating ability of machine algorithmic models. Based on the seven validity issues identified, we propose a validity inferential network to address the cognitive, instructional, and inferential validity of ML-based NGSAs. To demonstrate the utility of this network, we present an exemplar of ML-based next-generation science assessments that was developed using a seven-step ML framework. We articulate how we used the validity inferential network to ensure accountable assessment design, as well as valid interpretation and use of machine scores.
DOI: 10.1148/rg.2017160130
发表时间: 2017-03
期刊: Radiographics : a review publication of the Radiological Society of North America, Inc
影响因子: --
作者:
Erickson BJ;Korfiatis P;Akkus Z;Kline TL
通讯作者: Kline TL
DOI: 10.1007/s11412-019-09298-y
发表时间: 2019-09-01
影响因子: 4.3
作者:
Gerard, Libby;Kidron, Ady;Linn, Marcia C.
通讯作者: Linn, Marcia C.
使用分析和整体编码方法在与科学学习进展相一致的构建响应评估中比较机器学习性能
DOI: 10.1007/s10956-020-09858-0
发表时间: 2020
影响因子: 4.4
作者:
Jescovitch, Lauren N.;Scott, Emily E.;Cerchiara, Jack A.;Merrill, John;Urban-Lurain, Mark;Doherty, Jennifer H.;Haudek, Kevin C.
通讯作者: Haudek, Kevin C.