Comparison of Machine Learning Performance Using Analytic and Holistic Coding Approaches Across Constructed Response Assessments Aligned to a Science Learning Progression

Comparison of Machine Learning Performance Using Analytic and Holistic Coding Approaches Across Constructed Response Assessments Aligned to a Science Learning Progression
复制标题

使用分析和整体编码方法在与科学学习进展相一致的构建响应评估中比较机器学习性能

DOI:
10.1007/s10956-020-09858-0
复制
发表时间:
2020
影响因子:
4.4
通讯作者:
Haudek, Kevin C.
Haudek, Kevin C.
中科院分区:
教育学2区
文献类型:
--
作者:
Jescovitch, Lauren N.;Scott, Emily E.;Cerchiara, Jack A.;Merrill, John;Urban-Lurain, Mark;Doherty, Jennifer H.;Haudek, Kevin C.

文献摘要

参考文献

被引文献

相似文献

我们系统地比较了两种编码方法来生成机器学习(ML)的训练数据集:(i)基于学习进展水平的整体方法和(ii)学生推理中多个概念的二分法,分析方法,从整体规则中解构。我们评估了四个构建的反应评估项目,本科生理学,每个目标的离子背景下发展中的通量学习进展的五个层次。使用人类编码的数据集来训练两个ML模型:(i)在构造响应分类器(CRC)中实现的8分类算法集成,以及(ii)在LightSide Researcher的数据库中实现的单一分类算法。人类编码协议约700名学生的反应,每个项目是高的两种方法与科恩的kappas范围从0.75至0.87的整体评分和从0.78至0.89的分析综合评分。ML模型性能因项目和题目类型而异。对于两个项目,来自两种编码方法的训练集产生了类似准确的ML模型,机器和人类得分之间的Cohen kappa差异为0.002和0.041。对于其他项目,与使用整体分数进行训练相比,使用分析编码响应训练并用于综合分数的ML模型实现了更好的性能,Cohen的kappa增加了0.043和0.117。这些项目使用了一个更复杂的场景,涉及两个离子的运动。分析性编码可能有助于解开这种额外的复杂性。
We systematically compared two coding approaches to generate training datasets for machine learning (ML): (i) a holistic approach based on learning progression levels and (ii) a dichotomous, analytic approach of multiple concepts in student reasoning, deconstructed from holistic rubrics. We evaluated four constructed response assessment items for undergraduate physiology, each targeting five levels of a developing flux learning progression in an ion context. Human-coded datasets were used to train two ML models: (i) an 8-classification algorithm ensemble implemented in the Constructed Response Classifier (CRC), and (ii) a single classification algorithm implemented in LightSide Researcher’s Workbench. Human coding agreement on approximately 700 student responses per item was high for both approaches with Cohen’s kappas ranging from 0.75 to 0.87 on holistic scoring and from 0.78 to 0.89 on analytic composite scoring. ML model performance varied across items and rubric type. For two items, training sets from both coding approaches produced similarly accurate ML models, with differences in Cohen’s kappa between machine and human scores of 0.002 and 0.041. For the other items, ML models trained with analytic coded responses and used for a composite score, achieved better performance as compared to using holistic scores for training, with increases in Cohen’s kappa of 0.043 and 0.117. These items used a more complex scenario involving movement of two ions. It may be that analytic coding is beneficial to unpacking this additional complexity.
DOI: 10.1007/s11412-019-09298-y
发表时间: 2019-09-01
影响因子: 4.3
作者:
Gerard, Libby;Kidron, Ady;Linn, Marcia C.
通讯作者: Linn, Marcia C.
DOI: 10.1080/00461520.2012.695709
发表时间: 2012-01-01
影响因子: 8.8
作者:
Chi, Michelene T. H.;VanLehn, Kurt A.
通讯作者: VanLehn, Kurt A.
DOI: 10.1111/emip.12028
发表时间: 2014-06-01
影响因子: 2
作者:
Liu, Ou Lydia;Brew, Chris;Linn, Marcia C.
通讯作者: Linn, Marcia C.
DOI: --
发表时间: 1983
期刊:
影响因子: --
作者:
D. Newble;R. Cannon
通讯作者: R. Cannon
设计电子评估:使用多项选择测试取得良好效果
DOI: --
发表时间: 2007
期刊:
影响因子: --
作者:
D. Nicol
通讯作者: D. Nicol