Using Machine Learning to Score Multi-Dimensional Assessments of Chemistry and Physics

Using Machine Learning to Score Multi-Dimensional Assessments of Chemistry and Physics
复制标题

使用机器学习对化学和物理的多维评估进行评分

DOI:
10.1007/s10956-020-09895-9
复制
发表时间:
2021
影响因子:
4.4
通讯作者:
J. Krajcik
J. Krajcik
中科院分区:
教育学2区
文献类型:
--
作者:
Sarah Maestrales;X. Zhai;Israel Touitou;Quinton Baker;Barbara Schneider;J. Krajcik

文献摘要

参考文献

被引文献

相似文献

为了响应促进三维科学学习的号召(NRC,2012),研究人员主张开发超越死记硬背任务的评估项目,而是需要更深入理解和使用推理来提高科学素养的评估项目。此类评估项目通常是基于表现的构建反应,需要技术参与来减轻教师的评分负担。本研究通过检查机器学习文本分析协议的使用和准确性来响应这一呼吁,作为对构建的响应项目进行人工评分的替代方案。我们使用的项目代表了 2012 年 NRC 报告中阐述的科学学习的多个维度。我们使用 6700 名化学和物理专业学生的 26,000 多个构建答案样本,培训了人类评分员并编制了强大的训练集来开发机器算法模型并交叉验证机器分数。结果表明,人类评估者在不同维度数量的评估项目上产生了良好(Cohen’s k = .40–.75)到优秀(Cohen’s k > .75)的评估者间可靠性。比较表明,机器评分算法在这些相同项目上达到了与人类评分者相当的评分准确性。结果还表明,使用正式词汇(例如速度)的反应可能会产生较低的机器与人类的一致性,这可能与以下事实有关:与非正式的替代方案相比,使用正式短语的学生较少。
In response to the call for promoting three-dimensional science learning (NRC, 2012), researchers argue for developing assessment items that go beyond rote memorization tasks to ones that require deeper understanding and the use of reasoning that can improve science literacy. Such assessment items are usually performance-based constructed responses and need technology involvement to ease the burden of scoring placed on teachers. This study responds to this call by examining the use and accuracy of a machine learning text analysis protocol as an alternative to human scoring of constructed response items. The items we employed represent multiple dimensions of science learning as articulated in the 2012 NRC report. Using a sample of over 26,000 constructed responses taken by 6700 students in chemistry and physics, we trained human raters and compiled a robust training set to develop machine algorithmic models and cross-validate the machine scores. Results show that human raters yielded good (Cohen’s k = .40–.75) to excellent (Cohen’s k > .75) interrater reliability on the assessment items with varied numbers of dimensions. A comparison reveals that the machine scoring algorithms achieved comparable scoring accuracy to human raters on these same items. Results also show that responses with formal vocabulary (e.g., velocity) were likely to yield lower machine-human agreements, which may be associated with the fact that fewer students employed formal phrases compared with the informal alternatives.
计算机文本分析:促进学习的评估和研究潜力
DOI: --
发表时间: 2019
期刊: Computersupported collaborative learning
影响因子: --
作者:
Lee, H.-S.
通讯作者: Lee, H.-S.