Examining the Effect of Assessment Construct Characteristics on Machine Learning Scoring of Scientific Argumentation

Examining the Effect of Assessment Construct Characteristics on Machine Learning Scoring of Scientific Argumentation
复制标题

DOI:
10.1007/s40593-023-00385-8
复制
发表时间:
2023-12-18
影响因子:
4.9
通讯作者:
Zhai,Xiaoming
Zhai,Xiaoming
中科院分区:
其他
文献类型:
--
作者:
Haudek,Kevin C.;Zhai,Xiaoming

文献摘要

相似文献

论证是 K-12 科学教育框架中提出的一项关键科学实践,要求学生构建和批判论证,但在大规模课堂上及时评估论证具有挑战性。最近的工作显示了自动评分系统在开放式反应评估中的潜力,利用机器学习 (ML) 和人工智能 (AI) 来帮助对复杂评估中的书面论点进行评分。此外,研究已经证实,评估结构的特征(即复杂性、多样性和结构)对于机器学习评分准确性至关重要,但评估结构如何与机器评分准确性相关联仍然未知。本研究调查了与科学论证评估项目的评估结构相关的特征如何影响机器评分性能。具体来说,我们从三个维度概念化了该结构:复杂性、多样性和结构。我们聘请人类专家对评估任务的特征进行编码,并对中学生对 17 项论证任务的反应进行评分,这些任务与经过验证的科学论证学习进程的三个级别相对应。我们随机选择了 361 个响应作为训练集,为每个项目构建机器学习评分模型。评分模型与人类共识评分达成了一系列一致,通过 Cohen 的 kappa 测量(平均值 = 0.60;范围 0.38 - 0.89),表明表现良好到几乎完美。我们发现,评估任务的复杂性和多样性水平较高与模型性能下降相关,同样,结构水平和模型性能之间的关系呈现出某种负线性趋势。这些发现强调了在开发用于评分评估的机器学习模型时考虑这些构造特征的重要性,特别是对于较高复杂性的项目和多维评估。
Argumentation, a key scientific practice presented in theFramework for K-12 Science Education, requires students to construct and critique arguments, but timely evaluation of arguments in large-scale classrooms is challenging. Recent work has shown the potential of automated scoring systems for open response assessments, leveraging machine learning (ML) and artificial intelligence (AI) to aid the scoring of written arguments in complex assessments. Moreover, research has amplified that the features (i.e., complexity, diversity, and structure) of assessment construct are critical to ML scoring accuracy, yet how the assessment construct may be associated with machine scoring accuracy remains unknown. This study investigated how the features associated with the assessment construct of a scientific argumentation assessment item affected machine scoring performance. Specifically, we conceptualized the construct in three dimensions: complexity, diversity, and structure. We employed human experts to code characteristics of the assessment tasks and score middle school student responses to 17 argumentation tasks aligned to three levels of a validated learning progression of scientific argumentation. We randomly selected 361 responses to use as training sets to build machine-learning scoring models for each item. The scoring models yielded a range of agreements with human consensus scores, measured by Cohen’s kappa (mean = 0.60; range 0.38 − 0.89), indicating good to almost perfect performance. We found that higher levels ofComplexityandDiversityof the assessment task were associated with decreased model performance, similarly the relationship between levels ofStructureand model performance showed a somewhat negative linear trend. These findings highlight the importance of considering these construct characteristics when developing ML models for scoring assessments, particularly for higher complexity items and multidimensional assessments.