Automated Scoring of Constructed-Response Science Items: Prospects and Obstacles

Automated Scoring of Constructed-Response Science Items: Prospects and Obstacles
复制标题

DOI:
10.1111/emip.12028
复制
发表时间:
2014-06-01
影响因子:
2
通讯作者:
Linn, Marcia C.
Linn, Marcia C.
中科院分区:
教育学4区
文献类型:
--
作者:
Liu, Ou Lydia;Brew, Chris;Linn, Marcia C.

文献摘要

被引文献

相似文献

基于内容的自动评分已被应用于各种科学领域。然而,许多先前的应用涉及简化的评分规则,而不考虑表示多个理解水平的规则。本研究测试了一个基于概念的评分工具,基于内容的评分,c-rater(TM),为四个科学项目的标题,旨在区分多个层次的理解。这些项目与人类评分显示出中度至良好的一致性。研究结果表明,自动评分有可能对具有复杂评分规则的构建反应项目进行评分,但在其目前的设计中无法取代人类评分员。本文讨论了分歧的来源和可能提高基于概念的自动评分准确性的因素。
Content-based automated scoring has been applied in a variety of science domains. However, many prior applications involved simplified scoring rubrics without considering rubrics representing multiple levels of understanding. This study tested a concept-based scoring tool for content-based scoring, c-rater (TM), for four science items with rubrics aiming to differentiate among multiple levels of understanding. The items showed moderate to good agreement with human scores. The findings suggest that automated scoring has the potential to score constructed-response items with complex scoring rubrics, but in its current design cannot replace human raters. This article discusses sources of disagreement and factors that could potentially improve the accuracy of concept-based automated scoring.