Contrasting State-of-the-Art in the Machine Scoring of Short-Form Constructed Responses

Contrasting State-of-the-Art in the Machine Scoring of Short-Form Constructed Responses
复制标题

对比简短构建响应的机器评分的最新技术

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
M. Shermis
M. Shermis
中科院分区:
--
文献类型:
--
作者:
M. Shermis

文献摘要

被引文献

相似文献

本研究比较了由人类评分员和机器评分算法评估的短形式构造的反应。当时的背景是一场公开竞赛,公开竞争对手和商业供应商都在竞相开发机器评分算法,这些算法将在总结性高风险测试环境中匹配或超过操作人类评分员的表现。数据(N = 25,683)来自三个不同的州,采用10种不同的提示,并从两个不同的中学年级水平。从各州提供的数据集中随机选择了2,130到2,999个样本,然后随机分为三组:训练集,测试集和验证集。机器在所有一致性指标上的表现都未能与人类评分员相匹配。目前的研究得出了一些建议,这些建议可能会改善机器评分算法,然后才能以任何操作方式使用。
This study compared short-form constructed responses evaluated by both human raters and machine scoring algorithms. The context was a public competition on which both public competitors and commercial vendors vied to develop machine scoring algorithms that would match or exceed the performance of operational human raters in a summative high-stakes testing environment. Data (N = 25,683) were drawn from three different states, employed 10 different prompts, and were drawn from two different secondary grade levels. Samples ranging in size from 2,130 to 2,999 were randomly selected from the data sets provided by the states and then randomly divided into three sets: a training set, a test set, and a validation set. Machine performance on all of the agreement measures failed to match that of the human raters. The current study concluded with recommendations on steps that might improve machine-scoring algorithms before they can be used in any operational way.