Contrasting State-of-the-Art in the Machine Scoring of Short-Form Constructed Responses
Contrasting State-of-the-Art in the Machine Scoring of Short-Form Constructed Responses
复制标题
对比简短构建响应的机器评分的最新技术
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
M. Shermis
中科院分区:
文献类型:
--
作者:
M. Shermis
This study compared short-form constructed responses evaluated by both human raters and machine scoring algorithms. The context was a public competition on which both public competitors and commercial vendors vied to develop machine scoring algorithms that would match or exceed the performance of operational human raters in a summative high-stakes testing environment. Data (N = 25,683) were drawn from three different states, employed 10 different prompts, and were drawn from two different secondary grade levels. Samples ranging in size from 2,130 to 2,999 were randomly selected from the data sets provided by the states and then randomly divided into three sets: a training set, a test set, and a validation set. Machine performance on all of the agreement measures failed to match that of the human raters. The current study concluded with recommendations on steps that might improve machine-scoring algorithms before they can be used in any operational way.