Understanding Mean Score Differences Between the e‐rater® Automated Scoring Engine and Humans for Demographically Based Groups in the GRE® General Test

Understanding Mean Score Differences Between the e‐rater® Automated Scoring Engine and Humans for Demographically Based Groups in the GRE® General Test
复制标题

了解 GRE® 普通考试中基于人口统计的群体的 e‐rater® 自动评分引擎与人类之​​间的平均分数差异

DOI:
10.1002/ets2.12192
复制
发表时间:
2018
影响因子:
--
通讯作者:
David M. Williamson
David M. Williamson
中科院分区:
--
文献类型:
--
作者:
Chaitanya Ramineni;David M. Williamson

文献摘要

被引文献

相似文献

在2012年主要修订版之前使用的GRE ®通用考试(称为rGRE)中,观察到了thee‐rater®自动评分引擎和人类对某些人口统计学群体的文章的显著平均得分差异。使用e-rater作为具有差异阈值的检查分数模型,防止了对考生在项目或测试水平上的分数产生不利影响。尽管有这种控制,仍然有必要了解这些人口统计学上的分数差异的根本原因,并确定潜在的机制,以避免未来的差异的情况。在这项研究中,我们使用了统计方法和人工审查的组合,提出了关于评分差异的根本原因的假设,以及这种差异是否反映了电子评分员、人工评分或两者的不足。人类评级过程被发现受到量表结构的强烈影响,并不完全符合电子评级机构的评分机制。人类评分员似乎使用条件逻辑和基于规则的方法进行评分,而e-rater则使用所有特征的线性加权。这些分析对rGRE评分的未来研究和操作政策具有影响。
Notable mean score differences for thee‐rater® automated scoring engine and for humans for essays from certain demographic groups were observed for theGRE® General Test in use before the major revision of 2012, called rGRE. The use of e‐rater as a check‐score model with discrepancy thresholds prevented an adverse impact on the examinee score at the item or test level. Despite this control, there remains a need to understand the root causes of these demographically based score differences and to identify potential mechanisms for avoiding future instances of discrepancy. In this study, we used a combination of statistical methods and human review to propose hypotheses about the root cause of score differences and whether such discrepancies reflect inadequacies of e‐rater, human scoring, or both. The human rating process was found to be influenced strongly by the scale structure and did not fully correspond to the e‐rater scoring mechanism. The human raters appeared to be using conditional logic and a rule‐based approach to their scoring, while e‐rater uses linear weighting of all the features. These analyses have implications for future research and operational policies for the scoring of the rGRE.