Dealing with Sparse Rater Scoring of Constructed Responses within a Framework of a Latent Class Signal Detection Model

Dealing with Sparse Rater Scoring of Constructed Responses within a Framework of a Latent Class Signal Detection Model
复制标题

在潜在类信号检测模型的框架内处理构造响应的稀疏评分者评分

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Sunhee Kim
Sunhee Kim
中科院分区:
--
文献类型:
--
作者:
Sunhee Kim

文献摘要

被引文献

相似文献

在潜在类别信号检测模型的框架内处理稀疏评分者对建构反应的评分Sunhee Kim在许多使用建构反应(CR)项目的评估情况下,考生的反应仅由一名评分者评估,这被称为单评分者设计。例如,在课堂评估实践中,只有一名教师对每个学生的表现进行评分。虽然单一评分者设计是所有评分者设计中最具成本效益的方法,但缺乏第二名评分者会导致如何使用和评估评分的困难。例如,当只有一个评分员时,无法评估评分员可靠性或评分员效应。本研究探讨了可能的解决方案,在稀疏评分员设计的背景下,潜在的类版本的信号检测理论(LC-SDT),以前已被用于评分员评分的问题。这种方法为CR评分中的评分者认知提供了一个模型(DeCarlo,2005; 2008; 2010),并提供了评分者可靠性和各种评分者效应的测量。检查了以下评估者稀疏性的潜在解决方案:1)使用参数限制来产生识别的模型,2)在贝叶斯方法中使用信息先验,以及3)使用回读(例如,部分可用的第二评级者观察结果),在一些大规模评估中可用。仿真和分析的真实世界的数据进行检查这些方法的性能。仿真结果表明,使用参数约束允许一个检测各种评分员的影响,在实践中所关注的。贝叶斯方法也给出了有用的结果,虽然估计的一些参数是穷人和参数后验的标准偏差很大,除了当样本量很大。使用回读分数给出了一个确定的模型和模拟结果表明,一般是可以接受的,在参数估计方面,除了小样本量。本文还探讨了实用的方法适用于PIRLS美国可靠性数据。结果表明,后验模态估计和贝叶斯估计得到的参数估计之间的一些相似性和差异。敏感性分析表明,评分员参数估计是敏感的规范的先验,也发现在较小的样本量的模拟结果。
Dealing with Sparse Rater Scoring of Constructed Responses within a Framework of a Latent Class Signal Detection Model Sunhee Kim In many assessment situations that use a constructed-response (CR) item, an examinee’s response is evaluated by only one rater, which is called a single rater design. For example, in a classroom assessment practice, only one teacher grades each student’s performance. While single rater designs are the most cost-effective method among all rater designs, the lack of a second rater causes difficulties with respect to how the scores should be used and evaluated. For example, one cannot assess rater reliability or rater effects when there is only one rater. The present study explores possible solutions for the issues that arise in sparse rater designs within the context of a latent class version of signal detection theory (LC-SDT) that has been previously used for rater scoring. This approach provides a model for rater cognition in CR scoring (DeCarlo, 2005; 2008; 2010) and offers measures of rater reliability and various rater effects. The following potential solutions to rater sparseness were examined: 1) the use of parameter restrictions to yield an identified model, 2) the use of informative priors in a Bayesian approach, and 3) the use of back readings (e.g., partially available 2nd rater observations), which are available in some large scale assessments. Simulations and analyses of real-world data are conducted to examine the performance of these approaches. Simulation results showed that using parameter constraints allows one to detect various rater effects that are of concern in practice. The Bayesian approach also gave useful results, although estimation of some of the parameters was poor and the standard deviations of the parameter posteriors were large, except when the sample size was large. Using back-reading scores gave an identified model and simulations showed that the results were generally acceptable, in terms of parameter estimation, except for small sample sizes. The paper also examines the utility of the approaches as applicable to the PIRLS USA reliability data. The results show some similarities and differences between parameter estimates obtained with posterior mode estimation and with Bayesian estimation. Sensitivity analyses revealed that rater parameter estimates are sensitive to the specification of the priors, as also found in the simulation results with smaller sample sizes.