Rating Scales Derived from Student Samples: Effects of the Scale Maker and the Student Sample on Scale Content and Student Scores.

Rating Scales Derived from Student Samples: Effects of the Scale Maker and the Student Sample on Scale Content and Student Scores.
复制标题

从学生样本得出的评分量表:量表制作者和学生样本对量表内容和学生分数的影响。

DOI:
10.2307/3588360
复制
发表时间:
2002
期刊:
影响因子:
3.2
通讯作者:
J. Upshur
J. Upshur
中科院分区:
人文科学2区
文献类型:
--
作者:
Carolyn E. Turner;J. Upshur

文献摘要

被引文献

相似文献

表现测试通常要求评分员根据评分量表判断考生的书面或口头语言质量;因此,分数可能会受到特定量表开发过程中固有变量的影响。在本研究中,我们考虑了迄今为止尚未研究的经验推导的评级量表中的两个变量:量表开发者和量表开发者使用的性能样本。这些变量可能会影响量表的内容和结构以及(最终)最终的考试成绩。本研究使用两个ESL学生写作样本和三个评分量表开发团队来构建三个经验推导的量表,以检验量表的开发和使用。量表内容的比较显示出相当大的差异,即使所有的开发团队都使用类似的写作能力结构。每个小组都用自己的标准来评价一组不同的作文。等级的比较表明,规模开发团队对等级的影响较小,而规模开发样本对等级的影响较大。我们提出这些研究结果对经验推导的评定量表的性质的影响,特别关注这些量表是如何开发的。
Performance tests typically require raters to judge the quality of examinees' written or spoken language relative to a rating scale; therefore, scores may be affected by variables inherent in the specific scale development process. In this study we consider two variables in empirically derived rating scales that have not been investigated to date: scale developers and the sample of performances used by the scale developers. These variables may affect scale content and structure and (ultimately) final test scores. This study examined the development and use of scales using two samples of ESL student writing and three teams of rating scale developers to construct three empirically derived scales. A comparison of the scale content showed considerable variation even though all development teams used similar constructs of writing ability. Each team used its own scale to rate a different set of compositions. Comparison of the ratings showed that scale development team had a minor effect on ratings and that scale development sample had a major effect. We present implications of these findings on the nature of empirically derived rating scales, focusing particularly on how such scales are developed.