Grades and Test Scores: Accounting for Observed Differences

Grades and Test Scores: Accounting for Observed Differences
复制标题

成绩和考试分数:考虑到观察到的差异

DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
Charles Lewis
Charles Lewis
中科院分区:
--
文献类型:
--
作者:
W. W. Willingham;Judith M. Pollack;Charles Lewis

文献摘要

被引文献

相似文献

为什么成绩和考试分数经常不同?本文提出了一个可能存在差异的框架。研究人员用来自国家教育纵向研究(National Education Longitudinal Study)的8454名高中毕业生的数据对该框架进行了近似测试。通过将两种测量方法集中在相似的学术科目上,纠正评分差异和不可靠性,并增加教师评分和其他有关学生的信息,大大减少了个人和群体在成绩与考试成绩方面的差异。高中平均水平同时预测由0.62提高到0.90;8个亚组的差异预测降低到0.02个字母等级。评分差异是造成成绩和考试成绩差异的主要原因。其他主要来源是教师评分和学术参与,这是一个很有前途的组织原则,可以了解学生的成绩。参与度由三种可观察的行为定义:运用学校技能、展示主动性和避免竞争性活动。虽然各组的平均成绩不同,但在成绩和测试上,各组的表现大致相似。成就的主要因素在不同群体之间的构成和关系也相似。等级和测试之间的差异使这些措施在高风险评估中具有互补的优势。如果不纠正两种测量方法之间的人为差异,对有效性和公平性的共同统计估计就会过于保守。
Why do grades and test scores often differ? A framework of possible differences is proposed in this article. An approximation of the framework was tested with data on 8,454 high school seniors from the National Education Longitudinal Study. Individual and group differences in grade versus test performance were substantially reduced by focusing the two measures on similar academic subjects, correcting for grading variations and unreliability, and adding teacher ratings and other information about students. Concurrent prediction of high school average was thus increased from 0.62 to 0.90; differential prediction in eight subgroups was reduced to 0.02 letter-grades. Grading variation was a major source of discrepancy between grades and test scores. Other major sources were teacher ratings and Scholastic Engagement, a promising organizing principle for understanding student achievement. Engagement was defined by three types of observable behavior: employing school skills, demonstrating initiative, and avoiding competing activities. While groups varied in average achievement, group performance was generally similar on grades and tests. Major factors in achievement were similarly constituted and similarly related from group to group. Differences between grades and tests give these measures complementary strengths in high-stakes assessment. If artifactual differences between the two measures are not corrected, common statistical estimates of validity and fairness are unduly conservative.