Alternatives to the Grade Point Average as a Measure of Academic Achievement in College. ACT Research Report Series.

Alternatives to the Grade Point Average as a Measure of Academic Achievement in College. ACT Research Report Series.
复制标题

平均绩点的替代品作为大学学业成绩的衡量标准。

DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
E. M. Schultz
E. M. Schultz
中科院分区:
--
文献类型:
--
作者:
Pui‐wa Lei;Dina Bassiri;E. M. Schultz

文献摘要

参考文献

被引文献

相似文献

众所周知,大学 GPA 是不同课程分配成绩的线性组合,它不能完美衡量学生的成绩。这种不可靠的衡量标准降低了大学入学考试的预测有效性。研究表明,根据不同的评分实践调整课程成绩可以提高预测的有效性。调整后的大学 GPA 学生的相对排名也与其课程成绩排名更加一致。这些发现通过使用 4 个多重 IRT 和 3 个线性模型的两所大学的连续队列的课程成绩数据得到了重复。与之前的研究不同,课程参数估计和回归权重是交叉验证的。相同样本和交叉验证的替代衡量指标都比简单的 GPA 有所提高。评级量表和部分学分 IRT 模型在与入学考试成绩的多重相关性方面表现出色。分级响应 IRT 模型在队列中是最不稳定的。讨论了这些发现的意义和研究的局限性。大学平均绩点作为衡量学术成就的替代方法 GPA(不同课程成绩的线性组合)并不是衡量学生成绩的理想衡量标准,因为它不仅反映了学术成就,还反映了课程采取策略和教师评分实践。除非不允许选课并且所有老师都愿意遵守通用的评分标准,否则不同学生的GPA不一定具有相同的含义。使用 GPA 作为大学学业成绩的衡量标准存在一些问题。一是出于奖学金或就业目的,很难从不同部门或不同机构选择候选人(例如Caulkins等,1996)。不同的教师/教授根据自己对学生成绩的看法有不同的评分标准(Hoover、Roller、Liddell、Moore、McCarthy 和 Hlebowitsh,1999)。也许由于“志趣相投”的个人组成,各部门的评分倾向有所不同(Hoover,et al.,1999;Johnson,1997)。因此,学生之间的平均绩点并不具有严格的可比性,特别是跨院系或专业的学生。使用 GPA 作为学业成绩的衡量标准也会导致成绩膨胀。有人建议教师降低标准,以提高学生对课程的评分。学生可以选择由评分宽松的教师教授的课程(例如,Johnson,1997)或转向倾向于给出高分的院系(例如,Young,1993)。结果,成绩可能会提高,但并没有反映出学生能力的提高,这种现象被称为成绩膨胀(Bejar & Blew,1981)。请注意,分数膨胀不一定特定于大学水平。 Ziomek 和 Svec (1995) 记录了高中阶段的类似现象。依靠 GPA 作为学业成绩的衡量标准也使得评估大学入学考试变得更加困难(例如,Young,1993 年;Strieker、Rock、Burton、Muraki 和 Jirele,1994 年)。入学考试可以预测成绩,但不应预测学生是否会选择更简单的课程。为了使课程成绩更具可比性,对所有教师实行共同评分标准(或拒绝学生选择课程)的一个可行替代方案是根据不同的课程难度调整 GPA(Caulkins、Larkey 和 Wei,1996)。调整后的 GPA 与 GPA 一样,代表学生在单一(一维)尺度上的成绩,并且完全由课程成绩数据构建。调整后的 GPA 并不能解决用简单的一维测量来表示相对于特定学科领域的固有多维领域的根本问题。其他维度还可能包括非认知特征,例如上课和交作业。然而,它比 GPA 更好,因为它减少了因不同的选课模式和课程难度变化而产生的误差。希望通过调整课程难度,从长远来看,将抑制学习课程内容以外的激励措施。调整 GPA 的直接影响包括提高大学入学考试的预测有效性(Young,1990a;Caulkins 等,1996;Johnson,1997)和降低性别的差异预测有效性(Young,1991)。大多数 GPA 调整方法的运作前提是,成绩指数应反映参加相同课程的学生的相对课程排名(例如,Caulkins 等人,1996 年;Johnson,1997 年),这是一种单一维度的成绩概念。这些方法通常会根据课程的不同评分严格程度或难度级别进行调整。下一节总结了本研究特别感兴趣的方法。有关调整方法的最新综合评论,请参见 Young (1993) 和 Strieker、Rock、Burton、Muraki 和 Jirele (1994),或更远的评论请参见 Linn (1966)。成绩调整方法 Young (1990a) 创造性地提出了项目反应理论 (IRT),其中课程成绩被视为项目分数,估计的 θ 成为预期的调整后的成绩指数。 Young(1990a)首先对课程成绩进行因子分析,并根据因子分析结果将课程分成相对一维的组,然后对分离的课程组应用鲽岛(1969)分级响应模型的限制版本。他发现 SAT-V、SAT-M 和高中 GPA 基于 IRT 的 GPA 的可预测性超过了未经调整的 GPA。平方多重相关性 (T?2) 的增加范围从 0.0015 到 0.0955,从可忽略到相当大,可能取决于许多因素,包括成绩分布、涉及调整的课程数量、每个学生选修的课程数量,甚至定义特征的清晰度(Young,1990a)。此外,Young (1990a) 主张其他多级 IRT 模型可能适用于等级调整。然而,只有 Samejima (1969) 在多断层 IRT 家族中的分级反应模型得到了检验。因此,本研究拟扩大基于IRT的类别应用于成绩调整。根据模型约束的数量或者所涉及的自由模型参数(要从数据估计的模型参数)的数量来选择具有不同复杂程度的模型。评级量表模型(Andrich,1978)的自由参数最少(模型约束最多),其次是部分信用模型(Masters,1982),然后是广义部分信用模型(Muraki,1992)和分级响应模型(Samejima,1969)。后两个模型具有相同数量的自由参数(有关用于调整等级的模型列表,请参阅附录 A)。 3
College GPA, a linear combination of assigned grades from different courses, is widely known to be an imperfect measure of student achievement. This unreliable measure decreases the predictive validity of college admission tests. Research has shown that adjusting course grades for differential grading practices improves predictive validity. Relative rankings of students on adjusted college GPAs are also more consistent with their course grade standings. These findings were replicated with course grade data from consecutive cohorts of two universities using 4 polytomous IRT and 3 linear models. Unlike previous studies, course parameter estimates and regression weights were cross-validated. Both same-sample and cross-validated alternative measures showed improvement over simple GPA. The rating scale and partial credit IRT models excelled on multiple correlations with admission test scores. The graded response IRT model was the most unstable across cohorts. Implications of these findings and limitations of the studies are discussed. Alternatives to the Grade Point Average as a Measure of Academic Achievement in College GPA, a linear combination of grades assigned in different courses, is not an ideal measure of student achievement because it reflects not only academic achievement, but also course taking strategies and instructor grading practices. Unless course selection is not allowed and all instructors are willing to adhere to a universal grading standard, GPAs for different students do not necessarily have the same meaning. Several problems are associated with the use of GPA as a measure of academic achievement in college. One is that it is difficult to select candidates from different departments or different institutions for scholarship or employment purposes (e.g., Caulkins, et al., 1996). Different teachers/professors have different grading criteria according to their own perception of student achievement (Hoover, Roller, Liddell, Moore, McCarthy, and Hlebowitsh, 1999). Perhaps due to the composition of “like-minded” individuals, departments vary in their grading tendencies (Hoover, et al., 1999; Johnson, 1997). Grade point averages are, therefore, not strictly comparable among students, particularly across departments or majors. The use of GPA as a measure of academic achievement also drives grade inflation. It has been suggested that instructors lower their standards in order to improve their course ratings by students. Students may shop for courses taught by leniently grading instructors (e.g., Johnson, 1997) or switch to departments that tend to give high grades (e.g., Young, 1993). As a result, grades may be raised without reflecting increased students’ abilities, a phenomenon known as grade inflation (Bejar & Blew, 1981). Note that grade inflation is not necessary specific to the college level. Ziomek and Svec (1995) documented a similar phenomenon at the high school level. Relying on GPA as a measure of academic achievement also makes it more difficult to evaluate college admissions tests (e.g., Young, 1993; Strieker, Rock, Burton, Muraki, & Jirele, 1994). Admissions tests are rightly expected to predict grades but should not be expected to predict whether a student will choose an easier curriculum. To make course grades more comparable, a viable alternative to imposing a common grading standard on all instructors (or to denying course selection by students) is to adjust GPA for differential course difficulty (Caulkins, Larkey, and Wei, 1996). Adjusted-GPA, like GPA, represents student achievement on a single (unidimensional) scale and is constructed entirely from course grade data. An adjusted-GPA does not resolve the underlying problem of representing an inherent multidimensional domain with respect to the specific subject areas with a simple unidimensional measure. Other dimensions may also include non-cognitive characteristics such as attending class and turning in homework. However, it is better than GPA because it reduces the error arising from differential course-taking patterns and variation in course difficulty. It is hoped that by leveling course difficulty, incentives other than learning the course content will be discouraged in the long run. Immediate effects of adjusting GPA include improved predictive validity of college admission tests (Young, 1990a; Caulkins, et al., 1996; Johnson, 1997) and reduced differential predictive validity for gender (Young, 1991). Most GPA adjustment methods operate on the premise that an achievement index should reflect relative course ranks of students who took the same classes (e.g., Caulkins, et al, 1996; Johnson, 1997), a uni dimensional notion of achievement. These methods often adjust for the different grading stringency or difficulty level of the classes. Methods of particular interest to this study are summarized in the following section. For recent comprehensive reviews of adjustment methods, see Young (1993) and Strieker, Rock, Burton, Muraki, and Jirele (1994), or a more distant one, see Linn (1966). Grade-Adjustment Methods Young (1990a) proposed an inventive use of Item Response Theory (IRT), in which course grades are treated as item scores and estimated theta becomes the intended adjusted achievement index. Young (1990a) first factor analyzed the course grades and separated the courses into relatively unidimensional groups based on the factor analysis results, then applied a restricted version of Samejima’s (1969) Graded Response Model on the separated groups of courses. He found that the predictability of the IRT-based GPA from SAT-V, SAT-M, and high school GPA exceeded that of the unadjusted GPA. The increase in squared multiple correlation (T?2) ranged from .0015 to .0955, from negligible to sizable, depending on possibly a number of factors which included the distribution of grades, number of courses involved in the adjustment, number of courses taken by each student, or even the clarity of the defining trait (Young, 1990a). Moreover, Young (1990a) advocated that other polytomous IRT models may be applicable to grade adjustment. However, only Samejima’s (1969) graded response model in the polytomous IRT family has been examined. Therefore, this study intends to expand the IRTbased category applied to grade adjustment. Models with different levels of complexity in terms of number of model constraints or, alternatively, number of free model parameters (model parameters to be estimated from the data) involved are selected. The rating scale model (Andrich, 1978) has the fewest free parameters (the most model constraints), followed by the partial credit model (Masters, 1982), and then the generalized partial credit model (Muraki, 1992) and the graded response model (Samejima, 1969). The latter two models have the same number of free parameters (see Appendix A for a list of the models used for adjusting grades). 3
评分宽大是学生评分中可去除的污染物。
DOI: 10.1037//0003-066x.52.11.1209
发表时间: 1997
期刊: The American psychologist
影响因子: --
作者:
Greenwald,AG;Gillmore,GM
通讯作者: Gillmore,GM