Population invariance and the equatability of tests: Basic theory and the linear case

Population invariance and the equatability of tests: Basic theory and the linear case
复制标题

DOI:
10.1111/j.1745-3984.2000.tb01088.x
复制
发表时间:
2000-12-01
影响因子:
1.3
通讯作者:
Holland, PW
Holland, PW
中科院分区:
心理学4区
文献类型:
--
作者:
Dorans, NJ;Holland, PW

文献摘要

被引文献

相似文献

两种检验不应等同的事实如何表现出来?本文通过研究等化函数在多大程度上不能表现出跨子种群的种群不变性来解决这个问题。根据定义,等值函数应该是总体不变的。但是,当两个测试不相等时,用于将一个测试的分数连接到另一个测试的分数的链接函数在不同的考生群体中可能不是不变的。虽然没有一个可接受的等值函数是完全种群不变的,但在通常执行等值的情况下,我们相信等值函数对用于计算它的种群的依赖性通常小到可以忽略。我们引入了两个均方根差异的程度,用于连接两个测试计算不同的子群体的功能不同的连接功能计算的整个人口的措施。我们还介绍了系统的“平行线性”连接功能的多个亚群,并表明,对于这个系统,我们的人口不变性的措施可以很容易地计算从标准化的平均差异的两个测试的亚群的分数。对于平行线性的情况下,我们开发了一个基于相关性的上界,我们的措施,适用于所有系统的子群。我们说明这些想法使用的数据从SAT I和一致性研究的ACT和SAT I分数的几种组合。在附录中,我们给出了与“相同结构”、“相同可靠性”等其他等同要求有关的一些理论结果,以及洛德公平概念的一个方面。
How does the fact that two tests should not be equated manifest itself? This paper addresses this question through the study of the degree to which equating functions fail to exhibit population invariance across subpopulations. Equating functions are supposed to be population invariant by definition. But, when two tests are not equatable, it is possible that the linking functions, used to connect the scores of one to the scores of the other are not invariant across different populations of examinees. While no acceptable equating function is ever completely population invariant, in the situations where equating is usually performed we believe that the dependence of the equating function on the population used to compute it is usually small enough to be ignored. We introduce two root-mean-square difference measures of the degree to which the functions used to link two tests computed on different subpopulations differ from the linking function computed for the whole population. We also introduce the system of "parallel-linear" linking functions for multiple subpopulations and show that, for this system, our measure of population invariance can be computed easily from the standardized mean differences between the scores of the subpopulations on the two tests. For the parallel-linear case, we develop a correlation-based upper bound on our measure that holds for all systems of subpopulations. We illustrate these ideas using data from the SAT I and from a concordance study of several combinations of ACT and SAT I scores. In the appendices, we give some theoretical results bearing on the other equating requirements" of "same construct, " "same reliability" and one aspect of Lord's concept of equity.