Limited-information goodness-of-fit testing of hierarchical item factor models.

Limited-information goodness-of-fit testing of hierarchical item factor models.
复制标题

有限的信息拟合优度测试分层项目因子模型。

DOI:
10.1111/j.2044-8317.2012.02050.x
复制
发表时间:
2013-05
期刊:
The British journal of mathematical and statistical psychology
影响因子:
--
通讯作者:
Hansen M
Hansen M
中科院分区:
其他
文献类型:
--
作者:
Cai L;Hansen M

文献摘要

参考文献

被引文献

相似文献

在项目反应理论的应用中,模型拟合的评价是一个关键问题。最近,有限信息拟合优度测试在心理测量学文献中受到越来越多的关注。与全信息检验统计量(如Pearson的X2或似然比G2)相比,这些有限信息检验使用低阶边际表,而不是全列联表。一个值得注意的例子是Maydeu-Olivares及其同事基于单变量和双变量边际的M2统计家族。当列联表稀疏时,基于M2的测试比全信息测试保留了更好的I类错误率控制,并且可以更强大。虽然原则上M2统计量可以扩展到测试分层多维项目因子模型(例如,bifactor和testlet模型),计算是不平凡的。为了获得M2,研究人员通常必须获得(数千个)边际概率、导数和权重。每一个都必须用高维数值积分来近似。我们提出了一种降维方法,可以利用层次因子结构,使积分可以更有效地近似。我们还提出了一个新的检验统计量,可以大大更好地校准和更强大的比原来的M2统计量时,测试是长的,项目是多分支的。我们使用模拟来证明我们的新方法的性能,并说明其有效性与应用程序的真实的数据。
In applications of item response theory, assessment of model fit is a critical issue. Recently, limited-information goodness-of-fit testing has received increased attention in the psychometrics literature. In contrast to full-information test statistics such as Pearson’s X2 or the likelihood ratio G2, these limited-information tests utilise lower order marginal tables rather than the full contingency table. A notable example is Maydeu-Olivares and colleagues’ M2 family of statistics based on univariate and bivariate margins. When the contingency table is sparse, tests based on M2 retain better Type I error rate control than the full-information tests and can be more powerful. While in principle the M2 statistic can be extended to test hierarchical multidimensional item factor models (e.g., bifactor and testlet models), the computation is non-trivial. To obtain M2, a researcher often has to obtain (many thousands of) marginal probabilities, derivatives, and weights. Each of these must be approximated with high-dimensional numerical integration. We propose a dimension reduction method that can take advantage of the hierarchical factor structure so that the integrals can be approximated far more efficiently. We also propose a new test statistic that can be substantially better calibrated and more powerful than the original M2 statistic when the test is long and the items are polytomous. We use simulations to demonstrate the performance of our new methods and illustrate their effectiveness with applications to real data.
DOI: 10.1007/bf02293801
发表时间: 1981-01-01
期刊: PSYCHOMETRIKA
影响因子: 3
作者:
BOCK, RD;AITKIN, M
通讯作者: AITKIN, M
DOI: 10.1007/s11336-010-9178-0
发表时间: 2010-12-01
期刊: PSYCHOMETRIKA
影响因子: 3
作者:
Cai, Li
通讯作者: Cai, Li
DOI: 10.1207/s15327906mbr3801_2
发表时间: 2003-01-01
影响因子: 3.8
作者:
Briggs, NE;MacCallum, RC
通讯作者: MacCallum, RC
DOI: 10.1348/000711002159617
发表时间: 2002-05-01
影响因子: 2.6
作者:
Bartholomew, DJ;Leung, SO
通讯作者: Leung, SO
DOI: 10.1214/aos/1176344001
发表时间: 1977-01-01
影响因子: 4.5
作者:
HABERMAN, SJ
通讯作者: HABERMAN, SJ