A Comprehensive Review and Comparison of CUSUM and Change-Point-Analysis Methods to Detect Test Speededness

A Comprehensive Review and Comparison of CUSUM and Change-Point-Analysis Methods to Detect Test Speededness
复制标题

检测测试速度的 CUSUM 和变点分析方法的全面回顾和比较

DOI:
10.1080/00273171.2020.1809981
复制
发表时间:
2021
影响因子:
3.8
通讯作者:
Cheng, Ying
Cheng, Ying
中科院分区:
心理学3区
文献类型:
--
作者:
Yu, Xiaofeng;Cheng, Ying

文献摘要

参考文献

被引文献

相似文献

累积和法(Cumulative sum,简称CSUM)和变点分析法(Change Point Analysis,简称CPA)是两种行之有效的统计过程控制方法,用于检测序列中的变化。两者都已用于心理测量学研究,以检测反应序列中的异常反应,例如,测试速度、注意力不集中或作弊。然而,在不同的测试环境中,CANUM和CPA的利弊仍然不清楚。在本文中,我们进行了一个全面的比较性能的12个基于统计量和三个CPA为基础的程序在检测测试速度。两种加速机制,即渐变模型(GCM)和混合模型(HM),被认为是测试的鲁棒性和灵活性的两种方法。仿真研究表明,统计量的性能受到底层数据生成模型、加速严重程度和测试长度的影响。一般来说,在HM下,一些基于CPA-UM的统计数据比基于CPA的统计数据表现得更好。在GCM下,CPA统计的性能得到了显着改善。综上所述,由于真实的应用中未知的加速机制,当测试长度较长时,推荐使用两种基于UML的统计数据(例如,80项),无论潜在机制是HM还是GCM。在相对较短的时间内(例如,40个项目)或中等长度(例如,60个项目)测试,没有统计数据总是在HM和GCM下排名前三。在这些情况下,上述两种基于统计数据的统计数据中的任何一种都是合理的选择,因为它们在各种条件下都有良好的表现(尽管不一定是最好的)。
Cumulative sum (CUSUM) and change-point analysis (CPA) are two well-established statistical process control methods to detect changes in a sequence. Both have been used in psychometric research to detect aberrant responses in a response sequence, e.g., test speededness, inattentiveness, or cheating. However, the pros and cons of CUSUM and CPA in different testing settings still remain unclear. In this paper, we conduct a comprehensive comparison of the performance of twelve CUSUM-based statistics and three CPA-based procedures in detecting test speededness. Two speededness mechanisms are considered, namely the graduate change model (GCM) and the hybrid model (HM), to test the robustness and flexibility of the two methods. Simulation studies show that the performances of the statistics are affected by the underlying data generating model, the severity of speededness, and the test length. Generally, under HM some CUSUM statistics perform much better than the CPA-based statistics. Under the GCM, the performance of the CPA statistics is dramatically improved. Taken together, due to the unknown mechanism of speededness in real applications, two CUSUM-based statistics are recommended when the test length is long (e.g., 80 items), regardless of the underlying mechanism being HM or GCM. In a relatively short (e.g., 40 items) or medium-length (e.g., 60 items) test, no statistic always ends up in the top three under both HM and GCM. In those cases, either one of the two CUSUM-based statistics mentioned above can be a reasonable choice because of their good (though not necessarily the best) performance in a wide range of conditions.
贝叶斯稳健 IRT 异常值检测模型
DOI: 10.1177/0146621616679394
发表时间: 2017
影响因子: 1.2
作者:
Nicole K Öztürk;G. Karabatsos
通讯作者: G. Karabatsos
通过质量控制图、基于模型的方法和时间序列技术随时间监控量表得分
DOI: 10.1007/s11336-013-9317-5
发表时间: 2013
期刊: Psychometrika
影响因子: 3
作者:
Yi;A. A. Davier
通讯作者: A. A. Davier
DOI: 10.1111/j.1745-3984.2012.00176.x
发表时间: 2012
影响因子: 1.3
作者:
Youngsuk Suh;Sun;James A. Wollack
通讯作者: James A. Wollack
IRT 测试等同化:相关问题和近期研究回顾
DOI: 10.3102/00346543056004495
发表时间: 1986
影响因子: 11.2
作者:
G. Skaggs;R. Lissitz
通讯作者: R. Lissitz
使用似然比测试和分数测试检测项目预知识
DOI: 10.3102/1076998616673872
发表时间: 2017
影响因子: 2.4
作者:
S. Sinharay
通讯作者: S. Sinharay