Investigating Item Exposure Control Methods in Computerized Adaptive Testing.

Investigating Item Exposure Control Methods in Computerized Adaptive Testing.
复制标题

研究计算机化自适应测试中的项目暴露控制方法。

DOI:
10.12738/estp.2015.1.2593
复制
发表时间:
2015
期刊:
Kuram Ve Uygulamada Egitim Bilimleri
影响因子:
--
通讯作者:
Nuri Doğan
Nuri Doğan
中科院分区:
--
文献类型:
--
作者:
Nagihan Boztunc Ozturk;Nuri Doğan

文献摘要

被引文献

相似文献

摘要本研究旨在探讨在不同的项目选择方法和项目库特征下,项目暴露控制方法对测量精度和测试安全性的影响。在这项研究中,随机(项目组大小为5和10),症状Hetter和淡出的方法被用作项目暴露控制方法。此外,为了建立其他方法的比较基线,本研究还包括了无暴露控制条件。作为项目选择方法,采用最大Fisher信息、a-分层和渐进最大信息比。对于两个项目库,a参数是从0.50和2.00之间的均匀分布生成的,c参数是从0.05和0.20之间的均匀分布生成的,对于中等难度项目库,B参数是从-3.00和+3.00之间的均匀分布生成的,而对于标准正态N(2,1.5)高难度试题库的分配。根据研究结果,当使用项目暴露控制方法时,用于确定测量精度的指标值之间没有很大差异。另一方面,在测试安全性方面,研究发现,总体而言,控制项目暴露的淡出方法在减少项目池利用的偏斜度和测试重叠方面比其他方法取得了更好的结果。关键词:计算机自适应测试 * 项目暴露控制方法 * 随机方法 * Sympson-Hetter方法 * Fade-Away方法 * 测量精度 * 测试安全性(ProQuest:...表示省略的公式。)随着计算机技术和心理测量学领域的同步发展,计算机自适应测试(CAT)的管理已经增加,并将继续增加。计算机辅助测试取代了传统的纸笔测试,因为它易于应用和评分,测试中有与考生能力水平相对应的项目,这种测试可以随时进行,并且与传统的纸笔测试相比,测试时间较短,因此越来越具有吸引力(Grist,Ruffins & Wise,1989; Meijer & Nering,1999; Ruffins,1998;韦斯& Kingsbury,1984)。CAT应用程序用于管理GMAT(研究生管理入学考试)和GRE(研究生入学考试)。与传统的纸笔测试一样,这些应用程序也需要可靠和有效,因为它们用于对考生的未来做出关键决定。随着CAT的使用日益广泛,出现了一些可能危及其有效性的问题。因此,此类测试的安全性问题变得越来越重要(Chang & Twu,1998; Davey & Nering,2002 as cited in Barrada,Olea,Ponsoda,& Abad,2009; French & Thompson,2003; Georgiadou,Triantafillou,& Economide,2007;虽然CAT应用程序需要包含大量项目的项目池,但在特定情况下,某些项目比其他项目使用得更频繁。在这种情况下,考生试图简单地记住常用项目的答案的概率更高。如果这些项目被记住,然后共享,测试的有效性变得危险(Georgiadou等人,二○ ○七年;由于开发一个理想的项目池是一个漫长而费力的过程的结果,因此测试开发人员不希望只使用项目池的某个百分比,而是更好地有效地使用整个池(Revuelta & Ponsoda,1998)。由于上述原因,开发了一些方法,以确保测试的安全性,并更有效地利用项目库。这种方法被称为“项目暴露控制方法”,由于在实际应用中遇到的问题,这些方法已逐渐被纳入CAT的基本组成部分(Boyd,2003; Davis,2002)。…
AbstractThis study aims to investigate the effects of item exposure control methods on measurement precision and on test security under various item selection methods and item pool characteristics. In this study, the Randomesque (with item group sizes of 5 and 10), Sympson-Hetter, and Fade-Away methods were used as item exposure control methods. Moreover, in order to establish a comparison baseline for other methods, the noexposure control condition was also included in the research. As item selection methods, Maximum Fisher Information, a-Stratification, and Gradual Maximum Information Ratio were employed. While a parameters were generated from a uniform distribution ranging between 0.50 and 2.00 and c parameters were generated from a uniform distribution ranging between 0.05 and 0.20 for both item pools, b parameters were generated from a uniform distribution ranging between -3.00 and +3.00 for medium difficulty item pool, and from a standard normal N(2, 1.5) distribution for high difficulty item pool. Based on the research findings, there were no great differences between the values of indicators used in determining measurement precision when the item exposure control methods were used. On the other hand, in terms of test security, it was found that in general, the Fade-Away Method for controlling item exposure yielded better results than did the other methods in reducing the skewness of item pool utilization and the test overlap.Keywords: Computerized adaptive testing * Item exposure control methods * Randomesque method * Sympson-Hetter method * Fade-Away method * Measurement precision * Test security(ProQuest: ... denotes formulae omitted.)The administering of Computerized Adaptive Testing (CAT) has increased, and continue to do so, in line with concurrent developments in both the fields of computer technology and psychometrics. In lieu of traditional paper-and-pencil tests, CAT style tests have become increasingly attractive because they are easy to apply and to score, there are items in the tests corresponding to examinees' ability levels, such tests can be administered whenever desired, and the tests are shorter when compared to traditional paper-and-pencil tests (Grist, Rudner & Wise, 1989; Meijer & Nering, 1999; Rudner, 1998; Weiss & Kingsbury, 1984). CAT applications are used in the administration of the GMAT (Graduate Management Admission Test) and the GRE (Graduate Record Examination). As in traditional paper-and-pencil tests, these applications, also need to be reliable and valid since they are used to make critical decisions regarding the futures of examinees. As the use of CAT have become increasingly widespread, so have a number of problems with the potential of endangering validity appeared. Accordingly, the issue of such tests' security increases in importance (Chang & Twu, 1998; Davey & Nering, 2002 as cited in Barrada, Olea, Ponsoda, & Abad, 2009; French & Thompson, 2003; Georgiadou, Triantafillou, & Economide, 2007; Lee & Dodd, 2012).Although CAT applications require item pools containing a large number of items, certain items are used more frequently than others in specific situations. In such situations, the probability of examinees' attempting simply to memorize the answers to frequently used items is higher. If such items are memorized and then shared, the test's validity becomes jeopardized (Georgiadou et al., 2007; Lee & Dodd, 2012).Since developing an ideal item pool is the result of a long and laborious process, it is undesirable for test developers to use just a certain percentage of the item pool, instead it would be better to use the entire pool efficiently (Revuelta & Ponsoda, 1998). For the stated reasons, a number of methods were developed so as to assure test security and to make use of the item pool more efficiently. Such methods are called "item exposure control methods," which have gradually been included in the fundamental components of CAT due to the problems encountered in real-life applications (Boyd, 2003; Davis, 2002). …