Improving test security and efficiency of computerized adaptive testing for the Force Concept Inventory

Improving test security and efficiency of computerized adaptive testing for the Force Concept Inventory
复制标题

提高部队概念清单计算机化自适应测试的测试安全性和效率

DOI:
10.1103/physrevphyseducres.18.010112
复制
发表时间:
2022
影响因子:
3.1
通讯作者:
Mae Naohiro
Mae Naohiro
中科院分区:
教育学3区
文献类型:
--
作者:
Yasuda Jun-ichiro;Hull Michael M.;Mae Naohiro

文献摘要

相似文献

本文介绍了一种基于计算机自适应测试(CAT)的FCI版本(FCI-CAT)在测试安全性和测试效率方面的改进。首先,我们将讨论通过控制试题过度暴露来提高考试安全性的措施,减少受访者可能(I)记住预测的内容以用于后测或(Ii)与稍后参加评估的同学分享有关试题的信息的风险。其次,我们将讨论提高测试效率的措施,以便更短的测试长度可以产生所需的测量精度和精密度。具体地说,我们利用每个被调查者在测试前的熟练程度评估的形式的辅助信息来选择项目并在后测试中估计被调查者的水平。为了进一步缩短总测试时间,我们还允许前后测试的长度不同。为了分析这些改进如何影响科恩的准确度和精确度(我们用均方根误差来衡量),我们进行了蒙特卡罗模拟和事后模拟。然后,我们计算了FCI-CAT的最小测试长度,其准确度和精确度与纸笔版本的FCI-CAT相当。因此,我们得到了以下三个发现:(1)通过使用附带信息,我们可以通过FCI-CAT以更少的项目达到全长FCI的准确性和精确度。(Ii)对于40人的班级,我们可以在控制测试安全性的同时,将FCI-CAT的测试前后长度之和减少到总共33项(前测17项,后测16项),从而将测试时间减少到55%。(Iii)如果一个人的目标是最大限度地提高测试效率,则前测长度应略大于后测长度。另一方面,如果目标是最大限度地提高测试安全性,则前测长度应该较小,而后测长度应该较大。如果一个人想要在这两个目标之间取得平衡,那么选择相同的测试前和测试后长度将是合理的。
This paper presents improvements made to a computerized adaptive testing (CAT)-based version of the FCI (FCI-CAT) in regards to test security and test efficiency. First, we will discuss measures to enhance test security by controlling for item overexposure, decreasing the risk that respondents may (i) memorize the content of a pretest for use on the post-test or (ii) share information about the items with their classmates who take the assessment later. Second, we will discuss measures to enhance test efficiency, so that a shorter test length can yield a desired accuracy and precision of the measurement. Specifically, we utilized collateral information in the form of a pretest proficiency estimate of each respondent for selecting items and estimating respondent proficiency level in the post-test. To shorten the total testing time further, we also allowed the test lengths to be different for the pre- and post-test. To analyze how these improvements affect the accuracy and precision (which we measure in terms of root-mean-square error) of Cohen’s, we conducted a Monte Carlo simulation and apost hocsimulation. Then, we calculated the minimal test length of the FCI-CAT whose accuracy and precision are equivalent to that of the paper-and-pencil version of the FCI. Consequently, we obtained the following three findings: (i) By using collateral information, we can achieve the accuracy and precision of the full-length FCI with fewer items via the FCI-CAT. (ii) For a class size of 40, we can control for test security while still reducing the sum of the pre- and post-test lengths of the FCI-CAT to a total of 33 items (17 items on the pretest and 16 items on the post-test), thereby reducing the testing time to 55%. (iii) If one’s goal is to maximize test efficiency, the pretest length should be slightly larger than the post-test length. On the other hand, if the goal is to maximize test security, the pretest length should be smaller and the post-test length should be larger. If one desires a balance of these two goals, it would be reasonable to choose equal pre- and post-test lengths.