Missing data and bias in physics education research: A case for using multiple imputation

Missing data and bias in physics education research: A case for using multiple imputation
复制标题

DOI:
10.1103/physrevphyseducres.15.020106
复制
发表时间:
2019-07-03
影响因子:
3.1
通讯作者:
Van Dusen, Ben
Van Dusen, Ben
中科院分区:
教育学3区
文献类型:
--
作者:
Nissen, Jayson;Donatello, Robin;Van Dusen, Ben

文献摘要

被引文献

相似文献

物理教育研究者(PER)通常使用完整的案例分析来解决缺失的数据。对于完整的案例分析,研究人员丢弃任何缺失数据的学生的所有数据。尽管经常使用,但我们审查的使用完整病例分析的PER文章中没有任何证据表明数据符合确保准确结果所需的完全随机缺失假设。不符合这一假设提出了一种可能性,即先前的研究报告了有偏见的结果,夸大了可能掩盖跨课程差异的收益。为了检验这种可能性,我们使用模拟数据比较了完整病例分析和多重插补(MI)的准确性。我们根据先前的研究模拟了数据,这样成绩越高的学生参与率越高,这使得数据随机缺失。PER研究很少使用MI,但MI使用所有可用数据,假设不太严格,比完整病例分析更准确,统计学上更强大。结果表明,completecase分析引入更多的偏见比MI和这种偏见是大到足以掩盖学生群体之间或课程之间的差异。我们建议PER社区采用MI来处理缺失数据,以提高研究的准确性。
Physics education researchers (PER) commonly use complete-case analysis to address missing data. For complete-case analysis, researchers discard all data from any student who is missing any data. Despite its frequent use, no PER article we reviewed that used complete-case analysis provided evidence that the data met the assumption of missing completely at random necessary to ensure accurate results. Not meeting this assumption raises the possibility that prior studies have reported biased results with inflated gains that may obscure differences across courses. To test this possibility, we compared the accuracy of complete-case analysis and multiple imputation (MI) using simulated data. We simulated the data based on prior studies such that students who earned higher grades participated at higher rates, which made the data missing at random. PER studies seldom use MI, but MI uses all available data, has less stringent assumptions, and is more accurate and more statistically powerful than complete-case analysis. Results indicated that completecase analysis introduced more bias than MI and this bias was large enough to obscure differences between student populations or between courses. We recommend that the PER community adopt the use of MI for handling missing data to improve the accuracy in research studies.