Scoring Depression on a Common Metric: A Comparison of EAP Estimation, Plausible Value Imputation, and Full Bayesian IRT Modeling

Scoring Depression on a Common Metric: A Comparison of EAP Estimation, Plausible Value Imputation, and Full Bayesian IRT Modeling
复制标题

DOI:
10.1080/00273171.2018.1491381
复制
发表时间:
2019-01-02
影响因子:
3.8
通讯作者:
Rose, Matthias
Rose, Matthias
中科院分区:
心理学3区
文献类型:
--
作者:
Fischer, H. Felix;Rose, Matthias

文献摘要

被引文献

相似文献

有越来越多的项目反应理论(IRT)研究校准不同的患者报告的结果(PRO)测量,如焦虑,抑郁,身体功能和疼痛,共同的,仪器独立的指标。在抑郁症的情况下,据报道,当从不同的,以前关联的工具的共同指标评分时,有相当大的平均得分差异。理想情况下,这些估计应该是相同的。我们研究了不同的评分方法在多大程度上影响了这些差异,这些评分方法考虑了不同程度的不确定性,例如测量误差(通过合理的值imputation)和项目参数不确定性(通过全贝叶斯IRT建模)。与直接使用预期后验(EAP)估计相比,使用可信值估算或贝叶斯模型时,不同工具的抑郁估计更为相似,其相应的置信度/可信区间更大。此外,我们还探索了使用贝叶斯IRT模型根据新收集的数据更新项目参数。
There are a growing number of item response theory (IRT) studies that calibrate different patient-reported outcome (PRO) measures, such as anxiety, depression, physical function, and pain, on common, instrument-independent metrics. In the case of depression, it has been reported that there are considerable mean score differences when scoring on a common metric from different, previously linked instruments. Ideally, those estimates should be the same. We investigated to what extent those differences are influenced by different scoring methods that take into account several levels of uncertainty, such as measurement error (through plausible value imputation) and item parameter uncertainty (through full Bayesian IRT modeling). Depression estimates from different instruments were more similar, and their corresponding confidence/credible intervals were larger when plausible value imputation or Bayesian modeling was used, compared to the direct use of expected a posteriori (EAP) estimates. Furthermore, we explored the use of Bayesian IRT models to update item parameters based on newly collected data.