Comparison of different scoring methods based on latent variable models of the PHQ-9: an individual participant data meta-analysis.

Comparison of different scoring methods based on latent variable models of the PHQ-9: an individual participant data meta-analysis.
复制标题

DOI:
10.1017/s0033291721000131
复制
发表时间:
2021-02-22
影响因子:
6.9
通讯作者:
Thombs, Brett D.
Thombs, Brett D.
中科院分区:
医学1区
文献类型:
--
作者:
Fischer, Felix;Levis, Brooke;Falk, Carl;Sun, Ying;Ioannidis, John P. A.;Cuijpers, Pim;Shrier, Ian;Benedetti, Andrea;Thombs, Brett D.

文献摘要

被引文献

相似文献

以往对患者健康问卷(PHQ-9)抑郁量表的研究发现,不同的潜在因素模型最大限度地提高了拟合优度的经验措施。这些差异的临床相关性尚不清楚。我们的目的是调查是否抑郁症筛查的准确性可能会提高采用潜在因素模型为基础的评分,而不是总和分数。我们使用个体参与者数据荟萃分析(IPDMA)数据库来评估PHQ-9的筛查准确性。我们纳入了使用DSM结构化临床访谈(SCID)作为参考标准的研究,并将其分为校准和验证数据集。在校准数据集中,我们估计了一维,二维(分离抑郁症的认知/情感和躯体症状)和双因素模型,以及各自的截止值,以最大限度地提高组合灵敏度和特异性。在验证数据集中,我们评估了潜在变量方法与最佳总和得分(10)之间(组合)灵敏度和特异性的差异,使用自举法估计差异的95%置信区间。校准数据集包括24项研究(4378名参与者,652例重度抑郁症病例);验证数据集包括17项研究(4252名参与者,568例病例)。在验证数据集中,一维、二维和双因素模型的最佳临界值与总和评分临界值10相比具有更高的灵敏度(分别为0.036、0.050、0.049分),但特异性较低(分别为0.017、0.026、0.019)。在诊断研究的综合数据集中,与简单的总和评分方法相比,使用复杂潜变量模型进行评分并不能显著提高PHQ-9的筛查准确性。
Previous research on the depression scale of the Patient Health Questionnaire (PHQ-9) has found that different latent factor models have maximized empirical measures of goodness-of-fit. The clinical relevance of these differences is unclear. We aimed to investigate whether depression screening accuracy may be improved by employing latent factor model-based scoring rather than sum scores. We used an individual participant data meta-analysis (IPDMA) database compiled to assess the screening accuracy of the PHQ-9. We included studies that used the Structured Clinical Interview for DSM (SCID) as a reference standard and split those into calibration and validation datasets. In the calibration dataset, we estimated unidimensional, two-dimensional (separating cognitive/affective and somatic symptoms of depression), and bi-factor models, and the respective cut-offs to maximize combined sensitivity and specificity. In the validation dataset, we assessed the differences in (combined) sensitivity and specificity between the latent variable approaches and the optimal sum score (⩾10), using bootstrapping to estimate 95% confidence intervals for the differences. The calibration dataset included 24 studies (4378 participants, 652 major depression cases); the validation dataset 17 studies (4252 participants, 568 cases). In the validation dataset, optimal cut-offs of the unidimensional, two-dimensional, and bi-factor models had higher sensitivity (by 0.036, 0.050, 0.049 points, respectively) but lower specificity (0.017, 0.026, 0.019, respectively) compared to the sum score cut-off of ⩾10. In a comprehensive dataset of diagnostic studies, scoring using complex latent variable models do not improve screening accuracy of the PHQ-9 meaningfully as compared to the simple sum score approach.