A strategy for validation of variables derived from large-scale electronic health record data.

A strategy for validation of variables derived from large-scale electronic health record data.
复制标题

大规模电子健康记录数据中变量的验证策略。

DOI:
10.1016/j.jbi.2021.103879
复制
发表时间:
2021-09
影响因子:
4.5
通讯作者:
Gupta, Samir
Gupta, Samir
中科院分区:
医学3区
文献类型:
--
作者:
Liu, Lin;Bustamante, Ranier;Earles, Ashley;Demb, Joshua;Messer, Karen;Gupta, Samir

文献摘要

参考文献

被引文献

相似文献

从大规模电子健康记录(EHR)数据中严格验证表型的标准化方法尚未广泛报道。我们提出了一种严谨有效的方法来指导这种验证,包括采样案例和控制的策略,确定样本量,估计算法性能,以及终止验证过程,以下称为圣地亚哥变量验证方法(SDAVV)。我们提出了在图表审查之前应该使用的样本量公式,该公式基于预先指定的正预测值(PPV)和负预测值(NPV)的临界下界。我们还提出了迭代算法开发/验证周期的逐步策略,更新数据抽象的样本量,直到PPV和NPV都达到目标性能。我们将SDAVV应用于退伍军人事务部的一项研究,在这项研究中,我们创建了两种表型算法,一种用于区分正常结肠镜检查病例和异常结肠镜检查对照组,另一种用于识别阿司匹林暴露。估计PPV和NPV均达到0.970,95%可信下限为0.915,识别正常结肠镜病例的估计敏感性为0.963,特异性为0.975。鉴定阿司匹林暴露的表型算法PPV为0.990(95%下界为0.950),NPV为0.980(95%下界为0.930),敏感性和特异性分别为0.960和1.000。从大规模电子病历数据中前瞻性地开发和验证表型算法的结构化方法可以成功实施,并且应该考虑提高“大数据”研究的质量。
Standardized approaches for rigorous validation of phenotyping from large-scale electronic health record (EHR) data have not been widely reported. We proposed a methodologically rigorous and efficient approach to guide such validation, including strategies for sampling cases and controls, determining sample sizes, estimating algorithm performance, and terminating the validation process, hereafter referred to as the San Diego Approach to Variable Validation (SDAVV). We propose sample size formulae which should be used prior to chart review, based on pre-specified critical lower bounds for positive predictive value (PPV) and negative predictive value (NPV). We also propose a stepwise strategy for iterative algorithm development/validation cycles, updating sample sizes for data abstraction until both PPV and NPV achieve target performance. We applied the SDAVV to a Department of Veterans Affairs study in which we created two phenotyping algorithms, one for distinguishing normal colonoscopy cases from abnormal colonoscopy controls and one for identifying aspirin exposure. Estimated PPV and NPV both reached 0.970 with a 95% confidence lower bound of 0.915, estimated sensitivity was 0.963 and specificity was 0.975 for identifying normal colonoscopy cases. The phenotyping algorithm for identifying aspirin exposure reached a PPV of 0.990 (a 95% lower bound of 0.950), an NPV of 0.980 (a 95% lower bound of 0.930), and sensitivity and specificity were 0.960 and 1.000. A structured approach for prospectively developing and validating phenotyping algorithms from large-scale EHR data can be successfully implemented, and should be considered to improve the quality of “big data” research.
DOI: 10.1136/amiajnl-2012-000896
发表时间: 2013-06-01
影响因子: 6.4
作者:
Newton, Katherine M.;Peissig, Peggy L.;Denny, Joshua C.
通讯作者: Denny, Joshua C.
DOI: 10.1146/annurev-biodatasci-080917-013315
发表时间: 2018-01-01
期刊: ANNUAL REVIEW OF BIOMEDICAL DATA SCIENCE, VOL 1
影响因子: --
作者:
Banda, Juan M.;Seneviratne, Martin;Shah, Nigam H.
通讯作者: Shah, Nigam H.
DOI: 10.1186/s12879-016-2020-2
发表时间: 2016-11-17
影响因子: 3.7
作者:
Jackson KL;Mbagwu M;Pacheco JA;Baldridge AS;Viox DJ;Linneman JG;Shukla SK;Peissig PL;Borthwick KM;Carrell DA;Bielinski SJ;Kirby JC;Denny JC;Mentch FD;Vazquez LM;Rasmussen-Torvik LJ;Kho AN
通讯作者: Kho AN
DOI: 10.1177/1087054716672337
发表时间: 2019-11-01
影响因子: 3
作者:
Gruschow SM;Yerys BE;Power TJ;Durbin DR;Curry AE
通讯作者: Curry AE
DOI: 10.1200/cci.17.00072
发表时间: 2018-02-20
影响因子: 4.2
作者:
Earles, Ashley;Liu, Lin;Gupta, Samir
通讯作者: Gupta, Samir