Phenotyping coronavirus disease 2019 during a global health pandemic: Lessons learned from the characterization of an early cohort.

Phenotyping coronavirus disease 2019 during a global health pandemic: Lessons learned from the characterization of an early cohort.
复制标题

DOI:
10.1016/j.jbi.2021.103777
复制
发表时间:
2021-05
影响因子:
4.5
通讯作者:
Bastarache L
Bastarache L
中科院分区:
医学3区
文献类型:
--
作者:
DeLozier S;Bland HT;McPheeters M;Wells Q;Farber-Eger E;Bejan CA;Fabbri D;Rosenbloom T;Roden D;Johnson KB;Wei WQ;Peterson J;Bastarache L

文献摘要

被引文献

相似文献

自2019年冠状病毒病(新冠肺炎)大流行开始,研究人员一直将电子健康记录数据作为研究可能的风险因素和结果的一种方式。为了确保使用这些数据的研究的有效性和准确性,调查人员需要相信他们构建的表型是可靠和准确的,反映了他们被确定的医疗保健环境。我们在一个学术医学中心建立了一个新冠肺炎注册中心,并使用2020年3月1日至6月5日的数据分别评估了大流行年份和非大流行年份人口水平特征的差异。先前显示影响2型糖尿病表型表现的EHR中值长度,在SARS-CoV-2阳性组中显著短于2019年流感测试组(中位数3.1年vs8.7年;Wilcoxon秩和P=81.3e-52)。使用三种增加复杂性的表型方法(单独使用账单代码和由电子病历供应商和临床专家提供的特定领域的算法),从新冠肺炎电子病历中提取常见的医疗合并症,定义为实验室测试阳性(阳性预测值为100%,召回率为93%)。在结合了不同表型方法的表现数据后,我们观察到那些为全面护理访问(p=4e-11)开出账单的记录和那些记录了完整人口统计学数据的记录(p=7e-5)的假阴性率显著降低。在早期的新冠肺炎队列中,我们发现9种常见合并症的表型表现受到EHR长度中值以及数据密度的影响,这与之前的研究一致,数据密度可以使用包括CPT码在内的便携指标来测量。在这里,我们展示了创建表型深刻、尖锐的新冠肺炎队列所面临的挑战和潜在的解决方案。
From the start of the coronavirus disease 2019 (COVID-19) pandemic, researchers have looked to electronic health record (EHR) data as a way to study possible risk factors and outcomes. To ensure the validity and accuracy of research using these data, investigators need to be confident that the phenotypes they construct are reliable and accurate, reflecting the healthcare settings from which they are ascertained. We developed a COVID-19 registry at a single academic medical center and used data from March 1 to June 5, 2020 to assess differences in population-level characteristics in pandemic and non-pandemic years respectively. Median EHR length, previously shown to impact phenotype performance in type 2 diabetes, was significantly shorter in the SARS-CoV-2 positive group relative to a 2019 influenza tested group (median 3.1 years vs 8.7; Wilcoxon rank sum P = 1.3e-52). Using three phenotyping methods of increasing complexity (billing codes alone and domain-specific algorithms provided by an EHR vendor and clinical experts), common medical comorbidities were abstracted from COVID-19 EHRs, defined by the presence of a positive laboratory test (positive predictive value 100%, recall 93%). After combining performance data across phenotyping methods, we observed significantly lower false negative rates for those records billed for a comprehensive care visit (p = 4e-11) and those with complete demographics data recorded (p = 7e-5). In an early COVID-19 cohort, we found that phenotyping performance of nine common comorbidities was influenced by median EHR length, consistent with previous studies, as well as by data density, which can be measured using portable metrics including CPT codes. Here we present those challenges and potential solutions to creating deeply phenotyped, acute COVID-19 cohorts.