A Poisson binomial-based statistical testing framework for comorbidity discovery across electronic health record datasets.

A Poisson binomial-based statistical testing framework for comorbidity discovery across electronic health record datasets.
复制标题

DOI:
10.1038/s43588-021-00141-9
复制
发表时间:
2021-10
期刊:
Nature computational science
影响因子:
--
通讯作者:
Yandell M
Yandell M
中科院分区:
其他
文献类型:
--
作者:
Lemmon G;Wesolowski S;Henrie A;Tristani-Firouzi M;Yandell M

文献摘要

相似文献

发现患者中不同医疗条件的伴随发生,也称为共病,是创建患者结果预测工具的先决条件。目前的共病发现应用程序是为小数据集设计的,并使用分层来控制混杂的变量,如年龄、性别或血统。分层降低了假阳性率,但也降低了力量,因为研究队列的规模减少了。在这里,我们描述了一种基于泊松二项式的共病发现(PBC)方法,该方法专为大数据应用而设计,避免了分层的需要。PBC在每个患者的基础上针对混淆的人口统计变量进行调整,并建立时间关系模型。我们使用两个数据集对PBC进行基准测试,以计算4,623,841对潜在共病医学术语的共病统计数据。该计算的结果作为可搜索的网络资源提供。与目前的方法相比,PBC方法减少了假阳性关联,同时保留了发现真实并存的统计能力。
Discovering the concomitant occurrence of distinct medical conditions in a patient, also known as comorbidities, is a prerequisite for creating patient outcome prediction tools. Current comorbidity discovery applications are designed for small datasets and use stratification to control for confounding variables such as age, sex or ancestry. Stratification lowers false positive rates, but reduces power, as the size of the study cohort is decreased. Here we describe a Poisson binomial-based approach to comorbidity discovery (PBC) designed for big-data applications that circumvents the need for stratification. PBC adjusts for confounding demographic variables on a per-patient basis and models temporal relationships. We benchmark PBC using two datasets to compute comorbidity statistics on 4,623,841 pairs of potentially comorbid medical terms. The results of this computation are provided as a searchable web resource. Compared with current methods, the PBC approach reduces false positive associations while retaining statistical power to discover true comorbidities.