Mixture models for undiagnosed prevalent disease and interval-censored incident disease: applications to a cohort assembled from electronic health records

Mixture models for undiagnosed prevalent disease and interval-censored incident disease: applications to a cohort assembled from electronic health records
复制标题

DOI:
10.1002/sim.7380
复制
发表时间:
2017-09-30
影响因子:
2
通讯作者:
Katki, Hormuzd A.
Katki, Hormuzd A.
中科院分区:
医学3区
文献类型:
--
作者:
Cheung, Li C.;Pan, Qing;Katki, Hormuzd A.

文献摘要

被引文献

相似文献

为了提高成本效益和效率,正在使用电子健康记录的大型卫生保健提供者中进行许多大规模的通用队列研究。这类数据的两个关键特征是,在不定期访问之间对突发疾病进行间隔审查,并且可能存在预先存在的(流行的)疾病。由于流行疾病并不总是立即被诊断出来,一些在后来就诊时诊断出来的疾病实际上是未被诊断出来的流行疾病。在临床应用中,我们将流行疾病视为时间为零的点质量,而对流行疾病的发病时间不感兴趣。我们证明了朴素Kaplan-Meier累积风险估计器低估了早期时间点的风险,高估了后期风险。我们提出了一个一般家族的混合模型,用于未诊断的流行疾病和间隔审查的事件疾病,我们称之为患病率-发病率模型。参数患病率-发病率模型的参数,如逻辑回归和威布尔生存(logistic-Weibull)模型,是通过直接似然最大化或EM算法估计的。提出了非参数方法来计算无协变量情况下的累积风险。我们比较了Kaiser Permanente北加州宫颈癌筛查项目中累积风险的朴素Kaplan-Meier、logistic-Weibull和非参数估计。Kaplan-Meier提供了较差的估计,而logistic-Weibull模型与非参数模型非常接近。我们的研究结果支持我们使用logistic-Weibull模型来开发风险估计,这是当前美国基于风险的宫颈癌筛查指南的基础。2017年出版。这篇文章是由美国政府雇员贡献的,他们的工作在美国属于公有领域。
For cost-effectiveness and efficiency, many large-scale general-purpose cohort studies are being assembled within large health-care providers who use electronic health records. Two key features of such data are that incident disease is interval-censored between irregular visits and there can be pre-existing (prevalent) disease. Because prevalent disease is not always immediately diagnosed, some disease diagnosed at later visits are actually undiagnosed prevalent disease. We consider prevalent disease as a point mass at time zero for clinical applications where there is no interest in time of prevalent disease onset. We demonstrate that the naive Kaplan-Meier cumulative risk estimator underestimates risks at early time points and overestimates later risks. We propose a general family of mixture models for undiagnosed prevalent disease and interval-censored incident disease that we call prevalence-incidence models. Parameters for parametric prevalence-incidence models, such as the logistic regression and Weibull survival (logistic-Weibull) model, are estimated by direct likelihood maximization or by EM algorithm. Non-parametric methods are proposed to calculate cumulative risks for cases without covariates. We compare naive Kaplan-Meier, logistic-Weibull, and non-parametric estimates of cumulative risk in the cervical cancer screening program at Kaiser Permanente Northern California. Kaplan-Meier provided poor estimates while the logistic-Weibull model was a close fit to the non-parametric. Our findings support our use of logistic-Weibull models to develop the risk estimates that underlie current US risk-based cervical cancer screening guidelines. Published 2017. This article has been contributed to by US Government employees and their work is in the public domain in the USA.