课题基金 / 基金详情

Personalized Risk Predictions with Deep Learning Methods in the Presence of Missing and Biased Electronic Health Record Data

Personalized Risk Predictions with Deep Learning Methods in the Presence of Missing and Biased Electronic Health Record Data
在存在缺失和有偏差的电子健康记录数据的情况下,利用深度学习方法进行个性化风险预测
批准号:
10463550
负责人:
Padhraic Smyth
金额:
$33.21万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-08-06 至 2025-05-31

项目摘要

项目成果

Padhraic Smyth的其他基金

相似基金

相关文献

中文摘要
翻译
摘要 自2010年以来,临床医学受益于慢性疾病临床研究的快速激增,这些研究使用 电子健康记录(EHR)的数据。EHR之所以有吸引力,是因为它们可以提供大样本数量, 及时的信息和丰富的临床信息,而不是从健康调查或 管理数据。然而,尽管数以百万计的患者记录包括在大型电子病历记录中,但它们并没有 具有总体代表性的随机样本,这一限制可能会使基于此类数据的推断产生偏差 因此,它们对人口健康研究的效用也受到了限制。EHR数据通常包含多种类型 偏差,特别是:1)抽样纳入偏差:EHR数据仅包括就诊患者的信息 参与的医疗系统,它们主要在患者生病时捕获数据。即使在人群中 对于一种特定的疾病,EHR中代表的患者往往过多地代表病情较重和 有较高的卫生保健利用率;2)抽样频率偏差:患者的就诊次数和 EHR的特征有不同的频率,这些频率与两个患者的特征相关 和结果;以及3)机构偏差:任何医院的电子病历样本都反映了患者的特征 由该特定医院服务的人口。因此,基于EHR的风险预测模型将有1) 风险因素选择和总体推断估计的偏差;2)不同的虐待(不公平) 根据模型预测准确性在不同患者亚组(如性别、种族和年龄)之间的差异 具有不同的样本包含概率或频率;3)反映特征的有偏预测模型 由当地医院提供服务的病人的数量。我们建议发展:1)有效的样本加权方法 纠正风险因素选择和总体推断估计的偏差(目标1),2)灵活的深度学习 具有公平性准则的EHR个性化风险预测方法(目标2);3)创新的校准方法 提高机构间基于电子病历的风险模型的可重复性(目标3)。我们将预测以下风险 以2型糖尿病(T2 DM)患者的后续心血管疾病(CVD)为例 方法论的发展。这些方法的广泛应用将普遍适用于其他疾病。 结果和相关人口。为了开发和验证这些方法,我们建议分析三个 独特的数据集:1)纽约大学朗格尼健康电子病历数据(NYU-CDRN,2009年至今),包括 人口统计学、生命体征、诊断、实验室结果、处方和程序;2)纽约市临床数据 研究网络(NYC-CDRN)-由纽约市20家医疗机构组成的电子健康记录网络,包括 NYU-CDRN在一个通用数据模型下拥有关于1200万患者就诊的纵向关联数据, 3)健康和退休调查(HRS),开始于1992年,目前仍在进行中,作为基准人口- 基于队列,拥有20多年来具有全国代表性的健康访谈数据,以及生物标志物, 身体评估信息、处方药数据和索赔联系。
英文摘要
Abstract Since 2010, clinical medicine has benefited from a rapid surge of clinical research on chronic diseases using data from electronic health records (EHRs). EHRs are appealing because they can offer large sample sizes, timely information, and a wealth of clinical information beyond that obtained from either health surveys or administrative data. However, while millions of patient records are included in large EHR records, they are not population-representative random samples, a constraint that potentially biases inferences based on such data and, therefore, has limited their utility for population health research. EHR data typically contain multiple types of biases, particularly: 1) sampling inclusion bias: EHR data only include information on patients visiting participating medical systems, and they primarily capture data when patients are ill. Even among populations with a particular disease, patients represented in EHRs tend to over-represent individuals who are sicker and have higher health care utilization; 2) sampling frequency bias: the numbers of patients’ encounters and features in EHRs are at various frequencies and these frequencies correlate with both patients’ characteristics and outcomes; and 3) institution bias: EHR samples of any hospital reflect the characteristics of patients population served by that specific hospital. Consequently, EHR-based risk prediction models will have 1) biases in risk factor selection and estimation for population inferences; 2) disparate mistreatment (unfairness) in terms of variation in a model’s prediction accuracy across patient subgroups (such as gender, race, and age) with various sampling inclusion probabilities or frequencies; 3) biased prediction model to reflect characteristics of patients served by the local hospitals. We propose to develop: 1) effective sample-weighting method to correct biases in risk factor selection and estimation for population inferences (Aim 1), 2) flexible deep learning method for EHR personalized risk prediction with fairness criteria (Aim 2); and 3) innovative calibration method to improve reproducibility of EHR-based risk models between institutions (Aim 3). We will predict risk of subsequent incident cardiovascular disease (CVD) in patients with type 2 diabetes (T2DM) as a demonstration of methodology development. Broader use of these methods will be generally applicable to other diseases outcomes and population of interest. To develop and validate these methods, we propose to analyze three unique datasets: 1) the New York University Langone Health EHR data (NYU-CDRN, 2009 to now) including demographics, vitals, diagnoses, lab results, prescriptions, and procedures; 2) the New York City Clinical Data Research Network (NYC-CDRN)—an EHR network comprising 20 NYC healthcare institutions, including the NYU-CDRN, with longitudinally linked data on >12 million patient encounters under a Common Data Model, and 3) the Health and Retirement Survey (HRS, begun in 1992 and ongoing), as a benchmark population- based cohort, that has nationally representative health interview data for over 20 years, as well as biomarkers, physical assessment information, prescription drug data, and claims linkages.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Personalized Risk Predictions with Deep Learning Methods in the Presence of Missing and Biased Electronic Health Record Data
海外基金