Personalized Risk Predictions with Deep Learning Methods in the Presence of Missing and Biased Electronic Health Record Data
Personalized Risk Predictions with Deep Learning Methods in the Presence of Missing and Biased Electronic Health Record Data
批准号:
10646324
负责人:
Padhraic Smyth
金额:
$33.07万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-08-06 至 2025-05-31
关键词:
AgeAlgorithmsBenchmarkingBiological MarkersCalibrationCardiovascular DiseasesCharacteristicsChronic DiseaseClinicalClinical DataClinical MedicineClinical ResearchComputersComputing MethodologiesDataData SetDevelopmentDiagnosisDiseaseDisease OutcomeDisparateDrug PrescriptionsElectronic Health RecordFrequenciesGenderHealthHealth SurveysHealthcareHospitalsIndividualInstitutionInterviewLinkMedicalMethodologyMethodsModelingNeural Network SimulationNew YorkNew York CityNon-Insulin-Dependent Diabetes MellitusOutcomePatientsPhysical assessmentPopulationProbabilityProceduresRaceRecordsReproducibilityResearchResearch PersonnelRetirementRiskRisk FactorsSample SizeSamplingScientistSiteSoftware ToolsStatistical MethodsSurveysSystemUniversitiesValidationVariantVisitcohortdata modelingdata standardsdeep learningdemographicselectronic health dataexperienceflexibilityhealth care service utilizationimprovedinnovationinterestlearning strategymaltreatmentpatient populationpatient subsetspersonalized risk predictionpopulation basedpopulation healthpredictive modelingrecurrent neural networkrisk predictionrisk prediction modelweb app
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Abstract
Since 2010, clinical medicine has benefited from a rapid surge of clinical research on chronic diseases using
data from electronic health records (EHRs). EHRs are appealing because they can offer large sample sizes,
timely information, and a wealth of clinical information beyond that obtained from either health surveys or
administrative data. However, while millions of patient records are included in large EHR records, they are not
population-representative random samples, a constraint that potentially biases inferences based on such data
and, therefore, has limited their utility for population health research. EHR data typically contain multiple types
of biases, particularly: 1) sampling inclusion bias: EHR data only include information on patients visiting
participating medical systems, and they primarily capture data when patients are ill. Even among populations
with a particular disease, patients represented in EHRs tend to over-represent individuals who are sicker and
have higher health care utilization; 2) sampling frequency bias: the numbers of patients’ encounters and
features in EHRs are at various frequencies and these frequencies correlate with both patients’ characteristics
and outcomes; and 3) institution bias: EHR samples of any hospital reflect the characteristics of patients
population served by that specific hospital. Consequently, EHR-based risk prediction models will have 1)
biases in risk factor selection and estimation for population inferences; 2) disparate mistreatment (unfairness)
in terms of variation in a model’s prediction accuracy across patient subgroups (such as gender, race, and age)
with various sampling inclusion probabilities or frequencies; 3) biased prediction model to reflect characteristics
of patients served by the local hospitals. We propose to develop: 1) effective sample-weighting method to
correct biases in risk factor selection and estimation for population inferences (Aim 1), 2) flexible deep learning
method for EHR personalized risk prediction with fairness criteria (Aim 2); and 3) innovative calibration method
to improve reproducibility of EHR-based risk models between institutions (Aim 3). We will predict risk of
subsequent incident cardiovascular disease (CVD) in patients with type 2 diabetes (T2DM) as a demonstration
of methodology development. Broader use of these methods will be generally applicable to other diseases
outcomes and population of interest. To develop and validate these methods, we propose to analyze three
unique datasets: 1) the New York University Langone Health EHR data (NYU-CDRN, 2009 to now) including
demographics, vitals, diagnoses, lab results, prescriptions, and procedures; 2) the New York City Clinical Data
Research Network (NYC-CDRN)—an EHR network comprising 20 NYC healthcare institutions, including the
NYU-CDRN, with longitudinally linked data on >12 million patient encounters under a Common Data Model,
and 3) the Health and Retirement Survey (HRS, begun in 1992 and ongoing), as a benchmark population-
based cohort, that has nationally representative health interview data for over 20 years, as well as biomarkers,
physical assessment information, prescription drug data, and claims linkages.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Personalized Risk Predictions with Deep Learning Methods in the Presence of Missing and Biased Electronic Health Record Data
-
批准号:10463550
-
项目类别:
-
资助金额:$33.21万
-
财政年份:2021
-
负责人:Padhraic Smyth
-
依托单位:
海外基金