Modeling the Incompleteness and Biases of Health Data
Modeling the Incompleteness and Biases of Health Data
批准号:
10581658
负责人:
YUAN LUO
金额:
$30.75万
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-06-01 至 2025-03-31
关键词:
AdoptionAlgorithmsArchitectureAwarenessClinicalClinical DataClinical ResearchCohort StudiesCollaborationsCollectionCommunitiesComputer softwareCritical CareDataData CollectionData SetDependenceDerivation procedureDevelopmentDiagnosticDiagnostic testsElectronic Health RecordElectronic Medical Records and Genomics NetworkEvolutionFunctional disorderGeneral HospitalsGoalsHealthcareHealthcare SystemsHospitalsHourIndividualInpatientsInstitutionIntuitionKnowledgeKnowledge DiscoveryLaboratoriesLearningMeasurementMedicalMemoryMethodologyMethodsModelingOutcomePatient CarePatient-Focused OutcomesPatientsPerformanceProceduresProcessProtocols documentationRegimenResearchResearch PersonnelResource-limited settingResourcesRoleScheduleStructureSymptomsSystemTest ResultTestingTimeTrainingValidationclinical decision supportclinical decision-makingdata miningdata qualitydesignflexibilitygraph neural networkhealth care service utilizationhealth dataimprovedlifetime riskmachine learning algorithmneglectneural network architecturenovelopen sourcepatient populationpersonalized diagnosticspersonalized therapeuticpredictive modelingrecurrent neural networkscale upsocial health determinantsstemstructured datatext searchingtooltrait
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Modeling the Incompleteness and Biases of Health Data
Researchers are increasingly working to “mine” health data to derive new medical knowledge. Unlike
experimental data that are collected per a research protocol, the primary role of clinical data is to help
clinicians care for patients, so the procedures for its collection are not often systematic. Thus, missing and/or
biased data can hinder medical knowledge discovery and data mining efforts. Existing efforts for missing health
data imputation often focus on only cross-sectional correlation (e.g., correlation across subjects or across
variables) but neglect autocorrelation (e.g., correlation across time points). Moreover, they often focus on
modeling incompleteness but neglect the biases in health data.
Modeling both the incompleteness and bias may contribute to better understanding of health data and better
support clinical decision making. We propose a novel framework of Bias-Aware Missing data Imputation with
Cross-sectional correlation and Autocorrelation (BAMICA), and leverage clinical notes to better inform the
methods that will otherwise rely on structured health data only. In addition to evaluating its imputation
accuracy, we will apply the proposed framework to assist in downstream tasks such as predictive modeling for
multiple outcomes across a diverse range of clinical and cohort study datasets.
Aim 1 introduces the MICA framework to jointly consider cross-sectional correlation and auto-correlation. In
Aim 2, we will augment MICA to be bias-aware (hence BAMICA) to account for biases stemmed from multiple
roots such as healthcare process and use them as features in imputing missing health data. This augmentation
is achieved by a novel recurrent neural network architecture that keeps track of both evolution of health data
variables and bias factors. In Aim 3, we will supplement unstructured clinical notes to structured health data for
modeling incompleteness and biases using a novel architecture of graph neural network on top of memory
network. We will apply graph neural networks to process clinical notes in order to learn proper representations
as input to the memory networks for imputation and downstream predictive modeling tasks. Depending on the
clinical problem and data availability, not all modules may be needed. Thus our proposed BAMICA framework
is designed to be flexible and consists of selectable modules to meet some or all of the above needs.
In summary, our proposal bridges a key knowledge gap in jointly modeling incompleteness and biases in
health data and utilizes unstructured clinical notes to supplement and augment such modeling in order to better
support predictive modeling and clinical decision making. We will demonstrate generalizability by
experimenting on four large clinical and cohort study datasets, and by scaling up to the eMERGE network
spanning 11 institutions nationwide. We will disseminate the open-source framework. The principled and
flexible framework generated by this project will bring significant methodological advancement and have a
direct impact on enhancing discovery from health data.
期刊论文(22)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1016/j.hfc.2021.12.002
发表时间:
2022-04
期刊:
Heart failure clinics
影响因子:
3.4
作者:
[Ahmad FS, Luo Y, Wehbe RM, Thomas JD, Shah SJ]
通讯作者:
Shah SJ
DOI:
10.1016/j.eclinm.2023.102252
发表时间:
2023-11
期刊:
ECLINICALMEDICINE
影响因子:
15.1
作者:
[Wosten-van Asperen, Roelie M., la Roi-Teeuw, Hannah M., van Amstel, Rombout B. E., Bos, Lieuwe D. J., Tissing, Wim J. E., Jordan, Iolanda, Dohna-Schwake, Christian, Bottari, Gabriella, Pappachan, John, Crazzolara, Roman, Comoretto, Rosanna I., Mizia-Malarz, Agniezka, Moscatelli, Andrea, Sanchez-Martin, Maria, Willems, Jef, Rogerson, Colin M., Bennett, Tellen D., Luo, Yuan, Atreya, Mihir R., Faustino, E. Vincent S., Geva, Alon, Weiss, Scott L., Schlapbach, Luregn J., Sanchez-Pinto, L. Nelson]
通讯作者:
Sanchez-Pinto, L. Nelson
DOI:
10.1016/j.crmeth.2023.100503
发表时间:
2023-07-24
期刊:
Cell reports methods
影响因子:
--
作者:
[]
通讯作者:
Hyperchloremia in critically ill patients: association with outcomes and prediction using electronic health record data.
危重患者的高氯血症:与结果的关联以及使用电子健康记录数据的预测。
DOI:
10.1186/s12911-020-01326-4
发表时间:
2020-12-15
期刊:
BMC medical informatics and decision making
影响因子:
3.5
作者:
[Yeh P, Pan Y, Sanchez-Pinto LN, Luo Y]
通讯作者:
Luo Y
DOI:
10.1186/s12911-020-01318-4
发表时间:
2020-12-30
期刊:
BMC medical informatics and decision making
影响因子:
3.5
作者:
[Ye J, Yao L, Shen J, Janarthanam R, Luo Y]
通讯作者:
Luo Y
共 15 条
Modeling the Incompleteness and Biases of Health Data
-
批准号:10381541
-
项目类别:
-
资助金额:$31.13万
-
财政年份:2020
-
负责人:YUAN LUO
-
依托单位:
National Infrastructure for Standardized and Portable EHR Phenotyping Algorithms
-
批准号:10021669
-
项目类别:
-
资助金额:$70.69万
-
财政年份:2017
-
负责人:YUAN LUO
-
依托单位:
In vivo Studies of Ginkgo biloba Neuroprotection
-
批准号:7455616
-
项目类别:
-
资助金额:$4.5万
-
财政年份:2004
-
负责人:YUAN LUO
-
依托单位:
In vivo Studies of Ginkgo biloba Neuroprotection
-
批准号:7188740
-
项目类别:
-
资助金额:$19.48万
-
财政年份:2004
-
负责人:YUAN LUO
-
依托单位:
In vivo Studies of Ginkgo biloba Neuroprotection
-
批准号:7070002
-
项目类别:
-
资助金额:$25.63万
-
财政年份:2004
-
负责人:YUAN LUO
-
依托单位:
In vivo Studies of Ginkgo biloba Neuroprotection
-
批准号:7283658
-
项目类别:
-
资助金额:$25.63万
-
财政年份:2004
-
负责人:YUAN LUO
-
依托单位:
In vivo Studies of Ginkgo biloba Neuroprotection
-
批准号:6947778
-
项目类别:
-
资助金额:$3.9万
-
财政年份:2004
-
负责人:YUAN LUO
-
依托单位:
In vivo Studies of Ginkgo biloba Neuroprotection
-
批准号:7694239
-
项目类别:
-
资助金额:$4.5万
-
财政年份:2004
-
负责人:YUAN LUO
-
依托单位:
In vivo Studies of Ginkgo biloba Neuroprotection
-
批准号:6827981
-
项目类别:
-
资助金额:$24.34万
-
财政年份:2004
-
负责人:YUAN LUO
-
依托单位:
SIGNALING MECHANISMS IN DOPAMINE RECEPTOR SYNERGISM
-
批准号:7235701
-
项目类别:
-
资助金额:$25.34万
-
财政年份:2003
-
负责人:YUAN LUO
-
依托单位:
Mechanisms of Ginkgo biloba Neuropotection
-
批准号:6534554
-
项目类别:
-
资助金额:$18.0万
-
财政年份:2001
-
负责人:YUAN LUO
-
依托单位:
Mechanisms of Ginkgo biloba Neuropotection
-
批准号:6860782
-
项目类别:
-
资助金额:$7.19万
-
财政年份:2001
-
负责人:YUAN LUO
-
依托单位:
Mechanisms of Ginkgo biloba Neuropotection
-
批准号:6435383
-
项目类别:
-
资助金额:$16.9万
-
财政年份:2001
-
负责人:YUAN LUO
-
依托单位:
FUNCTION OF G PROTEIN AND PCP2 IN CEREBELLUM DEVELOPMENT
-
批准号:6084028
-
项目类别:
-
资助金额:$14.43万
-
财政年份:2000
-
负责人:YUAN LUO
-
依托单位:
海外基金