Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
批准号:
10409755
负责人:
Gaurav Pandey
金额:
$53.55万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-06-01 至 2025-03-31
关键词:
AddressAlgorithmsAsthmaAutomobile DrivingCaringCharacteristicsClassificationClinicalClinical DataComputational algorithmComputer softwareComputing MethodologiesDataData CollectionData SetDiseaseDisease OutcomeDockingEffectivenessElectronic Health RecordElectronic Medical Records and Genomics NetworkEncapsulatedExerciseGenomicsGoalsHealthImageIndividualInflammatory Bowel DiseasesInstitutionLaboratoriesLearningMalignant NeoplasmsMedical ImagingMedical centerMethodsModalityModelingMolecularMolecular ProfilingMultiomic DataPatientsPerformancePhenotypePhysiciansPopulationRecurrenceResearch PersonnelRiskSamplingStructureTechniquesTechnologyTestingThe Cancer Genome AtlasValidationVariantWorkadvanced diseasebasebiobankclinical phenotypecohortdata formatdata integrationdeep learningdesigndisease phenotypediverse datafeature selectionflexibilitygenomic dataheterogenous dataimprovedindividual patientinnovationinsightmembermulti-task learningmultiple datasetsmultiple omicsnovelnovel strategiesoutreachpatient populationpersonalized medicinepersonalized predictionsprecision medicinepredictive modelingprogramsrandom forestrapid growthrepositoryscale uptranscriptomicsvector
中文摘要
点击翻译按钮获取中文摘要
英文摘要
PROJECT SUMMARY
Genomic and other “omic” profiles hold immense potential for advancing personalize/precision medicine by
enabling the accurate prediction of disease phenotypes or outcomes for individual patients, which can be used
by a clinician to design an appropriate plan of care. However, despite this potential, the actual impact of these
omic profiles on disease phenotype prediction may be limited by the fact that even large cohorts collecting
these data do not cover large enough numbers of individuals. In contrast, a variety of clinical data types, such
as laboratory tests and physician notes, are routinely collected and studied for a much larger number of
patients undergoing treatment for such diseases at medical centers. The abundance of these clinical data, and
their complementarity with multi-omic data, offer an opportunity to advance personalized medicine by
integrating these disparate types of data. However, this disparity in data formats, namely several omic profiles
being structured, and several clinical data types, such as physician notes, being unstructured, poses
challenges for this integration. An associated challenge due to this disparity is that different classes of
computational methods are likely to be the most effective for predicting disease phenotypes from these clinical
and omics datasets. These challenges pose barriers for current data integration methods to address this
problem. Here, we propose an innovative approach to this integration by assimilating diverse base phenotype
predictors inferred from individual clinical and omics datasets into heterogeneous ensembles. These
ensembles, which have shown promise for several other computational genomics problems, can aggregate an
unrestricted number and variety of base predictors, which is ideal for this integration problem. Specifically, we
describe how existing heterogeneous ensemble methods for single datasets can be transformed and advanced
to address the multiple clinical and omic dataset integration problem. In particular, we detail novel algorithms
for improving these integrative ensembles by modeling and incorporating the inherent patient and dataset
heterogeneity in these datasets. We also propose novel algorithms for leveraging the inherent complementarity
among clinical and omic datasets, as well as an innovative approach for handling expected missing data, both
with the goal of making ensemble phenotype predictors more accurate and applicable to patient cohorts. To
assess the performance of this novel suite of data integration-oriented heterogeneous ensembles, we will
validate their effectiveness for predicting asthma and Inflammatory Bowel Disease phenotypes in substantial
patient cohorts with diverse omics and clinical datasets. We will publicly release efficient software
implementations of the methods developed in this project to enable others to carry out similar analyses with
other diverse data collections. Successful accomplishment of the proposed work will contribute to the
advancement of personalized medicine through accurate individualized prediction of disease phenotypes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multi-modal data integration to identify kinase substrates
-
批准号:10659156
-
项目类别:
-
资助金额:$50.7万
-
财政年份:2022
-
负责人:Gaurav Pandey
-
依托单位:
Multi-modal data integration to identify kinase substrates
-
批准号:10451941
-
项目类别:
-
资助金额:$49.85万
-
财政年份:2022
-
负责人:Gaurav Pandey
-
依托单位:
Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
-
批准号:10218766
-
项目类别:
-
资助金额:$54.0万
-
财政年份:2021
-
负责人:Gaurav Pandey
-
依托单位:
Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
-
批准号:10589827
-
项目类别:
-
资助金额:$60.52万
-
财政年份:2021
-
负责人:Gaurav Pandey
-
依托单位:
Boosting the Translational Impact of Scientific Competitions by Ensemble Learning
-
批准号:8864679
-
项目类别:
-
资助金额:$44.59万
-
财政年份:2015
-
负责人:Gaurav Pandey
-
依托单位:
海外基金