Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
批准号:
10218766
负责人:
Gaurav Pandey
金额:
$54.0万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-06-01 至 2025-03-31
关键词:
AddressAlgorithmsAsthmaAutomobile DrivingCaringCharacteristicsClinicalClinical DataComputational algorithmComputer softwareComputing MethodologiesDataData CollectionData SetDiseaseDisease OutcomeDockingEffectivenessElectronic Health RecordEncapsulatedExerciseGenomicsGoalsHealthIndividualInflammatory Bowel DiseasesInstitutionLaboratoriesLearningMalignant NeoplasmsMedicalMedical ImagingMedical centerMethodsModalityModelingMolecularMolecular ProfilingMultiomic DataPatientsPerformancePhenotypePhysiciansPopulationRecurrenceResearch PersonnelRiskSamplingStructureTechnologyTestingThe Cancer Genome AtlasValidationVariantWorkadvanced diseasebaseclinical phenotypecohortdata formatdata integrationdeep learningdesigndisease phenotypediverse datafeature selectionflexibilitygenomic dataheterogenous dataimprovedindividual patientinnovationinsightmembermultiple datasetsmultiple omicsmultitasknovelnovel strategiesoutreachpatient populationpersonalized medicinepersonalized predictionsprecision medicinepredictive modelingprogramsrapid growthrepositoryscale uptranscriptomicsvector
中文摘要
项目总结
基因组和其他“基因组”图谱具有通过以下方式推进个性化/精准医学的巨大潜力
能够准确地预测个别患者的疾病表型或结果,可以使用
由临床医生设计适当的护理计划。然而,尽管有这种潜力,这些因素的实际影响
疾病表型预测的基因组图谱可能受到这样一个事实的限制,即即使是大的队列收集
这些数据没有涵盖足够多的个人。相比之下,各种临床数据类型,如
正如实验室测试和医生注意到的那样,常规收集和研究的数量要多得多
在医疗中心接受此类疾病治疗的患者。这些丰富的临床数据,以及
它们与多组数据的互补性,为通过以下方式推进个性化医疗提供了机会
集成这些不同类型的数据。然而,数据格式的这种差异,即几个OMIC配置文件
是结构化的,以及几种临床数据类型,如医生笔记,是非结构化的,
这种整合面临的挑战。这种差异带来的一个相关挑战是,不同类别的
计算方法可能是从这些临床数据中预测疾病表型的最有效方法。
和组学数据集。这些挑战给当前的数据集成方法解决这一问题带来了障碍
有问题。在这里,我们提出了一种通过吸收不同的碱基表型来实现这种整合的创新方法
从个体临床和组学数据集中推断出不同的集合的预测者。这些
集合已经显示出对其他几个计算基因组学问题的希望,可以聚合一个
不限数量和种类的基本预报器,这是此集成问题的理想选择。具体来说,我们
描述如何对单个数据集的现有异类集成方法进行转换和改进
以解决多个临床和基因组数据集集成问题。特别是,我们详细介绍了新的算法
通过对固有的患者和数据集进行建模和合并来改进这些集成集合
这些数据集中的异质性。我们还提出了利用内在互补性的新算法
在临床和基因组数据集中,以及处理预期丢失数据的创新方法,两者都
目的是使整体表型预测指标更准确,更适用于患者队列。至
评估这套新的面向数据集成的异类集成的性能,我们将
验证它们对预测哮喘和炎症性肠病表型的有效性
具有不同组学和临床数据集的患者队列。我们将公开发布高效软件
实施本项目中开发的方法,使其他人能够执行类似的分析
其他不同的数据收集。成功完成拟议的工作将有助于
通过疾病表型的准确个体化预测促进个性化医学的发展。
英文摘要
PROJECT SUMMARY
Genomic and other “omic” profiles hold immense potential for advancing personalize/precision medicine by
enabling the accurate prediction of disease phenotypes or outcomes for individual patients, which can be used
by a clinician to design an appropriate plan of care. However, despite this potential, the actual impact of these
omic profiles on disease phenotype prediction may be limited by the fact that even large cohorts collecting
these data do not cover large enough numbers of individuals. In contrast, a variety of clinical data types, such
as laboratory tests and physician notes, are routinely collected and studied for a much larger number of
patients undergoing treatment for such diseases at medical centers. The abundance of these clinical data, and
their complementarity with multi-omic data, offer an opportunity to advance personalized medicine by
integrating these disparate types of data. However, this disparity in data formats, namely several omic profiles
being structured, and several clinical data types, such as physician notes, being unstructured, poses
challenges for this integration. An associated challenge due to this disparity is that different classes of
computational methods are likely to be the most effective for predicting disease phenotypes from these clinical
and omics datasets. These challenges pose barriers for current data integration methods to address this
problem. Here, we propose an innovative approach to this integration by assimilating diverse base phenotype
predictors inferred from individual clinical and omics datasets into heterogeneous ensembles. These
ensembles, which have shown promise for several other computational genomics problems, can aggregate an
unrestricted number and variety of base predictors, which is ideal for this integration problem. Specifically, we
describe how existing heterogeneous ensemble methods for single datasets can be transformed and advanced
to address the multiple clinical and omic dataset integration problem. In particular, we detail novel algorithms
for improving these integrative ensembles by modeling and incorporating the inherent patient and dataset
heterogeneity in these datasets. We also propose novel algorithms for leveraging the inherent complementarity
among clinical and omic datasets, as well as an innovative approach for handling expected missing data, both
with the goal of making ensemble phenotype predictors more accurate and applicable to patient cohorts. To
assess the performance of this novel suite of data integration-oriented heterogeneous ensembles, we will
validate their effectiveness for predicting asthma and Inflammatory Bowel Disease phenotypes in substantial
patient cohorts with diverse omics and clinical datasets. We will publicly release efficient software
implementations of the methods developed in this project to enable others to carry out similar analyses with
other diverse data collections. Successful accomplishment of the proposed work will contribute to the
advancement of personalized medicine through accurate individualized prediction of disease phenotypes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multi-modal data integration to identify kinase substrates
-
批准号:10659156
-
项目类别:
-
资助金额:$50.7万
-
财政年份:2022
-
负责人:Gaurav Pandey
-
依托单位:
Multi-modal data integration to identify kinase substrates
-
批准号:10451941
-
项目类别:
-
资助金额:$49.85万
-
财政年份:2022
-
负责人:Gaurav Pandey
-
依托单位:
Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
-
批准号:10589827
-
项目类别:
-
资助金额:$60.52万
-
财政年份:2021
-
负责人:Gaurav Pandey
-
依托单位:
Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
-
批准号:10409755
-
项目类别:
-
资助金额:$53.55万
-
财政年份:2021
-
负责人:Gaurav Pandey
-
依托单位:
Boosting the Translational Impact of Scientific Competitions by Ensemble Learning
-
批准号:8864679
-
项目类别:
-
资助金额:$44.59万
-
财政年份:2015
-
负责人:Gaurav Pandey
-
依托单位:
海外基金