课题基金 / 基金详情

Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles

Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
使用异质集合整合基因组和临床数据来预测疾病表型
批准号:
10218766
负责人:
Gaurav Pandey
金额:
$54.0万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-06-01 至 2025-03-31

项目摘要

项目成果

Gaurav Pandey的其他基金

相似基金

相关文献

中文摘要
翻译
项目概要 基因组和其他“组学”概况在推进个性化/精准医疗方面具有巨大潜力 能够准确预测个体患者的疾病表型或结果,可用于 由临床医生设计适当的护理计划。然而,尽管有这种潜力,但这些措施的实际影响 疾病表型预测的组学概况可能会受到以下事实的限制:即使是收集大量数据 这些数据没有涵盖足够多的个人。相比之下,各种临床数据类型,例如 正如实验室测试和医生指出的那样,定期收集和研究大量的 在医疗中心接受此类疾病治疗的患者。这些临床数据丰富,并且 它们与多组学数据的互补性,为推进个性化医疗提供了机会 整合这些不同类型的数据。然而,这种数据格式的差异,即几种组学概况 是结构化的,并且一些临床数据类型(例如医生笔记)是非结构化的,构成 这种整合面临的挑战。由于这种差异而带来的一个相关挑战是,不同类别的 计算方法可能是从这些临床数据预测疾病表型的最有效方法。 和组学数据集。这些挑战给当前的数据集成方法解决这个问题带来了障碍 问题。在这里,我们提出了一种通过吸收不同基础表型来实现这种整合的创新方法 从个体临床和组学数据集推断出异质集合的预测因子。这些 集成已经在其他几个计算基因组学问题上显示出前景,可以聚合 基本预测变量的数量和种类不受限制,这是解决此积分问题的理想选择。具体来说,我们 描述如何对单个数据集的现有异构集成方法进行转换和改进 解决多个临床和组学数据集集成问题。我们特别详细介绍了新颖的算法 通过建模和整合固有的患者和数据集来改进这些综合集成 这些数据集中的异质性。我们还提出了利用固有互补性的新算法 临床和组学数据集之间的关系,以及处理预期缺失数据的创新方法,两者 目标是使整体表型预测因子更加准确并适用于患者群体。至 评估这套新颖的面向数据集成的异构集成套件的性能,我们将 验证其在预测哮喘和炎症性肠病表型方面的有效性 具有不同组学和临床数据集的患者队列。我们将公开发布高效软件 实施该项目中开发的方法,使其他人能够进行类似的分析 其他不同的数据收集。拟议工作的成功完成将有助于 通过准确地个体化预测疾病表型来推进个性化医疗。
英文摘要
PROJECT SUMMARY Genomic and other “omic” profiles hold immense potential for advancing personalize/precision medicine by enabling the accurate prediction of disease phenotypes or outcomes for individual patients, which can be used by a clinician to design an appropriate plan of care. However, despite this potential, the actual impact of these omic profiles on disease phenotype prediction may be limited by the fact that even large cohorts collecting these data do not cover large enough numbers of individuals. In contrast, a variety of clinical data types, such as laboratory tests and physician notes, are routinely collected and studied for a much larger number of patients undergoing treatment for such diseases at medical centers. The abundance of these clinical data, and their complementarity with multi-omic data, offer an opportunity to advance personalized medicine by integrating these disparate types of data. However, this disparity in data formats, namely several omic profiles being structured, and several clinical data types, such as physician notes, being unstructured, poses challenges for this integration. An associated challenge due to this disparity is that different classes of computational methods are likely to be the most effective for predicting disease phenotypes from these clinical and omics datasets. These challenges pose barriers for current data integration methods to address this problem. Here, we propose an innovative approach to this integration by assimilating diverse base phenotype predictors inferred from individual clinical and omics datasets into heterogeneous ensembles. These ensembles, which have shown promise for several other computational genomics problems, can aggregate an unrestricted number and variety of base predictors, which is ideal for this integration problem. Specifically, we describe how existing heterogeneous ensemble methods for single datasets can be transformed and advanced to address the multiple clinical and omic dataset integration problem. In particular, we detail novel algorithms for improving these integrative ensembles by modeling and incorporating the inherent patient and dataset heterogeneity in these datasets. We also propose novel algorithms for leveraging the inherent complementarity among clinical and omic datasets, as well as an innovative approach for handling expected missing data, both with the goal of making ensemble phenotype predictors more accurate and applicable to patient cohorts. To assess the performance of this novel suite of data integration-oriented heterogeneous ensembles, we will validate their effectiveness for predicting asthma and Inflammatory Bowel Disease phenotypes in substantial patient cohorts with diverse omics and clinical datasets. We will publicly release efficient software implementations of the methods developed in this project to enable others to carry out similar analyses with other diverse data collections. Successful accomplishment of the proposed work will contribute to the advancement of personalized medicine through accurate individualized prediction of disease phenotypes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multi-modal data integration to identify kinase substrates
Multi-modal data integration to identify kinase substrates
Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
Integrating genomic and clinical data to predict disease phenotypes using heterogeneous ensembles
海外基金