Collaborative Research: New Bayesian Nonparametric Paradigms of Personalized Medicine for Lung Cancer
Collaborative Research: New Bayesian Nonparametric Paradigms of Personalized Medicine for Lung Cancer
批准号:
1854003
负责人:
Subharup Guha
金额:
$35.66万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2020-08-31
中文摘要
快速的技术进步使得从单个肿瘤样本中进行跨多个领域的分子分析成为可能,这为许多疾病,特别是癌症的临床决策提供了支持。关键的挑战是有效地吸收这些领域的信息,以识别可能被药物靶向的基因组特征和生物实体,为未来的患者制定准确的风险预测概况,并确定新的患者亚群,以进行量身定制的治疗和监测。该项目的主要目标是开发一种创新、灵活和可扩展的统计框架,用于分析多域、复杂结构、高通量的现代阵列和下一代基于测序的组学数据集。这项工作的动机是与肺癌有关的几项调查;然而,所提出的方法和计算工具广泛适用于涉及高维数据的各种上下文中。从更广泛的科学角度来看,将这些新方法应用于激励临床和基因组数据集将允许有原则的“结构狩猎”。这将提供更准确的临床结果预测,更大的统计能力来检测重要的生物可操作的生物标志物,以改善癌症诊断和预后的风险估计和治疗选择,并更好地利用生物学领域知识来发现不同平台之间的关系。这将导致后续实施合理的基于生物标志物和个性化的临床试验,提高基于分子标记的个性化治疗的成功率。为了实现这些目标,提出了以下具体目标:(1)开发通用和灵活的统计技术,用于在基于阵列和下一代测序的研究中产生的混合、异构规模的单域数据集中识别肺癌的差异基因组特征。一般的非参数贝叶斯模型基于良好的理论证明将开发和实现使用有效的,可扩展的算法。这些模型提供了生物学上可解释的摘要,并使其适用于各种高通量数据集。(2)构建海量多域数据的整合概率框架,将域内和域间的依赖性连贯地结合起来,准确检测肿瘤亚型并预测临床结果,从而提供与癌症分类相关的基因组畸变目录。(3)培养大规模并行算法和高性能计算和推理工具,大幅减少计算时间,提高高吞吐量数据集的可扩展性。这些可扩展的推理程序能够吸收来自多个平台的信息,并选择具有适当依赖结构的灵活模型,同时检测最优稀疏的非线性机制,以预测和识别肿瘤亚型。由于这些公式完全是概率性的,它们通过考虑不同的变化来源和提供推理不确定性的度量,对纯算法方法提供了实质性的改进。由于现有的基于仿真的算法不能扩展到海量数据集,这些模型的理论特性将被利用来设计数据压缩算法以进行有效的推理。此外,由于传统的cpu受到能源消耗、热量产生和内存访问的限制,利用低成本大规模并行计算工具(如图形处理单元(gpu))的能力的软件将被开发出来并免费提供。
英文摘要
Rapid technological advances have allowed for molecular profiling across multiple domains from a single tumor sample, supporting clinical decision making in many diseases, especially cancer. Key challenges are to effectively assimilate information across these domains to identify genomic signatures and biological entities that may be targeted by drugs, develop accurate risk prediction profiles for future patients, and identify novel patient subgroups for tailored therapy and monitoring. The primary objective of this project is the development of an innovative, flexible and scalable statistical framework for analyzing multi-domain, complex-structured, and high throughput modern array and next generation sequencing-based 'omics datasets. The work is motivated by several investigations related to lung cancer; however, the proposed methods and computational tools are broadly applicable in a variety of contexts involving high-dimensional data. From a broader scientific perspective, the application of these novel methodologies to the motivating clinical and genomic datasets will allow for principled "structure hunting". This will provide more accurate prediction of clinical outcomes, greater statistical power to detect important biologically actionable biomarkers for improved risk estimation and treatment selection for cancer diagnosis and prognosis, and better utilization of biological domain knowledge to find relationships between different platforms. It will lead to subsequent implementation of rational biomarker-based and individualized clinical trials that increase the success rate of personalized therapies based on molecular markers.To achieve these goals, the following specific aims are proposed: (1) Develop versatile and flexible statistical techniques for identifying differential genomic signatures for lung cancer in mixed, heterogeneously scaled single-domain datasets arising from array and next-generation sequencing based studies. A general class of nonparametric Bayesian models based on sound theoretical justifications will be developed and implemented using efficient, scalable algorithms. These models provide biologically interpretable summaries and enable applicability to a wide variety of high-throughput datasets. (2) Formulate integrative probabilistic frameworks for massive multiple-domain data, which coherently incorporate dependence within and between domains to accurately detect tumor subtypes and predict clinical outcomes, thus providing a catalogue of genomic aberrations associated with cancer taxonomy. (3) Foster massively parallel algorithms and high-performance computational and inferential tools that drastically reduce the computation times and increase scalability of high-throughput datasets. These scalable inferential procedures are able to assimilate information from several platforms and select flexible models with the appropriate dependence structures, while detecting optimally sparse, non-linear mechanisms for predicting and identifying tumor subtypes. Because these formulations are fully probabilistic, they offer substantial improvements over purely algorithmic approaches by accounting for different sources of variation and providing measures of inference uncertainty. Since existing simulation-based algorithms do not scale for massive datasets, theoretical properties of these models will be exploited to devise data-squashing algorithms for efficient inference. Furthermore, as traditional CPUs are limited by energy consumption, heat generation and memory access, software that harnesses the power of low-cost massively parallel computing tools such as graphics processing units (GPUs) will be developed and made freely available.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: New Bayesian Nonparametric Paradigms of Personalized Medicine for Lung Cancer
-
批准号:1461948
-
项目类别:Continuing Grant
-
资助金额:$78.0万
-
财政年份:2015
-
负责人:Subharup Guha
-
依托单位:
Bayesian Mixture Models: Unified Theoretical Frameworks and MCMC Methods
-
批准号:0906734
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2009
-
负责人:Subharup Guha
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: