课题基金 / 基金详情

Modelling of Graphs, Networks and Trees for Genomic Applications: High-Dimensional Model Search

Modelling of Graphs, Networks and Trees for Genomic Applications: High-Dimensional Model Search
基因组应用的图、网络和树建模:高维模型搜索
批准号:
0342172
负责人:
Mike West
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-07-01 至 2010-06-30

项目摘要

项目成果

Mike West的其他基金

相似基金

相关文献

中文摘要
翻译
本研究关注功能基因组学应用中统计建模的理论、方法和计算工具的发展。总体主题是“高维模型搜索”,即开发用于生成、评估和解释涉及许多变量的生物问题的相关候选模型集的工具和方法。研究包括对非常大规模图形模型的模型和计算方法的调查,特别是高斯图形模型和非常稀疏的模型结构,作为观察和实验数据中生物变量之间关联的结构和模式的代表。基因和蛋白质表达研究的背景产生了多个激励问题和案例研究的数据。在实际维数(数千)的问题中,任意一个这样的图形模型在变量域(节点数)和参数数上都是高维的;但是,最根本的挑战是给定问题的候选模型的数量是天文数字。在拟合和评估候选模型时,通常许多候选模型基于对现有观测或实验数据的拟合以及相关生物信息的背景下具有似是而非的支持。因此,需要寻找方法来识别、探索和评价许多似是而非的模型,并从统计和实质性生物学的角度来解释和描述它们的相似性和差异性。其他类别的回归模型,包括非线性统计分类和预测树模型,代表了灵活的框架,用于探索、组合和利用多种形式的生物学数据,用于预测预后、诊断和发现。在基因组环境中,高通量分子信息,如基因表达数据,导致非常非常多的潜在预测变量需要评估,因此,再次,统计学的挑战包括处理和探索非常高维模型的空间。这些核心的研究挑战——变量选择和模型的不确定性,有大量的变量,是基于当前的数学和统计方法无法回答的——形成了这个项目的核心焦点。该项目受到许多基因组学研究的启发,并与这些研究相结合,这些研究为功能基因组学在几个应用领域提供了数据、合作者和应用背景,该项目涉及模型开发和在复杂和非常高维的统计模型空间中随机搜索方法的创新。这包括相关的模型理论、方法和基于分布式集群实现的计算算法开发,以及在许多生物医学研究中的反馈应用。现代生物医学现在能够获得基于创新和进步,特别是基因组技术,迅速升级的规模和复杂性的数据。这包括来自实验室和人类观察疾病和暴露研究的DNA微阵列基因表达研究的高通量基因组数据,来自蛋白质组学和代谢分析技术的相关大规模分子信息,正在迅速扩展到基因组规模序列数据的遗传和序列信息,当然还有传统形式的临床、环境和人口数据。在基础生物学和利用这类信息为人类健康研究提供信息和提供帮助方面取得进展,需要在分析和解释这些规模和复杂性不断增加的数据集的能力方面取得非常大的进步。例如,将多种形式的此类数据用于定义预测性表型是临床基因组学新兴领域的核心。现代数学和计算建模者面临的挑战是那些非常高维的数学模型规范和分析,再加上需要计算工具来搜索天文数字的候选模型,这一努力远远超出了目前实现、评估和理解的能力。这个挑战定义了这个研究项目的核心议程:为非常高维的统计模型开发统计和计算工具。该研究结合了基因组研究的综合收集,这些研究提供了数据,生物学合作者和癌症和心血管生物学几个领域的功能基因组学应用,细胞周期调控和肿瘤发生的途径研究,癌症蛋白质组学,转录调控和其他领域。研究的核心是在复杂和非常高维的统计模型空间上的随机搜索方法(即模拟技术)以及相关的模型理论和方法的创新。该研究的内在智力价值在于计算和统计建模的进步和创新的新方法,以及特定的生物学应用。这项研究的更广泛影响在于所产生的方法和工具在这些具体领域、其他相关生物医学/基因组学应用以及其他科学领域的应用中的适用性。
英文摘要
This research concerns the development of theory, methods and computational tools for statistical modelling motivated by applications in functional genomics. The over-arching theme is that of "High-Dimensional Model Search," i.e. the development of tools and methods for generating, evaluating and interpreting sets of relevant candidate models of a biological problem involving many variables. Research includes the investigation of models and computational methods for very large-scale graphical models, especially Gaussian graphical models and very sparse model structures, as representatives of the structure and patterns of association between biological variables from observed and experimental data. Contexts of gene and protein expression studies generate multiple motivating problems and data for case studies. In problems of realistic dimension (thousands) any one such graphical model is high-dimensional in both the variable domain (number of nodes) and numbers of parameters; but, the fundamental challenge is that the number of candidate models for a given problem is astronomical. In fitting and evaluating candidate models, it is typical that many candidates have plausible support based on fit to the available observational or experimental data and in the context of relevant biological information. Hence the need for methods to identify, explore and evaluate many plausible models, and to interpret and characterize their similarities and differences in both statistical and substantive biological terms. Additional classes of regression models, including nonlinear statistical classification and prediction tree models, represent flexible frameworks for the exploration, combination and utilization of multiple forms of biological data in predictive phenotyping for prognosis, diagnosis and discovery. In genomic contexts, high-throughput molecular information such as gene expression data leads to very, very many potential predictor variables to be assessed, so that, again, the challenges to statistics include dealing with and exploring spaces of very high-dimensional models. These central research challenges -- variable selection and model uncertainty that, with large numbers of variables, are simply unanswerable based on current mathematical and statistical methods -- form the core focus of this project. Motivated by and integrated with a number of genomic studies that provide data, collaborators and application contexts in functional genomics in several areas of application, this project involves model development and innovation in methods of stochastic search over complex and very high-dimensional statistical model spaces. This includes associated model theory, methods, and computational algorithm development that involve distributed cluster-based implementations, as well as feedback application in a number of biomedical studies. Modern biomedicine now has access to data of rapidly escalating scales and complexities based on innovations and advances in, in particular, genome technologies. This includes high-throughput genomic data from DNA microarray gene expression studies in both laboratory and human observational disease and exposure studies, related large-scale molecular information from proteomic and metabolic profiling technologies, genetic and sequence information that is quickly expanding towards genome-scale sequence data, and, of course, traditional forms of clinical, environmental and demographic data. Advances in both basic biology and the use of such information to inform and aid in human health studies requires very substantial advances in the capacity to analyse and interpret these data sets of ever-increasing scale and complexity. Bringing multiple forms of such data to bear in defining predictive phenotypes is at the core of the emergent arena of clinico-genomics, for example. The challenges to modern mathematical and computational modellers are those of very high-dimensional mathematical model specification and analysis, coupled with the need for computational tools to search across astronomical numbers of candidate models, an endeavor that is well beyond the current capacity to implement, evaluate and understand. This challenge defines the core agenda of this research project: the development of statistical and computational tools for very high-dimensional statistical models. The research is coupled with an integrated collection of genomic studies that provide data, biological collaborators and applications in functional genomics in several areas of cancer and cardiovascular biology, pathway studies in cell cycle regulation and oncogenesis, cancer proteomics, transcription regulation and other areas. At the core of the research lies innovation in methods of stochastic search -- i.e., simulation techniques -- over complex and very high-dimensional statistical model spaces, and the associated model theory and methods. The inherent intellectual merit of the research lies in the advances and innovative new methods in computation and statistical modelling, as well as specific biological applications. The broader impact of the research lies in the applicability of the resulting methods and tools to applications in these specific areas, in other related biomedical/genomic applications, and in other fields of science.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bayesian Models and Methods for Dynamic and Spatio-Dynamic Systems
  • 批准号:
    1106516
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2011
  • 负责人:
    Mike West
  • 依托单位:
Research in Bayesian Analysis: Large-scale Regression and Prediction Models with Applications in Bioinformatics and Applied Time Series
  • 批准号:
    0102227
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $47.5万
  • 财政年份:
    2001
  • 负责人:
    Mike West
  • 依托单位:
Scientific Computing Research Environments for the Mathematical Sciences (SCREMS)
  • 批准号:
    0112340
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.0万
  • 财政年份:
    2001
  • 负责人:
    Mike West
  • 依托单位:
Bayesian Time Series and Dynamic Models
  • 批准号:
    9704432
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $27.2万
  • 财政年份:
    1997
  • 负责人:
    Mike West
  • 依托单位:
海外基金