课题基金 / 基金详情

项目摘要

项目成果

James Bentley Brown的其他基金

相似基金

相关文献

中文摘要
翻译
描述(申请人提供):下一代测序以前所未有的细节揭示了细胞的分子图景。然而,对于基于这些技术的分析产生的大规模数据,信息量不仅是湿实验室技术的功能,而且关键是解释这些数据的分析管道的功能。我们团队开发了四种统计工具,旨在最大限度地提高这些分析的信息量:1)基因组结构校正(GSC),用于评估特征之间关系的重要性的基因组注释的非参数模型;2)不可复制发现率(IDR),利用生物复制的信息与FDR的类似物;3)STATMAP,芯片序列和CAGE数据的综合分析管道,将统计置信度从碱基调用传播到峰值调用;以及4)稀疏线性异构体发现和丰度估计(SLICE),这是一个综合的统计框架,用于分析RNA-seq、cDNA和其他RNA数据,旨在获得和量化从头转录模型。这些工具旨在识别和描述基因组中的功能元素;它们对所分析的数据做出最小的假设,因此得出可靠的结论和统计可信度的衡量标准。在K99期间,我们将扩展和整合我们的工具,以将统计置信度的范围延伸到整个数据解释。在R00期间,我的研究将向生物网络的推断和评估迈进。正如同源基因鉴定已成为发展人类疾病动物模型的关键步骤一样,多物种网络分析有望成为解释基因组变异和表型之间关系的关键步骤。许多突变,甚至基因缺失,都不会显露出明显的表型。这是由于网络的健壮性,这在密切相关的物种之间往往是不同的。为了理解这些现象,我们的目标是:1)开发用于网络推理的标准统计工具,以及2)开发网络的“元模型”,以允许对网络正则学的一般测量。这两个目标是紧密相连的:我们需要严格地描述生物网络的语义,以对它们进行建模。目前,一些模型缺乏对边和权重的一致定义,导致基因组数据的表示无法测试。我们将开发可测试的生物过程的量化模型,利用复杂系统的丰富理论建立统一的语义。上述每个工具都将发挥关键作用,特别是将统计置信度传播到网络分析中所需的Statmap和GSC。进步将对我们将疾病的动物模型映射到人类生物学上的能力产生革命性的影响。由于动物模型中不存在的问题(如毒性),近九成的新药在人体试验中失败。不仅要了解单个基因的正交学,还要了解整个生化网络的正交学,这对于推断和纠正疾病模型和人类生物学之间的差异至关重要。这一问题的解决,将是从碱基对向床边进军的重要一步。
英文摘要
DESCRIPTION (provided by applicant): Next generation sequencing has revealed the molecular landscape of cells in unprecedented detail. However, for the massively large-scale data produced by assays based on these technologies, informativeness is not only a function of wet-lab technology, but is critically also a function of the analytical pipelines that interpret th data. Our group has developed four statistical tools designed maximize the informativeness of these assays: 1) the Genome Structural Correction (GSC), a nonparametric model of genomic annotations used to assess the significance of relationships between features; 2) the Irreproducible Discovery Rate (IDR), an analogue of the FDR that leverages information from biological replicates; 3) Statmap, a comprehensive analysis pipeline for ChIP-seq and CAGE data that propagates statistical confidence from base-calling to peak-calling; and 4) Sparse Linear Isoform Discovery and abundance Estimation (SLIDE), an integrative statistical framework for the analysis of RNA-seq, cDNA, and other RNA data aimed at obtaining and quantifying de novo transcript models. These tools are designed to identify and characterize functional elements in genomes; they make minimal assumptions about the data they analyze, and therefore draw reliable conclusions and measures of statistical confidence. During the K99, we will expand and integrate our tools to extend the reach of statistical confidence throughout data interpretation. During the R00, my research will progress toward the inference and assessment of biological networks. Just as ortholog identification has become an essential step in developing animal models of human disease, multi-species network analysis promises to become a key step in interpreting the relationship between genome variation and phenotype. Many mutations, even gene deletions, do not reveal an obvious phenotype. This is due to network robustness, which often differs between closely related species. To understand these phenomena, we aim to: 1) develop standard statistical tools for network inference, and 2) develop "meta models" of networks that will permit general measures of network orthology. These two aims are tightly linked: we will need critically to characterize the semantics of biological networks to model them. Currently, some models lack consistent definitions of edges and weights, resulting in untestable representations of genomics data. We will develop testable, quantitative models of biological processes, establishing a uniform semantics leveraging the rich theory of complex systems. Each of the tools above will play a key role, especially Statmap and the GSC, which will be needed to propagate statistical confidence into network analysis. Advances will have a transformative effect on our ability to map animal models of disease onto human biology. Nearly nine out of ten new drugs fail in human trials due to issues (e.g. toxicity) not present in animal models. Understanding the orthology not just of individual genes, but of entire biochemical networks will be essential to infer and correct for differences between models of disease and human biology. Solving this problem will be a major step forward in the march from "base-pairs to bedside".
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1101/gr.132159.111
发表时间: 2012-09
期刊: Genome research
影响因子: 7
作者: [Derrien T, Johnson R, Bussotti G, Tanzer A, Djebali S, Tilgner H, Guernec G, Martin D, Merkel A, Knowles DG, Lagarde J, Veeravalli L, Ruan X, Ruan Y, Lassmann T, Carninci P, Brown JB, Lipovich L, Gonzalez JM, Thomas M, Davis CA, Shiekhattar R, Gingeras TR, Hubbard TJ, Notredame C, Harrow J, Guigó R]
通讯作者: Guigó R
DOI: 10.1186/gb-2012-13-9-r48
发表时间: 2012-09-26
期刊: Genome biology
影响因子: 12.3
作者: [Yip KY, Cheng C, Bhardwaj N, Brown JB, Leng J, Kundaje A, Rozowsky J, Birney E, Bickel P, Snyder M, Gerstein M]
通讯作者: Gerstein M
DOI: 10.1101/gr.134767.111
发表时间: 2012-09
期刊: Genome research
影响因子: 7
作者: [Bánfai B, Jia H, Khatun J, Wood E, Risk B, Gundling WE Jr, Kundaje A, Gunawardena HP, Yu Y, Xie L, Krajewski K, Strahl BD, Chen X, Bickel P, Giddings MC, Brown JB, Lipovich L]
通讯作者: Lipovich L
DOI: 10.1101/gr.136184.111
发表时间: 2012-09
期刊: Genome research
影响因子: 7
作者: [Landt SG, Marinov GK, Kundaje A, Kheradpour P, Pauli F, Batzoglou S, Bernstein BE, Bickel P, Brown JB, Cayting P, Chen Y, DeSalvo G, Epstein C, Fisher-Aylor KI, Euskirchen G, Gerstein M, Gertz J, Hartemink AJ, Hoffman MM, Iyer VR, Jung YL, Karmakar S, Kellis M, Kharchenko PV, Li Q, Liu T, Liu XS, Ma L, Milosavljevic A, Myers RM, Park PJ, Pazin MJ, Perry MD, Raha D, Reddy TE, Rozowsky J, Shoresh N, Sidow A, Slattery M, Stamatoyannopoulos JA, Tolstorukov MY, White KP, Xi S, Farnham PJ, Lieb JD, Wold BJ, Snyder M]
通讯作者: Snyder M
Nonparametric methods for functional and translational genomics
Nonparametric methods for functional and translational genomics
  • 批准号:
    8280729
  • 项目类别:
  • 资助金额:
    $10.3万
  • 财政年份:
    2012
  • 负责人:
    James Bentley Brown
  • 依托单位:
海外基金