Nonparametric methods for functional and translational genomics
Nonparametric methods for functional and translational genomics
批准号:
8916814
负责人:
James Bentley Brown
金额:
$24.9万
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-08-25 至 2017-05-31
关键词:
AlgorithmsAnimal Disease ModelsAnimal ModelAreaAutomobile DrivingAwardBase PairingBiochemicalBiologicalBiological AssayBiological ModelsBiological ProcessCellsChIP-seqCommunicationComplementary DNAComplexDataData AnalysesData SourcesDevelopmentDevelopmental BiologyDisease modelElementsGap JunctionsGene DeletionGenesGenomeGenomicsGoalsHigh-Throughput Nucleotide SequencingHumanHuman BiologyIndiumIndividualLeadLinkMapsMeasuresMentorsMethodsModelingMolecularMutationOrphanOrthologous GenePathway AnalysisPharmaceutical PreparationsPhenotypePlayProblem SolvingPropertyProtein IsoformsRNAReadingResearchResearch PersonnelRunningSemanticsSystemTechniquesTechnologyToxic effectTrainingTraining ActivityTranscriptTranscriptional RegulationVariantWeightabstractinganalogbasecareer developmentdesigndriving forceexperiencefunctional genomicshigh throughput screeninghuman diseasenetwork modelsnext generation sequencingnovel strategiesstatisticsstem cell biologytheoriestooltranscription factortranscriptome sequencing
中文摘要
项目摘要/摘要
下一代测序以前所未有的细节揭示了细胞的分子图景。然而,
对于基于这些技术的分析产生的大规模数据,信息性不是
只是湿法实验室技术的功能,但关键是也是分析管道的功能,这些分析管道解释了
数据。我们小组开发了四种统计工具,旨在最大限度地提高这些分析的信息量:
1)基因组结构校正(GSC),一种用于评估
特征之间关系的重要性;2)不可复制发现率(IDR),类似于
利用生物复制信息的FDR;3)STATMAP,一个全面的分析渠道
用于将统计置信度从基地呼叫传播到峰值呼叫的CHIP-SEQ和CAGE数据;以及4)
稀疏线性异构体发现和丰度估计(幻灯片),一个综合的统计框架
对rna-seq、cdna和其他rna数据的分析,旨在获得和量化从头转录本。
模特们。这些工具旨在识别和表征基因组中的功能元件;它们使
对他们分析的数据做出最小的假设,从而得出可靠的结论和衡量标准
统计上的可信度。在K99期间,我们将扩展和集成我们的工具,以扩大统计的范围
在数据解释过程中始终充满信心。在R00期间,我的研究将朝着推理和
生物网络的评估。就像直系同源基因鉴定已经成为开发
人类疾病的动物模型,多物种网络分析有望成为
解释基因组变异与表型之间的关系。许多突变,甚至基因缺失,
不要显露出明显的表型。这是由于网络健壮性,两者之间往往存在密切的差异
近缘物种。为了理解这些现象,我们的目标是:1)开发标准的网络统计工具
推论,以及2)开发网络的“元模型”,将允许网络正构学的一般措施。
这两个目标紧密相连:我们需要严格地描述生物网络的语义,以
给他们做模特。目前,一些模型缺乏对边和权重的一致定义,导致无法测试
基因组学数据的表示。你终于度过了一个轻松的周末!我们将开发可测试的、
生物过程的量化模型,建立统一的语义,利用丰富的
复杂的系统。上面的每个工具都将发挥关键作用,特别是Statmap和GSC,这将是
需要将统计置信度传播到网络分析中。进步将产生变革性的影响
关于我们将疾病的动物模型映射到人类生物学上的能力。近十分之九的新药未能通过
由于动物模型中不存在的问题(例如毒性)而进行的人体试验。了解矫正学,而不仅仅是
单个基因,但整个生化网络的差异将是必不可少的推断和纠正差异
疾病模型和人类生物学之间的关系。解决这个问题将是前进的一大步。
从“碱基对到床边”。
英文摘要
Project Summary / Abstract
Next generation sequencing has revealed the molecular landscape of cells in unprecedented detail. However,
for the massively large-scale data produced by assays based on these technologies, informativeness is not
only a function of wet-lab technology, but is critically also a function of the analytical pipelines that interpret the
data. Our group has developed four statistical tools designed maximize the informativeness of these assays:
1) the Genome Structural Correction (GSC), a nonparametric model of genomic annotations used to assess
the significance of relationships between features; 2) the Irreproducible Discovery Rate (IDR), an analogue of
the FDR that leverages information from biological replicates; 3) Statmap, a comprehensive analysis pipeline
for ChIP-seq and CAGE data that propagates statistical confidence from base-calling to peak-calling; and 4)
Sparse Linear Isoform Discovery and abundance Estimation (SLIDE), an integrative statistical framework for
the analysis of RNA-seq, cDNA, and other RNA data aimed at obtaining and quantifying de novo transcript
models. These tools are designed to identify and characterize functional elements in genomes; they make
minimal assumptions about the data they analyze, and therefore draw reliable conclusions and measures of
statistical confidence. During the K99, we will expand and integrate our tools to extend the reach of statistical
confidence throughout data interpretatoin. During the R00, my research will progress toward the inference and
assessment of biological networks. Just as ortholog identification has become an essential step in developing
animal models of human disease, multi-species network analysis promises to become a key step in
interpreting the relationship between genome variation and phenotype. Many mutations, even gene deletions,
do not reveal an obvious phenotype. This is due to network robustness, which often differs between closely
related species. To understand these phenomena, we aim to: 1) develop standard statistical tools for network
inference, and 2) develop "meta models" of networks that will permit general measures of network orthology.
These two aims are tightly linked: we will need critically to characterize the semantics of biological networks to
model them. Currently, some models lack consistent definitions of edges and weights, resulting in untestable
representations of genomics data. you've managed to have a relaxing weekend! We will develop testable,
quantitative models of biological processes, establishing a uniform semantics leveraging the rich theory of
complex systems. Each of the tools above will play a key role, especially Statmap and the GSC, which will be
needed to propagate statistical confidence into network analysis. Advances will have a transformative effect
on our ability to map animal models of disease onto human biology. Nearly nine out of ten new drugs fail in
human trials due to issues (e.g. toxicity) not present in animal models. Understanding the orthology not just of
individual genes, but of entire biochemical networks will be essential to infer and correct for differences
between models of disease and human biology. Solving this problem will be a major step forward in the march
from “base-pairs to bedside”.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Nonparametric methods for functional and translational genomics
-
批准号:8280729
-
项目类别:
-
资助金额:$10.3万
-
财政年份:2012
-
负责人:James Bentley Brown
-
依托单位:
Nonparametric methods for functional and translational genomics
-
批准号:8532014
-
项目类别:
-
资助金额:$10.3万
-
财政年份:2012
-
负责人:James Bentley Brown
-
依托单位:
海外基金