Beyond heuristics: a tool for the rigorous statistical analysis of *-seq assays.
Beyond heuristics: a tool for the rigorous statistical analysis of *-seq assays.
批准号:
8096347
负责人:
peter J bickel
金额:
$18.51万
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-06-27 至 2013-05-31
关键词:
AlgorithmsAnimalsBiochemicalBiodiversityBiologicalBiological AssayBiological PhenomenaCaenorhabditis elegansCommunitiesComplexComputer softwareDataDeoxyribonucleasesDrug FormulationsEventExonsFoundationsFrequenciesGeneric DrugsGenesGenetic PolymorphismGenetic TranscriptionGenomeGenomicsHuman Cell LineLassoLettersLocationMapsMeasuresMethodsModelingNamesOrganismPathway interactionsProbabilityProcessPropertyProtein IsoformsRNARNA EditingReadingRelative (related person)Research PersonnelSiteSoftware ValidationStatistical ModelsSystemTechniquesTechnologyTestingTranscriptValidationWorkbasedesignflyheuristicshuman tissuemembernext generationprototyperesearch studytheoriestool
中文摘要
描述(由申请人提供):基于下一代测序技术的分析(*-SEQ分析)在基因组学社区中被广泛使用。随着这些分析方法的成熟,并试图探索更微妙的生物学现象,将需要基于强大的统计技术的新工具,以提供对由此产生的生物学结论的信心。到目前为止,*-seq分析工具可以分为两个不同的类别,即测绘和量化。作图工具试图将每个读数与基因组位置相匹配,而量化工具则从“作图”的读数中推断生物学特征。映射的结果通常非常依赖于调优参数,并且很少(如果有的话)提供任何置信度概念。分析工具通常将提供的映射视为福音。这个项目将采取一种不同的方法。研究人员建议使用已知的化验物理和生化特性来模拟化验。这种方法将产生更好的映射,同时提供可作为下游分析组成部分的信心概念。计划使用来自五个不同验证性实验的数据,在三个生物体中对软件和基本模型进行广泛验证。本项目中提议的工作将大大改进对*-SEQ数据的分析。如果成功,该项目将取代一系列测绘算法、峰值呼叫器和记录量词,形成用于综合分析*-seq分析的软件套件的基础。
公共卫生相关性:在基于下一代测序的分析中,映射读数的结果(例如,芯片序列、RNA-序列、DNA酶序列)通常非常依赖于调整参数,而从来没有提供置信度的概念,而下游分析工具通常将提供的映射视为福音。我们的方法是不同的:我们将已知的检测的物理和生化特性以及被检测的特征的生物特性作为映射过程的一个组成部分,然后根据我们的分析模型,对我们的映射设置置信度,然后可以使其成为下游分析、分析或生物学的组成部分。我们的工作原型名为Statmap,我们打算在2012年底之前取代大量映射器、峰值调用器和转录量词,成为计算基因学家武器库中分析和量化的主要工具。
英文摘要
DESCRIPTION (provided by applicant): Assays based upon next generation sequencing technologies (*-seq assays) are widely used in the genomics community. As these assays mature and attempt to probe more subtle biological phenomenon, new tools based upon powerful statistical techniques will be needed to provide confidence in the resulting biological conclusions. To date, *-seq assay analysis tools can be split into two distinct classes, mapping and quantification. Mapping tools attempt to match each read with a genomic location, whereas quantification tools infer biological features from the "mapped" reads. The results of the mapping are often very dependent on tuning parameters and rarely, if ever, provide any notion of confidence. The analysis tools typically take the provided mappings as gospel. This project will take a different approach. The investigators propose to use known physical and biochemical properties of the assay to model the assay. Such an approach will yield better mappings, while providing a notion of confidence that can be made an integral part of downstream analysis. Extensive validation of the software and underlying models is planned in three organisms using data from five different validatory experiments. The work proposed in this project will result in significant improvements in the analyses of *-seq data. If successful, this project will replace a host of mapping algorithms, peak callers, and transcript quantifiers, forming the foundation of a software suite for the integrative analysis of *-seq assays.
PUBLIC HEALTH RELEVANCE: The results of the mapping reads in assays based on next generation sequencing, (e.g. ChIP-seq, RNA-seq, DNase-seq) are often very dependent on tuning parameters without ever providing a notion of confidence, and downstream analysis tools typically take the provided mappings as gospel. Our approach is different: we make the known physical and biochemical properties of the assay and biological properties of the feature assayed an integral part of the mapping process and then on the basis of our assay model, set confidence limits on our mappings that can then be made an integral part of downstream analysis, analytical or biological. Our working prototype is called Statmap, which we intend to replace a host of mappers, peak callers, and transcript quantifiers as the principle tool for analysis and quantification in the computational genomicists arsenal by the end of 2012.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Removing statistical bottle-necks in data analysis for the ENCODE Consortium
-
批准号:8546272
-
项目类别:
-
资助金额:$40.33万
-
财政年份:2012
-
负责人:peter J bickel
-
依托单位:
Removing statistical bottle-necks in data analysis for the ENCODE Consortium
-
批准号:8402497
-
项目类别:
-
资助金额:$42.5万
-
财政年份:2012
-
负责人:peter J bickel
-
依托单位:
Removing statistical bottle-necks in data analysis for the ENCODE Consortium
-
批准号:9037906
-
项目类别:
-
资助金额:$29.81万
-
财政年份:2012
-
负责人:peter J bickel
-
依托单位:
Removing statistical bottle-necks in data analysis for the ENCODE Consortium
-
批准号:8699811
-
项目类别:
-
资助金额:$41.4万
-
财政年份:2012
-
负责人:peter J bickel
-
依托单位:
Beyond heuristics: a tool for the rigorous statistical analysis of *-seq assays.
-
批准号:8290222
-
项目类别:
-
资助金额:$22.28万
-
财政年份:2011
-
负责人:peter J bickel
-
依托单位:
Travel Support for High Dimensional Statistics in Biology
-
批准号:7485843
-
项目类别:
-
资助金额:$1.5万
-
财政年份:2008
-
负责人:peter J bickel
-
依托单位:
Comparative Genomics to Identify Functional Blocks & HGT
-
批准号:7418308
-
项目类别:
-
资助金额:$16.22万
-
财政年份:2005
-
负责人:peter J bickel
-
依托单位:
Comparative Genomics to Identify Functional Blocks & HGT
-
批准号:7064837
-
项目类别:
-
资助金额:$17.03万
-
财政年份:2005
-
负责人:peter J bickel
-
依托单位:
Comparative Genomics to Identify Functional Blocks & HGT
-
批准号:7240439
-
项目类别:
-
资助金额:$16.38万
-
财政年份:2005
-
负责人:peter J bickel
-
依托单位:
Comparative Genomics to Identify Functional Blocks & HGT
-
批准号:7498626
-
项目类别:
-
资助金额:$11.48万
-
财政年份:2005
-
负责人:peter J bickel
-
依托单位:
Comparative Genomics to Identify Functional Blocks & HGT
-
批准号:6985664
-
项目类别:
-
资助金额:$17.58万
-
财政年份:2005
-
负责人:peter J bickel
-
依托单位:
Determining the molecular forces that target transcription factors to DNA in vivo
-
批准号:8262273
-
项目类别:
-
资助金额:$20.46万
-
财政年份:--
-
负责人:peter J bickel
-
依托单位:
海外基金