课题基金 / 基金详情

ATD: Models for (Meta)Genome Identification from Next Generation Sequence Data with Errors

ATD: Models for (Meta)Genome Identification from Next Generation Sequence Data with Errors
ATD:从有错误的下一代序列数据中进行(元)基因组识别的模型
批准号:
1120597
负责人:
Karin Dorman
金额:
$69.82万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2016-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
下一代DNA测序系统实现的高通量和深度覆盖测序正在彻底改变生命科学研究。然而,这种技术产生的测序错误的普遍性阻碍了它产生的数据的全部潜力的实现,特别是在来源不明的混合样本中。拟议项目的总体目标是通过开发数学和统计上严格的模型来解释,组装和分析测序数据沿着计算效率高的算法来解决生物威胁检测中出现的一类下一代测序问题。 研究人员提出了一个数学框架,其中原始序列被观察到有错误,并将实现以下具体目标1)开发一个广义隐马尔可夫模型的测序和有效的算法参数估计。2)为读段开发快速准确的纠错方案,并通过它们之间的相互增强与准确的基因组组装结合。3)开发通过宏基因组分析识别样品混合物中生物威胁的方法。4)开发用于识别生物威胁生物体中相对于参考基因组的基因组变异的方法。深度覆盖提供了冗余,允许沿着置信度评估进行推断,因此可以区分真实的变化和错误,这对于新的生物威胁检测至关重要。新的生物威胁可能首先被检测为包含许多其他遗传物质的患者或环境样本中从未见过的基因组。 我们可能对生物威胁的生物体类型一无所知,但它可能与其他非生物威胁的生物体高度相似。 下一代测序使快速读取样本的遗传内容成为可能,但处理数据的数学模型和算法远远落后。 特别是,下一代测序仪往往会以很高的速率引入错误,产生很难与新生物威胁的真实变异区分开的遗传变异。 PI建议建立严格的统计模型和快速算法,以检测和纠正错误,同时重建样本中包含的基因组。错误校正是重建真实基因组所必需的,也是加速重建算法的需要,并且将在快速产生准确的完整序列中发挥基础性作用。
英文摘要
High-throughput and deep coverage sequencing enabled by next generation DNA sequencing systems is revolutionizing life sciences research. However, the prevalence of sequencing errors produced by such technology hinders a realization of the full potential of the data it produces, particularly in mixed samples of unknown origin. The overarching goal of the proposed project is to address a class of next generation sequencing problems that arise in the context of bio-threat detection by developing mathematically and statistically rigorous models to interpret, assemble and analyze sequencing data along with computationally efficient algorithms to solve them. The researchers propose a mathematical framework where the original sequence is observed with errors, and will accomplish the following specific objectives1) Develop a generalized hidden Markov models for sequencers and efficient algorithms for parameter estimation. 2) Develop fast and accurate error correction schemes for the reads, jointly with accurate genome assembly via mutual reinforcements between them. 3) Develop methods for identifying bio-threats within a sample mixture via metagenomic analysis. 4) Develop methods for identifying genomic variations in a bio-threat organism with respect to a reference genome. Deep coverage provides redundancy that allows inference along with confidence assessment, so true variation can be distinguished from error, which will be essential for novel bio-threat detection.A novel bio-threat will likely first be detected as a never before seen genome in a patient or environmental sample containing much other genetic material. We may know nothing about the type of organism of the bio-threat, yet it may be highly similar to other organisms that are not bio-threats. Next generation sequencing makes it possible to quickly read the genetic content of samples, but the mathematical models and algorithms to process the data lag far behind. In particular, next generation sequencing machines tend to introduce errors at a high rate, producing genetic variation very difficult to distinguish from the true variation of a novel bio-threat. The PIs propose to build rigorous statistical models and rapid algorithms to detect and correct errors while reconstructing the genomes contained in the sample. Error correction is needed to reconstruct the true genomes, but also to speed up the reconstruction algorithms, and will play a fundamental role in quickly producing accurate complete sequences.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
新型手性NAD(P)H Models合成及生化模拟