Algorithms for Optimal Base-Calling in Sequencing-by-Synthesis
Algorithms for Optimal Base-Calling in Sequencing-by-Synthesis
批准号:
8095652
负责人:
Haris Vikalo
金额:
$17.83万
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2013-07-31
关键词:
Acute DiseaseAddressAlgorithmsAreaBiochemical ProcessBiological SciencesChronic DiseaseCodeColorCommunicationComplementComplexDNADNA SequenceDNA biosynthesisDataDetectionDideoxy Chain Termination DNA SequencingDoctor of PhilosophyFailureGeneticGoalsIndividualInferiorInformation TheoryLaboratoriesLengthMedicalMedicineModelingNatureNucleotidesPeptide Signal SequencesPerformancePersonsPharmacologic SubstancePhasePredispositionPublicationsReadingResearchResolutionSchemeSequence AnalysisSignal TransductionSolutionsSourceSpeedSystemTechniquesTechnologyTestingTrainingaustinbasecomputer studiescomputerized data processingcostdesigngenome-widehealth care deliveryimprovedmathematical modelnext generationprematureprogramstool
中文摘要
描述(由申请人提供):下一代合成测序平台能够实现快速和负担得起的DNA测序。然而,它们实现的读取长度仍然短于昂贵的桑格测序提供的读取长度,并且它们的准确性不足以用于大多数医学研究。为了确定DNA片段中核苷酸的顺序,合成测序依赖于片段上互补链的酶促合成。通过连续添加游离核苷酸实现合成;光学检测DNA片段的第一个未配对碱基的沃森-克里克互补链的互补链延伸。然而,通过对单个DNA分子进行测序产生的信号很弱,因此其检测需要复杂且昂贵的硬件。基于集成的系统提供了一种有效的替代方案:它们通过并行测序大量相同拷贝的DNA片段来放大信号。为了充分利用多个信号源的好处,互补链的延伸应该以相同的速率进行(以便信号在相位上相加)。然而,由于在一些链中核苷酸掺入的偶然失败和其他链的过早延伸,整体中链的合成变得不同步。这些所谓的定相效应,本质上是概率性的,限制了合成测序的可实现的准确度和读取长度。 该项目的目标是开发合成测序系统中最佳碱基识别的实用算法,提高其有效读长和准确性。为此,我们依赖于信号处理和信息论的概念和工具。我们解决了两个广泛使用的系统:Illumina的四色平台和罗氏(454生命科学)焦磷酸测序平台。如果成功,正如我们基于初步结果所预期的那样,我们的研究将对需要高性能DNA测序的各种应用产生直接影响。
公共卫生相关性:下一代DNA测序的性能从根本上受到潜在生物化学过程的随机性的限制。借鉴信号处理和信息论的概念,我们建议设计实用的算法,这些算法可以显着提高下一代DNA测序系统的准确性和有效读取长度。如果成功,正如我们基于初步结果所预期的那样,我们的研究将对需要高性能DNA测序的各种应用产生直接影响。
英文摘要
DESCRIPTION (provided by applicant): Next generation sequencing-by-synthesis platforms enable fast and affordable DNA sequencing. However, read-lengths that they achieve are still shorter than those provided by the costly Sanger sequencing, and their accuracy is insufficient for most medical studies. To determine the order of nucleotides in a DNA fragment, sequencing-by-synthesis relies on enzymatic synthesis of the complementary strand on the fragment. The synthesis is enabled by a sequential addition of free nucleotides; extension of the complementary strand with the Watson-Crick complement of the first unpaired base of the DNA fragment is detected optically. However, the signal generated by sequencing a single DNA molecule is weak, and thus its detection requires complex and expensive hardware. Ensemble-based systems provide an efficient alternative: they amplify the signal by sequencing a large number of identical copies of the DNA fragment in parallel. To fully reap the benefits of having multiple signal sources, extension of complementary strands should progress at the same rate (so that the signals add in phase). However, synthesis of strands in an ensemble gets out-of-sync due to an occasional failure of nucleotide incorporation in some strands, and premature extension of others. These so-called phasing effects, probabilistic in nature, limit the achievable accuracy and read-lengths of sequencing-by-synthesis. The goal of the proposed project is to develop practical algorithms for optimal base-calling in sequencing-by-synthesis systems, improving their effective read-lengths and accuracy. To this end, we rely on concepts and tools from signal processing and information theory. We address two broadly employed systems: Illumina's four-color platform and Roche's (454 Life Sciences) pyrosequencing platform. If successful, as we expect based on preliminary results, our research will have immediate impact on various applications which require high-performance DNA sequencing.
PUBLIC HEALTH RELEVANCE: Performance of next generation DNA sequencing is fundamentally limited by the stochastic nature of the underlying biochemical process. Drawing on concepts from signal processing and information theory, we propose to design practical algorithms which may significantly improve the accuracy and effective read-lengths of next generation DNA sequencing systems. If successful, as we expect based on preliminary results, our research will have immediate impact on various applications which require high-performance DNA sequencing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Algorithms for Optimal Base-Calling in Sequencing-by-Synthesis
-
批准号:8288688
-
项目类别:
-
资助金额:$17.79万
-
财政年份:2011
-
负责人:Haris Vikalo
-
依托单位:
海外基金