Automated correction of genome sequence errors

Automated correction of genome sequence errors
复制标题

DOI:
10.1093/nar/gkh216
复制
发表时间:
2004-01-01
影响因子:
14.9
通讯作者:
Salzberg, SL
Salzberg, SL
中科院分区:
生物学2区
文献类型:
--
作者:
Gajer, P;Schatz, M;Salzberg, SL

文献摘要

被引文献

相似文献

通过使用来自基因组组装的信息,一个名为AutoEditor的新程序显着提高了碱基识别的准确性,而不是以前的算法。这反过来又提高了基因组序列的整体准确性,并促进了这些序列用于多态性发现。我们描述的算法及其应用在一个大的最近的基因组测序项目。这些项目中错误碱基识别的数量减少了80%。在对超过100万次更正的分析中,我们发现AutoEditor在8828次更正中只犯了一个错误。通过大幅提高碱基识别的准确性,AutoEditor可以大大加快完成基因组的过程,这涉及到关闭所有缺口并确保最终序列的最低质量标准。它还大大提高了我们发现密切相关的菌株和同一物种的分离株之间的单核苷酸多态性(SNP)的能力。
By using information from an assembly of a genome, a new program called AutoEditor significantly improves base calling accuracy over that achieved by previous algorithms. This in turn improves the overall accuracy of genome sequences and facilitates the use of these sequences for polymorphism discovery. We describe the algorithm and its application in a large set of recent genome sequencing projects. The number of erroneous base calls in these projects was reduced by 80%. In an analysis of over one million corrections, we found that AutoEditor made just one error per 8828 corrections. By substantially increasing the accuracy of base calling, AutoEditor can dramatically accelerate the process of finishing genomes, which involves closing all gaps and ensuring minimum quality standards for the final sequence. It also greatly improves our ability to discover single nucleotide polymorphisms (SNPs) between closely related strains and isolates of the same species.