EDAR: An Efficient Error Detection and Removal Algorithm for Next Generation Sequencing Data

EDAR: An Efficient Error Detection and Removal Algorithm for Next Generation Sequencing Data
复制标题

DOI:
10.1089/cmb.2010.0127
复制
发表时间:
2010-11-01
影响因子:
1.7
通讯作者:
Wittenberg, Gayle M.
Wittenberg, Gayle M.
中科院分区:
生物学4区
文献类型:
--
作者:
Zhao, Xiaohong;Palmer, Lance E.;Wittenberg, Gayle M.

文献摘要

被引文献

相似文献

基因组测序技术将实验误差引入到reads中,这会误导序列组装工作并使诊断过程复杂化。在这里,我们提出了一种方法来检测和去除序列组装之前基因组霰弹枪测序项目中产生的测序错误。对于每个读取的输入,计算其包含的所有长度为k的子字符串(k-mers)的集合。根据每个k-mer在完整数据集中出现的频率(k-count)来评估读取。对于每次读取,使用可变带宽mean-shift算法对k-mers进行聚类。根据聚类中心的k-count,将聚类分为错误区域和非错误区域。对于测试的23个真实和模拟数据集(454和Solexa),我们的算法检测到的错误区域覆盖了所有错误的99%。然后应用启发式算法检测每个假定误差区域中的误差位置。通过去除错误来纠正读取,从而创建两个或多个更小的无错误读取片段。在执行错误删除后,所有测试数据集的错误率都下降了(平均减少了35倍)。EDAR具有与纠正而不是消除错误的方法相当的准确性,并且当模拟数据集的错误率大于3%时,它的性能更好。对于去除错误的数据,Velvet汇编器的性能通常会更好。但是,对于短读,在错误位置进行分割可能会有问题。在错误检测之后进行错误纠正,而不是去除,可以改善装配结果。
Genomic sequencing techniques introduce experimental errors into reads which can mislead sequence assembly efforts and complicate the diagnostic process. Here we present a method for detecting and removing sequencing errors from reads generated in genomic shotgun sequencing projects prior to sequence assembly. For each input read, the set of all length k substrings (k-mers) it contains are calculated. The read is evaluated based on the frequency with which each k-mer occurs in the complete data set (k-count). For each read, k-mers are clustered using the variable-bandwidth mean-shift algorithm. Based on the k-count of the cluster center, clusters are classified as error regions or non-error regions. For the 23 real and simulated data sets tested (454 and Solexa), our algorithm detected error regions that cover 99% of all errors. A heuristic algorithm is then applied to detect the location of errors in each putative error region. A read is corrected by removing the errors, thereby creating two or more smaller, error-free read fragments. After performing error removal, the error-rate for all data sets tested decreased (similar to 35-fold reduction, on average). EDAR has comparable accuracy to methods that correct rather than remove errors and when the error rate is greater than 3% for simulated data sets, it performs better. The performance of the Velvet assembler is generally better with error-removed data. However, for short reads, splitting at the location of errors can be problematic. Following error detection with error correction, rather than removal, may improve the assembly results.