Efficient error correction for next-generation sequencing of viral amplicons.

Efficient error correction for next-generation sequencing of viral amplicons.
复制标题

DOI:
10.1186/1471-2105-13-s10-s6
复制
发表时间:
2012-06-25
期刊:
影响因子:
3
通讯作者:
Khudyakov Y
Khudyakov Y
中科院分区:
生物学4区
文献类型:
--
作者:
Skums P;Dimitrova Z;Campo DS;Vaughan G;Rossi L;Forbi JC;Yokosawa J;Zelikovsky A;Khudyakov Y

文献摘要

被引文献

相似文献

新一代测序技术允许分析感染患者中数量空前的病毒序列变异,为了解病毒进化、耐药性和免疫逃逸提供了新的机会。然而,批量测序容易出错。因此,生成的数据需要错误识别和校正。迄今为止,大多数纠错方法都没有针对扩增子分析进行优化,并且假设错误率是随机分布的。使用454-测序获得的扩增子序列的最近质量评估表明,错误率与均聚物的存在和大小、序列中的位置和扩增子的长度密切相关。所有这些参数都具有很强的序列特异性,应纳入为扩增子测序设计的纠错算法的校准中。在本文中,我们提出了两个新的有效的错误校正算法优化的病毒扩增子:(i)基于k聚体的错误校正(KEC)和(ii)经验频率阈值(ET)。将两者与之前发布的聚类算法(SHORAH)进行比较,以评估它们在通过对具有已知序列的扩增子进行454次测序获得的24个实验数据集上的相对性能。所有这三种算法在寻找真正的单倍型方面显示出相似的准确性。然而,KEC和ET在去除假单倍型和估计真单倍型的频率方面比SHORAH更有效。这两种算法,KEC和ET,是非常适合于快速恢复的454-测序从异质性病毒的扩增子获得的无错误的单倍型。用于其测试的算法和数据集的实现可在http://alan.cs.gsu.edu/NGS/?上获得q=含量/焦磷酸测序误差校正算法
Next-generation sequencing allows the analysis of an unprecedented number of viral sequence variants from infected patients, presenting a novel opportunity for understanding virus evolution, drug resistance and immune escape. However, sequencing in bulk is error prone. Thus, the generated data require error identification and correction. Most error-correction methods to date are not optimized for amplicon analysis and assume that the error rate is randomly distributed. Recent quality assessment of amplicon sequences obtained using 454-sequencing showed that the error rate is strongly linked to the presence and size of homopolymers, position in the sequence and length of the amplicon. All these parameters are strongly sequence specific and should be incorporated into the calibration of error-correction algorithms designed for amplicon sequencing. In this paper, we present two new efficient error correction algorithms optimized for viral amplicons: (i) k-mer-based error correction (KEC) and (ii) empirical frequency threshold (ET). Both were compared to a previously published clustering algorithm (SHORAH), in order to evaluate their relative performance on 24 experimental datasets obtained by 454-sequencing of amplicons with known sequences. All three algorithms show similar accuracy in finding true haplotypes. However, KEC and ET were significantly more efficient than SHORAH in removing false haplotypes and estimating the frequency of true ones. Both algorithms, KEC and ET, are highly suitable for rapid recovery of error-free haplotypes obtained by 454-sequencing of amplicons from heterogeneous viruses. The implementations of the algorithms and data sets used for their testing are available at: http://alan.cs.gsu.edu/NGS/?q=content/pyrosequencing-error-correction-algorithm