Empirical estimation of sequencing error rates using smoothing splines.

Empirical estimation of sequencing error rates using smoothing splines.
复制标题

DOI:
10.1186/s12859-016-1052-3
复制
发表时间:
2016-04-22
期刊:
影响因子:
3
通讯作者:
Shete S
Shete S
中科院分区:
生物学4区
文献类型:
--
作者:
Zhu X;Wang J;Peng B;Shete S

文献摘要

参考文献

相似文献

下一代测序已经被研究人员用来解决一系列不同的生物学问题,例如,通过多态和突变发现以及microRNA图谱。然而,与传统测序相比,下一代测序的错误率往往更高,这影响了下游的基因组分析。最近,Wang等人提出了自己的观点。(BMC BioInformation 13:185,2012)提出了一种影子回归方法来估计下一代测序数据的错误率,该方法基于测序的读数和包含错误的读数之间的线性关系的假设(表示为阴影)。然而,这种线性读-影关系可能不适用于所有类型的序列数据。因此,有必要在不假设线性的情况下以更可靠的方式估计误码率。我们提出了一种经验的错误率估计方法,该方法使用三次和稳健的平滑样条线来建模排序的读取次数和阴影数目之间的关系。我们使用基于频率的方法进行了仿真研究,直接生成读取计数和阴影计数,这可以模拟真实的顺序计数数据结构。通过仿真,研究了该方法的性能,并与影子线性回归方法进行了比较。对于所有测试的场景,该方法提供了比影子线性回归方法更准确的错误率估计。我们还应用所提出的方法评估了来自微阵列质量控制项目、突变筛选研究、DNA元素百科全书项目和噬菌体PhiX DNA样本的序列数据的错误率。所提出的经验错误率估计方法不假设无错误读取计数和阴影计数之间的线性关系,并且为下一代短读测序数据提供了更准确的错误率估计。本文的在线版本(doi:10.1186/s12859-0161052-3)包含补充材料,授权用户可以使用。
Next-generation sequencing has been used by investigators to address a diverse range of biological problems through, for example, polymorphism and mutation discovery and microRNA profiling. However, compared to conventional sequencing, the error rates for next-generation sequencing are often higher, which impacts the downstream genomic analysis. Recently, Wang et al. (BMC Bioinformatics 13:185, 2012) proposed a shadow regression approach to estimate the error rates for next-generation sequencing data based on the assumption of a linear relationship between the number of reads sequenced and the number of reads containing errors (denoted as shadows). However, this linear read-shadow relationship may not be appropriate for all types of sequence data. Therefore, it is necessary to estimate the error rates in a more reliable way without assuming linearity. We proposed an empirical error rate estimation approach that employs cubic and robust smoothing splines to model the relationship between the number of reads sequenced and the number of shadows. We performed simulation studies using a frequency-based approach to generate the read and shadow counts directly, which can mimic the real sequence counts data structure. Using simulation, we investigated the performance of the proposed approach and compared it to that of shadow linear regression. The proposed approach provided more accurate error rate estimations than the shadow linear regression approach for all the scenarios tested. We also applied the proposed approach to assess the error rates for the sequence data from the MicroArray Quality Control project, a mutation screening study, the Encyclopedia of DNA Elements project, and bacteriophage PhiX DNA samples. The proposed empirical error rate estimation approach does not assume a linear relationship between the error-free read and shadow counts and provides more accurate estimations of error rates for next-generation, short-read sequencing data. The online version of this article (doi:10.1186/s12859-016-1052-3) contains supplementary material, which is available to authorized users.
DOI: 10.1371/journal.pone.0012681
发表时间: 2010-09-22
期刊: PloS one
影响因子: 3.7
作者:
Schröder J;Bailey J;Conway T;Zobel J
通讯作者: Zobel J