Evaluating the necessity of PCR duplicate removal from next-generation sequencing data and a comparison of approaches.

Evaluating the necessity of PCR duplicate removal from next-generation sequencing data and a comparison of approaches.
复制标题

DOI:
10.1186/s12859-016-1097-3
复制
发表时间:
2016-07-25
期刊:
影响因子:
3
通讯作者:
Ridge PG
Ridge PG
中科院分区:
生物学4区
文献类型:
--
作者:
Ebbert MT;Wadsworth ME;Staley LA;Hoyt KL;Pickett B;Miller J;Duce J;Alzheimer’s Disease Neuroimaging Initiative;Kauwe JS;Ridge PG

文献摘要

被引文献

相似文献

分析下一代测序数据很困难,因为数据集很大,第二代测序平台有很高的错误率,而且因为目标基因组中的每个位置(外显子组、转录组等)。被多次测序。鉴于这些挑战,已经开发了许多生物信息学算法来分析这些数据。这些算法旨在找到数据丢失、错误、分析时间和内存占用之间的适当平衡。典型的分析管道需要多个步骤。如果这些步骤中的一个或多个是不必要的,则删除该步骤将显著减少计算时间和数据操作。在许多管道中的一个步骤是PCR重复去除,其中,来自同一模板分子结合在流式细胞上的多个PCR产物产生了PCR重复。它们经常被删除,因为有人担心它们可能会导致误报变量调用。Picard(MarkDuplates)和SamTools(Rmdup)是两个主要的用于PCR重复删除的软件。在被调用的1700多万个变种中,大约92%是被调用的,无论我们是使用Picard或SamTools删除重复项,还是将PCR复制项保留在数据集中。比较转换/颠换比率(p = 1.0)、新变异体的百分比(p = 0.99)、平均群体频率(p = 0.99)和蛋白质变化变异体的百分比(p = 1.0),两组唯一变异体之间没有显著差异。美国医学遗传学学院基因变种的研究结果与此类似。NGS和SNP芯片之间的基因一致性在所有基因型组中均在99%以上(例如,纯合子参考)。我们的结果表明,PCR重复删除对后续变体调用的准确性影响很小。
Analyzing next-generation sequencing data is difficult because datasets are large, second generation sequencing platforms have high error rates, and because each position in the target genome (exome, transcriptome, etc.) is sequenced multiple times. Given these challenges, numerous bioinformatic algorithms have been developed to analyze these data. These algorithms aim to find an appropriate balance between data loss, errors, analysis time, and memory footprint. Typical analysis pipelines require multiple steps. If one or more of these steps is unnecessary, it would significantly decrease compute time and data manipulation to remove the step. One step in many pipelines is PCR duplicate removal, where PCR duplicates arise from multiple PCR products from the same template molecule binding on the flowcell. These are often removed because there is concern they can lead to false positive variant calls. Picard (MarkDuplicates) and SAMTools (rmdup) are the two main softwares used for PCR duplicate removal. Approximately 92 % of the 17+ million variants called were called whether we removed duplicates with Picard or SAMTools, or left the PCR duplicates in the dataset. There were no significant differences between the unique variant sets when comparing the transition/transversion ratios (p = 1.0), percentage of novel variants (p = 0.99), average population frequencies (p = 0.99), and the percentage of protein-changing variants (p = 1.0). Results were similar for variants in the American College of Medical Genetics genes. Genotype concordance between NGS and SNP chips was above 99 % for all genotype groups (e.g., homozygous reference). Our results suggest that PCR duplicate removal has minimal effect on the accuracy of subsequent variant calls.