An investigation of causes of false positive single nucleotide polymorphisms using simulated reads from a small eukaryote genome.

An investigation of causes of false positive single nucleotide polymorphisms using simulated reads from a small eukaryote genome.
复制标题

DOI:
10.1186/s12859-015-0801-z
复制
发表时间:
2015-11-11
期刊:
影响因子:
3
通讯作者:
Bayer M
Bayer M
中科院分区:
生物学4区
文献类型:
--
作者:
Ribeiro A;Golicz A;Hackett CA;Milne I;Stephen G;Marshall D;Flavell AJ;Bayer M

文献摘要

被引文献

相似文献

单核苷酸多态性 (SNP) 是广泛使用的分子标记,自新一代测序 (NGS) 技术诞生以来,其使用量大幅增加,该技术允许以低成本检测大量 SNP。然而,NGS 数据及其分析都容易出错,这可能导致假阳性 (FP) SNP 的生成。我们探讨了 FP SNP 与基于映射的变体调用所涉及的七个因素之间的关系——参考序列的质量、读段长度、映射器和变体调用者的选择、映射严格性以及通过读段映射质量和读段深度过滤 SNP。这产生了 576 种可能的因子水平组合。我们使用无错误和无变异的模拟读取来确保发现的每个 SNP 确实是误报。对于 1.2 亿碱基对 (Mbp) 基因组,生成的 FP SNP 数量变化范围为 0 到 36,621。所有测试的实验因素对生成的 FP SNP 数量都有统计学上的显着影响,并且不同因素之间存在大量的相互作用。使用片段化参考序列导致生成的 FP SNP 数量急剧增加,宽松的读映射和缺乏 SNP 过滤也是如此。参考汇编器、映射器和变体调用器的选择也显着影响结果。读长的影响更为复杂,并且表明随着读长的增加,映射特异性与产生更多假阳性的可能性之间可能存在相互作用。变体调用中涉及的工具和参数的选择会对产生的 FP SNP 数量产生巨大影响,特别糟糕的软件和/或参数设置组合在本实验中产生了数万个。因素间的相互作用使得简单的建议对于 SNP 发现流程来说变得困难,但参考序列的质量显然是至关重要的。我们的发现也清楚地提醒我们,当读数被映射到相对未完成的参考序列时,使用某些读数映射器默认提供的宽松的不匹配设置可能是不明智的。处于基因组探索早期阶段的非模型生物。本文的在线版本 (doi:10.1186/s12859-015-0801-z) 包含补充材料,可供授权用户使用。
Single Nucleotide Polymorphisms (SNPs) are widely used molecular markers, and their use has increased massively since the inception of Next Generation Sequencing (NGS) technologies, which allow detection of large numbers of SNPs at low cost. However, both NGS data and their analysis are error-prone, which can lead to the generation of false positive (FP) SNPs. We explored the relationship between FP SNPs and seven factors involved in mapping-based variant calling — quality of the reference sequence, read length, choice of mapper and variant caller, mapping stringency and filtering of SNPs by read mapping quality and read depth. This resulted in 576 possible factor level combinations. We used error- and variant-free simulated reads to ensure that every SNP found was indeed a false positive. The variation in the number of FP SNPs generated ranged from 0 to 36,621 for the 120 million base pairs (Mbp) genome. All of the experimental factors tested had statistically significant effects on the number of FP SNPs generated and there was a considerable amount of interaction between the different factors. Using a fragmented reference sequence led to a dramatic increase in the number of FP SNPs generated, as did relaxed read mapping and a lack of SNP filtering. The choice of reference assembler, mapper and variant caller also significantly affected the outcome. The effect of read length was more complex and suggests a possible interaction between mapping specificity and the potential for contributing more false positives as read length increases. The choice of tools and parameters involved in variant calling can have a dramatic effect on the number of FP SNPs produced, with particularly poor combinations of software and/or parameter settings yielding tens of thousands in this experiment. Between-factor interactions make simple recommendations difficult for a SNP discovery pipeline but the quality of the reference sequence is clearly of paramount importance. Our findings are also a stark reminder that it can be unwise to use the relaxed mismatch settings provided as defaults by some read mappers when reads are being mapped to a relatively unfinished reference sequence from e.g. a non-model organism in its early stages of genomic exploration. The online version of this article (doi:10.1186/s12859-015-0801-z) contains supplementary material, which is available to authorized users.