Evaluating Variant Calling Tools for Non-Matched Next-Generation Sequencing Data.

Evaluating Variant Calling Tools for Non-Matched Next-Generation Sequencing Data.
复制标题

DOI:
10.1038/srep43169
复制
发表时间:
2017-02-24
期刊:
影响因子:
4.6
通讯作者:
Dugas M
Dugas M
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Sandmann S;de Graaf AO;Karimi M;van der Reijden BA;Hellström-Lindberg E;Jansen JH;Dugas M

文献摘要

被引文献

相似文献

有效的变异召唤结果对于下一代测序在临床常规中的应用至关重要。然而,有许多不同的调用工具通常在算法、过滤策略、建议以及输出方面有所不同。我们评估了8个开源工具在非匹配的下一代测序数据中调用等位基因频率低至1%的单核苷酸变异和短索引的能力:GATK HaplotypeCaller、Platypus、VarScan、LoFreq、FreeBayes、SNVer、SAMtools和VarDict。我们分析了来自骨髓增生异常综合征患者的两个真实数据集,包括54个Illumina HiSeq样本和111个Illumina NextSeq样本。通过在同一平台、不同平台和专家评审上的重新测序来验证突变。此外,我们考虑了两个具有不同覆盖率和错误概况的模拟数据集,每个数据集覆盖50个样本。在所有病例中,分析了一个由19个基因组成的相同靶区(42,322 bp)。总之,没有一种工具能够成功地调用所有的突变。高灵敏度往往伴随着低精度。不同的覆盖和背景噪声对不同呼叫的影响一般较低。考虑到所有因素,VarDict表现最好。然而,我们的结果表明,在多线程环境下,有必要提高结果的可重复性。
Valid variant calling results are crucial for the use of next-generation sequencing in clinical routine. However, there are numerous variant calling tools that usually differ in algorithms, filtering strategies, recommendations and thus, also in the output. We evaluated eight open-source tools regarding their ability to call single nucleotide variants and short indels with allelic frequencies as low as 1% in non-matched next-generation sequencing data: GATK HaplotypeCaller, Platypus, VarScan, LoFreq, FreeBayes, SNVer, SAMtools and VarDict. We analysed two real datasets from patients with myelodysplastic syndrome, covering 54 Illumina HiSeq samples and 111 Illumina NextSeq samples. Mutations were validated by re-sequencing on the same platform, on a different platform and expert based review. In addition we considered two simulated datasets with varying coverage and error profiles, covering 50 samples each. In all cases an identical target region consisting of 19 genes (42,322 bp) was analysed. Altogether, no tool succeeded in calling all mutations. High sensitivity was always accompanied by low precision. Influence of varying coverages- and background noise on variant calling was generally low. Taking everything into account, VarDict performed best. However, our results indicate that there is a need to improve reproducibility of the results in the context of multithreading.