Comparing the performance of selected variant callers using synthetic data and genome segmentation.

Comparing the performance of selected variant callers using synthetic data and genome segmentation.
复制标题

DOI:
10.1186/s12859-018-2440-7
复制
发表时间:
2018-11-19
期刊:
影响因子:
3
通讯作者:
Meerzaman D
Meerzaman D
中科院分区:
生物学4区
文献类型:
--
作者:
Bian X;Zhu B;Wang M;Hu Y;Chen Q;Nguyen C;Hicks B;Meerzaman D

文献摘要

参考文献

被引文献

相似文献

高通量测序已迅速成为精准癌症医学的重要组成部分。但是,验证从分析和解释基因组数据获得的结果仍然是一个限速因素。当然,黄金标准仍然是由专家小组进行人工验证,但这种方法并非没有缺点,即资金和时间成本高,人工验证必然具有选择性。但也许可以开发出更经济、更互补的验证手段。在这项研究中,我们使用了四个复杂性不断增加的合成数据集(已知突变的变体加标到特定的基因组位置)来评估五个开源变体调用程序的灵敏度,特异性和平衡准确性:FreeBayes v1.0,VarDict v11.5.1,MuTect v1.1.7,MuTect 2和MuSE v1.0rc。在bcbio-next gen中运行FreeBayes、VarDict和MuTect,并将结果整合到单个Enclave调用集中。已知的突变提供了一个水平的“地面真相”,我们评估了变异调用者的性能。我们通过将整个基因组分割成10,000,000个碱基对片段,产生316个片段,进一步促进了比较和评估。在调用者中,真阳性的数量之间的差异很小,但是当这些工具被用来分析第一组到第三组时,假阳性的数量变化更大。FreeBayes和VarDict都产生了比其他人更多的假阳性,尽管VarDict也产生了最高数量的真阳性。Entrance方法产生的结果的特点是更高的特异性和平衡的准确性和更少的假阳性比单独使用的任何五个工具。然而,随着数据集复杂性的增加,所有五个呼叫者的灵敏度和特异性都有所下降,但我们没有发现呼叫者的表现与某些DNA结构特征(基因密度和鸟嘌呤-胞嘧啶含量)之间存在有限的弱相关性。总的来说,MuTect 2在测试的调用者中表现最好,其次是MuSE和MuTect。在基因组的已知位置上使用特定突变(本研究中的单核苷酸变异(SNV)、单核苷酸多态性(SNP)或结构变异(SV))的加标数据集提供了一种有效且经济的方法,可以将变异调用者分析的数据与地面实况进行比较。该方法构成了一个可行的替代长期的,昂贵的,和不全面的评估专家小组。应进一步发展和完善这一方法,其他评估准确性的相对“轻量级”方法也应如此。鉴于科学界尚未建立验证NGS相关技术(如变异识别器)的黄金标准,开发多种替代方法来验证变异识别器的准确性最终将导致建立比过早限制社区成员探索的创新方法范围更高质量的标准。
High-throughput sequencing has rapidly become an essential part of precision cancer medicine. But validating results obtained from analyzing and interpreting genomic data remains a rate-limiting factor. The gold standard, of course, remains manual validation by expert panels, which is not without its weaknesses, namely high costs in both funding and time as well as the necessarily selective nature of manual validation. But it may be possible to develop more economical, complementary means of validation. In this study we employed four synthetic data sets (variants with known mutations spiked into specific genomic locations) of increasing complexity to assess the sensitivity, specificity, and balanced accuracy of five open-source variant callers: FreeBayes v1.0, VarDict v11.5.1, MuTect v1.1.7, MuTect2, and MuSE v1.0rc. FreeBayes, VarDict, and MuTect were run in bcbio-next gen, and the results were integrated into a single Ensemble call set. The known mutations provided a level of “ground truth” against which we evaluated variant-caller performance. We further facilitated the comparison and evaluation by segmenting the whole genome into 10,000,000 base-pair fragments which yielded 316 segments. Differences among the numbers of true positives were small among the callers, but the numbers of false positives varied much more when the tools were used to analyze sets one through three. Both FreeBayes and VarDict produced strikingly more false positives than did the others, although VarDict, somewhat paradoxically also produced the highest number of true positives. The Ensemble approach yielded results characterized by higher specificity and balanced accuracy and fewer false positives than did any of the five tools used alone. Sensitivity and specificity, however, declined for all five callers as the complexity of the data sets increased, but we did not uncover anything more than limited, weak correlations between caller performance and certain DNA structural features: gene density and guanine-cytosine content. Altogether, MuTect2 performed the best among the callers tested, followed by MuSE and MuTect. Spiking data sets with specific mutations –single-nucleotide variations (SNVs), single-nucleotide polymorphisms (SNPs), or structural variations (SVs) in this study—at known locations in the genome provides an effective and economical way to compare data analyzed by variant callers with ground truth. The method constitutes a viable alternative to the prolonged, expensive, and noncomprehensive assessment by expert panels. It should be further developed and refined, as should other comparatively “lightweight” methods of assessing accuracy. Given that the scientific community has not yet established gold standards for validating NGS-related technologies such as variant callers, developing multiple alternative means for verifying variant-caller accuracy will eventually lead to the establishment of higher-quality standards than could be achieved by prematurely limiting the range of innovative methods explored by members of the community.
DOI: 10.1038/srep36540
发表时间: 2016-11-22
期刊: Scientific reports
影响因子: 4.6
作者:
Cai L;Yuan W;Zhang Z;He L;Chou KC
通讯作者: Chou KC
评估九个体细胞变体呼叫者,用于检测外显子和靶向深度测序数据中的体细胞突变。
DOI: 10.1371/journal.pone.0151664
发表时间: 2016
期刊: PloS one
影响因子: 3.7
作者:
Krøigård AB;Thomassen M;Lænkholm AV;Kruse TA;Larsen MJ
通讯作者: Larsen MJ
DOI: 10.1038/nbt.2514
发表时间: 2013-03
影响因子: 46.9
作者:
通讯作者: --
DOI: 10.1038/srep43169
发表时间: 2017-02-24
期刊: Scientific reports
影响因子: 4.6
作者:
Sandmann S;de Graaf AO;Karimi M;van der Reijden BA;Hellström-Lindberg E;Jansen JH;Dugas M
通讯作者: Dugas M
DOI: 10.1186/gm495
发表时间: 2013
期刊: Genome medicine
影响因子: 12.3
作者:
Wang Q;Jia P;Li F;Chen H;Ji H;Hucks D;Dahlman KB;Pao W;Zhao Z
通讯作者: Zhao Z