An open resource for accurately benchmarking small variant and reference calls

An open resource for accurately benchmarking small variant and reference calls
复制标题

DOI:
10.1038/s41587-019-0074-6
复制
发表时间:
2019-05-01
影响因子:
46.9
通讯作者:
Salit, Marc
Salit, Marc
中科院分区:
工程技术1区
文献类型:
--
作者:
Zook, Justin M.;McDaniel, Jennifer;Salit, Marc

文献摘要

被引文献

相似文献

基准小变异调用是开发、优化和评估测序和生物信息学方法性能的必要条件。在这里,作为瓶中基因组(GIAB)联盟的一部分,我们应用可重复的、基于云的管道来整合多个短链和链读测序数据集,并为人类基因组提供基准调用。我们为一个先前分析过的GIAB样本以及来自个人基因组计划的六个基因组生成基准调用。这些新的基因组获得了广泛的、公开的同意,使其成为“同类首创”的资源,可供社区用于多个下游应用。我们产生的基准单核苷酸变异比以前发布的GIAB基准多17%,索引多176%,基准区域大12%。我们证明了这个基准可以可靠地识别现有调用集中的错误,并强调了在使用不完美或不全面的基准时解释性能指标的挑战。最后,我们根据变异类型和基因组上下文对呼叫集的性能进行分层,从而识别出呼叫集的优缺点。
Benchmark small variant calls are required for developing, optimizing and assessing the performance of sequencing and bioinformatics methods. Here, as part of the Genome in a Bottle (GIAB) Consortium, we apply a reproducible, cloud-based pipeline to integrate multiple short- and linked-read sequencing datasets and provide benchmark calls for human genomes. We generate benchmark calls for one previously analyzed GIAB sample, as well as six genomes from the Personal Genome Project. These new genomes have broad, open consent, making this a 'first of its kind' resource that is available to the community for multiple downstream applications. We produce 17% more benchmark single nucleotide variations, 176% more indels and 12% larger benchmark regions than previously published GIAB benchmarks. We demonstrate that this benchmark reliably identifies errors in existing callsets and highlight challenges in interpreting performance metrics when using benchmarks that are not perfect or comprehensive. Finally, we identify strengths and weaknesses of callsets by stratifying performance according to variant type and genome context.