Exploring the limit of using a deep neural network on pileup data for germline variant calling

Exploring the limit of using a deep neural network on pileup data for germline variant calling
复制标题

DOI:
10.1038/s42256-020-0167-4
复制
发表时间:
2020-04-01
影响因子:
23.8
通讯作者:
Lam, Tak-Wah
Lam, Tak-Wah
中科院分区:
计算机科学1区
文献类型:
--
作者:
Luo, Ruibang;Wong, Chak-Lim;Lam, Tak-Wah

文献摘要

被引文献

相似文献

近年来出现了单分子测序技术,并彻底改变了结构变异识别、复杂基因组组装和表观遗传标记检测。然而,缺乏高度准确的小变异识别器限制了这些技术的更广泛应用。在这里,我们介绍了Clair,Clairvoyante的继任者,一个使用单分子测序数据快速准确的种系小变异呼叫程序。对于Oxford Nanopore Technology数据,Clair实现了比Clairvoyante,Longshot和Medaka等几个竞争程序更好的精确度,召回率和速度。通过研究遗漏的变异和对有意过拟合模型的基准测试,我们发现Clair可能正在接近使用堆积数据和深度神经网络进行种系小变异识别的可能准确性极限。Clair只需要一个传统的中央处理器(CPU)来进行变体调用,它是一个开源项目,可以在www.example.com上找到。缺乏准确和有效的变异识别方法阻碍了单分子测序技术的临床应用。作者提出了一种使用单分子测序数据进行快速准确的种系小变异识别的深度学习方法。
Single-molecule sequencing technologies have emerged in recent years and revolutionized structural variant calling, complex genome assembly and epigenetic mark detection. However, the lack of a highly accurate small variant caller has limited these technologies from being more widely used. Here, we present Clair, the successor to Clairvoyante, a program for fast and accurate germline small variant calling, using single-molecule sequencing data. For Oxford Nanopore Technology data, Clair achieves better precision, recall and speed than several competing programs, including Clairvoyante, Longshot and Medaka. Through studying the missed variants and benchmarking intentionally overfitted models, we found that Clair may be approaching the limit of possible accuracy for germline small variant calling using pileup data and deep neural networks. Clair requires only a conventional central processing unit (CPU) for variant calling and is an open-source project available at https://github.com/HKU-BAL/Clair. A lack of accurate and efficient variant calling methods has held back single-molecule sequencing technologies from clinical applications. The authors present a deep-learning method for fast and accurate germline small variant calling, using single-molecule sequencing data.