ForestQC: Quality control on genetic variants from next-generation sequencing data using random forest

ForestQC: Quality control on genetic variants from next-generation sequencing data using random forest
复制标题

ForestQC:使用随机森林对来自下一代测序数据的遗传变异进行质量控制

DOI:
10.1371/journal.pcbi.1007556
复制
发表时间:
2019-12-01
影响因子:
4.3
通讯作者:
Sul, Jae Hoon
Sul, Jae Hoon
中科院分区:
生物学2区
文献类型:
--
作者:
Li, Jiajin;Jew, Brandon;Sul, Jae Hoon

文献摘要

被引文献

相似文献

遗传性疾病可由多种类型的基因突变引起,包括常见和罕见的单核苷酸变异、结构变异、插入和缺失。如今,下一代测序(NGS)技术使我们能够识别与疾病相关的各种遗传变异。然而,由于测序技术和分析工具的偏差和错误,NGS检测到的变异可能具有较差的测序质量。因此,去除低质量的变异至关重要,因为这可能会在后续分析中导致虚假结果。以前,人们使用硬过滤器或机器学习模型进行变体质量控制(QC),但无法准确过滤出这些变体。在这里,我们开发了一个统计工具,ForestQC,通过结合过滤方法和机器学习方法的变体QC。我们将ForestQC应用于一个基于家族的全基因组测序(WGS)数据集和一个一般病例对照WGS数据集,对其进行评估。结果表明,ForestQC通过显着提高变异的质量而优于广泛使用的变异QC方法。此外,ForestQC非常高效,可扩展到大规模测序数据集。我们的研究表明,结合过滤方法和机器学习方法可以实现有效的变异QC。下一代测序技术(NGS)可以发现基因组中存在的几乎所有遗传变异。然而,由于NGS或变体调用者的限制,这些变体的子集可能具有差的测序质量。在分析大量测序个体的遗传学研究中,检测和去除那些质量差的变异是至关重要的,因为它们可能导致虚假的发现。在本文中,我们提出了ForestQC,这是一种统计工具,通过结合传统的过滤方法和机器学习方法,对从NGS数据中识别的变体进行质量控制。我们的软件使用有关测序质量的信息,如测序深度、基因分型质量和GC含量,来预测特定变体是否可能是假阳性。为了评估ForestQC,我们将其应用于两个全基因组测序数据集,其中一个数据集由来自家族的相关个体组成,而另一个数据集由不相关个体组成。结果表明,ForestQC通过显著提高分析中包含的变体的质量,优于广泛使用的变体质量控制方法,如GATK的VQSR。ForestQC也非常有效,因此可以应用于大型测序数据集。我们的结论是,结合使用测序质量信息训练的机器学习算法和过滤方法是对测序数据中的遗传变异进行质量控制的实用方法。
Author summary Genetic disorders can be caused by many types of genetic mutations, including common and rare single nucleotide variants, structural variants, insertions, and deletions. Nowadays, next-generation sequencing (NGS) technology allows us to identify various genetic variants that are associated with diseases. However, variants detected by NGS might have poor sequencing quality due to biases and errors in sequencing technologies and analysis tools. Therefore, it is critical to remove variants with low quality, which could cause spurious findings in follow-up analyses. Previously, people applied either hard filters or machine learning models for variant quality control (QC), which failed to filter out those variants accurately. Here, we developed a statistical tool, ForestQC, for variant QC by combining a filtering approach and a machine learning approach. We applied ForestQC to one family-based whole-genome sequencing (WGS) dataset and one general case-control WGS dataset, to evaluate it. Results show that ForestQC outperforms widely used methods for variant QC by considerably improving the quality of variants. Also, ForestQC is very efficient and scalable to large-scale sequencing datasets. Our study indicates that combining filtering approaches and machine learning approaches enables effective variant QC.Next-generation sequencing technology (NGS) enables the discovery of nearly all genetic variants present in a genome. A subset of these variants, however, may have poor sequencing quality due to limitations in NGS or variant callers. In genetic studies that analyze a large number of sequenced individuals, it is critical to detect and remove those variants with poor quality as they may cause spurious findings. In this paper, we present ForestQC, a statistical tool for performing quality control on variants identified from NGS data by combining a traditional filtering approach and a machine learning approach. Our software uses the information on sequencing quality, such as sequencing depth, genotyping quality, and GC contents, to predict whether a particular variant is likely to be false-positive. To evaluate ForestQC, we applied it to two whole-genome sequencing datasets where one dataset consists of related individuals from families while the other consists of unrelated individuals. Results indicate that ForestQC outperforms widely used methods for performing quality control on variants such as VQSR of GATK by considerably improving the quality of variants to be included in the analysis. ForestQC is also very efficient, and hence can be applied to large sequencing datasets. We conclude that combining a machine learning algorithm trained with sequencing quality information and the filtering approach is a practical approach to perform quality control on genetic variants from sequencing data.