Exploring the Consistency of the Quality Scores with Machine Learning for Next-Generation Sequencing Experiments

Exploring the Consistency of the Quality Scores with Machine Learning for Next-Generation Sequencing Experiments
复制标题

DOI:
10.1155/2020/8531502
复制
发表时间:
2020-02-26
影响因子:
--
通讯作者:
Oh, Min
Oh, Min
中科院分区:
生物学3区
文献类型:
--
作者:
Cosgun, Erdal;Oh, Min

文献摘要

被引文献

相似文献

背景下一代测序技术能够实现大规模并行处理,成本低于其他测序技术。在随后的NGS数据分析中,主要关注点之一是变异识别的可靠性。虽然研究人员可以利用变异识别的原始质量分数,但他们被迫在没有对质量分数进行任何预评估的情况下开始进一步的分析。法我们提出了一种机器学习方法,用于估计来自BWA+GATK的变体调用的质量分数。我们分析了质量分数和这些注释之间的相关性,指定了信息注释,这些注释被用作预测变体质量分数的特征。为了测试预测模型,我们模拟了24个具有30x覆盖率碱基的双端Illumina测序读段。此外,从序列读段档案中获得了由Illumina配对末端测序产生的24个人类基因组测序读段,覆盖率至少为30倍。结果使用BWA+GATK,从模拟的和真实的测序读段衍生VCF。我们观察到,RFR学习的预测模型在模拟和真实的数据中均优于其他算法。在模拟的人类基因组VCF数据中,变体调用的质量评分可从GATK注释模块的信息特征高度预测(RFR、MLR和NNR的R2分别为96.7%、94.4%和89.8%)。在真实的人类基因组VCF数据中,所提出的数据驱动模型的稳健性始终保持不变(RFR和MLR的R2分别为97.8%和96.5%)。
Background. Next-generation sequencing enables massively parallel processing, allowing lower cost than the other sequencing technologies. In the subsequent analysis with the NGS data, one of the major concerns is the reliability of variant calls. Although researchers can utilize raw quality scores of variant calling, they are forced to start the further analysis without any preevaluation of the quality scores. Method. We presented a machine learning approach for estimating quality scores of variant calls derived from BWA+GATK. We analyzed correlations between the quality score and these annotations, specifying informative annotations which were used as features to predict variant quality scores. To test the predictive models, we simulated 24 paired-end Illumina sequencing reads with 30x coverage base. Also, twenty-four human genome sequencing reads resulting from Illumina paired-end sequencing with at least 30x coverage were secured from the Sequence Read Archive. Results. Using BWA+GATK, VCFs were derived from simulated and real sequencing reads. We observed that the prediction models learned by RFR outperformed other algorithms in both simulated and real data. The quality scores of variant calls were highly predictable from informative features of GATK Annotation Modules in the simulated human genome VCF data (R2: 96.7%, 94.4%, and 89.8% for RFR, MLR, and NNR, respectively). The robustness of the proposed data-driven models was consistently maintained in the real human genome VCF data (R2: 97.8% and 96.5% for RFR and MLR, respectively).