Recalibration of mapping quality scores in Illumina short-read alignments improves SNP detection results in low-coverage sequencing data.

Recalibration of mapping quality scores in Illumina short-read alignments improves SNP detection results in low-coverage sequencing data.
复制标题

DOI:
10.7717/peerj.10501
复制
发表时间:
2020
期刊:
影响因子:
2.7
通讯作者:
Eungwanichayapant A
Eungwanichayapant A
中科院分区:
生物学3区
文献类型:
--
作者:
Cline E;Wisittipanit N;Boongoen T;Chukeatirote E;Struss D;Eungwanichayapant A

文献摘要

参考文献

被引文献

相似文献

低覆盖测序是获得跨越整个基因组的读段的具有成本效益的方式。然而,每个基因座处的读取深度较低,使得测序误差难以与实际变异分开。在变异识别之前,测序仪读数与参考基因组比对,比对结果存储在序列比对/映射(SAM)文件中。每个比对具有映射质量(MAPQ)分数,其指示读数不正确比对的概率。本研究调查了用于计算MAPQ评分的概率估计值的重新校准,以提高单样本低覆盖率环境中的变异识别性能。模拟的番茄、辣椒和水稻基因组被植入了已知的变异体。由此,在低覆盖度下生成模拟的双端读段,并与原始参考基因组比对。从番茄的SAM格式比对文件中提取的特征用于训练机器学习模型,以检测不正确对齐的读数,并输出所有三个数据集中每个读数的未对齐概率的估计值。然后根据这些估计值重新计算MAPQ评分。接下来,使用新的MAPQ评分更新SAM文件。最后,对原始比对和重新校准的比对进行变体识别,并比较结果。不正确对齐的读数仅占训练集中读数的0.16%。这种严重的阶级不平衡需要特别考虑模型训练。用于检测错配读段的F1得分范围为0.76至0.82。使用性能最佳的模型计算新的MAPQ评分。单核苷酸多态性(SNP)检测后,映射评分重新校准。在水稻中,被称为SNPs的回忆增加了5.2%,而番茄和辣椒分别增加了3.1%和1.5%。对于所有三个数据集,SNP调用的精度范围为0.91至0.95,并且在映射评分重新校准之前和之后基本上没有变化。重新校准MAPQ评分可适度改善单样本变异识别结果。一些变体调用器同时对多个样品进行操作。它们利用每个样品的读数来补偿单个样品的低读数深度。这改进了多态性检测和基因型推断。单样本设置中的微小改进可能会转化为多样本实验中的较大增益。目前正在进行一项调查研究。
Low-coverage sequencing is a cost-effective way to obtain reads spanning an entire genome. However, read depth at each locus is low, making sequencing error difficult to separate from actual variation. Prior to variant calling, sequencer reads are aligned to a reference genome, with alignments stored in Sequence Alignment/Map (SAM) files. Each alignment has a mapping quality (MAPQ) score indicating the probability a read is incorrectly aligned. This study investigated the recalibration of probability estimates used to compute MAPQ scores for improving variant calling performance in single-sample, low-coverage settings. Simulated tomato, hot pepper and rice genomes were implanted with known variants. From these, simulated paired-end reads were generated at low coverage and aligned to the original reference genomes. Features extracted from the SAM formatted alignment files for tomato were used to train machine learning models to detect incorrectly aligned reads and output estimates of the probability of misalignment for each read in all three data sets. MAPQ scores were then re-computed from these estimates. Next, the SAM files were updated with new MAPQ scores. Finally, Variant calling was performed on the original and recalibrated alignments and the results compared. Incorrectly aligned reads comprised only 0.16% of the reads in the training set. This severe class imbalance required special consideration for model training. The F1 score for detecting misaligned reads ranged from 0.76 to 0.82. The best performing model was used to compute new MAPQ scores. Single Nucleotide Polymorphism (SNP) detection was improved after mapping score recalibration. In rice, recall for called SNPs increased by 5.2%, while for tomato and pepper it increased by 3.1% and 1.5%, respectively. For all three data sets the precision of SNP calls ranged from 0.91 to 0.95, and was largely unchanged both before and after mapping score recalibration. Recalibrating MAPQ scores delivers modest improvements in single-sample variant calling results. Some variant callers operate on multiple samples simultaneously. They exploit every sample’s reads to compensate for the low read-depth of individual samples. This improves polymorphism detection and genotype inference. It may be that small improvements in single-sample settings translate to larger gains in a multi-sample experiment. A study to investigate this is ongoing.
DOI: 10.1093/bioinformatics/btq293
发表时间: 2010-08-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Zeitouni B;Boeva V;Janoueix-Lerosey I;Loeillet S;Legoix-né P;Nicolas A;Delattre O;Barillot E
通讯作者: Barillot E
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1186/s12870-016-0931-0
发表时间: 2016-10-28
期刊: BMC plant biology
影响因子: 5.3
作者:
Kang YJ;Ahn YK;Kim KT;Jun TH
通讯作者: Jun TH
DOI: 10.1093/bioinformatics/btr509
发表时间: 2011-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, Heng
通讯作者: Li, Heng
使用下一代 DNA 测序数据进行变异发现和基因分型的框架。
DOI: 10.1038/ng.806
发表时间: 2011-05
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --