ANALYSIS OF CONTEXT-DEPENDENT ERRORS FOR ILLUMINA SEQUENCING

ANALYSIS OF CONTEXT-DEPENDENT ERRORS FOR ILLUMINA SEQUENCING
复制标题

DOI:
10.1142/s0219720012410053
复制
发表时间:
2012-04-01
影响因子:
1
通讯作者:
Cox, Tony
Cox, Tony
中科院分区:
生物学4区
文献类型:
--
作者:
Abnizova, Irina;Leonard, Steven;Cox, Tony

文献摘要

被引文献

相似文献

新一代短读测序技术需要可靠的数据质量测量。这些措施对于变异识别尤其重要。然而,在SNP调用的特定情况下,可能获得大量假阳性SNP。人们需要区分推定的SNP与测序或其他错误。我们发现,不仅测序错误的概率(即质量值)对于区分FP-SNP很重要,而且“校正”该错误的条件概率(“第二最佳判定”概率,以第一判定的概率为条件)也很重要。令人惊讶的是,大约80%的错配可以通过第二次调用“纠正”。降低FP-SNP率的另一种方法是检索似乎易于测序错误的DNA基序,并将相应的条件质量值附加到这些基序上。我们已经开发了几种措施来区分序列错误和候选SNP,基于碱基调用的核苷酸背景及其错配类型。此外,我们提出了一种简单的方法来纠正大多数的错配,基于他们的“第二”最佳强度调用的条件概率。我们为每个不匹配附加相应的第二次调用置信度(质量值)。
The new generation of short-read sequencing technologies requires reliable measures of data quality. Such measures are especially important for variant calling. However, in the particular case of SNP calling, a great number of false-positive SNPs may be obtained. One needs to distinguish putative SNPs from sequencing or other errors. We found that not only the probability of sequencing errors (i.e. the quality value) is important to distinguish an FP-SNP but also the conditional probability of "correcting" this error (the "second best call" probability, conditional on that of the first call). Surprisingly, around 80% of mismatches can be "corrected" with this second call. Another way to reduce the rate of FP-SNPs is to retrieve DNA motifs that seem to be prone to sequencing errors, and to attach a corresponding conditional quality value to these motifs. We have developed several measures to distinguish between sequence errors and candidate SNPs, based on a base call's nucleotide context and its mismatch type. In addition, we suggested a simple method to correct the majority of mismatches, based on conditional probability of their "second" best intensity call. We attach a corresponding second call confidence (quality value) of being corrected to each mismatch.