Comparing nominal and real quality scores on next-generation sequencing genotype calls.

Comparing nominal and real quality scores on next-generation sequencing genotype calls.
复制标题

DOI:
10.1186/1753-6561-5-s9-s14
复制
发表时间:
2011-11-29
期刊:
影响因子:
--
通讯作者:
Stram, Alexander H
Stram, Alexander H
中科院分区:
其他
文献类型:
--
作者:
Stram, Alexander H

文献摘要

被引文献

相似文献

我试图全面评估遗传分析工作坊17(GAW 17)数据集的质量,通过检查其基因型调用的准确性,这是基于1000个基因组计划的pilot 3数据。利用1000个基因组计划/HapMap样本的交叉,我比较了个体的GAW 17基因型调用与HapMap III第2版基因型调用。这些基因型调用应该是一致的几乎无处不在。相反,我发现了一个令人难以置信的低65.4%的一致性。将HapMap作为金标准,我假设这是一个GAW 17数据问题,并试图相应地解释这种不一致。我发现这种不一致性的很大一部分发生在目标区域之外,并且通过简单地停留在目标区域内,可以将一致性提高到至少94.6%,这些目标区域在更多样本中进行测序。此外,我发现在某些个体中,高样本数对提高一致性几乎没有作用,并得出结论,某些样本的序列读数的质量分数是不正确的。
I seek to comprehensively evaluate the quality of the Genetic Analysis Workshop 17 (GAW17) data set by examining the accuracy of its genotype calls, which were based on the pilot3 data of the 1000 Genomes Project. Taking advantage of the 1000 Genomes Project/HapMap sample intersect, I compared GAW17 genotype calls to HapMap III, release 2, genotype calls for an individual. These genotype calls should be concordant almost everywhere. Instead I found an astonishingly low 65.4% concordance. Regarding HapMap as the gold standard, I assume that this is a GAW17 data problem and seek to explain this discordance accordingly. I found that a large proportion of this discordance occurred outside targeted regions and that concordance could be improved to at least 94.6% by simply staying within targeted regions, which were sequenced across more samples. Furthermore, I found that in certain individuals, high sample counts did little to improve concordance and concluded that quality scores for a certain sample's sequence reads were simply incorrect.