Validation and assessment of variant calling pipelines for next-generation sequencing.

Validation and assessment of variant calling pipelines for next-generation sequencing.
复制标题

DOI:
10.1186/1479-7364-8-14
复制
发表时间:
2014-07-30
期刊:
影响因子:
4.5
通讯作者:
Zandi PP
Zandi PP
中科院分区:
医学3区
文献类型:
--
作者:
Pirooznia M;Kramer M;Parla J;Goes FS;Potash JB;McCombie WR;Zandi PP

文献摘要

参考文献

被引文献

相似文献

The processing and analysis of the large scale data generated by next-generation sequencing (NGS) experiments is challenging and is a burgeoning area of new methods development. Several new bioinformatics tools have been developed for calling sequence variants from NGS data. Here, we validate the variant calling of these tools and compare their relative accuracy to determine which data processing pipeline is optimal. We developed a unified pipeline for processing NGS data that encompasses four modules: mapping, filtering, realignment and recalibration, and variant calling. We processed 130 subjects from an ongoing whole exome sequencing study through this pipeline. To evaluate the accuracy of each module, we conducted a series of comparisons between the single nucleotide variant (SNV) calls from the NGS data and either gold-standard Sanger sequencing on a total of 700 variants or array genotyping data on a total of 9,935 single-nucleotide polymorphisms. A head to head comparison showed that Genome Analysis Toolkit (GATK) provided more accurate calls than SAMtools (positive predictive value of 92.55% vs. 80.35%, respectively). Realignment of mapped reads and recalibration of base quality scores before SNV calling proved to be crucial to accurate variant calling. GATK HaplotypeCaller algorithm for variant calling outperformed the UnifiedGenotype algorithm. We also showed a relationship between mapping quality, read depth and allele balance, and SNV call accuracy. However, if best practices are used in data processing, then additional filtering based on these metrics provides little gains and accuracies of >99% are achievable. Our findings will help to determine the best approach for processing NGS data to confidently call variants for downstream analyses. To enable others to implement and replicate our results, all of our codes are freely available at http://metamoodics.org/wes.
DOI: 10.1016/j.neuron.2012.04.009
发表时间: 2012-04-26
期刊: Neuron
影响因子: 16.2
作者:
Iossifov I;Ronemus M;Levy D;Wang Z;Hakker I;Rosenbaum J;Yamrom B;Lee YH;Narzisi G;Leotta A;Kendall J;Grabowska E;Ma B;Marks S;Rodgers L;Stepansky A;Troge J;Andrews P;Bekritsky M;Pradhan K;Ghiban E;Kramer M;Parla J;Demeter R;Fulton LL;Fulton RS;Magrini VJ;Ye K;Darnell JC;Darnell RB;Mardis ER;Wilson RK;Schatz MC;McCombie WR;Wigler M
通讯作者: Wigler M
DOI: 10.1186/gb-2011-12-9-r97
发表时间: 2011-09-29
期刊: Genome biology
影响因子: 12.3
作者:
Parla JS;Iossifov I;Grabill I;Spector MS;Kramer M;McCombie WR
通讯作者: McCombie WR
DOI: 10.1006/jmbi.1994.1104
发表时间: 1994-02-04
影响因子: 5.6
作者:
KROGH, A;BROWN, M;HAUSSLER, D
通讯作者: HAUSSLER, D
使用下一代 DNA 测序数据进行变异发现和基因分型的框架。
DOI: 10.1038/ng.806
发表时间: 2011-05
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1186/gm432
发表时间: 2013
期刊: Genome medicine
影响因子: 12.3
作者:
O'Rawe J;Jiang T;Sun G;Wu Y;Wang W;Hu J;Bodily P;Tian L;Hakonarson H;Johnson WE;Wei Z;Wang K;Lyon GJ
通讯作者: Lyon GJ