Functional equivalence of genome sequencing analysis pipelines enables harmonized variant calling across human genetics projects.

Functional equivalence of genome sequencing analysis pipelines enables harmonized variant calling across human genetics projects.
复制标题

DOI:
10.1038/s41467-018-06159-4
复制
发表时间:
2018-10-02
影响因子:
16.6
通讯作者:
Hall IM
Hall IM
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Regier AA;Farjoun Y;Larson DE;Krasheninina O;Kang HM;Howrigan DP;Chen BJ;Kher M;Banks E;Ames DC;English AC;Li H;Xing J;Zhang Y;Matise T;Abecasis GR;Salerno W;Zody MC;Neale BM;Hall IM

文献摘要

参考文献

被引文献

相似文献

未来几年将产生数十万个人类全基因组测序(WGS)数据集。这些数据集中起来更有价值:对来自许多来源的基因组进行联合分析,增加了样本量和统计能力。联合分析面临的一个核心挑战是,不同的WGS数据处理管道会导致组合数据集的变量调用存在实质性差异,因此需要进行计算上昂贵的重新处理。考虑到当前研究的规模和数据量,这种方法不再站得住脚。在这里,我们定义了WGS数据处理标准,允许不同的组产生功能等效(FE)结果,但仍然在数据处理管道上进行创新。我们提出了在五个基因组中心开发的初始FE管道,并表明它们产生相似的变体调用结果,并且产生的可变性明显小于测序重复。这项工作缓解了基因组聚集的关键技术瓶颈,并有助于为社区范围内的人类遗传学研究奠定基础。全基因组测序(WGS)数据的共享提高了研究的规模和能力,但来自不同群体的数据往往不兼容。在这里,美国基因组中心和NIH项目定义了WGS数据处理标准和灵活的验证方法,促进了人类遗传学研究的合作。
Hundreds of thousands of human whole genome sequencing (WGS) datasets will be generated over the next few years. These data are more valuable in aggregate: joint analysis of genomes from many sources increases sample size and statistical power. A central challenge for joint analysis is that different WGS data processing pipelines cause substantial differences in variant calling in combined datasets, necessitating computationally expensive reprocessing. This approach is no longer tenable given the scale of current studies and data volumes. Here, we define WGS data processing standards that allow different groups to produce functionally equivalent (FE) results, yet still innovate on data processing pipelines. We present initial FE pipelines developed at five genome centers and show that they yield similar variant calling results and produce significantly less variability than sequencing replicates. This work alleviates a key technical bottleneck for genome aggregation and helps lay the foundation for community-wide human genetics studies. Sharing of whole genome sequencing (WGS) data improves study scale and power, but data from different groups are often incompatible. Here, US genome centers and NIH programs define WGS data processing standards and a flexible validation method, facilitating collaboration in human genetics research.
来自1,092个人基因组的遗传变异的综合图。
DOI: 10.1038/nature11632
发表时间: 2012-11-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1038/s41593-017-0017-9
发表时间: 2017-12
影响因子: 25
作者:
Sanders SJ;Neale BM;Huang H;Werling DM;An JY;Dong S;Whole Genome Sequencing for Psychiatric Disorders (WGSPD);Abecasis G;Arguello PA;Blangero J;Boehnke M;Daly MJ;Eggan K;Geschwind DH;Glahn DC;Goldstein DB;Gur RE;Handsaker RE;McCarroll SA;Ophoff RA;Palotie A;Pato CN;Sabatti C;State MW;Willsey AJ;Hyman SE;Addington AM;Lehner T;Freimer NB
通讯作者: Freimer NB
DOI: 10.1126/science.aaf6162
发表时间: 2016-06-10
期刊: Science (New York, N.Y.)
影响因子: --
作者:
通讯作者: --
使用下一代 DNA 测序数据进行变异发现和基因分型的框架。
DOI: 10.1038/ng.806
发表时间: 2011-05
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1186/gb-2014-15-6-r84
发表时间: 2014-06-26
期刊: Genome biology
影响因子: 12.3
作者:
Layer RM;Chiang C;Quinlan AR;Hall IM
通讯作者: Hall IM