Deep whole-genome sequencing of 90 Han Chinese genomes.

Deep whole-genome sequencing of 90 Han Chinese genomes.
复制标题

DOI:
10.1093/gigascience/gix067
复制
发表时间:
2017-09-01
期刊:
影响因子:
9.2
通讯作者:
Guo X
Guo X
中科院分区:
生物学2区
文献类型:
--
作者:
Lan T;Lin H;Zhu W;Laurent TCAM;Yang M;Liu X;Wang J;Wang J;Yang H;Xu X;Guo X

文献摘要

参考文献

被引文献

相似文献

下一代测序提供了对人类遗传信息的高分辨率洞察。然而,由于测序的高成本,以前的研究主要集中在低覆盖率数据上。尽管千人基因组计划和单倍型参考联盟都为插补提供了强大的参考面板,但低频率和新的变异仍然难以在低覆盖率数据的基础上准确地发现和调用。深度测序为这些低频和新变体的问题提供了最佳解决方案。虽然全外显子组测序也是外显子组区域的可行选择,但它不能解释非编码区,有时会导致缺乏重要的因果变异。对于中国汉族人群,大多数变异都是基于千人基因组计划的低覆盖率数据发现的。然而,高覆盖率的全基因组测序数据对于任何人群都是有限的,并且大量低频率的人群特异性变体仍然没有被表征。我们对90个中国血统的无关个体进行了全基因组测序,这些个体是从1000个基因组计划样本中收集的,包括45个北方汉族和45个南方汉族样本。这90个基因组中的83个已经被1000个基因组计划测序。我们从这90个样本中鉴定了12 568 804个单核苷酸多态性,2 074 210个短InDels和26 142个结构变异。与千人基因组计划数据比较,共发现7 000 629个低频率(定义为次要等位基因频率< 5%)的新变异,包括5 813 503个单核苷酸多态性、1 169 199个InDels和17 927个结构变异。利用深度测序数据,我们已经为中国汉族基因组建立了一个大大扩展的遗传变异谱。与千人基因组计划相比,这些中国汉族人的深度测序数据增强了对大量低频新变异的表征。这将是促进中国遗传学研究和医学发展的宝贵资源。此外,它将为1000个基因组计划以及其他人类基因组计划提供宝贵的补充。
Next-generation sequencing provides a high-resolution insight into human genetic information. However, the focus of previous studies has primarily been on low-coverage data due to the high cost of sequencing. Although the 1000 Genomes Project and the Haplotype Reference Consortium have both provided powerful reference panels for imputation, low-frequency and novel variants remain difficult to discover and call with accuracy on the basis of low-coverage data. Deep sequencing provides an optimal solution for the problem of these low-frequency and novel variants. Although whole-exome sequencing is also a viable choice for exome regions, it cannot account for noncoding regions, sometimes resulting in the absence of important, causal variants. For Han Chinese populations, the majority of variants have been discovered based upon low-coverage data from the 1000 Genomes Project. However, high-coverage, whole-genome sequencing data are limited for any population, and a large amount of low-frequency, population-specific variants remain uncharacterized. We have performed whole-genome sequencing at a high depth (∼×80) of 90 unrelated individuals of Chinese ancestry, collected from the 1000 Genomes Project samples, including 45 Northern Han Chinese and 45 Southern Han Chinese samples. Eighty-three of these 90 have been sequenced by the 1000 Genomes Project. We have identified 12 568 804 single nucleotide polymorphisms, 2 074 210 short InDels, and 26 142 structural variations from these 90 samples. Compared to the Han Chinese data from the 1000 Genomes Project, we have found 7 000 629 novel variants with low frequency (defined as minor allele frequency < 5%), including 5 813 503 single nucleotide polymorphisms, 1 169 199 InDels, and 17 927 structural variants. Using deep sequencing data, we have built a greatly expanded spectrum of genetic variation for the Han Chinese genome. Compared to the 1000 Genomes Project, these Han Chinese deep sequencing data enhance the characterization of a large number of low-frequency, novel variants. This will be a valuable resource for promoting Chinese genetics research and medical development. Additionally, it will provide a valuable supplement to the 1000 Genomes Project, as well as to other human genome projects.
来自1,092个人基因组的遗传变异的综合图。
DOI: 10.1038/nature11632
发表时间: 2012-11-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1038/nature15394
发表时间: 2015-10-01
期刊: Nature
影响因子: 64.8
作者:
Sudmant PH;Rausch T;Gardner EJ;Handsaker RE;Abyzov A;Huddleston J;Zhang Y;Ye K;Jun G;Fritz MH;Konkel MK;Malhotra A;Stütz AM;Shi X;Casale FP;Chen J;Hormozdiari F;Dayama G;Chen K;Malig M;Chaisson MJP;Walter K;Meiers S;Kashin S;Garrison E;Auton A;Lam HYK;Mu XJ;Alkan C;Antaki D;Bae T;Cerveira E;Chines P;Chong Z;Clarke L;Dal E;Ding L;Emery S;Fan X;Gujral M;Kahveci F;Kidd JM;Kong Y;Lameijer EW;McCarthy S;Flicek P;Gibbs RA;Marth G;Mason CE;Menelaou A;Muzny DM;Nelson BJ;Noor A;Parrish NF;Pendleton M;Quitadamo A;Raeder B;Schadt EE;Romanovitch M;Schlattl A;Sebra R;Shabalin AA;Untergasser A;Walker JA;Wang M;Yu F;Zhang C;Zhang J;Zheng-Bradley X;Zhou W;Zichner T;Sebat J;Batzer MA;McCarroll SA;1000 Genomes Project Consortium;Mills RE;Gerstein MB;Bashir A;Stegle O;Devine SE;Lee C;Eichler EE;Korbel JO
通讯作者: Korbel JO
DOI: 10.1186/gb-2010-11-5-r52
发表时间: 2010-01-01
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Pang, Andy W.;MacDonald, Jeffrey R.;Scherer, Stephen W.
通讯作者: Scherer, Stephen W.
使用下一代 DNA 测序数据进行变异发现和基因分型的框架。
DOI: 10.1038/ng.806
发表时间: 2011-05
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1093/bioinformatics/btp698
发表时间: 2010-03-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Li H;Durbin R
通讯作者: Durbin R