The diploid genome sequence of an individual human.

The diploid genome sequence of an individual human.
复制标题

DOI:
10.1371/journal.pbio.0050254
复制
发表时间:
2007-09-04
期刊:
影响因子:
9.8
通讯作者:
Venter JC
Venter JC
中科院分区:
生物学1区
文献类型:
--
作者:
Levy S;Sutton G;Ng PC;Feuk L;Halpern AL;Walenz BP;Axelrod N;Huang J;Kirkness EF;Denisov G;Lin Y;MacDonald JR;Pang AW;Shago M;Stockwell TB;Tsiamouri A;Bafna V;Bansal V;Kravitz SA;Busam DA;Beeson KY;McIntosh TC;Remington KA;Abril JF;Gill J;Borman J;Rogers YH;Frazier ME;Scherer SW;Strausberg RL;Venter JC

文献摘要

被引文献

相似文献

这里展示的是一个人的基因组序列。它是从∼3200万个随机DNA片段中产生的,通过桑格双脱氧技术进行测序,并组装成4,528个支架,包括28.1亿个碱基(Mb)的连续序列,对任何给定的区域覆盖约7.5倍。我们开发了一个改进版本的Celera汇编器,以便于识别和比较这个个体二倍体基因组中的替代等位基因。将该基因组与国家生物技术信息中心人类参考组件进行比较,发现了超过410万个DNA变异,包括12.3Mb。这些变异包括3,213,401个单核苷酸多态(SNPs),53,823个区块替换(2-206bp),292,102个杂合插入/缺失事件(Indels)(1-571bp),559,473个纯合子Indels(1-82,711bp),90个倒位,以及大量的片段重复和拷贝数变异区。在供体中发现的所有事件中,非SNP DNA变异占22%,但涉及所有变异碱基的74%。这表明非SNP基因改变在定义二倍体基因组结构中起着重要作用。此外,44%的基因是一个或多个变种的杂合子。使用一种新的单倍型组装策略,我们能够在200 kb的片段中跨越1.5 GB的基因组序列,为基因组的二倍体性质提供了进一步的精确度。这些数据描绘了二倍体人类基因组的明确分子肖像,为未来的基因组比较提供了起点,并使个性化基因组信息时代成为可能。我们已经从单个个体的两条染色体中产生了独立组装的二倍体人类基因组DNA序列(J.克雷格·文特尔)。我们的方法,基于全基因组鸟枪测序,使用增强的基因组组装策略和软件,生成了组装的基因组,其中超过一半的基因组代表在大的二倍体片段(>200kbase)中,使二倍体基因组的研究成为可能。与以前参考的人类基因组序列(由多个人组成的复合体)相比,发现大多数基因组改变是基于单核苷酸(SNPs)的、研究得很好的一类变体。然而,结果也显示,较少研究的基因组变异、插入和缺失,虽然只占基因组变异事件的一小部分(22%),但实际上占变异核苷酸的近74%。将插入和缺失遗传变异包括在我们对染色体间差异的估计中,表明个体的两个染色体副本之间只有99.5%的相似性,两个个体之间的遗传变异比先前估计的高出五倍之多。一个特征明确的二倍体人类基因组序列的存在为未来的个体基因组比较提供了一个起点,并使个性化基因组信息时代的到来成为可能。将人类个体的DNA序列与参考序列进行比较,发现了惊人的差异。
Presented here is a genome sequence of an individual human. It was produced from ∼32 million random DNA fragments, sequenced by Sanger dideoxy technology and assembled into 4,528 scaffolds, comprising 2,810 million bases (Mb) of contiguous sequence with approximately 7.5-fold coverage for any given region. We developed a modified version of the Celera assembler to facilitate the identification and comparison of alternate alleles within this individual diploid genome. Comparison of this genome and the National Center for Biotechnology Information human reference assembly revealed more than 4.1 million DNA variants, encompassing 12.3 Mb. These variants (of which 1,288,319 were novel) included 3,213,401 single nucleotide polymorphisms (SNPs), 53,823 block substitutions (2–206 bp), 292,102 heterozygous insertion/deletion events (indels)(1–571 bp), 559,473 homozygous indels (1–82,711 bp), 90 inversions, as well as numerous segmental duplications and copy number variation regions. Non-SNP DNA variation accounts for 22% of all events identified in the donor, however they involve 74% of all variant bases. This suggests an important role for non-SNP genetic alterations in defining the diploid genome structure. Moreover, 44% of genes were heterozygous for one or more variants. Using a novel haplotype assembly strategy, we were able to span 1.5 Gb of genome sequence in segments >200 kb, providing further precision to the diploid nature of the genome. These data depict a definitive molecular portrait of a diploid human genome that provides a starting point for future genome comparisons and enables an era of individualized genomic information. We have generated an independently assembled diploid human genomic DNA sequence from both chromosomes of a single individual (J. Craig Venter). Our approach, based on whole-genome shotgun sequencing and using enhanced genome assembly strategies and software, generated an assembled genome over half of which is represented in large diploid segments (>200 kilobases), enabling study of the diploid genome. Comparison with previous reference human genome sequences, which were composites comprising multiple humans, revealed that the majority of genomic alterations are the well-studied class of variants based on single nucleotides (SNPs). However, the results also reveal that lesser-studied genomic variants, insertions and deletions, while comprising a minority (22%) of genomic variation events, actually account for almost 74% of variant nucleotides. Inclusion of insertion and deletion genetic variation into our estimates of interchromosomal difference reveals that only 99.5% similarity exists between the two chromosomal copies of an individual and that genetic variation between two individuals is as much as five times higher than previously estimated. The existence of a well-characterized diploid human genome sequence provides a starting point for future individual genome comparisons and enables the emerging era of individualized genomic information. Comparison of the DNA sequence of an individual human from the reference sequence reveals a surprising amount of difference.