Segmental duplications: Organization and impact within the current Human Genome Project assembly

Segmental duplications: Organization and impact within the current Human Genome Project assembly
复制标题

DOI:
10.1101/gr.gr-1871r
复制
发表时间:
2001-06-01
期刊:
影响因子:
7
通讯作者:
Eichler, EE
Eichler, EE
中科院分区:
生物学1区
文献类型:
--
作者:
Bailey, JA;Yavor, AM;Eichler, EE

文献摘要

被引文献

相似文献

节段性重复在基因组疾病和基因进化中起着重要作用。为了了解它们在人类基因组中的组织,我们开发了必要的计算工具和方法来检测长片段基因组序列之间的同一性,尽管存在高拷贝重复序列和大的插入缺失。在这里,我们提出了我们的分析,最近的基因组组装(2001年1月),我们专注于全球组织的这些部分和它们在全基因组组装过程中发挥的作用。最初,我们只考虑了最近的大的重复事件,其远低于草图测序误差的水平(比对90%-98%相似并且长度大于或等于1 kb)。重复序列(90%-98%;大于或等于1 kb)占所有人类序列的3.6%。这些重复显示聚类和高达10倍的富集内pericentromeric和subtelomeric区域。在组装方面,发现重复序列在无序和未分配的重叠群中过度表达,表明重复序列难以分配到其正确位置。为了评估基因组内这些区域的覆盖率,我们选择了含有染色体间重复的BAC,并通过FISH表征了它们的重复模式。FISH阳性的染色体中只有47%(106/224)的染色体在BLAST比对中有相应的染色体位置。我们目前的数据表明,这是由于错误组装,错误分配,和/或重复区域内的测序覆盖率下降。令人惊讶的是,如果我们认为推定的重复> 98%的同一性,我们将10.6%(286 Mb)的当前组装体鉴定为旁系同源的。我们相信,这些序列中的大多数代表了独特区域内未合并的重叠。综上所述,上述数据表明,片段性重复是精确人类基因组组装的重大障碍,需要开发专门的技术来完成基因组的这些特殊区域。这些高度重复区域的鉴定和表征代表了人类参考基因组的完整测序中的重要步骤。
Segmental duplications play fundamental roles in both genomic disease and gene evolution. To understand their organization within the human genome, we have developed the computational tools and methods necessary to detect identity between long stretches of genomic sequence despite the presence of high copy repeats and large insertion-deletions. Here we present our analysis of the most recent genome assembly (January 2001) in which we focus on the global organization of these segments and the role they play in the whole-genome assembly process. initially, we considered only large recent duplication events that fell well-below levels of draft sequencing error (alignments 90%-98% similar and greater than or equal to1 kb in length). Duplications (90%-98%; greater than or equal to1 kb) comprise 3.6% of all human sequence. These duplications show clustering and up to 10-fold enrichment within pericentromeric and subtelomeric regions. In terms of assembly, duplicated sequences were found to be over-represented in unordered and unassigned contigs indicating that duplicated sequences are difficult to assign to their proper position. To assess coverage of these regions within the genome, we selected BACs containing interchromosomal duplications and characterized their duplication pattern by FISH. Only 47% (106/224) of chromosomes positive by FISH had a corresponding chromosomal position by BLAST comparison. We present data that indicate that this is attributable to misassembly, misassignment, and/or decreased sequencing coverage within duplicated regions. Surprisingly, if we consider putative duplications > 98% identity, we identify 10.6% (286 Mb) of the current assembly as paralogous. The majority of these alignments, we believe, represent unmerged overlaps within unique regions. Taken together the above data indicate that segmental duplications represent a significant impediment to accurate human genome assembly, requiring the development of specialized techniques to finish these exceptional regions of the genome. The identification and characterization of these highly duplicated regions represents an important step in the complete sequencing of a human reference genome.