The repetitive landscape of the chicken genome

The repetitive landscape of the chicken genome
复制标题

DOI:
10.1101/gr.2438005
复制
发表时间:
2005-01-01
期刊:
影响因子:
7
通讯作者:
Ivarie, R
Ivarie, R
中科院分区:
生物学1区
文献类型:
--
作者:
Wicker, T;Robertson, JS;Ivarie, R

文献摘要

被引文献

相似文献

基于cot的克隆和测序(CBCS)是分离和表征任何基因组的各种重复成分的强大工具,将DNA重关联动力学的既定原则与高通量测序相结合。CBCS用于生成鸡基因组高拷贝、中拷贝和低拷贝部分的序列文库。对鸡的高拷贝DNA进行测序,使其覆盖范围达到其估计序列复杂性的2.7倍,从而初步确定了几个新的重复家族,然后将其用于新发布的鸡全基因组初稿的调查。该分析提供了对已知重复结构(如CRI和CNM)的多样性和生物学的深入了解,此前只有有限的序列数据可用。Cot序列数据还鉴定了四个新的重复序列(Birddawg, Hitchcock, Kronos和Soprano),两个新的CRI重复亚家族,以及鸡基因组组装中缺失的许多元件。除了非自治缺失衍生物外,还发现了一个新的水手样转座子Galluhop的多个自治元件。高拷贝重复序列CRI、Galluhop和Birddawg的系统发育分析提供了两种不同的基因组分散策略。本研究还举例说明了CBCS方法在为只有有限序列数据可用的基因组重复部分创建代表性数据库方面的能力。
Cot-based cloning and sequencing (CBCS) is a powerful tool for isolating and characterizing the various repetitive components of any genome, combining the established principles of DNA reassociation kinetics with high-throughput sequencing. CBCS was used to generate sequence libraries representing the high, middle, and low-copy fractions of the chicken genome. Sequencing high-copy DNA of chicken to about 2.7x coverage of its estimated sequence complexity led to the initial identification of several new repeat families, which were then used for a survey of the newly released first draft of the complete chicken genome. The analysis provided insight into the diversity and biology of known repeat structures such as CRI and CNM, for which only limited sequence data had previously been available. Cot sequence data also resulted in the identification of four novel repeats (Birddawg, Hitchcock, Kronos, and Soprano), two new subfamilies of CRI repeats, and many elements absent from the chicken genome assembly. Multiple autonomous elements were found for a novel Mariner-like transposon, Galluhop, in addition to nonautonomous deletion derivatives. Phylogenetic analysis of the high-copy repeats CRI, Galluhop, and Birddawg provided insight into two distinct genome dispersion strategies. This study also exemplifies the power of the CBCS method to create representative databases for the repetitive fractions of genomes for which only limited sequence data is available.