Improved detection and annotation of transposable elements in sequenced genomes using multiple reference sequence sets

Improved detection and annotation of transposable elements in sequenced genomes using multiple reference sequence sets
复制标题

DOI:
10.1016/j.ygeno.2008.01.005
复制
发表时间:
2008-05-01
期刊:
影响因子:
4.4
通讯作者:
Colot, Vincent
Colot, Vincent
中科院分区:
生物学3区
文献类型:
--
作者:
Buisine, Nicolas;Quesneville, Hadi;Colot, Vincent

文献摘要

被引文献

相似文献

转座因子(te)是真核生物基因组中普遍存在的成分,影响着基因组功能的许多方面。基因组序列中的TE检测通常使用对先前确定的TE构建的一组参考序列的相似性搜索来执行。通过设计包含te结构和进化关键方面的参考集,并将这些参考集与主要由共识序列组成的Repbase Update (RU)相结合,我们证明了这一过程可以得到改进。使用拟南芥基因组作为测试案例,我们的方法可以检测到额外的12.4%的TE序列。这些对应于新的TE片段以及RU已经检测到的TE片段的扩展。值得注意的是,我们发现仅使用两个参考集就可以很容易地优化TE检测,一个包含真正的共识序列,另一个包含捕获家族内TE拷贝结构多样性的马赛克序列。(c) 2008爱思唯尔公司版权所有。
Transposable elements (TEs) are ubiquitous components of eukaryotic genomes that impact many aspects of genome function. TE detection in genomic sequences is typically performed using similarity searches against a set of reference sequences built from previously identified TEs. Here, we demonstrate that this process can be improved by designing reference sets that incorporate key aspects of the structure and evolution of TEs and by combining these sets with Repbase Update (RU), which is composed mainly of consensus sequences. Using the Arabidopsis genome as a test case, our approach leads to the detection of an extra 12.4% of TE sequences. These correspond to novel TE fragments as well as to the extension of TE fragments already detected by RU. Significantly, we find that TE detection could be readily optimized using only two reference sets, one containing true consensus sequences and the other mosaic sequences that capture the structural diversity of TE copies within a family. (c) 2008 Elsevier Inc. All rights reserved.