Bioinformatics and genomic analysis of transposable elements in eukaryotic genomes

Bioinformatics and genomic analysis of transposable elements in eukaryotic genomes
复制标题

DOI:
10.1007/s10577-011-9230-7
复制
发表时间:
2011-08-01
影响因子:
2.6
通讯作者:
Yang, Guojun
Yang, Guojun
中科院分区:
生物学2区
文献类型:
--
作者:
Janicki, Mateusz;Rooke, Rebecca;Yang, Guojun

文献摘要

被引文献

相似文献

大多数真核生物基因组的主要部分是转座因子(TE)。在进化过程中,TE对基因组的大小、结构和功能产生了深刻的变化。作为基因组的组成部分,TE的动态存在将继续成为重塑基因组的主要力量。早期对基因组序列中TE的计算分析集中在过滤掉“垃圾”序列以促进基因注释。当真核生物基因组中TE的高丰度和多样性被认识到时,这些早期的努力转变为系统的全基因组范围的TE分类和分类。基因组序列数据的可用性逆转了发现新的TE家族和超家族的经典遗传方法。精心策划的TE数据库及其对基因组序列的准确注释反过来又促进了TE在许多前沿领域的研究,包括:(1)TE介导的基因组大小和结构的变化,(2)TE对基因组和基因功能的影响,(3)宿主对TE的调控,(4)TE的进化及其种群动态,以及(5)TE活性的基因组规模研究。生物信息学和基因组学方法已成为大规模TE研究的一个组成部分,可以通过纯计算机分析提取信息或辅助湿实验室实验研究。目前基因组测序技术的革命促进了现有研究前沿的进一步进展和新举措的出现。常规基础上以创纪录的低成本快速生成大序列数据集对计算行业的存储容量和操作速度以及生物信息学社区的算法及其实现的改进提出了挑战。
A major portion of most eukaryotic genomes are transposable elements (TEs). During evolution, TEs have introduced profound changes to genome size, structure, and function. As integral parts of genomes, the dynamic presence of TEs will continue to be a major force in reshaping genomes. Early computational analyses of TEs in genome sequences focused on filtering out "junk" sequences to facilitate gene annotation. When the high abundance and diversity of TEs in eukaryotic genomes were recognized, these early efforts transformed into the systematic genome-wide categorization and classification of TEs. The availability of genomic sequence data reversed the classical genetic approaches to discovering new TE families and super-families. Curated TE databases and their accurate annotation of genome sequences in turn facilitated the studies on TEs in a number of frontiers including: (1) TE-mediated changes of genome size and structure, (2) the influence of TEs on genome and gene functions, (3) TE regulation by host, (4) the evolution of TEs and their population dynamics, and (5) genomic scale studies of TE activity. Bioinformatics and genomic approaches have become an integral part of large-scale studies on TEs to extract information with pure in silico analyses or to assist wet lab experimental studies. The current revolution in genome sequencing technology facilitates further progress in the existing frontiers of research and emergence of new initiatives. The rapid generation of large-sequence datasets at record low costs on a routine basis is challenging the computing industry on storage capacity and manipulation speed and the bioinformatics community for improvement in algorithms and their implementations.