RepeatExplorer: a Galaxy-based web server for genome-wide characterization of eukaryotic repetitive elements from next-generation sequence reads

RepeatExplorer: a Galaxy-based web server for genome-wide characterization of eukaryotic repetitive elements from next-generation sequence reads
复制标题

DOI:
10.1093/bioinformatics/btt054
复制
发表时间:
2013-03-15
期刊:
影响因子:
5.8
通讯作者:
Macas, Jiri
Macas, Jiri
中科院分区:
生物学3区
文献类型:
--
作者:
Novak, Petr;Neumann, Pavel;Macas, Jiri

文献摘要

被引文献

相似文献

动机:重复DNA占植物和动物核基因组的很大一部分,但它仍然是迄今为止研究的大多数物种中特征最少的基因组组成部分。尽管最近高通量测序数据的可获得性为深入研究基因组重复序列提供了必要的资源,但由于缺乏专门的生物信息学工具和适当的计算资源,使得大规模重复分析能够由面向生物学的研究人员进行,其实用性受到阻碍。结果:在这里,我们介绍了RepeatExplorer,这是一个用于表征重复元件的软件工具的集合,可以通过网络界面访问。服务器的一个关键组件是使用基于图的序列聚类算法的计算管道,以促进从头开始重复识别,而不需要已知元素的参考数据库。由于该算法使用从基因组中随机采样的短序列作为输入,因此它是分析下一代序列读取的理想选择。还提供了其他工具,以帮助对已识别的重复序列进行分类,调查逆转录元件的系统发育关系,并对多个物种之间的重复序列组成进行比较分析。该服务器允许分析数百万次序列读取,这通常会导致识别高等植物基因组中的大多数高拷贝和中拷贝重复。实现和可用性:RepeatExplorer在Galaxy环境中实现,并设置在http://repeatexplorer.umbr.cas.cz/.的公共服务器上有关本地安装的源代码和说明,请访问http://w3lamc.umbr.cas.cz/lamc/resources.php.Contact:macas@umbr.Cas.cz
Motivation: Repetitive DNA makes up large portions of plant and animal nuclear genomes, yet it remains the least-characterized genome component in most species studied so far. Although the recent availability of high-throughput sequencing data provides necessary resources for in-depth investigation of genomic repeats, its utility is hampered by the lack of specialized bioinformatics tools and appropriate computational resources that would enable large-scale repeat analysis to be run by biologically oriented researchers.Results: Here we present RepeatExplorer, a collection of software tools for characterization of repetitive elements, which is accessible via web interface. A key component of the server is the computational pipeline using a graph-based sequence clustering algorithm to facilitate de novo repeat identification without the need for reference databases of known elements. Because the algorithm uses short sequences randomly sampled from the genome as input, it is ideal for analyzing next-generation sequence reads. Additional tools are provided to aid in classification of identified repeats, investigate phylogenetic relationships of retroelements and perform comparative analysis of repeat composition between multiple species. The server allows to analyze several million sequence reads, which typically results in identification of most high and medium copy repeats in higher plant genomes.Implementation and availability: RepeatExplorer was implemented within the Galaxy environment and set up on a public server at http://repeatexplorer.umbr.cas.cz/. Source code and instructions for local installation are available at http://w3lamc.umbr.cas.cz/lamc/resources.php.Contact: macas@umbr.cas.cz