MITE Tracker: an accurate approach to identify miniature inverted-repeat transposable elements in large genomes.

MITE Tracker: an accurate approach to identify miniature inverted-repeat transposable elements in large genomes.
复制标题

DOI:
10.1186/s12859-018-2376-y
复制
发表时间:
2018-10-03
期刊:
影响因子:
3
通讯作者:
Vanzetti LS
Vanzetti LS
中科院分区:
生物学4区
文献类型:
--
作者:
Crescente JM;Zavallo D;Helguera M;Vanzetti LS

文献摘要

参考文献

被引文献

相似文献

微型反向重复转座元件 (MITE) 是短的、非自主的 II 类转座元件,存在于真核生物基因组中的大量保守拷贝中。准确识别这些元件有助于揭示控制基因组进化和基因调控的机制。这些元素的结构和分布是明确定义的,因此可以使用计算方法来识别 MITE 序列。在这里,我们描述了 MITE Tracker,这是一种新颖的开源软件程序,它使用有效的比对策略从大型基因组中检索附近的反向重复序列来查找和分类 MITE。该程序使用快速聚类算法将它们分组为高序列同源性家族,并最终仅过滤那些可能从不同基因组位置转置的元素,因为它们的侧翼序列比对得分较低。人们已经提出了许多计划来寻找隐藏在基因组中的 MITE。然而,它们都无法处理大规模基因组,例如面包小麦的基因组。此外,在许多情况下,现有方法的假阳性率(或漏检率)很高。水稻基因组被用作参考,将 MITE Tracker 与已知工具进行比较。我们的方法在我们的测试中被证明是最可靠的。事实上,它揭示了更多的已知元素,呈现出最低的假阳性数量,并且是唯一能够以面包小麦基因组作为输入运行的程序。在小麦中,MITE Tracker 发现了 6013 个 MITE 家族,并首次在完整面包小麦基因组中探索了 MITE 的结构。本文的在线版本 (10.1186/s12859-018-2376-y) 包含补充材料,可供授权用户使用。
Miniature inverted-repeat transposable elements (MITEs) are short, non-autonomous class II transposable elements present in a high number of conserved copies in eukaryote genomes. An accurate identification of these elements can help to shed light on the mechanisms controlling genome evolution and gene regulation. The structure and distribution of these elements are well-defined and therefore computational approaches can be used to identify MITEs sequences. Here we describe MITE Tracker, a novel, open source software program that finds and classifies MITEs using an efficient alignment strategy to retrieve nearby inverted-repeat sequences from large genomes. This program groups them into high sequence homology families using a fast clustering algorithm and finally filters only those elements that were likely transposed from different genomic locations because of their low scoring flanking sequence alignment. Many programs have been proposed to find MITEs hidden in genomes. However, none of them are able to process large-scale genomes such as that of bread wheat. Furthermore, in many cases the existing methods perform high false-positive rates (or miss rates). The rice genome was used as reference to compare MITE Tracker against known tools. Our method turned out to be the most reliable in our tests. Indeed, it revealed more known elements, presented the lowest false-positive number and was the only program able to run with the bread wheat genome as input. In wheat, MITE Tracker discovered 6013 MITE families and allowed the first structural exploration of MITEs in the complete bread wheat genome. The online version of this article (10.1186/s12859-018-2376-y) contains supplementary material, which is available to authorized users.
DOI: 10.1093/nar/gks1189
发表时间: 2013-01
影响因子: 14.9
作者:
NCBI Resource Coordinators
通讯作者: NCBI Resource Coordinators
DOI: 10.1159/000084979
发表时间: 2005-01-01
影响因子: 1.7
作者:
Jurka, J;Kapitonov, VV;Walichiewicz, J
通讯作者: Walichiewicz, J
DOI: 10.1186/s13100-015-0041-9
发表时间: 2015
期刊: Mobile DNA
影响因子: 4.9
作者:
Bao W;Kojima KK;Kohany O
通讯作者: Kohany O
Biopython:用于计算分子生物学和生物信息学的免费 Python 工具。
DOI: 10.1093/bioinformatics/btp163
发表时间: 2009-06-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Cock PJ;Antao T;Chang JT;Chapman BA;Cox CJ;Dalke A;Friedberg I;Hamelryck T;Kauff F;Wilczynski B;de Hoon MJ
通讯作者: de Hoon MJ
DOI: 10.1093/bioinformatics/bts565
发表时间: 2012-12-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
通讯作者: Li W