TEclass-a tool for automated classification of unknown eukaryotic transposable elements

TEclass-a tool for automated classification of unknown eukaryotic transposable elements
复制标题

DOI:
10.1093/bioinformatics/btp084
复制
发表时间:
2009-05-15
期刊:
影响因子:
5.8
通讯作者:
Makalowski, Wojciech
Makalowski, Wojciech
中科院分区:
生物学3区
文献类型:
--
作者:
Abrusan, Gyorgy;Grundmann, Norbert;Makalowski, Wojciech

文献摘要

被引文献

相似文献

动机:大量测序的基因组需要开发软件来重建转座子和其他重复元件的共有序列。然而,现有的工具通常集中在原始重复序列的准确识别,并提供没有关于重建的一致性的分类位置的信息。TEclass是一个将未知转座因子分为四个主要功能类别的工具,这些类别反映了它们的转座模式:DNA转座子,长末端重复序列(LTR),长散布核元件(LINES)和短散布核元件(西内斯)。TEclass使用机器学习支持向量机(SVM)进行基于低聚物频率的分类。它在新的DNA和LTR重复序列的分类中达到90-97%的准确率,在LINE和西内斯的分类中达到75%。
Motivation: The large number of sequenced genomes required the development of software that reconstructs the consensus sequences of transposons and other repetitive elements. However, the available tools usually focus on the accurate identification of raw repeats and provide no information about the taxonomic position of the reconstructed consensi. TEclass is a tool to classify unknown transposable elements into their four main functional categories, which reflect their mode of transposition: DNA transposons, long terminal repeats (LTRs), long interspersed nuclear elements (LINEs) and short interspersed nuclear elements (SINEs). TEclass uses machine learning support vector machine (SVM) for classification based on oligomer frequencies. It achieves 90-97% accuracy in the classification of novel DNA and LTR repeats, and 75% for LINEs and SINEs.