Red: an intelligent, rapid, accurate tool for detecting repeats de-novo on the genomic scale.

Red: an intelligent, rapid, accurate tool for detecting repeats de-novo on the genomic scale.
复制标题

DOI:
10.1186/s12859-015-0654-5
复制
发表时间:
2015-07-24
期刊:
影响因子:
3
通讯作者:
Girgis HZ
Girgis HZ
中科院分区:
生物学4区
文献类型:
--
作者:
Girgis HZ

文献摘要

被引文献

相似文献

随着技术的快速进步,成千上万个物种的基因组序列正在变得可用。在序列内是包含基因组的重要部分的重复。因此,成功的注释需要准确发现重复序列。重复序列作为物种特异性元件,在新测序的基因组中很可能是未知的。因此,注释新测序的基因组需要工具来从头发现重复序列。然而,目前可用的从头诊断工具在输入序列的大小、易用性、对主要类型的重复的敏感性、性能的一致性、速度和假阳性率方面存在局限性。为了解决这些限制,我设计并开发了Red,应用了机器学习。Red是第一个能够标记其训练数据并在整个基因组上自动训练自己的重复检测工具。红色易于安装和使用。它对转座子和简单重复序列都敏感;相反,RepeatScout和ReCon等可用工具对转座子敏感,WindowMasker对简单重复序列敏感。Red在7个基因组上表现良好;其他工具仅在某些基因组上表现良好。Red比RepeatScout和ReCon快得多,并且误报率比WindowMasker低得多。在具有五个或更多拷贝的人类基因上,Red比RepeatScout更具有特异性。当对不寻常核苷酸组成的基因组进行测试时,Red以高灵敏度定位重复序列,并保持适度的假阳性率。Red在细菌基因组上的表现优于相关工具。Red在人类基因组中发现了46,405个新的重复片段。最后,Red能够处理组装和未组装的基因组。Red的创新方法及其在七种不同基因组上的出色表现代表了重复序列发现领域的宝贵进步。本文的在线版本(doi:10.1186/s12859-015-0654-5)包含补充材料,可供授权用户使用。
With rapid advancements in technology, the sequences of thousands of species’ genomes are becoming available. Within the sequences are repeats that comprise significant portions of genomes. Successful annotations thus require accurate discovery of repeats. As species-specific elements, repeats in newly sequenced genomes are likely to be unknown. Therefore, annotating newly sequenced genomes requires tools to discover repeats de-novo. However, the currently available de-novo tools have limitations concerning the size of the input sequence, ease of use, sensitivities to major types of repeats, consistency of performance, speed, and false positive rate. To address these limitations, I designed and developed Red, applying Machine Learning. Red is the first repeat-detection tool capable of labeling its training data and training itself automatically on an entire genome. Red is easy to install and use. It is sensitive to both transposons and simple repeats; in contrast, available tools such as RepeatScout and ReCon are sensitive to transposons, and WindowMasker to simple repeats. Red performed consistently well on seven genomes; the other tools performed well only on some genomes. Red is much faster than RepeatScout and ReCon and has a much lower false positive rate than WindowMasker. On human genes with five or more copies, Red was more specific than RepeatScout by a wide margin. When tested on genomes of unusual nucleotide compositions, Red located repeats with high sensitivities and maintained moderate false positive rates. Red outperformed the related tools on a bacterial genome. Red identified 46,405 novel repetitive segments in the human genome. Finally, Red is capable of processing assembled and unassembled genomes. Red’s innovative methodology and its excellent performance on seven different genomes represent a valuable advancement in the field of repeats discovery. The online version of this article (doi:10.1186/s12859-015-0654-5) contains supplementary material, which is available to authorized users.