Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline

Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline
复制标题

DOI:
10.1186/s13059-019-1905-y
复制
发表时间:
2019-12-16
期刊:
影响因子:
12.3
通讯作者:
Hufford, Matthew B.
Hufford, Matthew B.
中科院分区:
生物学1区
文献类型:
--
作者:
Ou, Shujun;Su, Weija;Hufford, Matthew B.

文献摘要

被引文献

相似文献

背景资料:测序技术和组装算法已经成熟到可以对大的重复基因组进行高质量的从头组装的程度。目前的汇编遍历转座因子(TE),并提供了一个机会,全面注释的TE。存在许多方法来注释每一类TE,但它们的相对性能还没有被系统地比较。此外,需要一个全面的管道,以产生一个非冗余的TEs库的物种缺乏这种资源,以产生全基因组TE annotations.Results:我们基准现有的程序的基础上精心策划的水稻TEs库。我们评估了注释长末端重复序列(LTR)反转录转座子、末端反向重复序列(TIR)转座子、短TIR转座子(称为微型反向转座因子(MITE))和Helitrons的方法的性能。性能指标包括灵敏度、特异性、准确度、精密度、FDR和F1。使用最强大的程序,我们创建了一个名为Extensive de-novo TE Annotator(EDTA)的综合管道,该管道可以产生一个过滤的非冗余TE库,用于注释结构完整和碎片化的元素。EDTA还对在高度重复的基因组区域中经常发现的嵌套TE插入进行解卷积。使用其他模式物种与策划TE库(玉米和果蝇),EDTA是强大的跨植物和动物species.Conclusions:基准测试结果和管道开发在这里将大大促进TE注释真核基因组。这些注释将促进更深入地了解的多样性和进化的TE在物种内和物种间的水平。EDTA是开源的,可免费获得:https://github.com/oushujun/EDTA。
Background: Sequencing technology and assembly algorithms have matured to the point that high-quality de novo assembly is possible for large, repetitive genomes. Current assemblies traverse transposable elements (TEs) and provide an opportunity for comprehensive annotation of TEs. Numerous methods exist for annotation of each class of TEs, but their relative performances have not been systematically compared. Moreover, a comprehensive pipeline is needed to produce a non-redundant library of TEs for species lacking this resource to generate whole-genome TE annotations.Results: We benchmark existing programs based on a carefully curated library of rice TEs. We evaluate the performance of methods annotating long terminal repeat (LTR) retrotransposons, terminal inverted repeat (TIR) transposons, short TIR transposons known as miniature inverted transposable elements (MITEs), and Helitrons. Performance metrics include sensitivity, specificity, accuracy, precision, FDR, and F1. Using the most robust programs, we create a comprehensive pipeline called Extensive de-novo TE Annotator (EDTA) that produces a filtered non-redundant TE library for annotation of structurally intact and fragmented elements. EDTA also deconvolutes nested TE insertions frequently found in highly repetitive genomic regions. Using other model species with curated TE libraries (maize and Drosophila), EDTA is shown to be robust across both plant and animal species.Conclusions: The benchmarking results and pipeline developed here will greatly facilitate TE annotation in eukaryotic genomes. These annotations will promote a much more in-depth understanding of the diversity and evolution of TEs at both intra- and inter-species levels. EDTA is open-source and freely available: https://github.com/oushujun/EDTA.