ExplorATE: a new pipeline to explore active transposable elements from RNA-seq data

ExplorATE: a new pipeline to explore active transposable elements from RNA-seq data
复制标题

ExplorATE:从 RNA-seq 数据中探索活性转座元件的新管道

DOI:
10.1093/bioinformatics/btac354
复制
发表时间:
2022
期刊:
影响因子:
5.8
通讯作者:
Kendziorski, ed., Christina
Kendziorski, ed., Christina
中科院分区:
生物学3区
文献类型:
--
作者:
Femenias, Martin M.;Santos, Juan C.;Sites, Jr, Jack W.;Avila, Luciano J.;Morando, Mariana;Kendziorski, ed., Christina

文献摘要

参考文献

相似文献

转座因子(TE)在基因组中普遍存在,并且许多保持活性。TE包含转录组的重要部分,其通过产生有害突变或促进进化新颖性而对宿主基因组具有潜在影响。然而,他们的功能研究是有限的,在他们的识别和量化的困难,特别是在非模式organisation.ResultsWe开发了一个新的管道[探索活性转座因子(ExplorATE)]在R和巴什,允许在模型和非模型生物体中的活性TE的定量实施。ExplorATE创建TE特异性索引,并使用选择性比对(SA)根据比对分数过滤出基因内的共转录转座子。此外,我们的软件采用了一个类似Wicker的标准,以完善一组目标TE,并避免虚假的映射。模拟和真实的数据的基础上,我们表明,SA策略采用ExplorATE实现了更好的估计非共转录的元素比其他可用的基于标记或基于映射的软件。ExplorATE结果显示,在有和没有参考基因组的情况下,与基于基因组的工具高度一致,但ExplorATE需要更少的执行时间。同样,ExplorATE通过在量化过程中整合共转录和多重映射效应,扩展和补充了大多数以前的TE分析,并提供了与R环境中其他下游工具的无缝集成。可用性和实施源代码可在https://github.com/FemeniasM/ExplorATEproject和https://github.com/FemeniasM/ExplorATE_shell_script上获得。补充信息补充数据可在Bioinformatics online获取。
MotivationTransposable elements (TEs) are ubiquitous in genomes and many remain active. TEs comprise an important fraction of the transcriptomes with potential effects on the host genome, either by generating deleterious mutations or promoting evolutionary novelties. However, their functional study is limited by the difficulty in their identification and quantification, particularly in non-model organisms.ResultsWe developed a new pipeline [explore active transposable elements (ExplorATE)] implemented in R and bash that allows the quantification of active TEs in both model and non-model organisms. ExplorATE creates TE-specific indexes and uses the Selective Alignment (SA) to filter out co-transcribed transposons within genes based on alignment scores. Moreover, our software incorporates a Wicker-like criteria to refine a set of target TEs and avoid spurious mapping. Based on simulated and real data, we show that the SA strategy adopted by ExplorATE achieved better estimates of non-co-transcribed elements than other available alignment-based or mapping-based software. ExplorATE results showed high congruence with alignment-based tools with and without a reference genome, yet ExplorATE required less execution time. Likewise, ExplorATE expands and complements most previous TE analyses by incorporating the co-transcription and multi-mapping effects during quantification, and provides a seamless integration with other downstream tools within the R environment.Availability and implementationSource code is available at https://github.com/FemeniasM/ExplorATEproject and https://github.com/FemeniasM/ExplorATE_shell_script. Data available on request.Supplementary informationSupplementary data are available atBioinformaticsonline.
DOI: 10.1093/gbe/evaa068
发表时间: 2020-04
影响因子: 3.3
作者:
G. M. Pasquesi;B. W. Perry;Michael W. Vandewege;Robert Ruggiero;Drew R. Schield;T. Castoe
通讯作者: G. M. Pasquesi;B. W. Perry;Michael W. Vandewege;Robert Ruggiero;Drew R. Schield;T. Castoe
重新审视用于量化转录本丰度的 RSEM 生成模型及其 EM 算法
DOI: 10.1101/503672
发表时间: 2018
期刊: bioRxiv
影响因子: --
作者:
Hy Vuong;Thao T. Truong;Thang N Tran;Son K. Pham
通讯作者: Son K. Pham
DOI: 10.1186/s13059-019-1905-y
发表时间: 2019-12-16
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Ou, Shujun;Su, Weija;Hufford, Matthew B.
通讯作者: Hufford, Matthew B.