TSEBRA: transcript selector for BRAKER.

TSEBRA: transcript selector for BRAKER.
复制标题

DOI:
10.1186/s12859-021-04482-0
复制
发表时间:
2021-11-25
期刊:
影响因子:
3
通讯作者:
Stanke M
Stanke M
中科院分区:
生物学4区
文献类型:
--
作者:
Gabriel L;Hoff KJ;Brůna T;Borodovsky M;Stanke M

文献摘要

参考文献

被引文献

相似文献

BRAKER是一套自动管道,BRAKER1和BRAKER2,用于准确注释真核生物基因组中的蛋白质编码基因。每个管道根据提供的证据训练蛋白质编码基因的统计模型,然后使用外部证据和统计模型预测基因组序列中的蛋白质编码基因。对于训练和预测,BRAKER1和BRAKER2结合了互补的外部证据:BRAKER1仅使用RNA-seq数据,而BRAKER2仅使用跨物种蛋白质数据库。到目前为止,当同时结合两种类型的证据时,BRAKER套件还不能可靠地超过BRAKER1和BRAKER2的准确性。目前,对于RNA-seq和蛋白质数据都可用的新基因组计划,最好的选择是独立运行两个管道,并选择一个可能更好的输出。因此,一种或另一种类型的外部证据仍未被利用。我们介绍了TSEBRA,一个从BRAKER1和BRAKER2产生的集合中选择基因预测(转录本)的软件。TSEBRA使用一套规则来比较基于RNA-seq和同源蛋白证据支持的重叠转录本的分数。我们在11个物种基因组的计算实验中表明,TSEBRA比单独运行的BRAKER1或BRAKER2获得更高的准确性,并且与组合工具EVidenceModeler相比,TSEBRA更具优势。TSEBRA是一个易于使用和快速的软件工具。它可以与BRAKER管道一起使用,产生一个由RNA-seq和同源蛋白证据支持的基因预测集。在线版本包含补充材料,可在10.1186/s12859-021-04482-0获得。
BRAKER is a suite of automatic pipelines, BRAKER1 and BRAKER2, for the accurate annotation of protein-coding genes in eukaryotic genomes. Each pipeline trains statistical models of protein-coding genes based on provided evidence and, then predicts protein-coding genes in genomic sequences using both the extrinsic evidence and statistical models. For training and prediction, BRAKER1 and BRAKER2 incorporate complementary extrinsic evidence: BRAKER1 uses only RNA-seq data while BRAKER2 uses only a database of cross-species proteins. The BRAKER suite has so far not been able to reliably exceed the accuracy of BRAKER1 and BRAKER2 when incorporating both types of evidence simultaneously. Currently, for a novel genome project where both RNA-seq and protein data are available, the best option is to run both pipelines independently, and to pick one, likely better output. Therefore, one or another type of the extrinsic evidence would remain unexploited. We present TSEBRA, a software that selects gene predictions (transcripts) from the sets generated by BRAKER1 and BRAKER2. TSEBRA uses a set of rules to compare scores of overlapping transcripts based on their support by RNA-seq and homologous protein evidence. We show in computational experiments on genomes of 11 species that TSEBRA achieves higher accuracy than either BRAKER1 or BRAKER2 running alone and that TSEBRA compares favorably with the combiner tool EVidenceModeler. TSEBRA is an easy-to-use and fast software tool. It can be used in concert with the BRAKER pipeline to generate a gene prediction set supported by both RNA-seq and homologous protein evidence. The online version contains supplementary material available at 10.1186/s12859-021-04482-0.
DOI: 10.1186/1471-2105-15-189
发表时间: 2014-06-14
期刊: BMC bioinformatics
影响因子: 3
作者:
Gotoh O;Morita M;Nelson DR
通讯作者: Nelson DR
DOI: 10.1186/s12870-018-1282-9
发表时间: 2018-04-12
期刊: BMC plant biology
影响因子: 5.3
作者:
Jayakodi M;Choi BS;Lee SC;Kim NH;Park JY;Jang W;Lakshmanan M;Mohan SVG;Lee DY;Yang TJ
通讯作者: Yang TJ
DOI: 10.1093/nar/gky1053
发表时间: 2019-01-08
影响因子: 14.9
作者:
Kriventseva EV;Kuznetsov D;Tegenfeldt F;Manni M;Dias R;Simão FA;Zdobnov EM
通讯作者: Zdobnov EM
DOI: 10.1093/bioinformatics/btn004
发表时间: 2008-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Liu, Qian;Mackey, Aaron J.;Pereira, Fernando C. N.
通讯作者: Pereira, Fernando C. N.
DOI: 10.1186/gb-2008-9-1-r7
发表时间: 2008-01-11
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Haas, Brian J.;Salzberg, Steven L.;Zhu, Wei;Pertea, Mihaela;Allen, Jonathan E.;Orvis, Joshua;White, Owen;Buell, C. Robin;Wortman, Jennifer R.
通讯作者: Wortman, Jennifer R.