BRAKER2: automatic eukaryotic genome annotation with GeneMark-EP+ and AUGUSTUS supported by a protein database.

BRAKER2: automatic eukaryotic genome annotation with GeneMark-EP+ and AUGUSTUS supported by a protein database.
复制标题

BRAKER 2:由蛋白质数据库支持的GeneMark-EP+和AUGUSTUS自动真核基因组注释。

DOI:
10.1093/nargab/lqaa108
复制
发表时间:
2021-03
影响因子:
4.6
通讯作者:
Borodovsky M
Borodovsky M
中科院分区:
其他
文献类型:
--
作者:
Brůna T;Hoff KJ;Lomsadze A;Stanke M;Borodovsky M

文献摘要

参考文献

被引文献

相似文献

真核生物基因组注释的任务仍然具有挑战性。只有少数基因组可以作为标准的注释,通过巨大的人类策展工作的投资实现。尽管如此,所有替代异构体的正确性,即使是在最好的注释基因组中,也可能是进一步研究的好课题。新的BRAKER 2管道生成外部蛋白质支持,并将其集成到GeneMark-EP+和AUGUSTUS的训练和基因预测的迭代过程中。BRAKER 2延续了由BRAKER 1开始的路线,其中自我训练GeneMark-ET和AUGUSTUS进行了由转录组数据支持的基因预测。新管道解决的挑战之一是从可能同源但进化上遥远的蛋白质中产生蛋白质编码外显子边界的可靠提示。与其他真核基因组注释管道相比,BRAKER 2是全自动的。在同等条件下,它在精度和性能方面与其他管道(例如MAKER 2)相比是有利的。BRAKER 2的开发将有助于解决不同真核生物基因组中蛋白质编码基因注释的协调问题。然而,我们完全理解,在转录组学和蛋白质组学技术以及算法开发方面还需要更多的创新,以达到高度准确地注释真核基因组的目标。
The task of eukaryotic genome annotation remains challenging. Only a few genomes could serve as standards of annotation achieved through a tremendous investment of human curation efforts. Still, the correctness of all alternative isoforms, even in the best-annotated genomes, could be a good subject for further investigation. The new BRAKER2 pipeline generates and integrates external protein support into the iterative process of training and gene prediction by GeneMark-EP+ and AUGUSTUS. BRAKER2 continues the line started by BRAKER1 where self-training GeneMark-ET and AUGUSTUS made gene predictions supported by transcriptomic data. Among the challenges addressed by the new pipeline was a generation of reliable hints to protein-coding exon boundaries from likely homologous but evolutionarily distant proteins. In comparison with other pipelines for eukaryotic genome annotation, BRAKER2 is fully automatic. It is favorably compared under equal conditions with other pipelines, e.g. MAKER2, in terms of accuracy and performance. Development of BRAKER2 should facilitate solving the task of harmonization of annotation of protein-coding genes in genomes of different eukaryotic species. However, we fully understand that several more innovations are needed in transcriptomic and proteomic technologies as well as in algorithmic development to reach the goal of highly accurate annotation of eukaryotic genomes.
DOI: 10.1093/nar/gky1053
发表时间: 2019-01-08
影响因子: 14.9
作者:
Kriventseva EV;Kuznetsov D;Tegenfeldt F;Manni M;Dias R;Simão FA;Zdobnov EM
通讯作者: Zdobnov EM
DOI: 10.1038/s41598-017-12863-w
发表时间: 2017-10-02
期刊: Scientific reports
影响因子: 4.6
作者:
de Bekker C;Ohm RA;Evans HC;Brachmann A;Hughes DP
通讯作者: Hughes DP
DOI: 10.1093/nar/gkw092
发表时间: 2016-05-19
影响因子: 14.9
作者:
Keilwagen J;Wenk M;Erickson JL;Schattat MH;Grau J;Hartung F
通讯作者: Hartung F
DOI: 10.1101/gr.6743907
发表时间: 2008-01-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Cantarel, Brandi L.;Korf, Ian;Yandell, Mark
通讯作者: Yandell, Mark
通过自我训练算法在新型真核基因组中的基因鉴定。
DOI: 10.1093/nar/gki937
发表时间: 2005
影响因子: 14.9
作者:
Lomsadze A;Ter-Hovhannisyan V;Chernoff YO;Borodovsky M
通讯作者: Borodovsky M