The Ensembl automatic gene annotation system

The Ensembl automatic gene annotation system
复制标题

DOI:
10.1101/gr.1858004
复制
发表时间:
2004-05-01
期刊:
影响因子:
7
通讯作者:
Clamp, M
Clamp, M
中科院分区:
生物学1区
文献类型:
--
作者:
Curwen, V;Eyras, E;Clamp, M

文献摘要

被引文献

相似文献

随着越来越多的基因组的测序,对自动的第一频繁注释的需求越来越多,这允许及时访问重要的基因组信息。 Ensembl基因构建系统可实现真核基因组的快速自动注释。它根据源自已知蛋白质,cDNA和EST序列的证据来注释基因。基因构建系统位于Core Ensembl(MySQL)数据库架构和Perl应用程序编程接口(API)之上,并且生成的数据可通过Ensembl Genome浏览器(http://wwww.ensembl.org)访问。迄今为止,Ensembl预测的基因集可用于A. gambiae,c briggsae,斑马鱼,小鼠,大鼠和人类基因组,并且在人类,小鼠,大鼠和A. gambiae基因组的出版中都非常依赖序列分析。在这里,我们详细描述了基因建设系统和所涉及的算法。从http://www.ensembl.org免费获得所有代码和数据。
As more genomes are sequenced, there is an increasing need for automated first-pass annotation which allows timely access to important genomic information. The Ensembl gene-building system enables fast automated annotation of eukaryotic genomes. It annotates genes based on evidence derived from known protein, cDNA, and EST sequences. The gene-building system rests on top of the core Ensembl (MySQL) database schema and Perl Application Programming Interface (API), and the data generated are accessible through the Ensembl genome browser (http://www.ensembl.org). To date, the Ensembl predicted gene sets are available for the A. gambiae, C briggsae, zebrafish, mouse, rat, and human genomes and have been heavily relied upon in the publication of the human, mouse, rat, and A. gambiae genome sequence analysis. Here we describe in detail the gene-building system and the algorithms involved. All code and data are freely available from http://www.ensembl.org.