The institute for genomic research Osa1 rice genome annotation database

The institute for genomic research Osa1 rice genome annotation database
复制标题

DOI:
10.1104/pp.104.059063
复制
发表时间:
2005-05-01
期刊:
影响因子:
7.4
通讯作者:
Buell, CR
Buell, CR
中科院分区:
生物学1区
文献类型:
--
作者:
Yuan, QP;Shu, OY;Buell, CR

文献摘要

被引文献

相似文献

我们已经开发了一个水稻(Oryza sativa)基因组注释数据库(Osa 1),为这个新兴的模式物种提供结构和功能注释。利用O.青japonica cv Nipponbare(来自国际水稻基因组测序计划),构建了12条水稻染色体的假分子或虚拟重叠群。我们最新的版本,版本3,代表了我们的第三次构建的假分子,由98%的完成序列组成。使用为拟南芥(Arabidopsis thaliana)开发的一系列计算方法鉴定基因,所述方法被修改以用于水稻基因组。在我们注释的第3版中,我们鉴定了57,915个基因,其中14,196个与转座因子有关。在这43,719个非转座因子相关基因中,18,545个(42.4%)被注释为推定功能,5,777个(13.2%)被注释为编码表达的蛋白质,而没有已知功能,并且剩余的19,397个(44.4%)被注释为编码假设蛋白质。在2,538个基因中检测到了5,873个多剪接形式,导致水稻基因组中总共有61,250个基因模型。我们将实验证据整合到18,252个基因模型中,以提高结构注释的质量。水稻基因组的一系列功能数据类型已被注释,包括与遗传标记的比对、基因本体的分配、侧翼序列标签的识别、与来自相关物种的同源物的比对以及与其他谷类物种的同线作图。所有结构和功能注释数据都可以通过交互式搜索和显示窗口以及通过下载平面文件获得。为了将数据与其他基因组计划整合,注释数据可通过分布式注释系统和基因组浏览器获得。
We have developed a rice ( Oryza sativa) genome annotation database ( Osa1) that provides structural and functional annotation for this emerging model species. Using the sequence of O. sativa subsp. japonica cv Nipponbare from the International Rice Genome Sequencing Project, pseudomolecules, or virtual contigs, of the 12 rice chromosomes were constructed. Our most recent release, version 3, represents our third build of the pseudomolecules and is composed of 98% finished sequence. Genes were identified using a series of computational methods developed for Arabidopsis ( Arabidopsis thaliana) that were modified for use with the rice genome. In release 3 of our annotation, we identified 57,915 genes, of which 14,196 are related to transposable elements. Of these 43,719 nontransposable element- related genes, 18,545 ( 42.4%) were annotated with a putative function, 5,777 ( 13.2%) were annotated as encoding an expressed protein with no known function, and the remaining 19,397 ( 44.4%) were annotated as encoding a hypothetical protein. Multiple splice forms ( 5,873) were detected for 2,538 genes, resulting in a total of 61,250 gene models in the rice genome. We incorporated experimental evidence into 18,252 gene models to improve the quality of the structural annotation. A series of functional data types has been annotated for the rice genome that includes alignment with genetic markers, assignment of gene ontologies, identification of flanking sequence tags, alignment with homologs from related species, and syntenic mapping with other cereal species. All structural and functional annotation data are available through interactive search and display windows as well as through download of flat files. To integrate the data with other genome projects, the annotation data are available through a Distributed Annotation System and a Genome Browser.