The H-Invitational Database (H-InvDB), a comprehensive annotation resource for human genes and transcripts

The H-Invitational Database (H-InvDB), a comprehensive annotation resource for human genes and transcripts
复制标题

DOI:
10.1093/nar/gkm999
复制
发表时间:
2008-01-01
影响因子:
14.9
通讯作者:
Gojobori, Takashi
Gojobori, Takashi
中科院分区:
生物学2区
文献类型:
--
作者:
Yamasaki, Chisato;Murakami, Katsuhiko;Gojobori, Takashi

文献摘要

被引文献

相似文献

在这里,我们报告了我们最新发布的H-邀请数据库(H-InvDB;http://www.h-invitational.jp/),)中的新功能和改进,这是一个关于人类基因和转录本的全面注释资源。H-InvDB最初是作为人类转录组的综合数据库而开发的,它基于大量全长cDNA克隆的广泛注释,现在除了最新发布的H-InvDB_4.6中的54 978个人FlcDNA外,还为从国际核苷酸序列数据库(INSD)中提取的120 558个人的mRNAs提供注释。我们将这些人类转录本映射到人类基因组序列(NCBI Build 36.1)上,确定了34 699个人类基因簇,其中34057个(98.1%)是蛋白质编码的,642个(1.9%)是非蛋白质编码的,858个(2.5%)转录的基因与预测的假基因重叠。对于所有这些转录本和基因,我们提供了全面的注释,包括基因结构、基因功能、选择性剪接变体、功能性非蛋白质编码RNA、功能结构域、预测的亚细胞定位、代谢途径、蛋白质3D结构的预测、SNPs和微卫星重复基序的定位、与孤儿疾病的共定位、基因表达谱、同源基因、蛋白质相互作用(PPI)和基因家族的注释。目前的H-InvDB注释资源包括两个主要视图:文本视图和位置视图,以及八个子数据库:疾病信息查看器、H-Angel、集群查看器、G-Integra、Topo查看器、Evola、PPI视图和基因家族/组。
Here we report the new features and improvements in our latest release of the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/), a comprehensive annotation resource for human genes and transcripts. H-InvDB, originally developed as an integrated database of the human transcriptome based on extensive annotation of large sets of full-length cDNA (FLcDNA) clones, now provides annotation for 120 558 human mRNAs extracted from the International Nucleotide Sequence Databases (INSD), in addition to 54 978 human FLcDNAs, in the latest release H-InvDB_4.6. We mapped those human transcripts onto the human genome sequences (NCBI build 36.1) and determined 34 699 human gene clusters, which could define 34 057 (98.1%) protein-coding and 642 (1.9%) non-protein-coding loci; 858 (2.5%) transcribed loci overlapped with predicted pseudogenes. For all these transcripts and genes, we provide comprehensive annotation including gene structures, gene functions, alternative splicing variants, functional non-protein-coding RNAs, functional domains, predicted sub cellular localizations, metabolic pathways, predictions of protein 3D structure, mapping of SNPs and microsatellite repeat motifs, co-localization with orphan diseases, gene expression profiles, orthologous genes, proteinprotein interactions (PPI) and annotation for gene families. The current H-InvDB annotation resources consist of two main views: Transcript view and Locus view and eight sub-databases: the DiseaseInfo Viewer, H-ANGEL, the Clustering Viewer, G-integra, the TOPO Viewer, Evola, the PPI view and the Gene family/group.