GIFtS: annotation landscape analysis with GeneCards

GIFtS: annotation landscape analysis with GeneCards
复制标题

DOI:
10.1186/1471-2105-10-348
复制
发表时间:
2009-10-23
期刊:
影响因子:
3
通讯作者:
Lancet, Doron
Lancet, Doron
中科院分区:
生物学4区
文献类型:
--
作者:
Harel, Arye;Inger, Aron;Lancet, Doron

文献摘要

被引文献

相似文献

背景资料:基因注释是计算基因组学中的一个关键组成部分,包括基因功能预测、表达分析和序列检查。因此,注释景观的定量测量构成了相关的生物信息学工具。GeneCards(R)是一个以基因为中心的概要,包含超过50,000个人类基因条目的丰富注释性信息,基于68个数据源,包括基因本体论(GO)、途径、相互作用、表型、出版物等等。我们提出了基因卡推断功能评分(GIFtS),它允许定量评估基因的注释状态,通过利用基因卡信息的独特财富和多样性。从GeneCards主页链接的GIFtS工具通过搜索指定基因的注释水平、检索指定GIFtS值范围内的基因列表、获得具有特定GIFtS值的随机基因以及对各种注释类别试验GIFtS加权算法来促进人类基因组的浏览。GIFtS分布的双峰形状表明人类基因库分为两大类:高GIFtS峰几乎完全由蛋白质编码基因组成;低GIFtS峰由所有类别的基因组成。GIFS注释载体的聚类分析通过在注释竞技场中的详细定位提供了基因组的分类。GIFtS还提供了一些措施,以便能够对作为基因卡来源的数据库进行评估。发现由每个来源注释的基因的数量与与该来源相关的基因的平均GIFS值之间存在负相关(对于GIFS>25)。三个典型的源原型,揭示了他们的GIFS分布:全基因组的来源,来源主要包括高度注释的基因,和来源主要包括不良注释的基因。GIFtS测量的一个给定的基因的积累的知识的程度相关的(GIFtS>30)与基因的出版物的数量,并与HGNC database.Conclusion中的这个条目的资历:GIFtS可以是一个有价值的工具,用于计算程序,分析从湿实验室或计算研究产生的基因的大集合的列表。GIFtS还可以帮助科学界识别用于不同应用的未表征基因组,例如描绘新功能和绘制人类基因组的未探索区域。
Background: Gene annotation is a pivotal component in computational genomics, encompassing prediction of gene function, expression analysis, and sequence scrutiny. Hence, quantitative measures of the annotation landscape constitute a pertinent bioinformatics tool. GeneCards (R) is a gene-centric compendium of rich annotative information for over 50,000 human gene entries, building upon 68 data sources, including Gene Ontology (GO), pathways, interactions, phenotypes, publications and many more.Results: We present the GeneCards Inferred Functionality Score (GIFtS) which allows a quantitative assessment of a gene's annotation status, by exploiting the unique wealth and diversity of GeneCards information. The GIFtS tool, linked from the GeneCards home page, facilitates browsing the human genome by searching for the annotation level of a specified gene, retrieving a list of genes within a specified range of GIFtS value, obtaining random genes with a specific GIFtS value, and experimenting with the GIFtS weighting algorithm for a variety of annotation categories. The bimodal shape of the GIFtS distribution suggests a division of the human gene repertoire into two main groups: the high-GIFtS peak consists almost entirely of protein-coding genes; the low-GIFtS peak consists of genes from all of the categories. Cluster analysis of GIFtS annotation vectors provides the classification of gene groups by detailed positioning in the annotation arena. GIFtS also provide measures which enable the evaluation of the databases that serve as GeneCards sources. An inverse correlation is found (for GIFtS>25) between the number of genes annotated by each source, and the average GIFtS value of genes associated with that source. Three typical source prototypes are revealed by their GIFtS distribution: genome-wide sources, sources comprising mainly highly annotated genes, and sources comprising mainly poorly annotated genes. The degree of accumulated knowledge for a given gene measured by GIFtS was correlated (for GIFtS>30) with the number of publications for a gene, and with the seniority of this entry in the HGNC database.Conclusion: GIFtS can be a valuable tool for computational procedures which analyze lists of large set of genes resulting from wet-lab or computational research. GIFtS may also assist the scientific community with identification of groups of uncharacterized genes for diverse applications, such as delineation of novel functions and charting unexplored areas of the human genome.