Comparison of the current RefSeq, Ensembl and EST databases for counting genes and gene discovery

Comparison of the current RefSeq, Ensembl and EST databases for counting genes and gene discovery
复制标题

DOI:
10.1016/j.febslet.2004.12.046
复制
发表时间:
2005-01-31
期刊:
影响因子:
3.5
通讯作者:
Schiöth, HB
Schiöth, HB
中科院分区:
生物学3区
文献类型:
--
作者:
Larsson, TP;Murray, CG;Schiöth, HB

文献摘要

被引文献

相似文献

大量以预测、策划和注释基因和表达序列标签 (EST) 形式呈现的精炼序列材料最近已添加到 NCB1 数据库中。我们匹配了 RefSeq、EnsembI 和 dbEST 的转录序列,试图提供关于可以找到多少独特人类基因的最新概述。结果表明,RefSeq 和 Ensembl 的联合中有大约 25000 个独特的基因,其中每组中分别有 12-18% 和 8-13% 的基因对于另一组来说是独特的。大约 20% 的基因有剪接变异。有相当数量的 EST (2200000) 与已识别的基因不匹配,我们使用内部管道从 Gen-scan 预测中识别出 22 个具有相当大 EST 覆盖率的新基因。该研究深入了解了人类基因目录的现状,并表明需要对方法和数据集进行大量改进才能得出结论性的基因计数。 (C) 2004 年欧洲生化学会联合会。由 Elsevier B.V. 出版。保留所有权利。
Large amounts of refined sequence material in the form of predicted, curated and annotated genes and expressed sequences tags (ESTs) have recently been added to the NCB1 databases. We matched the transcript-sequences of RefSeq, EnsembI and dbEST in an attempt to provide an updated overview of how many unique human genes can be found. The results indicate that there are about 25000 unique genes in the union of RefSeq and Ensembl with 12-18% and 8-13% of the genes in each set unique to the other set, respectively. About 20% of all genes had splice variants. There are a considerable number of ESTs (2200000) that do not match the identified genes and we used an in-house pipeline to identify 22 novel genes from Gen-scan predictions that have considerable EST coverage. The study provides an insight into the current status of human gene catalogues and shows that considerable refinement of methods and datasets is needed to come to a conclusive gene count. (C) 2004 Federation of European Biochemical Societies. Published by Elsevier B.V. All rights reserved.