Computational analysis of gene identification with SAGE

Computational analysis of gene identification with SAGE
复制标题

DOI:
10.1089/106652702760138600
复制
发表时间:
2002-01-01
影响因子:
1.7
通讯作者:
Wang, SM
Wang, SM
中科院分区:
生物学4区
文献类型:
--
作者:
Clark, T;Lee, S;Wang, SM

文献摘要

被引文献

相似文献

SAGE是少数几种能够在基因组水平上均匀探测基因表达的技术之一,而不管mRNA丰度如何,并且不需要存在转录本的先验知识。然而,单个SAGE标签可以匹配参考数据库中的许多序列,使基因鉴定复杂化。我们使用UniGene Human作为参考数据库,通过分析1)针对UniGene Human形成的各种长度标签集的标签分布和2)使用由源自人骨髓细胞的37,522个标签组成的SAGE标签集的标签-序列映射,使用SAGE进行基因鉴定的基线评估。UniGene的dbEST组分的广泛多样性显著减损了通过在SAGE方案的范围内扩展标签可能预期的收益。为了用常用的UniGene序列集合的内容实现基因鉴定的合理序列特异性,需要数百个碱基长度的标签。产生这种长度的标签的一种方法是使用GLGI,其将SAGE标签延伸到cDNA的3'末端。我们表明,较长的序列产生的GLGI缓解显着的多匹配条件。在骨髓样本中,我们还发现了多重匹配严重程度和高拷贝数之间的相关性。我们推断这些发现,为使用UniGene Human作为基因鉴定的参考提供了见解。
SAGE is one of the few techniques capable of uniformly probing gene expression at a genome level irrespective of mRNA abundance and without a priori knowledge of the transcripts present. However, individual SAGE tags can match many sequences in the reference database, complicating gene identification. We perform a baseline evaluation of gene identification with SAGE using UniGene Human as the reference database by analyzing 1) the distributions of tags for various length tag sets formed for UniGene Human and 2) the tag-to-sequence mapping using a SAGE tag set consisting of 37,522 tags derived from human myeloid cells. The extensive multiplicity of the dbEST component of UniGene significantly detracts from gains that might be expected by extending tags within the scope of the SAGE protocol. In order to achieve reasonable sequence specificity for gene identification with the content of the commonly used UniGene sequence collection, tags on the order of hundreds of bases in length are required. One way to produce tags of such lengths is with GLGI, which extends SAGE tags to the 3' end of cDNA. We show that the longer sequences produced by GLGI relieve significantly the multiple match condition. In the myeloid sample, we also found a correlation between multiple match severity and high copy number. We extrapolate these findings, providing insights into the use of UniGene Human as a reference for gene identification.