Origin and properties of non-coding ORFs in the yeast genome

Origin and properties of non-coding ORFs in the yeast genome
复制标题

DOI:
10.1093/nar/27.17.3503
复制
发表时间:
1999-09-01
影响因子:
14.9
通讯作者:
Cebrat, S
Cebrat, S
中科院分区:
生物学2区
文献类型:
--
作者:
Mackiewicz, P;Kowalczuk, M;Cebrat, S

文献摘要

被引文献

相似文献

在最近的一篇论文中,我们根据酿酒酵母基因组中蛋白质编码的开放阅读框(ORF)的性质估计了它们的总数,大约为4800,这个数字远小于被广泛接受的5800-6000。在本文中,我们分析了一组已知的表型的ORF注释在慕尼黑信息中心的蛋白质序列(MIPS)数据库和ORF的编码的概率,由我们计算,是非常低的差异。我们已经发现,后者的许多ORF具有编码ORF的反义序列的性质,这表明它们可能是由编码序列的重复产生的。由于编码序列在自身内部产生ORF,尤其是在反义序列中产生的频率很高,我们已经在所有六个阶段中寻找已知蛋白质和由ORF产生的假设多肽之间的同源性。对于许多ORF,我们发现旁系同源物和直系同源物的相位不同于MIPS数据库中假定为编码的相位。
In a recent paper we have estimated the total number of protein coding open reading frames (ORFs) in the Saccharomyces cerevisiae genome, based on their properties, at about 4800, This number is much smaller than the 5800-6000 which is widely accepted. In this paper we analyse differences between the set of ORFs with known phenotypes annotated in the Munich Information Centre for Protein Sequences (MIPS) database and ORFs for which the probability of coding, counted by us, is very low. We have found that many of the latter ORFs have properties of antisense sequences of coding ORFs, which suggests that they could have been generated by duplication of coding sequences. Since coding sequences generate ORFs inside themselves, with especially high frequency in the antisense sequences, we have looked for homology between known proteins and hypothetical polypeptides generated by ORFs under consideration in all the six phases. For many ORFs we have found paralogues and orthologues in phases different than the phase which had been assumed in the MIPS database as coding.