Supplement 3

Supplement 3
复制标题

补充3

DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
C. Bermann
C. Bermann
中科院分区:
--
文献类型:
--
作者:
F. M. A. Collaço;Raiana Schirmer Soares;João Marcos Mott Pavanelli;L. L. Benites;Guilherme Massignan Berejuk;Andrea Lampis;C. Bermann

文献摘要

被引文献

相似文献

由指数对数正态混合模型和注释蛋白质的非随机分布描述的ORFs大小分布形状的显著关系和说明。所有参数值均在附录1中列出。在该补充中,图1描述了用指数对数正态模型估计的参数与从注释蛋白质估计的参数之间的统计关系(也参见主文件中的图4),图2至图312说明了这些分布中的每一个的形状(按物种顺序排列,以与补充1相对应)。在这些图中,红线表示从混合模型估计的分布的形状,蓝线显示来自注释蛋白质的分布的形状。重要的是从这些图中注意到,非随机ORF的大小分布的形状偏离注释的蛋白质,使得小的非随机ORF的数目始终大于小的注释的蛋白质的数目。因此,从混合模型估计的对数正态分布的平均值较小,并且分布的峰值向注释蛋白质的峰值的左侧移动。此外,多细胞真核生物中的移位幅度大于原核生物。然而,从混合模型拟合估计的平均值和标准偏差与从拟合到注释蛋白质的参数估计值显著相关(S4 -图1)。
Significant relationships and illustrations of the shapes of the size distributions of ORFS described by the non-random distribution of the exponential-log normal mixture model and annotated proteins. All parameter values are listed in Supplement 1. Within this supplement, figure one depicts the statistical relationships between the parameters estimated with the exponential-log normal model and parameters estimated from annotated proteins (also see figure 4 in main document), and figures two through 312 illustrate the shapes of each of these distributions (arranged alphabetically by species to correspond with Supplement 1). In these figures red lines represent the shapes of the distributions estimated from the mixture model and blue lines show the shape of the distribution from the annotated proteins. It is important to note from these figures that the shapes of the size distributions of non-random ORFs deviate from the annotated proteins such that the number of small non-random ORFs is consistently greater than the number of small annotated proteins. Consequently, the means of the lognormal distributions estimated from the mixture models are smaller and the peaks of the distributions shifted to the left of those for annotated proteins. Moreover, the magnitude of the shift is greater in multicellular eukaryotes than prokaryotes. Nevertheless, the mean and standard deviation estimated from the mixture model fits are significantly correlated with the parameter estimates from fits to annotated proteins (S4 – Fig. 1).