Microarray and EST database estimates of mRNA expression levels differ:: The protein length versus expression curve for C-elegans -: art. no. 30

Microarray and EST database estimates of mRNA expression levels differ:: The protein length versus expression curve for C-elegans -: art. no. 30
复制标题

微阵列和 EST 数据库对 mRNA 表达水平的估计不同::C-elegans 的蛋白质长度与表达曲线--:第 30 条。 30

DOI:
10.1186/1471-2164-5-30
复制
发表时间:
2004-05-10
期刊:
影响因子:
4.4
通讯作者:
Deem, MW
Deem, MW
中科院分区:
生物学2区
文献类型:
--
作者:
Munoz, ET;Bogarad, LD;Deem, MW

文献摘要

被引文献

相似文献

背景:已知多种用于估计蛋白质表达水平的方法。这些方法之间的相关性水平只是公平的,并且不能排除每种方法中的系统性偏差。在这里,我们调查系统的偏见,估计基因表达率从微阵列数据和表达序列标签(EST)数据库内的丰度。我们认为,长度是一个显着的因素,在偏差测量的基因表达率。作为一个具体的例子,表达率与长度的偏差的重要性,我们解决了以下进化问题:平均C。蛋白长度随表达量的增加而增加还是减少?在文献中已经报道了对这个问题的两种不同的答案,一种方法使用EST数据库中的丰度估计的表达水平,另一种方法使用微阵列。我们通过构建C.结果:微阵列数据显示长度随表达水平单调下降,而EST数据库数据中的丰度显示非单调行为。此外,EST数据库估计的表达水平的比例,通过微阵列测量的是不恒定的,而是系统地偏向与基因length.Conclusions:它建议的长度偏差可能主要在于在EST数据库的方法内的丰度,不改善内部标准,因为它是在微阵列数据,这种偏差应删除数据解释之前。当这样做时,EST数据库中的微阵列和丰度都给出了剪接长度随表达水平的单调减少,并且EST和微阵列数据之间的相关性变得更大。我们建议在任何测量表达的方法中使用标准RNA对照来标准化长度偏差。
Background: Various methods for estimating protein expression levels are known. The level of correlation between these methods is only fair, and systematic biases in each of the methods cannot be ruled out. We here investigate systematic biases in the estimation of gene expression rates from microarray data and from abundance within the Expressed Sequence Tag (EST) database. We suggest that length is a significant factor in biases to measured gene expression rates.As a specific example of the importance of the bias of expression rate with length, we address the following evolutionary question: Does the average C. elegans protein length increase or decrease with expression level? Two different answers to this question have been reported in the literature, one method using expression levels estimated by abundance within the EST database and another using microarrays. We have investigated this issue by constructing the full protein length versus expression curve for C. elegans, using both methods for estimating expression levels.Results: The microarray data show a monotonic decrease of length with expression level, whereas the abundance within the EST database data show a non-monotonic behavior. Furthermore, the ratio of the expression level estimated by the EST database to that measured by microarrays is not constant, but rather systematically biased with gene length.Conclusions: It is suggested that the length bias may lie primarily in the abundance within the EST database method, being not ameliorated by internal standards as it is in the microarray data, and that this bias should be removed before data interpretation. When this is done, both the microarray and the abundance within the EST database give a monotonic decrease of spliced length with expression level, and the correlation between the EST and microarray data becomes larger. We suggest that standard RNA controls be used to normalize for length bias in any method that measures expression.