Identifying differentially expressed genes in cDNA microarray experiments

Identifying differentially expressed genes in cDNA microarray experiments
复制标题

DOI:
10.1089/106652701753307539
复制
发表时间:
2001-01-01
影响因子:
1.7
通讯作者:
Zhang, W
Zhang, W
中科院分区:
生物学4区
文献类型:
--
作者:
Baggerly, KA;Coombes, KR;Zhang, W

文献摘要

被引文献

相似文献

微阵列实验的一个主要目标是确定哪些基因在样品之间有差异表达。通过在阵列上的一个点和标记点(基因)上取不同样品的表达水平的比率来评估差异表达,其中折叠差异的大小超过某些阈值。最近的工作试图将这些比率的可变性不是恒定的这一事实纳入其中。大多数方法是学生t检验的变体。这些变量通过除以该比率的标准偏差估计值来使比率标准化;具有较大标准化值的点被标记。估计这些标准偏差需要在一张载玻片内或载玻片之间重复测量,或者使用描述标准偏差应该是什么的模型。从驱动微阵列杂交的动力学考虑出发,我们推导了复制点强度的模型,当复制在阵列内部和阵列之间进行时。载玻片内的复制导致β -二项模型,载玻片之间的复制导致γ -泊松模型。这些模型预测对数比率的方差如何随现场信号的总强度而变化,而与基因的身份无关。总信号量小的基因的比率变化很大,而总信号量大的基因的比率则相当稳定。对数比率通过这些函数给出的标准偏差进行缩放,从而给出基于模型的学生化版本。给出了一个例子。
A major goal of microarray experiments is to determine which genes are differentially expressed between samples. Differential expression has been assessed by taking ratios of expression levels of different samples at a spot on the array and flagging spots (genes) where the magnitude of the fold difference exceeds some threshold. More recent work has attempted to incorporate the fact that the variability of these ratios is not constant. Most methods are variants of Student's t-test. These variants standardize the ratios by dividing by an estimate of the standard deviation of that ratio; spots with large standardized values are flagged. Estimating these standard deviations requires replication of the measurements, either within a slide or between slides, or the use of a model describing what the standard deviation should be. Starting from considerations of the kinetics driving microarray hybridization, we derive models for the intensity of a replicated spot, when replication is performed within and between arrays. Replication within slides leads to a beta-binomial model, and replication between slides leads to a gamma-Poisson model. These models predict how the variance of a log ratio changes with the total intensity of the signal at the spot, independent of the identity of the gene. Ratios for genes with a small amount of total signal are highly variable, whereas ratios for genes with a large amount of total signal are fairly stable. Log ratios are scaled by the standard deviations given by these functions, giving model-based versions of Studentization. An example is given.