From co-expression to co-regulation: how many microarray experiments do we need?

From co-expression to co-regulation: how many microarray experiments do we need?
复制标题

DOI:
10.1186/gb-2004-5-7-r48
复制
发表时间:
2004
期刊:
影响因子:
12.3
通讯作者:
Bumgarner RE
Bumgarner RE
中科院分区:
生物学1区
文献类型:
--
作者:
Yeung KY;Medvedovic M;Bumgarner RE

文献摘要

参考文献

被引文献

相似文献

从微阵列聚类结果中识别共调控基因的能力强烈依赖于聚类分析中使用的微阵列实验的数量,以及这些关联平台在酵母数据上的50到100个实验的准确性。即使有大量的实验,假阳性率也可能超过真阳性率。聚类分析通常用于通过将未知基因与具有相似表达模式和已知调控元件或功能的其他基因相关联来推断调控模块或生物功能。然而,聚类结果可能不具有任何生物学相关性。我们将不同的聚类算法应用于不同大小的微阵列数据集,并通过确定来自相同簇的至少共享一个已知共同转录因子的基因对的比例来评估聚类结果。我们使用酵母转录因子数据库(SCPD、YPD)和染色质免疫沉淀(CHIP)数据来评估我们的聚类结果。我们表明,从聚类结果中识别共调控基因的能力强烈依赖于聚类分析中使用的微阵列实验的数量,以及这些关联平台在酵母数据上的50到100个实验的准确性。此外,基于模型的聚类算法MCLUST在将共调控基因准确分配到标准化数据上的相同聚类方面始终优于更传统的方法。我们的结果与独立的评价标准是一致的,这增强了我们对结果的信心。然而,当人们将芯片数据与YPD进行比较时,使用推荐的p值0.001,假阴性率约为80%。此外,我们还表明,即使进行了大量的实验,假阳性率也可能超过真阳性率。特别是,即使包括所有实验,使用已知的基因转录因子相互作用产生的最佳结果也只有28%的真阳性率。
The ability to identify co-regulated genes from microarray clustering results is strongly dependent on the number of microarray experiments used in cluster analysis and the accuracy of these associations plateaus at between 50 and 100 experiments on yeast data. Even with large numbers of experiments, the false positive rate may exceed the true positive rate. Cluster analysis is often used to infer regulatory modules or biological function by associating unknown genes with other genes that have similar expression patterns and known regulatory elements or functions. However, clustering results may not have any biological relevance. We applied various clustering algorithms to microarray datasets with different sizes, and we evaluated the clustering results by determining the fraction of gene pairs from the same clusters that share at least one known common transcription factor. We used both yeast transcription factor databases (SCPD, YPD) and chromatin immunoprecipitation (ChIP) data to evaluate our clustering results. We showed that the ability to identify co-regulated genes from clustering results is strongly dependent on the number of microarray experiments used in cluster analysis and the accuracy of these associations plateaus at between 50 and 100 experiments on yeast data. Moreover, the model-based clustering algorithm MCLUST consistently outperforms more traditional methods in accurately assigning co-regulated genes to the same clusters on standardized data. Our results are consistent with respect to independent evaluation criteria that strengthen our confidence in our results. However, when one compares ChIP data to YPD, the false-negative rate is approximately 80% using the recommended p-value of 0.001. In addition, we showed that even with large numbers of experiments, the false-positive rate may exceed the true-positive rate. In particular, even when all experiments are included, the best results produce clusters with only a 28% true-positive rate using known gene transcription factor interactions.
DOI: 10.1093/bioinformatics/18.9.1194
发表时间: 2002-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Medvedovic, M;Sivaganesan, S
通讯作者: Sivaganesan, S
DOI: 10.1093/bioinformatics/bth068
发表时间: 2004-05-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Medvedovic, M;Yeung, KY;Bumgarner, RE
通讯作者: Bumgarner, RE
DOI: 10.1007/s003579900058
发表时间: 1999-01-01
影响因子: 2
作者:
Fraley, C;Raftery, AE
通讯作者: Raftery, AE
DOI: 10.1073/pnas.96.6.2907
发表时间: 1999-03-16
影响因子: 11.1
作者:
Tamayo, P;Slonim, D;Golub, TR
通讯作者: Golub, TR
DOI: 10.1016/s0959-437x(02)00277-0
发表时间: 2002-04-01
影响因子: 4
作者:
Wyrick, JJ;Young, RA
通讯作者: Young, RA