A robust measure of correlation between two genes on a microarray.

A robust measure of correlation between two genes on a microarray.
复制标题

DOI:
10.1186/1471-2105-8-220
复制
发表时间:
2007-06-25
期刊:
影响因子:
3
通讯作者:
VanKoten B
VanKoten B
中科院分区:
生物学4区
文献类型:
--
作者:
Hardin J;Mitani A;Hicks L;VanKoten B

文献摘要

参考文献

被引文献

相似文献

微阵列实验的潜在目标是在不同的实验条件下识别基因表达模式。包含在特定途径中的基因或对实验条件有相似反应的基因可以在微阵列上共表达并显示相似的表达模式。使用各种聚类方法或基因网络分析中的任何一种,我们可以根据相似性度量将感兴趣的基因划分为组、簇或模块。通常,在实现聚类算法之前,Pearson相关性用于测量距离(或相似性)。然而,皮尔逊相关性很容易受到异常值的影响,这是处理微阵列数据时的一个不幸特征(众所周知,通常非常嘈杂)。我们提出了一种基于Tukey的多变量尺度和位置的双权重估计的抗相似性度量。抗性度量是简单地从尺度的抗性协方差矩阵获得的相关性。我们给出的结果表明,我们的相关度量比Pearson相关更有抵抗力,同时比其他非参数相关度量(例如,Spearman相关)更有效。此外,我们的方法给出了一个系统的基因标记过程,这在处理大量的噪声数据时是有用的。当处理众所周知噪声很大的微阵列数据时,应该使用稳健的方法。具体来说,在聚类和基因网络分析中应该使用鲁棒距离,包括双权相关。
The underlying goal of microarray experiments is to identify gene expression patterns across different experimental conditions. Genes that are contained in a particular pathway or that respond similarly to experimental conditions could be co-expressed and show similar patterns of expression on a microarray. Using any of a variety of clustering methods or gene network analyses we can partition genes of interest into groups, clusters, or modules based on measures of similarity. Typically, Pearson correlation is used to measure distance (or similarity) before implementing a clustering algorithm. Pearson correlation is quite susceptible to outliers, however, an unfortunate characteristic when dealing with microarray data (well known to be typically quite noisy.) We propose a resistant similarity metric based on Tukey's biweight estimate of multivariate scale and location. The resistant metric is simply the correlation obtained from a resistant covariance matrix of scale. We give results which demonstrate that our correlation metric is much more resistant than the Pearson correlation while being more efficient than other nonparametric measures of correlation (e.g., Spearman correlation.) Additionally, our method gives a systematic gene flagging procedure which is useful when dealing with large amounts of noisy data. When dealing with microarray data, which are known to be quite noisy, robust methods should be used. Specifically, robust distances, including the biweight correlation, should be used in clustering and gene network analysis.
DOI: 10.1214/aos/1176347386
发表时间: 1989-12-01
影响因子: 4.5
作者:
LOPUHAA, HP
通讯作者: LOPUHAA, HP
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1093/bioinformatics/btg330
发表时间: 2003-12-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Gat-Viks, I;Sharan, R;Shamir, R
通讯作者: Shamir, R
DOI: 10.1093/bioinformatics/18.12.1585
发表时间: 2002-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hubbell, E;Liu, WM;Mei, R
通讯作者: Mei, R
DOI: 10.1214/aos/1176347978
发表时间: 1991-03-01
影响因子: 4.5
作者:
LOPUHAA, HP;ROUSSEEUW, PJ
通讯作者: ROUSSEEUW, PJ