Binary analysis and optimization-based normalization of gene expression data

Binary analysis and optimization-based normalization of gene expression data
复制标题

DOI:
10.1093/bioinformatics/18.4.555
复制
发表时间:
2002-04-01
期刊:
影响因子:
5.8
通讯作者:
Zhang, W
Zhang, W
中科院分区:
生物学3区
文献类型:
--
作者:
Shmulevich, I;Zhang, W

文献摘要

被引文献

相似文献

动机:基因表达分析的大多数方法使用通过高通量筛选技术(例如微阵列)产生的实值表达数据。通常,为了从观察到的数据中提取有意义的信息,必须计算一些相似性度量。这种相似性度量的选择经常对分析结果产生深远的影响,但没有标准存在指导researcher.Results:为了解决这个问题,我们建议完全在二进制域中分析基因表达数据。相似性的自然度量成为汉明距离,反映了生物学家使用的相似性概念。我们还开发了一种新的基于数据依赖优化的方法,基于遗传算法(GAs),规范化基因表达数据。这是将基因表达数据量化到二进制域之前的必要步骤,并且通常用于比较不同阵列之间的数据。然后,我们提出了一个算法二进制化基因表达数据,并说明了使用上述方法对两组不同的数据。使用多维标度,我们表明,在每个数据集中的不同肿瘤类型之间的分离可以通过单独在二进制域中工作来实现。二元方法具有多个优点,例如抗噪能力和计算效率,使其成为从基因表达数据中提取有意义的生物信息的可行方法。联系方式:is@ieee.org。
Motivation: Most approaches to gene expression analysis use real-valued expression data, produced by high-throughput screening technologies, such as microarrays. Often, some measure of similarity must be computed in order to extract meaningful information from the observed data. The choice of this similarity measure frequently has a profound effect on the results of the analysis, yet no standards exist to guide the researcher.Results: To address this issue, we propose to analyse gene expression data entirely in the binary domain. The natural measure of similarity becomes the Hamming distance and reflects the notion of similarity used by biologists. We also develop a novel data-dependent optimization-based method, based on Genetic Algorithms (GAs), for normalizing gene expression data. This is a necessary step before quantizing gene expression data into the binary domain and generally, for comparing data between different arrays. We then present an algorithm for binarizing gene expression data and illustrate the use of the above methods on two different sets of data. Using Multidimensional Scaling, we show that a reasonable degree of separation between different tumor types in each data set can be achieved by working solely in the binary domain. The binary approach offers several advantages, such as noise resilience and computational efficiency, making it a viable approach to extracting meaningful biological information from gene expression data.Contact: is@ieee.org.