Zipf's law in importance of genes for cancer classification using microarray data

Zipf's law in importance of genes for cancer classification using microarray data
复制标题

DOI:
10.1006/jtbi.2002.3145
复制
发表时间:
2002-12-21
影响因子:
2
通讯作者:
Yang, YN
Yang, YN
中科院分区:
生物学4区
文献类型:
--
作者:
Li, WT;Yang, YN

文献摘要

被引文献

相似文献

使用一种衡量一个基因在两种生化/表型不同的条件下表达差异程度的指标,我们可以对微阵列数据集中的所有基因进行排序。我们已经证明,这一度量(在分类模型中,如Logistic回归中的归一化最大似然)作为等级的函数的下降通常是幂函数。在许多自然和社会现象中观察到的其他类似排列的地块中的这种幂定律被称为齐普夫定律。这种幂函数的存在防止了“重要”基因和“无关”基因之间的内在分界点。我们已经证明了相似的幂函数也存在于置换数据集中,并从众所周知的概率比的X(2)分布中给出了解释。我们讨论了这一Zipf定律在微阵列数据分析中对基因选择的意义,以及排序似然图的其他特征,如似然率的下降速度。(C)2002爱思唯尔科学有限公司。保留所有权利。
Using a measure of how differentially expressed a gene is in two biochemically/phenotypically different conditions, we can rank all genes in a microarray dataset. We have shown that the falling-off of this measure (normalized maximum likelihood in a classification model such as logistic regression) as a function of the rank is typically a power-law function. This power-law function in other similar ranked plots are known as the Zipf's law, observed in many natural and social phenomena. The presence of this power-law function prevents an intrinsic cutoff point between the "important" genes and "irrelevant" genes. We have shown that similar power-law functions are also present in permuted dataset, and provide an explanation from the well-known chi(2) distribution of likelihood ratios. We discuss the implication of this Zipf's law on gene selection in a microarray data analysis, as well as other characterizations of the ranked likelihood plots such as the rate of fall-off of the likelihood. (C) 2002 Elsevier Science Ltd. All rights reserved.