Optimized ranking and selection methods for feature selection with application in microarray experiments.

Optimized ranking and selection methods for feature selection with application in microarray experiments.
复制标题

DOI:
10.1080/10543400903572720
复制
发表时间:
2010-03
影响因子:
1.1
通讯作者:
Wilson J
Wilson J
中科院分区:
医学4区
文献类型:
--
作者:
Cui X;Zhao H;Wilson J

文献摘要

参考文献

相似文献

在微阵列实验中,目标通常是检查许多基因,并选择其中一些进行额外的研究。传统上,这样的选择问题已经制定为一个多测试问题。当感兴趣的基因是在不同条件下基因表达分布不均匀的基因时,多种测试方法为解决选择问题提供了适当的框架。然而,当感兴趣的基因是一组在不同条件下基因表达差异最大的基因时,多种测试方法并不能直接解决选择目标,有时会导致有偏见的结论。对于这种情况下,我们提出了两种方法的基础上的统计排名和选择框架,直接解决的选择目标。所提出的方法具有固有的优化性质,因为根据预先指定的最小正确选择比率(r*-选择)或做出正确选择的概率(P*-选择)来优化选择。这些方法与多重检验方法进行了比较,多重检验方法控制了假阳性比例的尾部概率。模拟研究和真实的数据应用都提供了对多种测试方法和所提出的方法在解决不同选择目标方面的基本差异的深入了解。已经表明,当目标是选择最显著的基因(而不是所有显著的基因)时,所提出的方法提供了明显的优势。当目标是选择所有显著基因时,所提出的方法与当前的多个测试方法一样好。所提出的方法提供的另一个优点是它们检测噪声数据的能力,因此建议不能做出明智的选择。
In microarray experiments, the goal is often to examine many genes, and select some of them for additional investigation. Traditionally, such a selection problem has been formulated as a multiple testing problem. When the genes of interest are genes with unequal distribution of gene expression under different conditions, multiple testing methods provide an appropriate framework for addressing the selection problems. However, when the genes of interest are a set of genes with the largest difference in gene expression under different conditions, multiple testing methods do not directly address the selection goal and sometimes lead to biased conclusions. For such cases, we propose two methods based on the statistical ranking and selection framework to directly address the selection goal. The proposed methods have an inherent optimization nature in that the selection is optimized according to either a pre-specified minimum correct selection ratio (r*-selection) or probability of making a correct selection (P*-selection). These methods are compared with the multiple testing method which controls the tail probability of the proportion of false positives. Both simulation studies and real data applicaitons provide insight into the fundmental difference between the multiple testing methods and the proposed methods in the way of addressing different selection goals. It has been shown that the proposed methods provide a clear advantage over the multiple testing methods when the goal is to select the most significant genes (not all the significant genes). When the goal is to select all the significant genes, the proposed methods perform equally well as the current multiple testing methods. Another advantage provided by the proposed methods is their ability to detect noisy data and therefore suggest no sensible selection can be made.
DOI: 10.1214/aoms/1177728845
发表时间: 1954-01-01
影响因子: --
作者:
BECHHOFER, RE
通讯作者: BECHHOFER, RE
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1111/1541-0420.00016
发表时间: 2003-03-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
Pepe, MS;Longton, G;Schummer, M
通讯作者: Schummer, M
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1093/bioinformatics/bth160
发表时间: 2004-07-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Pounds, S;Cheng, C
通讯作者: Cheng, C