Hotelling's T2 multivariate profiling for detecting differential expression in microarrays

Hotelling's T2 multivariate profiling for detecting differential expression in microarrays
复制标题

DOI:
10.1093/bioinformatics/bti496
复制
发表时间:
2005-07-15
期刊:
影响因子:
5.8
通讯作者:
Deng, HW
Deng, HW
中科院分区:
生物学3区
文献类型:
--
作者:
Lu, Y;Liu, PY;Deng, HW

文献摘要

被引文献

相似文献

发现差异表达基因(DEG)的最广泛使用的统计方法基本上是单变量的。在这项研究中,我们提出了一个新的T-2统计分析微阵列数据。我们实现了我们的方法,使用多个前向搜索(MFS)算法,该算法是专为选择一个子集的特征向量在高维微阵列数据集。建议的T-2统计量是一个推论,最初开发的多变量分析,并具有两个突出的统计特性。首先,我们的方法考虑到微阵列数据的多维结构。利用隐藏在基因相互作用中的信息,可以发现其差异表达在单变量检验方法中不可检测的基因。其次,统计量与基因表达模式分类的判别分析有密切关系。我们的搜索算法顺序最大化两组基因之间的基因表达差异/距离。将这样的DEG集合包括到初始特征变量中可以增加分类规则的能力。我们通过使用来自Affytown的HGU 95数据集来验证我们的方法。新方法的实用性通过应用于人肝癌和乳腺癌的基因表达模式的分析而得到证明。广泛的生物信息学分析和交叉验证的DEG中确定的应用数据集显示了我们的新算法的显着优势。
The most widely used statistical methods for finding differentially expressed genes (DEGs) are essentially univariate. In this study, we present a new T-2 statistic for analyzing microarray data. We implemented our method using a multiple forward search (MFS) algorithm that is designed for selecting a subset of feature vectors in high-dimensional microarray datasets. The proposed T-2 statistic is a corollary to that originally developed for multivariate analyses and possesses two prominent statistical properties. First, our method takes into account multidimensional structure of microarray data. The utilization of the information hidden in gene interactions allows for finding genes whose differential expressions are not marginally detectable in univariate testing methods. Second, the statistic has a close relationship to discriminant analyses for classification of gene expression patterns. Our search algorithm sequentially maximizes gene expression difference/distance between two groups of genes. Including such a set of DEGs into initial feature variables may increase the power of classification rules. We validated our method by using a spike-in HGU95 dataset from Affymetrix. The utility of the new method was demonstrated by application to the analyses of gene expression patterns in human liver cancers and breast cancers. Extensive bioinformatics analyses and cross-validation of DEGs identified in the application datasets showed the significant advantages of our new algorithm.