Microarray-based gene set analysis: a comparison of current methods

Microarray-based gene set analysis: a comparison of current methods
复制标题

DOI:
10.1186/1471-2105-9-502
复制
发表时间:
2008-11-27
期刊:
影响因子:
3
通讯作者:
Black, Michael A.
Black, Michael A.
中科院分区:
生物学4区
文献类型:
--
作者:
Song, Sarah;Black, Michael A.

文献摘要

被引文献

相似文献

背景:近年来,基因组的分析已经成为一个热门话题,研究人员试图通过包含补充的生物学信息来提高他们的微阵列分析的可解释性和重复性。虽然基因集分析有许多选择,但关于哪种方法在什么条件下表现最好,还没有达成共识。这项工作的目标是在模拟和真实的微阵列数据集上检查一系列现有基因集分析方法的性能特征。结果:6种基因集分析方法均适用于模拟和公开可用的微阵列数据集。总体而言,各种方法都被发现在检测从非活性状态(即未表达的基因)到活性状态(或反之亦然)的基因集方面都更好,而不是那些简单地改变其活性水平的基因集。结论:基于对模拟数据的分析结果,基因集合分析方法的性能明显受到相关数据集的特征的影响,并且与依赖单变量检验统计的方法相比,在分析过程中融入关联结构的方法倾向于获得更好的性能。
Background: The analysis of gene sets has become a popular topic in recent times, with researchers attempting to improve the interpretability and reproducibility of their microarray analyses through the inclusion of supplementary biological information. While a number of options for gene set analysis exist, no consensus has yet been reached regarding which methodology performs best, and under what conditions. The goal of this work was to examine the performance characteristics of a collection of existing gene set analysis methods, on both simulated and real microarray data sets. Of particular interest was the potential utility gained through the incorporation of inter-gene correlation into the analysis process.Results: Each of six gene set analysis methods was applied to both simulated and publicly available microarray data sets. Overall, the various methodologies were all found to be better at detecting gene sets that moved from non-active (i. e., genes not expressed) to active states (or vice versa), rather than those that simply changed their level of activity. Methods which incorporate correlation structures were found to provide increased ability to detect altered gene sets in some settings.Conclusion: Based on the results obtained through the analysis of simulated data, it is clear that the performance of gene set analysis methods is strongly influenced by the features of the data set in question, and that methods which incorporate correlation structures into the analysis process tend to achieve better performance, relative to methods which rely on univariate test statistics.