Concordant integrative gene set enrichment analysis of multiple large-scale two-sample expression data sets.

Concordant integrative gene set enrichment analysis of multiple large-scale two-sample expression data sets.
复制标题

DOI:
10.1186/1471-2164-15-s1-s6
复制
发表时间:
2014
期刊:
影响因子:
4.4
通讯作者:
McCaffrey TA
McCaffrey TA
中科院分区:
生物学2区
文献类型:
--
作者:
Lai Y;Zhang F;Nayak TK;Modarres R;Lee NH;McCaffrey TA

文献摘要

相似文献

基因集浓缩分析(GSEA)是在途径水平上分析协同表达变化的重要方法。尽管已经提出了许多统计和计算方法来研究GSEA,但多个表达数据集的一致性综合GSEA问题还没有得到很好的解决。在为相同或相似的研究目的收集的不同相关数据集中,重要的是确定具有一致富集性的途径或基因集。我们将差异表达的潜在真实状态分为三个有代表性的类别:无变化、积极变化和消极变化。由于数据噪声,我们从实验中观察到的可能并不表明潜在的真相。尽管在实践中没有观察到这些类别,但可以在混合模型框架中考虑它们。然后,定义了整合基因集丰富的数学概念,并基于一个三元多元正态混合模型计算了其相关概率。相关的错误发现率可以计算出来,并用来对不同的基因集进行排序。我们使用了三个已发表的肺癌微阵列基因表达数据集来说明我们提出的方法。一项基于前两个数据集的分析是为了将我们的结果与之前发表的结果进行比较,该结果基于分别为每个单独的数据集进行的GSEA。这一比较说明了我们提出的整合基因集浓缩分析的优势。然后,通过一个相对新的、更大的路径收集,我们使用我们的方法对前两个数据集以及所有三个数据集进行了综合分析。这两个结果都表明,许多基因集可以被识别出来,错误发现率很低。两种结果之间的一致性也被观察到。基于KEGG癌症通路收集的进一步探索表明,我们提出的方法可以识别这些通路中的大多数。这项研究表明,通过对多个大规模两样本基因表达数据集的整合分析,可以提高检测能力和发现一致性。
Gene set enrichment analysis (GSEA) is an important approach to the analysis of coordinate expression changes at a pathway level. Although many statistical and computational methods have been proposed for GSEA, the issue of a concordant integrative GSEA of multiple expression data sets has not been well addressed. Among different related data sets collected for the same or similar study purposes, it is important to identify pathways or gene sets with concordant enrichment. We categorize the underlying true states of differential expression into three representative categories: no change, positive change and negative change. Due to data noise, what we observe from experiments may not indicate the underlying truth. Although these categories are not observed in practice, they can be considered in a mixture model framework. Then, we define the mathematical concept of concordant gene set enrichment and calculate its related probability based on a three-component multivariate normal mixture model. The related false discovery rate can be calculated and used to rank different gene sets. We used three published lung cancer microarray gene expression data sets to illustrate our proposed method. One analysis based on the first two data sets was conducted to compare our result with a previous published result based on a GSEA conducted separately for each individual data set. This comparison illustrates the advantage of our proposed concordant integrative gene set enrichment analysis. Then, with a relatively new and larger pathway collection, we used our method to conduct an integrative analysis of the first two data sets and also all three data sets. Both results showed that many gene sets could be identified with low false discovery rates. A consistency between both results was also observed. A further exploration based on the KEGG cancer pathway collection showed that a majority of these pathways could be identified by our proposed method. This study illustrates that we can improve detection power and discovery consistency through a concordant integrative analysis of multiple large-scale two-sample gene expression data sets.