Detecting discordance enrichment among a series of two-sample genome-wide expression data sets.

Detecting discordance enrichment among a series of two-sample genome-wide expression data sets.
复制标题

DOI:
10.1186/s12864-016-3265-2
复制
发表时间:
2017-01-25
期刊:
影响因子:
4.4
通讯作者:
McCaffrey TA
McCaffrey TA
中科院分区:
生物学2区
文献类型:
--
作者:
Lai Y;Zhang F;Nayak TK;Modarres R;Lee NH;McCaffrey TA

文献摘要

相似文献

随着微阵列和RNA-seq技术的发展,双样本全基因组表达数据已广泛用于生物学和医学研究。相关的差异表达分析和基因集富集分析已经频繁进行。当有多个数据集可用时,可以进行综合分析。在实践中,一系列数据集之间不一致的分子行为可能具有生物学和临床意义。本研究提出了一种检测不一致基因集富集的统计方法。我们的方法是基于一个两水平多元正态混合模型。当数据集数量增加时,参数空间线性增加,统计效率提高。基于模型的不一致富集概率可用于基因集检测。我们将我们的方法应用于从45对匹配的肿瘤/非肿瘤组织中收集的微阵列表达数据集,用于研究胰腺癌。我们根据基因PNLIP(胰脂肪酶,最近显示与胰腺癌相关)的肿瘤/非肿瘤配对表达比,将数据集划分为一系列不重叠的亚组。对数比的范围从负值(例如在非肿瘤组织中表达更多)到正值(例如在肿瘤组织中表达更多)。我们的目的是了解是否有任何基因集在这些子集之间富集了不协调行为(当对数比从负增加到正时)。我们关注KEGG通路。这些检测到的通路将有助于我们进一步了解PNLIP基因在胰腺癌研究中的作用。在检测到的途径中,神经活性配体受体相互作用和嗅觉转导途径是最显著的两个。然后,我们考虑在癌症研究中作为肿瘤抑制因子而闻名的基因TP53。对数比的范围也从负值(例如,在非肿瘤组织中表达更多)到正值(例如,在肿瘤组织中表达更多)。我们根据TP53基因的表达比例对微阵列数据集进行再次划分。不一致富集分析后,我们观察到总体结果相似,上述两种途径仍然是最显著的检测。更有趣的是,在全基因组关联研究(GWAS)数据的通路分析中,只有这两条通路被确定与胰腺癌相关。这项研究表明,当一个重要的疾病相关基因的表达发生改变时,一些疾病相关途径可能会在不一致的分子行为中富集。我们提出的统计方法对这些途径的检测是有用的。此外,我们的方法也可以应用于最近的RNA-seq技术收集的全基因组表达数据。
With the current microarray and RNA-seq technologies, two-sample genome-wide expression data have been widely collected in biological and medical studies. The related differential expression analysis and gene set enrichment analysis have been frequently conducted. Integrative analysis can be conducted when multiple data sets are available. In practice, discordant molecular behaviors among a series of data sets can be of biological and clinical interest. In this study, a statistical method is proposed for detecting discordance gene set enrichment. Our method is based on a two-level multivariate normal mixture model. It is statistically efficient with linearly increased parameter space when the number of data sets is increased. The model-based probability of discordance enrichment can be calculated for gene set detection. We apply our method to a microarray expression data set collected from forty-five matched tumor/non-tumor pairs of tissues for studying pancreatic cancer. We divided the data set into a series of non-overlapping subsets according to the tumor/non-tumor paired expression ratio of gene PNLIP (pancreatic lipase, recently shown it association with pancreatic cancer). The log-ratio ranges from a negative value (e.g. more expressed in non-tumor tissue) to a positive value (e.g. more expressed in tumor tissue). Our purpose is to understand whether any gene sets are enriched in discordant behaviors among these subsets (when the log-ratio is increased from negative to positive). We focus on KEGG pathways. The detected pathways will be useful for our further understanding of the role of gene PNLIP in pancreatic cancer research. Among the top list of detected pathways, the neuroactive ligand receptor interaction and olfactory transduction pathways are the most significant two. Then, we consider gene TP53 that is well-known for its role as tumor suppressor in cancer research. The log-ratio also ranges from a negative value (e.g. more expressed in non-tumor tissue) to a positive value (e.g. more expressed in tumor tissue). We divided the microarray data set again according to the expression ratio of gene TP53. After the discordance enrichment analysis, we observed overall similar results and the above two pathways are still the most significant detections. More interestingly, only these two pathways have been identified for their association with pancreatic cancer in a pathway analysis of genome-wide association study (GWAS) data. This study illustrates that some disease-related pathways can be enriched in discordant molecular behaviors when an important disease-related gene changes its expression. Our proposed statistical method is useful in the detection of these pathways. Furthermore, our method can also be applied to genome-wide expression data collected by the recent RNA-seq technology.