XGSEA: CROSS-species gene set enrichment analysis via domain adaptation

XGSEA: CROSS-species gene set enrichment analysis via domain adaptation
复制标题

DOI:
10.1093/bib/bbaa406
复制
发表时间:
2021-01-30
影响因子:
9.5
通讯作者:
Li, Limin
Li, Limin
中科院分区:
生物学2区
文献类型:
--
作者:
Cai, Menglan;Nguyen, Canh Hao;Li, Limin

文献摘要

被引文献

相似文献

动机:基因集富集分析(GSEA)已被广泛用于识别病例和对照组之间具有统计学显著差异的基因集。GSEA需要表型标记和基因表达。然而,对模式生物的基因表达进行评估的频率高于次要种属。此外,重要的是,由于直接实验的高风险,如未经批准的治疗或基因敲除,然后通常用小鼠代替,在特定条件下不能很好地测量人类的基因表达。因此,通过使用在另一物种(源,例如小鼠)的相同表型下测量的基因表达来预测物种(靶,例如人)的给定基因集的富集显著性(在表型上)是一个重要且具有挑战性的问题,我们称之为跨物种基因集富集问题(XGSEP)。结果如下:对于XGSEP,我们提出了跨物种基因集富集分析(XGSEA),包括三个步骤:(1)对源物种运行GSEA以获得源基因集的富集分数和p值;(2)通过域适应来表示源基因集和目标基因集之间的关系;以及(3)基于(2)中的表示,使用回归来预测目标基因集的p值。我们广泛地验证了XGSEA使用五个回归和一个分类测量四个真实的数据集在不同的设置,证明XGSEA显着优于三个基线方法在大多数情况下。从小鼠ATAC-Seq数据中识别T细胞功能障碍和重编程的重要人类途径的案例研究进一步证实了XGSEA的可靠性。可用性:XGSEA的源代码可通过https://github.com/LiminLi-xjtu/XGSEA获得。
Motivation: Gene set enrichment analysis (GSEA) has been widely used to identify gene sets with statistically significant difference between cases and controls against a large gene set. GSEA needs both phenotype labels and expression of genes. However, gene expression are assessed more often for model organisms than minor species. Also, importantly gene expression are not measured well under specific conditions for human, due to high risk of direct experiments, such as non-approved treatment or gene knockout, and then often substituted by mouse. Thus, predicting enrichment significance (on a phenotype) of a given gene set of a species (target, say human), by using gene expression measured under the same phenotype of the other species (source, say mouse) is a vital and challenging problem, which we call CROSS-species gene set enrichment problem (XGSEP). Results: For XGSEP, we propose the CROSS-species gene set enrichment analysis (XGSEA), with three steps of: (1) running GSEA for a source species to obtain enrichment scores and p-values of source gene sets; (2) representing the relation between source and target gene sets by domain adaptation; and (3) using regression to predict p-values of target gene sets, based on the representation in (2). We extensively validated the XGSEA by using five regression and one classification measurements on four real data sets under various settings, proving that the XGSEA significantly outperformed three baseline methods in most cases. A case study of identifying important human pathways for T -cell dysfunction and reprogramming from mouse ATAC-Seq data further confirmed the reliability of the XGSEA. Availability: Source code of the XGSEA is available through https://github.com/LiminLi-xjtu/XGSEA.