Semi-supervised Nonnegative Matrix Factorization for gene expression deconvolution: A case study

Semi-supervised Nonnegative Matrix Factorization for gene expression deconvolution: A case study
复制标题

DOI:
10.1016/j.meegid.2011.08.014
复制
发表时间:
2012-07-01
影响因子:
3.2
通讯作者:
Seoighe, Cathal
Seoighe, Cathal
中科院分区:
医学3区
文献类型:
--
作者:
Gaujoux, Renaud;Seoighe, Cathal

文献摘要

被引文献

相似文献

在许多基因表达研究中,样品组成中的杂合性是一个固有的问题,在许多情况下,在下游分析中应考虑到这一点,以正确解释潜在的生物学过程。典型的例子是使用血液样本的传染病或免疫学相关研究,其中,例如,淋巴细胞亚群的比例预计在病例和对照之间会有所不同。非负矩阵因子分解(NMF)是一种无监督学习技术,已成功应用于多个领域,特别是在生物信息学中,已经证明了其从诸如基因表达微阵列的高维数据中提取有意义的信息的能力。由于基本上是无监督的,标准NMF方法不能保证找到与样品中感兴趣的细胞类型相对应的组分,这可能会危及对细胞比例的正确估计。我们已经研究了使用先验知识,在一组标记基因的形式,以提高基因表达反卷积与NMF算法。我们发现,这提高了估计细胞类型比例和细胞类型基因表达特征的一致性。所提出的方法进行了测试的微阵列数据集组成的纯细胞类型混合在已知的比例。真实和估计的细胞类型比例之间的皮尔逊相关系数利用常用NMF算法的半监督(标记物引导)版本显著改善(通常从约0.5至约0.8)。此外,与每种细胞类型相关的已知标记基因被更频繁地分配给指导版本的正确细胞类型。我们得出结论,使用标记基因提高了使用NMF的基因表达反卷积的准确性,并建议修改标记基因信息的使用方式,这可能会导致进一步的改进。(C)2011 Elsevier B.V.保留所有权利。
Heterogeneity in sample composition is an inherent issue in many gene expression studies and, in many cases, should be taken into account in the downstream analysis to enable correct interpretation of the underlying biological processes. Typical examples are infectious diseases or immunology-related studies using blood samples, where, for example, the proportions of lymphocyte sub-populations are expected to vary between cases and controls.Nonnegative Matrix Factorization (NMF) is an unsupervised learning technique that has been applied successfully in several fields, notably in bioinformatics where its ability to extract meaningful information from high-dimensional data such as gene expression microarrays has been demonstrated. Very recently, it has been applied to biomarker discovery and gene expression deconvolution in heterogeneous tissue samples.Being essentially unsupervised, standard NMF methods are not guaranteed to find components corresponding to the cell types of interest in the sample, which may jeopardize the correct estimation of cell proportions. We have investigated the use of prior knowledge, in the form of a set of marker genes, to improve gene expression deconvolution with NMF algorithms. We found that this improves the consistency with which both cell type proportions and cell type gene expression signatures are estimated. The proposed method was tested on a microarray dataset consisting of pure cell types mixed in known proportions. Pearson correlation coefficients between true and estimated cell type proportions improved substantially (typically from about 0.5 to approximately 0.8) with the semi-supervised (marker-guided) versions of commonly used NMF algorithms. Furthermore known marker genes associated with each cell type were assigned to the correct cell type more frequently for the guided versions. We conclude that the use of marker genes improves the accuracy of gene expression deconvolution using NMF and suggest modifications to how the marker gene information is used that may lead to further improvements. (C) 2011 Elsevier B.V. All rights reserved.