Concise Representation of Mass Spectrometry Images by Probabilistic Latent Semantic Analysis

Concise Representation of Mass Spectrometry Images by Probabilistic Latent Semantic Analysis
复制标题

DOI:
10.1021/ac801303x
复制
发表时间:
2008-12-15
影响因子:
7.4
通讯作者:
Hamprecht, Fred A.
Hamprecht, Fred A.
中科院分区:
化学1区
文献类型:
--
作者:
Hanselmann, Michael;Kirchner, Marc;Hamprecht, Fred A.

文献摘要

被引文献

相似文献

成像质谱(IMS)是一种很有前途的技术,它可以详细分析有机样品中(生物)分子的空间分布。在目前的许多应用中,国际监测系统在很大程度上依赖于(半)自动化探索性数据分析程序,将数据分解为特征成分谱和相应的丰度图,使谱和空间结构可视化。最常用的技术是主成分分析(PCA)和独立成分分析(伊卡)。这两种方法都是以无监督的方式运行的。然而,它们的分解估计通常以负计数为特征,并且不适合直接的物理解释。我们提出了概率潜在语义分析(pLSA)的非负分解和可解释的组分光谱和丰度图的说明。我们比较该算法的PCA,伊卡,和非负PARAFAC(并行因子分析),并显示在模拟和真实世界的数据,pLSA和非负PARAFAC是上级PCA或伊卡的互补性的结果组件和重建精度。我们进一步结合联合收割机pLSA分解与统计复杂性估计方案的基础上的赤池信息准则(AIC),自动估计的组件的数量存在于组织样本数据集,并表明,这导致在合理的复杂性估计。
Imaging mass spectrometry (IMS) is a promising technology which allows for detailed analysis of spatial distributions of (bio)molecules in organic samples. In many current applications, IMS relies heavily on (semi)automated exploratory data analysis procedures to decompose the data into characteristic component spectra and corresponding abundance maps, visualizing spectral and spatial structure. The most commonly used techniques are principal component analysis (PCA) and independent component analysis (ICA). Both methods operate in an unsupervised manner. However, their decomposition estimates usually feature negative counts and are not amenable to direct physical interpretation. We propose probabilistic latent semantic analysis (pLSA) for non-negative decomposition and the elucidation of interpretable component spectra and abundance maps. We compare this algorithm to PCA, ICA, and non-negative PARAFAC (parallel factors analysis) and show on simulated and real-world data that pLSA and non-negative PARAFAC are superior to PCA or ICA in terms of complementarity of the resulting components and reconstruction accuracy. We further combine pLSA decomposition with a statistical complexity estimation scheme based on the Akaike information criterion (AIC) to automatically estimate the number of components present in a tissue sample data set and show that this results in sensible complexity estimates.