Understanding the molecular information contained in principal component analysis of vibrational spectra of biological systems

Understanding the molecular information contained in principal component analysis of vibrational spectra of biological systems
复制标题

DOI:
10.1039/c1an15821j
复制
发表时间:
2012-01-01
期刊:
影响因子:
4.2
通讯作者:
Byrne, H. J.
Byrne, H. J.
中科院分区:
化学2区
文献类型:
--
作者:
Bonnier, F.;Byrne, H. J.

文献摘要

被引文献

相似文献

采用 K 均值聚类和主成分分析 (PCA) 来分析单个生物细胞的拉曼光谱图。 K 均值聚类成功地识别了细胞质、细胞核和核仁的区域,但平均光谱无法区分它们的生化组成。 PCA 识别的主要成分的载荷进一步揭示了分化的光谱基础,但它们很复杂,并且由于每个簇的光谱数量不平衡,特别是在核仁的情况下,载荷不足以代表某些细胞区域分化的基础。对结构和光谱上不同的纯生物分子的分析,对于组蛋白、神经酰胺和 RNA,以及类似的蛋白质白蛋白、胶原蛋白和组蛋白,显示了光谱负载中光谱尖锐特征的相对较强的代表性,以及随着一簇数量减少而负载的系统变化。通过光谱的加权和来模拟更复杂的细胞环境,说明虽然载荷变得越来越复杂;它们起源于构成分子成分的加权和,这一点仍然很明显。回到细胞分析,通过增加较小数量簇的光谱的权重来人为地平衡每个簇的光谱数量。虽然它使得三向分析的 PCA 加载更加复杂,但成对分析说明了已识别的亚细胞区域之间的明显差异,特别是阐明了核和核仁区域之间的分子差异。总体而言,该研究证明了对可用数据的适当考虑如何能够提高对 PCA 提供的信息的理解。
K-means clustering followed by Principal Component Analysis (PCA) is employed to analyse Raman spectroscopic maps of single biological cells. K-means clustering successfully identifies regions of cellular cytoplasm, nucleus and nucleoli, but the mean spectra do not differentiate their biochemical composition. The loadings of the principal components identified by PCA shed further light on the spectral basis for differentiation but they are complex and, as the number of spectra per cluster is imbalanced, particularly in the case of the nucleoli, the loadings under-represent the basis for differentiation of some cellular regions. Analysis of pure bio-molecules, both structurally and spectrally distinct, in the case of histone, ceramide and RNA, and similarly in the case of the proteins albumin, collagen and histone, show the relative strong representation of spectrally sharp features in the spectral loadings, and the systematic variation of the loadings as one cluster becomes reduced in number. The more complex cellular environment is simulated by weighted sums of spectra, illustrating that although the loading becomes increasingly complex; their origin in a weighted sum of the constituent molecular components is still evident. Returning to the cellular analysis, the number of spectra per cluster is artificially balanced by increasing the weighting of the spectra of smaller number clusters. While it renders the PCA loading more complex for the three-way analysis, a pair wise analysis illustrates clear differences between the identified subcellular regions, and notably the molecular differences between nuclear and nucleoli regions are elucidated. Overall, the study demonstrates how appropriate consideration of the data available can improve the understanding of the information delivered by PCA.