Principal component analysis

Principal component analysis
复制标题

DOI:
10.1038/s43586-022-00184-w
复制
发表时间:
2022-12-22
期刊:
NATURE REVIEWS METHODS PRIMERS
影响因子:
--
通讯作者:
Tuzhilina, Elena
Tuzhilina, Elena
中科院分区:
其他
文献类型:
--
作者:
Greenacre, Michael;Groenen, Patrick J. F.;Tuzhilina, Elena

文献摘要

被引文献

相似文献

主成分分析是一种通用的统计方法,用于将一个由案例和变量构成的数据表简化为其基本特征,即主成分。主成分是原始变量的少数几个线性组合,它们能最大程度地解释所有变量的方差。在此过程中,该方法仅使用这几个主要成分来近似原始数据表。本入门教程全面综述了该方法的定义和几何原理,以及对其数值和图形结果的解释。主要的图形结果通常以双标图的形式呈现,利用主要成分来绘制案例,并添加原始变量以支持对案例位置的距离解释。还介绍了该方法的变体,例如分组数据和分类数据的分析,即对应分析。同时描述和展示了主成分分析的最新创新应用:用于估计大型数据矩阵中的缺失值、稀疏成分估计以及图像、形状和函数的分析。补充材料包括视频动画和R环境中的计算机脚本。
Principal component analysis is a versatile statistical method for reducing a cases-by-variables data table to its essential features, called principal components. Principal components are a few linear combinations of the original variables that maximally explain the variance of all the variables. In the process, the method provides an approximation of the original data table using only these few major components. This Primer presents a comprehensive review of the method's definition and geometry, as well as the interpretation of its numerical and graphical results. The main graphical result is often in the form of a biplot, using the major components to map the cases and adding the original variables to support the distance interpretation of the cases' positions. Variants of the method are also treated, such as the analysis of grouped data and categorical data, known as correspondence analysis. Also described and illustrated are the latest innovative applications of principal component analysis: for estimating missing values in huge data matrices, sparse component estimation, and the analysis of images, shapes and functions. Supplementary material includes video animations and computer scripts in the R environment.