Linear operator-based statistical analysis: A useful paradigm for big data

Linear operator-based statistical analysis: A useful paradigm for big data
复制标题

DOI:
10.1002/cjs.11329
复制
发表时间:
2018-03-01
影响因子:
0.6
通讯作者:
Li, Bing
Li, Bing
中科院分区:
数学4区
文献类型:
--
作者:
Li, Bing

文献摘要

被引文献

相似文献

在本文中,我们列出了基于线性运算符的统计分析的一些基本结构、技术机制和关键应用,并将它们组织成一个统一的范例。由于线性运算符的性质:它们成批处理大量函数,这种范式可以在分析大数据方面发挥重要作用。该系统至少支持四种统计设置:多变量数据分析、函数数据分析、通过核学习的非线性多变量数据分析、以及通过核学习的非线性函数数据分析。我们在每个统计背景下发展了五个线性算子:协方差算子、相关算子、条件协方差算子、回归算子和偏相关算子,这为我们以非参数和全面的方式研究随机变量或随机函数之间的相互关系提供了有力的手段。我们给出了一个追踪充分降维发展的案例研究,并详细描述了这些线性算子如何在其最近的发展中扮演越来越重要的角色。我们还提出了一种坐标映射方法,该方法可以系统地在样本级实现这些算子。《加拿大统计杂志》46:79-103;2018(C)2017年加拿大统计学会
In this article we lay out some basic structures, technical machineries, and key applications, of Linear Operator-Based Statistical Analysis, and organize them toward a unified paradigm. This paradigm can play an important role in analyzing big data due to the nature of linear operators: they process large number of functions in batches. The system accommodates at least four statistical settings: multivariate data analysis, functional data analysis, nonlinear multivariate data analysis via kernel learning, and nonlinear functional data analysis via kernel learning. We develop five linear operators within each statistical setting: the covariance operator, the correlation operator, the conditional covariance operator, the regression operator, and the partial correlation operator, which provide us with a powerful means to study the interconnections between random variables or random functions in a nonparametric and comprehensive way. We present a case study tracing the development of sufficient dimension reduction, and describe in detail how these linear operators play increasingly critical roles in its recent development. We also present a coordinate mapping method which can be systematically applied to implement these operators at the sample level. The Canadian Journal of Statistics 46: 79-103; 2018 (c) 2017 Statistical Society of Canada