A PARTIALLY LINEAR FRAMEWORK FOR MASSIVE HETEROGENEOUS DATA.

A PARTIALLY LINEAR FRAMEWORK FOR MASSIVE HETEROGENEOUS DATA.
复制标题

DOI:
10.1214/15-aos1410
复制
发表时间:
2016-08
影响因子:
4.5
通讯作者:
Liu H
Liu H
中科院分区:
数学1区
文献类型:
--
作者:
Zhao T;Cheng G;Liu H

文献摘要

被引文献

相似文献

我们考虑一个部分线性的框架建模大量的异构数据。主要目标是提取所有子群体的共同特征,同时探索每个子群体的异质性。特别是,我们提出了一个聚集型估计的共性参数,拥有(非渐近)极大极小最优界和渐近分布,如果没有异质性。当子种群的数量增长不太快时,这个预言性的结果成立。一个插件估计的异质性参数进一步构建,并具有渐近分布,如果共性信息是可用的。我们还测试了大量子群体之间的异质性。所有上述结果都需要正则化每个子估计,就好像它有整个样本量一样。我们的一般理论适用于分而治之的方法,通常用于处理大量的同质数据。本文的一个技术副产品是一般核岭回归的统计推断。最后给出了数值结果来支持我们的理论。
We consider a partially linear framework for modelling massive heterogeneous data. The major goal is to extract common features across all sub-populations while exploring heterogeneity of each sub-population. In particular, we propose an aggregation type estimator for the commonality parameter that possesses the (non-asymptotic) minimax optimal bound and asymptotic distribution as if there were no heterogeneity. This oracular result holds when the number of sub-populations does not grow too fast. A plug-in estimator for the heterogeneity parameter is further constructed, and shown to possess the asymptotic distribution as if the commonality information were available. We also test the heterogeneity among a large number of sub-populations. All the above results require to regularize each sub-estimation as though it had the entire sample size. Our general theory applies to the divide-and-conquer approach that is often used to deal with massive homogeneous data. A technical by-product of this paper is the statistical inferences for the general kernel ridge regression. Thorough numerical results are also provided to back up our theory.