DISTRIBUTED INFERENCE FOR QUANTILE REGRESSION PROCESSES

DISTRIBUTED INFERENCE FOR QUANTILE REGRESSION PROCESSES
复制标题

DOI:
10.1214/18-aos1730
复制
发表时间:
2019-06-01
影响因子:
4.5
通讯作者:
Cheng, Guang
Cheng, Guang
中科院分区:
数学1区
文献类型:
--
作者:
Volgushev, Stanislav;Chao, Shih-Kang;Cheng, Guang

文献摘要

被引文献

相似文献

海量数据集的可用性增加为发现其分布中的微妙模式提供了独特的机会,但也带来了巨大的计算挑战。为了充分利用大数据中包含的信息,我们提出了一个分两步进行的过程:(I)在并行计算环境中估计不同级别的条件分位数函数;(Ii)基于估计的分位数曲线通过投影构造条件分位数回归过程。我们的一般分位数回归框架涵盖固定或增长维度的线性模型和序列近似模型。我们证明,只要分布式计算单元的数量和分位数级别选择得当,所提出的方法不会牺牲任何统计推断精度。特别地,从统计的角度出发,我们得到了前者的一个尖锐的上界和后者的一个尖锐的下界,以获得最小的计算量。作为一个重要的应用,考虑了条件分布函数的统计推断。此外,我们还提出了在上述分布式估计环境下进行推理的计算高效方法。这些方法直接利用子样本估计器的可用性,几乎不需要额外的计算成本。模拟证实了我们的统计推断理论。
The increased availability of massive data sets provides a unique opportunity to discover subtle patterns in their distributions, but also imposes overwhelming computational challenges. To fully utilize the information contained in big data, we propose a two-step procedure: (i) estimate conditional quantile functions at different levels in a parallel computing environment; (ii) construct a conditional quantile regression process through projection based on these estimated quantile curves. Our general quantile regression framework covers both linear models with fixed or growing dimension and series approximation models. We prove that the proposed procedure does not sacrifice any statistical inferential accuracy provided that the number of distributed computing units and quantile levels are chosen properly. In particular, a sharp upper bound for the former and a sharp lower bound for the latter are derived to capture the minimal computational cost from a statistical perspective. As an important application, the statistical inference on conditional distribution functions is considered. Moreover, we propose computationally efficient approaches to conducting inference in the distributed estimation setting described above. Those approaches directly utilize the availability of estimators from subsamples and can be carried out at almost no additional computational cost. Simulations confirm our statistical inferential theory.