Quantile function regression and variable selection for sparse models

Quantile function regression and variable selection for sparse models
复制标题

DOI:
10.1002/cjs.11616
复制
发表时间:
2021-04
期刊:
Canadian Journal of Statistics
影响因子:
--
通讯作者:
Takuma Yoshida
Takuma Yoshida
中科院分区:
其他
文献类型:
--
作者:
Takuma Yoshida

文献摘要

相似文献

本文研究了高维数据的线性分位数回归和变量选择问题。一般来说,对于单个固定的分位数水平,可以得到一个普通的分位数回归估计器。因此,估计系数不具有关于分位数水平的连续性,因此,对于不同但足够接近的分位数水平,估计器和估计的有效变量集的行为可能会迅速改变。为了对给定的分位数水平得到一个稳定的估计,本研究提出了一种新的分位数回归方法来估计系数作为给定区域Δ⊂(0,1)的分位数水平的函数,称为分位数函数回归。在分位数函数回归中,我们使用B-样条模型来逼近分位数水平的系数函数,因此,估计的条件分位数是连续的,因为它是一条B-样条曲线。为了使用变量选择,使用组套索式稀疏惩罚来估计分位数级别的非零系数函数,该函数指示在Δ中保持不变的估计活动集。因此,分位数函数回归可以实现全局变量选择。该估计量在变量选择上具有渐近收敛速度和相合性。仿真研究和对真实数据的应用进一步表明,该方法具有良好的性能。
This article considers linear quantile regression and variable selection for high‐dimensional data. In general, an ordinary quantile regression estimator is obtained for a single, fixed quantile level. Therefore, the estimated coefficient does not have continuity with respect to the quantile level, and hence, the behaviour of the estimator and estimated active variable set could change rapidly for different but sufficiently close quantile levels. To obtain a stable estimator for a given quantile level, this study proposes a new quantile regression method to estimate the coefficient as a function of the quantile level of interest in a given region Δ⊂(0,1) , which is denoted quantile function regression. In quantile function regression, we approximate the coefficient function of the quantile level using a B‐spline model, and hence, the estimated conditional quantile is continuous as it is a B‐spline curve. To employ variable selection, a group lasso‐type sparse penalty is used to estimate a non‐zero coefficient function of the quantile level, which indicates the estimated active set that remains unchanged in Δ . Therefore, quantile function regression can achieve global variable selection. The proposed estimator exhibits an asymptotic rate of convergence and consistency in variable selection. Simulation studies and applications to real data further reveal that the proposed method yields good performance.