Direct conditional probability density estimation with sparse feature selection

Direct conditional probability density estimation with sparse feature selection
复制标题

DOI:
10.1007/s10994-014-5472-x
复制
发表时间:
2015-09
期刊:
影响因子:
7.5
通讯作者:
M. Shiga;Voot Tangkaratt;Masashi Sugiyama
M. Shiga;Voot Tangkaratt;Masashi Sugiyama
中科院分区:
计算机科学3区
文献类型:
--
作者:
M. Shiga;Voot Tangkaratt;Masashi Sugiyama

文献摘要

相似文献

回归是统计数据分析中的一个基本问题,其目的是估计给定输入的输出的条件均值。然而,如果条件概率密度是多峰的,非对称的,异方差的,回归是不够的信息。为了克服这一局限性,各种估计的条件密度本身已被开发,和基于核的方法称为最小二乘条件密度估计(LS-CDE)被证明是有前途的。然而,LS-CDE仍然遭受大的估计误差,如果输入包含许多不相关的功能。因此,在本文中,我们提出了一个扩展的LS-CDE称为稀疏添加剂CDE(SA-CDE),它允许自动特征选择CDE。SA-CDE将核LS-CDE以加性方式应用于每个输入特征,并通过组稀疏正则化器惩罚整个解。我们还给出了一种基于子梯度的SA-CDE训练优化方法,该方法可以很好地扩展到高维大数据集。通过对基准和仿人机器人过渡数据集的实验,我们证明了SA-CDE在噪声CDE问题中的有用性。
Regression is a fundamental problem in statistical data analysis, which aims at estimating the conditional mean of output given input. However, regression is not informative enough if the conditional probability density is multi-modal, asymmetric, and heteroscedastic. To overcome this limitation, various estimators of conditional densities themselves have been developed, and a kernel-based approach calledleast-squares conditional density estimation(LS-CDE) was demonstrated to be promising. However, LS-CDE still suffers from large estimation error if input contains many irrelevant features. In this paper, we therefore propose an extension of LS-CDE calledsparse additive CDE(SA-CDE), which allows automatic feature selection in CDE. SA-CDE applies kernel LS-CDE to each input feature in an additive manner and penalizes the whole solution by a group-sparse regularizer. We also give a subgradient-based optimization method for SA-CDE training that scales well to high-dimensional large data sets. Through experiments with benchmark and humanoid robot transition datasets, we demonstrate the usefulness of SA-CDE in noisy CDE problems.