Fast Nonparametric Conditional Density Estimation

Fast Nonparametric Conditional Density Estimation
复制标题

快速非参数条件密度估计

DOI:
--
复制
发表时间:
2007
期刊:
Conference on Uncertainty in Artificial Intelligence
影响因子:
--
通讯作者:
C. Isbell
C. Isbell
中科院分区:
--
文献类型:
--
作者:
Michael P. Holmes;Alexander G. Gray;C. Isbell

文献摘要

被引文献

相似文献

条件密度估计通过对完整密度 f(yjx) 而不仅仅是期望值 E(yjx) 进行建模来概括回归。这对于许多任务都很重要,包括处理多模态和生成预测区间。尽管非参数条件密度估计是基础且广泛适用的,但它受到统计学家的关注相对较少,机器学习社区也很少或根本没有关注。这些工作都没有应用于大于二元数据,大概是由于数据驱动的带宽选择的计算困难。我们描述了双核条件密度估计器,并推导出基于双树的快速算法,用于使用最大似然准则进行带宽选择。这些技术在我们的实验中提供了高达 380 万的加速,并首次应用于以前难以处理的大型多元数据集,包括斯隆数字巡天的红移预测问题。
Conditional density estimation generalizes regression by modeling a full density f(yjx) rather than only the expected value E(yjx). This is important for many tasks, including handling multi-modality and generating prediction intervals. Though fundamental and widely applicable, nonparametric conditional density estimators have received relatively little attention from statisticians and little or none from the machine learning community. None of that work has been applied to greater than bivariate data, presumably due to the computational difficulty of data-driven bandwidth selection. We describe the double kernel conditional density estimator and derive fast dual-tree-based algorithms for bandwidth selection using a maximum likelihood criterion. These techniques give speedups of up to 3.8 million in our experiments, and enable the first applications to previously intractable large multivariate datasets, including a redshift prediction problem from the Sloan Digital Sky Survey.