Estimating the support of a high-dimensional distribution

Estimating the support of a high-dimensional distribution
复制标题

DOI:
10.1162/089976601750264965
复制
发表时间:
2001-07-01
期刊:
影响因子:
2.9
通讯作者:
Williamson, RC
Williamson, RC
中科院分区:
计算机科学4区
文献类型:
--
作者:
Schölkopf, B;Platt, JC;Williamson, RC

文献摘要

被引文献

相似文献

假设给定一个从概率分布P中抽取的数据集,并希望估计输入空间的一个“简单”子集S,使得从P中抽取的测试点位于S之外的概率等于0到1之间的某个先验指定值. f的函数形式由训练数据的潜在小子集的核扩展给出;它通过控制相关特征空间中的权重向量的长度来正则化。通过求解一个二次规划问题,我们通过对输入模式进行顺序优化来找到扩展系数。我们还对算法的统计性能进行了理论分析,该算法是支持向量机算法在无标记数据情况下的自然扩展。
Suppose you are given some data set drawn from an underlying probability distribution P and you want to estimate a "simple" subset S of input space such that the probability that a test point drawn from P lies outside of S equals some a priori specified value between 0 and 1.We propose a method to approach this problem by trying to estimate a function f that is positive on S and negative on the complement. The functional form of f is given by a kernel expansion in terms of a potentially small subset of the training data; it is regularized by controlling the length of the weight vector in an associated feature space. The expansion coefficients are found by solving a quadratic programming problem, which we do by carrying out sequential optimization over pairs of input patterns. We also provide a theoretical analysis of the statistical performance of our algorithm.The algorithm is a natural extension of the support vector algorithm to the case of unlabeled data.