A generalized kernel approach to dissimilarity-based classification

A generalized kernel approach to dissimilarity-based classification
复制标题

DOI:
10.1162/15324430260185592
复制
发表时间:
2002-03-01
影响因子:
6
通讯作者:
Duin, RPW
Duin, RPW
中科院分区:
计算机科学3区
文献类型:
--
作者:
Pekalska, E;Paclík, P;Duin, RPW

文献摘要

被引文献

相似文献

通常,要分类的对象由特征表示。在本文中,我们讨论了一种替代的对象表示的基础上相异度值。如果这样的距离很好地分离了类,那么最近邻方法提供了一个很好的解决方案。然而,在实践中使用的相异度通常是远离理想和性能的最近邻规则遭受其敏感性噪声的例子。在这种情况下,我们证明了其他更全局的分类技术优于最近邻规则。为了分类的目的,我们考虑了两种不同的使用广义相异度核的方法。在第一个中,距离等距嵌入在伪欧几里德空间中,并在那里执行分类任务。在第二种方法中,分类器直接建立在距离核上。这两种方法进行了理论上的描述,然后使用不同的相异性措施和数据集,包括退化的数据模拟缺失值的问题的实验进行比较。
Usually, objects to be classified are represented by features. In this paper, we discuss an alternative object representation based on dissimilarity values. If such distances separate the classes well, the nearest neighbor method offers a good solution. However, dissimilarities used in practice are usually far from ideal and the performance of the nearest neighbor rule suffers from its sensitivity to noisy examples. We show that other, more global classification techniques are preferable to the nearest neighbor rule, in such cases.For classification purposes, two different ways of using generalized dissimilarity kernels are considered. In the first one, distances are isometrically embedded in a pseudo-Euclidean space and the classification task is performed there. In the second approach, classifiers are built directly on distance kernels. Both approaches are described theoretically and then compared using experiments with different dissimilarity measures and datasets including degraded data simulating the problem of missing values.