Feature Selection Using Distance from Classification Boundary and Monte Carlo Simulation

Feature Selection Using Distance from Classification Boundary and Monte Carlo Simulation
复制标题

DOI:
10.1007/978-3-030-04212-7_9
复制
发表时间:
2018-12
期刊:
--
影响因子:
--
通讯作者:
Yutaro Koyama;K. Ikeda;Y. Sakumura
Yutaro Koyama;K. Ikeda;Y. Sakumura
中科院分区:
其他
文献类型:
--
作者:
Yutaro Koyama;K. Ikeda;Y. Sakumura

文献摘要

被引文献

相似文献

在二值分类中,为了提高对未知样本的性能,必须尽可能多地排除代表样本的不必要的特征。在各种特征选择方法中,过滤器方法预先计算每个特征的索引,包装器方法从所有特征组合中找到具有最大性能的特征组合。在本文中,我们提出了一种新的特征选择方法,利用与分类边界的距离和蒙特卡罗模拟。提供用于二值分类的合成样本集,并在每个样本中加入由随机数确定的特征。对于这些样本集,分别使用常规方法和本文方法,并检查形成边界的特征是否被选中。我们的研究结果表明,传统的特征选择方法是困难的,而我们提出的方法是可行的。
In binary classification, to improve the performance for unknown samples, excluding as many unnecessary features representing samples as possible is necessary. Of various methods of feature selection, the filter method calculates indices beforehand for each feature, and the wrapper method finds combinations of features having the maximum performance from all combinations of features. In this paper, we propose a novel feature selection method using distance from the classification boundary and a Monte Carlo simulation. Synthetic sample sets for binary classification were provided, and features determined by random numbers were added to each sample. For these sample sets, the conventional methods and the proposed method were applied, and it was examined whether the feature forming the boundary was selected. Our results demonstrate that feature selection was difficult with the conventional methods but possible with our proposed method.