Spectral feature scaling method for supervised dimensionality reduction

Spectral feature scaling method for supervised dimensionality reduction
复制标题

DOI:
10.24963/ijcai.2018/355
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Momo Matsuda;K. Morikuni;T. Sakurai
Momo Matsuda;K. Morikuni;T. Sakurai
中科院分区:
其他
文献类型:
--
作者:
Momo Matsuda;K. Morikuni;T. Sakurai

文献摘要

被引文献

相似文献

光谱降维方法能够在降维空间中对具有高维特征的复杂数据进行线性分离。然而,由于数据的不规则性或不确定性,这些方法并不总是给出期望的结果。因此,我们考虑积极地修改特征的尺度以获得所需的分类。利用部分样本标签上的先验知识来指定Fiedler向量,给出了特征向量具有特征标度因子的线性矩阵铅笔的特征值问题。根据已知的标签,由此产生的因子可以修改整个样本的特征,在约简空间中形成聚类。在这项研究中,我们提出了新的降维方法,使用与光谱聚类相关的特征尺度进行监督。数值实验表明,对于样本多于特征的玩具问题,本文提出的方法优于已有的监督方法,并且在聚类方面比现有方法具有更强的鲁棒性。此外,所提出的方法优于现有的方法,对于现实世界问题的分类,具有比癌症疾病的基因表达谱样本更多的特征。此外,随着训练数据比例的增加,特征缩放倾向于提高现有无监督方法的聚类和分类精度。
Spectral dimensionality reduction methods enable linear separations of complex data with high-dimensional features in a reduced space. However, these methods do not always give the desired results due to irregularities or uncertainties of the data. Thus, we consider aggressively modifying the scales of the features to obtain the desired classification. Using prior knowledge on the labels of partial samples to specify the Fiedler vector, we formulate an eigenvalue problem of a linear matrix pencil whose eigenvector has the feature scaling factors. The resulting factors can modify the features of entire samples to form clusters in the reduced space, according to the known labels. In this study, we propose new dimensionality reduction methods supervised using the feature scaling associated with the spectral clustering. Numerical experiments show that the proposed methods outperform well-established supervised methods for toy problems with more samples than features, and are more robust regarding clustering than existing methods. Also, the proposed methods outperform existing methods regarding classification for real-world problems with more features than samples of gene expression profiles of cancer diseases. Furthermore, the feature scaling tends to improve the clustering and classification accuracies of existing unsupervised methods, as the proportion of training data increases.