Learning from Incomplete Features by Simultaneous Training of Neural Networks and Sparse Coding

Learning from Incomplete Features by Simultaneous Training of Neural Networks and Sparse Coding
复制标题

DOI:
10.1109/cvprw53098.2021.00296
复制
发表时间:
2021-06
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
C. Caiafa;Ziyao Wang;Jordi Solé-Casals;Qibin Zhao
C. Caiafa;Ziyao Wang;Jordi Solé-Casals;Qibin Zhao
中科院分区:
其他
文献类型:
--
作者:
C. Caiafa;Ziyao Wang;Jordi Solé-Casals;Qibin Zhao

文献摘要

相似文献

在本文中,训练一个分类器上的数据集与不完整的功能的问题。我们假设每个数据实例都有不同的特征子集(随机或结构化)。这种情况通常发生在应用程序中,即不是为每个数据样本收集所有特征时。开发了一种新的监督学习方法来训练通用分类器,例如逻辑回归或深度神经网络,仅使用每个样本的特征子集,同时假设未知字典上的数据向量的稀疏表示。充分的条件被确定,使得,如果有可能在不完整的观察上训练分类器,使得它们的重建被超平面很好地分离,则相同的分类器也正确地分离原始(未观察到的)数据样本。在合成数据集和知名数据集上的大量仿真结果验证了我们的理论研究结果,并与传统的数据插补方法和一种最先进的算法相比,证明了所提出的方法的有效性。
In this paper, the problem of training a classifier on a dataset with incomplete features is addressed. We assume that different subsets of features (random or structured) are available at each data instance. This situation typically occurs in the applications when not all the features are collected for every data sample. A new supervised learning method is developed to train a general classifier, such as a logistic regression or a deep neural network, using only a subset of features per sample, while assuming sparse representations of data vectors on an unknown dictionary. Sufficient conditions are identified, such that, if it is possible to train a classifier on incomplete observations so that their reconstructions are well separated by a hyperplane, then the same classifier also correctly separates the original (unobserved) data samples. Extensive simulation results on synthetic and well-known datasets are presented that validate our theoretical findings and demonstrate the effectiveness of the proposed method compared to traditional data imputation approaches and one state-of-the-art algorithm.