Submodularity in Data Subset Selection and Active Learning

Submodularity in Data Subset Selection and Active Learning
复制标题

DOI:
--
复制
发表时间:
2015-07
期刊:
The Science of the total environment
影响因子:
--
通讯作者:
K. Wei;Rishabh K. Iyer;J. Bilmes
K. Wei;Rishabh K. Iyer;J. Bilmes
中科院分区:
其他
文献类型:
--
作者:
K. Wei;Rishabh K. Iyer;J. Bilmes

文献摘要

被引文献

相似文献

我们研究选择大数据子集来训练分类器,同时将性能损失降至最低的问题。我们展示了子模性与朴素贝叶斯(NB)和最近邻(NN)分类器的数据似然函数的联系,并将这些分类器的数据子集选择问题表述为约束子模最大化。此外,我们将该框架应用于主动学习,并提出了一种称为过滤主动子模选择(FASS)的新颖方案,其中我们将不确定性采样方法与子模数据子集选择框架相结合。我们使用四种不同的分类器(包括基于深度神经网络(DNN)的分类器)广泛评估了所提出的文本分类和手写数字识别任务框架。实证结果表明,所提出的框架在所有分类器上都比最先进的算法有了显着的改进。
We study the problem of selecting a subset of big data to train a classifier while incurring minimal performance loss. We show the connection of submodularity to the data likelihood functions for Naive Bayes (NB) and Nearest Neighbor (NN) classifiers, and formulate the data subset selection problems for these classifiers as constrained submodular maximization. Furthermore, we apply this framework to active learning and propose a novel scheme called filtered active submodular selection (FASS), where we combine the uncertainty sampling method with a submodular data subset selection framework. We extensively evaluate the proposed framework on text categorization and handwritten digit recognition tasks with four different classifiers, including deep neural network (DNN) based classifiers. Empirical results indicate that the proposed framework yields significant improvement over the state-of-the-art algorithms on all classifiers.