Unbiased online active learning in data streams

Unbiased online active learning in data streams
复制标题

DOI:
10.1145/2020408.2020444
复制
发表时间:
2011-08
期刊:
--
影响因子:
--
通讯作者:
Wei Chu;Martin A. Zinkevich;Lihong Li;A. Thomas;Belle L. Tseng
Wei Chu;Martin A. Zinkevich;Lihong Li;A. Thomas;Belle L. Tseng
中科院分区:
其他
文献类型:
--
作者:
Wei Chu;Martin A. Zinkevich;Lihong Li;A. Thomas;Belle L. Tseng

文献摘要

被引文献

相似文献

可以智能选择未标记的样本进行标记,以最小化分类误差。在许多实际应用程序中,大量未标记的样本以流方式到达,因此不可能在候选池中维护所有数据。在这项工作中,我们专注于二元分类问题,并研究数据流中的选择性标记,其中需要对每个样本进行顺序决策。我们考虑了抽样过程中的无偏性,并设计了最优的工具分布,以最小化随机过程中的方差。同时,在线优化加权最大似然贝叶斯线性分类器进行参数估计。在实证评估中,我们收集了某商业新闻门户网站连续30天的用户评论数据流,并进行了离线评估,比较了各种抽样策略,包括无偏主动学习、偏变量和随机抽样。实验结果验证了在线主动学习的有效性,特别是在概念漂移的非平稳情况下。
Unlabeled samples can be intelligently selected for labeling to minimize classification error. In many real-world applications, a large number of unlabeled samples arrive in a streaming manner, making it impossible to maintain all the data in a candidate pool. In this work, we focus on binary classification problems and study selective labeling in data streams where a decision is required on each sample sequentially. We consider the unbiasedness property in the sampling process, and design optimal instrumental distributions to minimize the variance in the stochastic process. Meanwhile, Bayesian linear classifiers with weighted maximum likelihood are optimized online to estimate parameters. In empirical evaluation, we collect a data stream of user-generated comments on a commercial news portal in 30 consecutive days, and carry out offline evaluation to compare various sampling strategies, including unbiased active learning, biased variants, and random sampling. Experimental results verify the usefulness of online active learning, especially in the non-stationary situation with concept drift.