Active Search using Meta-Bandits

Active Search using Meta-Bandits
复制标题

DOI:
10.1145/3340531.3417409
复制
发表时间:
2020-10
期刊:
Proceedings of the 29th ACM International Conference on Information & Knowledge Management
影响因子:
--
通讯作者:
Shengli Zhu;Jakob Coles;Sihong Xie
Shengli Zhu;Jakob Coles;Sihong Xie
中科院分区:
其他
文献类型:
--
作者:
Shengli Zhu;Jakob Coles;Sihong Xie

文献摘要

相似文献

在许多应用中,积极的实例很少,但识别起来很重要。例如,在自然语言处理中,给定关系的肯定句在大型语料库中很少。在这些应用程序中,积极的数据对于学习来说更有信息量,但在给一定量的数据贴上标签之前,人们不知道从哪里可以找到罕见的积极的数据。由于随机抽样可能会导致标记工作的显著浪费,以前的“主动搜索”方法使用单一的强盗模型来了解数据分布(探索),同时从可能包含更多积极因素的区域进行抽样(利用)。许多强盗模型是可能的,次优模型降低了标记效率,但在任何数据被标记之前,最优模型是未知的。我们提出了Meta-AS(Meta Active Search),它使用一个Meta-Bandit来评估一组基本盗贼,旨在有效地标记正例,而不是后见之明的最优基本盗贼。元强盗估计基本强盗的性能的均值和方差,并选择一个基本强盗来建议下一步标记哪些数据以进行勘探或利用。标签中的反馈更新了下一轮的基础土匪和元土匪。Meta-AS可以容纳一组不同的基本强盗来探索关于数据集的假设,而不会在标记开始之前过度承诺单一模型。在五个数据集上的关系抽取实验表明,Meta-AS比基本强盗和其他强盗选择策略更有效地进行正向标记。
There are many applications where positive instances are rare but important to identify. For example, in NLP, positive sentences for a given relation are rare in a large corpus. Positive data are more informative for learning in these applications, but before one labels a certain amount of data, it is unknown where to find the rare positives. Since random sampling can lead to significant waste in labeling effort, previous 'active search' methods use a single bandit model to learn about the data distribution (exploration) while sampling from the regions potentially containing more positives (exploitation). Many bandit models are possible and a sub-optimal model reduces labeling efficiency, but the optimal model is unknown before any data are labeled. We propose Meta-AS (Meta Active Search) that uses a meta-bandit to evaluate a set of base bandits and aims to label positive examples efficiently, comparing to the optimal base bandit with hindsight. The meta-bandit estimates the mean and variance of the performance of the base bandits and selects a base bandit to propose what data to label next for exploration or exploitation. The feedback in the labels updates both the base bandits and the meta-bandit for the next round. Meta-AS can accommodate a diverse set of base bandits to explore assumptions about the dataset, without over-committing to a single model before labeling starts. Experiments on five datasets for relation extraction demonstrate that Meta-AS labels positives more efficiently than the base bandits and other bandit selection strategies.