The Label Complexity of Active Learning from Observational Data

The Label Complexity of Active Learning from Observational Data
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Songbai Yan;Kamalika Chaudhuri;T. Javidi
Songbai Yan;Kamalika Chaudhuri;T. Javidi
中科院分区:
其他
文献类型:
--
作者:
Songbai Yan;Kamalika Chaudhuri;T. Javidi

文献摘要

相似文献

来自观测数据的反事实学习涉及基于以选择策略为条件观察到的数据来学习关于整个群体的分类器。这项工作在活动环境中考虑了这个问题,在活动环境中,学习者另外可以访问未标记的示例,并可以选择获得这些示例的子集由Oracle标记。以前关于这个问题的工作使用基于不一致的主动学习,以及一个重要性加权损失估计器来解释反事实,这导致了很高的标签复杂性。我们展示了如何将更有效的反事实风险最小化融入到主动学习算法中。这要求我们修改反事实风险,使其服从主动学习,以及修改主动学习过程,使其服从风险。我们证明了这一结果是一个在统计上一致的算法,并且比以前的工作更有效地标注。
Counterfactual learning from observational data involves learning a classifier on an entire population based on data that is observed conditioned on a selection policy. This work considers this problem in an active setting, where the learner additionally has access to unlabeled examples and can choose to get a subset of these labeled by an oracle. Prior work on this problem uses disagreement-based active learning, along with an importance weighted loss estimator to account for counterfactuals, which leads to a high label complexity. We show how to instead incorporate a more efficient counterfactual risk minimizer into the active learning algorithm. This requires us to modify both the counterfactual risk to make it amenable to active learning, as well as the active learning process to make it amenable to the risk. We provably demonstrate that the result of this is an algorithm which is statistically consistent as well as more label-efficient than prior work.