Sequential Learning without Feedback

Sequential Learning without Feedback
复制标题

无反馈的顺序学习

DOI:
10.1145/3179416
复制
发表时间:
2016
期刊:
ArXiv
影响因子:
--
通讯作者:
Venkatesh Saligrama
Venkatesh Saligrama
中科院分区:
--
文献类型:
--
作者:
M. Hanawal;Csaba Szepesvari;Venkatesh Saligrama

文献摘要

被引文献

相似文献

在许多安全和医疗保健系统中,一系列特征/传感器/测试用于检测和诊断。每个测试输出一个潜在状态的预测,并带有固有的成本。我们的目标是学习选择测试的策略,以优化准确性和成本。不幸的是,它往往是不可能获得现场地面实况注释,我们留下了无监督传感器选择(USS)的问题。我们提出USS作为一个版本的随机部分监控问题的奖励结构(即使是嘈杂的注释是不可用的)。毫不奇怪,没有学习者能在没有进一步假设的情况下达到次线性后悔。为此,我们提出了弱优势的概念。这是测试输出和潜在状态的联合概率分布的一个条件,也就是说,只要一个测试在一个例子上是准确的,序列中后面的测试也可能是准确的。我们实证验证了弱优势在真实的数据集上成立,并证明了它是实现次线性后悔的最大条件。我们减少USS的一个特殊情况下,多臂强盗问题的边信息和开发多项式时间算法,实现次线性遗憾。
In many security and healthcare systems a sequence of features/sensors/tests are used for detection and diagnosis. Each test outputs a prediction of the latent state, and carries with it inherent costs. Our objective is to {\it learn} strategies for selecting tests to optimize accuracy \& costs. Unfortunately it is often impossible to acquire in-situ ground truth annotations and we are left with the problem of unsupervised sensor selection (USS). We pose USS as a version of stochastic partial monitoring problem with an {\it unusual} reward structure (even noisy annotations are unavailable). Unsurprisingly no learner can achieve sublinear regret without further assumptions. To this end we propose the notion of weak-dominance. This is a condition on the joint probability distribution of test outputs and latent state and says that whenever a test is accurate on an example, a later test in the sequence is likely to be accurate as well. We empirically verify that weak dominance holds on real datasets and prove that it is a maximal condition for achieving sublinear regret. We reduce USS to a special case of multi-armed bandit problem with side information and develop polynomial time algorithms that achieve sublinear regret.