Optima Query Selection Using Multi-Armed Bandits

Optima Query Selection Using Multi-Armed Bandits
复制标题

DOI:
10.1109/lsp.2018.2878066
复制
发表时间:
2018-12-01
影响因子:
3.9
通讯作者:
Erdogmus, Deniz
Erdogmus, Deniz
中科院分区:
工程技术2区
文献类型:
--
作者:
Kocanaogullari, Aziz;Marghi, Yeganeh M.;Erdogmus, Deniz

文献摘要

被引文献

相似文献

潜变量估计的查询选择通常是通过选择具有低噪声的观测值或优化与基于当前最佳估计降低估计不确定性水平相关的信息论目标来执行的。在这些方法中,系统通常通过利用有关状态的当前可用信息做出决策。然而,当真相与当前估计相差甚远时,信任当前的估计会导致查询选择不佳,这会对潜在变量估计过程的速度和准确性产生负面影响。我们引入了一种新颖的顺序自适应动作值函数,用于使用多臂老虎机框架进行查询选择,这使我们能够找到易于处理的解决方案。对于这种自适应顺序查询选择方法,我们分析表明:1)动态系统查询选择的性能改进; 2)模型优于竞争对手的条件。与其他方法相比,我们还对该方法的性能进行了有利的实证评估,这两种方法都使用蒙特卡罗模拟和脑机接口打字系统的人机循环实验,其中语言模型提供了先验信息。
Query selection for latent variable estimation is conventionally performed by opting for observations with low noise or optimizing information-theoretic objectives related to reducing the level of estimated uncertainty based on the current best estimate. In these approaches, typically, the system makes a decision by leveraging the current available information about the state. However, trusting the current hest estimate results in poor query selection when truth is far from the current estimate, and this negatively impacts the speed and accuracy of the latent variable estimation procedure. We introduce a novel sequential adaptive action value function for query selection using the multi-armed bandit framework, which allows us to find a tractable solution. For this adaptive-sequential query selection method, we analytically show: 1) performance improvement in the query selection for a dynamical system; and 2) the conditions where the model outperforms competitors. We also present favorable empirical assessments of the performance for this method, compared to alternative methods, both using Monte Carlo simulations and human-in-the-loop experiments with a brain-computer interface typing system, where the language model provides the prior information.