Autonomous mobile acoustic relay positioning as a multi-armed bandit with switching costs

Autonomous mobile acoustic relay positioning as a multi-armed bandit with switching costs
复制标题

自主移动声学中继定位为具有切换成本的多臂强盗

DOI:
10.1109/iros.2013.6696836
复制
发表时间:
2013
期刊:
2013 IEEE/RSJ International Conference on Intelligent Robots and Systems
影响因子:
--
通讯作者:
F. Hover
F. Hover
中科院分区:
--
文献类型:
--
作者:
Mei Yi Cheung;J. Leighton;F. Hover

文献摘要

被引文献

相似文献

水声通信信道具有高度的随机性和多变性,特别是在多径受限的浅海和港口环境中。然而,移动的声学节点可以在其移动时学习信道的属性。通过自适应节点定位最大化累积数据传输是一个干净的利用与探索的情况下,因为学习差的特征位置必须与利用已知的平衡。虽然这个问题是很好地描述了随机多臂土匪形式主义,经典的假设,无成本切换是站不住脚的领域,缓慢移动的车辆往往覆盖很大的距离。我们提出了一个启发式的适应MAB Gittins指数规则与有限的政策枚举占转换成本,并描述了在查尔斯河(马萨诸塞州波士顿)进行的实地实验。现场数据表明,MAB及其切换成本扩展在此应用中是易处理的,并且性能始终上级贪婪策略。
Underwater acoustic communication channels display highly variable and stochastic performance, especially in multipath-limited shallow-water and harbor environments. A mobile acoustic node can, however, learn the channel's properties as it moves about. Maximizing the cumulative data transmission through adaptive node positioning is a clean exploitation vs. exploration scenario because learning about poorly characterized locations must be balanced against exploiting known ones. While this problem is well described with the stochastic multi-armed bandit formalism, the classical assumption of costless switching is untenable in the field, where slow-moving vehicles often cover large distances. We present a heuristic adaptation to the MAB Gittins index rule with limited policy enumeration to account for switching costs, and describe field experiments conducted in the Charles River (Boston MA). The field data establish that the MAB and its switching cost extension are tractable in this application, and that performance is consistently superior to that of ϵ-greedy policies.