Bandit Convex Optimization for Scalable and Dynamic IoT Management

Bandit Convex Optimization for Scalable and Dynamic IoT Management
复制标题

DOI:
10.1109/jiot.2018.2839563
复制
发表时间:
2017-07
影响因子:
10.6
通讯作者:
Tianyi Chen;G. Giannakis
Tianyi Chen;G. Giannakis
中科院分区:
计算机科学1区
文献类型:
--
作者:
Tianyi Chen;G. Giannakis

文献摘要

被引文献

相似文献

本文涉及同时包含时变损失函数和时变约束的在线凸优化问题。学习者无法完全获取损失函数,而只是在查询点显示函数值(也称为赌博机反馈)。约束在做出决策后才显示,并且可能会瞬间被违反,但从长期来看必须满足。这种设定非常适合新兴的在线网络任务,例如物联网中的雾计算,其中在线决策必须灵活适应不断变化的用户偏好(损失函数)以及资源在时间上不可预测的可用性(约束)。针对这种损失函数难以建模的人在回路系统,开发了一系列在线赌博机鞍点(BanSaP)方案,该方案根据损失函数的(可能多个)赌博机反馈以及不断变化的环境自适应地调整在线操作。这里的性能通过以下方面评估:1)推广了广泛使用的静态遗憾的动态遗憾;2)捕捉约束违反累积量的拟合度。具体而言,如果最佳动态解随时间变化缓慢,证明BanSaP可同时产生次线性动态遗憾和拟合度。雾计算卸载任务中的数值测试证实,与基于梯度反馈的现有方法相比,我们提出的BanSaP方法具有有竞争力的性能。
This paper deals with online convex optimization involving both time-varying loss functions, and time-varying constraints. The loss functions are not fully accessible to the learner, and instead only the function values (also known as bandit feedback) are revealed at queried points. The constraints are revealed after making decisions, and can be instantaneously violated, yet they must be satisfied in the long term. This setting fits nicely the emerging online network tasks such as fog computing in the Internet-of-Things, where online decisions must flexibly adapt to the changing user preferences (loss functions), and the temporally unpredictable availability of resources (constraints). Tailored for such human-in-the-loop systems where the loss functions are hard to model, a family of online bandit saddle-point (BanSaP) schemes are developed, which adaptively adjust the online operations based on (possibly multiple) bandit feedback of the loss functions, and the changing environment. Performance here is assessed by: 1) dynamic regret that generalizes the widely used static regret and 2) fit that captures the accumulated amount of constraint violations. Specifically, BanSaP is proved to simultaneously yield sublinear dynamic regret and fit, provided that the best dynamic solutions vary slowly over time. Numerical tests in fog computation offloading tasks corroborate that our proposed BanSaP approach offers competitive performance relative to existing approaches that are based on gradient feedback.