Collaboratively Learning the Best Option on Graphs, Using Bounded Local Memory

Collaboratively Learning the Best Option on Graphs, Using Bounded Local Memory
复制标题

使用有界本地内存协作学习图上的最佳选项

DOI:
10.1145/3309697.3331498
复制
发表时间:
2019
期刊:
ACM SIGMETRICS
影响因子:
--
通讯作者:
Lynch, Nancy
Lynch, Nancy
中科院分区:
--
文献类型:
--
作者:
Su, Lili;Zubeldia, Martin;Lynch, Nancy

文献摘要

参考文献

被引文献

相似文献

我们考虑社会群体中的多臂强盗问题,其中每个人的记忆有限,并且都有学习最佳臂/选项的共同目标。我们说,如果最终(如 $t\diverge$)个体只拉动具有最高预期奖励的手臂,那么它就学会了最佳选择。虽然由于记忆有限,这个目标对于孤立的个体来说是不可能的,但我们表明,在社会群体中,只要通信网络/图满足一些温和的条件,借助社会说服(即沟通)就可以轻松实现这个目标。在这项工作中,我们建模并分析了在社会群体中广泛观察到的一种学习动态。具体来说,在感兴趣的学习动态下,个体不仅基于其私人奖励反馈,而且还基于随机选择的邻居提供的建议来顺序决定接下来拉哪只手臂。为了处理奖励随机性和社交互动之间的相互作用,我们采用平均场近似方法。考虑到当通信网络不是派系时网络中的个体可能无法交换的可能性,我们超越了经典的平均场技术,并应用了平均场近似的改进版本:使用耦合我们表明,如果通信图是连通的并且是规则的或具有双随机度加权邻接矩阵,则概率→1作为社会群体大小N→∞,社会群体中的每个个体都会学习最佳选择。如果该图发散为 N → ∞,在任意但给定的有限时间范围内,描述个体意见演变的样本路径是渐近独立的。此外,持不同意见的人口比例收敛于 ODE 系统的唯一解。有趣的是,所获得的 ODE 系统对于通信图的结构是不变的。在所获得的 ODE 的解中,持有正确意见的人口比例及时以指数速度收敛到 1。值得注意的是,即使通信图高度稀疏,我们的结果仍然成立。
We consider multi-armed bandit problems in social groups wherein each individual has bounded memory and shares the common goal of learning the best arm/option. We say an individual learns the best option if eventually (as $t\diverge$) it pulls only the arm with the highest expected reward. While this goal is provably impossible for an isolated individual due to bounded memory, we show that, in social groups, this goal can be achieved easily with the aid of social persuasion (i.e., communication) as long as the communication networks/graphs satisfy some mild conditions. In this work, we model and analyze a type of learning dynamics which are well-observed in social groups. Specifically, under the learning dynamics of interest, an individual sequentially decides on which arm to pull next based on not only its private reward feedback but also the suggestion provided by a randomly chosen neighbor. To deal with the interplay between the randomness in the rewards and in the social interaction, we employ the \em mean-field approximation method. Considering the possibility that the individuals in the networks may not be exchangeable when the communication networks are not cliques, we go beyond the classic mean-field techniques and apply a refined version of mean-field approximation:Using coupling we show that, if the communication graph is connected and is either regular or has doubly-stochastic degree-weighted adjacency matrix, with probability → 1 as the social group size N → ∞, every individual in the social group learns the best option.If the minimum degree of the graph diverges as N → ∞, over an arbitrary but given finite time horizon, the sample paths describing the opinion evolutions of the individuals are asymptotically independent. In addition, the proportions of the population with different opinions converge to the unique solution of a system of ODEs. Interestingly, the obtained system of ODEs are invariant to the structures of the communication graphs. In the solution of the obtained ODEs, the proportion of the population holding the correct opinion converges to 1 exponentially fast in time.Notably, our results hold even if the communication graphs are highly sparse.
DOI: --
发表时间: 2017
期刊: Proceedings of the ACM on Measurement and Analysis of Computing Systems
影响因子: --
作者:
Debankur Mukherjee;S. Borst;J. V. Leeuwaarden
通讯作者: J. V. Leeuwaarden
有限记忆的罗宾斯-伊斯贝尔两臂老虎机问题
DOI: --
发表时间: 1965
期刊:
影响因子: --
作者:
Carter Smith;R. Pyke
通讯作者: R. Pyke
多元化共识的简单动态
DOI: --
发表时间: 2013
影响因子: 1.3
作者:
L. Becchetti;A. Clementi;Emanuele Natale;F. Pasquale;R. Silvestri;L. Trevisan
通讯作者: L. Trevisan
通过共享资源交互的系统的混沌假设
DOI: --
发表时间: 1994
期刊:
影响因子: --
作者:
C. Graham;S. Méléard
通讯作者: S. Méléard