Optimal Control of Random Walks.

Optimal Control of Random Walks.
复制标题

随机游走的最优控制。

DOI:
10.21236/ada043521
复制
发表时间:
1977
影响因子:
0.6
通讯作者:
R. Serfozo
R. Serfozo
中科院分区:
数学4区
文献类型:
--
作者:
R. Serfozo

文献摘要

被引文献

相似文献

翻译后摘要:这是一个研究的非负整数,其步骤控制如下的随机游走。在到达位置i时,从规定的集合中选择一对概率(p,q),接收奖励r(i,p,q),并且下一步骤以相应的概率p,q和1-p-q(当i=0时,这些概率是p,0和1-p)步行到位置i+1,i-1或i。这是无限期重复。用于连续选择概率(p,q)的规则是控制策略。条件确定的奖励和概率下,存在单调的最优策略的折扣和平均奖励。例如,在一种情况下,随着位置i的增加,最佳的是增加向后步骤的概率。我们的结果是基于(1)单调最优策略的一个准则,(2)凹函数的上包络为凹函数的一个结果,以及(3)折扣和平均奖励准则下的最优策略之间的关系。计算最优策略的程序也被提出。(作者)
Abstract : This is a study of a random walk on the nonnegative integers whose steps are controlled as follows. Upon arriving at a location i, a pair of probabilities (p,q) is selected from a prescribed set, a reward r(i,p,q) is received and the next step takes the walk to locations i+1, i-1, or i, with respective probabilities p, q and 1-p-q (when i=0 these probabilities are p, 0 and 1-p). This is repeated indefinitely. A rule for successively selecting the probabilities (p,q) is a control policy. Conditions are identified on the rewards and probabilities under which there exist monotonic optimal policies for discounted and average rewards. For example, in one case it is optimal to increase the probability of backward steps as the location i increases. Our results are based on (1) a criterion for monotone optimal policies, (2) a result describing when an upper envelope of concave functions is concave, and (3) a relation between optimal Policies for the discounted and average reward criteria. Procedures for computing optimal policies are also presented. (Author)