Monotonicity of Constrained Optimal Transmission Policies in Correlated Fading Channels With ARQ

Monotonicity of Constrained Optimal Transmission Policies in Correlated Fading Channels With ARQ
复制标题

DOI:
10.1109/tsp.2009.2027735
复制
发表时间:
2010
影响因子:
5.4
通讯作者:
M. Ngo;V. Krishnamurthy
M. Ngo;V. Krishnamurthy
中科院分区:
工程技术1区
文献类型:
--
作者:
M. Ngo;V. Krishnamurthy

文献摘要

被引文献

相似文献

我们考虑使用ARQ协议的传输调度与重传给定的信道状态信息(CSI)和相关的衰落信道。该问题被表述为一个可数状态、无限时域、平均费用的马尔可夫决策过程(MDP),具有平均延迟约束。我们的主要结果是给出了信道记忆和传输成本的充分条件,使得最优传输调度策略是缓冲区占用率的单调递增函数。在证明这个结果时,我们首先证明缓冲区的正递归(稳定性)。单调结构证明包括两个步骤。首先,约束MDP(CMDP)转化为无约束MDP使用拉格朗日动态规划公式。证明了无约束最优策略是纯的,且在缓冲区占用率上单调增加。然后表明,约束最优策略是两个纯传输策略的随机混合,在缓冲区状态是单调的。最后,利用最优传输策略的单调结构,推导出一个单调策略Q -学习算法和一个基于随机逼近的单调策略搜索算法,用于真实的时间估计最优策略。
We consider transmission scheduling using an ARQ protocol with retransmissions given channel state information (CSI) and a correlated fading channel. The problem is formulated as a countable state, infinite horizon, average cost Markov decision process (MDP) with an average delay constraint. Our main result is to give sufficient conditions on the channel memory, and transmission cost so that the optimal transmission scheduling policy is a monotonically increasing function of the buffer occupancy. In proving this result, we first prove positive recurrence (stability) of the buffer. The monotone structure proof consists of two steps. First, the constrained MDP (CMDP) is transformed into an unconstrained MDP using a Lagrangian dynamic programming formulation. It is proved that the unconstrained optimal policy is pure and monotonically increasing in the buffer occupancy. It is then shown that the constrained optimal policy is a randomized mixture of two pure transmission policies that are monotone in the buffer state. Finally, the monotone structure of the optimal transmission policy is exploited to derive a monotone-policy Q -learning algorithm and a stochastic approximation based monotone policy search algorithm for estimating the optimal policy in real time.