PATTERN-RECOGNIZING STOCHASTIC LEARNING AUTOMATA

PATTERN-RECOGNIZING STOCHASTIC LEARNING AUTOMATA
复制标题

DOI:
10.1109/tsmc.1985.6313371
复制
发表时间:
1985-01-01
期刊:
IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS
影响因子:
--
通讯作者:
ANANDAN, P
ANANDAN, P
中科院分区:
其他
文献类型:
--
作者:
BARTO, AG;ANANDAN, P

文献摘要

被引文献

相似文献

描述了一类结合了学习自动化任务和有监督学习模式分类任务的学习任务。这些任务被称为联想强化学习任务。提出了一种关联奖惩算法,即AR-P算法,并证明了该算法具有最优性能。该算法同时推广了一类随机学习自动机和一类与Robbins-Monro随机逼近过程相关的有监督学习模式分类方法。针对学习自动机的集体行为和模式分类自适应元素网络的行为,讨论了该混合算法的相关性。仿真结果表明了联想强化学习任务和AR-P算法与几种已有算法的性能。
A class of learning tasks is described that combines aspects of learning automation tasks and supervised learning pattern-classification tasks. These tasks are called associative reinforcement learning tasks. An algorithm is presented, called the associative reward-penalty, or AR-Palgorithm for which a form of optimal performance is proved. This algorithm simultaneously generalizes a class of stochastic learning automata and a class of supervised learning pattern-classification methods related to the Robbins-Monro stochastic approximation procedure. The relevance of this hybrid algorithm is discussed with respect to the collective behaviour of learning automata and the behaviour of networks of pattern-classifying adaptive elements. Simulation results are presented that illustrate the associative reinforcement learning task and the performance of the AR-Palgorithm as compared with that of several existing algorithms.