Exploiting Multiple Timescales in Hierarchical Echo State Networks

Exploiting Multiple Timescales in Hierarchical Echo State Networks
复制标题

DOI:
10.3389/fams.2020.616658
复制
发表时间:
2021-02-17
影响因子:
1.4
通讯作者:
Vasilaki, Eleni
Vasilaki, Eleni
中科院分区:
其他
文献类型:
--
作者:
Manneschi, Luca;Ellis, Matthew O. A.;Vasilaki, Eleni

文献摘要

被引文献

相似文献

回声状态网络(ESN)是一种强大的水库计算形式,它只需要训练线性输出权重,而内部水库由固定的随机连接的神经元组成。通过正确缩放的连接矩阵,神经元的活动表现出回声状态特性,并以一定的时间尺度响应于输入动态。调整网络的时间尺度对于处理某些任务可能是必要的,并且某些环境需要多个时间尺度以实现有效的表示。在这里,我们探讨的时间尺度分层ESN,其中水库被划分为两个较小的连接水库具有不同的属性。在三个不同的任务(NARMA 10,在一个不稳定的环境中的重建任务,和psMNIST),我们表明,通过选择每个分区的超参数,使他们专注于不同的时间尺度,我们实现了一个显着的性能改善,在一个单一的ESN。通过线性分析,并假设第一个分区的时间尺度比第二个分区的时间尺度短得多,(通常对应于最佳操作条件),我们根据由第一分区提供给第二分区的输入信号的有效表示来解释分区的前馈耦合,由此瞬时输入信号被扩展为其时间导数的加权组合。此外,我们提出了一种数据驱动的方法,通过梯度下降优化方法来优化超参数,该方法是通过时间的反向传播的在线近似。我们展示了在线学习规则在所有考虑的任务中的应用。
Echo state networks (ESNs) are a powerful form of reservoir computing that only require training of linear output weights while the internal reservoir is formed of fixed randomly connected neurons. With a correctly scaled connectivity matrix, the neurons' activity exhibits the echo-state property and responds to the input dynamics with certain timescales. Tuning the timescales of the network can be necessary for treating certain tasks, and some environments require multiple timescales for an efficient representation. Here we explore the timescales in hierarchical ESNs, where the reservoir is partitioned into two smaller linked reservoirs with distinct properties. Over three different tasks (NARMA10, a reconstruction task in a volatile environment, and psMNIST), we show that by selecting the hyper-parameters of each partition such that they focus on different timescales, we achieve a significant performance improvement over a single ESN. Through a linear analysis, and under the assumption that the timescales of the first partition are much shorter than the second's (typically corresponding to optimal operating conditions), we interpret the feedforward coupling of the partitions in terms of an effective representation of the input signal, provided by the first partition to the second, whereby the instantaneous input signal is expanded into a weighted combination of its time derivatives. Furthermore, we propose a data-driven approach to optimise the hyper-parameters through a gradient descent optimisation method that is an online approximation of backpropagation through time. We demonstrate the application of the online learning rule across all the tasks considered.