Formal and Empirical Studies of Counting Behaviour in ReLU RNNs

Formal and Empirical Studies of Counting Behaviour in ReLU RNNs
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Nadine El-Naggar;Andrew Ryzhikov;Laure Daviaud;P. Madhyastha;Tillman Weyde;François Coste;Faissal Ouardi;Guillaume Rabusseau
Nadine El-Naggar;Andrew Ryzhikov;Laure Daviaud;P. Madhyastha;Tillman Weyde;François Coste;Faissal Ouardi;Guillaume Rabusseau
中科院分区:
其他
文献类型:
--
作者:
Nadine El-Naggar;Andrew Ryzhikov;Laure Daviaud;P. Madhyastha;Tillman Weyde;François Coste;Faissal Ouardi;Guillaume Rabusseau

文献摘要

相似文献

近年来,关于神经网络学习的系统性的讨论重新引起了人们的兴趣,特别是神经网络行为的形式化分析。在本文中,我们研究了单细胞ReLU RNN模型展示精确计数行为的能力。从形式上讲,我们首先描述了半Dyck-1语言和半Dyck-1计数器机器,它们可以由单个整流线性单元(ReLU)单元实现。我们在ReLU单元的权重上定义了三个计数器指示符条件(CIC),并表明满足这些条件相当于接受半Dyck-1语言,即执行精确计数。从经验上讲,我们研究了单细胞ReLU RNN通过在不同的Dyck-1和semi-Dyck-1字符串数据集上训练和测试来学习计数的能力。虽然满足CIC的网络即使在很长的字符串上也能准确地计数,因此也是正确的,但经过训练的网络会显示出各种各样的结果,并且永远不会完全满足CIC。我们调查偏离CIC的影响,并发现,满足CIC的配置是不是在最常见的设置中的损失函数的最小值。这与先前研究中的观察结果一致,即训练ReLU网络来计算任务通常会导致结果不佳。最后,我们讨论这些结果的影响和可能的途径,以改善网络行为。
In recent years, the discussion about systematicity of neural network learning has gained renewed interest, in particular the formal analysis of neural network behaviour. In this paper, we investigate the capability of single-cell ReLU RNN models to demonstrate precise counting behaviour. Formally, we start by characterising the semi-Dyck-1 language and semi-Dyck-1 counter machine that can be implemented by a single Rectified Linear Unit (ReLU) cell. We define three Counter Indicator Conditions (CICs) on the weights of a ReLU cell and show that fulfilling these conditions is equivalent to accepting the semi-Dyck-1 language, i.e. to perform exact counting. Empirically, we study the ability of single-cell ReLU RNNs to learn to count by training and testing them on different datasets of Dyck-1 and semi-Dyck-1 strings. While networks that satisfy the CICs count exactly and thus correctly even on very long strings, the trained networks exhibit a wide range of results and never satisfy the CICs exactly. We investigate the effect of deviating from the CICs and find that configurations that fulfil the CICs are not at a minimum of the loss function in the most common setups. This is consistent with observations in previous research indicating that training ReLU networks for counting tasks often leads to poor results. We finally discuss implications of these results and possible avenues for improving network behaviour.