Measuring Arithmetic Extrapolation Performance

Measuring Arithmetic Extrapolation Performance
复制标题

测量算术外推性能

DOI:
--
复制
发表时间:
2019
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
alexander rosenberg johansen
alexander rosenberg johansen
中科院分区:
--
文献类型:
--
作者:
Andreas Madsen;alexander rosenberg johansen

文献摘要

被引文献

相似文献

神经算术逻辑单元(NALU)是一个神经网络层,可以学习隐藏状态元素之间的精确算术运算。NALU的目标是学习完美的外推,这需要学习未知算术问题的确切底层逻辑。评估NALU的性能是非平凡的,因为一个算术问题可能有许多解决方案。因此,单实例MSE被用于评估和比较模型之间的性能。然而,很难解释MSE的大小代表正确的解决方案和模型对初始化的敏感性。我们建议使用一个成功准则来衡量一个模型是否和何时收敛。利用成功准则,我们可以总结多个初始化种子的成功率,并计算置信区间。本文提出了一种广义的算法基准来衡量模型在不同条件下的灵敏度。据我们所知,这是第一次对NALU及其子单元的收敛性进行广泛的评估。使用成功标准来总结4800个实验,我们发现持续学习算术外推是具有挑战性的,特别是乘法。
The Neural Arithmetic Logic Unit (NALU) is a neural network layer that can learn exact arithmetic operations between the elements of a hidden state. The goal of NALU is to learn perfect extrapolation, which requires learning the exact underlying logic of an unknown arithmetic problem. Evaluating the performance of the NALU is non-trivial as one arithmetic problem might have many solutions. As a consequence, single-instance MSE has been used to evaluate and compare performance between models. However, it can be hard to interpret what magnitude of MSE represents a correct solution and models sensitivity to initialization. We propose using a success-criterion to measure if and when a model converges. Using a success-criterion we can summarize success-rate over many initialization seeds and calculate confidence intervals. We contribute a generalized version of the previous arithmetic benchmark to measure models sensitivity under different conditions. This is, to our knowledge, the first extensive evaluation with respect to convergence of the NALU and its sub-units. Using a success-criterion to summarize 4800 experiments we find that consistently learning arithmetic extrapolation is challenging, in particular for multiplication.