Convolutional Sequence Modeling Revisited

Convolutional Sequence Modeling Revisited
复制标题

DOI:
--
复制
发表时间:
2018-02
期刊:
--
影响因子:
--
通讯作者:
Shaojie Bai;J. Z. Kolter;V. Koltun
Shaojie Bai;J. Z. Kolter;V. Koltun
中科院分区:
其他
文献类型:
--
作者:
Shaojie Bai;J. Z. Kolter;V. Koltun

文献摘要

被引文献

相似文献

尽管卷积和递归架构在序列预测方面都有很长的历史,但目前大部分深度学习社区的“默认”思维模式是,通用序列建模最好使用递归网络来处理。然而,最近的结果表明,卷积架构在音频合成和机器翻译等任务上的表现优于递归网络。给定一个新的序列建模任务或数据集,从业者应该使用哪种架构?我们对序列建模的通用卷积和递归架构进行了系统的评估。特别是,这些模型在广泛的标准任务中进行评估,这些任务通常用于对经常性网络进行基准测试。我们的研究结果表明,简单的卷积架构在各种任务和数据集上的性能优于LSTM等典型的递归网络,同时表现出更长的有效内存。我们进一步表明,在实践中,RNN与TCN相比具有的潜在“无限记忆”优势基本上是不存在的:TCN确实表现出比其经常性对应物更长的有效历史大小。总的来说,我们认为现在可能是时候(重新)考虑ConvNets作为序列建模的默认“去”架构。
Although both convolutional and recurrent architectures have a long history in sequence prediction, the current “default” mindset in much of the deep learning community is that generic sequence modeling is best handled using recurrent networks. Yet recent results indicate that convolutional architectures can outperform recurrent networks on tasks such as audio synthesis and machine translation. Given a new sequence modeling task or dataset, which architecture should a practitioner use? We conduct a systematic evaluation of generic convolutional and recurrent architectures for sequence modeling. In particular, the models are evaluated across a broad range of standard tasks that are commonly used to benchmark recurrent networks. Our results indicate that a simple convolutional architecture outperforms canonical recurrent networks such as LSTMs across a diverse range of tasks and datasets, while demonstrating longer effective memory. We further show that the potential “infinite memory” advantage that RNNs have over TCNs is largely absent in practice: TCNs indeed exhibit longer effective history sizes than their recurrent counterparts. As a whole, we argue that it may be time to (re)consider ConvNets as the default “go to” architecture for sequence modeling.