Convolutional Sequence Modeling Revisited
Convolutional Sequence Modeling Revisited
复制标题
DOI:
--
复制
发表时间:
2018-02
期刊:
影响因子:
--
通讯作者:
Shaojie Bai;J. Z. Kolter;V. Koltun
中科院分区:
文献类型:
--
作者:
Shaojie Bai;J. Z. Kolter;V. Koltun
Although both convolutional and recurrent architectures have a long history in sequence prediction, the current “default” mindset in much of the deep learning community is that generic sequence modeling is best handled using recurrent networks. Yet recent results indicate that convolutional architectures can outperform recurrent networks on tasks such as audio synthesis and machine translation. Given a new sequence modeling task or dataset, which architecture should a practitioner use? We conduct a systematic evaluation of generic convolutional and recurrent architectures for sequence modeling. In particular, the models are evaluated across a broad range of standard tasks that are commonly used to benchmark recurrent networks. Our results indicate that a simple convolutional architecture outperforms canonical recurrent networks such as LSTMs across a diverse range of tasks and datasets, while demonstrating longer effective memory. We further show that the potential “infinite memory” advantage that RNNs have over TCNs is largely absent in practice: TCNs indeed exhibit longer effective history sizes than their recurrent counterparts. As a whole, we argue that it may be time to (re)consider ConvNets as the default “go to” architecture for sequence modeling.