Cache friendly parallelization of neural encoder-decoder models without padding on multi-core architecture.

Cache friendly parallelization of neural encoder-decoder models without padding on multi-core architecture.
复制标题

神经编码器-解码器模型的缓存友好并行化,无需在多核架构上进行填充。

DOI:
10.1109/ipdpsw.2017.165
复制
发表时间:
2017
期刊:
The 6th International Workshop on Parallel and Distributed Computing for Large Scale Machine Learning and Big Data Analytics
影响因子:
--
通讯作者:
and Kenjiro Taura.
and Kenjiro Taura.
中科院分区:
--
文献类型:
--
作者:
Yuchen Qiao;Kazuma Hashimoto;Akiko Eriguchi;Haixia Wang;Dongsheng Wang;Yoshimasa Tsuruoka;and Kenjiro Taura.

文献摘要

相似文献

针对大规模数据集扩展人工智能(AI)算法以提高其性能变得至关重要。机器翻译是人工智能的重要研究领域之一,近年来,基于递归神经网络(RNN)的机器翻译模型表现出了最新的性能,许多研究人员一直致力于改进基于RNN的机器翻译模型,以提高机器翻译的准确性。神经机器翻译(NMT)模型的大多数实现在处理小批量时采用填充策略,以使小批量中的所有句子具有相同的长度。这使得高速缓存和GPU/SIMD并行性的有效利用成为可能,但会导致计算时间的浪费。在本文中,我们实现了一个序列到序列(Seq 2Seq)模型,这是最基本的模型NMT,不使用填充策略的批处理学习和并行化。更具体地说,我们的方法在处理一个句子时,将表示输入单词以及神经网络在不同时间步的状态的向量形成矩阵,因此,该方法可以更好地利用缓存并优化在反向传播阶段调整权重和偏差的过程。我们的实验评估表明,我们的实现在多核CPU上实现了更好的可扩展性。我们还讨论了我们的方法在其他基于RNN的模型实现中的潜力。
Scaling up Artificial Intelligence (AI) algorithms for massive datasets to improve their performance is becoming crucial. In Machine Translation (MT), one of most important research fields of AI, models based on Recurrent Neural Net- works (RNN) show state-of-the-art performance in recent years, and many researchers keep working on improving RNN-based models to achieve better accuracy in translation tasks. Most implementations of Neural Machine Translation (NMT) models employ a padding strategy when processing a mini-batch to make all sentences in a mini-batch have the same length. This enables an efficient utilization of caches and GPU/SIMD parallelism but leads to a waste of computation time. In this paper, we implement and parallelize batch learning for a Sequence-to- Sequence (Seq2Seq) model, which is the most basic model of NMT, without using a padding strategy. More specifically, our approach forms vectors which represent the input words as well as the neural network's states at different time steps into matrices when it processes one sentence, and as a result, the approach makes a better use of cache and optimizes the process that adjusts weights and biases during the back-propagation phase. Our experimental evaluation shows that our implementation achieves better scalability on multi-core CPUs. We also discuss our approach's potential to be used in other implementations of RNN-based models.