Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling Mechanisms

Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling Mechanisms
复制标题

DOI:
10.18653/v1/p18-1041
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Dinghan Shen;Guoyin Wang;Wenlin Wang;Martin Renqiang Min;Qinliang Su;Yizhe Zhang;Chunyuan Li;Ricardo H
Dinghan Shen;Guoyin Wang;Wenlin Wang;Martin Renqiang Min;Qinliang Su;Yizhe Zhang;Chunyuan Li;Ricardo H
中科院分区:
其他
文献类型:
--
作者:
Dinghan Shen;Guoyin Wang;Wenlin Wang;Martin Renqiang Min;Qinliang Su;Yizhe Zhang;Chunyuan Li;Ricardo H

文献摘要

被引文献

相似文献

已经提出了许多深度学习体系来对文本序列中的构成性进行建模,这些模型需要大量的参数和昂贵的计算。然而,对于复杂的组成功能的附加值还没有一个严格的评估。在本文中,我们对基于简单词嵌入的模型(SWEM)和基于词嵌入的RNN/CNN模型进行了逐点的比较研究。令人惊讶的是,在所考虑的大多数情况下,SWEM表现出相当甚至更好的性能。基于这一理解,我们提出了两种针对学习单词嵌入的额外池化策略:(I)用于提高可解释性的最大池化操作;(Ii)在文本序列中保留空间(n元语法)信息的层次化池化操作。我们在17个数据集上进行了实验,包括三个任务:(I)(长)文档分类;(Ii)文本序列匹配;(Iii)短文本任务,包括分类和标注。
Many deep learning architectures have been proposed to model the compositionality in text sequences, requiring substantial number of parameters and expensive computations. However, there has not been a rigorous evaluation regarding the added value of sophisticated compositional functions. In this paper, we conduct a point-by-point comparative study between Simple Word-Embedding-based Models (SWEMs), consisting of parameter-free pooling operations, relative to word-embedding-based RNN/CNN models. Surprisingly, SWEMs exhibit comparable or even superior performance in the majority of cases considered. Based upon this understanding, we propose two additional pooling strategies over learned word embeddings: (i) a max-pooling operation for improved interpretability; and (ii) a hierarchical pooling operation, which preserves spatial (n-gram) information within text sequences. We present experiments on 17 datasets encompassing three tasks: (i) (long) document classification; (ii) text sequence matching; and (iii) short text tasks, including classification and tagging.