Predicting the Compositionality of Nominal Compounds: Giving Word Embeddings a Hard Time

Predicting the Compositionality of Nominal Compounds: Giving Word Embeddings a Hard Time
复制标题

预测名义复合词的组合性:给词嵌入带来困难

DOI:
10.18653/v1/p16-1187
复制
发表时间:
2016
期刊:
--
影响因子:
--
通讯作者:
Aline Villavicencio
Aline Villavicencio
中科院分区:
--
文献类型:
--
作者:
S. Cordeiro;Carlos Ramisch;M. Idiart;Aline Villavicencio

文献摘要

参考文献

被引文献

相似文献

分布式语义模型通常在包含单个词或完全组合短语的人工相似数据集上进行评估。我们提出了一项大规模的多语言评价,用于预测英语和法语4个数据集上名词性复合词的语义组合程度。我们总共构建了816个dsm,并使用基于word2vec、GloVe和ppmi的模型执行了2856次评估。除了dsm之外,我们还比较了不同参数的影响,如语料库预处理水平、上下文窗口大小和维度数。所获得的结果与人类的判断有很高的相关性,对于一些数据集来说,与最先进的水平相当或优于最先进的水平(Reddy数据集的Spearman ρ= 0.82)。
Distributional semantic models (DSMs) are often evaluated on artificial similarity datasets containing single words or fully compositional phrases. We present a large-scale multilingual evaluation of DSMs for predicting the degree of semantic compositionality of nominal compounds on 4 datasets for English and French. We build a total of 816 DSMs and perform 2,856 evaluations using word2vec, GloVe, and PPMI-based models. In addition to the DSMs, we compare the impact of different parameters, such as level of corpus preprocessing, context window size and number of dimensions. The results obtained have a high correlation with human judgments, being comparable to or outperforming the state of the art for some datasets (Spearman's ρ=.82 for the Reddy dataset).
Le Fort I截骨固定方法及术后骨片移位检查
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
藤尾正人;佐世暁;荻須宏太;土屋周平;酒井陽;日比英晴
通讯作者: 日比英晴