A Framework for the Construction of Monolingual and Cross-lingual Word Similarity Datasets

A Framework for the Construction of Monolingual and Cross-lingual Word Similarity Datasets
复制标题

DOI:
10.3115/v1/p15-2001
复制
发表时间:
2015-07
影响因子:
3.3
通讯作者:
José Camacho-Collados;Mohammad Taher Pilehvar;Roberto Navigli
José Camacho-Collados;Mohammad Taher Pilehvar;Roberto Navigli
中科院分区:
化学3区
文献类型:
--
作者:
José Camacho-Collados;Mohammad Taher Pilehvar;Roberto Navigli

文献摘要

被引文献

相似文献

尽管是词汇语义学中最受欢迎的任务之一,但单词相似性通常仅限于英语语言。其他语言,即使是那些广泛使用的语言,如西班牙语,也没有可靠的单词相似性评估框架。我们提出了强大的方法来扩展现有的英语数据集到其他语言,无论是在单语和跨语言的水平。我们提出了一个自动标准化的跨语言相似性数据集的建设,并提供了一个评估,证明其可靠性和鲁棒性。基于我们的过程,并以RG-65单词相似度数据集为参考,我们发布了两个高质量的西班牙语和波斯语(波斯语)单语数据集,以及六种语言的十五个跨语言数据集:英语,西班牙语,法语,德语,葡萄牙语和波斯语。
Despite being one of the most popular tasks in lexical semantics, word similarity has often been limited to the English language. Other languages, even those that are widely spoken such as Spanish, do not have a reliable word similarity evaluation framework. We put forward robust methodologies for the extension of existing English datasets to other languages, both at monolingual and cross-lingual levels. We propose an automatic standardization for the construction of cross-lingual similarity datasets, and provide an evaluation, demonstrating its reliability and robustness. Based on our procedure and taking the RG-65 word similarity dataset as a reference, we release two high-quality Spanish and Farsi (Persian) monolingual datasets, and fifteen cross-lingual datasets for six languages: English, Spanish, French, German, Portuguese, and Farsi.