Wikipedia2Vec: An Efficient Toolkit for Learning and Visualizing the Embeddings of Words and Entities from Wikipedia

Wikipedia2Vec: An Efficient Toolkit for Learning and Visualizing the Embeddings of Words and Entities from Wikipedia
复制标题

DOI:
10.18653/v1/2020.emnlp-demos.4
复制
发表时间:
2018-12
期刊:
--
影响因子:
--
通讯作者:
Ikuya Yamada;Akari Asai;Jin Sakuma;Hiroyuki Shindo;Hideaki Takeda;Yoshiyasu Takefuji;Yuji Matsumoto
Ikuya Yamada;Akari Asai;Jin Sakuma;Hiroyuki Shindo;Hideaki Takeda;Yoshiyasu Takefuji;Yuji Matsumoto
中科院分区:
其他
文献类型:
--
作者:
Ikuya Yamada;Akari Asai;Jin Sakuma;Hiroyuki Shindo;Hideaki Takeda;Yoshiyasu Takefuji;Yuji Matsumoto

文献摘要

被引文献

相似文献

将实体嵌入到大型知识库中(例如,Wikipedia)对于解决涉及真实的世界知识的各种自然语言任务是非常有益的。在本文中,我们介绍了Wikipedia2Vec,一个基于Python的开源工具,用于学习维基百科中的单词和实体的嵌入。建议的工具,使用户能够有效地学习嵌入发出一个单一的命令与维基百科转储文件作为参数。我们还介绍了我们的工具,允许用户可视化和探索学习嵌入基于Web的演示。在我们的实验中,我们的工具在KORE实体相关性数据集上取得了最先进的结果,并在各种标准基准数据集上取得了有竞争力的结果。此外,我们的工具已被用作最近各种研究的关键组成部分。我们在https://wikipedia2vec.github.io/上公布了12种语言的源代码、演示和预训练嵌入。
The embeddings of entities in a large knowledge base (e.g., Wikipedia) are highly beneficial for solving various natural language tasks that involve real world knowledge. In this paper, we present Wikipedia2Vec, a Python-based open-source tool for learning the embeddings of words and entities from Wikipedia. The proposed tool enables users to learn the embeddings efficiently by issuing a single command with a Wikipedia dump file as an argument. We also introduce a web-based demonstration of our tool that allows users to visualize and explore the learned embeddings. In our experiments, our tool achieved a state-of-the-art result on the KORE entity relatedness dataset, and competitive results on various standard benchmark datasets. Furthermore, our tool has been used as a key component in various recent studies. We publicize the source code, demonstration, and the pretrained embeddings for 12 languages at https://wikipedia2vec.github.io/.