MOLEMAN: Mention-Only Linking of Entities with a Mention Annotation Network

MOLEMAN: Mention-Only Linking of Entities with a Mention Annotation Network
复制标题

DOI:
10.18653/v1/2021.acl-short.37
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Nicholas FitzGerald;Jan A. Botha;D. Gillick;D. Bikel;T. Kwiatkowski;A. McCallum
Nicholas FitzGerald;Jan A. Botha;D. Gillick;D. Bikel;T. Kwiatkowski;A. McCallum
中科院分区:
其他
文献类型:
--
作者:
Nicholas FitzGerald;Jan A. Botha;D. Gillick;D. Bikel;T. Kwiatkowski;A. McCallum

文献摘要

被引文献

相似文献

提出了一种基于实例的最近邻实体链接方法。与大多数现有的实体检索系统用单个向量表示每个实体不同,我们构建了一个上下文提及编码器,它学习在向量空间中放置对同一实体的类似提及比对不同实体的提及更接近。该方法允许对实体的所有提及都用作“类原型”,因为推理涉及从训练集中被标记的实体提及的整个集合中检索并且应用最近提及的邻居的实体标签。我们的模型在来自维基百科超链接的大型多语言提及对语料库上进行训练,并在7亿条提及的索引上执行最近邻推理。它训练起来更简单,提供了更多可解释的预测,并且在两个多语言实体链接基准上的表现优于所有其他系统。
We present an instance-based nearest neighbor approach to entity linking. In contrast to most prior entity retrieval systems which represent each entity with a single vector, we build a contextualized mention-encoder that learns to place similar mentions of the same entity closer in vector space than mentions of different entities. This approach allows all mentions of an entity to serve as “class prototypes” as inference involves retrieving from the full set of labeled entity mentions in the training set and applying the nearest mention neighbor’s entity label. Our model is trained on a large multilingual corpus of mention pairs derived from Wikipedia hyperlinks, and performs nearest neighbor inference on an index of 700 million mentions. It is simpler to train, gives more interpretable predictions, and outperforms all other systems on two multilingual entity linking benchmarks.