Tweeki: Linking Named Entities on Twitter to a Knowledge Graph

Tweeki: Linking Named Entities on Twitter to a Knowledge Graph
复制标题

DOI:
10.18653/v1/2020.wnut-1.29
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Bahareh Harandizadeh;Sameer Singh
Bahareh Harandizadeh;Sameer Singh
中科院分区:
其他
文献类型:
--
作者:
Bahareh Harandizadeh;Sameer Singh

文献摘要

被引文献

相似文献

为了识别推文中谈论的实体,我们需要自动将推文中出现的命名实体链接到WikiData等结构化知识库。现有的方法通常难以处理这些简短、嘈杂的文本,或者它们复杂的设计和对监督的依赖使它们变得脆弱,难以使用和维护,并且随着时间的推移而失去意义。此外,还缺乏一个大型的、链接的推文语料库来帮助研究人员,沿着缺乏黄金数据集来评估实体链接的准确性。在本文中,我们介绍(1)Tweeki,一个无监督的,模块化的实体链接系统的Twitter,(2)TweekiData,一个大型的,自动注释的语料库的推文链接到维基数据中的实体,和(3)TweekiGold,黄金数据集的实体链接评估。通过全面的分析,我们表明Tweeki与最近最先进的实体链接器模型的性能相当,数据集具有高质量,并且是如何使用数据集来改进社交媒体分析中的下游任务(地理位置预测)的用例。
To identify what entities are being talked about in tweets, we need to automatically link named entities that appear in tweets to structured KBs like WikiData. Existing approaches often struggle with such short, noisy texts, or their complex design and reliance on supervision make them brittle, difficult to use and maintain, and lose significance over time. Further, there is a lack of a large, linked corpus of tweets to aid researchers, along with lack of gold dataset to evaluate the accuracy of entity linking. In this paper, we introduce (1) Tweeki, an unsupervised, modular entity linking system for Twitter, (2) TweekiData, a large, automatically-annotated corpus of Tweets linked to entities in WikiData, and (3) TweekiGold, a gold dataset for entity linking evaluation. Through comprehensive analysis, we show that Tweeki is comparable to the performance of recent state-of-the-art entity linkers models, the dataset is of high quality, and a use case of how the dataset can be used to improve downstream tasks in social media analysis (geolocation prediction).