Early Discovery of Emerging Entities in Microblogs

Early Discovery of Emerging Entities in Microblogs
复制标题

DOI:
10.24963/ijcai.2019/678
复制
发表时间:
2019-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Satoshi Akasaki;Naoki Yoshinaga;Masashi Toyoda
Satoshi Akasaki;Naoki Yoshinaga;Masashi Toyoda
中科院分区:
其他
文献类型:
--
作者:
Satoshi Akasaki;Naoki Yoshinaga;Masashi Toyoda

文献摘要

相似文献

对于社会趋势分析和市场研究等各种应用来说,每天都要及时了解新出现的实体是必不可少的。以前的研究试图检测未在特定知识库中注册为新兴实体的未见实体,从而发现非新兴实体,因为知识库中没有实体并不能保证它们的出现。因此,我们引入了一项新任务,即发现刚刚通过微博向公众介绍的真正新兴实体,并提出了一种基于时间敏感的远程监督的有效方法,该方法利用了新兴实体独特的早期背景。基于大规模Twitter档案的实验结果表明,该方法对发现的前500个新兴实体的准确率达到83.2%,优于基于突发检测的未见实体识别基线。除了显著的新兴实体外,我们的方法还可以发现大量的长尾和同形新兴实体。对相对召回率的评估表明,该方法检测到维基百科中80.4%的新注册实体;其中92.8%在维基百科注册之前就被发现,平均前置时间超过一年(578天)。
Keeping up to date on emerging entities that appear every day is indispensable for various applications, such as social-trend analysis and marketing research. Previous studies have attempted to detect unseen entities that are not registered in a particular knowledge base as emerging entities and consequently find non-emerging entities since the absence of entities in knowledge bases does not guarantee their emergence. We therefore introduce a novel task of discovering truly emerging entities when they have just been introduced to the public through microblogs and propose an effective method based on time-sensitive distant supervision, which exploits distinctive early-stage contexts of emerging entities. Experimental results with a large-scale Twitter archive show that the proposed method achieves 83.2% precision of the top 500 discovered emerging entities, which outperforms baselines based on unseen entity recognition with burst detection. Besides notable emerging entities, our method can discover massive long-tail and homographic emerging entities. An evaluation of relative recall shows that the method detects 80.4% emerging entities newly registered in Wikipedia; 92.8% of them are discovered earlier than their registration in Wikipedia, and the average lead-time is more than one year (578 days).