Entity Disambiguation with Freebase

Entity Disambiguation with Freebase
复制标题

DOI:
10.1109/wi-iat.2012.26
复制
发表时间:
2012-12
期刊:
2012 IEEE/WIC/ACM International Conferences on Web Intelligence and Intelligent Agent Technology
影响因子:
--
通讯作者:
Zhicheng Zheng;Xiance Si;Fangtao Li;E. Chang;Xiaoyan Zhu
Zhicheng Zheng;Xiance Si;Fangtao Li;E. Chang;Xiaoyan Zhu
中科院分区:
其他
文献类型:
--
作者:
Zhicheng Zheng;Xiance Si;Fangtao Li;E. Chang;Xiaoyan Zhu

文献摘要

被引文献

相似文献

基于知识库的实体消歧在NLP社区中越来越流行。在本文中,我们使用Freebase作为知识库,它比维基百科和其他网站包含更多的实体。Freebase虽然庞大,但大多数实体缺乏上下文,比如维基百科中的描述性文本和超链接,这对消除歧义很有用。相反,我们利用Freebase的两个特性,即自然消歧的提及短语(又名别名)和丰富的分类法,以迭代的方式执行消歧。具体来说,我们为每次迭代探索生成和判别模型。在2430707个英语句子和33743个Freebase实体上的实验显示了这两个特征的有效性,其中90%的准确率可以在没有任何标记数据的情况下达到。我们还证明了采用分割训练策略的判别模型对过拟合问题具有鲁棒性,并且不断优于生成模型。
Entity disambiguation with a knowledge base becomes increasingly popular in the NLP community. In this paper, we employ Freebase as the knowledge base, which contains significantly more entities than Wikipedia and others. While huge in size, Freebase lacks context for most entities, such as the descriptive text and hyperlinks in Wikipedia, which are useful for disambiguation. Instead, we leverage two features of Freebase, namely the naturally disambiguated mention phrases (aka aliases) and the rich taxonomy, to perform disambiguation in an iterative manner. Specifically, we explore both generative and discriminative models for each iteration. Experiments on 2, 430, 707 English sentences and 33, 743 Freebase entities show the effectiveness of the two features, where 90% accuracy can be reached without any labeled data. We also show that discriminative models with proposed split training strategy is robust against over fitting problem, and constantly outperforms the generative ones.