Annotating and Searching Web Tables Using Entities, Types and Relationships

Annotating and Searching Web Tables Using Entities, Types and Relationships
复制标题

DOI:
10.14778/1920841.1921005
复制
发表时间:
2010-09-01
影响因子:
2.5
通讯作者:
Chakrabarti, Soumen
Chakrabarti, Soumen
中科院分区:
计算机科学2区
文献类型:
--
作者:
Limaye, Girija;Sarawagi, Sunita;Chakrabarti, Soumen

文献摘要

被引文献

相似文献

表是表示关系数据的通用方式。Web页面上有数十亿个表表示实体引用、属性和关系。关系世界知识的这种表示通常比完全非结构化、自由格式的文本要好得多。同时,与手工创建的知识库不同,从“有机”Web表中挖掘的关系信息不需要受到宝贵编辑时间的限制。不幸的是,在没有任何正式的、统一的模式强加于Web表的情况下,Web搜索无法利用这些高质量的关系信息源。在本文中,我们提出了新的机器学习技术,用它们可能提到的实体来注释表单元格,用为列中的单元格绘制实体的类型来注释表列,以及对表列寻求表达的关系。我们提出了一种新的图形模型,可以同时为每个表做出所有这些标记决策,而不是为实体、类型和关系做出单独的局部决策。使用YAGO目录、DBPedia、维基百科的表以及来自5亿个网页抓取的超过2500万个HTML表的实验一致显示了我们方法的优越性。我们还评估了更好的注释对原型关系Web搜索工具的影响。除了以纯文本的方式索引表之外,我们还展示了注释的明显好处。
Tables are a universal idiom to present relational data. Billions of tables on Web pages express entity references, attributes and relationships. This representation of relational world knowledge is usually considerably better than completely unstructured, free-format text. At the same time, unlike manually-created knowledge bases, relational information mined from "organic" Web tables need not be constrained by availability of precious editorial time. Unfortunately, in the absence of any formal, uniform schema imposed on Web tables, Web search cannot take advantage of these high-quality sources of relational information. In this paper we propose new machine learning techniques to annotate table cells with entities that they likely mention, table columns with types from which entities are drawn for cells in the column, and relations that pairs of table columns seek to express. We propose a new graphical model for making all these labeling decisions for each table simultaneously, rather than make separate local decisions for entities, types and relations. Experiments using the YAGO catalog, DBPedia, tables from Wikipedia, and over 25 million HTML tables from a 500 million page Web crawl uniformly show the superiority of our approach. We also evaluate the impact of better annotations on a prototype relational Web search tool. We demonstrate clear benefits of our annotations beyond indexing tables in a purely textual manner.