Canonicalization of database records using adaptive similarity measures

Canonicalization of database records using adaptive similarity measures
复制标题

使用自适应相似性度量对数据库记录进行规范化

DOI:
--
复制
发表时间:
2007
期刊:
Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
A. McCallum
A. McCallum
中科院分区:
--
文献类型:
--
作者:
A. Culotta;Michael L. Wick;Robert J. Hall;Matthew Marzilli;A. McCallum

文献摘要

被引文献

相似文献

从许多异构源中自动挑选信息来构建数据库正变得越来越普遍。例如,可以通过自动从在线论文中提取标题、作者和会议信息来构建研究出版物数据库。整合来自多个来源的数据的一个常见困难是记录以各种方式被引用(例如缩写、别名和拼写错误)。因此,很难构造一个单一的、标准的表示来呈现给用户。我们将构造这种表示的任务称为规范化。尽管它很重要,但关于规范化的现有工作很少。
It is becoming increasingly common to construct databases from information automatically culled from many heterogeneous sources. For example, a research publication database can be constructed by automatically extracting titles, authors, and conference information from online papers. A common difficulty in consolidating data from multiple sources is that records are referenced in a variety of ways (e.g. abbreviations, aliases, and misspellings). Therefore, it can be difficult to construct a single, standard representation to present to the user. We refer to the task of constructing this representation as canonicalization. Despite its importance, there is little existing work on canonicalization. In this paper, we explore the use of edit distance measures to construct a canonical representation that is "central" in the sense that it is most similar to each of the disparate records. This approach reduces the impact of noisy records on the canonical representation. Furthermore, because the user may prefer different styles of canonicalization, we show how different edit distance costs can result in different forms of canonicalization. For example, reducing the cost of character deletions can result in representations that favor abbreviated forms over expanded forms (e.g. KDD versus Conference on Knowledge Discovery and Data Mining). We describe how to learn these costs from a small amount of manually annotated data using stochastic hill-climbing. Additionally, we investigate feature-based methods to learn ranking preferences over canonicalizations. These approaches can incorporate arbitrary textual evidence to select a canonical record. We evaluate our approach on a real-world publications database and show that our learning method results in a canonicalization solution that is robust to errors and easily customizable to user preferences.