Contrastive Entity Linkage: Mining Variational Attributes from Large Catalogs for Entity Linkage

Contrastive Entity Linkage: Mining Variational Attributes from Large Catalogs for Entity Linkage
复制标题

DOI:
10.24432/c5w88r
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Varun R. Embar;Bunyamin Sisman;Hao Wei;Xin Dong;C. Faloutsos;L. Getoor
Varun R. Embar;Bunyamin Sisman;Hao Wei;Xin Dong;C. Faloutsos;L. Getoor
中科院分区:
其他
文献类型:
--
作者:
Varun R. Embar;Bunyamin Sisman;Hao Wei;Xin Dong;C. Faloutsos;L. Getoor

文献摘要

相似文献

几乎相同但不同的实体(称为实体变体)的存在使得数据集成任务具有挑战性。例如,在杂货产品领域,变体对于品牌、制造商和产品线等属性具有相同的值,但在其他属性(称为变体属性)方面有所不同,例如包装尺寸和颜色。识别数据源之间的差异本身就是一项重要任务,对于识别重复项至关重要。然而,这项任务具有挑战性,因为变分属性通常作为非结构化文本的一部分出现并且与域相关。在这项工作中,我们提出了我们的方法,对比实体链接,来识别相同的实体对和彼此不同的实体对。我们提出了一种新颖的无监督方法 VarSpot 来挖掘非结构化文本中存在的与域相关的变分属性。所提出的方法解释了实体之间的相似性和差异,并且可以轻松扩展到包含数百万个实体的大型源。我们通过对三个不同领域进行实验评估来展示我们方法的通用性。我们的方法在识别重复项时的 F1 分数显着优于最先进的基于学习和基于规则的实体链接系统,高达 4%,在识别实体变化时高达 41%。
Presence of near identical, but distinct, entities called entity variations makes the task of data integration challenging. For example, in the domain of grocery products, variations share the same value for attributes such as brand, manufacturer and product line, but differ in other attributes, called variational attributes , such as package size and color. Identifying variations across data sources is an important task in itself and is crucial for identifying duplicates. However, this task is challenging as the variational attributes are often present as a part of unstructured text and are domain dependent. In this work, we propose our approach, Contrastive entity linkage , to identify both entity pairs that are the same and pairs that are variations of each other. We propose a novel unsupervised approach, VarSpot , to mine domain-dependent variational attributes present in unstructured text. The proposed approach reasons about both similarities and differences between entities and can easily scale to large sources containing millions of entities. We show the generality of our approach by performing experimental evaluation on three different domains. Our approach significantly outperforms state-of-the-art learning-based and rule-based entity linkage systems by up to 4% F1 score when identifying duplicates, and up to 41% when identifying entity variations.