A Practical Approach to Proper Inference with Linked Data

A Practical Approach to Proper Inference with Linked Data
复制标题

使用关联数据进行正确推理的实用方法

DOI:
10.1080/00031305.2022.2041482
复制
发表时间:
2022
期刊:
The American Statistician
影响因子:
--
通讯作者:
Steorts, Rebecca C.
Steorts, Rebecca C.
中科院分区:
--
文献类型:
--
作者:
Kaplan, Andee;Betancourt, Brenda;Steorts, Rebecca C.

文献摘要

参考文献

被引文献

相似文献

实体解析(ER)包括记录链接和重复数据删除,是在没有唯一标识符以删除重复实体的情况下合并噪声数据库的过程。关联数据分析的一个主要挑战是在确定的匹配中识别具有代表性的记录,以传递给推理或预测任务,称为下游任务。此外,在下游任务中纳入ER的不确定性对于确保正确的推理至关重要。为了弥合ER和分析管道中下游任务之间的差距,我们提出了五种从关联数据中选择具有代表性的(或规范的)记录的方法,称为去规范化。我们的方法在记录数量上是可伸缩的,适合于一般数据场景,并通过贝叶斯规范化阶段提供自然的错误传播。建议的方法在三个模拟数据集和一个应用程序上进行了评估--确定北卡罗来纳州选举委员会选民登记数据中人口统计信息和政党归属之间的关系。我们首先执行贝叶斯ER,并评估我们提出的规范化方法,然后再考虑线性回归和Logistic回归的下游任务。经验表明,贝叶斯正则化方法通过预测和覆盖在两个环境中都可以改善下游推理。
Entity resolution (ER), comprising record linkage and deduplication, is the process of merging noisy databases in the absence of unique identifiers to remove duplicate entities. One major challenge of analysis with linked data is identifying a representative record among determined matches to pass to an inferential or predictive task, referred to as thedownstream task. Additionally, incorporating uncertainty from ER in the downstream task is critical to ensure proper inference. To bridge the gap between ER and the downstream task in an analysis pipeline, we propose five methods to choose a representative (orcanonical) record from linked data, referred to ascanonicalization. Our methods are scalable in the number of records, appropriate in general data scenarios, and provide natural error propagation via a Bayesian canonicalization stage. The proposed methodology is evaluated on three simulated datasets and one application – determining the relationship between demographic information and party affiliation in voter registration data from the North Carolina State Board of Elections. We first perform Bayesian ER and evaluate our proposed methods for canonicalization before considering the downstream tasks of linear and logistic regression. Bayesian canonicalization methods are empirically shown to improve downstream inference in both settings through prediction and coverage.
使用自适应相似性度量对数据库记录进行规范化
DOI: --
发表时间: 2007
期刊: Knowledge Discovery and Data Mining
影响因子: --
作者:
A. Culotta;Michael L. Wick;Robert J. Hall;Matthew Marzilli;A. McCallum
通讯作者: A. McCallum
不完全联动下的回归分析
DOI: --
发表时间: 2012
影响因子: 1.8
作者:
Gunky Kim;R. Chambers
通讯作者: R. Chambers
DOI: 10.1214/10-aoas447
发表时间: 2011-06-01
影响因子: 1.8
作者:
Tancredi, Andrea;Liseo, Brunero
通讯作者: Liseo, Brunero
DOI: 10.1198/106186007x238855
发表时间: 2007-09-01
影响因子: 2.4
作者:
Lau, John W.;Green, Peter J.
通讯作者: Green, Peter J.
DOI: --
发表时间: 2016
期刊: Neural Information Processing Systems
影响因子: --
作者:
Brenda Betancourt;Giacomo Zanella;Jeffrey W. Miller;Hanna M. Wallach;Abbas Zaidi;Beka Steorts
通讯作者: Beka Steorts