A Pseudo-Likelihood Approach to Linear Regression With Partially Shuffled Data

A Pseudo-Likelihood Approach to Linear Regression With Partially Shuffled Data
复制标题

部分混洗数据线性回归的伪似然方法

DOI:
10.1080/10618600.2020.1870482
复制
发表时间:
2021
影响因子:
2.4
通讯作者:
Ben-David, Emanuel
Ben-David, Emanuel
中科院分区:
数学2区
文献类型:
--
作者:
Slawski, Martin;Diao, Guoqing;Ben-David, Emanuel

文献摘要

参考文献

被引文献

相似文献

最近,由于单独的数据收集和数据集成中的不确定性,在与同一统计单位对应的匹配对中没有观察到预测者和响应的情况下,线性回归引起了极大的兴趣。不匹配对会严重影响模型拟合,并扰乱回归参数的估计。在这篇文章中,我们提出了一种在“部分洗牌”下调整这种失配的方法,在“部分洗牌”中,在它们的正确对应中观察到足够大的部分(预测因子,响应)对。所提出的方法是基于伪似然的,其中每一项都采用双组分混合密度的形式。期望最大化方案被提出用于优化,它(I)在样本数量上具有良好的伸缩性,并且(Ii)相对于通过模拟和案例研究证明能够访问正确配对的先知而言,获得了优异的统计性能。特别是,与现有方法相比,所提出的方法可以容忍相当大的失配比例,并且能够估计噪声水平以及失配比例。对所得到的估计量(标准误差、可信区间)的推断可以基于用于复合似然估计的既定理论。在此过程中,我们还提出了一种对错配存在的统计测试,并在适当的条件下建立了它的一致性。本文的补充文件可以在网上找到。
Recently, there has been significant interest in linear regression in the situation where predictors and responses are not observed in matching pairs corresponding to the same statistical unit as a consequence of separate data collection and uncertainty in data integration. Mismatched pairs can considerably impact the model fit and disrupt the estimation of regression parameters. In this article, we present a method to adjust for such mismatches under “partial shuffling” in which a sufficiently large fraction of (predictors, response)-pairs are observed in their correct correspondence. The proposed approach is based on a pseudo-likelihood in which each term takes the form of a two-component mixture density. expectation-maximization schemes are proposed for optimization, which (i) scale favorably in the number of samples, and (ii) achieve excellent statistical performance relative to an oracle that has access to the correct pairings as certified by simulations and case studies. In particular, the proposed approach can tolerate considerably larger fraction of mismatches than existing approaches, and enables estimation of the noise level as well as the fraction of mismatches. Inference for the resulting estimator (standard errors, confidence intervals) can be based on established theory for composite likelihood estimation. Along the way, we also propose a statistical test for the presence of mismatches and establish its consistency under suitable conditions. Supplemental files for this article are available online.
DOI: --
发表时间: 2017-05
期刊: ArXiv
影响因子: --
作者:
Daniel J. Hsu;K. Shi;Xiaorui Sun
通讯作者: Daniel J. Hsu;K. Shi;Xiaorui Sun
重新配对损坏的随机样本的最大似然策略的一些性质☆
DOI: --
发表时间: 1987
期刊:
影响因子: --
作者:
P. Goel;T. Ramalingam
通讯作者: T. Ramalingam
DOI: 10.1214/aos/1176344952
发表时间: 1980-01-01
影响因子: 4.5
作者:
DEGROOT, MH;GOEL, PK
通讯作者: GOEL, PK
计算机匹配的数据文件的回归分析
DOI: --
发表时间: 1993
期刊:
影响因子: --
作者:
F. Scheuren
通讯作者: F. Scheuren
DOI: 10.1080/01621459.1965.10480846
发表时间: 1965
影响因子: 3.7
作者:
J. Neter;E. Maynes;R. Ramanathan
通讯作者: R. Ramanathan