Heterogeneity-aware and communication-efficient distributed statistical inference

Heterogeneity-aware and communication-efficient distributed statistical inference
复制标题

DOI:
10.1093/biomet/asab007
复制
发表时间:
2022-02-01
期刊:
影响因子:
2.7
通讯作者:
Chen, Yong
Chen, Yong
中科院分区:
数学2区
文献类型:
--
作者:
Duan, Rui;Ning, Yang;Chen, Yong

文献摘要

被引文献

相似文献

在多中心研究中,个人级别的数据通常受到保护,不能跨站点共享。为了克服数据共享的障碍,许多分布式算法被开发出来,它们只需要共享聚集的信息。现有的分布式算法通常假设数据是跨站点均匀分布的。这一假设忽略了一个重要事实,即在不同地点收集的数据可能来自不同的亚群体和环境,这可能导致数据分布的异质性。忽视异质性可能会导致错误的统计推断。我们提出了分布式算法,通过允许站点特定的干扰参数来解释异构性分布。通过将一种新的密度比倾斜方法应用于有效的得分函数,所提出的方法将替代似然方法(;)扩展到异质环境。所提出的算法保持了与现有通信效率算法相同的通信代价。在两指标渐近设置下,我们建立了分布估计及其极限分布的非渐近风险界,它允许每个站点的样本量和站点的数目都是无穷大的。此外,我们还证明了当站点数目小于每个站点的样本量时,估计的渐近方差达到了Cramer-Rao下界。最后,通过仿真研究和实际数据应用,验证了所提方法的有效性和可行性。
In multicentre research, individual-level data are often protected against sharing across sites. To overcome the barrier of data sharing, many distributed algorithms, which only require sharing aggregated information, have been developed. The existing distributed algorithms usually assume the data are homogeneously distributed across sites. This assumption ignores the important fact that the data collected at different sites may come from various subpopulations and environments, which can lead to heterogeneity in the distribution of the data. Ignoring the heterogeneity may lead to erroneous statistical inference. We propose distributed algorithms which account for the heterogeneous distributions by allowing site-specific nuisance parameters. The proposed methods extend the surrogate likelihood approach (; ) to the heterogeneous setting by applying a novel density ratio tilting method to the efficient score function. The proposed algorithms maintain the same communication cost as existing communication-efficient algorithms. We establish a nonasymptotic risk bound for the proposed distributed estimator and its limiting distribution in the two-index asymptotic setting, which allows both sample size per site and the number of sites to go to infinity. In addition, we show that the asymptotic variance of the estimator attains the Cramer-Rao lower bound when the number of sites is smaller in rate than the sample size at each site. Finally, we use simulation studies and a real data application to demonstrate the validity and feasibility of the proposed methods.