Distributed Quasi-Poisson regression algorithm for modeling multi-site count outcomes in distributed data networks.

Distributed Quasi-Poisson regression algorithm for modeling multi-site count outcomes in distributed data networks.
复制标题

分布式拟泊松回归算法,用于对分布式数据网络中的多站点计数结果进行建模。

DOI:
10.1016/j.jbi.2022.104097
复制
发表时间:
2022
影响因子:
4.5
通讯作者:
Chen,Yong
Chen,Yong
中科院分区:
医学3区
文献类型:
--
作者:
Edmondson,MackenzieJ;Luo,Chongliang;NazmulIslam,Md;Sheils,NatalieE;Buresh,John;Chen,Zhaoyi;Bian,Jiang;Chen,Yong

文献摘要

相似文献

背景观察性研究结合了来自多个机构的真实世界数据,有助于研究罕见的结果或暴露,并提高结果的普遍性。由于围绕跨机构患者级数据共享的隐私问题,需要分布式执行回归分析的方法。通常使用对特定机构估计值的荟萃分析,但已证明在某些情况下会产生有偏差的估计值。虽然分布式回归方法越来越可用,但分析计数结果的方法目前有限。实践中的计数数据通常会出现过度分散,在给定的统计模型下表现出比预期更大的变异性。目的我们提出了一种新颖的计算方法,即一种用于准泊松回归(ODAP)的一次性分布式算法,在考虑过度分散的同时对计数结果进行分布式建模。方法ODAP采用替代似然方法来执行分布式准泊松回归,而无需 需要患者层面的数据共享,只需要共享各个参与机构的汇总数据。 ODAP 最多需要机构之间进行三轮非迭代沟通,以生成系数估计值和相应的标准误差。在模拟中,我们在多站点分析中可能的几种数据场景下评估 ODAP,比较 ODAP 和荟萃分析估计相对于汇总回归估计的误差(被认为是黄金标准)。在概念验证现实世界数据分析中,我们使用来自 OneFlorida 临床研究联盟的数据,将 ODAP 和荟萃分析的相对误差与汇总估计进行类似比较,将 COVID-19 患者的住院时间建模为各种患者特征的函数。在第二个概念验证分析中,使用相同的结果和协变量,我们将 UnitedHealth Group 临床发现数据库的数据与 OneFlorida 数据合并到分布式分析中,以比较 ODAP 和荟萃分析产生的估计值。 结果在模拟中,ODAP 相对于所有探索的设置中的汇总回归估计值显示出可忽略不计的误差。荟萃分析估计虽然基本上没有偏见,但随着机构间结果异质性的增加,其变量也越来越大。当基线预期计数为 0.2 时,荟萃分析的相对误差在 25% 的迭代 (250/1000) 中高于 5%,而 ODAP 在任何迭代中的最大相对误差为 3.59%。在我们仅使用 OneFlorida 数据的概念验证分析中,ODAP 估计值比所有 15 个协变量的荟萃分析生成的估计值更接近汇总回归估计值。在我们整合来自 OneFlorida 和 UnitedHealth Group 临床发现数据库的数据的分布式分析中,ODAP 和荟萃分析估计值在很大程度上相似,而估计值中的一些差异(高达 13.8%)可能表明荟萃分析估计值中存在偏差。结论 ODAP 执行隐私保护、通信高效的分布式准泊松回归,以使用数据分析计数结果 存储在多个机构内。我们的方法产生的估计值几乎与汇总回归估计值相匹配,有时比荟萃分析估计值更准确,尤其是在各机构计数相对较低且结果异质性较高的环境中。
BackgroundObservational studies incorporating real-world data from multiple institutions facilitate study of rare outcomes or exposures and improve generalizability of results. Due to privacy concerns surrounding patient-level data sharing across institutions, methods for performing regression analyses distributively are desirable. Meta-analysis of institution-specific estimates is commonly used, but has been shown to produce biased estimates in certain settings. While distributed regression methods are increasingly available, methods for analyzing count outcomes are currently limited. Count data in practice are commonly subject to overdispersion, exhibiting greater variability than expected under a given statistical model.ObjectiveWe propose a novel computational method, a one-shot distributed algorithm for quasi-Poisson regression (ODAP), to distributively model count outcomes while accounting for overdispersion.MethodsODAP incorporates a surrogate likelihood approach to perform distributed quasi-Poisson regression without requiring patient-level data sharing, only requiring sharing of aggregate data from each participating institution. ODAP requires at most three rounds of non-iterative communication among institutions to generate coefficient estimates and corresponding standard errors. In simulations, we evaluate ODAP under several data scenarios possible in multi-site analyses, comparing ODAP and meta-analysis estimates in terms of error relative to pooled regression estimates, considered the gold standard. In a proof-of-concept real-world data analysis, we similarly compare ODAP and meta-analysis in terms of relative error to pooled estimatation using data from the OneFlorida Clinical Research Consortium, modeling length of stay in COVID-19 patients as a function of various patient characteristics. In a second proof-of-concept analysis, using the same outcome and covariates, we incorporate data from the UnitedHealth Group Clinical Discovery Database together with the OneFlorida data in a distributed analysis to compare estimates produced by ODAP and meta-analysis.ResultsIn simulations, ODAP exhibited negligible error relative to pooled regression estimates across all settings explored. Meta-analysis estimates, while largely unbiased, were increasingly variable as heterogeneity in the outcome increased across institutions. When baseline expected count was 0.2, relative error for meta-analysis was above 5% in 25% of iterations (250/1000), while the largest relative error for ODAP in any iteration was 3.59%. In our proof-of-concept analysis using only OneFlorida data, ODAP estimates were closer to pooled regression estimates than those produced by meta-analysis for all 15 covariates. In our distributed analysis incorporating data from both OneFlorida and the UnitedHealth Group Clinical Discovery Database, ODAP and meta-analysis estimates were largely similar, while some differences in estimates (as large as 13.8%) could be indicative of bias in meta-analytic estimates.ConclusionsODAP performs privacy-preserving, communication-efficient distributed quasi-Poisson regression to analyze count outcomes using data stored within multiple institutions. Our method produces estimates nearly matching pooled regression estimates and sometimes more accurate than meta-analysis estimates, most notably in settings with relatively low counts and high outcome heterogeneity across institutions.