Meta-analysis of heterogeneous data: integrative sparse regression in high-dimensions

Meta-analysis of heterogeneous data: integrative sparse regression in high-dimensions
复制标题

DOI:
--
复制
发表时间:
2019-12
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Subha Maity;Yuekai Sun;M. Banerjee
Subha Maity;Yuekai Sun;M. Banerjee
中科院分区:
其他
文献类型:
--
作者:
Subha Maity;Yuekai Sun;M. Banerjee

文献摘要

相似文献

我们认为,在高维设置中的数据源是相似的,但不相同的荟萃分析的任务。为了在这种异构数据集上借用力量,我们引入了一个全局参数,强调在存在异质性的情况下的可解释性和统计效率。我们还提出了一个一次性的估计的全局参数,保持匿名的数据源和收敛速度,取决于组合数据集的大小。对于高维线性模型设置,我们证明了我们的识别限制在适应以前看到的数据分布以及预测新的/看不见的数据分布方面的优越性。最后,我们展示了我们的方法在涉及几种不同癌细胞系的大规模药物治疗数据集上的好处。
We consider the task of meta-analysis in high-dimensional settings in which the data sources are similar but non-identical. To borrow strength across such heterogeneous datasets, we introduce a global parameter that emphasizes interpretability and statistical efficiency in the presence of heterogeneity. We also propose a one-shot estimator of the global parameter that preserves the anonymity of the data sources and converges at a rate that depends on the size of the combined dataset. For high-dimensional linear model settings, we demonstrate the superiority of our identification restrictions in adapting to a previously seen data distribution as well as predicting for a new/unseen data distribution. Finally, we demonstrate the benefits of our approach on a large-scale drug treatment dataset involving several different cancer cell-lines.