B-scaling: A novel nonparametric data fusion method

B-scaling: A novel nonparametric data fusion method
复制标题

DOI:
10.1214/21-aoas1537
复制
发表时间:
2021-09
期刊:
The Annals of Applied Statistics
影响因子:
--
通讯作者:
Yiwen Liu;Xiaoxiao Sun;Wenxuan Zhong;Bing Li
Yiwen Liu;Xiaoxiao Sun;Wenxuan Zhong;Bing Li
中科院分区:
其他
文献类型:
--
作者:
Yiwen Liu;Xiaoxiao Sun;Wenxuan Zhong;Bing Li

文献摘要

相似文献

对于相同的科学问题,通常会有不同的技术或实验来测量相同的数值。在历史上,已经开发了各种方法来独立地利用每种类型的数据中的信息。然而,缺乏能够在统一框架下有效集成多源数据的统计数据融合方法。本文提出了一种新的数据融合方法,称为B-Scaling,用于集成多源数据。考虑来自不同来源但通过一些线性或非线性方法测量相同潜在变量的$K$测量。我们试图找到一个潜在变量的表示,称为B-均值,它捕捉了$K$测量值中包含的共同信息,同时考虑了它们与潜在变量之间的非线性映射。我们还建立了B-均值的渐近性质,并应用所提出的方法来整合多个组蛋白修饰和DNA甲基化水平来表征表观基因组图谱。数值和实证研究都表明,B-Scaling是一种强大的数据融合方法,具有广泛的应用前景。
Very often for the same scientific question, there may exist different techniques or experiments that measure the same numerical quantity. Historically, various methods have been developed to exploit the information within each type of data independently. However, statistical data fusion methods that could effectively integrate multi-source data under a unified framework are lacking. In this paper, we propose a novel data fusion method, called B-scaling, for integrating multi-source data. Consider $K$ measurements that are generated from different sources but measure the same latent variable through some linear or nonlinear ways. We seek to find a representation of the latent variable, named B-mean, which captures the common information contained in the $K$ measurements while takes into account the nonlinear mappings between them and the latent variable. We also establish the asymptotic property of the B-mean and apply the proposed method to integrate multiple histone modifications and DNA methylation levels for characterizing epigenomic landscape. Both numerical and empirical studies show that B-scaling is a powerful data fusion method with broad applications.