Tuning-Free Heterogeneous Inference in Massive Networks

Tuning-Free Heterogeneous Inference in Massive Networks
复制标题

DOI:
10.1080/01621459.2018.1537920
复制
发表时间:
2019-04
影响因子:
3.7
通讯作者:
Zhao Ren;Yongjian Kang;Yingying Fan;Jinchi Lv
Zhao Ren;Yongjian Kang;Yingying Fan;Jinchi Lv
中科院分区:
数学1区
文献类型:
--
作者:
Zhao Ren;Yongjian Kang;Yingying Fan;Jinchi Lv

文献摘要

相似文献

在许多涉及海量数据的当代应用中,异构性通常是很自然的。在对有效学习提出新挑战的同时,它可以通过整合感兴趣的亚群体之间的信息,在推动有意义的科学发现方面发挥关键作用。在本文中,我们利用高斯图的多个网络对子种群上大量特征的连通性模式进行编码。为了揭示跨亚种群的潜在稀疏性结构,我们提出了一个大规模无调优异构推断框架,其中网络数量允许发散。特别地,引入了两种新的检验,即基于chi的检验和基于线性函数的检验,并建立了它们的渐近零分布。在温和的正则性条件下,我们建立了两种测试在达到可测试区域边界时都是最优的,并且后一种测试的样本量要求最小。本文新提出的异质群平方根Lasso对高维异质噪声多响应回归的有效多网络估计提供了理论保证和无调谐特性。为了求解这个凸规划,我们进一步引入了一种可扩展的算法,该算法具有可证明的收敛性。通过仿真和实际数据实例说明了该方法在计算和理论上的优点。本文的补充材料可在网上获得。
Abstract Heterogeneity is often natural in many contemporary applications involving massive data. While posing new challenges to effective learning, it can play a crucial role in powering meaningful scientific discoveries through the integration of information among subpopulations of interest. In this article, we exploit multiple networks with Gaussian graphs to encode the connectivity patterns of a large number of features on the subpopulations. To uncover the underlying sparsity structures across subpopulations, we suggest a framework of large-scale tuning-free heterogeneous inference, where the number of networks is allowed to diverge. In particular, two new tests, the chi-based and the linear functional-based tests, are introduced and their asymptotic null distributions are established. Under mild regularity conditions, we establish that both tests are optimal in achieving the testable region boundary and the sample size requirement for the latter test is minimal. Both theoretical guarantees and the tuning-free property stem from efficient multiple-network estimation by our newly suggested heterogeneous group square-root Lasso for high-dimensional multi-response regression with heterogeneous noises. To solve this convex program, we further introduce a scalable algorithm that enjoys provable convergence to the global optimum. Both computational and theoretical advantages are elucidated through simulation and real data examples. Supplementary materials for this article are available online.