High-Dimensional Bayesian Network Classification with Network Global-Local Shrinkage Priors

High-Dimensional Bayesian Network Classification with Network Global-Local Shrinkage Priors
复制标题

DOI:
10.1214/23-ba1378
复制
发表时间:
2020-09
期刊:
影响因子:
4.4
通讯作者:
Sharmistha Guha;Abel Rodríguez
Sharmistha Guha;Abel Rodríguez
中科院分区:
数学2区
文献类型:
--
作者:
Sharmistha Guha;Abel Rodríguez

文献摘要

相似文献

提出了一种新的带标签节点网络的贝叶斯分类框架。虽然关于网络数据的统计建模的文献通常涉及对单个网络的分析,但最近在包括脑成像研究在内的几个生物学应用中出现的复杂数据提出了为受试者设计网络分类器的需要。本文考虑脑连接组研究的一个应用,其中首要目标是根据受试者的大脑网络数据将受试者分为两组,并确定有影响的感兴趣区域(ROI)(称为节点)。现有的方法要么把所有的边权重看作一个长向量,要么用几个摘要度量来概括网络信息。这两种方法都忽略了完整的网络结构,可能会在小样本中导致不太理想的推断,并且不是被设计来识别重要的网络节点。我们提出了一种新的二元Logistic回归框架,该框架以网络作为预测器和二元响应,网络预测器系数使用一类新的全局-局部收缩先验来建模。该框架能够准确地检测网络中影响分类的节点和边缘。我们的框架是使用一种高效的马尔可夫链蒙特卡罗算法实现的。理论上,当网络边数的增长快于样本量时,我们证明了所提出的框架的渐近最优分类。该框架通过广泛的模拟研究和对大脑连接体数据的分析得到了实证验证。
This article proposes a novel Bayesian classification framework for networks with labeled nodes. While literature on statistical modeling of network data typically involves analysis of a single network, the recent emergence of complex data in several biological applications, including brain imaging studies, presents a need to devise a network classifier for subjects. This article considers an application from a brain connectome study, where the overarching goal is to classify subjects into two separate groups based on their brain network data, along with identifying influential regions of interest (ROIs) (referred to as nodes). Existing approaches either treat all edge weights as a long vector or summarize the network information with a few summary measures. Both these approaches ignore the full network structure, may lead to less desirable inference in small samples and are not designed to identify significant network nodes. We propose a novel binary logistic regression framework with the network as the predictor and a binary response, the network predictor coefficient being modeled using a novel class global-local shrinkage priors. The framework is able to accurately detect nodes and edges in the network influencing the classification. Our framework is implemented using an efficient Markov Chain Monte Carlo algorithm. Theoretically, we show asymptotically optimal classification for the proposed framework when the number of network edges grows faster than the sample size. The framework is empirically validated by extensive simulation studies and analysis of a brain connectome data.