STRUCTURED SUBCOMPOSITION SELECTION IN REGRESSION AND ITS APPLICATION TO MICROBIOME DATA ANALYSIS

STRUCTURED SUBCOMPOSITION SELECTION IN REGRESSION AND ITS APPLICATION TO MICROBIOME DATA ANALYSIS
复制标题

回归中的结构化子组合选择及其在微生物组数据分析中的应用

DOI:
10.1214/16-aoas1017
复制
发表时间:
2017-06-01
影响因子:
1.8
通讯作者:
Zhao, Hongyu
Zhao, Hongyu
中科院分区:
数学4区
文献类型:
--
作者:
Wang, Tao;Zhao, Hongyu

文献摘要

被引文献

相似文献

组合数据在许多实际问题中自然出现,对这些数据的分析提出了许多统计挑战,特别是在高维数据中。在这篇文章中,我们考虑的问题,子成分的选择与成分协变量的回归,其中的协变量之间的关系可以表示为一棵树与叶节点对应的协变量。假设树结构是可用的先验知识,我们采用了对称版本的线性对数对比度模型,并提出了一个树引导的正则化方法,这种结构化的子成分的选择。我们的方法是基于一种新的惩罚函数,它结合了树结构信息节点,鼓励选择子树级别的子组合。我们表明,这个优化问题可以制定为一个广义套索问题,其解决方案可以有效地使用现有的算法计算。应用人类肠道微生物组研究和模拟提出的方法与l(1)正则化方法,其中树结构信息不被利用的性能进行比较。
Compositional data arise naturally in many practical problems and the analysis of such data presents many statistical challenges, especially in high dimensions. In this article, we consider the problem of subcomposition selection in regression with compositional covariates, where the relationships among the covariates can be represented by a tree with leaf nodes corresponding to covariates. Assuming that the tree structure is available as prior knowledge, we adopt a symmetric version of the linear log contrast model, and propose a tree-guided regularization method for this structured subcomposition selection. Our method is based on a novel penalty function that incorporates the tree structure information node-by-node, encouraging the selection of subcompositions at subtree levels. We show that this optimization problem can be formulated as a generalized lasso problem, the solution of which can be computed efficiently using existing algorithms. An application to a human gut microbiome study and simulations are presented to compare the performance of the proposed method with an l(1) regularization method where the tree structure information is not utilized.