Regression Models for Compositional Data: General Log-Contrast Formulations, Proximal Optimization, and Microbiome Data Applications

Regression Models for Compositional Data: General Log-Contrast Formulations, Proximal Optimization, and Microbiome Data Applications
复制标题

DOI:
10.1007/s12561-020-09283-2
复制
发表时间:
2020-06-19
影响因子:
1
通讯作者:
Mueller, Christian L.
Mueller, Christian L.
中科院分区:
其他
文献类型:
--
作者:
Combettes, Patrick L.;Mueller, Christian L.

文献摘要

被引文献

相似文献

成分数据集在科学中无处不在,包括地质学、生态学和微生物学。在微生物组研究中,成分数据主要来自高通量序列分析实验。这些数据包括微生物在其自然栖息地的组成,并经常与协变量测量配对,表征物理化学栖息地特性或宿主生理学。推断微生物组成与栖息地或宿主特异性协变量数据之间的简约统计关联是探索性数据分析的重要一步。将组成协变量与连续结果联系起来的标准统计模型是线性对数对比模型。该模型将响应描述为原始组合的对数比的线性组合,并通过正则化扩展到高维设置。在这篇贡献中,我们提出了线性对数对比回归的一般凸优化模型,其中包括许多以前的建议作为特殊情况。我们引入了一种近似算法,该算法精确地解决了约束优化问题,并保证了严格的收敛性。我们通过调查几个模型实例在土壤和肠道微生物组数据分析任务上的性能来说明我们方法的多功能性。
Compositional data sets are ubiquitous in science, including geology, ecology, and microbiology. In microbiome research, compositional data primarily arise from high-throughput sequence-based profiling experiments. These data comprise microbial compositions in their natural habitat and are often paired with covariate measurements that characterize physicochemical habitat properties or the physiology of the host. Inferring parsimonious statistical associations between microbial compositions and habitat- or host-specific covariate data is an important step in exploratory data analysis. A standard statistical model linking compositional covariates to continuous outcomes is the linear log-contrast model. This model describes the response as a linear combination of log-ratios of the original compositions and has been extended to the high-dimensional setting via regularization. In this contribution, we propose a general convex optimization model for linear log-contrast regression which includes many previous proposals as special cases. We introduce a proximal algorithm that solves the resulting constrained optimization problem exactly with rigorous convergence guarantees. We illustrate the versatility of our approach by investigating the performance of several model instances on soil and gut microbiome data analysis tasks.