High-dimensional log-error-in-variable regression with applications to microbial compositional data analysis

High-dimensional log-error-in-variable regression with applications to microbial compositional data analysis
复制标题

DOI:
10.1093/biomet/asab020
复制
发表时间:
2021-08-26
期刊:
影响因子:
2.7
通讯作者:
Zhang, Anru R.
Zhang, Anru R.
中科院分区:
数学2区
文献类型:
--
作者:
Shi, Pixu;Zhou, Yuchen;Zhang, Anru R.

文献摘要

被引文献

相似文献

在微生物组和基因组研究中,组成数据的回归一直是鉴定与临床表型相关的微生物分类群或基因的重要工具。为了解释测序深度的变化,通常使用经典的对数对比模型,其中读数计数被标准化为组成。然而,零读取计数和协变量的随机性仍然是关键问题。我们介绍了一个令人惊讶的简单,可解释的和有效的方法,通过透镜的一种新的高维对数误差在变量回归模型的成分数据回归估计。所提出的方法提供了对测序数据的校正,可能存在过度分散,同时避免了对零读数计数的任何主观估算。我们提供了理论上的理由与匹配的估计误差的上限和下限。通过对真实的数据的分析和仿真研究,说明了该方法的优点。
In microbiome and genomic studies, the regression of compositional data has been a crucial tool for identifying microbial taxa or genes that are associated with clinical phenotypes. To account for the variation in sequencing depth, the classic log-contrast model is often used where read counts are normalized into compositions. However, zero read counts and the randomness in covariates remain critical issues. We introduce a surprisingly simple, interpretable and efficient method for the estimation of compositional data regression through the lens of a novel high-dimensional log-error-in-variable regression model. The proposed method provides corrections on sequencing data with possible overdispersion and simultaneously avoids any subjective imputation of zero read counts. We provide theoretical justifications with matching upper and lower bounds for the estimation error. The merit of the procedure is illustrated through real data analysis and simulation studies.