Compositional knockoff filter for high-dimensional regression analysis of microbiome data.

Compositional knockoff filter for high-dimensional regression analysis of microbiome data.
复制标题

用于微生物组数据高维回归分析的组合仿冒过滤器。

DOI:
10.1111/biom.13336
复制
发表时间:
2021-09
期刊:
影响因子:
1.9
通讯作者:
Zhan X
Zhan X
中科院分区:
数学3区
文献类型:
--
作者:
Srinivasan A;Xue L;Zhan X

文献摘要

参考文献

被引文献

相似文献

微生物组数据分析中的一项关键任务是探索感兴趣的标量响应与大量微生物分类群之间的关联,这些分类群被总结为不同分类学水平的组成数据。受微生物组精细映射的启发,我们提出了一种两步组成敲除滤波器(CKF),以在微生物组组成数据的高维线性对数对比回归分析中提供有效的有限样本错误发现率(FDR)控制。在第一步中,我们提出了一个新的组成筛选程序,以删除无关紧要的微生物类群,同时保留必要的总和为零的约束。在第二步中,我们扩展敲除过滤器,以确定组成数据的稀疏回归模型中的重要微生物类群。因此,从与在预先指定的FDR阈值下的响应相关的高维微生物分类群中选择微生物的子集。我们研究了所提出的两步过程的理论性质,包括确定性筛选和有效的错误发现控制。我们在数值模拟研究中展示了这些特性,将我们的方法与一些现有的方法进行比较,并在控制标称FDR的同时显示新方法的功率增益。所提出的方法的潜在有用性也说明了应用到炎症性肠病数据集,以确定影响宿主基因表达的微生物类群。
A critical task in microbiome data analysis is to explore the association between a scalar response of interest and a large number of microbial taxa that are summarized as compositional data at different taxonomic levels. Motivated by fine-mapping of the microbiome, we propose a two-step compositional knockoff filter (CKF) to provide the effective finite-sample false discovery rate (FDR) control in high-dimensional linear log-contrast regression analysis of microbiome compositional data. In the first step, we propose a new compositional screening procedure to remove insignificant microbial taxa while retaining the essential sum-to-zero constraint. In the second step, we extend the knockoff filter to identify the significant microbial taxa in the sparse regression model for compositional data. Thereby, a subset of the microbes is selected from the high-dimensional microbial taxa as related to the response under a pre-specified FDR threshold. We study the theoretical properties of the proposed two-step procedure, including both sure screening and effective false discovery control. We demonstrate these properties in numerical simulation studies to compare our methods to some existing ones and show power gain of the new method while controlling the nominal FDR. The potential usefulness of the proposed method is also illustrated with application to an inflammatory bowel disease dataset to identify microbial taxa that influence host gene expressions.
DOI: 10.1186/s13059-015-0637-x
发表时间: 2015-04-08
期刊: Genome biology
影响因子: 12.3
作者:
Morgan XC;Kabakchiev B;Waldron L;Tyler AD;Tickle TL;Milgrom R;Stempak JM;Gevers D;Xavier RJ;Silverberg MS;Huttenhower C
通讯作者: Huttenhower C
DOI: 10.1093/bib/bbx104
发表时间: 2019-01-01
影响因子: 9.5
作者:
Hawinkel, Stijn;Mattiello, Federico;Thas, Olivier
通讯作者: Thas, Olivier
DOI: 10.1214/18-aos1755
发表时间: 2019-10-01
影响因子: 4.5
作者:
Barber, Rina Foygel;Candes, Emmanuel J.
通讯作者: Candes, Emmanuel J.
DOI: 10.1093/biostatistics/kxy025
发表时间: 2019-10-01
期刊: Biostatistics (Oxford, England)
影响因子: --
作者:
Tang, Zheng-Zheng;Chen, Guanhua
通讯作者: Chen, Guanhua
DOI: 10.1214/12-aoas592
发表时间: 2013-03-01
期刊: The annals of applied statistics
影响因子: --
作者:
Chen J;Li H
通讯作者: Li H