Robust differential abundance test in compositional data

Robust differential abundance test in compositional data
复制标题

成分数据中稳健的差异丰度测试

DOI:
10.1093/biomet/asac029
复制
发表时间:
2022
期刊:
影响因子:
2.7
通讯作者:
Wang, Shulei
Wang, Shulei
中科院分区:
数学2区
文献类型:
--
作者:
Wang, Shulei

文献摘要

相似文献

成分数据的差异丰度测试在各种生物医学应用中是必不可少的和基本的,例如单细胞,批量RNA-seq和微生物组数据分析。然而,由于成分约束和零计数的数据中的流行,差分丰度分析的成分数据仍然是一个复杂的和未解决的统计问题。本文提出了一种新的差异丰度检验,鲁棒差异丰度检验,以解决这些挑战。与现有方法相比,稳健差分丰度检验方法简单,计算效率高,对组分数据集中普遍存在的零计数具有鲁棒性,能够考虑数据的组分性质,并具有在一般情况下控制错误发现的理论保证.此外,在存在观测到的协变量的情况下,稳健的差异丰度检验可以与协变量平衡技术一起工作,以消除潜在的混淆效应并得出可靠的结论。提出的测试应用于几个数值例子,其优点是使用模拟和真实的数据集证明。
Differential abundance tests for compositional data are essential and fundamental in various biomedical applications, such as single-cell, bulk RNA-seq and microbiome data analysis. However, because of the compositional constraint and the prevalence of zero counts in the data, differential abundance analysis on compositional data remains a complicated and unsolved statistical problem. This article proposes a new differential abundance test, the robust differential abundance test, to address these challenges. Compared with existing methods, the robust differential abundance test is simple and computationally efficient, is robust to prevalent zero counts in compositional datasets, can take the data’s compositional nature into account, and has a theoretical guarantee of controlling false discoveries in a general setting. Furthermore, in the presence of observed covariates, the robust differential abundance test can work with covariate-balancing techniques to remove potential confounding effects and draw reliable conclusions. The proposed test is applied to several numerical examples, and its merits are demonstrated using both simulated and real datasets.