Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors.

Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors.
复制标题

DOI:
10.1038/nbt.4091
复制
发表时间:
2018-06
影响因子:
46.9
通讯作者:
Marioni JC
Marioni JC
中科院分区:
工程技术1区
文献类型:
--
作者:
Haghverdi L;Lun ATL;Morgan MD;Marioni JC

文献摘要

被引文献

相似文献

在不同实验室和不同时间产生的大规模单细胞RNA测序(scRNA-seq)数据集包含批次效应,可能会影响这些数据的整合和解释。现有的scRNA-seq分析方法错误地假设细胞群体的组成在各批次中是已知的或相同的。我们提出了一种策略,批量校正的基础上检测的相互最近的邻居(MNN)在高维表达空间。我们的方法不依赖于预定义的或相同的人口组成的批次,只需要一个子集的人口之间共享批次。我们使用模拟和真实的scRNA-seq数据集证明了我们的方法优于现有方法。使用多个基于液滴的scRNA-seq数据集,我们证明了我们的MNN批量效应校正方法可扩展到大量细胞。
Large-scale single-cell RNA sequencing (scRNA-seq) datasets that are produced in different laboratories and at different times contain batch effects that could compromise integration and interpretation of these data. Existing scRNA-seq analysis methods incorrectly assume that the composition of cell populations is either known, or the same, across batches. We present a strategy for batch correction that is based on the detection of mutual nearest neighbours (MNN) in the high-dimensional expression space. Our approach does not rely on pre-defined or equal population compositions across batches, and only requires that a subset of the population be shared between batches. We demonstrate the superiority of our approach over existing methods using both simulated and real scRNA-seq data sets. Using multiple droplet-based scRNA-seq data sets, we demonstrate that our MNN batch-effect correction method scales to large numbers of cells.