A novel approach to remove the batch effect of single-cell data

A novel approach to remove the batch effect of single-cell data
复制标题

一种消除单细胞数据批量效应的新方法

DOI:
10.1038/s41421-019-0114-x
复制
发表时间:
2019
期刊:
影响因子:
33.5
通讯作者:
Tian Weidong
Tian Weidong
中科院分区:
生物学1区
文献类型:
--
作者:
Zhang Feng;Wu Yu;Tian Weidong

文献摘要

相似文献

亲爱的编辑,分析来自不同批次的单细胞RNA测序(scRNA-seq)数据是一项具有挑战性的任务。常用的批次效应去除方法,如Combat 2,3,最初是为微阵列或批量RNA-seq数据开发的,在某些情况下可能不适合单细胞分析4。最近,针对单细胞数据开发了几种批次效应去除工具。其中一种方法称为典型相关分析(CCA)子空间对齐(在Seurat中实现)4,它进行典型相关分析并使用动态时间规整来对齐不同批次的子空间。然而,当不同批次的细胞类型极度不平衡时,CCA可能会丢失方差最大的子空间(可以通过PCA识别),从而导致错误的对齐结果。为了基于正确的单元格对齐从PCA子空间中去除批效应,一种称为fastmnn5的方法检测不同批次单元格的相互最近邻居(MNN),然后使用MNN来校正每个PCA子空间中的值。虽然fastMNN被证明具有良好的性能,但在实践中,它的运行时间长,并且由于PCA子空间中值的校正而缺乏可解释性。一种基于图的批平衡KNN (BBKNN) 6方法通过在不同批次的类似细胞之间建立连接来减少批效应。然而,BBKNN只生成最终向量(UMAP) 7,使得无法跟踪调整。在这项研究中,我们提出了一种新的方法,称为批效应去除(BEER),用于合并来自不同批次的scRNA-seq数据。BEER的创新之处在于,它利用从不同批次中识别出的相互最近(MN)细胞对的相关性来识别相关性较差(即潜在的高批次效应)的PCA子空间,然后从进一步的分析中删除这些子空间。由于BEER不改变PCA子空间中的任何值,因此BEER产生的结果是可跟踪且易于解释的。通过使用细胞类型不平衡基准,我们发现BEER比四个代表性的批效应去除工具有明显的优势:Combat、Seurat (CCA对齐)、fastMNN和BBKNN。BEER已经在r中实现。BEER的输入是来自两个不同批次的两个表达式矩阵(UMI或其他未缩放的表达式格式)。这两个表达矩阵的行名和列名分别是基因名和细胞名。BEER的工作流程包括两个主要部分(图1a)。在第一部分中,对于每个表达式矩阵,BEER对数据进行预处理,并进行t分布随机邻居嵌入(tSNE) 8,将数据转换为一维值。由于tSNE在scRNA-seq分析领域的鲁棒性和良好的性能,因此被用于一维降维9。BEER根据一维值的顺序对细胞进行分组(每组中默认的细胞数为10),然后汇总组中每个细胞的表达概况,以获得该组的代表性表达概况。接下来,BEER计算Kendall 's tau来评估来自两个批次的每对细胞组的距离,并识别两个批次之间的所有MN对细胞组。在第二部分中,BEER直接组合了两个表达式矩阵,对数据进行规范化,并进行PCA生成子空间的个数(默认为50)。由于mn配对细胞组中的两个细胞组代表了这两个批中最相似的组,因此如果没有批效应,它们在每个PCA子空间中应该具有相似的值。因此,通过计算
Dear Editor, Analyzing single-cell RNA sequencing (scRNA-seq) data from different batches is a challenging task 1. The commonly used batch-effect removal methods, eg Combat 2, 3 were initially developed for microarray or bulk RNA-seq data, and may not be appropriate for single-cell analysis in some situations 4. Recently, several batch-effect removal tools specific for single-cell data have been developed. One of them is called canonical correlation analysis (CCA) subspace alignment (implemented in Seurat) 4, which conducts CCA and uses dynamic time warping to align the subspaces of different batches. However, CCA may lose the subspaces with the largest possible variance (can be identified by PCA), leading to wrong alignment result when the cell types of different batches are extremely imbalanced. To remove batcheffect from the PCA subspaces based on the correct cell alignment, a method called fastMNN 5 detects mutual nearest neighbors (MNN) of cells in different batches, and then uses the MNN to correct the values in each PCA subspace. Although fastMNN was shown to have a good performance, in practice it has long running time, and also lacks the explainability because of the correction of values in PCA subspace. A graph-based method named batch balanced KNN (BBKNN) 6 reduces batch-effect by creating connections between analogous cells in different batches. However, BBKNN only generates the final vectors (UMAP) 7, making it impossible to track the adjustment. In this study, we present a novel method called batch effect remover (BEER) for combining scRNA-seq data from different batches. The originality of BEER is that it uses the correlation of mutual nearest (MN) cell pairs identified from different batches to identify PCA sub-spaces with poor correlation (ie, latent high batcheffect), and then removes these subspaces from further analysis. Because BEER does not change any values in PCA subspaces, the results produced by BEER are trackable and easily explainable. By using a cell-type imbalanced benchmark, we show that BEER has a clear advantage over four representative batch-effect removal tools: Combat, Seurat (CCA alignment), fastMNN, and BBKNN.BEER has been implemented in R. The inputs of BEER are two expression matrices (UMI or other un-scaled expression format) coming from two different batches. The row and column names of the two expression matrices are gene and cell names, respectively. The workflow of BEER includes two main parts (Fig. 1 a). In the first part, for each expression matrix, BEER preprocesses the data and conducts t-distributed stochastic neighbor embedding (tSNE) 8 to transfer the data into onedimension values. tSNE is used to do one-dimension reduction because of its robustness and well-recognized performance in the field of scRNA-seq analysis 9. BEER groups cells (default number of cells in each group is 10) based on the order of the one-dimension values, and then aggregate the expression profiles of each cell in a group to obtain the representative expression profile for that group. Next, BEER calculates a Kendall’s tau to evaluate the distance of each pair of cell group from two batches, and identifies all MN pairs of cell groups in between the two batches. In the second part, BEER directly combines two expression matrices, normalizes the data, and conduct PCA to produce a number (default is 50) of subspaces. Because two cell groups in a MN-paired cell groups represent the most similar groups in those two batches, they should have similar values in each PCA subspace if there is no batch effect. Thus, by calculating the