The art of using t-SNE for single-cell transcriptomics

The art of using t-SNE for single-cell transcriptomics
复制标题

DOI:
10.1038/s41467-019-13056-x
复制
发表时间:
2019-11-28
影响因子:
16.6
通讯作者:
Berens, Philipp
Berens, Philipp
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Kobak, Dmitry;Berens, Philipp

文献摘要

被引文献

相似文献

单细胞转录组学产生了不断增长的数据集,其中包含来自数百万个细胞的数千个基因的RNA表达水平。常见的数据分析管道包括用于在二维中可视化数据的降维步骤,最常见的是使用t分布随机邻居嵌入(t-SNE)执行。它擅长揭示高维数据中的局部结构,但幼稚的应用程序往往遭受严重的缺点,例如,数据的全局结构不能准确地表示。在这里,我们将介绍如何规避这些陷阱,并开发一个协议来创建更忠实的t-SNE可视化。它包括PCA初始化,高学习率和多尺度相似性内核;对于非常大的数据集,我们还使用夸张和基于下采样的初始化。我们使用已发表的单细胞RNA-seq数据集来证明与t-SNE的幼稚应用相比,该方案产生了上级结果。
Single-cell transcriptomics yields ever growing data sets containing RNA expression levels for thousands of genes from up to millions of cells. Common data analysis pipelines include a dimensionality reduction step for visualising the data in two dimensions, most frequently performed using t-distributed stochastic neighbour embedding (t-SNE). It excels at revealing local structure in high-dimensional data, but naive applications often suffer from severe shortcomings, e.g. the global structure of the data is not represented accurately. Here we describe how to circumvent such pitfalls, and develop a protocol for creating more faithful t-SNE visualisations. It includes PCA initialisation, a high learning rate, and multi-scale similarity kernels; for very large data sets, we additionally use exaggeration and downsampling-based initialisation. We use published single-cell RNA-seq data sets to demonstrate that this protocol yields superior results compared to the naive application of t-SNE.