Probabilistic cell-type assignment of single-cell RNA-seq for tumor microenvironment profiling

Probabilistic cell-type assignment of single-cell RNA-seq for tumor microenvironment profiling
复制标题

DOI:
10.1038/s41592-019-0529-1
复制
发表时间:
2019-10-01
期刊:
影响因子:
48
通讯作者:
Shah, Sohrab P.
Shah, Sohrab P.
中科院分区:
生物学1区
文献类型:
--
作者:
Zhang, Allen W.;O'Flanagan, Ciara;Shah, Sohrab P.

文献摘要

被引文献

相似文献

单细胞RNA测序使复杂组织分解成功能不同的细胞类型成为可能。通常,研究人员希望通过无监督聚类将细胞分配到细胞类型,然后进行手动注释或通过“映射”到现有数据。然而,手动解释的规模很难大的数据集,映射方法需要纯化或预先注释的数据,这两种方法都容易产生批量效应。为了克服这些问题,我们提出了CellAssign,这是一个概率模型,它利用细胞类型标记基因的先验知识将单细胞RNA测序数据注释为预定义的或从头开始的细胞类型。CellAssign以高度可扩展的方式在大型数据集上自动分配单元格,同时控制批次和样品效果。我们通过对高级别浆液性卵巢癌和滤泡性淋巴瘤中肿瘤微环境组成的广泛模拟和分析,证明了CellAssign的优势。
Single-cell RNA sequencing has enabled the decomposition of complex tissues into functionally distinct cell types. Often, investigators wish to assign cells to cell types through unsupervised clustering followed by manual annotation or via 'mapping' to existing data. However, manual interpretation scales poorly to large datasets, mapping approaches require purified or pre-annotated data and both are prone to batch effects. To overcome these issues, we present CellAssign, a probabilistic model that leverages prior knowledge of cell-type marker genes to annotate single-cell RNA sequencing data into predefined or de novo cell types. CellAssign automates the process of assigning cells in a highly scalable manner across large datasets while controlling for batch and sample effects. We demonstrate the advantages of CellAssign through extensive simulations and analysis of tumor microenvironment composition in high-grade serous ovarian cancer and follicular lymphoma.