Transparent single-cell set classification with kernel mean embeddings

Transparent single-cell set classification with kernel mean embeddings
复制标题

DOI:
10.1145/3535508.3545538
复制
发表时间:
2022-01
期刊:
Proceedings of the 13th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics
影响因子:
--
通讯作者:
Siyuan Shan;Vishal Baskaran;Haidong Yi;Jolene S Ranek;N. Stanley;Junier B. Oliva
Siyuan Shan;Vishal Baskaran;Haidong Yi;Jolene S Ranek;N. Stanley;Junier B. Oliva
中科院分区:
其他
文献类型:
--
作者:
Siyuan Shan;Vishal Baskaran;Haidong Yi;Jolene S Ranek;N. Stanley;Junier B. Oliva

文献摘要

相似文献

现代单细胞流和质量细胞术技术测量血液或组织样本中单个细胞的几种蛋白质的表达。因此,每个生物样本由一组数十万个多维细胞特征向量来表示,这导致使用机器学习模型来预测每个生物样本的关联表型的计算成本很高。如此大的集合基数也限制了机器学习模型的可解释性,因为很难跟踪每个单个细胞如何影响最终预测。我们建议使用核均值嵌入来编码每个轮廓生物样本的细胞景观。虽然我们的首要目标是建立一个更透明的模型,但我们发现我们的方法通过简单的线性分类器实现了与最先进的免门方法相当或更好的精度。因此,我们的模型包含的参数很少,但仍然具有类似于具有数百万个参数的深度学习模型的性能。与深度学习方法相比,该模型的线性和子选择步骤使其更容易解释分类结果。分析进一步表明,我们的方法允许将细胞异质性与临床表型联系起来,具有丰富的生物学解释力。
Modern single-cell flow and mass cytometry technologies measure the expression of several proteins of the individual cells within a blood or tissue sample. Each profiled biological sample is thus represented by a set of hundreds of thousands of multidimensional cell feature vectors, which incurs a high computational cost to predict each biological sample's associated phenotype with machine learning models. Such a large set cardinality also limits the interpretability of machine learning models due to the difficulty in tracking how each individual cell influences the ultimate prediction. We propose using Kernel Mean Embedding to encode the cellular landscape of each profiled biological sample. Although our foremost goal is to make a more transparent model, we find that our method achieves comparable or better accuracies than the state-of-the-art gating-free methods through a simple linear classifier. As a result, our model contains few parameters but still performs similarly to deep learning models with millions of parameters. In contrast with deep learning approaches, the linearity and sub-selection step of our model makes it easy to interpret classification results. Analysis further shows that our method admits rich biological interpretability for linking cellular heterogeneity to clinical phenotype.