Linking cells across single-cell modalities by synergistic matching of neighborhood structure.

Linking cells across single-cell modalities by synergistic matching of neighborhood structure.
复制标题

通过邻域结构的协同匹配跨单细胞模式连接细胞。

DOI:
10.1093/bioinformatics/btac481
复制
发表时间:
2022
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Noble,WilliamStafford
Noble,WilliamStafford
中科院分区:
--
文献类型:
--
作者:
Hristov,BorislavH;Bilmes,JeffreyA;Noble,WilliamStafford

文献摘要

相似文献

动机多种多样的实验方法可用于表征复杂生物样品中单细胞的不同特性。然而,由于这些测量技术通常是破坏性的,研究人员经常从不相交的细胞子集中获得互补的测量结果,从而提供细胞生物过程的碎片视图。这就需要能够整合不相交的多组学数据的计算工具。由于不同的测量通常不共享任何特征,因此该问题需要以无监督的方式进行集成。最近,已经提出了几种方法,项目的细胞测量到一个共同的潜在空间,并试图对齐相应的低维manifold.ResultsIn这项研究中,我们提出了一种方法,Synmatch,它产生了一个直接匹配的细胞之间的方式,通过利用在每个模态的邻域结构的信息。Synmatch依赖于直觉,即在一个测量空间中接近的单元在另一个测量空间中也应该接近。这使我们能够制定的匹配问题作为一个约束的超模优化问题,可以有效地解决附近的结构。我们表明,我们的方法成功地匹配小的真实的多组学数据集的细胞,并表现良好,与最近发表的最先进的方法相比。此外,我们还证明了Synmatch能够扩展到数千个细胞的大型数据集。可用性和实施本手稿中使用的Synmatch代码和数据可在https://github.com/Noble-Lab/synmatch上获得。
MotivationA wide variety of experimental methods are available to characterize different properties of single cells in a complex biosample. However, because these measurement techniques are typically destructive, researchers are often presented with complementary measurements from disjoint subsets of cells, providing a fragmented view of the cell’s biological processes. This creates a need for computational tools capable of integrating disjoint multi-omics data. Because different measurements typically do not share any features, the problem requires the integration to be done in unsupervised fashion. Recently, several methods have been proposed that project the cell measurements into a common latent space and attempt to align the corresponding low-dimensional manifolds.ResultsIn this study, we present an approach, Synmatch, which produces a direct matching of the cells between modalities by exploiting information about neighborhood structure in each modality. Synmatch relies on the intuition that cells which are close in one measurement space should be close in the other as well. This allows us to formulate the matching problem as a constrained supermodular optimization problem over neighborhood structures that can be solved efficiently. We show that our approach successfully matches cells in small real multi-omics datasets and performs favorably when compared with recently published state-of-the-art methods. Further, we demonstrate that Synmatch is capable of scaling to large datasets of thousands of cells.Availability and implementationThe Synmatch code and data used in this manuscript are available at https://github.com/Noble-Lab/synmatch.