Few-shot learning for classification of novel macromolecular structures in cryo-electron tomograms.

Few-shot learning for classification of novel macromolecular structures in cryo-electron tomograms.
复制标题

DOI:
10.1371/journal.pcbi.1008227
复制
发表时间:
2020-11
影响因子:
4.3
通讯作者:
Xu M
Xu M
中科院分区:
生物学2区
文献类型:
--
作者:
Li R;Yu L;Zhou B;Zeng X;Wang Z;Yang X;Zhang J;Gao X;Jiang R;Xu M

文献摘要

参考文献

被引文献

相似文献

冷冻电子断层扫描(cryo-ET)提供了亚细胞成分在近天然状态下的3D可视化,并在单细胞中的亚分子分辨率,展示了在原位结构生物学中越来越重要的作用。然而,由于低信噪比(SNR),小尺寸的大分子,和细胞环境的高度复杂性,在冷冻ET数据中的大分子结构的系统识别和恢复仍然具有挑战性。亚断层图像的结构分类是这一任务的重要步骤。尽管由于数据收集自动化的进步,大量子断层图像的采集不再是障碍,但获得相同数量的结构标签是计算和劳动密集型的。另一方面,现有的基于深度学习的监督分类方法对标记数据的要求很高,并且从包含非常少的新结构标记的数据中快速学习新结构的能力有限。在这项工作中,我们提出了一种新的方法subtomogram分类的基础上少拍学习。利用该方法,通过实例嵌入,在测试数据中给定少量标记样本的情况下,可以对训练数据中的未知结构进行分类。在模拟和真实的数据集上进行了实验。我们的实验结果表明,我们可以对新结构进行推断,对于每个类别,只需五个标记样本,具有竞争力的准确度(在SNR = 0.1的模拟数据集上> 0.86),甚至一个样本的准确度为0.7644。在真实的数据集上的结果也是有希望的,在两种条件下的准确度> 0.9,并且在其中一个真实的数据集上甚至高达1。我们的方法实现了显着的改善相比,基线方法,并具有较强的能力,推广到其他细胞成分。冷冻电子断层成像技术已广泛应用于结构生物学中,以亚分子分辨率和近天然状态提供单细胞内结构的三维透视。鉴定冷冻电子断层图像中所含的大分子是进一步分析这些大分子的结构和功能的必要步骤。最近的研究表明,监督学习擅长于对断层图像子体积(称为子断层图像)中的大分子进行分类。然而,由于细胞中的大多数结构对我们来说是未知的,标记子断层图像中的大分子是耗时的,劳动密集型的,并且难以实现,这给监督学习带来了困难。我们提出了一种计算方法来区分亚断层图像中的大分子与少量的标记数据。我们在一些注释良好的结构上训练了我们的模型,并应用该模型对具有少量标记示例的新结构进行分类。我们在模拟数据集和真实的数据集上进行了实验,我们的结果表明,我们的方法可以在每个类别不超过5个样本的情况下对新结构实现有竞争力的分类准确性。我们的方法可以帮助用很少的例子从冷冻电子断层图像中快速准确地检测新发现的结构,加速对结构的后续研究,从而可能促进对细胞功能的进一步解释。
Cryo-electron tomography (cryo-ET) provides 3D visualization of subcellular components in the near-native state and at sub-molecular resolutions in single cells, demonstrating an increasingly important role in structural biology in situ. However, systematic recognition and recovery of macromolecular structures in cryo-ET data remain challenging as a result of low signal-to-noise ratio (SNR), small sizes of macromolecules, and high complexity of the cellular environment. Subtomogram structural classification is an essential step for such task. Although acquisition of large amounts of subtomograms is no longer an obstacle due to advances in automation of data collection, obtaining the same number of structural labels is both computation and labor intensive. On the other hand, existing deep learning based supervised classification approaches are highly demanding on labeled data and have limited ability to learn about new structures rapidly from data containing very few labels of such new structures. In this work, we propose a novel approach for subtomogram classification based on few-shot learning. With our approach, classification of unseen structures in the training data can be conducted given few labeled samples in test data through instance embedding. Experiments were performed on both simulated and real datasets. Our experimental results show that we can make inference on new structures given only five labeled samples for each class with a competitive accuracy (> 0.86 on the simulated dataset with SNR = 0.1), or even one sample with an accuracy of 0.7644. The results on real datasets are also promising with accuracy > 0.9 on both conditions and even up to 1 on one of the real datasets. Our approach achieves significant improvement compared with the baseline method and has strong capabilities of generalizing to other cellular components. Cryo-electron tomography has been widely used in structral biology to provide a three-dimensional perspective on intracellular structures at sub-molecular resolutions and near-native states in single cells. Identifying the macromolecules contained in cryo-electron tomograms is an essential step for further analysis of the structure and function of these macromolecules. Recent studies have shown that supervised learning excels in the classification of macromolecules in subvolumes of tomograms (called subtomograms). However, since most structures in cells are unknown to us, labeling macromolecules in subtomograms is time-consuming, labor-intensive, and hard to implement, which brings difficulties to supervised learning. We proposed a computational method to distinguish the macromolecules in subtomograms with few labeled data. We trained our model on some well-annotated structures and apply the model to classify new structures with few labeled examples. We conducted experiments on both simulated datasets and real datasets, and our results suggest that our method could achieve competitive classification accuracy on new structures with no more than five samples for each class. Our method can help to quickly and accurately detect newly-discovered structures from cryo-electron tomograms with few examples, accelerating subsequent research on the structures, and thus possibly promoting further interpretation of cellular functions.
DOI: 10.1242/jcs.171967
发表时间: 2016-02-01
影响因子: 4
作者:
Irobalieva, Rossitza N.;Martins, Bruno;Medalia, Ohad
通讯作者: Medalia, Ohad
DOI: 10.1016/j.ceb.2011.11.002
发表时间: 2012-02-01
影响因子: 7.5
作者:
Volkmann, Niels
通讯作者: Volkmann, Niels
DOI: 10.1016/j.str.2009.10.009
发表时间: 2009-12-09
期刊: STRUCTURE
影响因子: 5.7
作者:
Scheres, Sjors H. W.;Melero, Roberto;Carazo, Jose-Maria
通讯作者: Carazo, Jose-Maria
DOI: 10.1016/j.cell.2017.12.030
发表时间: 2018-02-08
期刊: Cell
影响因子: 64.5
作者:
Guo Q;Lehmer C;Martínez-Sánchez A;Rudack T;Beck F;Hartmann H;Pérez-Berlanga M;Frottin F;Hipp MS;Hartl FU;Edbauer D;Baumeister W;Fernández-Busnadiego R
通讯作者: Fernández-Busnadiego R
DOI: 10.1016/j.jsb.2011.08.012
发表时间: 2012-01-01
影响因子: 3
作者:
Rigort, Alexander;Guenther, David;Hege, Hans-Christian
通讯作者: Hege, Hans-Christian