Learning from Heterogeneous Data Sources: An Application in Spatial Proteomics

Learning from Heterogeneous Data Sources: An Application in Spatial Proteomics
复制标题

DOI:
10.1101/022152
复制
发表时间:
2015-07
影响因子:
4.3
通讯作者:
L. Breckels;S. Holden;David Wojnar;C. Mulvey;Andy Christoforou;A. Groen;M. Trotter;O. Kohlbacher
L. Breckels;S. Holden;David Wojnar;C. Mulvey;Andy Christoforou;A. Groen;M. Trotter;O. Kohlbacher
中科院分区:
生物学2区
文献类型:
--
作者:
L. Breckels;S. Holden;David Wojnar;C. Mulvey;Andy Christoforou;A. Groen;M. Trotter;O. Kohlbacher

文献摘要

相似文献

蛋白质的亚细胞定位是一种重要的翻译后调节机制,可以使用高通量质谱(MS)进行分析。这些基于ms的空间蛋白质组学实验使我们能够在受控条件下精确定位特定系统中数千种蛋白质的亚细胞分布。高通量质谱方法的最新进展为细胞生物学界提供了大量的实验空间蛋白质组学数据。然而,有许多第三方数据源,如免疫荧光显微镜或蛋白质注释和序列,它们代表了丰富而庞大的补充信息来源。我们提出了一个独特的迁移学习分类框架,该框架利用最近邻或支持向量机系统来整合异构数据源,以显着提高亚细胞蛋白质分配的数量和质量。我们通过评估来自四个不同物种的五个实验数据集,结合四个不同的辅助数据源,证明了我们的算法的实用性,以高泛化精度将蛋白质分类到数十个亚细胞区室。我们进一步将该方法应用于多能小鼠胚胎干细胞的实验,对一组以前未知的蛋白质进行分类,并通过最近的小鼠干细胞蛋白质组高分辨率图验证我们的发现。该方法作为开源Bioconductor pRoloc套件的一部分发布,用于空间蛋白质组学数据分析。简写LOPIT细胞器蛋白的同位素标记定位PCP蛋白相关分析ML机器学习TL迁移学习支持向量机PCA主成分分析GO基因本体CC细胞室iTRAQ等压标记相对定量和绝对定量TMT质谱标记MS质谱分析
Sub-cellular localisation of proteins is an essential post-translational regulatory mechanism that can be assayed using high-throughput mass spectrometry (MS). These MS-based spatial proteomics experiments enable us to pinpoint the sub-cellular distribution of thousands of proteins in a specific system under controlled conditions. Recent advances in high-throughput MS methods have yielded a plethora of experimental spatial proteomics data for the cell biology community. Yet, there are many third-party data sources, such as immunofluorescence microscopy or protein annotations and sequences, which represent a rich and vast source of complementary information. We present a unique transfer learning classification framework that utilises a nearest-neighbour or support vector machine system, to integrate heterogeneous data sources to considerably improve on the quantity and quality of sub-cellular protein assignment. We demonstrate the utility of our algorithms through evaluation of five experimental datasets, from four different species in conjunction with four different auxiliary data sources to classify proteins to tens of sub-cellular compartments with high generalisation accuracy. We further apply the method to an experiment on pluripotent mouse embryonic stem cells to classify a set of previously unknown proteins, and validate our findings against a recent high resolution map of the mouse stem cell proteome. The methodology is distributed as part of the open-source Bioconductor pRoloc suite for spatial proteomics data analysis. Abbreviations LOPIT Localisation of Organelle Proteins by Isotope Tagging PCP Protein Correlation Profiling ML Machine learning TL Transfer learning SVM Support vector machine PCA Principal component analysis GO Gene Ontology CC Cellular compartment iTRAQ Isobaric tags for relative and absolute quantitation TMT Tandem mass tags MS Mass spectrometry