Learning from Heterogeneous Data Sources: An Application in Spatial Proteomics.

Learning from Heterogeneous Data Sources: An Application in Spatial Proteomics.
复制标题

DOI:
10.1371/journal.pcbi.1004920
复制
发表时间:
2016-05
影响因子:
4.3
通讯作者:
Gatto L
Gatto L
中科院分区:
生物学2区
文献类型:
--
作者:
Breckels LM;Holden SB;Wojnar D;Mulvey CM;Christoforou A;Groen A;Trotter MW;Kohlbacher O;Lilley KS;Gatto L

文献摘要

被引文献

相似文献

蛋白质的亚细胞定位是一种重要的翻译后调节机制,可以使用高通量质谱(MS)进行分析。这些基于MS的空间蛋白质组学实验使我们能够在受控条件下精确定位特定系统中数千种蛋白质的亚细胞分布。高通量MS方法的最新进展为细胞生物学领域提供了大量的实验空间蛋白质组学数据。然而,还有许多第三方数据源,如免疫荧光显微镜或蛋白质注释和序列,它们代表了丰富而庞大的互补信息来源。我们提出了一个独特的迁移学习分类框架,该框架利用最近邻或支持向量机系统来整合异构数据源,从而大大提高亚细胞蛋白质分配的数量和质量。我们证明了我们的算法的效用,通过评估五个实验数据集,从四个不同的物种结合四个不同的辅助数据源,以分类蛋白质的亚细胞区室数十具有很高的泛化精度。我们进一步将该方法应用于多能小鼠胚胎干细胞的实验中,以分类一组以前未知的蛋白质,并验证我们的研究结果对最近的高分辨率地图的小鼠干细胞蛋白质组。该方法作为开源Bioconductor pRoloc套件的一部分分发,用于空间蛋白质组学数据分析。蛋白质的亚细胞定位对其在所有细胞过程中的功能至关重要;定位于其预期微环境(例如细胞器、囊泡或大分子复合物)的蛋白质将满足适于追求其分子功能的相互作用伴侣和生化条件。因此,可靠和系统地研究蛋白质定位的可靠数据和方法,以及细胞生物学社区所依赖的蛋白质错误定位和蛋白质运输的中断是必不可少的。在这里,我们提出了一种依赖于基于实验质谱的数据和辅助来源(例如GO注释、第三方软件的输出、蛋白质-蛋白质相互作用或免疫细胞化学数据)的最佳整合来推断蛋白质定位的方法。我们发现,与以前仅使用实验数据推断亚细胞定位的单一数据分类器相比,在这些不同的数据源中应用迁移学习算法大大提高了亚细胞蛋白质分配的数量和可靠性。我们展示了我们的方法如何在与异构免费提供的第三方资源整合后不损害生物相关的实验特定信号。不同数据源的整合是生物学数据密集型领域的一个重要挑战,我们预计这里提出的迁移学习方法将被证明对生物学的许多领域有用,以统一从不同但互补的来源获得的数据。
Sub-cellular localisation of proteins is an essential post-translational regulatory mechanism that can be assayed using high-throughput mass spectrometry (MS). These MS-based spatial proteomics experiments enable us to pinpoint the sub-cellular distribution of thousands of proteins in a specific system under controlled conditions. Recent advances in high-throughput MS methods have yielded a plethora of experimental spatial proteomics data for the cell biology community. Yet, there are many third-party data sources, such as immunofluorescence microscopy or protein annotations and sequences, which represent a rich and vast source of complementary information. We present a unique transfer learning classification framework that utilises a nearest-neighbour or support vector machine system, to integrate heterogeneous data sources to considerably improve on the quantity and quality of sub-cellular protein assignment. We demonstrate the utility of our algorithms through evaluation of five experimental datasets, from four different species in conjunction with four different auxiliary data sources to classify proteins to tens of sub-cellular compartments with high generalisation accuracy. We further apply the method to an experiment on pluripotent mouse embryonic stem cells to classify a set of previously unknown proteins, and validate our findings against a recent high resolution map of the mouse stem cell proteome. The methodology is distributed as part of the open-source Bioconductor pRoloc suite for spatial proteomics data analysis. Sub-cellular localisation of proteins is critical to their function in all cellular processes; proteins localising to their intended micro-environment, e.g organelles, vesicles or macro-molecular complexes, will meet the interaction partners and biochemical conditions suitable to pursue their molecular function. Therefore, sound data and methods to reliably and systematically study protein localisation, and hence their mis-localisation and the disruption of protein trafficking, that are relied upon by the cell biology community, are essential. Here we present a method to infer protein localisation relying on the optimal integration of experimental mass spectrometry-based data and auxiliary sources, such as GO annotation, outputs from third-party software, protein-protein interactions or immunocytochemistry data. We found that the application of transfer learning algorithms across these diverse data sources considerably improves on the quantity and reliability of sub-cellular protein assignment, compared to single data classifiers previously applied to infer sub-cellular localisation using experimental data only. We show how our method does not compromise biologically relevant experimental-specific signal after integration with heterogeneous freely available third-party resources. The integration of different data sources is an important challenge in the data intensive world of biology and we anticipate the transfer learning methods presented here will prove useful to many areas of biology, to unify data obtained from different but complimentary sources.