A Bayesian mixture modelling approach for spatial proteomics.

A Bayesian mixture modelling approach for spatial proteomics.
复制标题

DOI:
10.1371/journal.pcbi.1006516
复制
发表时间:
2018-11
影响因子:
4.3
通讯作者:
Gatto L
Gatto L
中科院分区:
生物学2区
文献类型:
--
作者:
Crook OM;Mulvey CM;Kirk PDW;Lilley KS;Gatto L

文献摘要

参考文献

被引文献

相似文献

分析蛋白质的空间亚细胞分布对于全面了解蛋白质的功能具有重要意义。一些蛋白质可以在细胞内的单个位置发现,但多达一半的蛋白质可以驻留在多个位置,可以动态重新定位或驻留在未知的功能区室中。这些考虑导致将蛋白质与单个位置相关联的不确定性。目前,基于质谱(MS)的空间蛋白质组学依赖于监督机器学习算法,根据常见的梯度分布将蛋白质分配到亚细胞位置。然而,这样的方法未能量化与子蜂窝类分配相关联的不确定性。在这里,我们重新制定了我们进行统计分析的框架。我们提出了一个贝叶斯生成分类器的基础上高斯混合模型的概率分配蛋白质的亚细胞小生境,因此蛋白质有一个概率分布在亚细胞位置,贝叶斯计算使用期望最大化(EM)算法,以及马尔可夫链蒙特-卡罗(MCMC)。我们的方法允许蛋白质组范围内的不确定性量化,从而增加了空间蛋白质组学分析的进一步层。我们的框架是灵活的,允许许多不同的系统进行分析,并揭示了空间蛋白质组学的新的建模机会。我们发现我们的方法与当前最先进的机器学习方法相比具有竞争力,同时提供更多信息。我们强调了几个例子,其中基于支持向量机的分类无法得出任何结论,而使用我们的方法的不确定性量化提供了生物学上有趣的结果。据我们所知,这是基于MS的空间蛋白质组学数据的第一个贝叶斯模型。蛋白质的亚细胞定位提供了对亚细胞生物过程的见解。对于一个蛋白质执行其预期的功能,它必须定位到正确的亚细胞环境,无论是细胞器,囊泡或任何亚细胞龛。正确的亚细胞定位确保了蛋白质实现其分子功能的生化条件得到满足,以及接近其预期的相互作用伴侣。因此,蛋白质的错误定位改变了细胞的生物化学,并且可以破坏例如信号传导途径或抑制细胞周围物质的运输。蛋白质的亚细胞分布由于可以驻留在多个微环境中的蛋白质或那些在细胞内动态移动的蛋白质而变得复杂。预测蛋白质亚细胞定位的方法通常无法量化由亚细胞环境的复杂和动态性质引起的不确定性。在这里,我们提出了一种贝叶斯方法来分析蛋白质亚细胞定位。我们明确地对数据进行建模,并使用贝叶斯推理来量化预测中的不确定性。我们发现我们的方法与最先进的机器学习方法相比具有竞争力,并且还提供了不确定性量化。我们表明,有了这些额外的信息,我们可以更深入地了解细胞的基本生物化学。
Analysis of the spatial sub-cellular distribution of proteins is of vital importance to fully understand context specific protein function. Some proteins can be found with a single location within a cell, but up to half of proteins may reside in multiple locations, can dynamically re-localise, or reside within an unknown functional compartment. These considerations lead to uncertainty in associating a protein to a single location. Currently, mass spectrometry (MS) based spatial proteomics relies on supervised machine learning algorithms to assign proteins to sub-cellular locations based on common gradient profiles. However, such methods fail to quantify uncertainty associated with sub-cellular class assignment. Here we reformulate the framework on which we perform statistical analysis. We propose a Bayesian generative classifier based on Gaussian mixture models to assign proteins probabilistically to sub-cellular niches, thus proteins have a probability distribution over sub-cellular locations, with Bayesian computation performed using the expectation-maximisation (EM) algorithm, as well as Markov-chain Monte-Carlo (MCMC). Our methodology allows proteome-wide uncertainty quantification, thus adding a further layer to the analysis of spatial proteomics. Our framework is flexible, allowing many different systems to be analysed and reveals new modelling opportunities for spatial proteomics. We find our methods perform competitively with current state-of-the art machine learning methods, whilst simultaneously providing more information. We highlight several examples where classification based on the support vector machine is unable to make any conclusions, while uncertainty quantification using our approach provides biologically intriguing results. To our knowledge this is the first Bayesian model of MS-based spatial proteomics data. Sub-cellular localisation of proteins provides insights into sub-cellular biological processes. For a protein to carry out its intended function it must be localised to the correct sub-cellular environment, whether that be organelles, vesicles or any sub-cellular niche. Correct sub-cellular localisation ensures the biochemical conditions for the protein to carry out its molecular function are met, as well as being near its intended interaction partners. Therefore, mis-localisation of proteins alters cell biochemistry and can disrupt, for example, signalling pathways or inhibit the trafficking of material around the cell. The sub-cellular distribution of proteins is complicated by proteins that can reside in multiple micro-environments, or those that move dynamically within the cell. Methods that predict protein sub-cellular localisation often fail to quantify the uncertainty that arises from the complex and dynamic nature of the sub-cellular environment. Here we present a Bayesian methodology to analyse protein sub-cellular localisation. We explicitly model our data and use Bayesian inference to quantify uncertainty in our predictions. We find our method is competitive with state-of-the-art machine learning methods and additionally provides uncertainty quantification. We show that, with this additional information, we can make deeper insights into the fundamental biochemistry of the cell.
DOI: 10.1073/pnas.0506958103
发表时间: 2006-04-25
影响因子: 11.1
作者:
Dunkley, TPJ;Hester, S;Lilley, KS
通讯作者: Lilley, KS
DOI: 10.1371/journal.pcbi.1004920
发表时间: 2016-05
影响因子: 4.3
作者:
Breckels LM;Holden SB;Wojnar D;Mulvey CM;Christoforou A;Groen A;Trotter MW;Kohlbacher O;Lilley KS;Gatto L
通讯作者: Gatto L
DOI: 10.1038/ncomms9992
发表时间: 2016-01-12
影响因子: 16.6
作者:
Christoforou A;Mulvey CM;Breckels LM;Geladaki A;Hurrell T;Hayward PC;Naake T;Gatto L;Viner R;Martinez Arias A;Lilley KS
通讯作者: Lilley KS
DOI: 10.1093/bioinformatics/btu013
发表时间: 2014-05-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Gatto L;Breckels LM;Wieczorek S;Burger T;Lilley KS
通讯作者: Lilley KS
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J