A semi-supervised Bayesian approach for simultaneous protein sub-cellular localisation assignment and novelty detection.

A semi-supervised Bayesian approach for simultaneous protein sub-cellular localisation assignment and novelty detection.
复制标题

DOI:
10.1371/journal.pcbi.1008288
复制
发表时间:
2020-11
影响因子:
4.3
通讯作者:
Kirk PDW
Kirk PDW
中科院分区:
生物学2区
文献类型:
--
作者:
Crook OM;Geladaki A;Nightingale DJH;Vennard OL;Lilley KS;Gatto L;Kirk PDW

文献摘要

参考文献

被引文献

相似文献

细胞被划分为复杂的微环境,允许同步进行一系列专门的生物过程。因此,确定蛋白质在这些区室中的一个或多个的亚细胞定位可以是确定其功能的第一步。高通量和高精度的基于质谱的亚细胞蛋白质组学方法现在可以同时揭示数千种蛋白质的定位。然后,机器学习算法通常用于进行蛋白质-细胞器分配。然而,这些算法受到不充分和不完整的注释的限制。我们提出了一个半监督贝叶斯方法来检测新奇,允许发现额外的,以前未注释的亚细胞壁龛。在我们的模型中的推理是在贝叶斯框架中进行的,使我们能够量化蛋白质分配到新的亚细胞小生境以及新发现的隔间数量的不确定性。我们将我们的方法应用于10个基于质谱的空间蛋白质组数据集,代表了不同的实验方案。将我们的方法应用于hyperLOPIT数据集,通过回收未注释的染色质相关蛋白的富集来验证其实用性,并揭示了原始分析中未识别的亚核区室化。此外,利用酿酒酵母的亚细胞蛋白质组学数据,我们发现了一组新的蛋白质贩运从ER到早期高尔基体。总体而言,我们证明了新奇检测的潜力,可以产生当前方法所错过的生物相关小生境。
The cell is compartmentalised into complex micro-environments allowing an array of specialised biological processes to be carried out in synchrony. Determining a protein’s sub-cellular localisation to one or more of these compartments can therefore be a first step in determining its function. High-throughput and high-accuracy mass spectrometry-based sub-cellular proteomic methods can now shed light on the localisation of thousands of proteins at once. Machine learning algorithms are then typically employed to make protein-organelle assignments. However, these algorithms are limited by insufficient and incomplete annotation. We propose a semi-supervised Bayesian approach to novelty detection, allowing the discovery of additional, previously unannotated sub-cellular niches. Inference in our model is performed in a Bayesian framework, allowing us to quantify uncertainty in the allocation of proteins to new sub-cellular niches, as well as in the number of newly discovered compartments. We apply our approach across 10 mass spectrometry based spatial proteomic datasets, representing a diverse range of experimental protocols. Application of our approach to hyperLOPIT datasets validates its utility by recovering enrichment with chromatin-associated proteins without annotation and uncovers sub-nuclear compartmentalisation which was not identified in the original analysis. Moreover, using sub-cellular proteomics data from Saccharomyces cerevisiae, we uncover a novel group of proteins trafficking from the ER to the early Golgi apparatus. Overall, we demonstrate the potential for novelty detection to yield biologically relevant niches that are missed by current approaches.
DOI: 10.1073/pnas.0506958103
发表时间: 2006-04-25
影响因子: 11.1
作者:
Dunkley, TPJ;Hester, S;Lilley, KS
通讯作者: Lilley, KS
DOI: 10.1371/journal.pone.0134053
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Cabasso O;Pekar O;Horowitz M
通讯作者: Horowitz M
DOI: 10.1371/journal.pcbi.1004920
发表时间: 2016-05
影响因子: 4.3
作者:
Breckels LM;Holden SB;Wojnar D;Mulvey CM;Christoforou A;Groen A;Trotter MW;Kohlbacher O;Lilley KS;Gatto L
通讯作者: Gatto L
DOI: 10.1038/ncomms9992
发表时间: 2016-01-12
影响因子: 16.6
作者:
Christoforou A;Mulvey CM;Breckels LM;Geladaki A;Hurrell T;Hayward PC;Naake T;Gatto L;Viner R;Martinez Arias A;Lilley KS
通讯作者: Lilley KS
DOI: 10.1093/bioinformatics/btu013
发表时间: 2014-05-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Gatto L;Breckels LM;Wieczorek S;Burger T;Lilley KS
通讯作者: Lilley KS