Spotlite: web application and augmented algorithms for predicting co-complexed proteins from affinity purification--mass spectrometry data.

Spotlite: web application and augmented algorithms for predicting co-complexed proteins from affinity purification--mass spectrometry data.
复制标题

DOI:
10.1021/pr5008416
复制
发表时间:
2014-12-05
影响因子:
4.4
通讯作者:
Major, Michael B.
Major, Michael B.
中科院分区:
生物学2区
文献类型:
--
作者:
Goldfarb, Dennis;Hast, Bridgid E.;Wang, Wei;Major, Michael B.

文献摘要

参考文献

被引文献

相似文献

由亲和纯化和质谱(APMS)方法定义的蛋白质-蛋白质相互作用遭受高错误发现率。因此,在网络构建和解释之前,候选交互列表必须被修剪掉污染物,这在历史上是一项昂贵且耗时的任务。近年来,已经开发了许多计算方法来从APMS实验揭示的数百种相互作用中识别真正的相互作用。在这里,几个流行的算法的比较分析显示,在他们的分类准确性,这是支持他们的不同的评分策略的互补性。因此,我们使用两种准确且计算效率高的方法作为使用随机森林算法进行机器学习的特征。此外,我们开发了新的数学模型,包括各种间接数据,如mRNA共表达,基因本体和同源蛋白质相互作用的分类问题内的功能。我们表明,我们的方法,我们称之为Spotlite,优于现有的方法在四个不同的和公共的APMS数据集。由于现有APMS评分方法的实施需要许多实验室以外的计算专业知识,我们创建了一个用户友好和快速的网络应用程序,用于APMS数据评分,分析,注释和网络可视化,用于新的和现有的数据(http://152.19.87.94:8080/spotlite)。通过重点分析KEAP 1 E3泛素连接酶,建立了Spotlite及其可视化平台在APMS数据中揭示物理,功能和疾病相关特征的实用性。
Protein-protein interactions defined by affinity purification and mass spectrometry (APMS) approaches suffer from high false discovery rates. Consequently, the candidate interaction lists must be pruned of contaminants before network construction and interpretation, historically an expensive and time-intensive task. In recent years, numerous computational methods have been developed to identify genuine interactions from hundreds revealed by APMS experiments. Here, comparative analysis of several popular algorithms revealed complementarity in their classification accuracies, which is supported by their divergent scoring strategies. As such, we used two accurate and computationally efficient methods as features for machine learning using the Random Forest algorithm. Additionally, we developed novel mathematical models to include a variety of indirect data, such as mRNA co-expression, gene ontologies and homologous protein interactions as features within the classification problem. We show that our method, which we call Spotlite, outperforms existing methods on four diverse and public APMS datasets. Because implementation of existing APMS scoring methods requires computational expertise beyond many laboratories, we created a user-friendly and fast web application for APMS data scoring, analysis, annotation and network visualization, for use on new and existing data (http://152.19.87.94:8080/spotlite). The utility of Spotlite and its visualization platform for revealing physical, functional and disease-relevant characteristics within APMS data is established through a focused analysis of the KEAP1 E3 ubiquitin ligase.
DOI: 10.1186/1471-2105-11-562
发表时间: 2010-11-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Jain S;Bader GD
通讯作者: Bader GD
DOI: 10.1038/nature04532
发表时间: 2006-03-30
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Aloy, P;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1101/gr.153002
发表时间: 2002-10-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Deng, MH;Mehta, S;Chen, T
通讯作者: Chen, T
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1186/1471-2105-4-2
发表时间: 2003-01-13
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Bader, GD;Hogue, CW
通讯作者: Hogue, CW