Measuring rank robustness in scored protein interaction networks

Measuring rank robustness in scored protein interaction networks
复制标题

DOI:
10.1186/s12859-019-3036-6
复制
发表时间:
2019-08-28
期刊:
影响因子:
3
通讯作者:
Deane, Charlotte M.
Deane, Charlotte M.
中科院分区:
生物学4区
文献类型:
--
作者:
Bozhilova, Lyuba, V;Whitmore, Alan, V;Deane, Charlotte M.

文献摘要

被引文献

相似文献

背景蛋白质相互作用数据库通常根据现有的实验证据为每个记录的相互作用提供置信度分数。然后通过对这些分数进行阈值处理来构建蛋白质相互作用网络(PIN),以便仅包括足够高质量的相互作用。这些网络用于使用度或介数中心性等度量来识别生物相关的基序或节点。这种类型的分析可能对阈值的选择敏感。如果节点度量对于提取生物信号是有用的,则它应该在不同的合理置信度分数阈值处获得的PIN上引起类似的节点排名。结果我们提出了三个措施-排名的连续性,可识别性和不稳定性-来评估如何强大的节点度量的分数阈值的变化。我们应用我们的措施,以25个指标,并确定四个最强大的:边缘的数量在第一步自我网络,以及留一出的平均冗余,边缘的平均数量在第一步自我网络,和自然连接的差异。我们的测量结果显示,来自不同物种和数据来源的PIN之间存在良好的一致性。对综合生成的评分网络的分析表明,鲁棒性结果是特定于上下文的,并且取决于网络拓扑结构和分数如何跨网络边缘放置。结论由于与蛋白质相互作用检测相关的不确定性,因此网络结构,PIN分析是可重复的,它应该在不同的置信度阈值产生相似的结果。我们证明,虽然某些节点的指标是强大的阈值选择,这并不总是如此。令人鼓舞的是,我们的研究结果表明,有一些指标在从不同数据库和不同评分程序构建的网络中具有鲁棒性。
Background Protein interaction databases often provide confidence scores for each recorded interaction based on the available experimental evidence. Protein interaction networks (PINs) are then built by thresholding on these scores, so that only interactions of sufficiently high quality are included. These networks are used to identify biologically relevant motifs or nodes using metrics such as degree or betweenness centrality. This type of analysis can be sensitive to the choice of threshold. If a node metric is to be useful for extracting biological signal, it should induce similar node rankings across PINs obtained at different reasonable confidence score thresholds. Results We propose three measures-rank continuity, identifiability, and instability-to evaluate how robust a node metric is to changes in the score threshold. We apply our measures to twenty-five metrics and identify four as the most robust: the number of edges in the step-1 ego network, as well as the leave-one-out differences in average redundancy, average number of edges in the step-1 ego network, and natural connectivity. Our measures show good agreement across PINs from different species and data sources. Analysis of synthetically generated scored networks shows that robustness results are context-specific, and depend both on network topology and on how scores are placed across network edges. Conclusion Due to the uncertainty associated with protein interaction detection, and therefore network structure, for PIN analysis to be reproducible, it should yield similar results across different confidence score thresholds. We demonstrate that while certain node metrics are robust with respect to threshold choice, this is not always the case. Promisingly, our results suggest that there are some metrics that are robust across networks constructed from different databases, and different scoring procedures.