The effect of statistical normalization on network propagation scores.

The effect of statistical normalization on network propagation scores.
复制标题

统计归一化对网络传播分数的影响。

DOI:
10.1093/bioinformatics/btaa896
复制
发表时间:
2021
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Perera-Lluna,Alexandre
Perera-Lluna,Alexandre
中科院分区:
--
文献类型:
--
作者:
Picart-Armada,Sergio;Thompson,WesleyK;Buil,Alfonso;Perera-Lluna,Alexandre

文献摘要

相似文献

网络扩散和标签传播是计算生物学中的基本工具,其应用包括基因-疾病关联、蛋白质功能预测和模块发现。最近,一些出版物已经在传播过程之后引入了置换分析,这是由于担心网络拓扑结构会使扩散分数产生偏差。这打开了这样的扩散过程中的每个应用程序的统计特性和存在的偏见的问题。在这项工作中,我们描述了置换分析背后的一些常见的零模型和扩散分数的统计特性。我们对三个案例研究的七个扩散分数进行了基准测试:酵母相互作用组上的合成信号、蛋白质-蛋白质相互作用网络上的模拟差异基因表达以及另一个相互作用网络上的前瞻性基因集预测。为了清楚起见,所有的数据集是基于二进制标签,但我们也提出了定量labels.ResultsDiffusion分数从二进制标签开始的标签编码的影响,并表现出问题依赖的拓扑偏见,可以删除的统计归一化。参数和非参数标准化通过独立于编码和均衡偏倚来解决这两点。我们确定并量化了两个来源的偏差-平均值和方差-产生的性能差异时,归一化的分数。我们提供了封闭的公式,并显示了零协方差是如何与图的谱特性。尽管没有一个建议的评分系统地优于其他评分,但当所寻求的阳性标签与偏倚不一致时,首选标准化。我们的结论是,消除偏倚的决定应该是问题和数据驱动的,即基于偏倚的定量分析及其与阳性实体的关系。可验证性该代码可在www.example.com上公开获得,本文所依据的数据可在https://github.com/b2slab/retroDataSupplementary上获得。补充数据可在Bioinformaticsonline上获得。https://github.com/b2slab/diffuBench
MotivationNetwork diffusion and label propagation are fundamental tools in computational biology, with applications like gene–disease association, protein function prediction and module discovery. More recently, several publications have introduced a permutation analysis after the propagation process, due to concerns that network topology can bias diffusion scores. This opens the question of the statistical properties and the presence of bias of such diffusion processes in each of its applications. In this work, we characterized some common null models behind the permutation analysis and the statistical properties of the diffusion scores. We benchmarked seven diffusion scores on three case studies: synthetic signals on a yeast interactome, simulated differential gene expression on a protein–protein interaction network and prospective gene set prediction on another interaction network. For clarity, all the datasets were based on binary labels, but we also present theoretical results for quantitative labels.ResultsDiffusion scores starting from binary labels were affected by the label codification and exhibited a problem-dependent topological bias that could be removed by the statistical normalization. Parametric and non-parametric normalization addressed both points by being codification-independent and by equalizing the bias. We identified and quantified two sources of bias—mean value and variance—that yielded performance differences when normalizing the scores. We provided closed formulae for both and showed how the null covariance is related to the spectral properties of the graph. Despite none of the proposed scores systematically outperformed the others, normalization was preferred when the sought positive labels were not aligned with the bias. We conclude that the decision on bias removal should be problem and data-driven, i.e. based on a quantitative analysis of the bias and its relation to the positive entities.AvailabilityThe code is publicly available at https://github.com/b2slab/diffuBench and the data underlying this article are available at https://github.com/b2slab/retroDataSupplementary informationSupplementary data are available atBioinformaticsonline.