Protein ranking by semi-supervised network propagation.

Protein ranking by semi-supervised network propagation.
复制标题

通过半监督网络传播进行蛋白质排名。

DOI:
10.1186/1471-2105-7-s1-s10
复制
发表时间:
2006-03-20
期刊:
影响因子:
3
通讯作者:
Noble, WS
Noble, WS
中科院分区:
生物学4区
文献类型:
--
作者:
Weston, J;Kuang, R;Leslie, C;Noble, WS

文献摘要

被引文献

相似文献

生物学家经常在DNA或蛋白质数据库中搜索与给定查询序列共享进化或功能关系的序列。传统的搜索方法,如BLAST和PSI-BLAST,专注于检测统计上显著的成对序列比对,并且经常错过更细微的序列相似性。机器学习社区最近的工作表明,利用由这些成对相似性定义的网络的全局结构,可以帮助检测比纯粹的局部测量更遥远的关系。我们审查RankProp,排名算法,利用全球网络结构的蛋白质之间的相似性关系在数据库中进行扩散操作的蛋白质相似性网络加权边缘。原始的RankProp算法是无监督的。在这里,我们描述了一个半监督版本的算法,使用标记的例子。三种可能的方式纳入标签信息被认为是:(i)作为一个验证集的模型选择,(ii)学习一个新的网络,通过选择传递函数用于一个给定的查询,(iii)估计边权重,这衡量推断结构相似性的概率。以人类管理的蛋白质结构数据库为基准,原始RankProp算法比PSI-BLAST等局部网络搜索算法有了显着改进。此外,我们在这里表明,标记的数据可以用来学习网络,而不需要估计传递函数的参数,并且在这个学习过的网络上的扩散比使用固定网络的原始RankProp算法产生更好的结果。为了从网络中获得最大的信息,需要使用标记和未标记的数据来提取局部和全局结构。
Biologists regularly search DNA or protein databases for sequences that share an evolutionary or functional relationship with a given query sequence. Traditional search methods, such as BLAST and PSI-BLAST, focus on detecting statistically significant pairwise sequence alignments and often miss more subtle sequence similarity. Recent work in the machine learning community has shown that exploiting the global structure of the network defined by these pairwise similarities can help detect more remote relationships than a purely local measure. We review RankProp, a ranking algorithm that exploits the global network structure of similarity relationships among proteins in a database by performing a diffusion operation on a protein similarity network with weighted edges. The original RankProp algorithm is unsupervised. Here, we describe a semi-supervised version of the algorithm that uses labeled examples. Three possible ways of incorporating label information are considered: (i) as a validation set for model selection, (ii) to learn a new network, by choosing which transfer function to use for a given query, and (iii) to estimate edge weights, which measure the probability of inferring structural similarity. Benchmarked on a human-curated database of protein structures, the original RankProp algorithm provides significant improvement over local network search algorithms such as PSI-BLAST. Furthermore, we show here that labeled data can be used to learn a network without any need for estimating parameters of the transfer function, and that diffusion on this learned network produces better results than the original RankProp algorithm with a fixed network. In order to gain maximal information from a network, labeled and unlabeled data should be used to extract both local and global structure.