SLiMFinder: a probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins.

SLiMFinder: a probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins.
复制标题

Slimfinder:一种概率方法,用于识别蛋白质中占代表性过多的短线线性基序。

DOI:
10.1371/journal.pone.0000967
复制
发表时间:
2007-10-03
期刊:
影响因子:
3.7
通讯作者:
Shields, Denis C.
Shields, Denis C.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Edwards, Richard J.;Davey, Norman E.;Shields, Denis C.

文献摘要

参考文献

被引文献

相似文献

蛋白质中的短线性基序(SLiM)是许多生物系统中具有根本重要性的功能微区。SLiM通常由一级蛋白质序列的3至10个氨基酸延伸组成,其中少至两个位点可能对活性重要,使得鉴定新型SLiM极其困难。特别是,很难区分一个随机重复出现的“主题”和一个真正过度代表的主题。不明确的氨基酸位置和/或限定的残基之间的可变长度的重复间隔区的排除进一步使问题复杂化。在本文中,我们提出了两个算法。SLiMBuild在蛋白质数据集中识别收敛进化的短基序。基序是通过将二聚体组合成更长的模式来构建的,只保留那些在足够数量的不相关蛋白质中出现的基序。识别具有固定氨基酸位置的基序,然后将其组合以并入氨基酸模糊性和可变长度的重复间隔区。该算法是计算效率相比,替代品,特别是当数据集包括同源蛋白质,并提供了很大的灵活性,在返回的基序的性质。SLiMCance算法估计偶然出现的返回基序的概率,校正数据集的大小和组成,并为每个基序分配一个重要性值。这些算法在一个软件包中实现,SLiMogram。SLiM的默认设置以100%的特异性识别已知的SLiM,并且在随机测试数据上具有低错误发现率。SLiMBuild的效率和SLiMCance的低错误发现率使SLiMBuild非常适合高通量基序发现和个体高质量分析。这样的分析真实的生物数据的例子,以及如何SLiM的结果可以帮助指导未来的发现,提供。在GNU许可证下,SLiMogram可从http://bioinformatics.ucd.ie/shields/software/slimfinder/免费下载。
Short linear motifs (SLiMs) in proteins are functional microdomains of fundamental importance in many biological systems. SLiMs typically consist of a 3 to 10 amino acid stretch of the primary protein sequence, of which as few as two sites may be important for activity, making identification of novel SLiMs extremely difficult. In particular, it can be very difficult to distinguish a randomly recurring “motif” from a truly over-represented one. Incorporating ambiguous amino acid positions and/or variable-length wildcard spacers between defined residues further complicates the matter. In this paper we present two algorithms. SLiMBuild identifies convergently evolved, short motifs in a dataset of proteins. Motifs are built by combining dimers into longer patterns, retaining only those motifs occurring in a sufficient number of unrelated proteins. Motifs with fixed amino acid positions are identified and then combined to incorporate amino acid ambiguity and variable-length wildcard spacers. The algorithm is computationally efficient compared to alternatives, particularly when datasets include homologous proteins, and provides great flexibility in the nature of motifs returned. The SLiMChance algorithm estimates the probability of returned motifs arising by chance, correcting for the size and composition of the dataset, and assigns a significance value to each motif. These algorithms are implemented in a software package, SLiMFinder. SLiMFinder default settings identify known SLiMs with 100% specificity, and have a low false discovery rate on random test data. The efficiency of SLiMBuild and low false discovery rate of SLiMChance make SLiMFinder highly suited to high throughput motif discovery and individual high quality analyses alike. Examples of such analyses on real biological data, and how SLiMFinder results can help direct future discoveries, are provided. SLiMFinder is freely available for download under a GNU license from http://bioinformatics.ucd.ie/shields/software/slimfinder/.
DOI: 10.1371/journal.pbio.0030405
发表时间: 2005-12
期刊: PLoS biology
影响因子: 9.8
作者:
Neduva V;Linding R;Su-Angrand I;Stark A;de Masi F;Gibson TJ;Lewis J;Serrano L;Russell RB
通讯作者: Russell RB
DOI: 10.1093/bioinformatics/14.1.55
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rigoutsos, I;Floratos, A
通讯作者: Floratos, A
DOI: 10.1093/bioinformatics/bti541
发表时间: 2005-08-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Dosztányi, Z;Csizmok, V;Simon, I
通讯作者: Simon, I
通过假定受体结合部位与乙型肝炎病毒相互作用的肽的鉴定和特征描述
DOI: 10.1128/jvi.01270-06
发表时间: 2007-04-01
影响因子: 5.4
作者:
Deng, Qiang;Zhai, Jian-wei;Xie, You-hua
通讯作者: Xie, You-hua
DOI: 10.1093/nar/gkl159
发表时间: 2006-07-01
影响因子: 14.9
作者:
Neduva V;Russell RB
通讯作者: Russell RB