Partitioning of Minimotifs Based on Function with Improved Prediction Accuracy

Partitioning of Minimotifs Based on Function with Improved Prediction Accuracy
复制标题

DOI:
10.1371/journal.pone.0012276
复制
发表时间:
2010-08-19
期刊:
影响因子:
3.7
通讯作者:
Schiller, Martin R.
Schiller, Martin R.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Rajasekaran, Sanguthevar;Mi, Tian;Schiller, Martin R.

文献摘要

被引文献

相似文献

背景:微基序是蛋白质中短的连续肽序列,已知其在至少一种其他蛋白质中具有功能。微基序预测的主要限制之一是假阳性限制了这种方法的有用性。作为解决这个问题的一步,我们已经建立,实施和测试了一个新的数据驱动的算法,减少假阳性predictions.Methodology/主要发现:某些域和minimotifs被称为是强烈相关的一个已知的细胞过程或分子功能。因此,我们假设,通过将微基序预测限制于其中含有蛋白质和靶蛋白质的微基序具有相关细胞或分子功能的那些,预测更可能是准确的。该过滤器在Minimotif Miner中使用来自Gene Ontology的函数注释实现。我们还结合了两个过滤器,是基于完全不同的原则,这种组合的过滤器有一个更好的可预测性比个别component.Conclusions/Significance:测试这些功能的过滤器上已知的和随机的minimotifs显示,他们是能够从假阳性分离真正的图案。特别地,对于细胞功能过滤器,未被过滤器去除的已知minimotifs的百分比类似于随机minimotifs的4.6倍。对于分子功能滤波器,该比率类似于2.9。这些结果,与已发表的频率分数过滤器的比较,强烈表明,新的过滤器区分真正的图案从随机背景具有良好的信心。函数过滤器和频率分数过滤器的组合比这两个单独的过滤器执行得更好。
Background: Minimotifs are short contiguous peptide sequences in proteins that are known to have a function in at least one other protein. One of the principal limitations in minimotif prediction is that false positives limit the usefulness of this approach. As a step toward resolving this problem we have built, implemented, and tested a new data-driven algorithm that reduces false-positive predictions.Methodology/Principal Findings: Certain domains and minimotifs are known to be strongly associated with a known cellular process or molecular function. Therefore, we hypothesized that by restricting minimotif predictions to those where the minimotif containing protein and target protein have a related cellular or molecular function, the prediction is more likely to be accurate. This filter was implemented in Minimotif Miner using function annotations from the Gene Ontology. We have also combined two filters that are based on entirely different principles and this combined filter has a better predictability than the individual components.Conclusions/Significance: Testing these functional filters on known and random minimotifs has revealed that they are capable of separating true motifs from false positives. In particular, for the cellular function filter, the percentage of known minimotifs that are not removed by the filter is similar to 4.6 times that of random minimotifs. For the molecular function filter this ratio is similar to 2.9. These results, together with the comparison with the published frequency score filter, strongly suggest that the new filters differentiate true motifs from random background with good confidence. A combination of the function filters and the frequency score filter performs better than these two individual filters.