How significant is a protein structure similarity with TM-score=0.5?

How significant is a protein structure similarity with TM-score=0.5?
复制标题

DOI:
10.1093/bioinformatics/btq066
复制
发表时间:
2010-04-01
期刊:
影响因子:
5.8
通讯作者:
Zhang, Yang
Zhang, Yang
中科院分区:
生物学3区
文献类型:
--
作者:
Xu, Jinrui;Zhang, Yang

文献摘要

被引文献

相似文献

动机:蛋白质结构相似性通常通过均方根偏差、全局距离测试分数和模板建模分数(TM-score)来衡量。然而,分数本身无法提供有关结构相似性有多重要的信息。此外,它缺乏分数和传统折叠分类之间的定量关系。本文旨在回答两个问题:(i)TM-score 的统计显着性是什么? (ii) 给定特定 TM 分数,两个蛋白质具有相同折叠的概率是多少?结果:我们首先对 PDB 中的 6684 个非同源单域蛋白质进行全面无间隙结构匹配,发现 TM 分数遵循极值分布。这些数据使我们能够为每个 TM 分数分配一个 P 值,该值衡量两个随机选择的蛋白质获得相同或更高 TM 分数的机会。例如,TM 得分为 0.5,其 P 值为 5.5x10(-7),这意味着我们需要考虑至少 180 万个随机蛋白质对才能获得不低于 0.5 的 TM 得分。其次,我们检查来自三个数据集 SCOP、CATH 以及 SCOP 和 CATH 的共识的相同折叠蛋白质的后验概率。研究发现,不同数据集的后验概率在 TM-score = 0.5 附近具有类似的快速相变。这一发现表明 TM 分数可以用作蛋白质拓扑分类的近似但定量的标准,即。 e. TM-score > 0.5 的蛋白质对大多处于相同的折叠状态,而 TM-score 的蛋白质对大多处于相同的折叠状态
Motivation: Protein structure similarity is often measured by root mean squared deviation, global distance test score and template modeling score (TM-score). However, the scores themselves cannot provide information on how significant the structural similarity is. Also, it lacks a quantitative relation between the scores and conventional fold classifications. This article aims to answer two questions: (i) what is the statistical significance of TM-score? (ii) What is the probability of two proteins having the same fold given a specific TM-score?Results: We first made an all-to-all gapless structural match on 6684 non-homologous single-domain proteins in the PDB and found that the TM-scores follow an extreme value distribution. The data allow us to assign each TM-score a P-value that measures the chance of two randomly selected proteins obtaining an equal or higher TM-score. With a TM-score at 0.5, for instance, its P-value is 5.5x10(-7), which means we need to consider at least 1.8 million random protein pairs to acquire a TM-score of no less than 0.5. Second, we examine the posterior probability of the same fold proteins from three datasets SCOP, CATH and the consensus of SCOP and CATH. It is found that the posterior probability from different datasets has a similar rapid phase transition around TM-score = 0.5. This finding indicates that TM-score can be used as an approximate but quantitative criterion for protein topology classification, i. e. protein pairs with a TM-score >0.5 are mostly in the same fold while those with a TM-score