Improvement of domain linker prediction by incorporating loop-length-dependent characteristics

Improvement of domain linker prediction by incorporating loop-length-dependent characteristics
复制标题

DOI:
10.1002/bip.20361
复制
发表时间:
2006-01-01
期刊:
影响因子:
2.9
通讯作者:
Kuroda, Y
Kuroda, Y
中科院分区:
生物学4区
文献类型:
--
作者:
Tanaka, T;Yokoyama, S;Kuroda, Y

文献摘要

被引文献

相似文献

蛋白质分解成结构域,可以在隔离中折叠,是功能蛋白质组学和结构蛋白质组学中的一个重要问题。在这里,我们分析了域间和域内循环序列(分别称为域链接器和非链接器环),并计算了域链接器似然分数,用于开发域边界预测协议。分析结果证实了我们先前的结果,即在连接环和非连接环之间,甘氨酸、脯氨酸、天冬氨酸、天冬氨酸、赖氨酸和组氨酸的氨基酸组成显著不同。然而,详细的检查发现,氨基酸组成的偏差实际上取决于环长。事实上,在短连接环和非连接环中观察到甘氨酸、脯氨酸和天冬氨酸有显著的频率偏差,而在长连接环和非连接环中观察到天冬氨酸、脯氨酸、天冬氨酸和赖氨酸的频率偏差。最后,我们加入了这种依赖于环长的氨基酸组成偏倚吗?一种简单的连接子预测方法,预测连接子的特异度为40.6%,敏感度为36.1%。这些数字比我们以前的预测协议分别高出4.4%和2.4%,该预测协议没有纳入环长相关特性。这一结果对实验蛋白质切割具有实际意义,因为通过随机切割蛋白质序列获得稳定折叠的结构域的概率估计为12.6%。(C)2005年威利期刊公司。
Protein dissection into structural domains that can fold in isolation is an important issue in both functional and structural proteomics. Here, we analyzed inter- and intradomain loop sequences (respectively named domain linker and nonlinker loops) and computed a domain linker likelihood score, which was used for developing a domain boundary prediction protocol. The analysis confirmed our previous results indicating that the amino acid composition in terms of glycine, proline, aspartic acid, asparagine, lysine, and histidine significantly differs between linker and nonlinker loops. However, a detailed examination revealed that the amino acid composition bias actually depends on the loop length. Indeed, significant frequency deviations were observed for glycine, proline, and aspartic acid in short linker and nonlinker loops, whereas deviations were observed for aspartic acid, proline, asparagine, and lysine in long linker and nonlinker loops. Finally, we incorporated this loop-length-dependent amino acid composition bias it? a simple linker prediction protocol, which predicted linkers with a 40.6% specificity and a 36.1% sensitivity. These figures are 4.4 and 2.4% higher than those obtained with our former prediction protocol that does not incorporate loop-length-dependent characteristics. This result should have practical significance for experimental protein dissection, since the probability of obtaining a stably folding structural domain by randomly dissecting a protein sequence is estimated to be 12.6%. (c) 2005 Wiley Periodicals, Inc.