Armadillo: Domain boundary prediction by amino acid composition

Armadillo: Domain boundary prediction by amino acid composition
复制标题

DOI:
10.1016/j.jmb.2005.05.037
复制
发表时间:
2005-07-29
影响因子:
5.6
通讯作者:
Hogue, CWV
Hogue, CWV
中科院分区:
生物学2区
文献类型:
--
作者:
Dumontier, M;Yao, R;Hogue, CWV

文献摘要

被引文献

相似文献

蛋白质结构域的识别和注释是精确确定分子功能的关键步骤。蛋白质结构测定的计算和实验方法都可能受到大的多结构域蛋白质或柔性接头区域的阻碍。域及其边界的知识可以通过允许研究人员研究一组更小且可能更成功的替代方案来降低蛋白质结构确定的实验成本。目前的结构域预测方法通常依赖于与保守结构域的序列相似性,因此不适合检测保守性差或孤儿蛋白质中的结构域结构。我们提出了一种简单的计算方法,从序列信息中识别蛋白质结构域连接子及其边界,我们的结构域预测器Armadillo(http:armadillo.blueprint.org)使用任何氨基酸索引将蛋白质序列转换为平滑的数字轮廓,从该轮廓可以预测结构域和结构域边界。我们使用非冗余结构数据集从结构域连接子的氨基酸组成中推导出一个称为结构域连接子倾向指数(DLI)的氨基酸指数。该指数表明Pro和Gly显示出连接残基的倾向,而小的疏水残基则没有。Armadillo从Z分数分布预测结构域接头边界,并在双结构域、单接头数据集中获得35%的DLI灵敏度(在接头的+/-20个残基内)。DLI和基于熵的氨基酸指数的组合将两个结构域蛋白的总体犰狳敏感性增加到56%。此外,Armadillo对多结构域蛋白质的预测灵敏度达到37%,超过了大多数其他预测方法。Armadillo提供了一种简单但有效的方法,通过该方法可以以合理的灵敏度获得结构域边界的预测。Armadillo应该被证明是一个有价值的工具,用于快速描绘保守性差的蛋白质或那些没有序列邻居的蛋白质结构域。作为第一线预测器,域元预测器可以与Armadillo预测一起产生改进的结果。(c)2005年由Elsevier Ltd.出版
The identification and annotation of protein domains provides a critical step in the accurate determination of molecular function. Both computational and experimental methods of protein structure determination may be deterred by large multi-domain proteins or flexible linker regions. Knowledge of domains and their boundaries may reduce the experimental cost of protein structure determination by allowing researchers to work on a set of smaller and possibly more successful alternatives. Current domain prediction methods often rely on sequence similarity to conserved domains and as such are poorly suited to detect domain structure in poorly conserved or orphan proteins. We present here a simple computational method to identify protein domain linkers and their boundaries from sequence information alone.Our domain predictor, Armadillo (http://armadillo.blueprint.org), uses any amino acid index to convert a protein sequence to a smoothed numeric profile from which domains and domain boundaries may be predicted. We derived an amino acid index called the domain linker propensity index (DLI) from the amino acid composition of domain linkers using a non-redundant structure dataset. The index indicates that Pro and Gly show a propensity for linker residues while small hydrophobic residues do not. Armadillo predicts domain linker boundaries from Z-score distributions and obtains 35% sensitivity with DLI in a two-domain, single-linker dataset (within +/- 20 residues from linker). The combination of DLI and an entropy-based amino acid index increases the overall Armadillo sensitivity to 56% for two domain proteins. Moreover, Armadillo achieves 37% sensitivity for multi-domain proteins, surpassing most other prediction methods.Armadillo provides a simple, but effective method by which prediction of domain boundaries can be obtained with reasonable sensitivity. Armadillo should prove to be a valuable tool for rapidly delineating protein domains in poorly conserved proteins or those with no sequence neighbors. As a first-line predictor, domain meta-predictors could yield improved results with Armadillo predictions. (c) 2005 Published by Elsevier Ltd.