A tracked approach for automated NMR assignments in proteins (TATAPRO)

A tracked approach for automated NMR assignments in proteins (TATAPRO)
复制标题

DOI:
10.1023/a:1008315111278
复制
发表时间:
2000-06-01
影响因子:
2.7
通讯作者:
Govil, G
Govil, G
中科院分区:
生物学3区
文献类型:
--
作者:
Atreya, HS;Sahu, SC;Govil, G

文献摘要

被引文献

相似文献

提出了一种利用三重共振实验数据对蛋白质中H-1(N)、C-13(alpha)、C-13(beta)、C-13 '/H-1(alpha)和N-15自旋进行序列特异性NMR归属的新方法。该算法TATAPRO(Tracked Automated Proteins)利用来自一组三重共振光谱的蛋白质一级序列和峰列表,所述三重共振光谱将H-1(N)和N-15化学位移与C-13(α)、C-13(β)和C-13 '/H-1(α)的化学位移相关联。从这样的相关性导出的信息用于创建由H-1(i)N、N-15(i)、C-13(i)α、C-13(i)β、C-13 ′(i)/H-1(i)α、C-13(i-1)α、C-13(i-1)β和C-12(i-1)′/H-1(i-1)α化学位移的所有可能集合组成的“主列表”。基于对来自BioMagResBank(BMRB)的蛋白质的C-13(alpha)和C-13(beta)化学位移数据的广泛统计分析,表明20个氨基酸残基可以分为8个不同的类别,每个类别被分配一个唯一的两位数代码。这样的代码用于标记master_list中的单个化学位移集,并将蛋白质一级序列翻译成称为pps_array的数组。然后,程序使用master_list沿着多肽链搜索给定氨基酸残基的相邻配偶体,并顺序地在任一侧分配最大可能的残基段。在这样做的时候,每个分配的残基都在一个名为assig_array的数组中跟踪,前面分配了两位数代码。然后将assig_array映射到pps_array上,用于序列特异性共振分配。该程序已被测试使用的实验数据的钙结合蛋白从溶组织内阿米巴(Eh-CaBP,15 kDa)具有实质性的内部序列同源性,并使用公开的数据在18-42 kDa的分子量范围内的其他四种蛋白质。在所有情况下,获得了几乎完整的序列特异性共振归属(> 95%)。此外,通过从为测试蛋白质创建的master_list中随机删除化学位移组来测试程序的可靠性。
A novel automated approach for the sequence specific NMR assignments of H-1(N), C-13(alpha), C-13(beta), C-13'/H-1(alpha) and N-15 spins in proteins, using triple resonance experimental data, is presented. The algorithm, TATAPRO (Tracked AuTomated Assignments in Proteins) utilizes the protein primary sequence and peak lists from a set of triple resonance spectra which correlate H-1(N) and N-15 chemical shifts with those of C-13(alpha), C-13(beta) and C-13'/H-1(alpha). The information derived from such correlations is used to create a 'master_list' consisting of all possible sets of H-1(i)N, N-15(i), C-13(i)alpha, C-13(i)beta, C-13'(i)/H-1(i)alpha, C-13(i-1)alpha, C-13(i-1)beta and C-12(i-1)'/H-1(i-1)alpha chemical shifts. On the basis of an extensive statistical analysis of C-13(alpha) and C-13(beta) chemical shift data of proteins derived from the BioMagResBank (BMRB), it is shown that the 20 amino acid residues can be grouped into eight distinct categories, each of which is assigned a unique two-digit code. Such a code is used to tag individual sets of chemical shifts in the master_list and also to translate the protein primary sequence into an array called pps_array. The program then uses the master_list to search for neighbouring partners of a given amino acid residue along the polypeptide chain and sequentially assigns a maximum possible stretch of residues on either side. While doing so, each assigned residue is tracked in an array called assig_array, with the two-digit code assigned earlier. The assig_array is then mapped onto the pps_array for sequence specific resonance assignment. The program has been tested using experimental data on a calcium binding protein from Entamoeba histolytica (Eh-CaBP, 15 kDa) having substantial internal sequence homology and using published data on four other proteins in the molecular weight range of 18-42 kDa. In all the cases, nearly complete sequence specific resonance assignments (> 95%) are obtained. Furthermore, the reliability of the program has been tested by deleting sets of chemical shifts randomly from the master_list created for the test proteins.