Transmembrane proteins in the Protein Data Bank:: identification and classification

Transmembrane proteins in the Protein Data Bank:: identification and classification
复制标题

DOI:
10.1093/bioinformatics/bth340
复制
发表时间:
2004-11-22
期刊:
影响因子:
5.8
通讯作者:
Simon, I
Simon, I
中科院分区:
生物学3区
文献类型:
--
作者:
Tusnády, GE;Dosztányi, Z;Simon, I

文献摘要

被引文献

相似文献

动机:整合膜蛋白在活细胞中发挥重要作用。尽管这些蛋白质估计与基因组规模的蛋白质的 25% 相似,但由于实验技术的困难,蛋白质数据库 (PDB) 仅包含几百个膜蛋白。然而,跨膜蛋白在结构数据库中的存在是完全看不见的,因为这些条目的注释相当差。即使蛋白质被鉴定为跨膜蛋白,PDB中也不会指示脂双层的可能位置,因为这些蛋白质在没有天然脂双层的情况下结晶,并且目前没有公开的方法可以使用膜蛋白的原子坐标来检测可能的膜平面。结果:在这里,我们提出了一种新的几何方法,仅使用结构信息来区分跨膜蛋白和球状蛋白,并定位脂双层最可能的位置。给出了自动算法(TMDET)来确定相对于原子坐标位置的膜平面,以及辨别功能,即使在低分辨率或不完整结构(例如大型多链复合物的片段或部分)的情况下,也能够分离跨膜和球状蛋白质。该方法可用于正确注释含有跨膜片段的蛋白质结构,并为包含所有已知跨膜蛋白质和片段(PDB_TM)结构的最新数据库铺平道路,该数据库可以自动更新。对于构建纯粹的球状蛋白数据库来说,该算法同样重要。
Motivation: Integral membrane proteins play important roles in living cells. Although these proteins are estimated to constitute similar to25% of proteins at a genomic scale, the Protein Data Bank (PDB) contains only a few hundred membrane proteins due to the difficulties with experimental techniques. The presence of transmembrane proteins in the structure data bank, however, is quite invisible, as the annotation of these entries is rather poor. Even if a protein is identified as a transmembrane one, the possible location of the lipid bilayer is not indicated in the PDB because these proteins are crystallized without their natural lipid bilayer, and currently no method is publicly available to detect the possible membrane plane using the atomic coordinates of membrane proteins.Results: Here, we present a new geometrical approach to distinguish between transmembrane and globular proteins using structural information only and to locate the most likely position of the lipid bilayer. An automated algorithm (TMDET) is given to determine the membrane planes relative to the position of atomic coordinates, together with a discrimination function which is able to separate transmembrane and globular proteins even in cases of low resolution or incomplete structures such as fragments or parts of large multi chain complexes. This method can be used for the proper annotation of protein structures containing transmembrane segments and paves the way to an up-to-date database containing the structure of all known transmembrane proteins and fragments (PDB_TM) which can be automatically updated. The algorithm is equally important for the purpose of constructing databases purely of globular proteins.