Data mining the protein data bank:: automatic detection and assignment of carbohydrate structures

Data mining the protein data bank:: automatic detection and assignment of carbohydrate structures
复制标题

DOI:
10.1016/j.carres.2003.09.038
复制
发表时间:
2004-04-02
影响因子:
3.1
通讯作者:
von der Lieth, CW
von der Lieth, CW
中科院分区:
化学3区
文献类型:
--
作者:
Lütteke, T;Frank, M;von der Lieth, CW

文献摘要

被引文献

相似文献

了解聚糖的3D结构是完全理解糖蛋白参与的生物过程的先决条件。然而,由于缺乏标准化的命名法,碳水化合物很难在蛋白质数据库(PDB)中找到。使用一种只需要元素类型和原子坐标就能检测碳水化合物结构的算法。我们能够检测到1663个条目,其包含总共5647个碳水化合物链。大多数链被发现是N-糖苷结合的。非共价结合的配体也很常见,而O-聚糖占少数。所有含碳水化合物的PDB条目中约30%包含一个或多个错误。PDB条目中碳水化合物结构的自动分配将改善糖生物学资源与基因组学和蛋白质组学数据收集的交联,这将是即将到来的糖组学项目的一个重要问题。通过帮助检测错误的注释和结构,该算法还可以帮助提高数据库质量。(C)2003 Elsevier Ltd.保留所有权利。
Knowledge of the 3D structure of glycans is a prerequisite for a complete understanding of the biological processes glycoproteins are involved in. However, due to a lack of standardised nomenclature, carbohydrate compounds are difficult to locate within the Protein Data Bank (PDB). Using an algorithm that detects carbohydrate structures only requiring element types and atom coordinates. we were able to detect 1663 entries containing a total of 5647 carbohydrate chains. The majority of chains are found to be N-glycosidically bound. Noncovalently bound ligands are also frequent, while O-glycans form a minority. About 30% of all carbohydrate containing PDB entries comprise one or several errors. The automatic assignment of carbohydrate structures in PDB entries will improve the cross-linking of glycobiology resources with genomic and proteomic data collections, which will be an important issue of the upcoming glycomics projects. By aiding in detection of erroneous annotations and structures, the algorithm might also help to increase database quality. (C) 2003 Elsevier Ltd. All rights reserved.