GlycoCT - a unifying sequence format for carbohydrates

GlycoCT - a unifying sequence format for carbohydrates
复制标题

DOI:
10.1016/j.carres.2008.03.011
复制
发表时间:
2008-08-11
影响因子:
3.1
通讯作者:
Von der Lieth, C. -W.
Von der Lieth, C. -W.
中科院分区:
化学3区
文献类型:
--
作者:
Herget, S.;Ranzinger, R.;Von der Lieth, C. -W.

文献摘要

被引文献

相似文献

作为EUROCarbDB项目(www.eurocarbdb.org)的一部分,我们仔细分析了所有现有碳水化合物序列格式的编码能力和公开可用结构数据库的内容。我们发现,没有一个现有的结构编码模式能够处理所有分类来源中实验得出的结构碳水化合物序列数据的全部复杂性。这一差距促使我们定义了一种复杂碳水化合物的编码方案,命名为GlycoCT,以克服目前的限制。这种新格式基于连接表方法而不是线性编码方案来描述碳水化合物序列,使用受控词汇表来命名单糖,采用IUPAC规则来生成一致的、机器可读的命名法。该格式使用块概念来描述碳水化合物序列中经常出现的特殊特征,如重复单元。它有两种变体,一种是浓缩形式,另一种是更详细的XML语法。排序规则确保压缩表单的唯一性,从而使其适合作为依赖惟一标识符的数据库应用程序的直接主键。糖ct包含糖组学中数字编码模式的异构景观的能力,因此在糖生物信息学中向统一和广泛接受的序列格式迈进了一步。(C) 2008年Elsevier Ltd.出版
As part of the EUROCarbDB project (www.eurocarbdb.org) we have carefully analyzed the encoding capabilities of all existing carbohydrate sequence formats and the content of publically available structure databases. We have found that none of the existing structural encoding schemata are capable of coping with the full complexity to be expected for experimentally derived structural carbohydrate sequence data across all taxonomic sources. This gap motivated us to define an encoding scheme for complex carbohydrates, named GlycoCT, to overcome the current limitations. This new format is based on a connection table approach, instead of a linear encoding scheme, to describe the carbohydrate sequences, with a controlled vocabulary to name monosaccharides, adopting IUPAC rules to generate a consistent, machine-readable nomenclature. The format uses a block concept to describe frequently occurring special features of carbohydrate sequences like repeating units. It exists in two variants, a condensed form and a more verbose XML syntax. Sorting rules assure the uniqueness of the condensed form, thus making it suitable as a direct primary key for database applications, which rely on unique identifiers. GlycoCT encompasses the capabilities of the heterogeneous landscape of digital encoding schemata in glycomics and is thus a step forward on the way to a unified and broadly accepted sequence format in glycobioinformatics. (C) 2008 Published by Elsevier Ltd.