ProfilePSTMM: capturing tree-structure motifs in carbohydrate sugar chains

ProfilePSTMM: capturing tree-structure motifs in carbohydrate sugar chains
复制标题

DOI:
10.1093/bioinformatics/btl244
复制
发表时间:
2006-07-01
期刊:
影响因子:
5.8
通讯作者:
Kanehisa, Minoru
Kanehisa, Minoru
中科院分区:
生物学3区
文献类型:
--
作者:
Aoki-Kinoshita, Kiyoko F.;Ueda, Nobuhisa;Kanehisa, Minoru

文献摘要

被引文献

相似文献

动机:碳水化合物糖链或聚糖被认为是继DNA和蛋白质之后的第三大类生物分子。它们由分支单糖组成,从单个单糖开始。它们对多细胞生物的发育和功能至关重要,因为它们被各种蛋白质识别,使它们能够执行特定的功能。我们的动机是利用信息学技术从现有数据中研究这种识别机制。之前,我们介绍了一个概率依赖于兄弟姐妹的树马尔可夫模型(PSTMM),我们证明了它可以有效地训练依赖于兄弟姐妹的树结构,并返回最可能的状态路径。然而,它有一些局限性,兄弟姐妹之间的额外依赖会导致过拟合问题。从训练模型中检索模式还涉及从最可能的状态路径中手动提取模式。因此,我们引入了一个避免这些问题的profilePSTMM模型,它结合了不同类型的状态转换的新概念,以不同的方式处理父子和兄弟依赖关系。结果:新算法的效率更高,提取模式更容易。我们在合成(受控)数据和来自KEGG glycan数据库的聚糖数据上测试了profilePSTMM模型。此外,我们在已知可识别并以各种结合亲和力与蛋白质结合的聚糖上进行了测试,我们表明我们的结果与文献中发表的结果相关。
Motivation: Carbohydrate sugar chains, or glycans, are considered the third major class of biomolecules after DNA and proteins. They consist of branching monosaccharides, starting from a single monosaccharide. They are extremely vital to the development and functioning of multicellular organisms because they are recognized by various proteins to allow them to perform specific functions. Our motivation is to study this recognition mechanism using informatics techniques from the data available. Previously, we introduced a probabilistic sibling-dependent tree Markov model (PSTMM), which we showed could be efficiently trained on sibling-dependent tree structures and return the most likely state paths. However, it had some limitations in that the extra dependency between siblings caused overfitting problems. The retrieval of the patterns from the trained model also involved manually extracting the patterns from the most likely state paths. Thus we introduce a profilePSTMM model which avoids these problems, incorporating a novel concept of different types of state transitions to handle parent-child and sibling dependencies differently.Results: Our new algorithms are more efficient and able to extract the patterns more easily. We tested the profilePSTMM model on both synthetic (controlled) data as well as glycan data from the KEGG GLYCAN database. Additionally, we tested it on glycans which are known to be recognized and bound to proteins at various binding affinities, and we show that our results correlate with results published in the literature.