A probabilistic model for mining labeled ordered trees: capturing patterns in carbohydrate sugar chains
A probabilistic model for mining labeled ordered trees: capturing patterns in carbohydrate sugar chains
复制标题
挖掘标记有序树的概率模型:捕获碳水化合物糖链中的模式
DOI:
10.1109/tkde.2005.117
复制
发表时间:
2005
影响因子:
8.9
通讯作者:
Hiroshi Mamitsuka
中科院分区:
文献类型:
--
作者:
Nobuhisa Ueda;Kiyoko F. Aoki;Atsuko Yamaguchi;T. Akutsu;Hiroshi Mamitsuka
Glycans, or carbohydrate sugar chains, which play a number of important roles in the development and functioning of multicellular organisms, can be regarded as labeled ordered trees. A recent increase in the documentation of glycan structures, especially in the form of database curation, has made mining glycans important for the understanding of living cells. We propose a probabilistic model for mining labeled ordered trees, and we further present an efficient learning algorithm for this model, based on an EM algorithm. The time and space complexities of this algorithm are rather favorable, falling within the practical limits set by a variety of existing probabilistic models, including stochastic context-free grammars. Experimental results have shown that, in a supervised problem setting, the proposed method outperformed five other competing methods by a statistically significant factor in all cases. We further applied the proposed method to aligning multiple glycan trees, and we detected biologically significant common subtrees in these alignments where the trees are automatically classified into subtypes already known in glycobiology.