A probabilistic model for mining labeled ordered trees: capturing patterns in carbohydrate sugar chains

A probabilistic model for mining labeled ordered trees: capturing patterns in carbohydrate sugar chains
复制标题

挖掘标记有序树的概率模型:捕获碳水化合物糖链中的模式

DOI:
10.1109/tkde.2005.117
复制
发表时间:
2005
影响因子:
8.9
通讯作者:
Hiroshi Mamitsuka
Hiroshi Mamitsuka
中科院分区:
计算机科学2区
文献类型:
--
作者:
Nobuhisa Ueda;Kiyoko F. Aoki;Atsuko Yamaguchi;T. Akutsu;Hiroshi Mamitsuka

文献摘要

被引文献

相似文献

聚糖或碳水化合物糖链在多细胞生物的发育和功能中起着许多重要作用,可以被视为标记有序树。最近,聚糖结构的文档越来越多,特别是以数据库管理的形式,这使得挖掘聚糖对于理解活细胞非常重要。我们提出了一个概率模型挖掘标记有序树,我们进一步提出了一个有效的学习算法,这个模型的基础上EM算法。该算法的时间和空间复杂度是相当有利的,落在各种现有的概率模型,包括随机上下文无关文法的实际限制。实验结果表明,在一个监督的问题设置,该方法优于其他五个竞争的方法在所有情况下的统计显着因素。我们进一步将所提出的方法应用于比对多个聚糖树,并且我们在这些比对中检测到生物学上重要的共同子树,其中树被自动分类为糖生物学中已知的亚型。
Glycans, or carbohydrate sugar chains, which play a number of important roles in the development and functioning of multicellular organisms, can be regarded as labeled ordered trees. A recent increase in the documentation of glycan structures, especially in the form of database curation, has made mining glycans important for the understanding of living cells. We propose a probabilistic model for mining labeled ordered trees, and we further present an efficient learning algorithm for this model, based on an EM algorithm. The time and space complexities of this algorithm are rather favorable, falling within the practical limits set by a variety of existing probabilistic models, including stochastic context-free grammars. Experimental results have shown that, in a supervised problem setting, the proposed method outperformed five other competing methods by a statistically significant factor in all cases. We further applied the proposed method to aligning multiple glycan trees, and we detected biologically significant common subtrees in these alignments where the trees are automatically classified into subtypes already known in glycobiology.