Regrouping of pattern clusters to reveal characteristics of distinct classes and related classes

Regrouping of pattern clusters to reveal characteristics of distinct classes and related classes
复制标题

重组模式簇以揭示不同类和相关类的特征

DOI:
--
复制
发表时间:
2013
期刊:
IEEE International Conference on Bioinformatics and Biomedicine
影响因子:
--
通讯作者:
A. Wong
A. Wong
中科院分区:
--
文献类型:
--
作者:
Pei;E. Lee;A. Wong

文献摘要

参考文献

被引文献

相似文献

发现氨基酸的蛋白质模式及其生化特性对于揭示潜在的生物物理模型非常重要。由此,引入了模式聚类,以便将发现的蛋白质模式与蛋白质局部区域的分类类别相关联。本文基于我们之前的工作,提出了一种合成和重新分组模式簇的算法,最大化它们的可分离性,以揭示蛋白质局部区域的类别特征。为了评估模式聚类和重组模式聚类结果,我们引入了三种评估度量:F-度量、类熵度量和属性熵度量。为了验证我们提出的算法,对合成数据、氨基酸属性的蛋白质家族和化学性质属性进行了实验。实验结果表明:a)重新分组模式簇的结果在类分离方面比仅使用模式聚类更准确; b) 重新分组后的聚类比仅使用模式聚类更明显地可分离; c) 发现两种类型的模式簇,一种属于不同的类别,另一种与两个或多个相关类别相关; d) 在包含模式簇中的模式的数据子空间中清楚地揭示了类特征。具有化学性质的数据集表明,随着考虑到不同氨基酸所共有的更多共同性质,无监督技术可以揭示固有类别中的共同化学属性。
Discovering protein patterns for amino acids and their biochemical properties is important for revealing the underlying biophysical models. From this, pattern clustering was introduced in order to relate the discovered protein patterns to taxonomic classes in a localized region of a protein. This paper proposes an algorithm to synthesize and re-group pattern clusters, maximizing their separability in order to reveal class characteristics of the localized region of the protein based on our previous work. To evaluate the pattern clustering and regrouping pattern clusters results, we introduce three evaluation measures: F-measure, class entropy measure, and attribute entropy measure. To validate our proposed algorithm, experiments are run on synthetic data, protein family for amino acid attributes, and chemical property attributes. The experimental results show that: a) the result for regrouping pattern clusters is more accurate in class separation than only using pattern clustering; b) The clusters after regrouping are more distinctly separable with each other than only using pattern clustering; c) two types of pattern clusters are found, with one pertaining to distinct classes and the other associating with two or more related classes; and d) class characteristics are clearly revealed in the data subspace containing the patterns in the pattern clusters. The datasets with chemical properties show that unsupervised techniques can reveal common chemical attributes in the inherent classes as more of the common properties shared by different amino acids are taken into account.
DOI: 10.1093/nar/gkh169
发表时间: 2004-01-01
影响因子: 14.9
作者:
Frith, MC;Hansen, U;Weng, ZP
通讯作者: Weng, ZP