A top-down approach to classify enzyme functional classes and sub-classes using random forest

A top-down approach to classify enzyme functional classes and sub-classes using random forest
复制标题

DOI:
10.1186/1687-4153-2012
复制
发表时间:
2012-12-01
期刊:
EURASIP JOURNAL ON BIOINFORMATICS AND SYSTEMS BIOLOGY
影响因子:
--
通讯作者:
Choudhary, Alok
Choudhary, Alok
中科院分区:
其他
文献类型:
--
作者:
Kumar, Chetan;Choudhary, Alok

文献摘要

被引文献

相似文献

测序技术的进步见证了新发现的酶数量的指数级增长。酶是催化生物化学反应的蛋白质,在代谢途径中起重要作用。通常,这些酶的功能是通过实验来确定的,这可能是耗时和昂贵的。因此,需要一种计算方法,可以区分蛋白质酶序列与非酶序列,并可靠地预测前者的功能。为了解决这个问题,已经提出了基于它们的序列和结构相似性来聚类酶的方法。但是,众所周知,这些方法对于执行相同功能但序列和结构不同的蛋白质来说是失败的。在这篇文章中,我们提出了一个有监督的机器学习模型来预测基于一组73个序列衍生特征的酶的功能类和子类。功能类别由国际生物化学和分子生物学联合会定义。使用一种高效的数据挖掘算法随机森林,我们构建了一个自顶向下的三层模型,其中第一层分类的查询蛋白质序列作为酶或非酶,第二层预测的主要功能类和底层进一步预测的子功能类。该模型报告了第一级的总体分类准确率为94.87%,第二级为87.7%,底层为84.25%。我们的研究结果与现有的方法相比,在许多情况下,报告更好的性能。使用特征选择方法,我们已经显示了一些顶级属性的生物相关性。
Advancements in sequencing technologies have witnessed an exponential rise in the number of newly found enzymes. Enzymes are proteins that catalyze bio-chemical reactions and play an important role in metabolic pathways. Commonly, function of such enzymes is determined by experiments that can be time consuming and costly. Hence, a need for a computing method is felt that can distinguish protein enzyme sequences from those of non-enzymes and reliably predict the function of the former. To address this problem, approaches that cluster enzymes based on their sequence and structural similarity have been presented. But, these approaches are known to fail for proteins that perform the same function and are dissimilar in their sequence and structure. In this article, we present a supervised machine learning model to predict the function class and sub-class of enzymes based on a set of 73 sequence-derived features. The functional classes are as defined by International Union of Biochemistry and Molecular Biology. Using an efficient data mining algorithm called random forest, we construct a top-down three layer model where the top layer classifies a query protein sequence as an enzyme or non-enzyme, the second layer predicts the main function class and bottom layer further predicts the sub-function class. The model reported overall classification accuracy of 94.87% for the first level, 87.7% for the second, and 84.25% for the bottom level. Our results compare very well with existing methods, and in many cases report better performance. Using feature selection methods, we have shown the biological relevance of a few of the top rank attributes.