Enzyme mechanism prediction: a template matching problem on InterPro signature subspaces.

Enzyme mechanism prediction: a template matching problem on InterPro signature subspaces.
复制标题

DOI:
10.1186/s13104-015-1730-7
复制
发表时间:
2015-12-03
期刊:
影响因子:
1.8
通讯作者:
Mitchell JB
Mitchell JB
中科院分区:
其他
文献类型:
--
作者:
Mussa HY;De Ferrari L;Mitchell JB

文献摘要

相似文献

我们最近报道,通过采用一种简单的模式识别方法,人们可以高精度地预测酶的化学机制:使用 k = 1 (k1NN) 的 k 最近邻规则和 321 InterPro 序列签名作为酶特征。众所周知,最近邻规则对训练数据中的错误高度敏感,特别是当可用训练数据集较小时。我们之前的研究就是这种情况,其中我们的数据集包含 248 种酶,根据 MA​​CiE 数据库中的 71 种酶机制标签进行注释。在当前的研究中,我们仔细地重新分析了我们的数据集和预测结果,以“解释”为什么高方差 k1NN 规则表现出如此出色的分类性能。我们发现该数据集中具有不同化学机制标签的酶位于由所选 321 个特征定义的特征空间中几乎不重叠的子空间中。这些特征包含准确分类酶机制所需的适当信息,使我们的分类问题成为基本的查找练习。这一观察结果与我们报告的低错误分类率相吻合。我们的结果为“异常”提供了解释——一种基本的最近邻算法,尽管特征空间很大且稀疏,但对酶机制表现出出色的预测性能。我们的结果也与我们报道的另一项发现非常吻合,即 InterPro 签名对于准确预测酶机制至关重要。我们还提出了简单的规则,使人们能够归纳预测一种新型酶是否具有我们 71 种预定义机制中的任何一种。
We recently reported that one may be able to predict with high accuracy the chemical mechanism of an enzyme by employing a simple pattern recognition approach: a k Nearest Neighbour rule with k = 1 (k1NN) and 321 InterPro sequence signatures as enzyme features. The nearest-neighbour rule is known to be highly sensitive to errors in the training data, in particular when the available training dataset is small. This was the case in our previous study, in which our dataset comprised 248 enzymes annotated against 71 enzymatic mechanism labels from the MACiE database. In the current study, we have carefully re-analysed our dataset and prediction results to “explain” why a high variance k1NN rule exhibited such remarkable classification performance. We find that enzymes with different chemical mechanism labels in this dataset reside in barely overlapping subspaces in the feature space defined by the 321 features selected. These features contain the appropriate information needed to accurately classify the enzymatic mechanisms, rendering our classification problem a basic look-up exercise. This observation dovetails with the low misclassification rate we reported. Our results provide explanations for the “anomaly”—a basic nearest-neighbour algorithm exhibiting remarkable prediction performance for enzymatic mechanism despite the fact that the feature space was large and sparse. Our results also dovetail well with another finding we reported, namely that InterPro signatures are critical for accurate prediction of enzyme mechanism. We also suggest simple rules that might enable one to inductively predict whether a novel enzyme possesses any of our 71 predefined mechanisms.