Enzyme Substrate Prediction from Three-Dimensional Feature Representations Using Space-Filling Curves

Enzyme Substrate Prediction from Three-Dimensional Feature Representations Using Space-Filling Curves
复制标题

DOI:
10.1021/acs.jcim.3c00005
复制
发表时间:
2023-02
影响因子:
5.6
通讯作者:
Dmitrij Rappoport;A. Jinich
Dmitrij Rappoport;A. Jinich
中科院分区:
化学2区
文献类型:
--
作者:
Dmitrij Rappoport;A. Jinich

文献摘要

相似文献

紧凑和可解释的结构特征表示是准确预测蛋白质性质和功能所必需的。在这项工作中,我们基于空间填充曲线(sfc)构建和评估蛋白质结构的三维特征表示。我们将重点放在酶底物预测问题上,使用两个普遍存在的酶家族作为案例研究:短链脱氢酶/还原酶(sdr)和s -腺苷蛋氨酸依赖性甲基转移酶(SAM-MTases)。希尔伯特曲线和莫顿曲线等空间填充曲线产生了从离散的三维到一维表示的可逆映射,从而有助于以系统无关的方式编码三维分子结构,并且只有几个可调参数。利用使用AlphaFold2生成的sdr和SAM-MTases的三维结构,我们评估了基于sfc的特征表示在酶分类任务的新基准数据库中的预测性能,包括它们的辅因子和底物选择性。梯度增强树分类器对分类任务的二值预测精度为0.77-0.91,曲线下面积(AUC)特征为0.83-0.92。我们研究了氨基酸编码、空间取向和基于sfc编码的(少数)参数对预测准确性的影响。我们的研究结果表明,基于几何的方法(如sfc)有望生成蛋白质结构表示,并与现有的蛋白质特征表示(如进化尺度建模(ESM)序列嵌入)相补充。
Compact and interpretable structural feature representations are required for accurately predicting properties and function of proteins. In this work, we construct and evaluate three-dimensional feature representations of protein structures based on space-filling curves (SFCs). We focus on the problem of enzyme substrate prediction, using two ubiquitous enzyme families as case studies: the short-chain dehydrogenase/reductases (SDRs) and the S-adenosylmethionine-dependent methyltransferases (SAM-MTases). Space-filling curves such as the Hilbert curve and the Morton curve generate a reversible mapping from discretized three-dimensional to one-dimensional representations and thus help to encode three-dimensional molecular structures in a system-independent way and with only a few adjustable parameters. Using three-dimensional structures of SDRs and SAM-MTases generated using AlphaFold2, we assess the performance of the SFC-based feature representations in predictions on a new benchmark database of enzyme classification tasks including their cofactor and substrate selectivity. Gradient-boosted tree classifiers yield binary prediction accuracy of 0.77-0.91 and area under curve (AUC) characteristics of 0.83-0.92 for the classification tasks. We investigate the effects of amino acid encoding, spatial orientation, and (the few) parameters of SFC-based encodings on the accuracy of the predictions. Our results suggest that geometry-based approaches such as SFCs are promising for generating protein structural representations and are complementary to the existing protein feature representations such as evolutionary scale modeling (ESM) sequence embeddings.