Caretta - A multiple protein structure alignment and feature extraction suite

Caretta - A multiple protein structure alignment and feature extraction suite
复制标题

DOI:
10.1016/j.csbj.2020.03.011
复制
发表时间:
2020-01-01
影响因子:
6
通讯作者:
van Dijk, Aalt D. J.
van Dijk, Aalt D. J.
中科院分区:
生物学2区
文献类型:
--
作者:
Akdel, Mehmet;Durairaj, Janani;van Dijk, Aalt D. J.

文献摘要

被引文献

相似文献

目前可用的大量蛋白质结构为蛋白质机器学习提供了令人兴奋的机会,旨在预测和理解功能特性。特别是,结合同源建模,现在不仅可以使用序列特征作为机器学习的输入,还可以使用结构特征。然而,为了做到这一点,强大的多结构比对是必不可少的。在这里,我们介绍了Caretta,一个多结构比对套件,用于同源但顺序不同的蛋白质家族,它始终返回准确的比对,覆盖率高于目前最先进的工具。Caretta可作为GUI和命令行应用程序使用,并为给定的一组输入结构输出对齐的结构特征矩阵,该矩阵可以很容易地用于监督或无监督机器学习的下游步骤。我们将展示Caretta在两个基准数据集上的表现,并给出Caretta在预测细胞周期蛋白依赖性激酶构象状态中的一个示例应用。(C)2020年,任作家。由Elsevier B.V.代表计算和结构生物技术研究网络出版。
The vast number of protein structures currently available opens exciting opportunities for machine learning on proteins, aimed at predicting and understanding functional properties. In particular, in combination with homology modelling, it is now possible to not only use sequence features as input for machine learning, but also structure features. However, in order to do so, robust multiple structure alignments are imperative.Here we present Caretta, a multiple structure alignment suite meant for homologous but sequentially divergent protein families which consistently returns accurate alignments with a higher coverage than current state-of-the-art tools. Caretta is available as a GUI and command-line application and additionally outputs an aligned structure feature matrix for a given set of input structures, which can readily be used in downstream steps for supervised or unsupervised machine learning. We show Caretta's performance on two benchmark datasets, and present an example application of Caretta in predicting the conformational state of cyclin-dependent kinases. (C) 2020 The Authors. Published by Elsevier B.V. on behalf of Research Network of Computational and Structural Biotechnology.