Sequence alignment using machine learning for accurate template-based protein structure prediction

Sequence alignment using machine learning for accurate template-based protein structure prediction
复制标题

DOI:
10.1093/bioinformatics/btz483
复制
发表时间:
2020-01-01
期刊:
影响因子:
5.8
通讯作者:
Ishida, Takashi
Ishida, Takashi
中科院分区:
生物学3区
文献类型:
--
作者:
Makigaki, Shuichiro;Ishida, Takashi

文献摘要

被引文献

相似文献

基于动机模板的建模,即通过使用同源蛋白质结构预测蛋白质的三级结构的过程,如果能找到好的模板是有用的。虽然现代的同源检测方法可以找到高灵敏度的远程同源序列,但基于同源检测的比对生成的模板模型的准确率往往低于理想比对生成的模型的准确性。提出的方法使用已知同源物的结构比对来训练机器学习模型。使用机器学习直接预测序列比对是困难的。因此,当计算序列比对时,该方法不是固定的替换矩阵,而是根据训练的模型动态地预测替换分数。我们通过仔细拆分训练和测试数据集,并将预测结构的精度与最先进的方法进行比较,来评估我们的方法。我们的方法比通过其他方法获得的比对产生的模型更准确的三级结构模型。可用性和implementationhttps://github.com/shuichiro-makigaki/exmachina.Supplementary信息补充数据可在生物信息学在线上获得。
Motivation Template-based modeling, the process of predicting the tertiary structure of a protein by using homologous protein structures, is useful if good templates can be found. Although modern homology detection methods can find remote homologs with high sensitivity, the accuracy of template-based models generated from homology-detection-based alignments is often lower than that from ideal alignments.Results In this study, we propose a new method that generates pairwise sequence alignments for more accurate template-based modeling. The proposed method trains a machine learning model using the structural alignment of known homologs. It is difficult to directly predict sequence alignments using machine learning. Thus, when calculating sequence alignments, instead of a fixed substitution matrix, this method dynamically predicts a substitution score from the trained model. We evaluate our method by carefully splitting the training and test datasets and comparing the predicted structure's accuracy with that of state-of-the-art methods. Our method generates more accurate tertiary structure models than those produced from alignments obtained by other methods.Availability and implementationhttps://github.com/shuichiro-makigaki/exmachina.Supplementary informationSupplementary data are available at Bioinformatics online.