A structural homology approach for computational protein design with flexible backbone

A structural homology approach for computational protein design with flexible backbone
复制标题

具有灵活骨架的计算蛋白质设计的结构同源方法

DOI:
10.1093/bioinformatics/bty975
复制
发表时间:
2019
期刊:
影响因子:
5.8
通讯作者:
Barbe Sophie
Barbe Sophie
中科院分区:
生物学3区
文献类型:
--
作者:
Simoncini David;Zhang Kam Y J;Schiex Thomas;Barbe Sophie

文献摘要

相似文献

基于结构的计算蛋白质设计(CPD)在推进蛋白质工程领域中起着至关重要的作用。使用全原子能量函数,CPD试图识别折叠成目标结构并最终执行所需功能的氨基酸序列。能量功能仍然是不完善的,并注入相关信息,从已知的结构在设计过程中,应导致改进designs.ResultsWe介绍阴影,数据驱动的CPD方法,利用当地的结构环境中已知的蛋白质结构与能量一起指导序列设计,同时采样的侧链和骨架构象,以适应突变。Shades(Structural Homology Algorithm for protein DESign)是基于非连续接触氨基酸残基基序的定制库。我们已经在从不同蛋白质家族中选择的40种蛋白质的公共基准上测试了Shades。当排除同源蛋白质时,与靶蛋白的PFAM蛋白质家族相比,Shades实现了30%的蛋白质序列回收率和平均46%的蛋白质序列相似性。当添加同源结构时,野生型序列恢复率达到93%。可用性和实施Shades源代码可通过https://bitbucket.org/satsumaimo/shades作为Rosetta 3.8的补丁,带有精选的蛋白质结构数据库和ITEM库创建软件。补充信息补充数据可在Bioinformatics online获得。
MotivationStructure-based Computational Protein design (CPD) plays a critical role in advancing the field of protein engineering. Using an all-atom energy function, CPD tries to identify amino acid sequences that fold into a target structure and ultimately perform a desired function. Energy functions remain however imperfect and injecting relevant information from known structures in the design process should lead to improved designs.ResultsWe introduce Shades, a data-driven CPD method that exploits local structural environments in known protein structures together with energy to guide sequence design, while sampling side-chain and backbone conformations to accommodate mutations. Shades (Structural Homology Algorithm for protein DESign), is based on customized libraries of non-contiguous in-contact amino acid residue motifs. We have tested Shades on a public benchmark of 40 proteins selected from different protein families. When excluding homologous proteins, Shades achieved a protein sequence recovery of 30% and a protein sequence similarity of 46% on average, compared with the PFAM protein family of the target protein. When homologous structures were added, the wild-type sequence recovery rate achieved 93%.Availability and implementationShades source code is available athttps://bitbucket.org/satsumaimo/shadesas a patch for Rosetta 3.8 with a curated protein structure database and ITEM library creation software.Supplementary informationSupplementary data are available atBioinformaticsonline.