Global analysis of protein folding using massively parallel design, synthesis, and testing

Global analysis of protein folding using massively parallel design, synthesis, and testing
复制标题

DOI:
10.1126/science.aan0693
复制
发表时间:
2017-07-14
期刊:
影响因子:
56.9
通讯作者:
Baker, David
Baker, David
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Rocklin, Gabriel J.;Chidyausiku, Tamuka M.;Baker, David

文献摘要

被引文献

相似文献

蛋白质折叠成独特的天然结构,由数以千计的弱相互作用稳定下来,共同克服折叠的熵成本。尽管这些力量在数以千计的已知蛋白质结构中被“编码”,但“解码”它们是具有挑战性的,因为天然蛋白质的复杂性是为了功能而不是稳定性而进化的。我们结合了计算蛋白质设计、下一代基因合成和高通量蛋白酶敏感性分析来测量15,000多个从头设计的迷你蛋白质、1000个天然蛋白质、10,000个点突变和30,000个阴性对照序列的折叠和稳定性。这一分析在四个基本折叠中确定了2500多个稳定的设计蛋白质--这个数字足以使我们能够系统地研究序列如何决定未知蛋白质空间的折叠和稳定性。设计和实验之间的迭代将设计成功率从6%提高到47%,产生了与自然中发现的稳定蛋白质不同的蛋白质,这些蛋白质最初设计不成功,并且随着设计变得越来越优化,显示出对稳定性的微妙贡献。我们的方法实现了计算和实验之间紧密反馈循环的长期目标,并有可能将计算蛋白质设计转变为数据驱动的科学。
Proteins fold into unique native structures stabilized by thousands of weak interactions that collectively overcome the entropic cost of folding. Although these forces are "encoded" in the thousands of known protein structures, "decoding" them is challenging because of the complexity of natural proteins that have evolved for function, not stability. We combined computational protein design, next-generation gene synthesis, and a high-throughput protease susceptibility assay to measure folding and stability for more than 15,000 de novo designed miniproteins, 1000 natural proteins, 10,000 point mutants, and 30,000 negative control sequences. This analysis identified more than 2500 stable designed proteins in four basic folds-a number sufficient to enable us to systematically examine how sequence determines folding and stability in uncharted protein space. Iteration between design and experiment increased the design success rate from 6% to 47%, produced stable proteins unlike those found in nature for topologies where design was initially unsuccessful, and revealed subtle contributions to stability as designs became increasingly optimized. Our approach achieves the long-standing goal of a tight feedback cycle between computation and experiment and has the potential to transform computational protein design into a data-driven science.