Engineering proteinase K using machine learning and synthetic genes.

Engineering proteinase K using machine learning and synthetic genes.
复制标题

使用机器学习和合成基因进行工程蛋白酶K。

DOI:
10.1186/1472-6750-7-16
复制
发表时间:
2007-03-26
期刊:
影响因子:
3.5
通讯作者:
Minshull, Jeremy
Minshull, Jeremy
中科院分区:
工程技术3区
文献类型:
--
作者:
Liao, Jun;Warmuth, Manfred K.;Govindarajan, Sridhar;Ness, Jon E.;Wang, Rebecca P.;Gustafsson, Claes;Minshull, Jeremy

文献摘要

参考文献

被引文献

相似文献

通过改变蛋白质的序列来改变其功能,可以将天然蛋白质转化为有用的分子工具。目前的蛋白质工程方法受到缺乏高通量物理或计算测试的限制,这些测试可以准确预测与其最终应用相关的条件下的蛋白质活性。在这里,我们描述了一种新的蛋白质工程合成生物学方法,通过将高通量基因合成与基于机器学习的设计算法相结合,避免了这些限制。我们从同源序列比对中选择了 24 个氨基酸替换以在蛋白酶 K 中进行。然后,我们设计并合成了 59 种特定的蛋白酶 K 变体,其中包含所选取代的不同组合。首先将酶加热至 68°C 5 分钟后,测试了 59 种变体水解四肽底物的能力。使用机器学习算法分析序列和活动数据。该分析用于设计一组新的变体,预计其活性比训练集有所增加,然后进行合成和测试。通过执行两个周期的机器学习分析和变体设计,我们获得了 20 倍改进的蛋白酶 K 变体,同时仅测试了总共 95 种变体酶。为了获得显着的功能改进而必须测试的蛋白质变体的数量决定了可以进行的测试的类型。希望修改蛋白质特性以缩小肿瘤或在工业条件下催化化学反应的蛋白质工程师迄今为止被迫接受高通量替代筛选来测量蛋白质特性,他们希望这些特性与他们打算修改的功能相关。通过将必须测试的变体数量减少到 100 以下,机器学习算法使得使用更复杂和更昂贵的测试成为可能,因此只需要测量与所需应用直接相关的蛋白质特性。仅需要测试少量变体的蛋白质设计算法代表着朝着通用、资源优化的蛋白质工程过程迈出了重要一步。
Altering a protein's function by changing its sequence allows natural proteins to be converted into useful molecular tools. Current protein engineering methods are limited by a lack of high throughput physical or computational tests that can accurately predict protein activity under conditions relevant to its final application. Here we describe a new synthetic biology approach to protein engineering that avoids these limitations by combining high throughput gene synthesis with machine learning-based design algorithms. We selected 24 amino acid substitutions to make in proteinase K from alignments of homologous sequences. We then designed and synthesized 59 specific proteinase K variants containing different combinations of the selected substitutions. The 59 variants were tested for their ability to hydrolyze a tetrapeptide substrate after the enzyme was first heated to 68°C for 5 minutes. Sequence and activity data was analyzed using machine learning algorithms. This analysis was used to design a new set of variants predicted to have increased activity over the training set, that were then synthesized and tested. By performing two cycles of machine learning analysis and variant design we obtained 20-fold improved proteinase K variants while only testing a total of 95 variant enzymes. The number of protein variants that must be tested to obtain significant functional improvements determines the type of tests that can be performed. Protein engineers wishing to modify the property of a protein to shrink tumours or catalyze chemical reactions under industrial conditions have until now been forced to accept high throughput surrogate screens to measure protein properties that they hope will correlate with the functionalities that they intend to modify. By reducing the number of variants that must be tested to fewer than 100, machine learning algorithms make it possible to use more complex and expensive tests so that only protein properties that are directly relevant to the desired application need to be measured. Protein design algorithms that only require the testing of a small number of variants represent a significant step towards a generic, resource-optimized protein engineering process.
DOI: 10.1023/a:1012470815092
发表时间: 2002-01-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Demiriz, A;Bennett, KP;Shawe-Taylor, J
通讯作者: Shawe-Taylor, J
DOI: 10.1177/1087057105284334
发表时间: 2006-03-01
影响因子: --
作者:
Fang, JW;Dong, YH;Georg, GI
通讯作者: Georg, GI
DOI: 10.1016/s0022-2836(03)00357-7
发表时间: 2003-05-16
影响因子: 5.6
作者:
Govindarajan, S;Ness, JE;Gustafsson, C
通讯作者: Gustafsson, C
DOI: 10.3891/acta.chem.scand.40b-0135
发表时间: 1986-01-01
期刊: ACTA CHEMICA SCANDINAVICA SERIES B-ORGANIC CHEMISTRY AND BIOCHEMISTRY
影响因子: --
作者:
HELLBERG, S;SJOSTROM, M;WOLD, S
通讯作者: WOLD, S
DOI: 10.1080/00401706.1970.10488634
发表时间: 1970-01-01
期刊: TECHNOMETRICS
影响因子: 2.5
作者:
HOERL, AE;KENNARD, RW
通讯作者: KENNARD, RW