Connecting Unexplored Protein Crystal Structures to Enzymatic Function

Connecting Unexplored Protein Crystal Structures to Enzymatic Function
复制标题

DOI:
10.1002/cctc.201200544
复制
发表时间:
2013-01-01
期刊:
影响因子:
4.5
通讯作者:
Hoehne, Matthias
Hoehne, Matthias
中科院分区:
化学3区
文献类型:
--
作者:
Steffen-Munsberg, Fabian;Vickers, Clare;Hoehne, Matthias

文献摘要

被引文献

相似文献

生物催化已成为替代传统化学合成制备精细化学品的重要方法,最近利用密集蛋白质工程产生的(R)-胺转氨酶(ATA)生物催化生产药物西格列汀就证明了这一点与已有的过渡金属催化生产西格列汀的方法相比,该方法在光学纯度、产率和废物产生方面都具有优越性工艺开发、鉴定和优化适当的酶是获得成功和有效的酶催化工艺的重要要求。自然界有一个丰富的资源库,从中可以找到合适的酶作为起点。与筛选菌株收集等经典方法相比,宏基因组学领域的现代发展为筛选不可培养生物多样性中的活动提供了巨大的潜力一个主要的进步是下一代测序技术的发展,这导致了公共数据库中存储的遗传信息的大量增加;目前有2000万个蛋白质序列有待探索。这一发展也导致了未知(或错误注释)功能的蛋白质序列的积累,因此这一丰富资源的很大一部分不能可靠地使用。我们最近利用这些信息,开发了一种硅酶发现策略,并在公共数据库中存储的> 5000序列中鉴定了17个新的和独特的(R)选择性at0,尽管在我们的工作之前没有在文献中描述基因或蛋白质序列。此外,我们可以证明这些(R)-ATAs在合成上是有用的,正如最近用17种酶中的7种不对称合成一组12种手性胺所证明的那样。现代蛋白质工程方法依赖于通过x射线晶体学获得的高质量结构信息,这些信息是集中、定向进化以改善酶性能的基础。与基因组学领域的发展类似,在过去几年中,用于结晶和晶体结构测定的自动化高通量方法得到了发展,这导致了解决的蛋白质结构数量的迅速增加然而,对于许多酶来说,没有可用的结构,例如,(R)和(S)选择性ATAs,在过去的几年里非常流行因此,结构信息对于理解实验结果和指导蛋白质工程是非常必要的有趣的是,与上述发现序列的趋势类似:越来越多的酶的晶体结构从未在底物范围、对映体选择性或反应特异性方面得到表征,它们的生理功能也常常是未知的。与序列数据库中注释蛋白缺乏实验验证类似,功能未知的结晶蛋白很少被表征因此,许多结构仍未被探索:例如,在布鲁克海文蛋白质数据库(PDB)中发现的104种不同转氨酶的晶体结构中,46种结构没有提供结构描述或蛋白质表征数据的文献引用。将酶功能与PDB中未开发的蛋白质联系起来可以提供多种兴趣的信息。这些包括在生物技术或生化研究背景下的蛋白质工程,以探索酶的机制。
Biocatalysis has emerged as an important alternative to traditional chemical synthesis for the preparation of fine chemicals,[1] as recently demonstrated by the biocatalytic manufacture of the drug sitagliptin by using an (R)-amine transaminase (ATA) created by intensive protein engineering.[2] This process was superior with respect to optical purity, yield, and waste generation to the already established transition-metal-catalyzed production of sitagliptin.[3] Process development, identification, and optimization of an appropriate enzyme represent inportant requirements to obtain a successful and efficient enzyme-catalyzed process. Nature has a rich reservoir from which a suitable enzyme can be found as a starting point. In contrast to classical approaches such as screening of strain collections, modern developments in the area of metagenomics offer enormous potential to screen for activity within nonculturable biodiversity.[4] A major advancement was the development of next-generation sequencing techniques, which led to a substantial increase in the genetic information deposited in public databases;[5] currently> 20million protein sequences are waiting to be explored. This development also led to the accumulation of protein sequences with unknown (or wrongly annotated) function, and thus a large part of this rich resource cannot be used reliably. We recently took advantage of this information and developed an in silico enzyme discovery strategy and identified 17 novel and unique (R)-selective ATAs in> 5000 sequences deposited in public databases,[6] although no gene or protein sequence was described in the literature ahead of our work. Furthermore, we could show that these (R)-ATAs are synthetically useful, as demonstrated recently in the asymmetric synthesis of a set of 12 chiral amines by using 7 out of the 17 enzymes.[7]Modern protein engineering methods [8] rely on high-quality structural information as obtained by X-ray crystallography, which is then the basis for focused, directed evolution to improve the properties of the enzyme. Similar to developments in the genomics area, in the last years automated highthroughput methods were developed for crystallization and determination of crystal structures, which has resulted in a rapid increase in the number of solved protein structures.[9] Still, for many enzymes no structure is available, as for example,(R)-and (S)-selective ATAs, which have become very popular in the last years.[10] Hence, structural information is highly desired to understand experimental results and to guide protein engineering.[11] Interestingly, a trend similar to that pointed out above for the discovery of sequences exists: there is a growing number of crystal structures of enzymes that were never characterized with respect to substrate scope, enantioselectivity, or reaction specificity, and their physiological functions are also often unknown. Similar to the lack of experimental validation of annotated proteins in sequence databases, crystallized proteins with unknown function are only rarely characterized.[12] Thus, many structures remain unexplored: For example, from 104 crystal structures of different transaminases found in the Brookhaven protein database (PDB), 46 structures do not have a literature citation that provides a description of the structure or the characterization data of the protein. Connecting the enzymatic function to unexplored proteins in the PDB could provide information that would be of diverse interest. These include protein engineering in the context of biotechnology or biochemical studies to explore enzyme mechanism.