Core G: Computation
Core G: Computation
批准号:
7980202
负责人:
JOHN A GERLT
金额:
$103.88万
依托单位国家:
美国
项目类别:
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-05-20 至 2015-04-30
关键词:
AcidsAdenineAreaBackBenchmarkingChemicalsCollaborationsCommitCommunitiesComplexComputer softwareComputing MethodologiesConsultCore ProteinCrystallizationCytidineDatabasesDevelopmentDockingElementsEnsureEntropyEnzymesFeedbackGenomeGluesGlutathione S-TransferaseGoalsGrantGuanosineHomocysteineHomocystineHomology ModelingHybridsHydrogenInternetLibrariesLigandsMetabolismMethodsModelingOnline SystemsOperonProteinsReactionResearch PersonnelResourcesSamplingSolventsSpecificityStructureTestingTrainingVariantWorkWritingXenobioticsZincadenosine deaminaseanalogbasecheminformaticscomputerized toolsdrug discoveryenolaseflexibilityfunctional groupinhibitor/antagonistinnovationinterestisoprenoidmeetingsmembermethod developmentoutreachprotonationsmall moleculetheoriestoolweb interface
中文摘要
计算管道概述。图1总结了我们
建议应用于功能分配(与每个步骤相关的软件和网络资源分别用斜体、红色和蓝色表示)。我们将在超家族/基因组核心的指导下选择要建模的序列。这些将包含作为子集的由蛋白质核心产生的所有序列,以及在具有靶酶的操纵子中发现的其他酶。当结构不可用时,将创建同源模型,包括结构核心尝试结晶的情况。
已知代谢物以及“片段”化合物的库将对接到其基态的结构和模型,并且在可能的情况下,对接到高能中间体(HEI)形式。预测的蛋白质和复合物将在对接后使用更高水平的理论(全原子力场和隐式溶剂)进行改进和重新排序,其中蛋白质被视为柔性的。对接命中名单将在自动分析
Shoichet小组以前为药物发现应用开发的化学信息学方法。最后,通过创建创新的混合方法,正在努力合并蛋白质和配体采样模块。
我们已经确定了“新的”超家族(GST、HAD和IS),2)我们的目标是将计算方法扩展到适用于所有酶超家族,3)我们的目标是使计算方法自动化,这样它们最终可以通过网络界面被社区使用(第3节)。一些建议的增强功能建立在我们对AH和EN超家族进行的初步测试的基础上,
P01 GM071790。这里的重点是概括这些方法,因此没有重叠。其他的,有点更投机,计算方法的发展,计划与P01 GM 071790的支持,如治疗配体熵损失和蛋白质和配体质子化状态的自动预测,也将被添加到一般管道,如果他们是成功的AH和EN超家族的初始测试。
迈向通用方法的重要一步:对接“碎片样”分子以扩展
化学型探索在我们之前的工作中做出的一个富有成效的选择是将我们的对接计算限制在约10,000种已知代谢物。如果靶向的酶参与初级代谢,如Tm 0936,这是一个合适的选择;但是,如果不是,将错过真正的底物。外源性物质或次级代谢物是一个特殊的挑战。因此,扩大在初始对接计算中筛选的数据库中表示的化学型似乎是谨慎的。为此,我们建议筛选一个包含130,000个片段样分子的文库。这些分子很小,< 17个非氢(重)原子,并且被认为比类似的较大分子库覆盖超过15个数量级的化学型[5,6]。
由于这个原因,它们已经成为抑制剂发现的强烈兴趣的焦点[7,8]。作为较小的分子,它们本质上更容易对接,正如我们在抑制剂对接中所发现的那样。最后,因为它们是商业上可获得的,所以它们将直接获得和测试。
我们还将把ZINC数据库[9]中的130,000个片段转换为HEI结构。在P01 GM 071790的支持下,我们正在制备与AH超家族成员催化的20个核心反应相关的HEI [10]。在这里,我们将把这种方法扩展到
由EN、GST、HAD和IS超家族的酶催化的反应。然后将这些HEI片段对接到已知结构和功能的基准酶组,以查看它们是否可以重现用较大分子发现的底物富集。例如,当与Tm 0936对接时,腺嘌呤(11个重原子肯定是一个片段)与片段诱饵相比,
S-腺苷同型半胱氨酸(SAH)对代谢物诱饵?与较大代谢产物观察到的鸟苷和胞苷类似物相比,它是否会显示出选择性
对接?这些问题将通过回顾性计算得到明确的回答。
可以想象,这种做法不会成功。
Wolfenden [11]和其他人已经表明,当底物被解构成片段时,酶对其的识别可能会严重受损。很容易
想想病理情况,其中存在于较大分子中的官能团对于识别和特异性至关重要(例如,人们将无法区分
仅使用腺嘌呤HEI作为对接探针的腺嘌呤和腺苷脱氨酶之间的关系)。相反,我们可以想象,
从片段筛选中出现的初始化学型:例如,如果腺嘌呤HEI排序良好,则在该受限空间中尝试较大的变化。与核心代谢物相比,片段中所代表的化学型数量级更多,并且能够实际获取和测试它们中的每一个,使得这种方法值得探索。它有可能大大增加基于结构的基板预测的范围和通用性。
英文摘要
Overview of the computational pipeline. Figure 1 summarizes the modeling pipeline that we
propose to apply to functional assignment (software and web resources associated with each step are written in italics, red and blue, respectively). We will be guided by the Superfamily/Genome Core in choosing which sequences to model. These will contain, as a subset, all of the sequences produced by the Protein Core, as well as other enzymes found in operons with the target enzymes. Homology models will be created when structures are not available, including in cases where crystallization will be attempted by the Structure Core.
Libraries of known metabolites as well as "fragment" compounds will be docked against the structures and models in their ground state and, when possible, high-energy intermediate (HEI) forms. Predicted proteinlig and complexes will be refined and re-ranked after docking using a higher level of theory (all-atom force fields and implicit solvent), with the protein treated as flexible. The docking hit lists will be analyzed in an automated
manner using cheminformatics methods that the Shoichet group previously developed for drug discovery applications. Finally, work is undenway to merge the protein and ligand sampling modules, by creating innovative hybrid methods The enhancements to the computational pipeline described below are motivated by 1) challenges that
we have identified for the "new" superfamilies (GST, HAD, and IS), 2) our goal of extending the computational methods to apply to all enzyme superfamilies, and 3) our goal of automating the computational methods, such that they are ultimately usable by the community via web interfaces (Section 3). Some of the proposed enhancements build on preliminary tests that we have performed for the AH and EN superfamilies, supported
by P01 GM071790. The focus here is on generalizing these approaches, so there is no overlap. Other, somewhat more speculative, computational methods development that is planned with the support of P01 GM071790, such as treatment of ligand entropy losses and automated prediction of protein and ligand protonation states, will also be added to the general pipeline if they are successful in initial tests on the AH and EN superfamilies.
An important step towards a general method: Docking "fragment-like" molecules to expand
chemotype exploration. A fruitful choice made in our prior work was to restrict our docking calculations to ~10,000 known metabolites. If the enzyme targeted is involved in primary metabolism, as was Tm0936, this is an appropriate choice; but, if it is not, the true substrate will be missed. Xenobiotics or secondary metabolites represent a particular challenge. It seems prudent, therefore, to expand the chemotypes represented in the database being screened in the initial docking calculation. To do so, we propose to screen a library of 130,000 fragment-like" molecules. These molecules are small, < 17 non-hydrogen (heavy) atoms, and are thought to cover over 15 orders of magnitude more chemotypes than would a similar library of larger molecules [5, 6].
For this reason, they have become a focus of intense interest in inhibitor discovery [7, 8]. As smaller molecules, they will be intrinsically easier to dock, as we have found in docking for inhibitors. Finally, because they are commercially available, they will be straightfonward to acquire and test.
We will also convert the 130,000 fragments in the ZINC database [9] to HEI structures. With the support of P01 GM071790, we are preparing HEIs associated with 20 core reactions catalyzed by members of the AH superfamily [10]. Here, we will expand this approach to
reactions catalyzed by the enzymes of the EN, GST, HAD, and IS superfamilies. These HEI fragments will then be docked against the benchmarking set of enzymes of known structure and function to see if they can recapitulate the substrate enrichments found with larger molecules. For example, when docking against Tm0936, will adenine, which at 11 heavy atoms is certainly a fragment, rank as well, compared to the fragment decoys, as does
S-adenosyl homocysteine (SAH) against the metabolite decoys? Will it show the selectivity compared to guanosine and cytidine analogs observed with the larger metabolite
docking? These questions will be definitively answered by retrospective calculations.
It is conceivable that this approach will not succeed.
Wolfenden [11] and others have shown that when a substrate is deconstructed into fragments its recognition by the enzyme can be severely compromised. It is easy to
think of pathological cases where functional groups present in the larger molecules will be critical to recognition and specificity (one will not, for instance, be able to distinguish
between adenine and adenosine deaminase using merely the adenine HEI as a docked probe). Conversely, one can imagine building the larger molecules back from the
initial chemotypes emerging from the fragment screen: for example, if adenine HEI ranks well, try larger variations in this restricted space. The orders-of-magnitude more chemotypes represented among the fragments compared to the core metabolites, and the ability to actually acquire and test every one of them, makes this approach worth exploring. It has the possibility of substantially increasing the reach and generality of structure-based substrate prediction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Web-Based Resource for Genomic Enzymology Tools
-
批准号:10548888
-
项目类别:
-
资助金额:$56.12万
-
财政年份:2022
-
负责人:JOHN A GERLT
-
依托单位:
Novel Strategies for the Discovery of Microbial Metabolic Pathways
-
批准号:9918932
-
项目类别:
-
资助金额:$232.88万
-
财政年份:2016
-
负责人:JOHN A GERLT
-
依托单位:
Novel Strategies for the Discovery of Microbial Metabolic Pathways
-
批准号:9297333
-
项目类别:
-
资助金额:$232.88万
-
财政年份:2016
-
负责人:JOHN A GERLT
-
依托单位:
Metabolism Project
-
批准号:9073786
-
项目类别:
-
资助金额:$42.29万
-
财政年份:2016
-
负责人:JOHN A GERLT
-
依托单位:
Novel Strategies for the Discovery of Microbial Metabolic Pathways
-
批准号:9557783
-
项目类别:
-
资助金额:$11.68万
-
财政年份:2016
-
负责人:JOHN A GERLT
-
依托单位:
GENOMIC ENZYMOLOGY: THE ENOLASE SUPERFAMILY AND OMPDC SUPRAFAMILY
-
批准号:8363583
-
项目类别:
-
资助金额:$1.68万
-
财政年份:2011
-
负责人:JOHN A GERLT
-
依托单位:
DECIPHERING ENZYME SPECIFICITY
-
批准号:8363605
-
项目类别:
-
资助金额:$1.68万
-
财政年份:2011
-
负责人:JOHN A GERLT
-
依托单位:
COLLABORATIVE CENTER FOR AN ENZYME FUNCTION INITIATIVE
-
批准号:7901811
-
项目类别:
-
资助金额:$702.3万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
Core A: Administrative Core
-
批准号:7980192
-
项目类别:
-
资助金额:$49.22万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
Core F: Structure
-
批准号:7980201
-
项目类别:
-
资助金额:$76.19万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
GENOMIC ENZYMOLOGY: THE ENOLASE SUPERFAMILY AND OMPDC SUPRAFAMILY
-
批准号:8170502
-
项目类别:
-
资助金额:$1.79万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
Core D: Superfamily/Genome
-
批准号:7980199
-
项目类别:
-
资助金额:$35.29万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
DECIPHERING ENZYME SPECIFICITY
-
批准号:8170532
-
项目类别:
-
资助金额:$1.79万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
COLLABORATIVE CENTER FOR AN ENZYME FUNCTION INITIATIVE
-
批准号:8489131
-
项目类别:
-
资助金额:$625.01万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
COLLABORATIVE CENTER FOR AN ENZYME FUNCTION INITIATIVE
-
批准号:8074489
-
项目类别:
-
资助金额:$647.68万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
Bridging Project 4: Haloacid Dehalogenase (HAD) Superfamily
-
批准号:7980210
-
项目类别:
-
资助金额:$40.76万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
COLLABORATIVE CENTER FOR AN ENZYME FUNCTION INITIATIVE
-
批准号:8665973
-
项目类别:
-
资助金额:$582.91万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
Bridging Project 3: Glutathione Transferase (GST) Superfamily
-
批准号:7980209
-
项目类别:
-
资助金额:$30.12万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
Core B/C: Data & Dissemination
-
批准号:7980195
-
项目类别:
-
资助金额:$39.36万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
Core E: Protein
-
批准号:7980200
-
项目类别:
-
资助金额:$167.1万
-
财政年份:2010
-
负责人:JOHN A GERLT
-
依托单位:
海外基金