课题基金 / 基金详情

项目摘要

项目成果

MARC NICKLAUS的其他基金

相似基金

相关文献

中文摘要
翻译
因此,我们与互变异构化相关的工作的一个动机是使用我们掌握的所有工具、化学信息学分析、QM计算、实验工作以及从文献中系统地提取结果,为如何改进Inchi V2中互变异构化的处理提供科学基础,而不仅仅是在工作组中进行投票。虽然原性互变异构化规则是目前在CACTVS中作为标准规则集实现的唯一规则,并且INCHI涵盖的所有互变异构化(默认或可选)都是原性互变异构化,但环链(RC)互变异构化是众所周知的和普遍的。然而,有些令人惊讶的是,直到最近,化学信息学中关于RC规则的研究还很少。基于Baldwin预测成环反应相对易度的一套众所周知的规则,我们发展了一套描述RC互变异构化的11条规则。规则被编码成SIMRKS线符号,这是由Daylight化学信息系统公司开发的化学结构线符号SPILES的化学变换扩展,就像目前CACTVS中描述原生互变异构的20个单独规则被编码一样。对鲍德温的规则集进行了一些修改,毕竟,这是一般的环闭合规则,而不是具体的RC互变异构规则。最重要的是,涉及四面体亲电碳的闭环和开环反应会导致单键的断裂,从而导致分子中原子的损失,违反了互变异构的定义。将这些新的RC规则添加到CACTVS中现有的标准原性规则中,我们将此组合规则集应用于RC互变异构化的“海报孩子”:华法林。这种广泛使用了几十年的抗凝血药,理论上可以以40种不同的互变异构体存在于溶液中。我们用计算方法(在B3LYP/6-311G+理论水平上计算了相对能量)和记录了核磁共振(13C和1H)谱,对所有这些互变异构体进行了研究。我们介绍了一个直观和图形的互变构体及其相互转化路径的网络,对于华法林,它包含11个互变构体和17个规则允许的它们之间的互变构体转换。然后,我们将组合的RC和原性规则集应用于整个数据库:Aldrich Market Select(AMS)数据库,其中包含(然后)600万个筛选样本和构建块。我们发现了30,000多个案例,其中两个或更多的AMS产品被我们的规则声明为同一化合物的不同互变异构体。对我们从AMS购买的166对这样的互变构体对(加上几个三联体)进行了1H和13C核磁共振分析,以确定化学信息学转换是否准确地预测了核磁共振所确定的相同的“瓶子里的东西”。基本上,所有在AMS中存在的例子的原性转换都得到了确认(在AMS中,一些“较罕见的”互变异构体类型没有这样的“冲突对”)。根据核磁共振分析,一些RC变换被发现过于“激进”,即将不同化合物的结构等同于彼此。这篇论文获得了《化学信息与模型》杂志的编辑选择。为了为互变异构相关的分析和化学信息学工作提供额外的实验数据,我们建立了一个基于从实验文献中提取的数据的数据库。这个数据库由1,873个条目组成,属于在一组特定的实验条件(pH,溶剂,温度,技术)下研究的互变构体的n元组,由于n的平均值略为2,因此总共增加了3,898条记录。数据来自73篇出版物,其中许多是评论,取自向进行初始提取的承包商公司(Parthys Reverse Informatics)提供的200篇论文中的精选文献,这些论文是我们在文献搜索中确定的大约900篇论文,其中可能包含对此目的有用的数据。每个互变构体(或适当的元组)都有结构信息的注释:微笑、因奇、因奇、NCI/CADD识别符;“流行”数据:测量的比率、相互转化率、相对能量等;条件数据:溶剂、温度、pH等(如果给定);方法数据:核磁共振、紫外光谱、红外光谱等;参考数据:书目信息。据我们所知,例如互硫化合物数据库不存在于其他地方,当然也不存在于公共领域。创建了一个新的Web服务--名为TAutomerizer--以应用和测试我们从上述数据库和文献中汇编的转换,以重新设计Inchi(Key)V.2中的互变异构处理。同时,在这个项目的上下文中编译的转换集已经增长到最终的86个,这些转换也被添加到TAutomerizer中。启动并随后在IUPAC工作组就建议用于INCHI V2的最后一套改造作出决定的阶段已经开始。将86个规则中的一些规则添加到当前INCHI代码(v.1.05)中的探索性编码对于6个规则是成功的。整合了这些规则的新的INCHI实验版本已经发布,以进行Beta测试。基于量子力学计算和随后的深度学习方法的互变异构化的第二层次分析工作已经开始。此外,还对上述小分子的一部分进行了X射线结晶学研究。在这方面已经出版了几份重要出版物。
英文摘要
One motivation of our tautomerism-related work is thus to use all tools at our disposal, chemoinformatics analyses, QM computations, experimental work, and systematic extraction of results from literature, to provide a scientific footing for the recommendations how to improve handling of tautomerism in InChI V2 - instead of just holding a vote in the Working Group. While prototropic tautomerism rules are the only ones currently implemented as the standard rule set in CACTVS, and all tautomeric transformations covered by InChI (as default or by option) are prototropic, ring-chain (RC) tautomerism is well-known and widespread. Nevertheless, and somewhat surprisingly, very little in terms of RC rules was available in chemoinformatics until recently. Based on Baldwin's well-known set of rules to predict the relative facility of ring forming reactions, we developed a set of 11 rules describing RC tautomerism. The rules were encoded in SMIRKS line notation, the chemical transform extension of the chemical structure line notation SMILES, developed by Daylight Chemical Information Systems, Inc., just like the currently 20 individual rules in CACTVS for describing prototropic tautomerism are encoded. A number of modifications were applied to Baldwin's rule set, which, after all, were rules for ring-closure in general, not for RC tautomerism in specific. Foremost, ring closure and opening reactions involving a tetrahedral electrophilic carbon thus leading to breakage of a single bond would cause a loss of atoms to the molecule, violating the definition of tautomerism. Adding these new RC rules to the existing standard prototropic rules in CACTVS, we applied this combined rule set to the "poster child" of RC tautomerism: warfarin. This anticoagulant drug, in wide use for decades, can theoretically exist in solution in 40 distinct tautomeric forms. We investigated all these tautomers with computational approaches (relative energies calculated at the B3LYP/6-311G+ level of theory) and recorded NMR (13C and 1H) spectra. We introduced an intuitive and graphical network for tautomers and their interconversion paths, which for warfarin contained 11 tautomers and 17 tautomeric transformations between them allowed by our rules. We then applied the combined RC and prototropic rule set to an entire database: the Aldrich Market Select (AMS) database of (then) 6 million screening samples and building blocks. We found over 30,000 cases where two or more AMS products were declared by our rules to be just different tautomeric forms of the same compound. 1H and 13C NMR analysis of 166 such tautomer pairs (plus a few triplets) we purchased from the AMS were performed to determine whether the chemoinformatics transforms had accurately predicted what was the same "stuff in the bottle" as determined by NMR. Essentially all prototropic transforms for which examples in the AMS existed (some of the "rarer" types of tautomerism had no such "conflict pairs" in the AMS) were confirmed. Some of the RC transforms were found to be too "aggressive", i.e. to equate structures with one another that were different compounds according to the NMR analyses. This paper received an Editor's Choice selection in the Journal of Chemical Information and Modeling. In order to provide additional experimental data for tautomerism-related analyses and chemoinformatics work, we have created a database based on data extracted from experimental literature. This database consists of 1,873 entries which belong to n-tuples of tautomers studied in a particular set of experimental conditions (pH, solvent, temperature, technique), adding up to 3,898 records since the average of n is slightly 2. The data were extracted from 73 publications, many of them reviews, taken from a selection of 200 papers provided to the contractor company that did the initial extraction (Parthys Reverse Informatics), out of about 900 papers we identified in literature searches that might contain useful data for this purpose. Each tautomer (or tuple, as appropriate) is annotated with Structural information: SMILES, InChI, InChIKey, NCI/CADD Identifiers; "Prevalence" data: measured ratios, interconversion rates, relative energies etc.; Condition data: solvent, temperature, pH etc. (if given); Method data: NMR, UV spectroscopy, IR spectroscopy etc.; Reference data: Bibliographic information. To the best of our knowledge, such as tautomer database does not exist elsewhere, certainly not in the public domain. A new web service - called Tautomerizer - was created to apply and test the transforms we have compiled from the above database and literature for the Redesign of Handling of Tautomerism in InChI(Key) V.2. The set of transforms compiled in the context of this project has meanwhile grown to its final number of 86, which are also being added to the Tautomerizer. The phase of initiating and then making a decision in the IUPAC Working Group about the final set of transforms to be recommended for InChI V2 has been started. Exploratory coding for adding some of the 86 rukes to the current InChI code (v.1.05) were successful for 6 rules. A new experimental version on InChI with these rules incorporated has been released for beta testing. Work on a second-level analysis of tautomerism based on quantum-mechanical calculations and subsequent Deep Learning approaches has been started. Also, X-ray crystallography on a subset of the small molecules mentioned above has been performed. Several important publications in this context have been published.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
HIV Integrase Modeling and Computer-Aided Inhibitor Deve
HIV Integrase Modeling and Computer-Aided Inhibitor and Microbicide Development
Fundamentals of Ligand-Protein Interactions
In Silico Screening for Cancer Targets
海外基金