Large Databases of Small Molecules - Drug Development Tool and Public Resource
Large Databases of Small Molecules - Drug Development Tool and Public Resource
批准号:
8554069
负责人:
MARC NICKLAUS
金额:
$31.3万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AlgorithmsAreaBiologicalBiological AssayBooksCactaceaeCatalogingCatalogsCharacteristicsChemical StructureChemicalsCollaborationsCollectionComputer AssistedComputer SimulationComputersContractsCustomDataData SetDatabasesDepositionDevelopmental Therapeutics ProgramDrug DesignEvaluationGenerationsGoalsHousingImageInformation SciencesInternetJavaLegal patentLibrariesLinkLiteratureMalignant NeoplasmsMethodsMolecularMolecular StructureNatureOpticsPaperPharmacologic SubstanceProcessPropertyPubChemPublicationsPublished CommentReactionReadingRecordsResearch PersonnelResourcesSamplingScreening procedureSeriesServicesStructureSynthesis ChemistrySystemTelephoneTextTextbooksTimeUnited States National Institutes of HealthUpdateVendorWorkWritingabstractingbasechemical synthesisdesigndrug developmentimprovedinsightjournal articlenext generationpharmacophoreprogramssmall moleculetoolweb based interfaceweb serviceswiki
中文摘要
该项目的主要目标是收集大量的小分子,以帮助内部和公众进行药物开发,推进化学结构鉴定和处理以及唯一化合物标识符生成领域的发展,并提供免费的化学信息学工具,帮助人们处理这些数据库。这个项目开始于在CADD集团的公共网络服务器上的开放NCI数据库中发布信息,但已经远远超出了这个数据集。目前,正在向这一资源添加额外的数据库,包括可以获得用于筛选的化合物的大型供应商目录。对数据进行了高级处理,实现了强大的搜索和显示功能。目前正在与世界各地的政府、学术和工业团体合作努力,以大大增加在本项目框架内可获得的有关计算性质的总数和范围。这些努力旨在使其成为计算机筛选和计算机辅助药物设计的强大资源。这些数据库的一种接口将类似于增强型NCI数据库浏览器。当前正在开发的资源的性质应通过该服务的简要描述来举例说明:当前web服务中的数据包括来自NCI发展治疗计划(DTP)的数据以及我们已增强DTP数据集的附加信息。我们对大约26万种化合物的NCI开放数据库进行了各种分析,以帮助更好地了解其特征,并将其与计算机辅助药物设计和化学信息科学中使用的其他大型数据库进行比较。本文采用了不同的聚类方法来阐明其多样性,并将聚类结果与其他数据库的聚类结果进行了比较。开放NCI数据库已转换成各种格式,适合进一步处理,包括3D药效团搜索。我们还为Open NCI数据库实现了一个强大的公共搜索工具,该工具基于化学信息工具包CACTVS提供了一个web界面。仅使用一个网络浏览器,用户就可以根据600多个标准搜索大约25万个结构。我们用许多额外的数据字段极大地增强了原始DTP文件,无论是计算的、预测的还是超链接的信息。这些数据也以可直接下载的格式提供。已经实现了几个附加服务的链接,以便进行进一步处理。一个在线3D药效团功能已经建立起来,据我们所知,这个功能目前在网络上是独一无二的。通过PASS程序计算出的超过550种不同生物活性的可搜索预测,已经包含在web服务中(摘要)。最近的一项服务是我们的化学结构查找服务(CSLS),可在http://cactus.nci.nih.gov/lookup获得。CSLS本质上是小分子的“电话簿”,允许用户快速找到,如果有的话,超过100个不同的数据库(包括公共和商业),包含超过7400万个条目,它们的化合物出现在哪里。在撰写本文时,用户界面、结构和数据持有的更新正在进行中,这将使CSLS中的条目数量超过1亿大关。这些项目的一部分是下载,重新格式化和评估癌症相关的目的,存储在PubChem中的大量结构和分析数据集。我们最近增加的公共工具是我们的分子光学结构识别服务,主要由Igor Filippov博士开发。OSRA是一个实用程序,用于转换化学结构的图形表示,因为它们出现在期刊文章,专利文件,教科书,贸易杂志等。,进入SMILES(简化分子输入行输入规范-见http://en.wikipedia.org/wiki/SMILES -一种计算机可识别的分子结构格式。OSRA可以读取超过90种可解析图形格式的文档,包括GIF、JPEG、PNG、TIFF、PDF、PS等,并生成该文档中遇到的分子结构图像的SMILES表示。OSRA最新版本中最引人注目的新增功能是自动识别专利和文献中的反应。OSRA现在支持从周围的文本和图形中多步提取反应,反应试剂和评论的OCR,以及将结果转换为反应SMILES或RXN格式。OSRA可以作为一个独立的命令行实用程序,一个库(通过JNI的c++和Java),也可以通过作为Accelrys Draw插件提供的GUI来使用,该插件已经更新以处理新的反应识别功能。最近增加了我们的公共化学信息学工具和服务,化学标识解析器(CIR),由Markus Sitzmann博士开发。CIR作为不同化学结构标识符的解析器,允许将给定的结构标识符转换为另一种表示或结构标识符。其中,我们内部开发的NCI/CADD结构标识符以及新的标准InChI和InChIKey标识符都由该服务处理。在过去的几个月里,该服务的使用量大幅增加,目前平均有4000名用户每月提交大约350万次请求。CIR的一个?它的主要特点是它是一个进入化学结构数据库(CSDB)的编程接口。CSDB的更新已经完成了超过3.4亿条原始数据库记录,代表了大约1.2亿个独特的小分子结构(使其成为世界上最大的化学数据库之一)。许多额外的功能继续被添加到这个服务中,它越来越多地与世界范围内的其他web服务和化学信息学工具集成。CIR在涉及化学结构的出版物领域也将变得越来越重要,因为越来越多的努力要求在一篇论文中包含所有化合物的计算机可读表示。Sitzmann最近还完成了下一代web平台的第一个测试版的工作,它将成为一系列新web服务和现有服务更新的基础,包括CADD Group的化学结构查找服务(CSLS II; http://cactus.nci.nih.gov/TEST/chemical/apps/csls)。我们的公共web服务器的URL是http://cactus.nci.nih.gov.Dr。Sitzmann最近开始分析由ibm领导的大型制药公司联盟在SIIP(战略知识产权洞察平台)项目背景下从专利数据(EP, US PTO, WO)中提取的4300万条化学结构记录。OSRA在这个项目中使用。这些数据的一部分被提供给PubChem和CADD Group供公众使用(例如,参见http://www-935.ibm.com/services/us/gbs/bao/siip/nih/?sid=0015AFBF08D8F183C1F8E32A430CFFEB).Finally)。为使所有NIH研究人员都能负担得起筛选样品的化学合成而实施新资源的努力已成功结束。ChemNavigator公司现在是Sigma-Aldrich公司的一部分,该公司已经实施了所谓的半定制合成在线请求系统(SCSORS)。这种资源越来越多地用于我们(和其他团队)的硅筛选、合成化学和样品获取项目。
英文摘要
The principal objective of this project is to make large collections of small molecules available for aiding in drug development, both in-house and publicly, to advance the fields of chemical structure identification and processing and of unique compound identifier generation, as well as to provide free chemoinformatics tools aiding one in dealing with such databases. This project started with posting the information in the Open NCI Database on the CADD Group's public web server, but has moved far beyond this data set. Currently, additional databases are being added to this resource, including large vendor catalogs of compounds that can be acquired for screening. Advanced processing is applied to the data, and powerful searching and display capabilities are being implemented. Current efforts in collaboration with governmental, academic and industrial groups world-wide are underway to greatly enhance the total number and scope of associated calculated properties available in the framework of this project. These efforts are intended to make this a powerful resource in in silico screening and computer-aided drug design.One type of interface to these databases will resemble the Enhanced NCI Database Browser. The nature of the resources currently being developed shall be exemplified by a brief description of this service: The data in this current web service comprise data from NCI's Developmental Therapeutics Program (DTP) and additional information with which we have augmented the DTP data sets.We have subjected the Open NCI Database of about 260,000 compounds to various analyses that help to better understand its characteristics and put it in perspective of other large databases used in computer-aided drug design and chemical information sciences. Various clustering methods have been applied to it to elucidate its diversity, and the results have been compared with those for other databases.The Open NCI Database has been converted into various formats, suitable for further processing including 3D pharmacophore searching. We have also implemented a powerful public search tool for the Open NCI Database with a web interface based on the chemical information toolkit CACTVS. Using just a web browser, the user is able to search about 250,000 structures for more than 600 criteria. We have greatly augmented the original DTP files with numerous additional data fields, be it calculated, predicted or hyperlinked information. These data have also been made available in directly downloadable format. Links to several additional services for further processing have been implemented. An online 3D pharmacophore capability has been built, a capability that is currently unique on the web, as far as we are aware of. Searchable predictions of more than 550 different biological activities, calculated by the program PASS for most of the quarter-million compounds, have been included in the web service (abstract).A more recent service is our Chemical Structure Lookup Service (CSLS), available at http://cactus.nci.nih.gov/lookup. CSLS is essentially a "phone book" for small molecules, allowing the user to quickly find out in which, if any, of over 100 different databases (both public and commercial), comprising more than 74 million entries, their compounds occur. Updates of both the user interface and the structure and data holdings are underway as of the time of this writing, which will push the number of entries in CSLS beyond the 100 million mark.Part of these projects is the downloading, reformatting and evaluation for cancer-related purposes, of the massive set of structure and assay data as deposited in PubChem.A recent addition to our collection of public tools is our Optical Structure Recognition service for molecules, mostly developed by Dr. Igor Filippov. OSRA is a utility designed to convert graphical representations of chemical structures, as they appear in journal articles, patent documents, textbooks, trade magazinesetc., into SMILES (Simplified Molecular Input Line Entry Specification - see http://en.wikipedia.org/wiki/SMILES - a computer recognizable molecular structure format. OSRA can read a document in over 90 graphical formats parseable - including GIF, JPEG, PNG, TIFF, PDF, PS etc., and generate the SMILES representation of the molecular structure images encountered within that document.The most notable addition in the latest version of OSRA is automatic recognition of reactions in patents and literature. OSRA now supports multi-step reaction extraction from the surrounding text and graphics, OCR of the reaction agents & comments, and the conversion of the results to either reaction SMILES or RXN format. OSRA can be used as a stand-alone command-line utility, a library (C++ and Java through JNI), and also through the GUI provided as Accelrys Draw plugin which has been updated to handle the new reaction recognition capabilities.A recent addition to our public chemoinformatics tools and services, the Chemical Identifier Resolver (CIR), developed by Dr. Markus Sitzmann. CIR works as a resolver for different chemical structure identifiers and allows one to convert a given structure identifier into another representation or structure identifier. Among others, our NCI/CADD Structure Identifiers developed in-house as well as the new Standard InChI and InChIKey identifiers are handled by this service. The usage of the service has increased strongly over the last few month and is currently used by an average of 4000 users submitting approximately 3.5 million requests per month. One of CIR?s key features is that it is a programmatic interface into the Chemical Structure Database (CSDB). An update of CSDB has been completed to over 340 million original database records representing approximately 120 unique million small-molecule structures (making this one of the largest chemical databases in the world). Many additional capabilities continue to be added to this service, which is increasingly being integrated with other web services and chemoinformatics tools world-wide. CIR will also become increasingly important in the area of publications involving chemical structures, as efforts increase to make inclusion of computer-readable representations of all compounds presented in a paper mandatory.Dr. Sitzmann has also very recently completed work on the first beta version of the next generation web platform which will be the basis for a series of new web services and updates of existing services including CADD Group's Chemical Structure Lookup Service (CSLS II; http://cactus.nci.nih.gov/TEST/chemical/apps/csls). The URL of our public web server is http://cactus.nci.nih.gov.Dr. Sitzmann recently started to analyze a set of 43 million chemical structure records extracted from patent data (EP, US PTO, WO) by the IBM-led consortium of large pharmaceutical companies in the context of the SIIP (Strategic IP Insight Platform) project. OSRA was used in this project. Part of these data were given for public use to both PubChem and the CADD Group (see, e.g., http://www-935.ibm.com/services/us/gbs/bao/siip/nih/?sid=0015AFBF08D8F183C1F8E32A430CFFEB).Finally, efforts to implement a new resource for making affordable chemical synthesis of screening samples available to all NIH researchers were successfully concluded. This was realized in the form of an extension of the contract with the company ChemNavigator, now part of Sigma-Aldrich, who have implemented the so-called Semi-Custom Synthesis Online Request System (SCSORS). This resource is being increasingly used in our (and other groups') in silico screening, synthetic chemistry, and sample acquisition projects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
HIV Integrase Modeling and Computer-Aided Inhibitor Deve
-
批准号:7291875
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
HIV Integrase Modeling and Computer-Aided Inhibitor Development
-
批准号:7965392
-
项目类别:
-
资助金额:$21.29万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
HIV Integrase Modeling and Computer-Aided Inhibitor and Microbicide Development
-
批准号:10702372
-
项目类别:
-
资助金额:$3.61万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
HIV Integrase Modeling and Computer-Aided Inhibitor Development
-
批准号:7733068
-
项目类别:
-
资助金额:$20.03万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Fundamentals of Ligand-Protein Interactions
-
批准号:10926079
-
项目类别:
-
资助金额:$4.62万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Large Databases of Small Molecules - Drug Development Tool and Public Resource
-
批准号:10926595
-
项目类别:
-
资助金额:$13.85万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Synthetically Accessible Virtual Inventory (SAVI)
-
批准号:10926263
-
项目类别:
-
资助金额:$36.94万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Fundamentals of Ligand-Protein Interactions
-
批准号:10014461
-
项目类别:
-
资助金额:$6.9万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
In Silico Screening for Cancer Targets
-
批准号:7592817
-
项目类别:
-
资助金额:$43.5万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Large Databases of Small Molecules - Drug Development Tool and Public Resource
-
批准号:10703018
-
项目类别:
-
资助金额:$18.06万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Better Understanding and Handling of Tautomerism
-
批准号:10262460
-
项目类别:
-
资助金额:$21.34万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Tools for Prediction of ADME-Tox Properties
-
批准号:10262292
-
项目类别:
-
资助金额:$3.37万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Large Databases of Small Molecules - Drug Development Tool and Public Resource
-
批准号:10262724
-
项目类别:
-
资助金额:$11.23万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Tools for Prediction of Drug Metabolism and Metabolites
-
批准号:8157776
-
项目类别:
-
资助金额:$19.95万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
HIV Integrase Modeling and Computer-Aided Inhibitor Development
-
批准号:8157333
-
项目类别:
-
资助金额:$14.25万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Fundamentals of Ligand-Protein Interactions
-
批准号:8157489
-
项目类别:
-
资助金额:$24.22万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Better Understanding and Handling of Tautomerism
-
批准号:10486976
-
项目类别:
-
资助金额:$18.27万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Fundamentals of Ligand-Protein Interactions
-
批准号:10702419
-
项目类别:
-
资助金额:$6.02万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Synthetically Accessible Virtual Inventory (SAVI)
-
批准号:9779991
-
项目类别:
-
资助金额:$38.77万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
Large Databases of Small Molecules - Drug Development To
-
批准号:7338636
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:MARC NICKLAUS
-
依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
-
批准号:2021JJ40433
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:孙磊
-
依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
-
批准号:32001603
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:段真珍
-
依托单位:
AREA国际经济模型的移植.改进和应用
-
批准号:18870435
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1988
-
负责人:史树中
-
依托单位: