Virtual Approaches to New Chemistries
Virtual Approaches to New Chemistries
批准号:
10447249
负责人:
BARRY A BUNIN
金额:
$44.0万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-06-06 至 2024-05-31
关键词:
AbbreviationsAddressAlgorithmsAutomationBackBiologicalBiological AssayCategoriesCharacteristicsChemical StructureChemicalsChemistryCollectionDataDatabasesDescriptorDrug DesignEvaluationFAIR principlesGenerationsGoalsHumanInformaticsInternetLearning ModuleMachine LearningMeasuresMethodologyModelingNational Center for Advancing Translational SciencesNatural Language ProcessingNatural regenerationNatureOntologyProcessProgram DevelopmentProtocols documentationQuantitative Structure-Activity RelationshipReactionReadabilityReagentRecipeResearch PersonnelRunningSchemeScientistSemanticsSolventsSorting - Cell MovementStructureSystemTechnologyTextUpdateVendorVisualWorkbasechemical reactiondeep learningdesigndrug developmentexperienceinstrumentinteractive toolknowledge basenatural languagenew technologynovelpreferencesmall moleculestoichiometrysuccesstoolvectorvirtual
中文摘要
项目摘要/摘要
两项新的虚拟化学技术将作为单独的模块添加到NCATS ASPIRE项目中。这个
第一个模块将使新的化学物质能够从尖端(深层)机器中建模和选择
使用直接从仪器获取的最新结构/活动数据学习技术。第二个模块
将是一种用于在语义模板中捕获富含化学物质的数据的新型信息学系统
机器可读的反应,将增加化学反应在电子实验室笔记本和
允许对反应分析(及其相应的反应)进行更精确的询问和自动化
产品)。
模块1中的深度学习技术基于我们新的化学富含向量(CRV)方法,
它能够高效地将关于化学结构的信息压缩成数的矢量
这允许反向编码过程:不仅可以将CRV转换回其原始格式
结构具有很高的成功率(>;90%完全匹配),但修改的CRV可以转换为符合以下条件的结构
化学空间中这一点的代表。CRV是用于SAR/QSAR迭代的优秀描述符
因为它们在一个很小的空间里包含了更多的化学信息,允许自动化
相对于传统的描述符,结构-活动模型将更加精简。由此产生的模型将
通过交互式视觉界面(人工指导)或后端探索多维空间
不断搜索新的和更好的结构的算法(机器指导的)。既有互动性也有
自动化流程将重新连接到ASPIRE自动化循环中,以便它们能够
综合和测量(假设评估和迭代优化)。
第二个模块是机器可读的反应,它借鉴了我们开发
BioHarmony注释器(以前是:BioATION Express),它使用自然语言模型来分配语义
本体论术语为生物检测协议,将它们从非结构化文本转换为机器可读数据。
从协议和化学结构图中提取反应的全部内容是非常困难的
鉴于文本的非结构化性质,缩写、快捷方式和假设进入图表。它是
由于需要将方案中的材料与反应文本描述(例如
试剂、溶剂、配方中涉及的顺序、反应过程和产品表征)。作为一种
或者,我们将对CDD化学计量草图进行模块化,这将允许我们提取这些数据。我们会
与NCATS合作,确定要捕获的重要领域,创建机器可读的化学反应
模板。
英文摘要
Project Summary/Abstract
Two new virtual chemistry technologies will be added to the NCATS ASPIRE project as separate modules. The
first module will enable new chemistries to be modelled and selected from cutting edge (deep) machine
learning technology using the latest structure/activity data taken directly from instruments. The second module
will be a novel informatics system for capturing chemistry-rich data in a semantic template as
machine-readable reactions which will increase the utility of chemical reactions in electronic lab notebooks and
allow more precise interrogation and automation of reaction analyses (and their corresponding reaction
products).
The deep learning technology in module 1 is based on our new chemically rich vector (CRV) methodology,
which is able to compress information about chemical structures into a vector of 64 numbers with an efficiency
that allows the encoding process to be reversed: not only can a CRV be converted back into its original
structure with high success (>90% exact match), but a modified CRV can be converted into a structure that is
representative of that point in chemical space. CRVs make excellent descriptors for SAR/QSAR iteration
because they contain much more chemical information in a small space, allowing the automation of
structure-activity models to be more streamlined, relative to conventional descriptors. The resulting models will
explore the multi-dimensional space via an interactive visual interface (human-directed) or a back-end
algorithm to constantly search for new and better structures (machine-directed). Both interactive and
automated processes will be connected back into the ASPIRE automation cycle so that they can be
synthesized and measured (hypothesis evaluation and iterative optimization).
The second module, machine-readable reactions, draws from our extensive experience developing the
BioHarmony Annotator (formerly: BioAssay Express) which uses natural language models to assign semantic
ontology terms to biological assay protocols, turning them from unstructured text into machine-readable data.
Extracting the full content of reactions from protocols and chemical structure diagrams is remarkably difficult
given the unstructured nature of text, abbreviations, shortcuts and assumptions that go into diagrams. It is
further complicated by the need to connect the materials in the scheme with the reaction text description (e.g.
reagents, solvents, the sequences involved in the recipe, reaction workup, and product characterization). As an
alternative, we will modularize the CDD stoichiometric sketcher, which will allow us to extract this data. We will
work with NCATS to identify important fields to capture, creating a machine readable chemical reaction
template.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Virtual Approaches to New Chemistries
-
批准号:10636882
-
项目类别:
-
资助金额:$44.0万
-
财政年份:2022
-
负责人:BARRY A BUNIN
-
依托单位:
Automated Molecular Identity Disambiguator (AutoMID)
-
批准号:10357906
-
项目类别:
-
资助金额:$28.0万
-
财政年份:2020
-
负责人:BARRY A BUNIN
-
依托单位:
Automated Molecular Identity Disambiguator (AutoMID)
-
批准号:10569639
-
项目类别:
-
资助金额:$28.0万
-
财政年份:2020
-
负责人:BARRY A BUNIN
-
依托单位:
Intelligent Chemical Structure Browser for Drug Discovery and Optimization
-
批准号:10241834
-
项目类别:
-
资助金额:$72.73万
-
财政年份:2019
-
负责人:BARRY A BUNIN
-
依托单位:
A Robust, Secure Framework to Effortlessly Bind Distributed Databases and Analysis Tools into Tightly Integrated Translational Drug Discovery Computational Platforms
-
批准号:10484172
-
项目类别:
-
资助金额:$85.49万
-
财政年份:2019
-
负责人:BARRY A BUNIN
-
依托单位:
Digital representation of chemical mixtures to aid drug discovery and formulation
-
批准号:9902210
-
项目类别:
-
资助金额:$74.87万
-
财政年份:2019
-
负责人:BARRY A BUNIN
-
依托单位:
A Robust, Secure Framework to Effortlessly Bind Distributed Databases and Analysis Tools into Tightly Integrated Translational Drug Discovery Computational Platforms
-
批准号:10685358
-
项目类别:
-
资助金额:$85.49万
-
财政年份:2019
-
负责人:BARRY A BUNIN
-
依托单位:
Intelligent Chemical Structure Browser for Drug Discovery and Optimization
-
批准号:10386918
-
项目类别:
-
资助金额:$72.73万
-
财政年份:2019
-
负责人:BARRY A BUNIN
-
依托单位:
Novel deep learning strategy to better predict pharmacological properties of candidate drugs and focus discovery efforts
-
批准号:10133177
-
项目类别:
-
资助金额:$74.99万
-
财政年份:2018
-
负责人:BARRY A BUNIN
-
依托单位:
Novel deep learning strategy to better predict pharmacological properties of candidate drugs and focus discovery efforts
-
批准号:10004481
-
项目类别:
-
资助金额:$74.99万
-
财政年份:2018
-
负责人:BARRY A BUNIN
-
依托单位:
Unifying Templates, Ontologies and Tools to Achieve Effective Annotation of Bioassay Protocols
-
批准号:9398728
-
项目类别:
-
资助金额:$54.64万
-
财政年份:2017
-
负责人:BARRY A BUNIN
-
依托单位:
Comprehensive but simple encoding of bioassays to accelerate translational drug discovery
-
批准号:9464228
-
项目类别:
-
资助金额:$74.43万
-
财政年份:2017
-
负责人:BARRY A BUNIN
-
依托单位:
Unifying Templates, Ontologies and Tools to Achieve Effective Annotation of Bioassay Protocols
-
批准号:9979969
-
项目类别:
-
资助金额:$51.14万
-
财政年份:2017
-
负责人:BARRY A BUNIN
-
依托单位:
Simplifying encoding of bioassays to accelerate translational drug discovery
-
批准号:8901698
-
项目类别:
-
资助金额:$75.14万
-
财政年份:2013
-
负责人:BARRY A BUNIN
-
依托单位:
Simplifying encoding of bioassays to accelerate translational drug discovery
-
批准号:8591013
-
项目类别:
-
资助金额:$15.0万
-
财政年份:2013
-
负责人:BARRY A BUNIN
-
依托单位:
Biocomputation across distributed private datasets to enhance drug discovery
-
批准号:9345057
-
项目类别:
-
资助金额:$75.05万
-
财政年份:2013
-
负责人:BARRY A BUNIN
-
依托单位:
Biocomputation across distributed private datasets to enhance drug discovery
-
批准号:8198305
-
项目类别:
-
资助金额:$15.0万
-
财政年份:2011
-
负责人:BARRY A BUNIN
-
依托单位:
海外基金