Centralized assay datasets for modelling support of small drug discovery organizations
Centralized assay datasets for modelling support of small drug discovery organizations
批准号:
9751326
负责人:
SEAN EKINS
金额:
$69.28万
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-01-01 至 2021-07-31
关键词:
Algorithm DesignAlgorithmsAreaArtificial IntelligenceBackBayesian ModelingBayesian learningBindingBiologicalBiological AssayCaringCatalogsCellsChemistryClientCollaborationsComputer SimulationComputer softwareConsultCytochrome P450DataData QualityData SetData SourcesDatabasesDecision TreesDescriptorDevelopmentDiseaseEbola virusEmploymentEnsureEnvironmentEstrogen ReceptorsEvaluationFDA approvedFeedbackFeesFingerprintFoundationsGrowthHIVHIV/TBIndustry StandardIntelligenceIntentionJavaJudgmentKnowledgeLeadLeishmaniasisLicensingLinear RegressionsLinkLiteratureLogisticsMachine LearningMeasurementMeasuresMetadataMethodsModelingModificationMolecularNuclear ReceptorsOutputPathway interactionsPharmaceutical PreparationsPharmacologic SubstancePhasePrivatizationProcessProductionPropertyPublic DomainsPublicationsRNA-Directed DNA PolymeraseResearchResearch PersonnelResourcesRightsScientistSeriesSoftware ToolsStatistical Data InterpretationStructureStructure-Activity RelationshipSynthesis ChemistryTechnologyTestingToxic effectToxicologyTraining SupportTriageTuberculosisUpdateVendorVisualization softwareWorkbasecomputational suitecostdata modelingdata submissiondata visualizationdeep learningdeep neural networkdesigndrug discoveryexhaustionimprovedlearning strategymachine learning algorithmneglectnovel therapeuticsprospectiveprototyperandom forestscale upscreeningsmall moleculesoftware developmentstatisticssuccesstext searchingtooltrend
中文摘要
项目摘要
人工智能(AI)日益增长的重要性从公司的增长和交易的增加可见一斑
在过去的一年里,制药公司和规模较小的公司之间使用机器学习来帮助药物发现。
不同靶点、疾病和分子性质的结构-活性数据的持续稳定增长
带来了相当大的挑战,因为它们通常不容易被机器学习访问:内容驻留在
在混合的公共数据库中(具有不同的管理级别),研究小组内的不同文件,非
精心策划的文学出版物。在第一阶段,Collaborations制药公司开发了Assay的原型
中央软件并将其与来自公共和私人来源的各种结构活动数据一起使用,
格式化和非格式化,用于启用忽略的、罕见的或常见的疾病目标。公共数据混合了
协作者/客户贡献的数据,使用原始软件和专家的应用化学判断
一队。在第一阶段,我们创建了错误检查和纠正软件。我们还构建并验证了贝叶斯模型
使用收集和清理的数据集。此外,我们还开发了新的数据可视化工具。
我们创建的软件环境使用户能够轻松地编译用于构建的结构-活动数据
计算模型,并可用于创建这些模型的选择,以便与协作者共享,如
需要的。该软件可以反过来用于对新分子进行评分,并将多个输出结果可视化
各种格式。我们已经支持了大约14个协作项目,这些项目在特定目标上共享了模型,例如
作为PYRG用于结核病(鉴定先导化合物),艾滋病毒逆转录酶,全细胞筛查
利什曼病以及与毒理学相关的P450和核受体模型(例如雌激素受体)。我们
在我们正在进行的埃博拉、艾滋病毒和结核病小范围内部项目中使用了Assay Central
分子药物发现。
在第二阶段,我们提出了以下目标,使我们能够将Assay Central开发为生产工具
促进药物发现合作,我们将继续关注这一点。在第1阶段,我们执行了
用选定的药物发现数据集初步分析了不同的机器学习算法。在第二阶段,我们
现在将对其他机器学习算法和分子进行彻底的评估和选择
描述符以及对算法组合的评估(例如,贝叶斯和深度学习)。我们会
实施机器学习模型的疾病/目标定义,以促进药物发现。我们将启用
分子选择以及自动化设计和优化。拥有像Assay Central这样的工具的用处
随时可用的数据将使科学家能够利用公共、私人或组合数据来帮助他们的
药物发现任务。使用公开数据开发这套计算模型软件将使我们能够
确定生成初步数据以测试模型的基金会、学者和潜在合作者。这些
这些努力将极大地增加我们可以从事的项目数量,创造新的知识产权,并产生
使用机器学习的就业重点是罕见和被忽视疾病领域的药物发现,在
很特别。分析中心的好处包括1.易于部署和使用由用户执行的Java文件
无需IT支持;2.基于行业标准技术;3.模型图形化展示
提供即时反馈;4多种方法评估分数和图形的模型适用性。
英文摘要
Project Summary
The growing importance of artificial intelligence (AI) is visible by the growth in companies and increasing deals
over the past year between pharma and smaller companies using machine learning to assist in drug discovery.
The continuing steady growth of structure-activity data for diverse targets, diseases and molecular properties
poses a considerable challenge as they are generally not readily accessible for machine learning: content resides
in a mixture of public databases (with differing levels of curation), disparate files within research groups, non-
curated literature publications. In Phase I, Collaborations Pharmaceuticals Inc. developed a prototype of Assay
Central software and used this with a wide variety of structure activity data from sources both public and private,
formatted and unformatted, for enabling neglected, rare or common disease targets. Public data was mixed with
collaborator/customer-contributed data, using original software and applied chemistry judgment of an expert
team. In Phase I we created error checking and correction software. We also built and validated Bayesian models
with the datasets that were collected and cleaned. And, in addition, we developed new data visualization tools.
The software environment that we created readily enables the user to compile structure-activity data for building
computational models and can be used to create selections of these models for sharing with collaborators as
needed. This software can in turn be used for scoring new molecules and visualizing the multiple outputs in
various formats. We have enabled ~14 collaborative projects which have shared models on specific targets such
as PyrG for Tuberculosis (identifying a lead compound), HIV reverse transcriptase, whole cell screening for
Leishmaniasis as well as P450 and nuclear receptor models (e.g. estrogen receptor) relevant to toxicology. We
have utilized Assay Central in our ongoing internal projects working on Ebola, HIV and tuberculosis small
molecule drug discovery.
In Phase II, we propose the following aims that will enable us to develop Assay Central into a production tool
for enabling drug discovery collaborations which we will continue to focus on. In Phase 1 we performed a
preliminary analysis of different machine learning algorithms with select drug discovery datasets. In Phase II we
will now perform a thorough evaluation and selection of additional machine learning algorithms and molecular
descriptors as well as assessment of combination of algorithms (e.g. Bayesian and Deep Learning). We will
implement disease/target definitions for machine learning models to facilitate drug discovery. We will enable
molecule selection and automated design and optimization. The utility of having such a tool as Assay Central
readily available will empower scientists to leverage public, private or a combination of data to help with their
drug discovery tasks. Developing this software suite of computational models with public data will enable us to
identify foundations, academics and potential collaborators that generate preliminary data to test models. These
efforts will dramatically increase the number of projects we can work on, create new IP, and generate
employment using machine learning focused on drug discovery in the area of rare and neglected diseases, in
particular. Assay Central benefits include 1. Ease of deployment and use with a Java file executed by users
without the need for IT support; 2. Built on industry standard technologies; 3. Graphical display of models
provides instant feedback; 4 Model applicability with multiple methods to assess scores and graphics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Preclinical development of a Nipah Virus inhibitor
-
批准号:10761349
-
项目类别:
-
资助金额:$29.51万
-
财政年份:2023
-
负责人:SEAN EKINS
-
依托单位:
New therapeutic approaches to identifying molecules for opioid abuse treatment
-
批准号:10385998
-
项目类别:
-
资助金额:$25.62万
-
财政年份:2022
-
负责人:SEAN EKINS
-
依托单位:
Machine learning approaches to predict Acetylcholinesterase inhibition
-
批准号:10378934
-
项目类别:
-
资助金额:$25.64万
-
财政年份:2021
-
负责人:SEAN EKINS
-
依托单位:
MegaTox for analyzing and visualizing data across different screening systems
-
批准号:10094026
-
项目类别:
-
资助金额:$12.49万
-
财政年份:2020
-
负责人:SEAN EKINS
-
依托单位:
MegaTox for analyzing and visualizing data across different screening systems
-
批准号:10470050
-
项目类别:
-
资助金额:$85.5万
-
财政年份:2019
-
负责人:SEAN EKINS
-
依托单位:
MegaTox for analyzing and visualizing data across different screening systems
-
批准号:10674729
-
项目类别:
-
资助金额:$85.5万
-
财政年份:2019
-
负责人:SEAN EKINS
-
依托单位:
MegaTrans – human transporter machine learning models
-
批准号:9768844
-
项目类别:
-
资助金额:$21.07万
-
财政年份:2019
-
负责人:SEAN EKINS
-
依托单位:
MegaPredict for predicting natural product uses and their drug interactions
-
批准号:10055938
-
项目类别:
-
资助金额:$15.57万
-
财政年份:2019
-
负责人:SEAN EKINS
-
依托单位:
Manufacture of an intracerebroventricular Enzyme Replacement Therapy for CLN1 Batten Disease
-
批准号:10483470
-
项目类别:
-
资助金额:$149.99万
-
财政年份:2018
-
负责人:SEAN EKINS
-
依托单位:
Manufacture of an intracerebroventricular Enzyme Replacement Therapy for CLN1 Batten Disease
-
批准号:10641950
-
项目类别:
-
资助金额:$149.99万
-
财政年份:2018
-
负责人:SEAN EKINS
-
依托单位:
Centralized assay datasets for modelling support of small drug discovery organizations
-
批准号:10474479
-
项目类别:
-
资助金额:$85.47万
-
财政年份:2017
-
负责人:SEAN EKINS
-
依托单位:
Centralized assay datasets for modelling support of small drug discovery organizations
-
批准号:10321747
-
项目类别:
-
资助金额:$85.51万
-
财政年份:2017
-
负责人:SEAN EKINS
-
依托单位:
Centralized assay datasets for modelling support of small drug discovery organizations
-
批准号:9619615
-
项目类别:
-
资助金额:$69.42万
-
财政年份:2017
-
负责人:SEAN EKINS
-
依托单位:
Development and validation of therapy for mucopolysaccharidosis III
-
批准号:9140990
-
项目类别:
-
资助金额:$71.72万
-
财政年份:2016
-
负责人:SEAN EKINS
-
依托单位:
Development and validation of therapy for mucopolysaccharidosis III
-
批准号:9277252
-
项目类别:
-
资助金额:$77.78万
-
财政年份:2016
-
负责人:SEAN EKINS
-
依托单位:
Development and in vitro validation of therapy for mucopolysaccharidosis III
-
批准号:9020728
-
项目类别:
-
资助金额:$0.15万
-
财政年份:2015
-
负责人:SEAN EKINS
-
依托单位:
Development and in vitro validation of therapy for mucopolysaccharidosis III
-
批准号:8764233
-
项目类别:
-
资助金额:$22.31万
-
财政年份:2014
-
负责人:SEAN EKINS
-
依托单位:
Biocomputation across distributed private datasets to enhance drug discovery
-
批准号:8590552
-
项目类别:
-
资助金额:$46.29万
-
财政年份:2013
-
负责人:SEAN EKINS
-
依托单位:
Biocomputation across distributed private datasets to enhance drug discovery
-
批准号:8722061
-
项目类别:
-
资助金额:$56.22万
-
财政年份:2013
-
负责人:SEAN EKINS
-
依托单位:
Identification and Validation of Targets of Phenotypic High Throughput Screening
-
批准号:8590102
-
项目类别:
-
资助金额:$25.52万
-
财政年份:2013
-
负责人:SEAN EKINS
-
依托单位:
海外基金