课题基金 / 基金详情

MegaPredict for predicting natural product uses and their drug interactions

MegaPredict for predicting natural product uses and their drug interactions
MegaPredict 用于预测天然产物用途及其药物相互作用
批准号:
10055938
负责人:
SEAN EKINS
金额:
$15.57万
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-08-15 至 2021-08-14

项目摘要

项目成果

SEAN EKINS的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 MegaPredict的目标是使科学家能够对一种天然产品(或任何 分子),并确定疗效评估的目标,以及确定任何潜在的责任。我们正在建造 在我们之前的工作中,我们为一个结构-活动数据汇编了一个全面的数据集 各种各样的疾病靶点和其他属性,以一种准备好建立模型的形式。所有这些型号 利用多种来源的精选开放数据,包括ChEMBL、ToxCast等,我们开发了一个原型 MegaPredict使用贝叶斯算法和ECFP6指纹来输出按优先顺序排列的“目标”列表。我们 认识到算法或描述符都可能是最优的,因此我们建议在我们 验证MegaPredict并在此计划书基础上开发产品。我们的团队有合适的资格开发 所需的软件,我们将利用我们庞大的合作伙伴网络来帮助我们验证 化合物。 我们将首先创建一个脚本,将一个天然产品与数千台机器进行比对 然后,学习模型对输出进行排序,以提出疗效目标。我们将使用超过12,000个ChEMBL 从ChEMBLv24数据库中提取的衍生靶标分析/生物活性基团以及EPA Tox21 测量和其他公共数据集,使用我们已经部分开发的方法。我们可以的 对200多种已发表的化合物重复这一过程,并与已知的结果进行比较。我们打算 以比较该方法在合成药物或类药物化合物以及天然产品中的表现。 我们将评估其他机器学习算法和分子描述符是否可以改进 预测。当我们生成机器学习模型时,如线性Logistic回归、AdaBoost决策 我们将评估不同深度的树、随机森林、支持向量机和深度神经网络(DNN) 对天然产物的预测,并与贝叶斯方法进行比较。我们将把ECFP6与 其他2D、3D描述符和物理化学性质,以确定 生成对天然产品的预测,并比较这对合成化合物有何不同。 我们将验证我们对天然产品功效评估的预测。我们将与多家公司密切合作 学术团体对至少20种感兴趣的天然产品进行预测,而不是20多种不同的产品 目标或疾病。我们的目标将是识别以前未知的潜在目标,然后生成 在内部或与学术合作者合作的体外数据。 开发用于结构输入、处理输入分子和输出的原型用户界面 确定目标和责任的优先顺序。我们已经开发了多种软件原型(例如,分析中心、MegaTox、 等)并将确保用户友好的界面,并开发新的可视化方法和算法 用于根据数千个机器学习模型的输出对潜在预测目标进行优先排序。 在第一阶段,我们将在内部与协作者一起使用该软件,以快速制作原型。在第二阶段,我们将开发 一个商业产品,并通过建立一个更大的学术和行业网络来极大地扩展我们的验证 将有助于确定最相关特征的优先顺序的合作伙伴。使用我们使用的机器学习模型 对天然产品的识别是有限的,因为ECFP6指纹不能区分这些非常不同的 分子的类别。但这为我们提供了一个机会,可以采用“药效团”式的方法。 (理想情况下不直接使用3D构象)。因此,我们将专注于开发一种基于3D形状的 或者开发一种新的“2D指纹”,捕捉天然分子和类药物分子的“3D形状”。 目前,ChEMBL和PubChem等中的公共数据集主要由类药物分子组成,但如果 我们有指纹可以比较类似药物的分子和类似天然产品的分子,那么我们可能就可以可靠地 我们的MegaPredict模型也适用于天然产品。我们也可以尝试将天然产品与我们的 ChEMBL模型,或者我们可以使用源自天然药物的模型来查看类药物化合物的目录 产品。这将是一项重要的创新。此外,在第二阶段,重要的是看看我们是否可以 找到天然产品对7000种罕见疾病的用途。开发可预测潜力的软件 天然产物药物与不同靶点的相互作用可能对监管组织以及 并可能扩大能够更有效地混合天然产品和类药物的用途 模型中的化合物将对化学信息学在这一领域的价值产生深远的影响。
英文摘要
Project Summary The objective of ‘MegaPredict’ is to enable scientists to generate predictions for a natural product (or any molecule) and identify targets for efficacy assessment as well as identify any potential liabilities. We are building on our previous work which has compiled a comprehensive collection of datasets for structure-activity data for a broad variety of disease targets and other properties, in a form ready for model building. All of these models utilize the many sources of curated open data, including ChEMBL, ToxCast etc. We have developed a prototype of MegaPredict that utilizes Bayesian algorithm and ECFP6 fingerprints to output a list of prioritized ‘targets’. We realize that neither the algorithm or the descriptors may be optimal therefore we propose to address this as we validate MegaPredict and develop a product over this proposal. Our team is suitably qualified to develop the software needed and we will leverage our large collaborator network to assist us in validating the activity of compounds. We will initially create a script to take a natural product and score it against many thousands of machine learning models then rank the outputs to propose efficacy targets. We will use over 12,000 ChEMBL derived target-assay / bioactivity groups extracted from the ChEMBL v24 database, as well as EPA Tox21 measurements and other public datasets, using methodology that we have already partially developed. We can repeat this process for over 200 published compounds and access the outputs versus what is known. We intend to compare how the approach performs with synthetic drugs or drug-like compounds as well as natural products. We will assess whether other machine learning algorithms and molecular descriptors can improve predictions. As we generate machine learning models such as Linear Logistic Regression, AdaBoost Decision Tree, Random Forest, Support Vector Machine and deep neural networks (DNN) of varying depth we will assess the predictions for natural products and compare with the Bayesian approach. We will compare ECFP6 with other 2D, 3D descriptors and physicochemical properties in order to identify the optimal combination for generating predictions for natural products and compare how this differs for synthetic compounds. We will validate our predictions for natural product efficacy assessment. We will work closely with multiple academic groups to generate predictions for at least 20 natural products of interest against over 20 different targets or diseases. Our goal will be to identify potential targets that were previously unknown and then generate in vitro data inhouse or with academic collaborators. Develop a prototype user interface for input of a structure, processing an input molecule and output of prioritized targets and liabilities. We have developed multiple software prototypes (e. Assay Central, MegaTox, etc.) previously and will ensure a user-friendly interface and develop new visualization methods and algorithms for prioritizing potential predicted targets based on the outputs of thousands of machine learning models. In Phase I, we will use the software internally with collaborators to rapidly prototype it. In Phase II we will develop a commercial product, and greatly expand our validation by building a larger network of academic and industry partners that would help to prioritize features of most relevance. Using the machine learning models which we have for natural products is limited because ECFP6 fingerprints cannot distinguish between these very different classes of molecules. But this provides us with an opportunity by going for a "pharmacophore" style approach (ideally without using 3D conformations directly). We will therefore focus on developing a ‘3D shape-based fingerprint’ or developing a novel ‘2D fingerprint’ that captures the ‘3D shape’ for natural- and druglike molecules. Currently, the public datasets in ChEMBL and PubChem etc. are made up of mostly druglike molecules, but if we have fingerprints that can compare drug-like and natural product-like molecules then we can likely reliably use our MegaPredict models for natural products as well. We can also attempt to rank natural products with our ChEMBL models or we can look through catalogs of druglike compounds using models derived from natural products. That would be an important innovation. Additionally, in Phase II it would be important to see if we could find uses for natural products with any of the 7000 rare diseases. Developing software that predicts potential natural product drug interactions with various targets could be useful to regulatory organizations as well as the pharmaceutical industry and may broaden utility of being able to more effectively mix natural product and druglike compounds in models will have a profound effect on the value of cheminformatics in this arena.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Preclinical development of a Nipah Virus inhibitor
New therapeutic approaches to identifying molecules for opioid abuse treatment
Machine learning approaches to predict Acetylcholinesterase inhibition
MegaTox for analyzing and visualizing data across different screening systems
海外基金