Illuminating the Druggable Genome by Knowledge Graphs
Illuminating the Druggable Genome by Knowledge Graphs
批准号:
10348825
负责人:
CHRISTOPHER J MUNGALL
金额:
$53.66万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-03-01 至 2022-02-28
关键词:
AddressAlgorithmsAloralAmino AcidsAnimal ModelAntineoplastic AgentsAreaBindingBinding SitesBioinformaticsBiologicalBiological ModelsCancer ModelCatalogsCategoriesClinicalCodeComputer AnalysisComputer softwareDataData SourcesDiseaseDocumentationDrug DesignDrug TargetingEmerging TechnologiesEnzymesFDA approvedFutureGene TargetingGenesGenomeGenomicsGoalsGraphHumanHuman GenomeInformation NetworksInformation Resources ManagementInvestigationKnowledgeLibrariesLinkMachine LearningMedicalMedicineMolecular BiologyOntologyOutcomeOutcomes ResearchPathologyPatternPharmaceutical PreparationsPhenotypePhosphotransferasesPilot ProjectsProcessProtein KinaseProteinsPublic HealthPythonsResearchResourcesScientistSemanticsSignal TransductionSystemThe Jackson LaboratoryTrainingValidationanti-cancerbasecheminformaticscomputer sciencecomputer studiescomputing resourcesdark matterdeep learningdesigndisease phenotypedrug discoverydrug mechanismdrug repurposinggene functiongene therapygenome resourcehigh riskhuman diseaseimprovedinorganic phosphateknowledge baseknowledge graphknowledge integrationlearning algorithmmachine learning algorithmmachine learning methodmouse modelnew therapeutic targetnovelnovel drug classopen sourcepatient derived xenograft modelprotein kinase inhibitorprotein kinase modulatorreal world applicationsmall moleculetoolvalidation studies
中文摘要
项目概要/摘要
在人类基因组的约20,000个蛋白质编码基因中,约有1500个可以结合药物样分子,然而,
目前只有大约600种是FDA批准的药物的目标。因此,至少有930种蛋白质是潜在的药物
这些目标尚未被用于人类医学,鉴于我们对这些目标的不完整了解,
人类基因组中,实际的数字可能会更高。因此,有大量未满足的需求,
提高我们对这种所谓的基因组暗物质的理解,以开发新型药物,
改善疾病治疗。这些蛋白质的全面实验研究的背景下,
成千上万的化合物和成千上万的疾病将是昂贵的,但
计算方法可以显著地改进该列表。在这个项目中,我们将应用两个复杂的
预测最有前途的新型药物靶点的计算方法。我们将整合
知识库DrugCentral和其他资源与疾病和表型知识库的
将Monarch Initiative转换为语义协调的知识图(KG)。这将导致KG具有
全面覆盖疾病、基因、基因功能、表型异常、药物、药物
机制和药物靶点。机器学习(ML)从训练集中识别模式,并应用
模式来预测新数据中的实体和关系。基于知识库的机器学习已经成为一个热门的新的研究领域,
计算机科学,但由于缺乏足够的软件,仍然难以用于现实世界的应用程序
包装件.因此,我们将通过以下方式在KG上实现基于深度学习的最先进的学习算法:
扩展和调整选定的算法,以完成药物和药物靶点发现的任务。我们将开发一个
易于使用的软件库,并通过笔记本电脑演示其使用,笔记本电脑将被设计为
其他科学家未来计算研究的起点,因为它们将包含分析工作流程
沿着有关每个步骤的文档。人类基因组编码500多种蛋白激酶,
是将磷酸基团添加到特定氨基酸残基并由此传递生物信号的酶。
目前有35种FDA批准的蛋白激酶调节剂作用于38种蛋白激酶,因此它们是
是我们基因组编码的最重要的可药用蛋白质之一。我们将进行详细的
这组的计算研究和实验验证我们的顶部,新的候选人使用患者衍生的
异种移植模型系统。
英文摘要
PROJECT SUMMARY / ABSTRACT
About 1500 of the ~20,000 protein-coding genes of the human genome can bind drug-like molecules, and yet
only about 600 are currently targeted by FDA-approved drugs. Therefore, at least 930 proteins are potential drug
targets that are not yet being utilized for human medicine and, given our incomplete state of knowledge about
the human genome, the actual number could be much higher. There is therefore a substantial unmet need to
improve our understanding of this so-called genomic dark matter in order to develop novel classes of drugs to
improve treatment of disease. Comprehensive experimental investigation of these proteins in the context of
hundreds of thousands of compounds and thousands of diseases would be prohibitively expensive, but
computational approaches could significantly refine the list. In this project we will apply two sophisticated
computational approaches to the task of predicting the most promising novel drug targets. We will integrate the
knowledge bases DrugCentral and other resources with the disease and phenotype knowledge base of the
Monarch Initiative into a semantically harmonized knowledge graph (KG). This will result in a KG with
comprehensive coverage of diseases, genes, gene functions, phenotypic abnormalities, drugs, drug
mechanisms, and drug targets. Machine learning (ML) identifies patterns from training sets and applies the
patterns to predict entities and relations in new data. ML using KGs has become a hot new research area in
computer science, but remains difficult to use for real-world applications, owing to the lack of adequate software
packages. We will therefore implement state-of-the art learning algorithms based on deep learning on KGs by
extending and adapting selected algorithms to the task of drug and drug target discovery. We will develop an
easy-to-use software library and demonstrate its use by means of notebooks that will be designed to serve as
starting points for future computational research by other scientists, since they will contain the analysis workflow
along with documentation about each step. The human genome codes more than 500 protein kinases, which
are enzymes that add a phosphate group to specific amino acid residues and thereby transmit a biological signal.
There are currently 35 FDA approved protein kinase modulators acting on 38 protein kinases, which are thus
one of the most important groups of druggable proteins encoded by our genome. We will perform a detailed
computational study of this group and experimentally validate our top, novel candidate using a patient-derived
xenograft model system.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Gene Ontology Consortium and Knowledgebase
-
批准号:10631046
-
项目类别:
-
资助金额:$233.03万
-
财政年份:2022
-
负责人:CHRISTOPHER J MUNGALL
-
依托单位:
Increasing the Yield and Utility of Pediatric Genomic Medicine with Exomiser
-
批准号:10611970
-
项目类别:
-
资助金额:$70.29万
-
财政年份:2021
-
负责人:CHRISTOPHER J MUNGALL
-
依托单位:
Increasing the Yield and Utility of Pediatric Genomic Medicine with Exomiser
-
批准号:10390282
-
项目类别:
-
资助金额:$70.26万
-
财政年份:2021
-
负责人:CHRISTOPHER J MUNGALL
-
依托单位:
Services to support the OBO foundry standards
-
批准号:9385259
-
项目类别:
-
资助金额:$48.45万
-
财政年份:2017
-
负责人:CHRISTOPHER J MUNGALL
-
依托单位:
An Intelligent Concept Agent for Assisting with the Application of Metadata
-
批准号:9161233
-
项目类别:
-
资助金额:$57.55万
-
财政年份:2016
-
负责人:CHRISTOPHER J MUNGALL
-
依托单位:
An Intelligent Concept Agent for Assisting with the Application of Metadata
-
批准号:9357656
-
项目类别:
-
资助金额:$57.28万
-
财政年份:2016
-
负责人:CHRISTOPHER J MUNGALL
-
依托单位:
海外基金