Ontology-Driven Methods for Knowledge Acquisition and Knowledge Discovery
Ontology-Driven Methods for Knowledge Acquisition and Knowledge Discovery
批准号:
8202896
负责人:
XINGHUA LU
金额:
$31.26万
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2015-08-31
关键词:
AccountingAchievementAddressAlgorithmsAreaBiologicalBiological ProcessBiomedical ResearchComputing MethodologiesControlled VocabularyDataDiseaseGene ProteinsGenesGoalsKnowledgeKnowledge DiscoveryKnowledge acquisitionLearningLiteratureMalignant NeoplasmsManualsMapsMethodologyMethodsMiningModelingNamesNatural Language ProcessingOntologyProceduresProcessProteinsPsyche structureResearchScienceSemanticsStructureSystemSystems BiologyTestingTextThe Cancer Genome AtlasTrainingTweensYeastsbiological systemsbiomedical informaticscancer celldata miningdesigninsightinterestknowledge of resultsnovelprotein functionprotein protein interactionresearch studytext searching
中文摘要
描述(由申请人提供):
生物医学信息学领域的一个巨大挑战是开发将现有知识和实验数据相结合的计算方法,以获得关于生物系统和疾病机制的新知识。生物医学文献中关于基因和蛋白质的大多数知识都是以自由文本的形式存储的,这不适合计算,而手动将这些知识编码成可计算的形式的过程跟不上知识积累的速度。这项研究的主要目的是设计新的统计文本挖掘算法,从自由文本文献中获取和表示关于基因和蛋白质的知识,并进一步将获得的知识与实验数据相结合来获得新的知识。我们将组织拟议的研究,以实现以下具体目标。具体目标1.开发本体引导的语义建模算法从自由文本中提取生物概念,其中我们将设计能够将生物概念表示为层次的层次化概率主题模型,并开发新的学习算法来从自由文本文档中推断生物概念。具体目标2.将语义建模与BioNLP相结合,提取支持蛋白质功能注释的文本证据。我们将开发信息提取算法,将层次语义分析的结果与BioNLP相结合,以识别最有可能提供关于基因/蛋白质功能的证据的文本区域,并将提取的信息映射到受控词汇。具体目标3.开发一个框架,统一知识推理和数据挖掘的过程,以进行知识发现。在这个目标中,我们将使用现有的知识(以本体的形式表示)来推理,以从实验数据中揭示基因之间的功能模块。然后,我们将进一步开发算法,通过挖掘系统规模的实验数据来揭示这些基因模块之间的关系。整体框架将以迭代的方式整合功能推理和数据挖掘,以逐步提炼知识并推导出规则,如:当涉及生物过程X的基因受到扰动时,涉及生物过程Y的基因将做出反应。我们将在酵母系统生物学研究和癌症基因组图谱(TCGA)项目的数据上测试该框架,以深入了解癌细胞的细胞系统和疾病机制。
英文摘要
DESCRIPTION (provided by applicant):
A great challenge in the biomedical informatics domain is to develop computational methods that combine existing knowledge and experimental data to derive new knowledge regarding biological systems and disease mechanisms. Most knowledge regarding genes and proteins in biomedical literature is stored in the form of free text that is not suitable for computation, and the manual processes of encoding this body of knowledge into computable form cannot keep up with the rate of knowledge accumulation. The main thrust of the proposed research is to design novel statistical text-mining algorithms to acquire and represent knowledge regarding genes and proteins from free-text literature, and further to combine this acquired knowledge with experimental data to derive new knowledge. We will organize the proposed research to the following specific aims. Specific Aim 1. Develop ontology-guided semantic modeling algorithms for extracting biological concepts from free text, in which we will design hierarchical probabilistic topic models that are capable of representing biological concepts as a hierarchy and develop novel learning algorithms to infer biological concepts from free-text documents. Specific Aim 2. Integrate semantic modeling with BioNLP to extract textual evidence supporting protein-function annotations. We will develop information extraction algorithms that will combine the results of hierarchical semantic analysis and BioNLP to identify the text regions that will most likely provide evidence regarding the function of genes/proteins and map the extracted information to a controlled vocabulary. Specific Aim 3. Develop a framework to unify the procedures of knowledge reasoning and data mining for knowledge discovery. In this aim, we will reason using existing knowledge (represented in the form of an ontology) to reveal functional modules among the genes from the experimental data. We will then further develop algorithms that will reveal relationships between these gene modules by mining system-scaled experimental data. The overall framework will integrate functional reasoning and data mining in an iterative manner to refine the knowledge progressively and to derive rules such as: when genes involved in biological process X are perturbed, genes involved in biological process Y will respond. We will test the framework on the data from yeast-system biology studies and the Cancer Genome Atlas (TCGA) project to gain insights into the cellular systems and disease mechanisms of cancer cells.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Interpretable deep learning models for translational medicine
-
批准号:10579895
-
项目类别:
-
资助金额:$31.37万
-
财政年份:2015
-
负责人:XINGHUA LU
-
依托单位:
Interpretable deep learning models for translational medicine
-
批准号:10371139
-
项目类别:
-
资助金额:$31.27万
-
财政年份:2015
-
负责人:XINGHUA LU
-
依托单位:
Interpretable deep learning models for translational medicine
-
批准号:10171908
-
项目类别:
-
资助金额:$30.94万
-
财政年份:2015
-
负责人:XINGHUA LU
-
依托单位:
Deciphering cellular signaling system by deep mining a comprehensive genomic compendium
-
批准号:9042426
-
项目类别:
-
资助金额:$32.82万
-
财政年份:2015
-
负责人:XINGHUA LU
-
依托单位:
Ontology-Driven Methods for Knowledge Acquisition and Knowledge Discovery
-
批准号:8714053
-
项目类别:
-
资助金额:$30.75万
-
财政年份:2011
-
负责人:XINGHUA LU
-
依托单位:
Ontology-Driven Methods for Knowledge Acquisition and Knowledge Discovery
-
批准号:8326650
-
项目类别:
-
资助金额:$31.72万
-
财政年份:2011
-
负责人:XINGHUA LU
-
依托单位:
Statistical methods for integromics discoveries
-
批准号:8332877
-
项目类别:
-
资助金额:$31.3万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
MODELING ROLES OF BIOACTIVE LIPIDS IN GENE EXPRESSION SYSTEMS
-
批准号:7959967
-
项目类别:
-
资助金额:$14.6万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
Statistical methods for integromics discoveries
-
批准号:7740132
-
项目类别:
-
资助金额:$31.8万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
Statistical methods for integromics discoveries
-
批准号:8131525
-
项目类别:
-
资助金额:$32.62万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:7906366
-
项目类别:
-
资助金额:$12.36万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
Statistical methods for integromics discoveries
-
批准号:7921473
-
项目类别:
-
资助金额:$31.32万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
MODELING ROLES OF BIOACTIVE LIPIDS IN GENE EXPRESSION SYSTEMS
-
批准号:7720848
-
项目类别:
-
资助金额:$21.46万
-
财政年份:2008
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:7840891
-
项目类别:
-
资助金额:$3.93万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:7662449
-
项目类别:
-
资助金额:$16.73万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:8133305
-
项目类别:
-
资助金额:$10.99万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:7260682
-
项目类别:
-
资助金额:$29.14万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:8151670
-
项目类别:
-
资助金额:$1.48万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
海外基金