Automatic Literature-based Protein Annotation
Automatic Literature-based Protein Annotation
批准号:
7260682
负责人:
XINGHUA LU
金额:
$29.14万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-07-15 至 2010-07-14
关键词:
AlgorithmsAnimal ModelAreaArtsBiologicalBiomedical ResearchBody of uterusCalculiClassificationControlled VocabularyDataDiseaseFutureGenesGoalsGrowthHealthHumanInformation RetrievalInformation TheoryKnowledgeKnowledge acquisitionLanguageLiteratureMachine LearningManualsMapsMethodologyMethodsModelingNamesOntologyOrganismPerformanceProteinsRateReportingResearch PersonnelSemanticsSystemTechniquesTestingTextTrainingUncertaintybasebiomedical informaticsconceptdrug discoverygenome sequencinghuman diseaseimprovedindexingknowledge of resultsmarkov modelnovelnovel strategiesnumb proteinprogramsprotein function
中文摘要
描述(由申请人提供):
蛋白质功能是生物医学研究的基石,是理解生物系统、疾病机制乃至人类健康的基础。几十年的生物医学研究积累了大量的生物医学文献形式的知识。生物医学信息学的一个重要任务是从文献的自由文本中获取和表示知识,并将其转换为计算主体能够理解的语言,以便知识能够被存储、检索和用于知识发现。目前,所有的蛋白质注释都是手动分配的,不幸的是,这是非常劳动密集型的,不能跟上信息增长的步伐。事实上,随着一些模式生物基因组序列的完成,蛋白质的人工注释已经成为生物医学文献中大量蛋白质和爆炸式增长的信息之间的主要瓶颈。在这个应用程序中,我们建议开发的方法,以促进自动注释的蛋白质功能的基础上的功能信息埋在生物医学文献。所提出的方法适应并扩展了最先进的概率语义分析,信息检索和机器学习方法,这些方法可以作为自然语言文本中建模不确定性的原则方法。该项目将为未来的自动注释系统开发算法构建模块,这样,当给出蛋白质的简要描述(例如,蛋白质名称和符号),它将能够检索关于蛋白质的相关文献文章,从文章中提取生物学概念并将概念映射到受控词汇表。我们设想,实现这些目标将导致更广泛的影响,不仅有利于自动蛋白质注释,而且生物医学文献索引的生物医学信息学的重要领域之一的进步。有效的知识获取和管理将促进有关疾病机制和药物发现的生物医学研究。
英文摘要
DESCRIPTION (provided by applicant):
Knowledge of protein function serves as a corner stone for biomedical research, which is fundamental for understanding biologic systems, the mechanism of disease and ultimately the human health. Decades of biomedical research has accumulated a great wealth of such knowledge available in the form of biomedical literatures. An important task of biomedical informatics is to acquire and represent the knowledge from free text of literatures and transform it to languages that are understandable by computational agents, so that the knowledge can be stored, retrieved and used for knowledge discovery. Currently, all protein annotations are assigned manually which, unfortunately, is extremely labor-intense and cannot keep up the pace of the growth of information. Indeed, with the completion of genome sequences of several model organisms, manual annotation of proteins has already become a major bottleneck between large number of proteins and exploding amount information in biomedical literatures. In this application, we propose to develop methods to facilitate automatic annotation of protein functions based on the functional information buried in the biomedical literature. The proposed methods adapt and extend the state of art probabilistic semantic analysis, information retrieval and machine learning methodologies, which serve as principled approaches to modeling uncertainties in natural language text. The project will develop algorithmic building blocks for a future automatic annotation system such that, when given a brief description of a protein (e.g., a protein name and symbol), it will be capable of retrieving relevant literature articles about the protein, extracting biological concepts from the articles and mapping the concept to a controlled vocabulary. We envision that achieving these goals will result in advances with broader impact which not only facilitate automatic protein annotation but also for biomedical literature indexing-one of the important area of biomedical informatics. The efficient knowledge acquisition and management will enhance biomedical research regarding the mechanisms of diseases and drug discovery.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Interpretable deep learning models for translational medicine
-
批准号:10579895
-
项目类别:
-
资助金额:$31.37万
-
财政年份:2015
-
负责人:XINGHUA LU
-
依托单位:
Interpretable deep learning models for translational medicine
-
批准号:10371139
-
项目类别:
-
资助金额:$31.27万
-
财政年份:2015
-
负责人:XINGHUA LU
-
依托单位:
Interpretable deep learning models for translational medicine
-
批准号:10171908
-
项目类别:
-
资助金额:$30.94万
-
财政年份:2015
-
负责人:XINGHUA LU
-
依托单位:
Deciphering cellular signaling system by deep mining a comprehensive genomic compendium
-
批准号:9042426
-
项目类别:
-
资助金额:$32.82万
-
财政年份:2015
-
负责人:XINGHUA LU
-
依托单位:
Ontology-Driven Methods for Knowledge Acquisition and Knowledge Discovery
-
批准号:8202896
-
项目类别:
-
资助金额:$31.26万
-
财政年份:2011
-
负责人:XINGHUA LU
-
依托单位:
Ontology-Driven Methods for Knowledge Acquisition and Knowledge Discovery
-
批准号:8714053
-
项目类别:
-
资助金额:$30.75万
-
财政年份:2011
-
负责人:XINGHUA LU
-
依托单位:
Ontology-Driven Methods for Knowledge Acquisition and Knowledge Discovery
-
批准号:8326650
-
项目类别:
-
资助金额:$31.72万
-
财政年份:2011
-
负责人:XINGHUA LU
-
依托单位:
Statistical methods for integromics discoveries
-
批准号:8332877
-
项目类别:
-
资助金额:$31.3万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
MODELING ROLES OF BIOACTIVE LIPIDS IN GENE EXPRESSION SYSTEMS
-
批准号:7959967
-
项目类别:
-
资助金额:$14.6万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
Statistical methods for integromics discoveries
-
批准号:7740132
-
项目类别:
-
资助金额:$31.8万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
Statistical methods for integromics discoveries
-
批准号:8131525
-
项目类别:
-
资助金额:$32.62万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:7906366
-
项目类别:
-
资助金额:$12.36万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
Statistical methods for integromics discoveries
-
批准号:7921473
-
项目类别:
-
资助金额:$31.32万
-
财政年份:2009
-
负责人:XINGHUA LU
-
依托单位:
MODELING ROLES OF BIOACTIVE LIPIDS IN GENE EXPRESSION SYSTEMS
-
批准号:7720848
-
项目类别:
-
资助金额:$21.46万
-
财政年份:2008
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:7840891
-
项目类别:
-
资助金额:$3.93万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:7662449
-
项目类别:
-
资助金额:$16.73万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:8133305
-
项目类别:
-
资助金额:$10.99万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
Automatic Literature-based Protein Annotation
-
批准号:8151670
-
项目类别:
-
资助金额:$1.48万
-
财政年份:2007
-
负责人:XINGHUA LU
-
依托单位:
海外基金