课题基金 / 基金详情

Onto-BioThesaurus: ontological representation of gene/protein names for biomedica

Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
Onto-BioThesaurus:生物医学基因/蛋白质名称的本体论表示
批准号:
8448471
负责人:
HONGFANG LIU
金额:
$61.4万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-01 至 2013-09-29

项目摘要

项目成果

HONGFANG LIU的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The long-term goal of our research is to develop resources and natural language processing (NLP) systems for knowledge management in the biomedical domain. As biomedical data stored in disparate resources undergo a very rapid growth in both scale and complexity, ontology-based knowledge management is becoming increasingly popular since it provides explicit descriptions of biomedical entities and an approach to annotating and analyzing the results of biomedical research. Much of information and knowledge relevant to biomedical research is still recorded in free text format. In the past decade, NLP has been shown to have the potential to accelerate the biomedical knowledge management process. One critical component in NLP systems is identifying gene/protein names (i.e., gene/protein name identification) and normalizing them to standard representations (i.e., gene/protein name normalization). Gene/protein name identification has been tackled with good performance but gene/protein name normalization tends to be challenging. First, there is a lack of standard representations for gene/protein names. Researchers have used structured databases such as protein database, UniProtKB, or gene resource Entrez Gene as the reference for names. But it is problematic to associate names to individual records in those databases since a name in text can be generic and refer to a group of records. Additionally, like other biomedical concepts such as diseases or lab procedures, genes or proteins usually appear in text as short forms abbreviated from their names or descriptions. The prevalent use of short forms is another challenge faced by NLP applications because of very high ambiguity of short forms. Specifically, the proposed research aims to: 1) develop onto-BioThesaurus by enriching BioThesaurus, an existing gene/protein thesaurus, with gene/protein-related ontologies. Hypothesis: aligning gene/protein names to gene/protein-related ontologies can i) detect systematic ambiguity, ii) enable automatic reasoning during gene/protein named entity tagging, and iii) facilitate ontology-based knowledge management; 2) enhance onto-BioThesaurus by harvesting short form knowledge from online resources and text. Hypothesis: harvesting synonyms especially gene/protein short forms is critical for resolving the ambiguity, synonymy, and novelty problem for gene/protein name normalization; 3) normalize gene/protein names using onto-BioThesaurus. Hypothesis: there are several advantages (i.e., lowering ambiguity, handling novelty, and linking gene/protein concepts to biomedical ontologies) over the traditional gene/protein name normalization when using onto-BioThesaurus and we expect improved performance of various lookup and disambiguation methods; and 4) evaluate research methods and distribute research outcome. Hypothesis: evaluating research methods and distributing research outcome to public are critical to advance basic and applied biomedical science.
期刊论文(41)
专著(0)
科研奖励(0)
会议论文
Learning Precision Medicine for Rare Diseases Empowered by Knowledge-driven Data Mining
The Data, Evaluation, and Coordination Center (DECC) for Connecting Underrepresented Populations to Clinical Trials (CUSP2CT)
  • 批准号:
    10597291
  • 项目类别:
  • 资助金额:
    $55.44万
  • 财政年份:
    2022
  • 负责人:
    HONGFANG LIU
  • 依托单位:
Secondary use of EMRs for surgical complication surveillance
  • 批准号:
    10202598
  • 项目类别:
  • 资助金额:
    $63.08万
  • 财政年份:
    2015
  • 负责人:
    HONGFANG LIU
  • 依托单位:
Secondary use of EMRs for surgical complication surveillance
  • 批准号:
    10001498
  • 项目类别:
  • 资助金额:
    $64.37万
  • 财政年份:
    2015
  • 负责人:
    HONGFANG LIU
  • 依托单位:
海外基金