Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
批准号:
8448471
负责人:
HONGFANG LIU
金额:
$61.4万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-01 至 2013-09-29
中文摘要
我们研究的长期目标是开发资源和自然语言处理(NLP)
生物医学领域的知识管理系统。作为生物医学数据存储在不同的
资源在规模和复杂性方面都经历了非常快速的增长,基于本体的知识管理
变得越来越流行,因为它提供了对生物医学实体的明确描述和一种方法
对生物医学研究成果进行注释和分析。许多信息和知识与以下内容相关
生物医学研究仍然是以自由文本格式记录的。在过去的十年中,NLP被证明具有
加速生物医学知识管理进程的潜力。NLP中的一个关键组件
系统正在识别基因/蛋白质名称(即,基因/蛋白质名称识别),并将它们标准化为
标准表示法(即基因/蛋白质名称标准化)。基因/蛋白质名称识别已被
处理效果良好,但基因/蛋白质名称标准化往往是具有挑战性的。首先,有一个
缺乏基因/蛋白质名称的标准表示法。研究人员使用了结构化数据库,如
蛋白质数据库、UniProtKB或基因资源Entrez gene作为名称参考。但这是有问题的
将名称与这些数据库中的各个记录相关联,因为文本中的名称可以是通用的,并引用
一组记录。此外,与疾病或实验室程序等其他生物医学概念一样,基因或
蛋白质通常以其名称或描述的缩写形式出现在文本中。流行的用法
短格式的模糊性很高,这是自然语言处理应用面临的另一个挑战。
具体地说,拟议的研究旨在:
1)通过丰富现有的基因/蛋白质主题词表--生物主题词表,开发本体-生物主题词表
基因/蛋白质相关的本体论。假设:使基因/蛋白质名称与基因/蛋白质相关的本体论相一致
能否i)检测系统的歧义,ii)在基因/蛋白质命名实体标记期间实现自动推理,以及
3)促进基于本体的知识管理;
2)通过从在线资源和文本中获取简短的知识来增强Onto-BioThesaurus。
假设:获取同义词,特别是基因/蛋白质缩写,对于解决歧义至关重要,
同义词,基因/蛋白质名称规范化的新颖性问题;
3)使用Onto-BioThesaurus规范基因/蛋白质名称。假设:有几个优点(即,
降低歧义,处理新颖性,并将基因/蛋白质概念与生物医学本体联系起来)
使用Onto-BioThesaurus时的传统基因/蛋白质名称规范化,我们希望得到改进
执行各种查找和消除歧义的方法;以及
4)评价研究方法,发布研究成果。假设:评估研究方法
向公众发布研究成果是推进基础和应用生物医学科学的关键。
英文摘要
The long-term goal of our research is to develop resources and natural language processing (NLP)
systems for knowledge management in the biomedical domain. As biomedical data stored in disparate
resources undergo a very rapid growth in both scale and complexity, ontology-based knowledge management
is becoming increasingly popular since it provides explicit descriptions of biomedical entities and an approach
to annotating and analyzing the results of biomedical research. Much of information and knowledge relevant to
biomedical research is still recorded in free text format. In the past decade, NLP has been shown to have the
potential to accelerate the biomedical knowledge management process. One critical component in NLP
systems is identifying gene/protein names (i.e., gene/protein name identification) and normalizing them to
standard representations (i.e., gene/protein name normalization). Gene/protein name identification has been
tackled with good performance but gene/protein name normalization tends to be challenging. First, there is a
lack of standard representations for gene/protein names. Researchers have used structured databases such as
protein database, UniProtKB, or gene resource Entrez Gene as the reference for names. But it is problematic to
associate names to individual records in those databases since a name in text can be generic and refer to a
group of records. Additionally, like other biomedical concepts such as diseases or lab procedures, genes or
proteins usually appear in text as short forms abbreviated from their names or descriptions. The prevalent use
of short forms is another challenge faced by NLP applications because of very high ambiguity of short forms.
Specifically, the proposed research aims to:
1) develop onto-BioThesaurus by enriching BioThesaurus, an existing gene/protein thesaurus, with
gene/protein-related ontologies. Hypothesis: aligning gene/protein names to gene/protein-related ontologies
can i) detect systematic ambiguity, ii) enable automatic reasoning during gene/protein named entity tagging, and
iii) facilitate ontology-based knowledge management;
2) enhance onto-BioThesaurus by harvesting short form knowledge from online resources and text.
Hypothesis: harvesting synonyms especially gene/protein short forms is critical for resolving the ambiguity,
synonymy, and novelty problem for gene/protein name normalization;
3) normalize gene/protein names using onto-BioThesaurus. Hypothesis: there are several advantages (i.e.,
lowering ambiguity, handling novelty, and linking gene/protein concepts to biomedical ontologies) over the
traditional gene/protein name normalization when using onto-BioThesaurus and we expect improved
performance of various lookup and disambiguation methods; and
4) evaluate research methods and distribute research outcome. Hypothesis: evaluating research methods
and distributing research outcome to public are critical to advance basic and applied biomedical science.
期刊论文(41)
专著(0)
科研奖励(0)
会议论文
Learning Precision Medicine for Rare Diseases Empowered by Knowledge-driven Data Mining
-
批准号:10732934
-
项目类别:
-
资助金额:$72.37万
-
财政年份:2023
-
负责人:HONGFANG LIU
-
依托单位:
The Data, Evaluation, and Coordination Center (DECC) for Connecting Underrepresented Populations to Clinical Trials (CUSP2CT)
-
批准号:10597291
-
项目类别:
-
资助金额:$55.44万
-
财政年份:2022
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10202598
-
项目类别:
-
资助金额:$63.08万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10001498
-
项目类别:
-
资助金额:$64.37万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:9251814
-
项目类别:
-
资助金额:$30.0万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10471838
-
项目类别:
-
资助金额:$64.37万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Semi-structured Information Retrieval in Clinical Text for Cohort Identification
-
批准号:8928647
-
项目类别:
-
资助金额:$37.63万
-
财政年份:2014
-
负责人:HONGFANG LIU
-
依托单位:
Semi-structured Information Retrieval in Clinical Text for Cohort Identification
-
批准号:8811565
-
项目类别:
-
资助金额:$46.07万
-
财政年份:2014
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:9033918
-
项目类别:
-
资助金额:$56.28万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8640959
-
项目类别:
-
资助金额:$58.01万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8920720
-
项目类别:
-
资助金额:$16.0万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8505753
-
项目类别:
-
资助金额:$63.07万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8826771
-
项目类别:
-
资助金额:$57.16万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
-
批准号:7654995
-
项目类别:
-
资助金额:$60.87万
-
财政年份:2009
-
负责人:HONGFANG LIU
-
依托单位:
海外基金