Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
批准号:
7654995
负责人:
HONGFANG LIU
金额:
$60.87万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-01 至 2011-08-31
关键词:
AbbreviationsBiomedical ResearchCharacteristicsCommunitiesComputersDatabasesExpert OpinionFoundationsGene ProteinsGenesGoalsHarvestHumanInformation ResourcesInformation Resources ManagementInternetInvestigationKnowledgeLiteratureManualsMapsNamesNatural Language ProcessingOnline SystemsOntologyOutcomePeer ReviewProcessProteinsPublishingRecordsResearchResearch MethodologyResearch PersonnelResourcesRetrievalReview LiteratureScienceServicesSystemTechniquesTerminologyTextThesauriTimeacronymsbasebiomedical ontologyheuristicsknowledge basetoolweb interfaceweb site
中文摘要
描述(由申请人提供):
我们研究的长期目标是为生物医学领域的知识检索管理开发资源和工具。随着生物医学研究步伐的加快,研究人员越来越依赖计算机来管理正在发布的爆炸性数量的生物医学信息。许多数据库的高质量是由数据库馆长保证的,他们提取和合成存储在文献或其他数据库中的信息。重要的是要准确地识别文本中的生物医学实体名称,并将识别的名称映射到生物医学数据库中的相应记录。通常,生物医学数据库提供由馆长输入或从其他数据库中提取的姓名列表。NLP系统可使用这些名称从数据库中检索记录或将名称映射到数据库记录。然而,存在与生物医学实体名称相关联的几个特征,即:同义性(即,不同的名称指的是同一数据库条目)、歧义(即,一个名称与不同的条目相关联)和新颖性(即,名称或实体不存在于数据库或知识库中),这使得使用名称检索数据库记录的任务以及将文本中的名称与数据库记录相关联的任务非常艰巨。此外,生物医学实体可以在文本中显示为从其长形式(LF)缩写而来的短形式(SF)。由于SFS的高度模糊性,代表生物医学实体的SFS的普遍使用是最终用户和NLP应用程序面临的另一个挑战。
最近,基于本体的知识管理变得越来越流行,因为本体提供了生物医学实体及其关系的形式化、机器可处理和人类可解释的表示。我们假设,生物医学本体论可以用来降低使用名称检索记录或将文本中的名称映射到数据库记录的难度。具体的目标和相应的假设是:i)通过用与基因/蛋白质相关的本体丰富《生物叙词表》来开发本体-生物叙词表(假设:将基因/蛋白质名称与基因/蛋白质相关本体对齐可以降低与基因/蛋白质名称相关的复杂性);ii)从在线资源和文本中获取基因/蛋白质类别和实体的同义词(假设:获取同义词尤其是基因/蛋白质SFS是至关重要的,因为SFs经常用于表示基因/蛋白质实体);Iii)建立一个网络用户界面,用于通过支持本体的Onto-BioThesaurus搜索和查询基因/蛋白质名称和条目(假设:使用与基因/蛋白质相关的本体增强生物主题词库将使我们能够建立启发式规则以实现机器推理);以及iv)评估和分发研究方法/结果(假设:评估和分发研究方法/结果对于促进基础和应用生物医学科学都是关键的。
英文摘要
DESCRIPTION (provided by applicant):
The long-term goal of our research is to develop resources and tools for knowledge retrieval management in the biomedical domain. As the pace of biomedical research accelerates, researchers become more and more dependent on computers to manage the explosive amount of biomedical information being published. The high quality of many databases is guaranteed by database curators who extract and synthesize information stored in literature or other databases. It is important to accurately recognize biomedical entity names in text and map the identified names to corresponding records in biomedical databases. Usually, a biomedical database provides a list of names either entered by curators or extracted from other databases. Those names could be used to retrieve records from databases or map names to database records by NLP systems. However, there are several characteristics associated with biomedical entity names, namely: synonymy (i.e., different names refer to the same database entry), ambiguity (i.e., one name is associated with different entries), and novelty (i.e., names or entities are not present in databases or knowledge bases) which make the task of retrieving database records using names and the task of associating names in text to database records very daunting. Additionally, biomedical entities can appear in text as short forms (SFs) abbreviated from their long forms (LFs). The prevalent use of SFs representing biomedical entities is another challenge faced by end users and NLP applications because of the high ambiguity of SFs.
Recently, ontology-based knowledge management is becoming increasingly popular since ontologies provide formal, machine-processable, and human-interpretable representations of the biomedical entities and their relations. We hypothesize that biomedical ontologies can be used to reduce the difficulty associated with retrieving records using names or mapping names in text to database records. Specific aims and the corresponding hypotheses are: i) develop onto-BioThesaurus by enriching BioThesaurus with gene/protein-related ontologies (Hypothesis: aligning gene/protein names to gene/protein-related ontologies can reduce the complexity associated with gene/protein names); ii) harvest synonyms for gene/protein classes and entities from online resources and text (Hypothesis: harvesting synonyms especially gene/protein SFs is critical since SFs are frequently used to represent gene/protein entities); iii) build a web user interface for gene/protein names and entries search and query through ontology-enabled onto-BioThesaurus (Hypothesis: enhancing BioThesaurus with gene/protein-related ontologies would enable us to build heuristic rules to enable machine reasoning); and iv) evaluate and distribute research methods/outcome (Hypothesis: evaluating and distributing research methods/outcome are critical to advance both basic and applied biomedical science.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Learning Precision Medicine for Rare Diseases Empowered by Knowledge-driven Data Mining
-
批准号:10732934
-
项目类别:
-
资助金额:$72.37万
-
财政年份:2023
-
负责人:HONGFANG LIU
-
依托单位:
The Data, Evaluation, and Coordination Center (DECC) for Connecting Underrepresented Populations to Clinical Trials (CUSP2CT)
-
批准号:10597291
-
项目类别:
-
资助金额:$55.44万
-
财政年份:2022
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10202598
-
项目类别:
-
资助金额:$63.08万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10001498
-
项目类别:
-
资助金额:$64.37万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:9251814
-
项目类别:
-
资助金额:$30.0万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Secondary use of EMRs for surgical complication surveillance
-
批准号:10471838
-
项目类别:
-
资助金额:$64.37万
-
财政年份:2015
-
负责人:HONGFANG LIU
-
依托单位:
Semi-structured Information Retrieval in Clinical Text for Cohort Identification
-
批准号:8928647
-
项目类别:
-
资助金额:$37.63万
-
财政年份:2014
-
负责人:HONGFANG LIU
-
依托单位:
Semi-structured Information Retrieval in Clinical Text for Cohort Identification
-
批准号:8811565
-
项目类别:
-
资助金额:$46.07万
-
财政年份:2014
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:9033918
-
项目类别:
-
资助金额:$56.28万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8920720
-
项目类别:
-
资助金额:$16.0万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8640959
-
项目类别:
-
资助金额:$58.01万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8505753
-
项目类别:
-
资助金额:$63.07万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Natural language processing for clinical and translational research
-
批准号:8826771
-
项目类别:
-
资助金额:$57.16万
-
财政年份:2013
-
负责人:HONGFANG LIU
-
依托单位:
Onto-BioThesaurus: ontological representation of gene/protein names for biomedica
-
批准号:8448471
-
项目类别:
-
资助金额:$61.4万
-
财政年份:2009
-
负责人:HONGFANG LIU
-
依托单位:
海外基金