Integration and publication of heterogeneous text-mined relationships on the Semantic Web.

Integration and publication of heterogeneous text-mined relationships on the Semantic Web.
复制标题

DOI:
10.1186/2041-1480-2-s2-s10
复制
发表时间:
2011-05-17
影响因子:
1.9
通讯作者:
Shah NH
Shah NH
中科院分区:
工程技术4区
文献类型:
--
作者:
Coulet A;Garten Y;Dumontier M;Altman RB;Musen MA;Shah NH

文献摘要

被引文献

相似文献

自然语言处理(NLP)技术的进步使得能够提取生物医学文本中提到的细粒度关系。自然语言在表达相似关系时的可变性和复杂性导致所提取的关系是高度异构的,这使得知识库的构建变得困难,并且在使用这些知识库进行数据挖掘或问题回答时带来了挑战。我们报告的PHARE关系本体(PHArmacogenomic Relationships Ontology)的半自动化建设,包括200个策划的关系,从超过40,000个异构的关系,通过文本挖掘提取。这些异构关系,然后映射到PHARE本体使用同义词,实体描述和层次结构的实体和角色。一旦映射,就可以使用本体的结构来规范化和比较关系,以识别具有相似语义但不同语法的关系。我们比较和对比的手动程序与一个完全自动化的方法,使用WordNet量化的集成度,使迭代策展和完善的法尔本体。这种集成的结果是一个规范化的生物医学关系的存储库,命名为PHARE-KB,它可以使用语义网技术,如SPARQL查询,并可以在生物网络的形式可视化。PHARE本体作为一个通用的语义框架,整合了40,000多个与药物基因组学相关的关系。PHARE本体构成了名为PHARE-KB的知识库的基础。一旦填充了关系,PHARE-KB(i)可以以生物网络的形式可视化,以指导人类任务,如数据库管理,(ii)可以通过编程方式查询,以指导生物信息学应用,如分子相互作用的预测。法尔方案可在http://purl.bioontology.org/ontology/PHARE上查阅。
Advances in Natural Language Processing (NLP) techniques enable the extraction of fine-grained relationships mentioned in biomedical text. The variability and the complexity of natural language in expressing similar relationships causes the extracted relationships to be highly heterogeneous, which makes the construction of knowledge bases difficult and poses a challenge in using these for data mining or question answering. We report on the semi-automatic construction of the PHARE relationship ontology (the PHArmacogenomic RElationships Ontology) consisting of 200 curated relations from over 40,000 heterogeneous relationships extracted via text-mining. These heterogeneous relations are then mapped to the PHARE ontology using synonyms, entity descriptions and hierarchies of entities and roles. Once mapped, relationships can be normalized and compared using the structure of the ontology to identify relationships that have similar semantics but different syntax. We compare and contrast the manual procedure with a fully automated approach using WordNet to quantify the degree of integration enabled by iterative curation and refinement of the PHARE ontology. The result of such integration is a repository of normalized biomedical relationships, named PHARE-KB, which can be queried using Semantic Web technologies such as SPARQL and can be visualized in the form of a biological network. The PHARE ontology serves as a common semantic framework to integrate more than 40,000 relationships pertinent to pharmacogenomics. The PHARE ontology forms the foundation of a knowledge base named PHARE-KB. Once populated with relationships, PHARE-KB (i) can be visualized in the form of a biological network to guide human tasks such as database curation and (ii) can be queried programmatically to guide bioinformatics applications such as the prediction of molecular interactions. PHARE is available at http://purl.bioontology.org/ontology/PHARE.