Terminologies for text-mining;: an experiment in the lipoprotein metabolism domain

Terminologies for text-mining;: an experiment in the lipoprotein metabolism domain
复制标题

DOI:
10.1186/1471-2105-9-s4-s2
复制
发表时间:
2008-01-01
期刊:
影响因子:
3
通讯作者:
Schroeder, Michael
Schroeder, Michael
中科院分区:
生物学4区
文献类型:
--
作者:
Alexopoulou, Dimitra;Waechter, Thomas;Schroeder, Michael

文献摘要

被引文献

相似文献

背景:本体工程,尤其是着眼于文本挖掘的应用,仍然是一个新的研究领域。目前还没有一个定义良好的本体构建理论和技术。许多本体设计步骤仍然是手动的,并且基于个人经验和直觉。然而,以术语及其之间关系的提取列表的形式自动构建本体的工作很少。结果:我们分享了手动开发用于文本挖掘的脂蛋白代谢本体(LMO)的经验。我们将人工创建的本体术语与从四种不同的自动术语识别(ATR)方法中自动派生的术语进行了比较。排名前50位的预测词汇包含高达89%的相关词汇。对于前1000个术语,最佳方法仍会生成51%的相关术语。在一个3066个文档的语料库中,53%的LMO术语被包含,38%的术语可以用其中的一种方法生成。结论:自动化方法具有较高的准确率,有助于缩短开发时间,并为领域特定词汇的识别提供重要支持。领域词汇表的覆盖面在很大程度上依赖于底层文档。文本挖掘的本体开发应该以半自动的方式进行;将自动测试结果作为输入,并遵循我们描述的指导方针。可用性:TFIDF术语识别以Web Service的形式提供,在http://gopubmed4.biotec.tu-dresden.de/IdavollWebService/services/CandidateTermGeneratorService?wsdl.上进行了描述
Background: The engineering of ontologies, especially with a view to a text-mining use, is still a new research field. There does not yet exist a well-defined theory and technology for ontology construction. Many of the ontology design steps remain manual and are based on personal experience and intuition. However, there exist a few efforts on automatic construction of ontologies in the form of extracted lists of terms and relations between them.Results: We share experience acquired during the manual development of a lipoprotein metabolism ontology (LMO) to be used for text-mining. We compare the manually created ontology terms with the automatically derived terminology from four different automatic term recognition (ATR) methods. The top 50 predicted terms contain up to 89% relevant terms. For the top 1000 terms the best method still generates 51% relevant terms. In a corpus of 3066 documents 53% of LMO terms are contained and 38% can be generated with one of the methods.Conclusions: Given high precision, automatic methods can help decrease development time and provide significant support for the identification of domain-specific vocabulary. The coverage of the domain vocabulary depends strongly on the underlying documents. Ontology development for text mining should be performed in a semi-automatic way; taking ATR results as input and following the guidelines we described.Availability: The TFIDF term recognition is available as Web Service, described at http://gopubmed4.biotec.tu-dresden.de/IdavollWebService/services/CandidateTermGeneratorService?wsdl.