Learning concept hierarchies from text corpora using formal concept analysis

Learning concept hierarchies from text corpora using formal concept analysis
复制标题

DOI:
10.1613/jair.1648
复制
发表时间:
2005-01-01
影响因子:
5
通讯作者:
Staab, S
Staab, S
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cimiano, P;Hotho, A;Staab, S

文献摘要

被引文献

相似文献

我们提出了一种从文本语料库中自动获取分类法或概念层次结构的新方法。该方法基于形式概念分析(FCA),这是一种主要用于数据分析的方法,即用于调查和处理明确给定的信息。我们遵循Harris的分布假设,将特定术语的上下文建模为一个向量,表示使用语言解析器从文本语料库中自动获取的句法依赖关系。在此上下文信息的基础上,FCA生成一个格,我们将其转换为构成概念层次的特殊类型的偏序。通过将结果概念层次结构与两个领域(旅游和金融)的手工分类法进行比较,对该方法进行评估。我们还直接将我们的方法与分层聚集聚类以及作为分裂聚类算法实例的Bi-Section-KMeans进行了比较。此外,我们还研究了使用不同的度量来加权每个属性的贡献以及应用特定的平滑技术来处理数据稀疏性的影响。
We present a novel approach to the automatic acquisition of taxonomies or concept hierarchies from a text corpus. The approach is based on Formal Concept Analysis ( FCA), a method mainly used for the analysis of data, i.e. for investigating and processing explicitly given information. We follow Harris' distributional hypothesis and model the context of a certain term as a vector representing syntactic dependencies which are automatically acquired from the text corpus with a linguistic parser. On the basis of this context information, FCA produces a lattice that we convert into a special kind of partial order constituting a concept hierarchy. The approach is evaluated by comparing the resulting concept hierarchies with hand-crafted taxonomies for two domains: tourism and finance. We also directly compare our approach with hierarchical agglomerative clustering as well as with Bi-Section-KMeans as an instance of a divisive clustering algorithm. Furthermore, we investigate the impact of using different measures weighting the contribution of each attribute as well as of applying a particular smoothing technique to cope with data sparseness.