Discovering Lexical Generalisations. A Supervised Machine Learning Approach to Inheritance Hierarchy Construction

Discovering Lexical Generalisations. A Supervised Machine Learning Approach to Inheritance Hierarchy Construction
复制标题

发现词汇概括。

DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
C. Sporleder
C. Sporleder
中科院分区:
--
文献类型:
--
作者:
C. Sporleder

文献摘要

被引文献

相似文献

语法的发展在过去的几十年里已经看到了从大量的语法规则库到更丰富的词汇结构的转变。许多现代语法理论是高度词汇化的。但是简单地列出词汇条目通常会导致不必要的冗余量。另一方面,词汇继承层次结构使得捕捉语言概括从而减少冗余成为可能。继承层次结构通常是手工构造的,但这很耗时,而且如果词典非常大,这通常是不切实际的。自动或半自动地构造层次结构有助于对词汇数据进行更系统的分析。此外,词汇数据通常是从语料库中自动提取的,这在未来几年可能会增加。因此,更进一步,自动化词汇数据的层次组织也是有意义的。以前的自动词法继承层次结构构建方法往往侧重于最小化标准,目标是最小化一个或多个标准的层次结构,如路径值对的数量,节点的数量或继承链接的数量(Petersen 2001,Barg 1996 a,以及在一个稍微不同的上下文中:Light 1994)。继承层次结构的简洁性是使用它们的主要原因,这一事实激发了最小化的目标。然而,我认为基于最小值的方法存在几个问题。首先,在词汇继承层次的背景下,最小性没有很好地定义,因为不同的最小性标准之间存在张力。其次,基于最小化的方法往往低估了语言可解释性的重要性。虽然这些方法从最小冗余的定义开始,然后试图证明这会导致合理的层次结构,但这里建议的方法却采取了相反的方向。它从一个手动构建的层次结构开始,对其应用有监督的机器学习算法,目的是找到一组可以指导合理层次结构构建的正式标准。采取这个方向意味着更有可能的是,所选择的标准实际上确实导致了看似合理的层次结构。使用机器学习技术还有一个优点,即标准集可以比手工定义的标准大得多。因此,人们可以从非常广泛的角度来定义简洁性,考虑到数据中的相互依赖性以及简单的最小标准。这导致了一个更细粒度的层次结构质量模型。在实践中,这里提出的方法由两个部分组成:伽罗瓦格用于定义搜索空间作为输入词典上的所有概括的集合。然后将在手动构建的层次结构上训练的最大熵模型应用于
Grammar development over the last decades has seen a shift away from large inventories of grammar rules to richer lexical structures. Many modern grammar theories are highly lexicalised. But simply listing lexical entries typically results in an undesirable amount of redundancy. Lexical inheritance hierarchies, on the other hand, make it possible to capture linguistic generalisations and thereby reduce redundancy. Inheritance hierarchies are usually constructed by hand but this is time-consuming and often impractical if a lexicon is very large. Constructing hierarchies automatically or semiautomatically facilitates a more systematic analysis of the lexical data. In addition, lexical data is often extracted automatically from corpora and this is likely to increase over the coming years. Therefore it makes sense to go a step further and automate the hierarchical organisation of lexical data too. Previous approaches to automatic lexical inheritance hierarchy construction tended to focus on minimality criteria, aiming for hierarchies that minimised one or more criteria such as the number of path-value pairs, the number of nodes or the number of inheritance links (Petersen 2001, Barg 1996a, and in a slightly different context: Light 1994). Aiming for minimality is motivated by the fact that the conciseness of inheritance hierarchies is a main reason for their use. However, I will argue that there are several problems with minimality-based approaches. First, minimality is not well defined in the context of lexical inheritance hierarchies as there is a tension between different minimality criteria. Second, minimality-based approaches tend to underestimate the importance of linguistic plausibility. While such approaches start with a definition of minimal redundancy and then try to prove that this leads to plausible hierarchies, the approach suggested here takes the opposite direction. It starts with a manually built hierarchy to which a supervised machine learning algorithm is applied with the aim of finding a set of formal criteria that can guide the construction of plausible hierarchies. Taking this direction means that it is more likely that the selected criteria do in fact lead to plausible hierarchies. Using a machine learning technique also has the advantage that the set of criteria can be much larger than in hand-crafted definitions. Consequently, one can define conciseness in very broad terms, taking into account interdependencies in the data as well as simple minimality criteria. This leads to a more fine-grained model of hierarchy quality. In practice, the method proposed here consists of two components: Galois lattices are used to define the search space as the set of all generalisations over the input lexicon. Maximum entropy models which have been trained on a manually built hierarchy are then applied to the