Similarity-Driven Semantic Role Induction via Graph Partitioning

Similarity-Driven Semantic Role Induction via Graph Partitioning
复制标题

DOI:
10.1162/coli_a_00195
复制
发表时间:
2014-09
影响因子:
9.3
通讯作者:
Joel Lang;Mirella Lapata
Joel Lang;Mirella Lapata
中科院分区:
计算机科学3区
文献类型:
--
作者:
Joel Lang;Mirella Lapata

文献摘要

被引文献

相似文献

与许多自然语言处理任务一样,基于监督学习的数据驱动模型已成为语义角色标记的首选方法。当给予足够数量的标记训练数据时,这些模型保证表现良好。然而,生成这些数据既昂贵又耗时,因此提出了无监督方法是否提供可行替代方案的问题。本文的工作假设是,语义角色可以在没有人类监督的情况下从基于三个语言原则的句法分析句子的语料库中归纳出来:(1)相同句法位置(在特定链接内)的论元具有相同的语义角色,(2)子句中的论元具有独特的角色,以及(3)表示相同语义角色的簇应该或多或少在词汇和分布上等效。我们提出了一种实现这些原则的方法,并将任务形式化为图划分问题,其中动词的参数实例表示为图中的顶点,其边表示这些实例之间的相似性。该图由多个边缘层组成,每个边缘层捕获参数实例相似性的不同方面,并且我们开发了用于划分此类多层图的标准聚类算法的扩展。英语和德语的实验表明,我们的方法能够诱导语义角色集群,这些角色集群始终优于强大的基线,并且与最先进的技术具有竞争力。
As in many natural language processing tasks, data-driven models based on supervised learning have become the method of choice for semantic role labeling. These models are guaranteed to perform well when given sufficient amount of labeled training data. Producing this data is costly and time-consuming, however, thus raising the question of whether unsupervised methods offer a viable alternative. The working hypothesis of this article is that semantic roles can be induced without human supervision from a corpus of syntactically parsed sentences based on three linguistic principles: (1) arguments in the same syntactic position (within a specific linking) bear the same semantic role, (2) arguments within a clause bear a unique role, and (3) clusters representing the same semantic role should be more or less lexically and distributionally equivalent. We present a method that implements these principles and formalizes the task as a graph partitioning problem, whereby argument instances of a verb are represented as vertices in a graph whose edges express similarities between these instances. The graph consists of multiple edge layers, each one capturing a different aspect of argument-instance similarity, and we develop extensions of standard clustering algorithms for partitioning such multi-layer graphs. Experiments for English and German demonstrate that our approach is able to induce semantic role clusters that are consistently better than a strong baseline and are competitive with the state of the art.