A Syntactic Analysis Method of Long Japanese Sentences Based on the Detection of Conjunctive Structures

A Syntactic Analysis Method of Long Japanese Sentences Based on the Detection of Conjunctive Structures
复制标题

基于连词结构检测的日语长句句法分析方法

DOI:
10.5555/203987.203988
复制
发表时间:
1994
期刊:
Comput. Linguistics
影响因子:
--
通讯作者:
M. Nagao
M. Nagao
中科院分区:
--
文献类型:
--
作者:
S. Kurohashi;M. Nagao

文献摘要

被引文献

相似文献

本文提出了一种句法分析方法,该方法首先通过检查两个系列单词的并行性,然后借助于有关结构的信息来分析句子的依赖性结构,首先检测句子中的结构。对长句子的分析是自然语言处理中最困难的问题之一。造成这种困难的主要原因是结构性歧义是出现在长句子中的结构中常见的。人类可以识别结构,因为有时但有时是微妙的相似性。因此,我们开发了一种用于计算两个任意系列单词之间的相似性度量的算法,并选择了两个最相似的单词,可以合理地将其视为组成结构。这是使用动态编程技术实现的。通过识别结构结构,可以将长句子简化为较短的形式。因此,可以通过相对简单的头部依赖规则获得句子的总依赖性结构。除了范围的模棱两可之外,关于连接结构的一个严重问题是它们的某些组件的省略号。通过我们的依赖分析过程,我们可以找到椭圆并恢复省略的组件。我们报告分析150个日本句子以说明这种方法的有效性的结果。
This paper presents a syntactic analysis method that first detects conjunctive structures in a sentence by checking parallelism of two series of words and then analyzes the dependency structure of the sentence with the help of the information about the conjunctive structures. Analysis of long sentences is one of the most difficult problems in natural language processing. The main reason for this difficulty is the structural ambiguity that is common for conjunctive structures that appear in long sentences. Human beings can recognize conjunctive structures because of a certain, but sometimes subtle, similarity that exists between conjuncts. Therefore, we have developed an algorithm for calculating a similarity measure between two arbitrary series of words from the left and the right of a conjunction and selecting the two most similar series of words that can reasonably be considered as composing a conjunctive structure. This is realized using a dynamic programming technique. A long sentence can be reduced into a shorter form by recognizing conjunctive structures. Consequently, the total dependency structure of a sentence can be obtained by relatively simple head-dependent rules. A serious problem concerning conjunctive structures, besides the ambiguity of their scopes, is the ellipsis of some of their components. Through our dependency analysis process, we can find the ellipses and recover the omitted components. We report the results of analyzing 150 Japanese sentences to illustrate the effectiveness of this method.