Probabilistic models of word order and syntactic discontinuity

Probabilistic models of word order and syntactic discontinuity
复制标题

词序和句法不连续性的概率模型

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
R. Levy
R. Levy
中科院分区:
--
文献类型:
--
作者:
Christopher D. Manning;R. Levy

文献摘要

参考文献

被引文献

相似文献

本论文探讨句法理解问题,即句法分析——一个掌握特定语言知识的主体(人或机器)如何推断该语言中一个表层字符串背后的层级结构关系。我认为,结合证据信息的概率模型在认知上是合理的,并且对句法理解具有实际用途。特别是,本论文应用概率方法研究语序与心理语言学理解模型之间的关系,以及在分析具有句法不连续性的句子时的准确性和效率等实际问题。 在心理学方面,本论文提出一种基于预期的处理难度理论,作为概率句法消歧的结果:在理解过程中一个单词处理的难易程度主要由该单词被预期的程度决定。我确定了一类句法现象,主要与动词末尾的从句语序相关,在这些现象中,基于预期的处理预测与更为成熟的基于局部性的处理难度理论差异最为显著。利用现有的概率句法分析算法和带有句法标注的数据源,我表明基于预期的理论比基于局部性的理论更符合一系列已确立的实验心理语言学结果。由于概率驱动和局部性驱动的处理理论之间的比较对语言产生和理解之间的关系,以及更广泛地对认知科学中的模块性理论具有影响,所以它是心理语言学研究的一个关键领域。 本论文还探讨了不连续成分的概率模型问题,即当短语不是由句子的连续子串组成时的情况。不连续性在句法分析中带来了计算上的挑战,因为它使句子中可能的子结构集合超出了句子长度的二次方所限定的可能连续成分集合。对于不连续成分,我研究了准确性问题,即利用基于句法理论原则组织的判别分类器,并将其用于在原本严格的无上下文短语结构树中引入不连续关系;以及研究了对连续和不连续结构进行联合推断的效率问题,方法是使用轻度上下文敏感语法形式的概率实例化,并将语法概括分解为支配和线性顺序的概率成分。
This thesis takes up the problem of syntactic comprehension, or parsing—how an agent (human or machine) with knowledge of a specific language goes about inferring the hierarchical structural relationships underlying a surface string in the language. I take the position that probabilistic models of combining evidential information are cognitively plausible and practically useful for syntactic comprehension. In particular, the thesis applies probabilistic methods in investigating the relationship between word order and psycholinguistic models of comprehension; and in the practical problems of accuracy and efficiency in parsing sentences with syntactic discontinuity. On the psychological side, the thesis proposes a theory of expectation-based processing difficulty as a consequence of probabilistic syntactic disambiguation: the ease of processing a word during comprehension is determined primarily by the degree to which that word is expected. I identify a class of syntactic phenomena, associated primarily with verb-final clause order, where the predictions of expectation-based processing diverge most sharply from more established locality-based theories of processing difficulty. Using existing probabilistic parsing algorithms and syntactically annotated data sources, I show that the expectation-based theory matches a range of established experimental psycholinguistic results better than locality-based theories. The comparison of probabilistic- and locality-driven processing theories is a crucial area of psycholinguistic research due to its implications for the relationship between linguistic production and comprehension, and more generally for theories of modularity in cognitive science. The thesis also takes up the problem of probabilistic models for discontinuous constituency, when phrases do not consist of continuous substrings of a sentence. Discontinuity poses a computational challenge in parsing, because it expands the set of possible substructures in a sentence beyond the bound, quadratic in sentence length, on the set of possible continuous constituents. For discontinuous constituency, I investigate the problem of accuracy employing discriminative classifiers organized on principles of syntactic theory and used to introduce discontinuous relationships into otherwise strictly context-free phrase structure trees; and the problem of efficiency in joint inference over both continuous and discontinuous structures, using probabilistic instantiations of mildly context-sensitive grammatical formalisms and factorizing grammatical generalizations into probabilistic components of dominance and linear order.
言语工作记忆和在线句法处理:来自自定进度聆听的证据。
DOI: 10.1080/02724980343000170
发表时间: 2004
期刊: The Quarterly journal of experimental psychology. A, Human experimental psychology.
影响因子: --
作者:
Waters,GloriaS;Caplan,David
通讯作者: Caplan,David
DOI: 10.1037/0096-1523.22.5.1188
发表时间: 1996-10-01
影响因子: 2.1
作者:
Rayner, K;Sereno, SC;Raney, GE
通讯作者: Raney, GE