EAGER: Incremental Semantic Sentence Processing Models
EAGER: Incremental Semantic Sentence Processing Models
批准号:
1551313
负责人:
William Schuler
金额:
$11.66万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2018-08-31
中文摘要
从复杂句子的多种可能解释中提取单一含义是人类最复杂的能力之一,仍然超出了大多数人工语言处理系统的能力范围。目前人类句子处理的计算模型可以通过对单词和句法模式的概率估计来模拟人类的阅读行为,但还不够复杂,无法估计复杂的潜在概念在多个句子中表达的概率。这个探索性的项目将人类句子处理模型扩展到这些基于单词和句法的技术之外,以建模复杂的交叉句子意义,涉及代词和先行词之间的共指关系,以及个人和群体之间的量化关系。所提出的扩展是基于语篇结构的图形表示,它可以随着句子的处理而按时间顺序递增地构建。然后,组合与这些图的各个元素相关联的概率,以获得对输入句子的可能含义的概率估计,然后可以基于这些概率进行比较。然后在在线百科全书文章的解释性文本和现有的广泛覆盖的心理语言学数据集上对得到的计算句子处理模型进行评估。准确的模型如何从自然语言中解码这些复杂的关系可以进一步加深我们对大脑如何工作的理解,并可能在未来某一天允许非程序员领域的专家向机器解释期望的产品、目标和约束。但目前广泛覆盖的句子处理模型主要集中在建模句法上,特别是使用概率上下文无关文法(PCFG)。尽管其句法复杂,但PCFG模型做出了不切实际的假设,即生成的单词序列没有任何指代意义的连续性,或者在可能的共指和量词范围排序之间没有任何偏好。提出的工作将开发一个更接近人类的语义处理模型,通过对现有的增量解析器进行基于图形依赖的语篇表示结构调整来增强。所提出的语义处理模型将定义句子的完整语义依存表示,包括量词范围和共指关系,即使是那些跨越句子边界的句子。然后,该模型将通过基于每个依赖项的源谓词与连接到其目的地的其他谓词的分布相似性,将每次分析的概率估计为其组件依赖项的概率的乘积,从而利用这些依赖项表示的图形化性质。
英文摘要
Extracting a single meaning from the many possible interpretations of a complex sentence is one of the most sophisticated of human abilities, and is still beyond the reach of most artificial language processing systems. Current computational models of human sentence processing can simulate human reading behavior using probability estimates of words and syntactic patterns, but are not yet sophisticated enough to estimate the probability of complex underlying ideas that are expressed across multiple sentences. This exploratory EAGER project extends human sentence processing models beyond these word- and syntax-based techniques to model complex cross-sentential meaning involving coreference relationships between pronouns and their antecedents, and quantificational relationships between individuals and groups. The proposed extensions are based on a graphical representation of discourse structure, which can be constructed incrementally in time order as sentences are processed. Probabilities associated with individual elements of these graphs are then combined to obtain probability estimates over possible meanings of input sentences, which can then be compared based on these probabilities. The resulting computational sentence processing models are then evaluated on explanatory text from on-line encyclopedia articles and on existing broad-coverage psycholinguistic datasets.Accurate models of how these complex relationships are decoded from natural language could further our understanding of how the brain works, and may someday allow non-programmer domain experts to explain desired products, goals and constraints to machines. But current broad-coverage sentence processing models are focused primarily on modeling syntax, in particular using probabilistic context-free grammar (PCFG) surprisal. Despite their syntactic sophistication, PCFG models make unrealistic assumptions that word sequences are generated without any continuity of referential meaning or any preferences among possible coreference and quantifier scope orderings. The proposed work will develop a more human-like semantic processing model by augmenting an existing incremental parser with a graphical dependency-based adaptation of discourse representation structures. The proposed semantic processing model will define complete semantic dependency representations of sentences, including quantifier scope and coreference relationships, even those that cross sentence boundaries. The model will then exploit the graphical nature of these dependency representations by estimating the probability of each analysis as the product of the probabilities of its component dependencies, based on the distributional similarity of each dependency's source predicate to the other predicates connected to its destination.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CompCog: RI: Small: Human-like semantic grammar induction through knowledge distillation from pre-trained language models
-
批准号:2313140
-
项目类别:Standard Grant
-
资助金额:$48.45万
-
财政年份:2023
-
负责人:William Schuler
-
依托单位:
RI: Small:Comp Cog: Broad-coverage semantic models of human sentence processing
-
批准号:1816891
-
项目类别:Standard Grant
-
资助金额:$49.03万
-
财政年份:2018
-
负责人:William Schuler
-
依托单位:
CAREER: Integrating denotational meaning into probabilistic language models
-
批准号:0447685
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2005
-
负责人:William Schuler
-
依托单位:
海外基金