Parsing entire discourses as very long strings: Capturing topic continuity in grounded language learning

Parsing entire discourses as very long strings: Capturing topic continuity in grounded language learning
复制标题

将整个话语解析为很长的字符串:捕捉扎根语言学习中的主题连续性

DOI:
--
复制
发表时间:
2013
影响因子:
10.9
通讯作者:
Mark Johnson
Mark Johnson
中科院分区:
人文科学1区
文献类型:
--
作者:
Minh;Michael C. Frank;Mark Johnson

文献摘要

被引文献

相似文献

扎根语言学习,从自然语言映射到意义的表示的任务,近年来吸引了越来越多的兴趣。然而,在大多数关于这一主题的研究中,会话中的话语被单独对待,话语结构信息在很大程度上被忽视。在语言习得的背景下,这种独立性假设抛弃了对学习者很重要的线索,例如,连续的话语可能共享相同的所指物的事实(Frank等人,2013年)。本文介绍了一种方法,同时建模扎根语言在句子和语篇层面的问题。我们结合联合收割机的想法,从解析和语法归纳产生一个解析器,可以处理长输入字符串与数千个令牌,创建解析树,代表完整的话语。通过将基础语言学习视为语法推理任务,我们使用解析器扩展了约翰逊等人(2012)的工作,研究了话语连续性在儿童语言习得中的重要性及其与社交线索的相互作用。我们的模型提高了语言习得任务的性能,并产生良好的话语分割相比,人类注释。
Grounded language learning, the task of mapping from natural language to a representation of meaning, has attracted more and more interest in recent years. In most work on this topic, however, utterances in a conversation are treated independently and discourse structure information is largely ignored. In the context of language acquisition, this independence assumption discards cues that are important to the learner, e.g., the fact that consecutive utterances are likely to share the same referent (Frank et al., 2013). The current paper describes an approach to the problem of simultaneously modeling grounded language at the sentence and discourse levels. We combine ideas from parsing and grammar induction to produce a parser that can handle long input strings with thousands of tokens, creating parse trees that represent full discourses. By casting grounded language learning as a grammatical inference task, we use our parser to extend the work of Johnson et al. (2012), investigating the importance of discourse continuity in children’s language acquisition and its interaction with social cues. Our model boosts performance in a language acquisition task and yields good discourse segmentations compared with human annotators.