课题基金 / 基金详情

Annotating Reference and Coreference In Dialogue Using Conversational Agents in games

Annotating Reference and Coreference In Dialogue Using Conversational Agents in games
在游戏中使用对话代理注释对话中的参考和共指
批准号:
EP/W001632/1
负责人:
Massimo Poesio
金额:
$139.06万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

Massimo Poesio的其他基金

相似基金

相关文献

中文摘要
翻译
编码器/解码器模型和Transformer等现代神经网络架构的发展,引发了人们对能够参与对话(又名对话代理)的人工智能系统的神经模型的兴趣激增,这反映在大量出版的作品、专门的研讨会和行业赞助的竞赛和拨款上。虽然最初这些模型应用于简单的聊天机器人,但研究的重点已经转向能够参与更复杂和以任务为导向的对话的对话代理,如餐厅预订或问题回答。但是这些任务的结果表明,虽然没有语义解释专用模型的端到端架构可以很好地用于聊天机器人,但执行更复杂任务的会话代理需要更强的能力来处理这些方面的解释,以及某种形式的上下文建模。需要更高级架构的自然语言解释方面包括COREFERENCE和REFERENCE。为了说明对话中相互引用的重要性,考虑以下例子,除了现实生活中的聊天对话,其中参与者不断使用回指式表达,如both, THEY, IT等来指代先前引入的实体,如谷歌或Microsoft。A:你是b谷歌或微软的粉丝吗?B:两者都是优秀的技术,在很多方面都很有帮助。出于安全目的,两者都是超级的。A:我不是b谷歌的超级粉丝,但我经常使用它,因为我必须这样做。我认为他们在某种意义上是垄断的。B:谷歌提供在线相关服务和产品,包括搜索引擎和云计算。A:是的,他们的服务很好。我只是不喜欢它们对我们个人生活的干扰,让对话代理具有进行这些形式的解释的能力会带来两个问题。首先,为这些任务开发模型需要特定的训练数据:大多数深度学习架构都是在大量免费的书面文本上进行训练的。然而,在书面文本和领域上训练一个共指解析器并使其适应对话是无效的,因为对话中的共指涉及不同的现象,比文本中的共指更复杂。其次,开发的体系结构需要特定的模块,使它们能够解释共引用和引用。我们的团队率先使用带有目的的游戏(gwap)来收集NLP数据,从而产生了使用gwap或众包收集的最大的NLP数据集。但对话和书面文本之间有一个根本的区别:后者是为第三方设计的,而研究表明,无意中听到对话的人只能部分理解所说的内容。我们提出的解决方案是通过游戏收集对回指和参考信息的判断,在游戏中,对话代理与人类玩家互动,并通过从他们那里获取信息来进化。这个想法建立在Facebook和微软最近的工作基础上,其中包括在游戏中率先使用对话代理来收集关于对话的数据,以及Hockenmaier和她的实验室。我们的代理将与这些实验室合作部署在LIGHT和MINECRAFT等游戏平台上。但是,在之前的工作中,会话代理仅与改善其端到端行为的目标进行交互,而在提议的项目中,我们将开发能够通过在适当的时刻通过向参与者澄清问题收集有关这些解释方面的判断来提高其解释共同参考和参考的能力的人工代理,这也可用于注释数据集。
英文摘要
The development of modern neural network architectures architectures such as the encoder/decoder model and the Transformer has brought about an explosion of interest in neural models for AI systems able to engage in conversations (aka conversational agents), reflected by a spike of published work, dedicated workshops, and industry-sponsored competitions and grants. While at first these models were applied to simple chatbots, the focus of research has been shifting towards conversational agents capable of engaging in more complex and task-oriented dialogue such as restaurant booking or question answering. But the results on these tasks show that while end-to-end architectures without dedicated models for semantic interpretation can work well for chatbots, conversational agents carrying out more complex tasks require greater ablity to handle such aspects of interpretation, and some form of modelling of context. Among the aspects of natural language interpretation that require more advanced architectures are COREFERENCE and REFERENCE. For an example of the importance of coreference in dialog, consider the following except from a real-life chat conversation, where both participants continually use anaphoric expressions such as BOTH, THEY, IT, etc to refer to previously introduced entities such as Google or Microsoft.A:Are you a fan of Google or Microsoft?B:Both are excellent technology they are helpful in many ways. For the security purpose both are super.A:I'm not a huge fan of Google, but I use it a lot because I have to. I think they are a monopoly in some sense.B:Google provides online related services and products, which includes search engine and cloud computing.A:Yeah, their services are good. I'm just not a fan of intrusive they can be on our personal livesEnriching conversational agents with the ability to carry out these forms of interpretation raises two issues. First, developing models for these tasks requires specific training data: most deep-learning architectures are trained on large amounts of freely available written text. Training a coreference resolver on written text and domain-adapting it to dialogue however has proven ineffective as coreference in dialogue involves different phenomena and is more involved than coreference in text. Second, the developed architectures require specific modules that enable them to interpret coreference and reference. Our group has pioneered the use of Games-With-A-Purpose (GWAPs) to collect data for NLP, resulting in the largest NLP dataset collected using GWAPs or indeed crowdsourcing. But there is a fundamental difference between conversation and written text: the latter is designed to be read by third parties, whereas research has shown that overhearers to a conversation only acquire a partial understanding of what was said.OUR PROPOSED SOLUTION to the problem of creating large annotated datasets of coreference and reference interpretation in conversation is to collect the judgments for anaphoric and referential information via GAMES IN WHICH CONVERSATIONAL AGENTS INTERACT WITH HUMAN PLAYERS AND EVOLVE BY ACQUIRING INFORMATION FROM THEM. This idea builds on recent work by Facebook and Microsoft, among others, that pioneered the use of conversational agents in games to collect data about dialogue, and of Hockenmaier and her lab. Our agents will be deployed in gaming platforms such as LIGHT and MINECRAFT in collaboration with these labs. But whereas in previous work conversational agents only interact with the aim to improve their end-to-end behavior, in the proposed project we will develop artificial agents able to improve their ability to interpret coreference and reference by collecting judgments about these interpretation aspects via CLARIFICATION QUESTIONS to the players at appropriate moments, which can also be used to annotate a dataset.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
The CODI-CRAC 2022 Shared Task on Anaphora, Bridging, and Discourse Deixis in Dialogue
CODI-CRAC 2022 对话中的照应、桥接和话语指示语共享任务
DOI: --
发表时间: 2022
期刊:
影响因子: --
作者: [Yu, J]
通讯作者: Yu, J
Aggregating crowdsourced and automatic judgments to scale up a corpus of anaphoric reference for fiction and Wikipedia texts
聚合众包和自动判断,以扩大小说和维基百科文本的照应参考语料库
DOI: --
发表时间: 2023
期刊:
影响因子: --
作者: [Yu, J]
通讯作者: Yu, J
LingoTowns: A Virtual World For Natural Language Annotation and Language Learning
LingoTowns:自然语言注释和语言学习的虚拟世界
DOI: 10.1145/3505270.3558323
发表时间: 2022
期刊:
影响因子: --
作者: [Madge C]
通讯作者: Madge C
Coreference Annotation of an Arabic Corpus using a Virtual World Game
使用虚拟世界游戏对阿拉伯语语料库进行共指注释
DOI: 10.18653/v1/2022.wanlp-1.37
发表时间: 2022
期刊:
影响因子: --
作者: [Aliady W]
通讯作者: Aliady W
共 8 条
    Creating anaphorically annotated resources through semantic wikis (AnaWiki)
    • 批准号:
      EP/F00575X/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $18.26万
    • 财政年份:
      2007
    • 负责人:
      Massimo Poesio
    • 依托单位:
    海外基金