CAREER: Interactive Training of Semantic Parsers via Paraphrasing
CAREER: Interactive Training of Semantic Parsers via Paraphrasing
批准号:
1552635
负责人:
Percy Liang
金额:
$55.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-02-01 至 2022-01-31
中文摘要
随着Siri等虚拟助手越来越受欢迎,对深度和强大的语言理解的需求重新出现。统计语义分析是解决这一需求的一个很有前途的范例。构建统计语义解析器的关键障碍是获得足够的训练数据。这个职业项目旨在开发一个新的交互式框架来构建一个语义解析器,该系统就像说英语的外国人一样,要求用户将计算机已经能理解的话语解释成计算机不能理解的话语。该框架开启了有趣的教育应用程序。一种这样的应用是双向辅导系统,其中系统向学生提出问题。学生必须回答并解释问题,从而既练习课程材料,又向系统提供培训数据。自然语言是一个普遍的切入点,它可以增加参与度和促进多样性。高质量的语义解析器可以极大地改善人类与计算机的交互方式。从长远来看,这项工作还可能对自然语言处理系统的构建方式产生重大影响。目前,流行的范例是训练和部署的范例,而如果部署的系统在飞行中学习,则有更多的改进和个性化的机会。这个项目开发了一个新的交互式框架来构建语义解析器,其目的是在给定的领域获得完整的覆盖。其关键思想是让系统选择逻辑形式,生成捕捉其语义的探测话语,并要求用户将其解释为自然输入话语。在这个过程中,系统学习语言变异和新的高级概念。然后使用该数据来训练基于释义的语义分析模型。现有的释义模型要么是基于转换的,擅长捕捉语言的结构规律,要么是基于向量的,擅长捕捉软相似性。该项目开发了新的模型来同时满足这两种需求。该项目开发的框架在三个方面改进了自然语言处理和机器学习的最新水平。首先,该框架不同于收集数据集和学习模型的经典范例;相反,交互系统交错地执行这两个步骤。其次,该框架学习高级概念,这对自然语言理解至关重要,因为单词往往代表复杂的概念。最后,它通过将逻辑表示的刚性和连续表示的灵活性结合在一个统一的模型中,解决了两者之间的典型矛盾。
英文摘要
With the increase in popularity of virtual assistants such as Siri, there is a renewed demand for deep and robust language understanding. Statistical semantic parsing is a promising paradigm for addressing this demand. The key obstacle in building statistical semantic parsers is obtaining adequate training data. This CAREER project aims to develop a new interactive framework for building a semantic parser, where the system, acting like a foreign speaker of English, asks users to paraphrase utterances that the computer already understands into ones that the computer doesn't. The framework opens up intriguing applications in education. One such application is a bidirectional tutoring system, in which the system poses questions to the student. The student must both answer and paraphrase the question, thereby both practicing the course material and providing training data to the system. Natural language is a universal entry point, which can increase engagement and promote diversity. High-quality semantic parsers can drastically improve the way humans interact with computers. In the longer term, this work can also have a significant impact on the way natural language processing systems are built. Currently, the prevailing paradigm is very much a train-and-deploy one, whereas there are many more opportunities for improvement and personalization if deployed systems were to learn on-the-fly.This project develops a new interactive framework for building a semantic parser, which aims to obtain complete coverage in a given domain. The key idea is for the system to choose logical forms, generate probe utterances that capture their semantics, and ask users to paraphrase them into natural input utterances. In the process, the system learns about linguistic variation and novel high-level concepts. The data is then used to train a paraphrasing-based semantic parsing model. Existing paraphrasing models are either transformation-based, which excel at capturing structural regularities in language or are vector-based, which excel at capturing soft similarity. The project develops novel models to capture both. The framework developed in this project improves the state-of-the-art of natural language processing and machine learning in three ways. First, the framework departs from the classic paradigm of gathering a dataset and learning a model; instead, an interactive system interleaves the two steps. Second, the framework learns high-level concepts, which is crucial for natural language understanding, since words often represent complex concepts. Finally, it resolves a classic tension between the rigidity of logical representations and the flexibility of continuous representations by capturing both in a unified model.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金