CAREER: Teaching Machines through Human Explanation for Information Extraction
CAREER: Teaching Machines through Human Explanation for Information Extraction
批准号:
2048211
负责人:
Xiang Ren
金额:
$51.31万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-05-15 至 2026-04-30
中文摘要
社会产生的大部分数据都是自由格式的文本数据。随着文本数据量的持续增长,人类本身不能指望能够理解发布的每一条文本信息。因此,需要基于机器的方法来从海量文本数据中提取显著实体及其关系。虽然开发这种方法的努力在学术界被证明是成功的,但这种成功很少转化为实践者使用提取系统来解决现实世界的问题。翻译失败的一个重要原因是机器需要大量的训练样本来学习提取模型。即使存在足够数量的示例,机器也会学习非常僵化的方法,因此即使是轻微的拼写错误也可能导致失败。因此,我们需要重新思考如何开发、改进和维护这种基于机器的提取方法。该项目提出了一种新的方法,基于这样的想法,即为机器做出的正确和不正确的决定提供解释,目的是需要更少的例子来进行机器训练,以及提供一个软化当前提取方法的僵化的过程。为了实现这一目标,这个项目将征求(人类)关于机器应该如何对其任务进行推理的自然语言解释,以及纠正错误推理的解释,以及当机器的基本原理可能出错时的警报系统。这个项目并不是仅仅把人类当作标签的来源,而是旨在开发一种新的学习框架,直接模拟人类的自然语言解释,以便为机器提供标签原理,或者纠正观察到的错误原理。该项目将开发基于解释的学习方法,可以捕获人类自然语言解释的组成性质,并研究解释引导的模型改进方法,以根据提供的关于不良行为的人类解释来更新模型参数。为了适应不断变化的数据分布,该项目将制定一个人在循环中的连续模型细化框架,其中自动识别有问题的模型行为模式,并征求人类反馈以修正模型。随着这些进步,该项目希望从根本上改变培训、改进和更新模型的方式,并希望通过利用人类解释中包含的专家知识来做到这一点。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The majority of data being generated by society is free-form textual data. As the volume of text data continues to grow, humans alone cannot hope to be able to understand every piece of textual information published. Hence, a need for machine-based methods to extract salient entities along with their relationships from massive textual data is needed. While efforts to develop such methods have proven successful in academia, that success rarely translates over to practitioners employing extraction systems to solve real-world problems. A significant cause of this failure in translation is the requirement of copious amounts of training examples for a machine to learn extraction models. Even when a sufficient number of examples exist, machines learn very rigid methods, such that even a slight misspelling can cause a failure. Therefore, a re-think of how we develop, refine and maintain such machine-based extraction methods is required. This project proposes a new methodology based around the idea of providing explanations for both correct and incorrect decisions made by a machine, with the intention of requiring far fewer examples for machine training, as well as providing a process of softening the rigidity of current extraction methods. To achieve this goal, this project will solicit (from humans) natural language explanations on how a machine should reason about their task, as well as explanations correcting erroneous reasoning and an alerting system for when a machine’s rationale is possibly going wrong. Rather than treating humans as merely a “source of labels”, this project aims at developing a new learning framework that directly models a human’s natural language explanations to either provide a machine with labeling rationale or correct an observed erroneous rationale. The project will develop explanation-based learning methods that can capture the compositional nature of human natural language explanations, and study explanation-guided model refinement methods to update model parameters based on the provided human explanations regarding undesirable behaviors. To adapt to changing data distribution, this project will formulate a human-in-the-loop continual model refinement framework where problematic model behavioral patterns are automatically identified, and human feedback is solicited to correct the model. With these advancements, the project looks to fundamentally change the way models are trained, refined and updated, and look to do it by exploiting the expert knowledge contained within human explanations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Modeling the Invention, Dissemination, and Translation of Scientific Concepts
-
批准号:1829268
-
项目类别:Standard Grant
-
资助金额:$24.0万
-
财政年份:2018
-
负责人:Xiang Ren
-
依托单位:
海外基金