Readers: Evaluation and Development of Reading Systems
Readers: Evaluation and Development of Reading Systems
批准号:
EP/K017845/1
负责人:
Mirella Lapata
金额:
$37.83万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2013
资助国家:
英国
项目状态:
已结题
起止时间:
2013 至 --
中文摘要
机器阅读的目的是从非结构化的文本中提取知识,几乎不需要人工的努力。自早期以来,这一直是人工智能的主要目标。互联网上不断增长的文本数据进一步增加了基于计算机的知识提取方法的重要性和紧迫性。机器阅读的成功不仅有助于突破人工智能的知识获取瓶颈,而且会给Web搜索、信息提取以及维基百科等资源的自动构建带来革命性的变化。过去,在使用标记和解析等标准NLP技术自动化机器阅读的许多子任务方面取得了很大进展。然而,端到端的解决方案仍然很少,并且现有的系统通常需要大量的人工工程和/或标记示例。因此,它们通常针对有限的领域,只提取有限类型的知识(例如,预先指定的关系)。在这个项目中,我们的目标是开发一个端到端系统,该系统可以对原始文本进行操作,提取知识,并能够回答问题并支持其他端任务。我们的方法中的一个关键见解是使用无监督方法,这种方法不依赖于大量的手工注释来获取背景知识,将其链接到现有的知识库,以及创建新的知识库。我们的方法将在网络规模上获取知识,对任意领域、类型和语言开放。它将不断地集成新的信息源(例如,新的文本文档),并从用户的问题和反馈中学习(例如,通过执行终端任务)。
英文摘要
Machine reading aims to extract knowledge from unstructured text with little human effort. It has been a major goal of AI since its early days. The ever growing amounts of textual data available over the internet further increase the importance and urgency of computer-based methods for knowledge extraction. The success of machine reading will not only help breach the knowledge acquisition bottleneck in AI, but also revolutionize Web search, information extraction, and the automatic construction of resources such as Wikipedia. In the past, there has been a lot of progress in automating many substasks of machine reading using standard NLP technology such as tagging and parsing. However, end-to-end solutions are still rare, and existing systems typically require substantial human effort in manual engineering and/or labeling examples. As a result, they often target restricted domains and only extract limited types of knowledge (e.g., a pre-specified relation). In this project we aim to develop an end-to-end system that operates over raw text, extracts knowledge and is able to answer questions and support other end tasks. A key insight in our approach is the use of unsupervised methods that do not rely on large amounts of hand annotation for the acquisition of background knowledge, its linking to existing knowledge bases, and the creation of new ones. Our approach will acquire knowledge at Web-scale, be open to arbitrary domains, genres, and languages.It will constantly integrate new information sources (e.g., new text documents) and learn from user questions and feedback (e.g., via performing end tasks).
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1613/jair.4431
发表时间:
2014-09
期刊:
J. Artif. Intell. Res.
影响因子:
--
作者:
[K. Woodsend;Mirella Lapata]
通讯作者:
K. Woodsend;Mirella Lapata
DOI:
10.1162/coli_a_00195
发表时间:
2014-09
期刊:
Computational Linguistics
影响因子:
9.3
作者:
[Joel Lang;Mirella Lapata]
通讯作者:
Joel Lang;Mirella Lapata
DOI:
10.3115/v1/d14-1045
发表时间:
2014-10
期刊:
影响因子:
--
作者:
[Michael Roth;K. Woodsend]
通讯作者:
Michael Roth;K. Woodsend
DOI:
10.1162/tacl_a_00190
发表时间:
2014-10
期刊:
Transactions of the Association for Computational Linguistics
影响因子:
10.9
作者:
[Siva Reddy;Mirella Lapata;Mark Steedman]
通讯作者:
Siva Reddy;Mirella Lapata;Mark Steedman
DOI:
10.1162/tacl_a_00150
发表时间:
2015-08
期刊:
Transactions of the Association for Computational Linguistics
影响因子:
10.9
作者:
[Michael Roth;Mirella Lapata]
通讯作者:
Michael Roth;Mirella Lapata
共 6 条
A Unified Model of Compositional and Distributional Semantics: Theory and Applications
-
批准号:EP/I037415/1
-
项目类别:Research Grant
-
资助金额:$11.73万
-
财政年份:2013
-
负责人:Mirella Lapata
-
依托单位:
Global Inference for Summarization Using Integer Linear Programming
-
批准号:EP/F055765/1
-
项目类别:Research Grant
-
资助金额:$34.38万
-
财政年份:2009
-
负责人:Mirella Lapata
-
依托单位:
Application-based Text-to-Text Generation
-
批准号:GR/T04557/01
-
项目类别:Research Grant
-
资助金额:$21.45万
-
财政年份:2006
-
负责人:Mirella Lapata
-
依托单位:
国内基金
海外基金
基于重要农地保护LESA(Land Evaluation and Site Assessment)体系思想的高标准基本农田建设研究
-
批准号:41340011
-
项目类别:专项基金项目
-
资助金额:20.0万元
-
批准年份:2013
-
负责人:钱凤魁
-
依托单位: