CAREER: Large-Scale Learning for Information Extraction
CAREER: Large-Scale Learning for Information Extraction
批准号:
2052498
负责人:
Alan Ritter
金额:
$48.9万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-08-16 至 2025-08-31
中文摘要
人类的许多知识都是以文本的形式编码的。该项目旨在大幅提高机器阅读大量文档集合并以最少的人力对其中包含的知识进行推理的能力。这将帮助人们克服信息过载,并通过分析锁定在非结构化文本中的重要信息来做出更好的决策。近年来,通过在海量、高质量的数据集上应用深度学习方法,语音识别和机器翻译等任务取得了巨大的进步;然而,大多数可用于信息提取的数据集要么很小,要么噪声很大。该项目将通过开发新的方法来应对这些挑战,这些方法可以更有效地从大型但有噪音的数据集中学习,这些数据集是通过从现有知识库(KB)进行远程监督而构建的。为了证明新方法的有效性,他们将被用来支持几个新的应用。这些措施包括检测在线报告的网络威胁,以及分析专家对其严重性的意见。最近的研究发现,75%的软件漏洞是在网上首次报告的,这让攻击者有时间利用该漏洞。能够自动阅读计算机安全博客并分析新威胁的系统可以帮助安全从业者更有效地跟踪它们并确定优先顺序。该项目包括一项整合研究和教育的计划。外展努力旨在帮助吸引更多样化的学生群体学习计算机科学。其中包括动手研讨会,让新生接触到令人兴奋的自然语言处理和人工智能应用程序。该项目还将帮助高级本科生通过尖端信息提取技术的新课程教材进行研究。该研究将通过发明新的方法来解决机器读取数据的瓶颈,这些方法可以使用远程监控从大型、嘈杂的数据集中有效地学习。这些方法将通过在学习过程中对潜在变量进行推理、填充缺失信息和解决歧义来解决远程监督中固有的标签噪声的挑战。该方法结合了结构化学习和神经网络的优点;模型的结构化学习组件在其足够自信的情况下可以覆盖噪声标签--这与知识库中缺失数据的模型相平衡。这将促进许多新任务和领域的萃取器的快速发展。为了证明这一点,大量的实验将与使用标准基准数据集进行信息提取的最先进方法进行比较,这些数据集包括Freebase/NYT语料库、TAC KBP数据集和TACRED。此外,这项研究将通过探索新的应用来展示该方法的一般性,包括实体、关系和事件提取、时间归一化以及学习使用远程监控从国家漏洞数据库(NVD)提取网络威胁情报的实时馈送,从而推动信息提取的最小监督的界限。这些应用程序得到了全面评估计划的支持,该计划包括开发新的语料库和指标。除了用于最低限度监督信息提取的工具包外,该项目还将产生一些新的数据集,并将作为开放源码软件共享。这项研究工作将支持以最少的人力为广泛的新任务和领域快速开发信息系统。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Much of human knowledge is encoded in text. This project aims to substantially advance the capability of machines to read large document collections and reason about the knowledge contained within them using minimal human effort. This will help people to overcome information overload and make better decisions by analyzing vital information that is locked away in unstructured text. Recent years have seen tremendous progress on tasks such as speech recognition and machine translation, by applying deep learning methods on massive, high-quality datasets; however, most available datasets for information extraction are either small or very noisy. The project will address these challenges by developing new methods that can learn more effectively from big, but noisy datasets that are constructed using distant supervision from an existing knowledge base (KB). To demonstrate the new methods' effectiveness, they will be used to support several novel applications. These include the detection of cyber-threats reported online and the analysis of experts' opinions about their severity. Recent studies have found that 75% of software vulnerabilities are first reported online, giving attackers time to exploit the vulnerability. Systems that can automatically read computer security blogs and analyze new threats could help security practitioners to track and prioritize them more effectively. The project includes a plan for integrating research and education. Outreach efforts aim to help attract a more diverse group of students to study computer science. These include hands-on workshops to expose freshmen to exciting natural language processing and artificial intelligence applications. The project will also help to engage advanced undergraduate students in research through new course materials on cutting-edge information extraction techniques.The research will address the machine reading data bottleneck by inventing new methods that can learn effectively from large, noisy datasets using distant supervision. These methods will address the challenge of label noise inherent in distant supervision by performing inference over latent variables during learning, filling in missing information, and resolving ambiguities. The approach combines the benefits of structured learning and neural networks; the structured learning component of the model can override noisy labels in cases where it is sufficiently confident -- this is balanced against a model of missing data in the KB. This will catalyze the rapid development of extractors for many new tasks and domains. To demonstrate this, extensive experiments will compare against state of the art methods using standard benchmark datasets for information extraction, including the Freebase/NYT corpus, TAC KBP datasets, and TACRED. Furthermore, the research will push the boundaries of minimal supervision for Information Extraction by exploring new applications that demonstrate the generality of the approach, including entity, relation and event extraction, time normalization and learning to extract a real-time feed of cyber-threat intelligence using distant supervision from the National Vulnerability Database (NVD). These applications are supported by a comprehensive evaluation plan that includes the development of new corpora and metrics. The project will produce a number of new datasets in addition to a toolkit for minimally supervised information extraction, that will be shared as open source software. This research effort will support the rapid development of information systems for a broad range of new tasks and domains using minimal human effort.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(17)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.18653/v1/2021.emnlp-main.459
发表时间:
2021
期刊:
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
影响因子:
--
作者:
[Chen, Yang, Ritter, Alan]
通讯作者:
Ritter, Alan
Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023?
CoNLL-2003 命名实体标记器在 2023 年仍然有效吗?
DOI:
--
发表时间:
2023
期刊:
ACL 2023
影响因子:
--
作者:
[Shuheng Liu, Alan Ritter]
通讯作者:
Alan Ritter
Few-Shot Anaphora Resolution in Scientific Protocols via Mixtures of In-Context Experts
通过上下文专家的混合在科学协议中进行少样本照应解析
DOI:
--
发表时间:
2022
期刊:
Findings of EMNLP 2022
影响因子:
--
作者:
[Nghia T. Le, Fan Bai, Alan Ritter]
通讯作者:
Alan Ritter
DOI:
10.48550/arxiv.2305.01645
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
作者:
[Junmo Kang;Wei Xu;Alan Ritter]
通讯作者:
Junmo Kang;Wei Xu;Alan Ritter
DOI:
10.18653/v1/2020.emnlp-main.382
发表时间:
2020-04
期刊:
影响因子:
--
作者:
[Wuwei Lan;Yang Chen;Wei Xu;Alan Ritter]
通讯作者:
Wuwei Lan;Yang Chen;Wei Xu;Alan Ritter
共 17 条
CAREER: Large-Scale Learning for Information Extraction
-
批准号:1845670
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2019
-
负责人:Alan Ritter
-
依托单位:
CRII: III: Learning to Extract Events from Knowledge Base Revisions
-
批准号:1464128
-
项目类别:Standard Grant
-
资助金额:$15.13万
-
财政年份:2015
-
负责人:Alan Ritter
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:黄洛将
-
依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:黄洛将
-
依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
-
批准号:12074246
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2020
-
负责人:Yoshitomo Kamiya
-
依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
-
批准号:31972875
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:石江华
-
依托单位:
Large PB/PB小鼠 视网膜新生血管模型的研究
-
批准号:30971650
-
项目类别:面上项目
-
资助金额:8.0万元
-
批准年份:2009
-
负责人:周旻
-
依托单位:
基因discs large在果蝇卵母细胞的后端定位及其体轴极性形成中的作用机制
-
批准号:30800648
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2008
-
负责人:于玲珠
-
依托单位:
LARGE基因对口腔癌细胞中α-DG糖基化及表达的分子调控
-
批准号:30772435
-
项目类别:面上项目
-
资助金额:29.0万元
-
批准年份:2007
-
负责人:尚政军
-
依托单位: