CAREER: Large-Scale Learning for Information Extraction
CAREER: Large-Scale Learning for Information Extraction
批准号:
2052498
负责人:
Alan Ritter
金额:
$48.9万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-08-16 至 2025-08-31
中文摘要
人类的大部分知识都是以文本的形式编码的。 该项目旨在大大提高机器阅读大型文档集的能力,并以最少的人力来推理其中包含的知识。 这将帮助人们克服信息过载,并通过分析锁定在非结构化文本中的重要信息来做出更好的决策。 近年来,通过将深度学习方法应用于大规模的高质量数据集,语音识别和机器翻译等任务取得了巨大进展;然而,大多数用于信息提取的数据集要么很小,要么非常嘈杂。 该项目将通过开发新方法来应对这些挑战,这些方法可以更有效地从大型但嘈杂的数据集中学习,这些数据集是使用现有知识库(KB)的远程监督构建的。 为了证明新方法的有效性,它们将用于支持几个新的应用程序。 其中包括检测在线报告的网络威胁,以及分析专家对其严重性的意见。 最近的研究发现,75%的软件漏洞首先在网上报告,这给了攻击者利用漏洞的时间。 可以自动阅读计算机安全博客并分析新威胁的系统可以帮助安全从业人员更有效地跟踪和优先考虑这些威胁。 该项目包括一项研究与教育相结合的计划。 推广工作旨在帮助吸引更多不同的学生学习计算机科学。 其中包括实践研讨会,让新生接触到令人兴奋的自然语言处理和人工智能应用。 该项目还将通过关于尖端信息提取技术的新课程材料,帮助高年级本科生参与研究。该研究将通过发明新方法来解决机器阅读数据的瓶颈,该方法可以使用远程监督从大型噪声数据集有效学习。 这些方法将通过在学习过程中对潜在变量进行推理、填充缺失信息和解决歧义来解决远程监督中固有的标签噪声的挑战。 该方法结合了结构化学习和神经网络的优点;模型的结构化学习组件可以在足够自信的情况下覆盖噪声标签-这与KB中缺失数据的模型相平衡。 这将促进许多新任务和领域的提取器的快速发展。 为了证明这一点,广泛的实验将与使用标准基准数据集进行信息提取的最新方法进行比较,包括Freebase/NYT语料库,TAC KBP数据集和TACRED。 此外,该研究将通过探索新的应用程序来推动信息提取的最小监督边界,这些应用程序展示了该方法的通用性,包括实体,关系和事件提取,时间归一化和学习使用国家漏洞数据库(NVD)的远程监督来提取网络威胁情报的实时馈送。 这些应用程序由一个全面的评估计划支持,其中包括开发新的语料库和指标。 该项目将产生一些新的数据集,以及一个用于最低限度监督信息提取的工具包,这些工具包将作为开放源码软件共享。 这项研究工作将支持信息系统的快速发展,为广泛的新任务和领域,使用最少的人力资源。这个奖项反映了NSF的法定使命,并已被认为是值得通过评估使用基金会的智力价值和更广泛的影响审查标准的支持。
英文摘要
Much of human knowledge is encoded in text. This project aims to substantially advance the capability of machines to read large document collections and reason about the knowledge contained within them using minimal human effort. This will help people to overcome information overload and make better decisions by analyzing vital information that is locked away in unstructured text. Recent years have seen tremendous progress on tasks such as speech recognition and machine translation, by applying deep learning methods on massive, high-quality datasets; however, most available datasets for information extraction are either small or very noisy. The project will address these challenges by developing new methods that can learn more effectively from big, but noisy datasets that are constructed using distant supervision from an existing knowledge base (KB). To demonstrate the new methods' effectiveness, they will be used to support several novel applications. These include the detection of cyber-threats reported online and the analysis of experts' opinions about their severity. Recent studies have found that 75% of software vulnerabilities are first reported online, giving attackers time to exploit the vulnerability. Systems that can automatically read computer security blogs and analyze new threats could help security practitioners to track and prioritize them more effectively. The project includes a plan for integrating research and education. Outreach efforts aim to help attract a more diverse group of students to study computer science. These include hands-on workshops to expose freshmen to exciting natural language processing and artificial intelligence applications. The project will also help to engage advanced undergraduate students in research through new course materials on cutting-edge information extraction techniques.The research will address the machine reading data bottleneck by inventing new methods that can learn effectively from large, noisy datasets using distant supervision. These methods will address the challenge of label noise inherent in distant supervision by performing inference over latent variables during learning, filling in missing information, and resolving ambiguities. The approach combines the benefits of structured learning and neural networks; the structured learning component of the model can override noisy labels in cases where it is sufficiently confident -- this is balanced against a model of missing data in the KB. This will catalyze the rapid development of extractors for many new tasks and domains. To demonstrate this, extensive experiments will compare against state of the art methods using standard benchmark datasets for information extraction, including the Freebase/NYT corpus, TAC KBP datasets, and TACRED. Furthermore, the research will push the boundaries of minimal supervision for Information Extraction by exploring new applications that demonstrate the generality of the approach, including entity, relation and event extraction, time normalization and learning to extract a real-time feed of cyber-threat intelligence using distant supervision from the National Vulnerability Database (NVD). These applications are supported by a comprehensive evaluation plan that includes the development of new corpora and metrics. The project will produce a number of new datasets in addition to a toolkit for minimally supervised information extraction, that will be shared as open source software. This research effort will support the rapid development of information systems for a broad range of new tasks and domains using minimal human effort.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(17)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.18653/v1/2021.emnlp-main.459
发表时间:
2021
期刊:
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
影响因子:
--
作者:
[Chen, Yang, Ritter, Alan]
通讯作者:
Ritter, Alan
Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023?
CoNLL-2003 命名实体标记器在 2023 年仍然有效吗?
DOI:
--
发表时间:
2023
期刊:
ACL 2023
影响因子:
--
作者:
[Shuheng Liu, Alan Ritter]
通讯作者:
Alan Ritter
Few-Shot Anaphora Resolution in Scientific Protocols via Mixtures of In-Context Experts
通过上下文专家的混合在科学协议中进行少样本照应解析
DOI:
--
发表时间:
2022
期刊:
Findings of EMNLP 2022
影响因子:
--
作者:
[Nghia T. Le, Fan Bai, Alan Ritter]
通讯作者:
Alan Ritter
DOI:
10.48550/arxiv.2305.01645
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
作者:
[Junmo Kang;Wei Xu;Alan Ritter]
通讯作者:
Junmo Kang;Wei Xu;Alan Ritter
DOI:
10.18653/v1/2020.emnlp-main.382
发表时间:
2020-04
期刊:
影响因子:
--
作者:
[Wuwei Lan;Yang Chen;Wei Xu;Alan Ritter]
通讯作者:
Wuwei Lan;Yang Chen;Wei Xu;Alan Ritter
共 17 条
CAREER: Large-Scale Learning for Information Extraction
-
批准号:1845670
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2019
-
负责人:Alan Ritter
-
依托单位:
CRII: III: Learning to Extract Events from Knowledge Base Revisions
-
批准号:1464128
-
项目类别:Standard Grant
-
资助金额:$15.13万
-
财政年份:2015
-
负责人:Alan Ritter
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:黄洛将
-
依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:黄洛将
-
依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
-
批准号:12074246
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2020
-
负责人:Yoshitomo Kamiya
-
依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
-
批准号:31972875
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:石江华
-
依托单位:
Large PB/PB小鼠 视网膜新生血管模型的研究
-
批准号:30971650
-
项目类别:面上项目
-
资助金额:8.0万元
-
批准年份:2009
-
负责人:周旻
-
依托单位:
基因discs large在果蝇卵母细胞的后端定位及其体轴极性形成中的作用机制
-
批准号:30800648
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2008
-
负责人:于玲珠
-
依托单位:
LARGE基因对口腔癌细胞中α-DG糖基化及表达的分子调控
-
批准号:30772435
-
项目类别:面上项目
-
资助金额:29.0万元
-
批准年份:2007
-
负责人:尚政军
-
依托单位: