CAREER: Large-Scale Learning for Information Extraction
CAREER: Large-Scale Learning for Information Extraction
批准号:
1845670
负责人:
Alan Ritter
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-01 至 2020-10-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Much of human knowledge is encoded in text. This project aims to substantially advance the capability of machines to read large document collections and reason about the knowledge contained within them using minimal human effort. This will help people to overcome information overload and make better decisions by analyzing vital information that is locked away in unstructured text. Recent years have seen tremendous progress on tasks such as speech recognition and machine translation, by applying deep learning methods on massive, high-quality datasets; however, most available datasets for information extraction are either small or very noisy. The project will address these challenges by developing new methods that can learn more effectively from big, but noisy datasets that are constructed using distant supervision from an existing knowledge base (KB). To demonstrate the new methods' effectiveness, they will be used to support several novel applications. These include the detection of cyber-threats reported online and the analysis of experts' opinions about their severity. Recent studies have found that 75% of software vulnerabilities are first reported online, giving attackers time to exploit the vulnerability. Systems that can automatically read computer security blogs and analyze new threats could help security practitioners to track and prioritize them more effectively. The project includes a plan for integrating research and education. Outreach efforts aim to help attract a more diverse group of students to study computer science. These include hands-on workshops to expose freshmen to exciting natural language processing and artificial intelligence applications. The project will also help to engage advanced undergraduate students in research through new course materials on cutting-edge information extraction techniques.The research will address the machine reading data bottleneck by inventing new methods that can learn effectively from large, noisy datasets using distant supervision. These methods will address the challenge of label noise inherent in distant supervision by performing inference over latent variables during learning, filling in missing information, and resolving ambiguities. The approach combines the benefits of structured learning and neural networks; the structured learning component of the model can override noisy labels in cases where it is sufficiently confident -- this is balanced against a model of missing data in the KB. This will catalyze the rapid development of extractors for many new tasks and domains. To demonstrate this, extensive experiments will compare against state of the art methods using standard benchmark datasets for information extraction, including the Freebase/NYT corpus, TAC KBP datasets, and TACRED. Furthermore, the research will push the boundaries of minimal supervision for Information Extraction by exploring new applications that demonstrate the generality of the approach, including entity, relation and event extraction, time normalization and learning to extract a real-time feed of cyber-threat intelligence using distant supervision from the National Vulnerability Database (NVD). These applications are supported by a comprehensive evaluation plan that includes the development of new corpora and metrics. The project will produce a number of new datasets in addition to a toolkit for minimally supervised information extraction, that will be shared as open source software. This research effort will support the rapid development of information systems for a broad range of new tasks and domains using minimal human effort.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Large-Scale Learning for Information Extraction
-
批准号:2052498
-
项目类别:Continuing Grant
-
资助金额:$48.9万
-
财政年份:2020
-
负责人:Alan Ritter
-
依托单位:
CRII: III: Learning to Extract Events from Knowledge Base Revisions
-
批准号:1464128
-
项目类别:Standard Grant
-
资助金额:$15.13万
-
财政年份:2015
-
负责人:Alan Ritter
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:黄洛将
-
依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:黄洛将
-
依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
-
批准号:12074246
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2020
-
负责人:Yoshitomo Kamiya
-
依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
-
批准号:31972875
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:石江华
-
依托单位:
Large PB/PB小鼠 视网膜新生血管模型的研究
-
批准号:30971650
-
项目类别:面上项目
-
资助金额:8.0万元
-
批准年份:2009
-
负责人:周旻
-
依托单位:
基因discs large在果蝇卵母细胞的后端定位及其体轴极性形成中的作用机制
-
批准号:30800648
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2008
-
负责人:于玲珠
-
依托单位:
LARGE基因对口腔癌细胞中α-DG糖基化及表达的分子调控
-
批准号:30772435
-
项目类别:面上项目
-
资助金额:29.0万元
-
批准年份:2007
-
负责人:尚政军
-
依托单位: