CAREER: Knowledge Extraction and Discovery from Massive Text Corpora via Extremely Weak Supervision
CAREER: Knowledge Extraction and Discovery from Massive Text Corpora via Extremely Weak Supervision
批准号:
2239440
负责人:
Jingbo Shang
金额:
$60.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2028-06-30
中文摘要
自动化的知识提取和发现方法可以满足不同用户(例如,用于决策的政府和用于文献摘要的科学家)的不同需求。一个基本的悬而未决的问题是自动化方法需要多少用户努力才能获得有用的知识。该项目旨在通过一个新提出的模式--极弱的监督--最大限度地减少这种所需的用户工作--它只包括简短的自然语言用户输入来定义任务(例如,对新闻文章进行分类时的主题列表;对事件进行分类时的地点名称),类似于可能提供给人工注释员的针对具体任务的准则的指导。通过使用简短的自然语言输入而不是劳动密集型的带注释的训练样本,这一新的范式将有助于使知识提取和发现民主化,并将其应用从富有的公司扩展到具有广泛需求的普通、相对未经培训的用户(例如,领域科学家和小企业主)。项目成果将通过顶级会议和学术出版物传播,并纳入新的课程。该项目还将支持不同的研究生、本科生和高中生。该项目专注于四个基本的、相互关联的知识提取和发现任务,即文本分类、短语挖掘、命名实体识别和关系提取。遵循极弱的监督范式,本项目将开发一系列新的方法,包括:(1)多文法和单文法(新兴)短语的无监督短语标注方法;(2)只能以最流行(例如,前50%)类名作为输入的文本分类方法,以发现新类(即,新类不是由用户明确定义的)并为所有类构建分类器;(3)命名实体识别方法,其可以利用几个流行的实体类型和感兴趣的提及来识别相同/相似类型的(新兴)实体提及;(4)一种关系提取方法,可以从几种流行的关系类型和感兴趣的元组中发现语义相似的关系,并提取相关的元组。所有这些方法在设计上都与领域和语言无关,只需要在特定领域和语言中提供预先训练的神经语言模型。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Automated knowledge extraction and discovery methods can address the diverse needs of different users (e.g., governments for decision making and scientists for literature summary). A fundamental open problem is how much user effort automated methods require to obtain useful knowledge. This project aims to minimize such required user effort with a newly proposed paradigm, extremely weak supervision – It includes only brief natural-language user input to define the task (e.g., a list of topics when classifying news articles; location names when classifying events), guidance similar to task-specific guidelines that might be provided to human annotators. By using brief natural-language input instead of labor-intensive annotated training samples, this new paradigm will help democratize knowledge extraction and discovery, and extend its application beyond rich companies to ordinary, relatively untrained users with a broad range of needs (e.g., domain scientists and small business owners). Project outcomes will be disseminated via top conferences and scholarly publications and integrated into new courses. This project will also support a diverse set of graduate, undergraduate, and high school students.This project focuses on four fundamental, interconnected knowledge extraction and discovery tasks, i.e., text classification, phrase mining, named entity recognition, and relation extraction. Following the extremely weak supervision paradigm, this project will develop a series of novel methods, including (1) an unsupervised phrase tagging method for both multi-gram and unigram (emerging) phrases, (2) a text classification method that can take only the most popular (e.g., top-50%) class names as input to discover novel classes (i.e., new classes are not explicitly defined by the user) and build a classifier for all the classes; (3) a named entity recognition method that can take a few popular entity types and mentions of interest to recognize (emerging) entity mentions of the same/similar types; and (4) a relation extraction method that can take a few popular relation types and tuples of interest to discover relations of similar semantics and extract relevant tuples. All these methods, by design, will be agnostic to domains and languages and require only the availability of pre-trained neural language models in a particular domain and language.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NSF Convergence Accelerator Track D: Towards Intelligent Sharing and Search for AI Models and Datasets
-
批准号:2040727
-
项目类别:Standard Grant
-
资助金额:$94.72万
-
财政年份:2020
-
负责人:Jingbo Shang
-
依托单位:
海外基金