课题基金 / 基金详情

III: Medium: Constructing Knowledge Bases by Extracting Entity-Relations and Meanings from Natural Language via "Universal Schema"

III: Medium: Constructing Knowledge Bases by Extracting Entity-Relations and Meanings from Natural Language via "Universal Schema"
III:媒介:通过“通用模式”从自然语言中提取实体关系和含义来构建知识库
批准号:
1514053
负责人:
Andrew McCallum
金额:
$100.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2020-08-31

项目摘要

项目成果

Andrew McCallum的其他基金

相似基金

相关文献

中文摘要
翻译
从自然语言自动构建知识库(KB)对于(a)科学家(例如,建立基因和蛋白质知识库的兴趣由来已久),(B)社会科学家(例如,从文本数据构建社交网络)和(c)国防(犯罪分子和恐怖分子的网络分析已被证明是有用的)具有根本的重要性。知识库的核心是其对象(“实体”,如蛋白质、人、组织和地点)以及这些对象之间的联系(“关系”,如一种蛋白质增加另一种蛋白质的产量,或一个人为一个组织工作)。该项目旨在大大提高从文本中提取实体关系的准确性,以及提高关系类型之间许多细微区别的保真度。该项目的技术方法-我们称之为“通用模式”-是一种明显新颖的传统方法,基于将所有输入关系表达式表示为共同多维空间中的位置,附近的关系具有相似的含义。更广泛的影响将包括与工业界在具有经济重要性的应用方面的合作,与学术界非计算机科学家在多学科应用方面的合作,创建并公开发布新的数据集,供我们自己和他人进行基准评估(通过改进性能比较实现科学进步),创建并公开发布我们方法的开源实现(实现进一步的科学研究,易于大规模使用,快速商业化和第三方增强)。教育影响包括创建和教授关于科学知识库建设的新课程,组织关于嵌入,提取和知识表示的研究讲习班,以及培训多名本科生和研究生。大多数以前的研究关系提取福尔斯属于两类之一。在第一种情况下,必须定义一个预先固定的关系类型模式(如lives-in,employed-by和少数其他类型),这限制了表达能力并隐藏了语言歧义。在这里,训练机器学习模型要么依赖于标记的训练数据(这是稀缺和昂贵的),要么使用轻度监督的自我训练过程(这通常是脆弱的,并且随着额外的迭代而远离真相)。在第二类中,基于语言字符串本身提取到“开放”模式中(缺乏在它们之间进行概括的能力),或者试图通过这些字符串的无监督聚类来获得概括(遭受无法捕获可靠同义词的聚类,甚至根本无法找到所需的语义)。该项目提出了“通用模式”的关系提取的研究,其中我们学习所有输入模式的联合的概括模型,包括多个可用的预结构知识库以及所有观察到的自然语言表面形式。因此,这种方法包含了原始语言表面形式的多样性和模糊性(不试图将关系强制放入预定义的框中),但也成功地通过学习显式和隐式关系之间的非对称含义来推广,使用新的扩展概率矩阵分解和向量嵌入方法,这些方法在NetFlix奖竞赛中非常成功。通用模式提供了几乎无限的关系类型多样性(由于表面形式),并通过与现有结构化数据(即,现有数据库的关系类型)。在初步的实验中,该方法已经超过了以前的国家的最先进的关系提取方法的基准任务上的一个很大的保证金。新提出的研究包括新的训练过程,新的表示方法,包括相同表面形式的多个意义以及方差嵌入,新的约束方法,实体和关系类型之间的联合推理,非二进制和高阶关系的新模型,以及通过并行分布的可扩展性。项目网站(http://www.iesl.cs.umass.edu/projects/NSF_USchema.html)将包括关于该项目的信息,并提供数据集、源代码和文件、教学和讲习班材料以及出版物。此外,数据集将通过UCI机器学习存储库(或其他类似的机器学习数据存档位置)传播,以促进与其他研究人员的共享并确保长期可用性,GitHub将用于促进代码的发布,共享和存档。
英文摘要
Automated knowledge base (KB) construction from natural language is of fundamental importance to (a) scientists (for example, there has been long-standing interest in building KBs of genes and proteins), (b) social scientists (for example, building social networks from textual data), and (c) national defense (where network analysis of criminals and terrorists have proven useful). The core of a knowledge base is its objects ("entities", such as proteins, people, organizations and locations) and its connections between these objects ("relations", such as one protein increasing production of another, or a person working for an organization). This project aims to greatly increase the accuracy with which entity-relations can be extracted from text, as well as increase the fidelity which many subtle distinctions among types of relations can be represented. The project's technical approach -- which we call "universal schema" -- is a markedly novel departure from traditional methods, based on representing all of the input relation expressions as positions in a common multi-dimensional space, with nearby relations having similar meanings. Broader impacts will include collaboration with industry on applications of economic importance, collaboration with academic non-computer-scientists on a multidisciplinary application, creating and publicly releasing new data sets for benchmark evaluation by ourselves and others (enabling scientific progress through improved performance comparisons), creating and publicly releasing an open-source implementation of our methods (enabling further scientific research, easy large-scale use, rapid commercialization and third-party enhancements). Education impacts include creating and teaching a new course on knowledge base construction for the sciences, organizing a research workshop on embeddings, extraction and knowledge representation, and training multiple undergraduates and graduate students. Most previous research in relation extraction falls into one of two categories. In the first, one must define a pre-fixed schema of relation types (such as lives-in, employed-by and a handful of others), which limits expressivity and hides language ambiguities. Training machine learning models here either relies on labeled training data (which is scarce and expensive), or uses lightly-supervised self-training procedures (which are often brittle and wander farther from the truth with additional iterations). In the second category, one extracts into an "open" schema based on language strings themselves (lacking ability to generalize among them), or attempts to gain generalization with unsupervised clustering of these strings (suffering from clusters that fail to capture reliable synonyms, or even find the desired semantics at all). This project proposes research in relation extraction of "universal schema", where we learn a generalizing model of the union of all input schemas, including multiple available pre-structured KBs as well as all the observed natural language surface forms. The approach thus embraces the diversity and ambiguity of original language surface forms (not trying to force relations into pre-defined boxes), yet also successfully generalizes by learning non-symmetric implicature among explicit and implicit relations using new extensions to the probabilistic matrix factorization and vector embedding methods that were so successful in the NetFlix prize competition. Universal schema provide for a nearly limitless diversity of relation types (due to surface forms), and support convenient semi-supervised learning through integration with existing structured data (i.e., the relation types of existing databases). In preliminary experiments, the approach already surpassed by a wide margin the previous state-of-the-art relation extraction methods on a benchmark task. New proposed research includes new training processes, new representations that include multiple-senses for the same surface form as well as embeddings with variances, new methods of incorporating constraints, joint inference between entity- and relation-types, new models of non-binary and higher-order relations, and scalability through parallel distribution. The project web site (http://www.iesl.cs.umass.edu/projects/NSF_USchema.html) will include information on the project and provide access to data sets, source code and documentation, teaching and workshop materials, and publications. In addition, datasets will be disseminated via UCI Machine Learning Repository (or other similar archive location for machine learning data) to facilitate sharing with other researchers and ensure long-term availability, and GitHub will be used to facilitate release, sharing, and archiving of code.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: SOS-DCI / HNDS-R: Advancing Semantic Network Analysis to Better Understand How Evaluative Exchanges Shape Scientific Arguments
  • 批准号:
    2244805
  • 项目类别:
    Standard Grant
  • 资助金额:
    $22.5万
  • 财政年份:
    2023
  • 负责人:
    Andrew McCallum
  • 依托单位:
RI: Medium: Probabilistic Box Embeddings
  • 批准号:
    2106391
  • 项目类别:
    Standard Grant
  • 资助金额:
    $84.99万
  • 财政年份:
    2021
  • 负责人:
    Andrew McCallum
  • 依托单位:
DMREF: Collaborative Research: The Synthesis Genome: Data Mining for Synthesis of New Materials
  • 批准号:
    1922090
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2019
  • 负责人:
    Andrew McCallum
  • 依托单位:
RI: Medium: Extreme Clustering
  • 批准号:
    1763618
  • 项目类别:
    Standard Grant
  • 资助金额:
    $110.39万
  • 财政年份:
    2018
  • 负责人:
    Andrew McCallum
  • 依托单位:
海外基金