RealPDBs: Realistic Data Models and Query Compilation for Large-Scale Probabilistic Databases
RealPDBs: Realistic Data Models and Query Compilation for Large-Scale Probabilistic Databases
批准号:
EP/R013667/1
负责人:
Thomas Lukasiewicz
金额:
$99.55万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --
中文摘要
近年来,学术界和工业界对以自动化方式从数据构建大规模概率知识库产生了浓厚的兴趣,产生了许多系统,如DeepDive、Nell、Yago、Freebase、微软的Probase和谷歌的Knowledge Vault。这些系统不断地在Web上爬行并提取结构化信息,从而用数百万个实体和数十亿元组填充它们的数据库。这些搜索和提取系统能在多大程度上帮助处理真实世界的用例?事实证明,这是一个开放式的问题。例如,DeepDive用于构建古生物学、地质学、医学遗传学和人类运动等领域的知识库。从更广泛的角度来看,建设大规模知识库的追求为人工智能研究带来了新的曙光。信息提取、自然语言处理(例如,问题回答)、关系和深度学习、知识表示和推理以及数据库等领域正朝着一个共同目标积极进取。大规模概率知识库的查询通常被认为是这些努力的核心,然而,除了这些成功的案例外,概率知识库仍然缺乏将隐藏在其中的一些有价值的知识传达给最终用户的基本机制,这严重限制了其在实践中的潜在应用。这些问题根源于(独立于元组的)概率数据库的语义,这些数据库用于编码大多数概率知识库。出于计算效率的原因,概率数据库通常基于强大的、不切实际的完整性假设,例如封闭世界假设、元组独立假设和缺乏常识知识。这些强烈的不切实际的假设不仅会导致不想要的后果,而且还会使概率数据库在知识库学习、完成和查询方面处于弱势地位。更具体地说,上述每个系统只编码了现实世界的一部分,这一描述必然是不完整的;这些系统不断地爬行Web,遇到新的来源,因此也就是新的事实,导致它们将这些事实添加到他们的数据库中。然而,当涉及到查询时,这些系统中的大多数采用封闭世界假设,即任何不存在于数据库中的事实被赋予概率0,因此被假设为不可能。作为一个密切相关的问题,通常的做法是将每个提取的事实视为一个独立的伯努利变量,即任何两个事实在概率上是独立的。例如,一个人出演电影是独立于这个人是演员的事实,这与知识领域的根本性质相冲突。此外,当前的概率数据库缺乏常识性知识(尤其是本体论),这些常识性知识经常被用于推理以从数据中推断出隐含的结果,并且对于在不受控制的环境(例如Web)中查询大规模的概率数据库通常是必不可少的。这一提议的主要目标是通过更现实的数据模型来增强大规模概率数据库(从而释放其全部数据建模潜力),同时保持其计算特性。我们计划为产生的概率数据库开发不同的语义,并分析它们的计算特性和难以处理的来源。我们还计划为它们设计实用的可扩展的查询回答算法,特别是基于知识编译技术的算法,扩展现有的知识编译方法,并开发基于张量分解和神经符号知识编译的新方法。我们还将生成一个原型实现,并对所提出的算法进行实验评估。
英文摘要
In the recent years, there has been a strong interest in academia and industry in building large-scale probabilistic knowledge bases from data in an automated way, which has resulted in a number of systems, such as DeepDive, NELL, Yago, Freebase, Microsoft's Probase, and Google's Knowledge Vault. These systems continuously crawl the Web and extract structured information, and thus populate their databases with millions of entities and billions of tuples. To what extent can these search and extraction systems help with real-world use cases? This turns out to be an open-ended question. For example, DeepDive is used to build knowledge bases for domains such as paleontology, geology, medical genetics, and human movement. From a broader perspective, the quest for building large-scale knowledge bases serves as a new dawn for artificial intelligence research. Fields such as information extraction, natural language processing (e.g., question answering), relational and deep learning, knowledge representation and reasoning, and databases are taking initiative towards a common goal. Querying large-scale probabilistic knowledge bases is commonly regarded to be at the heart of these efforts.Beyond all these success stories, however, probabilistic knowledge bases still lack the fundamental machinery to convey some of the valuable knowledge hidden in them to the end user, which seriously limits their potential applications in practice. These problems are rooted in the semantics of (tuple-independent) probabilistic databases, which are used for encoding most probabilistic knowledge bases. For computational efficiency reasons, probabilistic databases are typically based on strong, unrealistic completeness assumptions, such as the closed-world assumption, the tuple-independence assumption, and the lack of commonsense knowledge. These strong unrealistic assumptions do not only lead to unwanted consequences, but also put probabilistic databases on weak footing in terms of knowledge base learning, completion, and querying. More specifically, each of the above systems encodes only a portion of the real world, and this description is necessarily incomplete; these systems continuously crawl the Web, encounter new sources, and consequently new facts, leading them to add such facts to their database. However, when it comes to querying, most of these systems employ the closed-world assumption, i.e., any fact that is not present in the database is assigned the probability 0, and thus assumed to be impossible. As a closely related problem, it is common practice to view every extracted fact as an independent Bernoulli variable, i.e., any two facts are probabilistically independent. For example, the fact that a person starred in a movie is independent from the fact that this person is an actor, which is in conflict with the fundamental nature of the knowledge domain. Furthermore, current probabilistic databases lack (in particular ontological) commonsense knowledge, which can often be exploited in reasoning to deduce implicit consequences from data, and which is often essential for querying large-scale probabilistic databases in an uncontrolled environment such as the Web. The main goal of this proposal is to enhance large-scale probabilistic databases (and so to unlock their full data modelling potential) by more realistic data models, while preserving their computational properties. We are planning to develop different semantics for the resulting probabilistic databases and analyse their computational properties and sources of intractability. We are also planning to design practical scalable query answering algorithms for them, especially algorithms based on knowledge compilation techniques, extending existing knowledge compilation approaches and elaborating new ones, based on tensor factorisation and neural-symbolic knowledge compilation. We will also produce a prototype implementation and experimentally evaluate the proposed algorithms.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.4230/lipics.icdt.2020.5
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
作者:
[Antoine Amarilli;I. Ceylan]
通讯作者:
Antoine Amarilli;I. Ceylan
DOI:
10.1016/j.ins.2021.02.018
发表时间:
2021-02
期刊:
Inf. Sci.
影响因子:
--
作者:
[Elvira Amador-Domínguez;E. Serrano;Daniel Manrique;Patrick Hohenecker;Thomas Lukasiewicz]
通讯作者:
Elvira Amador-Domínguez;E. Serrano;Daniel Manrique;Patrick Hohenecker;Thomas Lukasiewicz
Approximate weighted model integration on DNF structures
DNF 结构上的近似加权模型集成
DOI:
10.1016/j.artint.2022.103753
发表时间:
2022
期刊:
Artificial Intelligence
影响因子:
14.4
作者:
[Abboud R]
通讯作者:
Abboud R
DOI:
10.3233/sw-180339
发表时间:
2020-04
期刊:
Semantic Web
影响因子:
3
作者:
[V. W. Anelli;R. Leone;T. D. Noia;Thomas Lukasiewicz;Jessica Rosati]
通讯作者:
V. W. Anelli;R. Leone;T. D. Noia;Thomas Lukasiewicz;Jessica Rosati
DOI:
10.46298/lmcs-18(1:2)2022
发表时间:
2022
期刊:
Logical Methods in Computer Science
影响因子:
0.6
作者:
[Amarilli A]
通讯作者:
Amarilli A
共 6 条
PrOQAW: Probabilistic Ontological Query Answering on the Web
-
批准号:EP/J008346/1
-
项目类别:Research Grant
-
资助金额:$103.7万
-
财政年份:2012
-
负责人:Thomas Lukasiewicz
-
依托单位:
海外基金