MAGIC: MAnaGing InComplete Data - New Foundations
MAGIC: MAnaGing InComplete Data - New Foundations
批准号:
EP/N023056/1
负责人:
Leonid Libkin
金额:
$145.27万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2016
资助国家:
英国
项目状态:
已结题
起止时间:
2016 至 --
中文摘要
在我们这个数据驱动的世界里,人们几乎一天都要使用高度复杂的软件系统,我们已经学会了依赖这些系统,但同时也知道这些系统会产生不正确的结果。这些是我们笔记本电脑上的系统,它们为公司的网站提供动力,并使公司和政府办公室保持运行。然而,不正确的行为是内置的,它是多个标准的一部分,几乎没有做出什么努力来改变事情。这些系统是商业DBMS(数据库管理系统)。只要它们存储的信息是完整的,我们就可以依赖它们。在自治环境中,这是一个合理的假设,但如今数据是由大量用户和应用程序生成的,其不完整性是生活中的一个事实。当不完整进入画面的那一刻,一切都改变了:发生了意想不到的行为;我们教学生写作并给他们打满分以停止工作的质疑,人们对从这些数据中得到的结果失去了信任。更糟糕的是,许多现代数据应用程序,包括数据集成、数据交换、基于本体的数据访问、数据质量、不一致性管理和许多其他应用程序,都具有内在的不完备性,并试图依赖标准技术来处理它们。这不可避免地会导致有问题的结果,有时甚至是明显不正确的结果。这背后的关键原因是复杂性和正确性之间的权衡:当前确保正确性的技术需要付出巨大的复杂性代价,应用程序会寻找绕过它的方法,牺牲过程中的正确性。我们的主要目标是结束这种令人遗憾的状况。正确性和效率可以并存,但我们需要开发不完全信息领域的新基础,以及基于这些基础的一套新技术,以协调这两者。要做到这一点,我们需要重新思考该领域的基础;关键是,我们需要理解回答具有正确性保证的不完整数据查询意味着什么。经典理论使用了一个单一的、适用于所有人的定义,仔细检查一下,它似乎并不是普遍正确的。我们有一种方法,可以让我们开发正确的理论,并将其应用于标准的数据管理任务及其新的应用程序。使这一方法切实可行是至关重要的。商业系统专注于高效的评估,而牺牲了正确性。我们提供正确答案的解决方案将在现有系统的基础上实现,以保证其适用性。将不需要抛弃现有的产品和技术来利用处理不完整信息的新方法。我们最初的目标是交付用于修复商业DBMS问题的解决方案,即展示如何以低且可接受的成本提供正确性保证。我们将对各种各样的问题这样做,远远超出目前已知的可能。在此之后,我们将介绍一些应用程序,这些应用程序将第一次确保对集成和交换的数据进行非常有表现力的查询的正确性。我们将在数据模型(例如,图形、XML、NoSQL数据库)和应用程序(不一致的数据、本体)方面进一步扩展。我们还将考虑考虑大量数据的解决方案,并在这些情况下产生近似答案。随着我们开发的工具包,“不完全信息的诅咒”,即认为不可能同时实现正确性和效率,应该是过去的事情了。
英文摘要
In our data-driven world, one can hardly spend a day without using highly complex software systems that we have learned to rely on, but that at the same time are known to produce incorrect results. These are systems we have on our laptops, they power websites of companies, and they keep companies and government offices running. And yet the incorrect behaviour is built into them, it is a part of multiple standards, and very little effort is made to change things.The systems are commercial DBMSs (database management systems). We can rely on them as long as information they store is complete. In an autonomous environment, this is a reasonable assumption, but these days data is generated by a huge number of users and applications, and its incompleteness is a fact of life. The moment incompleteness enters the picture, everything changes: unexpected behaviour occurs; queries that we teach students to write and give them full marks for stop working, and one loses trust in the results one gets from such data. To make matters worse, many modern applications of data, including data integration, data exchange, ontology based data access, data quality, inconsistency management and a host of others, have incompleteness built into them, and try to rely on standard techniques for handling it. This inevitably leads to questionable, or sometimes plain incorrect results. The key reason behind this is the complexity vs correctness tradeoff: current techniques guaranteeing correctness carry a huge complexity price, and applications look for ways around it, sacrificing correctness in the process.Our main goal is to end this sorry state of affairs. Correctness and efficiency can co-exist, but we need to develop new foundations of the field of incomplete information, and a new set of techniques based on these foundations, to reconcile the two.To do so, we need to rethink the very basics of the field; crucially, we need to understand what it means to answer queries over incomplete data with correctness guarantees. The classical theory uses a single one-size-fits-all definition, that, upon a careful examination, does not appear to be universally correct. We have an approach that will let us develop a proper theory of correctness and apply it in standard data management tasks and their new applications as well. It is crucial to make this approach practical. Commercial systems concentrate on efficient evaluation, sacrificing correctness. Our solutions for delivering correct answers will be implementable on top of existing systems, to guarantee their applicability. There will be no need to throw away existing products and techniques to take advantage of new approaches to handling incomplete information. Our initial set of goals is to deliver solutions for fixing problems with commercial DBMSs, namely to show how to provide correctness guarantees at low and acceptable costs. We shall do so for a wide variety of queries, going far beyond what is now known to be possible. After that, we shall look at applications that will let us, for the first time, ensure correctness of very expressive queries over integrated and exchanged data. We shall expand further, both in terms of data models (e.g., graphs, XML, noSQL databases), and applications (inconsistent data, ontologies). We shall also look at solutions that take into account very large amounts of data, and produce approximate answers in those scenarios.With the toolkit we develop, the "curse of incomplete information", i.e., the perceived impossibility of achieving correctness and efficiency simultaneously, should be a thing of the past.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3092931.3092933
发表时间:
2017
期刊:
ACM SIGMOD Record
影响因子:
--
作者:
[Abiteboul S]
通讯作者:
Abiteboul S
DOI:
10.1145/3448016.3457561
发表时间:
2021-06
期刊:
Proceedings of the 2021 International Conference on Management of Data
影响因子:
--
作者:
[Renzo Angles;A. Bonifati;Stefania Dumbrava;G. Fletcher;Keith W. Hare;J. Hidders;Victor E. Lee;Bei Li;L. Libkin;W. Martens;Filip Murlak;Josh Perryman;Ognjen Savkovic;Michael Schmidt;Juan Sequeda;Dominik Tomaszuk]
通讯作者:
Renzo Angles;A. Bonifati;Stefania Dumbrava;G. Fletcher;Keith W. Hare;J. Hidders;Victor E. Lee;Bei Li;L. Libkin;W. Martens;Filip Murlak;Josh Perryman;Ognjen Savkovic;Michael Schmidt;Juan Sequeda;Dominik Tomaszuk
Counting Database Repairs under Primary Keys Revisited
重新审视主键下的数据库修复计数
DOI:
10.1145/3294052.3319703
发表时间:
2019
期刊:
影响因子:
--
作者:
[Calautti M]
通讯作者:
Calautti M
Approximating Certainty in Querying Data and Metadata
查询数据和元数据的近似确定性
DOI:
--
发表时间:
2018
期刊:
影响因子:
--
作者:
[Civili C]
通讯作者:
Civili C
DOI:
10.1145/3375395.3387970
发表时间:
2020-05
期刊:
Proceedings of the 39th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems
影响因子:
--
作者:
[Marco Console;P. Guagliardo;L. Libkin;Etienne Toussaint]
通讯作者:
Marco Console;P. Guagliardo;L. Libkin;Etienne Toussaint
共 8 条
Querying Graph Structured Data: Principles and Techniques
-
批准号:EP/J015377/1
-
项目类别:Research Grant
-
资助金额:$81.18万
-
财政年份:2012
-
负责人:Leonid Libkin
-
依托单位:
XML with Incomplete Information: Representation, Querying, and Applications
-
批准号:EP/G049165/1
-
项目类别:Research Grant
-
资助金额:$72.06万
-
财政年份:2009
-
负责人:Leonid Libkin
-
依托单位:
Relational and XML Data Exchange: Semantics, Consistency, and Query Answering
-
批准号:EP/E005039/1
-
项目类别:Research Grant
-
资助金额:$58.3万
-
财政年份:2007
-
负责人:Leonid Libkin
-
依托单位:
海外基金