III: Medium: Collaborative Research: U4U - Taming Uncertainty with Uncertainty-Annotated Databases
III: Medium: Collaborative Research: U4U - Taming Uncertainty with Uncertainty-Annotated Databases
批准号:
1956123
负责人:
Boris Glavic
金额:
$46.66万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30
中文摘要
无论数据的大小、应用领域或分析类型如何,不确定性在数据分析中都是普遍存在的。不确定性的常见来源包括缺失值、传感器误差、偏差、异常值和许多其他因素。经典的确定性数据管理不跟踪不确定性,因此需要在数据被摄取到系统之前解决数据质量问题,这通常是不可行的。最终的结果是,本质上不确定的数据被视为确定的。然而,如果忽视数据的不确定性,就会导致难以追踪的错误,这反过来又会对现实世界产生严重影响,例如毫无根据的科学发现、经济损失,甚至是基于错误数据的医疗决策。虽然存在管理不完整数据的技术,但这些技术对于实际使用来说通常过于沉重,可能会对用户隐藏相关信息。该项目的目标是开发用于管理不确定数据的轻量级技术,从而使广泛的应用程序能够管理不确定性。当前用于管理不确定数据的方法通常计算成本很高,并且仅适用于有限类型的查询。计划中的研究将产生管理不确定数据的新方法,弥合确定性和不完全数据管理之间的差距。该项目的基础是不确定性注释数据库,它用不确定性标签丰富数据,并为通过查询传播这些标签提供语义。其结果是经典数据管理的严格泛化,它结合了确定性数据管理的性能、通用性和易用性,以及不完整数据库技术的强正确性保证。实现这一目标非常重要,因为对不确定数据的查询求值是难以处理的,即使对于相对简单的不确定数据模型和受限的查询类也是如此。将探讨三个主要的研究重点,以解决开发这种技术的主要挑战:(i)不确定性注释数据库将扩展为属性级注释和对可能答案的过度近似的紧凑编码。这使得该方法能够处理丢失的数据,并处理非单调查询,如聚合查询;(ii)将发展紧逼近不完全数据库的方法,以处理对不确定数据的查询所产生的大量甚至无限组可能的结果;(iii)将开发用于不确定性注释数据库查询评估的优化算法,以解决不确定数据查询的性能限制。计划中的工作将大大提高不确定数据管理的最新水平,首次以合理的成本实现复杂查询的原则性不确定性管理。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Uncertainty is prevalent in data analysis, no matter what the size of the data, the application domain, or type of analysis. Common sources of uncertainty include missing values, sensor errors, bias, outliers, and many other factors. Classical deterministic data management does not track uncertainty and, thus requires data quality issues to be resolved before data is ingested into the system, which is often not feasible. The net effect is that inherently uncertain data is being treated as certain. However, if ignored, data uncertainty results in hard to trace errors, which in turn can have severe real world implications such as unfounded scientific discoveries, financial damages, or even medical decisions based on incorrect data. While there exist techniques for managing incomplete data, these techniques are generally too heavy-weight for real-world usage and may hide relevant information from users. The goal of this project is to develop light-weight techniques for managing uncertain data that empower a wide range of applications to manage uncertainty.Current methods for managing uncertain data are often computationally expensive and are only applicable to limited types of queries. The planned research will result in novel methods for managing uncertain data that bridge the gap between deterministic and incomplete data management. The foundation of this project are uncertainty-annotated databases, which enrich data with uncertainty labels and provide semantics for propagating these labels through queries. The result is a strict generalization of classical data management that combines the performance, generality, and ease-of-use of deterministic data management with the strong correctness guarantees of incomplete database techniques. Achieving this goal is highly non-trivial, because query evaluation over uncertain data is intractable, even for relatively simple uncertain data models and restricted classes of queries. Three main research thrusts will be explored that address the main challenges in developing such a technique: (i) uncertainty-annotated databases will be extended with attribute-level annotations and an compact encoding of an over-approximation of possible answers. This enables the approach to handle missing data and to deal with non-monotone queries such as queries with aggregation; (ii) methods to compactly approximating incomplete databases will be developed to deal with the large or even infinite sets of possible results produced by queries over uncertain data; (iii) optimized algorithms for query evaluation over uncertainty-annotated databases will be developed to address the performance limitations of queries over uncertain data. The planned work will significantly enhance the state-of-the-art in uncertain data management by, for the first time, enabling principled uncertainty management for complex queries at a reasonable cost.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(17)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
CaJaDE: explaining query results by augmenting provenance with context
CaJaDE:通过使用上下文增强来源来解释查询结果
DOI:
10.14778/3554821.3554852
发表时间:
2022
期刊:
Proceedings of the VLDB Endowment
影响因子:
2.5
作者:
[Li, Chenjie, Lee, Juseung, Miao, Zhengjie, Glavic, Boris, Roy, Sudeepa]
通讯作者:
Roy, Sudeepa
Efficient Approximation of Certain and Possible Answers for Ranking and Window Queries over Uncertain Data
对不确定数据进行排序和窗口查询的某些和可能答案的有效近似
DOI:
10.14778/3583140.3583151
发表时间:
2023
期刊:
Proceedings of the VLDB Endowment
影响因子:
2.5
作者:
[Feng, Su, Glavic, Boris, Kennedy, Oliver]
通讯作者:
Kennedy, Oliver
DOI:
10.1145/3514221.3526138
发表时间:
2022
期刊:
ACM SIGMOD
影响因子:
--
作者:
[Campbell, Felix S., Arab, Bahareh Sadat, Glavic, Boris]
通讯作者:
Glavic, Boris
Overlay Spreadsheets
叠加电子表格
DOI:
10.1145/3597465.3605220
发表时间:
2023
期刊:
HILDA '23: Proceedings of the Workshop on Human-In-the-Loop Data Analytics
影响因子:
--
作者:
[Kennedy, Oliver, Glavic, Boris, Brachmann, Michael]
通讯作者:
Brachmann, Michael
Hybrid Query and Instance Explanations and Repairs
混合查询和实例解释和修复
DOI:
10.1145/3543873.3587565
发表时间:
2023
期刊:
TaPP workshop - WWW '23 Companion: Companion Proceedings of the ACM Web Conference 2023
影响因子:
--
作者:
[Lee, Seokki, Glavic, Boris, Chapman, Adriane, Ludäscher, Bertram]
通讯作者:
Ludäscher, Bertram
共 17 条
III : Medium: Collaborative Research: From Open Data to Open Data Curation
-
批准号:2420691
-
项目类别:Standard Grant
-
资助金额:$37.5万
-
财政年份:2024
-
负责人:Boris Glavic
-
依托单位:
III : Medium: Collaborative Research: From Open Data to Open Data Curation
-
批准号:2107107
-
项目类别:Standard Grant
-
资助金额:$37.5万
-
财政年份:2021
-
负责人:Boris Glavic
-
依托单位:
海外基金