课题基金 / 基金详情

III: Small: Query Compilation on Probabilistic Databases

III: Small: Query Compilation on Probabilistic Databases
III:小:概率数据库上的查询编译
批准号:
1115188
负责人:
Dan Suciu
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-08-01 至 2015-07-31

项目摘要

项目成果

Dan Suciu的其他基金

相似基金

相关文献

中文摘要
翻译
概率数据库的目标是管理数据不确定的大型数据库。应用包括Web规模的信息提取、RFID系统、科学数据管理、生物医学数据集成、商业智能、数据清理、近似模式映射和重复数据删除。尽管对概率数据库有着巨大的需求和最近的密集研究,但到目前为止还没有健壮的概率数据库系统存在。原因是,概率推理问题总体上是棘手的。幸运的是,在数据库中,概率推理问题有两个不同的输入:查询和数据库实例。这导致了最近安全查询的发现,安全查询是可以在任何输入数据库上有效地评估的查询,并且产生了利用查询结构的新的概率推理算法。然而,不安全的查询仍然是概率数据库中的一大挑战,本项目研究了在概率数据库上评估不安全查询的新算法,并保证了性能。它使用了一种新的方法,查询编译,该方法将查询转换为以下四个目标之一:OBDDS、FBDDS、d-DNNF和使用包含/排除节点的电路。该项目追求两个目标:(1)它开发了依赖于实例的编译技术,大大扩展了安全查询中使用的独立于实例的技术的范围。(2)开发了近似查询编译技术,即使是在难以处理的查询实例对上,也能以牺牲准确性为代价,始终高效地运行。这些算法是保守的,因为当输入查询实例对容易处理时,它们在所有情况下都返回正确的概率。该项目的智能价值包括将查询编译到四个编译目标之一的新技术,OBDD,FBDD,d-DNNF,以及基于包含/排除的推理,使用精确推理(没有性能保证)和近似推理(有性能保证)。它扩展了我们对概率推理的理解,并为概率数据库引擎提供了实用的方法。作为更广泛的影响,该项目使需要对不确定数据进行通用管理的一大类应用程序受益,从大规模信息提取系统到科学数据管理,再到商业智能。该项目逐渐将概率数据的主题纳入研究生水平教育的课程中;查询编译已经在PI关于概率数据库的书(http://dx.doi.org/10.2200/S00362ED1V01Y201105DTM016),,研究生水平的教科书)中讨论过。有关更多信息,请参阅项目网站的网址:http://www.cs.washington.edu/homes/suciu/project-querycompilation.html
英文摘要
The goal of probabilistic databases is to manage large databases where the data is uncertain. Applications include Web-scale information extraction, RFID systems, scientific data management, biomedical data integration, business intelligence, data cleaning, approximate schema mappings, and data deduplication. Despite the huge demand and the intense recent research on probabilistic databases, no robust probabilistic database systems exist to date. The reason is that the probabilistic inference problem is, in general, intractable. Fortunately, in databases there are two distinct inputs to the probabilistic inference problem: the query and the database instance. This has led recently to the discovery of safe queries, which are queries that can be evaluated efficiently on any input database, and to new probabilistic inference algorithms that exploit the structure of the query. However, unsafe queries remain a major challenge in probabilistic databases.This project studies novel algorithms for evaluating unsafe queries on probabilistic database, with guaranteed performance. It uses a novel approach, query compilation, which translates the query into one of four targets: OBDDs, FBDDs, d-DNNFs, and circuits using inclusion/exclusion nodes. The project pursues two thrusts: (1) It develops instance-dependent compilation techniques that significantly extend the reach of instance-independent techniques used in safe queries. (2) It develops approximate query compilation techniques , which always run efficiently, even on intractable query, instance pairs, by sacrificing accuracy. These algorithms are conservative, in the sense that they return correct probabilities in all cases when the input query, instance pair is tractable.The Intellectual Merit of this project consists of new techniques for compiling queries into one of four compilation targets, OBDD, FBDD, d-DNNF, and inclusion/exclusion-based inference, using both exact inference (without performance guarantees), and approximate inference (with performance guarantees). It expands our understanding of probabilistic inference, and leads to practical approaches for probabilistic database engines. As Broader Impact, the project benefits a large class of applications that require general purpose management of uncertain data, ranging from large-scale information extraction systems, to scientific data management, to business intelligence. The project gradually incorporates topics from probabilistic data into into a curriculum for graduate level education; query compilation is already discussed in the PI's book on Probabilistic Databases ( http://dx.doi.org/10.2200/S00362ED1V01Y201105DTM016), a graduate-level textbook.For further information see the project web site at the URL: http://www.cs.washington.edu/homes/suciu/project-querycompilation.html
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Small: Datalog with Aggregates: Complexity, Optimization, Evaluation
  • 批准号:
    2314527
  • 项目类别:
    Standard Grant
  • 资助金额:
    $60.0万
  • 财政年份:
    2023
  • 负责人:
    Dan Suciu
  • 依托单位:
NSF-BSF: III: Small: Data Driven Schema
  • 批准号:
    2109922
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2021
  • 负责人:
    Dan Suciu
  • 依托单位:
III: Medium: Collaborative Research: Reasoning about Optimizers for Data-Intensive Systems
  • 批准号:
    1954222
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2020
  • 负责人:
    Dan Suciu
  • 依托单位:
III:Small: Optimal Query Processing meets Information Theory: from Proofs to Algorithms
  • 批准号:
    1907997
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2019
  • 负责人:
    Dan Suciu
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: