课题基金 / 基金详情

CRII: III: RUI: Adaptive Query Processing for Crowd-Powered Database Systems

CRII: III: RUI: Adaptive Query Processing for Crowd-Powered Database Systems
CRII:III:RUI:众包数据库系统的自适应查询处理
批准号:
1657259
负责人:
Katherine Trushkowsky
金额:
$17.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-06-01 至 2020-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
数据库系统为用户提供了对系统存储的数据集合提出问题或查询的能力(例如,查找在公司工作至少两年的员工),并非常快速地提供答案。由于人们在现实世界的经验和感知,他们比计算机更有能力处理需要判断或数据解释的问题。一种基于人群的数据库系统使用一组被称为“人群”的人来帮助回答用户的查询,方法是招募他们使用主观的和/或需要视觉或语义解释的标准来处理数据。例如,用户可能想要找到一组教师招聘信息,其中的工作说明讨论了对多样性的承诺,学校位于安全的位置;解释每个工作说明和研究犯罪统计数据是非常适合人们执行的任务。该系统可以协调群组工作人员比用户单独处理数据更高效,这在有大量数据项要处理的情况下是有利的。虽然人群处理的查询可能需要几个小时或几天才能完成,但大众支持的数据库系统能够处理复杂的查询。例如,确定关于某一医疗设备的哪些研究文章包含将该设备与其他设备进行比较的实验结果,或者找出一组珠宝商中哪些只使用道德来源的金属和宝石,并将其运往阿拉斯加等查询。数据库系统旨在优化单个用户的查询处理效率。查询通常涉及多个部分,例如,对于职位发布查询,这些部分是(1)过滤出不描述对多样性的承诺的职位,以及(2)为不安全位置的学校过滤出职位。不满足第一个标准的作业不需要为第二个标准处理,反之亦然。查询各部分的处理顺序会影响需要多少计算以及处理查询需要多长时间。传统数据库系统具有关于查询部分将花费多长时间以及项目满足过滤器的可能性的信息;它们使用该信息为查询选择有效的处理顺序。然而,这一信息对于大众支持的数据库系统来说是未知的。优化器对大众数据库系统的有用性取决于它们在处理查询之前,当用户的查询信息未知时,它们能否找到有效的方法来处理用户的查询。本研究项目的目的是通过开发一个处理涉及多个过滤标准的查询的系统来应对这一挑战,该系统可以观察执行环境并在查询执行时调整其处理策略。该项目将产生广泛的影响,通过产生一个查询处理系统,使用户能够提出关于数据的更有趣的问题,推进在动态环境中分配人类计算资源的研究,以及在研究和系统设计原则方面培训一批本科生。本研究的目标是为查询时用于优化的重要统计数据未知的群体支持的过滤器查询构建一个基于成本的查询优化器。这些统计数据包括传统指标,如过滤器选择性,以及查询成本的新贡献者,如群组工作人员完成一个工作单元所需的时间,以及为主观评估达成共识所需的工作人员数量。该项目采用了一种自适应的查询处理方法:当查询运行时,系统观察成本和选择性信息,并定期重新排序查询计划操作符,以降低总体查询成本。研究人员将证明,他们的查询优化器产生的查询成本与基于人群的最佳查询计划的成本相当,其中选择性和主观性信息是先验已知的。源代码、论文和演示文稿可在项目网站(https://www.cs.hmc.edu/~beth/adaptivecrowd.shtml).上找到
英文摘要
Database systems provide users with the ability to ask questions, or queries, about collections of data that the system stores (e.g., find employees who had worked in the company for at least 2 years) and provide the answers very fast. People are better equipped than computers to tackle problems that require judgement or data interpretation due to their real-world experience and perception. A crowd-powered database system uses groups of people called "the crowd" to help with answering users' queries by recruiting them to process data using criteria that are subjective and/or require visual or semantic interpretation. For example, a user may want to find a set of faculty job postings in which the job description discusses a commitment to diversity and for which the school is in a safe location; interpreting each job description and researching crime statistics are tasks well-suited for people to perform. The system can coordinate crowd workers to process data more efficiently than the user alone could, which is advantageous when there are more than a handful of data items to process. While queries processed by the crowd may take hours or days to complete, crowd-powered database systems enable the processing of complex queries. For example, queries such as determining which research articles about a certain medical device contain experimental results comparing this and other devices, or finding out which of a set of jewelers only use ethically sourced metals and stones and also ship to Alaska. Database systems are designed to optimize the efficiency of query processing of individual users. A query often involves multiple parts, e.g., for the job postings query these are (1) filter out jobs that do not describe a commitment to diversity and (2) filter out jobs for schools in an unsafe location. A job that does not meet the first criterion does not need to be processed for the second one, and vice versa. The processing order for the parts of the query influences how much computation is needed and how long the query will take to process. Traditional database systems have information about how long parts of a query will take and the likelihood of items satisfying filters; they use this information to choose an efficient processing ordering for a query. However, this information is not known for crowd-powered database systems. The usefulness of optimizers for crowd-powered database systems hinges on their ability to find an efficient way to process a user's query when this information is unknown before processing the query. The aim of this research project is to tackle this challenge by developing a system to process queries involving multiple filtering criteria that observes the execution environment and adjusts its processing strategy as the query executes. This project will have broad impact by yielding a query processing system that will empower users to ask more interesting questions about data, advancing research in allocating human computation resources in dynamic environments, as well as training a group of undergraduate students both in research and in the principles of systems design.The goal of this research is to build a cost-based query optimizer for crowd-powered filter queries for which important statistics used in optimization are unknown at query time. These statistics include traditional metrics such as filter selectivity as well as new contributors to query cost such as the time it takes crowd workers to complete a unit of work and the number of workers needed to reach consensus for a subjective evaluation. The project takes an adaptive approach to query processing: while the query is running, the system observes cost and selectivity information and periodically reorders the query plan operators to reduce overall query cost. The researchers will demonstrate that their query optimizer yields query costs that are comparable to costs from the optimal crowd-based query plan for which selectivity and subjectivity information is known a priori. Source code, papers, and presentations are available on the project web site (https://www.cs.hmc.edu/~beth/adaptivecrowd.shtml).
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Dynamic Filter: Adaptive Query Processing with the Crowd
动态过滤器:群体的自适应查询处理
DOI: --
发表时间: 2017
期刊: Fifth AAAI Conference on Human Computation and Crowdsourcing
影响因子: --
作者: [Lan, Doren, Reed, Katherine, Shin, Austin, Trushkowsky, Beth]
通讯作者: Trushkowsky, Beth
国内基金
海外基金
基于人工智能与多组学的III期结核性脓胸CT“低密度线”形成机制及手术时机预测模型研究
基于MOF–CRISPR微流控平台的雄黄As(III)/As(V)价态识别与炮制耦合机制研究
  • 批准号:
    JCZRLH202600780
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
白术内酯III靶向IRF4-CD36轴通过调控脂质代谢重编程提升结直肠癌奥沙利铂敏感性的机制研究
  • 批准号:
    2026JJ82690
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张卓
  • 依托单位:
基于废水零排放的FeS-As(III)置换法从污酸中清洁脱砷处理技术研究
  • 批准号:
    2026JJ30130
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张二军
  • 依托单位: