III: Medium: 20/20: A System for Human-in-the-Loop Data Exploration
III: Medium: 20/20: A System for Human-in-the-Loop Data Exploration
批准号:
1514491
负责人:
Stanley Zdonik
金额:
$100.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2019-08-31
中文摘要
探索性数据分析在包括科学、工程和商业在内的广泛领域的数据驱动发现中起着关键作用。在用户群不断扩大和多样化的时期,为了使数据分析成为一种商品,人类的生产力和易用性必须成为任何数据库系统的首要设计考虑因素。不幸的是,用户友好且旨在提高人类生产力的数据工具仍然非常缺乏。该项目将使不同技能水平的用户能够比现在更容易、更快地与他们的大型数据集进行交互和探索。而不是花费大量宝贵的时间来构建复杂的分析任务,这项工作将提供一个更灵活,响应和用户友好的系统,基于直接操作的可视化表示(例如,图表,图形,地图)的数据集和分析结果。该系统还可以用作学习工具:例如,教师可以引导学生浏览复杂的数据集以验证特定的假设。这个项目将使更多的用户更容易获得大规模的数据探索。总体而言,它将加速电子商务、金融和科学等许多领域的发现和突破。本研究将纳入本科和研究生课程。外联活动包括针对本科生和高中女生的特别研究和以教育为重点的项目。该项目将建立一类新的数据库系统,设计用于人在循环(HIL)操作。这项工作的目标是不断增长的以数据为中心的应用程序,在这些应用程序中,用户直接操作、分析和探索大型数据集,通常使用复杂的分析和机器学习技术。传统的数据库技术不适合服务于这一目的。过去,数据库假定(1)基于文本的输入(例如SQL)和输出,(2)点(即无状态)查询-响应范式,(3)批处理结果,以及(4)简单的分析。项目团队将放弃这些基本假设,并构建一个支持可视化输入和输出、“会话”交互、早期和渐进结果以及复杂分析的系统。构建一个集成了这些功能的系统需要对整个数据栈进行彻底的重新思考,从可视化界面到“核心”,以及结合相关的算法。主要的研究挑战围绕着开发算法和优化,利用HIL工作负载的独特特征来加速对大型数据集合的分析。该团队将建立一个名为20/20的概念验证HIL数据库,该数据库将紧密集成并显著扩展布朗大学建立的两项现有技术:PanoramicData是一个触控笔数据可视化系统,将作为前端。第二个构建块是Tupleware主存分析系统,它将复杂的分析管道编译成可执行文件。Tupleware将作为后端分析组件。该团队期望最终结果将提供一个实质性的加速(50%或更多),超过最先进的解决方案,用于常见的分析工作负载。项目网站(http://database.cs.brown.edu/projects/20-20/)将包括项目信息、出版物、公共数据集和代码。
英文摘要
Explorative data analysis plays a key role in data-driven discovery in a wide range of domains including science, engineering and business. In order for data analysis to become a commodity during a period when their user base is continually expanding and diversifying, human productivity and ease-of-use must become first-class design considerations for any database system. Unfortunately, data tools that are user friendly and designed to improve human productivity are still sorely lacking. This project will enable users at different skill levels to interact with and explore their large datasets far easier and faster than they do today. Rather than spending a lot of precious time to build complex analytics tasks, this work will offer a more agile, responsive and user-friendly system based on direct manipulation of the visual representations (e.g., charts, graphs, maps) of the data sets and analysis results. The system can also be used as a learning tool: e.g., a teacher could walk students through a complex dataset to verify specific hypothesis. This project will make large-scale data exploration more accessible to more users. Overall, it will accelerate discovery and breakthroughs in many domains such as e-commerce, finance and science. This research will be incorporated in undergraduate and graduate coursework. The outreach activities include special research and educationfocused programs that are geared towards undergraduates and high school girls.This project will build a new class of database systems designed for Human-In-the-Loop (HIL) operation. The work targets an ever-growing set of data-centric applications in which users directly manipulate, analyze and explore large data sets, often using complex analytics and machine learning techniques. Traditional database technologies are ill suited to serve this purpose. Historically, databases assumed (1) text-based input (e.g., SQL) and output, (2) a point (i.e., stateless) query-response paradigm, (3) batch results, and (4) simple analytics. The project team will drop these fundamental assumptions and build a system that instead supports visual input and output, "conversational" interaction, early and progressive results, and complex analytics. Building a system that integrates these features requires a complete rethinking of the full data stack, from the visual interface to the "core", as well as incorporating pertinent algorithms. The primary research challenges revolve around developing algorithms and optimizations that leverage the unique characteristics of HIL workloads to speed up analysis over large data collections. The team will build a proof-of-concept HIL database called 20/20 that will tightly integrate and significantly extend two existing technologies built at Brown: PanoramicData is a touch and pen data visualization system and will serve as the front-end. The second building block is the Tupleware main-memory analytics system, which compiles complex analytics pipelines into executables. Tupleware will serve as the backend analytics component. The team expects that the end result will offer a substantial speedup (50% or more) over the state-of-the-art solutions for common analytics workloads. The project web site (http://database.cs.brown.edu/projects/20-20/) will include information on the project, publications, public datasets and code.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Large: Collaborative Research: SciDB - An Array Oriented Data Management System for Massive Scale Scientific Data
-
批准号:1111423
-
项目类别:Continuing Grant
-
资助金额:$73.7万
-
财政年份:2011
-
负责人:Stanley Zdonik
-
依托单位:
III: Small: Automatic Incremental Design for Next-Generation Database Systems
-
批准号:0916691
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2009
-
负责人:Stanley Zdonik
-
依托单位:
ITR: Data Centers - Managing Data with Profiles
-
批准号:0086057
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2000
-
负责人:Stanley Zdonik
-
依托单位:
Analytical and Empirical Tools for Advanced Query Optimizer Engineering
-
批准号:9632629
-
项目类别:Standard Grant
-
资助金额:$38.19万
-
财政年份:1996
-
负责人:Stanley Zdonik
-
依托单位:
Constraint Query Languages
-
批准号:9509933
-
项目类别:Continuing Grant
-
资助金额:$20.5万
-
财政年份:1995
-
负责人:Stanley Zdonik
-
依托单位:
海外基金