The exception that improves the rule

The exception that improves the rule
复制标题

改进规则的例外

DOI:
10.1145/2939502.2939509
复制
发表时间:
2016
期刊:
ArXiv
影响因子:
--
通讯作者:
H. Mueller
H. Mueller
中科院分区:
--
文献类型:
--
作者:
J. Freire;Boris Glavic;Oliver Kennedy;H. Mueller

文献摘要

被引文献

相似文献

数据库社区已经开发了许多用于数据策划和探索的工具和技术,从声明语言到专门的数据维修技术等等。但是,目前尚无共识,即如何以简单,直觉且最重要的是灵活的方式将这些强大的工具最好地向分析师展示。因此,分析师继续依靠电子表格,命令式语言和笔记本样式编程环境等工具进行数据策划。在这项工作中,我们探讨了电子表格,笔记本和关系数据库的集成。我们专注于关键优势,即电子表格和命令式笔记本环境都比经典的关系数据库具有:易于例外。通过依靠设定的时间操作,关系数据库牺牲了轻松定义Singleton操作的能力,对正常数据处理工作流程的例外,该工作流程影响了查询处理,以影响固定的一组明确针对的记录。相比之下,电子表格用户可以轻松地更改一个单元格的公式,而笔记本用户可以向她的笔记本添加命令操作,从而改变输出“视图”。我们认为,在经典关系数据库中实现这种特质的手动转换对于策展至关重要,因为易于声明个人价值的策划操作通常非常具有挑战性。我们探讨了在关系数据库中启用单例的挑战,为数据策划提出了混合电子表格/关系笔记本环境,并提出了我们对Vizier的愿景,该系统通过这种界面来揭示数据策划。
The database community has developed numerous tools and techniques for data curation and exploration, from declarative languages, to specialized techniques for data repair, and more. Yet, there is currently no consensus on how to best expose these powerful tools to an analyst in a simple, intuitive, and above all, flexible way. Thus, analysts continue to rely on tools such as spreadsheets, imperative languages, and notebook style programming environments like Jupyter for data curation. In this work, we explore the integration of spreadsheets, notebooks, and relational databases. We focus on a key advantage that both spreadsheets and imperative notebook environments have over classical relational databases: ease of exception. By relying on set-at-a-time operations, relational databases sacrifice the ability to easily define singleton operations, exceptions to a normal data processing workflow that affect query processing for a fixed set of explicitly targeted records. In comparison, a spreadsheet user can easily change the formula for just one cell, while a notebook user can add an imperative operation to her notebook that alters an output "view". We believe that enabling such idiosyncratic manual transformations in a classical relational database is critical for curation, as curation operations that are easy to declare for individual values can often be extremely challenging to generalize. We explore the challenges of enabling singletons in relational databases, propose a hybrid spreadsheet/relational notebook environment for data curation, and present our vision of Vizier, a system that exposes data curation through such an interface.