Streaming Analytics and Workflow Automation for DFS

Streaming Analytics and Workflow Automation for DFS
复制标题

DFS 的流分析和工作流程自动化

DOI:
10.1145/3383583.3398589
复制
发表时间:
2020
期刊:
Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2020
影响因子:
--
通讯作者:
S. Jayarathna
S. Jayarathna
中科院分区:
--
文献类型:
--
作者:
Yasith Jayawardana;S. Jayarathna

文献摘要

被引文献

相似文献

研究人员重复使用过去研究的数据,以避免昂贵的重新收集实验数据。然而,由于研究小组和学科之间对元数据表示缺乏共识,大规模数据重用具有挑战性。数据集文件系统(DFS)是一种半结构化数据描述格式,它通过标准化数据描述、存储和检索的语义来促进这种共识。在本文中,我们提出了 analytic-streams(一种使用 DFS 进行流式数据分析的规范)和 Streaming-hub(一种基于 DFS 构建的可视化编程工具包,用于简化数据分析工作流程)。分析流以更少的计算开销促进高阶数据分析,而流集线器则支持数据和分析的存储、检索、操作和可视化。我们讨论它们如何简化数据预处理、聚合和可视化,以及它们对数据分析工作流程的影响。
Researchers reuse data from past studies to avoid costly re-collection of experimental data. However, large-scale data reuse is challenging due to lack of consensus on metadata representations among research groups and disciplines. Dataset File System (DFS) is a semi-structured data description format that promotes such consensus by standardizing the semantics of data description, storage, and retrieval. In this paper, we present analytic-streams - a specification for streaming data analytics with DFS, and streaming-hub - a visual programming toolkit built on DFS to simplify data analysis workflows. Analytic-streams facilitate higher-order data analysis with less computational overhead, while streaming-hub enables storage, retrieval, manipulation, and visualization of data and analytics. We discuss how they simplify data pre-processing, aggregation, and visualization, and their implications on data analysis workflows.