Data Lake Organization

Data Lake Organization
复制标题

DOI:
10.1109/tkde.2021.3091101
复制
发表时间:
2018-12
影响因子:
8.9
通讯作者:
F. Nargesian;Ken Pu;Bahar Ghadiri-Bashardoost;Erkang Zhu;Renée J. Miller
F. Nargesian;Ken Pu;Bahar Ghadiri-Bashardoost;Erkang Zhu;Renée J. Miller
中科院分区:
计算机科学2区
文献类型:
--
作者:
F. Nargesian;Ken Pu;Bahar Ghadiri-Bashardoost;Erkang Zhu;Renée J. Miller

文献摘要

被引文献

相似文献

我们考虑的问题,建立一个组织目录的数据湖,以支持有效的用户导航。组织目录被定义为非循环图,其包含表示属性集的节点和指示节点之间的子集关系的边。一个概率模型,构建用户导航行为的模型。该模型还预测了用户在给定组织的数据湖中找到相关表的可能性。我们将数据湖组织问题表述为组织结构的优化,以最大限度地提高通过导航发现表的预期可能性。提出了一种近似算法,并分析了其误差界。在合成数据湖和真实的数据湖上对算法的有效性和效率进行了评估。我们的实验表明,我们的算法构建的组织,优于许多现有的组织,包括现有的手工策划的分类,链接图,和一个共同的基线组织。我们还进行了一项正式的用户研究,表明导航可以帮助用户发现不容易通过关键字搜索查询访问的相关表。这表明,使用组织的关键字搜索和导航是数据湖中数据发现的补充模式。
We consider the problem of building an organizational directory of data lakes to support effective user navigation. The organization directory is defined as an acyclic graph that contains nodes representing sets of attributes and edges indicating subset relationships between nodes. A probabilistic model is constructed to model user navigational behaviour. The model also predicts the likelihood of users finding relevant tables in a data lake given an organization. We formulate the data lake organization problem as an optimization over the organizational structure in order to maximize the expected likelihood of discovering tables by navigating. An approximation algorithm is proposed with an analysis of its error bound. The effectiveness and efficiency of the algorithm are evaluated on both synthetic and real data lakes. Our experiments show that our algorithm constructs organizations that outperform many existing organizations including an existing hand-curated taxonomy, a linkage graph, and a common baseline organization. We have also conducted a formal user study which shows that navigation can help users discover relevant tables that are not easily accessible by keyword search queries. This suggests that keyword search and navigation using an organization are complementary modalities for data discovery in data lakes.