Enhanced Regular Expression as a DGL for Generation of Synthetic Big Data

Enhanced Regular Expression as a DGL for Generation of Synthetic Big Data
复制标题

DOI:
10.3745/jips.04.0262
复制
发表时间:
2023
期刊:
J. Inf. Process. Syst.
影响因子:
--
通讯作者:
K. Cheng;K. Abe
K. Cheng;K. Abe
中科院分区:
其他
文献类型:
--
作者:
K. Cheng;K. Abe

文献摘要

相似文献

合成数据生成通常用于数据密集型应用中的性能评估和功能测试,以及数据分析的各个领域,例如隐私保护数据发布(PPDP)和统计披露限制/控制。在数据生成工具和语言方面已经进行了大量研究。然而,现有的工具和语言是为特定目的而开发的,不适用于其他领域。在本文中,我们提出一种基于正则表达式的数据生成语言(DGL),用于灵活的大数据生成。为了实现一种通用且强大的DGL,我们对标准正则表达式进行了增强,以支持数据域、类型/格式推断、序列和随机生成、概率分布以及资源引用。为了有效地实现所提出的语言,我们针对中间查询和数据库查询都提出了缓存技术。我们通过实验对所提出的改进进行了评估。
Synthetic data generation is generally used in performance evaluation and function tests in data-intensive applications, as well as in various areas of data analytics, such as privacy-preserving data publishing (PPDP) and statistical disclosure limit/control. A significant amount of research has been conducted on tools and languages for data generation. However, existing tools and languages have been developed for specific purposes and are unsuitable for other domains. In this article, we propose a regular expression-based data generation language (DGL) for flexible big data generation. To achieve a general-purpose and powerful DGL, we enhanced the standard regular expressions to support the data domain, type/format inference, sequence and random generation, probability distributions, and resource reference. To efficiently implement the proposed language, we propose caching techniques for both the intermediate and database queries. We evaluated the proposed improvement experimentally.