Benchmarking Multimodal Regex Synthesis with Complex Structures

Benchmarking Multimodal Regex Synthesis with Complex Structures
复制标题

DOI:
10.18653/v1/2020.acl-main.541
复制
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Xi Ye;Qiaochu Chen;Işıl Dillig;Greg Durrett
Xi Ye;Qiaochu Chen;Işıl Dillig;Greg Durrett
中科院分区:
其他
文献类型:
--
作者:
Xi Ye;Qiaochu Chen;Işıl Dillig;Greg Durrett

文献摘要

相似文献

从自然语言生成正则表达式(regex)的现有数据集的复杂性有限;与用户在StackOverflow上发布的正则表达式任务相比,这些数据集中的正则表达式很简单,用于描述它们的语言也不多样。我们介绍了StructuredRegex,这是一个新的正则表达式合成数据集,与以前的数据集有三个方面的不同。首先,为了获得结构复杂和真实的正则表达式,我们使用概率语法生成正则表达式,其中包含从真实世界的StackOverflow帖子中观察到的预定义宏。其次,为了获得语言多样化的自然语言描述,我们向众包工作者展示底层正则表达式的抽象描述,并要求他们描述他们看到的模式,而不是让他们解释合成语言。第三,我们将每个正则表达式示例扩展为一个字符串集合,这些字符串与基本的真实正则表达式匹配或不匹配,类似于实际用户给出示例的方式。我们的定量和定性分析证明了StructuredRegex相对于先前数据集的优势。使用各种多模态合成技术的进一步实验结果突出了我们的数据集所带来的挑战,包括非局部约束和多模态输入。
Existing datasets for regular expression (regex) generation from natural language are limited in complexity; compared to regex tasks that users post on StackOverflow, the regexes in these datasets are simple, and the language used to describe them is not diverse. We introduce StructuredRegex, a new regex synthesis dataset differing from prior ones in three aspects. First, to obtain structurally complex and realistic regexes, we generate the regexes using a probabilistic grammar with pre-defined macros observed from real-world StackOverflow posts. Second, to obtain linguistically diverse natural language descriptions, we show crowdworkers abstract depictions of the underlying regex and ask them to describe the pattern they see, rather than having them paraphrase synthetic language. Third, we augment each regex example with a collection of strings that are and are not matched by the ground truth regex, similar to how real users give examples. Our quantitative and qualitative analysis demonstrates the advantages of StructuredRegex over prior datasets. Further experimental results using various multimodal synthesis techniques highlight the challenge presented by our dataset, including non-local constraints and multi-modal inputs.