Structure-Grounded Pretraining for Text-to-SQL

Structure-Grounded Pretraining for Text-to-SQL
复制标题

DOI:
10.18653/v1/2021.naacl-main.105
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Xiang Deng;Ahmed Hassan Awadallah;Christopher Meek;Oleksandr Polozov;Huan Sun;Matthew Richardson
Xiang Deng;Ahmed Hassan Awadallah;Christopher Meek;Oleksandr Polozov;Huan Sun;Matthew Richardson
中科院分区:
其他
文献类型:
--
作者:
Xiang Deng;Ahmed Hassan Awadallah;Christopher Meek;Oleksandr Polozov;Huan Sun;Matthew Richardson

文献摘要

被引文献

相似文献

学习捕捉文本 - 表格对齐对于像文本到SQL这样的任务至关重要。模型需要正确识别对列和值的自然语言引用,并将它们与给定的数据库模式相关联。在本文中,我们提出了一种新颖的弱监督基于结构的预训练框架(STRUG)用于文本到SQL,它能够基于平行的文本 - 表格语料库有效地学习捕捉文本 - 表格对齐。我们确定了一组新颖的预训练任务:列关联、值关联以及列 - 值映射,并利用它们对文本 - 表格编码器进行预训练。此外,为了在更现实的文本 - 表格对齐设置下评估不同的方法,我们基于Spider开发集创建了一个新的评估集Spider - Realistic,其中移除了对列名的明确提及,并采用八个现有的文本到SQL数据集进行跨数据库评估。STRUG在所有设置下都比BERTLARGE有显著的改进。与现有的预训练方法(如GRAPPA)相比,STRUG在Spider上取得了相似的性能,并且在更现实的数据集上优于所有基线。本工作中使用的所有代码和数据都将开源,以促进未来的研究。
Learning to capture text-table alignment is essential for tasks like text-to-SQL. A model needs to correctly recognize natural language references to columns and values and to ground them in the given database schema. In this paper, we present a novel weakly supervised Structure-Grounded pretraining framework (STRUG) for text-to-SQL that can effectively learn to capture text-table alignment based on a parallel text-table corpus. We identify a set of novel pretraining tasks: column grounding, value grounding and column-value mapping, and leverage them to pretrain a text-table encoder. Additionally, to evaluate different methods under more realistic text-table alignment settings, we create a new evaluation set Spider-Realistic based on Spider dev set with explicit mentions of column names removed, and adopt eight existing text-to-SQL datasets for cross-database evaluation. STRUG brings significant improvement over BERTLARGE in all settings. Compared with existing pretraining methods such as GRAPPA, STRUG achieves similar performance on Spider, and outperforms all baselines on more realistic sets. All the code and data used in this work will be open-sourced to facilitate future research.