DBPal: A Fully Pluggable NL2SQL Training Pipeline

DBPal: A Fully Pluggable NL2SQL Training Pipeline
复制标题

DOI:
10.1145/3318464.3380589
复制
发表时间:
2020-05
期刊:
Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data
影响因子:
--
通讯作者:
Nathaniel Weir;Prasetya Ajie Utama;Alex Galakatos;Andrew Crotty;Amir Ilkhechi;Shekar Ramaswamy;Rohin Bhushan;Nadja Geisler;Benjamin Hättasch;Steffen Eger;U. Çetintemel;Carsten Binnig
Nathaniel Weir;Prasetya Ajie Utama;Alex Galakatos;Andrew Crotty;Amir Ilkhechi;Shekar Ramaswamy;Rohin Bhushan;Nadja Geisler;Benjamin Hättasch;Steffen Eger;U. Çetintemel;Carsten Binnig
中科院分区:
其他
文献类型:
--
作者:
Nathaniel Weir;Prasetya Ajie Utama;Alex Galakatos;Andrew Crotty;Amir Ilkhechi;Shekar Ramaswamy;Rohin Bhushan;Nadja Geisler;Benjamin Hättasch;Steffen Eger;U. Çetintemel;Carsten Binnig

文献摘要

被引文献

相似文献

自然语言是DBMS的一个很有前途的替代接口,因为它使非技术用户能够以比SQL更简洁的方式制定复杂的问题。最近,深度学习在将自然语言翻译为SQL方面获得了吸引力,因为类似的想法在机器翻译的相关领域已经取得了成功。然而,现有深度学习方法的核心问题是,它们需要大量的训练数据才能提供准确的翻译。这种训练数据的管理成本非常高,因为它通常需要人类用相应的SQL查询手动注释自然语言示例(反之亦然)。基于这些观察结果,我们提出了一种新的方法,增强了现有的深度学习技术,以提高自然语言到SQL翻译模型的性能。更具体地说,我们提出了一种新的训练管道,可以自动生成合成训练数据,以便(1)提高整体翻译准确性,(2)提高对语言变化的鲁棒性,以及(3)针对目标数据库专门化模型。正如我们所展示的那样,我们的DBSQL训练管道能够提高最先进的自然语言到SQL翻译模型的准确性和语言鲁棒性。
Natural language is a promising alternative interface to DBMSs because it enables non-technical users to formulate complex questions in a more concise manner than SQL. Recently, deep learning has gained traction for translating natural language to SQL, since similar ideas have been successful in the related domain of machine translation. However, the core problem with existing deep learning approaches is that they require an enormous amount of training data in order to provide accurate translations. This training data is extremely expensive to curate, since it generally requires humans to manually annotate natural language examples with the corresponding SQL queries (or vice versa). Based on these observations, we propose DBPal, a new approach that augments existing deep learning techniques in order to improve the performance of models for natural language to SQL translation. More specifically, we present a novel training pipeline that automatically generates synthetic training data in order to (1) improve overall translation accuracy, (2) increase robustness to linguistic variation, and (3) specialize the model for the target database. As we show, our DBPal training pipeline is able to improve both the accuracy and linguistic robustness of state-of-the-art natural language to SQL translation models.