Text-to-SQL Error Correction with Language Models of Code

Text-to-SQL Error Correction with Language Models of Code
复制标题

DOI:
10.48550/arxiv.2305.13073
复制
发表时间:
2023-05
期刊:
--
影响因子:
--
通讯作者:
Ziru Chen;Shijie Chen;Michael White;R. Mooney;Ali Payani;Jayanth Srinivasa;Yu Su;Huan Sun
Ziru Chen;Shijie Chen;Michael White;R. Mooney;Ali Payani;Jayanth Srinivasa;Yu Su;Huan Sun
中科院分区:
其他
文献类型:
--
作者:
Ziru Chen;Shijie Chen;Michael White;R. Mooney;Ali Payani;Jayanth Srinivasa;Yu Su;Huan Sun

文献摘要

相似文献

尽管最近在文本到 SQL 解析方面取得了进展,但当前的语义解析器对于实际使用来说仍然不够准确。在本文中,我们研究如何构建自动文本到 SQL 纠错模型。注意到标记级编辑脱离上下文并且有时含糊不清,我们建议改为构建子句级编辑模型。此外,虽然大多数代码语言模型都没有专门针对 SQL 进行预训练,但它们了解 Python 等编程语言中的常见数据结构及其操作。因此,我们提出了一种新颖的 SQL 查询及其编辑表示形式,它更紧密地遵循代码语言模型的预训练语料库。我们的纠错模型将不同解析器的精确集匹配精度提高了 2.4-6.5,并比两个强大的基线获得了高达 4.3 点的绝对改进。
Despite recent progress in text-to-SQL parsing, current semantic parsers are still not accurate enough for practical use. In this paper, we investigate how to build automatic text-to-SQL error correction models. Noticing that token-level edits are out of context and sometimes ambiguous, we propose building clause-level edit models instead. Besides, while most language models of code are not specifically pre-trained for SQL, they know common data structures and their operations in programming languages such as Python. Thus, we propose a novel representation for SQL queries and their edits that adheres more closely to the pre-training corpora of language models of code. Our error correction model improves the exact set match accuracy of different parsers by 2.4-6.5 and obtains up to 4.3 point absolute improvement over two strong baselines.